Architecture
Budgets & bounded autonomy
Every agent has finite turn, tool call, and duration bounds. Policy also decides whether an exhausted run returns one constrained final answer or fails.
Normal work stops at a turn or tool call limit. Final-answer policy may allow one constrained
grace turn. Token and cost usage are checked at turn boundaries, so measured usage may cross the
threshold. maxDuration is a deadline and can interrupt in-flight work.
The stop policy
Section titled “The stop policy”The agent definition fixes its AgentPolicy:
AgentPolicy.make({ maxTurns: 12, maxToolCalls: 24, maxDuration: "5 minutes", toolConcurrency: 4, repeatedFailureLimit: 3, tokenBudget: 80_000, completionReserveTokens: 8_000, costBudgetMicrousd: 2_000_000, onExhaustion: "final-answer",});| Bound | What it limits | Result at the limit |
|---|---|---|
maxTurns |
model requests per run | follows onExhaustion |
maxToolCalls |
declared and programmatic tool calls | follows onExhaustion |
maxDuration |
logical run time, including durable suspension | typed failure |
tokenBudget |
cumulative input and output tokens | follows onExhaustion once |
completionReserveTokens |
capacity held for finalization | ends research before reserve use |
costBudgetMicrousd |
estimated run cost | typed failure |
repeatedFailureLimit |
consecutive terminal tool failures | typed failure |
toolConcurrency |
parallel tool handlers per batch | limits concurrency |
AgentPolicyError.limit names a failed bound. DN and DC persist the same value in
SubmissionSettled.policyLimit.
Transient provider failures
Section titled “Transient provider failures”modelRetries repeats a model call that failed before streaming any content, when the failure
is a retryable AiError or a stream error part with a numeric HTTP code or status of 429 or
5xx. Provider-defined tools exposed to the call must all be marked Tool.Readonly: hosted
execution can happen before any response part arrives. Calls that already streamed text,
reasoning, or a tool call are never repeated. The default is 0, which surfaces the first
failure unchanged.
Waits double from one second up to thirty seconds. A longer provider retryAfter takes
precedence, and all waiting counts toward maxDuration.
Exhaustion: final answer or failure
Section titled “Exhaustion: final answer or failure”onExhaustion: "final-answer" is the default for turns, tool calls, and tokens.
- An over-budget tool batch executes no handlers. Each call becomes a synthetic failed result that asks the model to answer from existing evidence. DN and DC do not durably declare the rejected batch.
- Later model requests forbid tools through
toolChoice: "none". If the definition owns a completion tool, only its singleton call remains allowed. - Turn exhaustion permits one grace turn. A second grace turn is impossible.
- Completion records
finishReason: "budget-exhausted"and anexhaustedvalue oftokens,tool-calls, orturns. - Any disallowed tool call fails with
ModelProtocolErrorbefore execution.
Synthetic rejections do not count toward repeatedFailureLimit because no handler ran.
onExhaustion: "fail" rejects work past the bound. The final permitted turn may still finish
through the definition’s singleton completion tool. Use this mode when a partial answer is invalid.
Duration, cost, and repeated tool failure always fail. Token exhaustion gets at most one constrained final answer.
Per-run allowances
Section titled “Per-run allowances”A caller may tighten turn and tool call limits for one run:
AgentRuntime.run(agent, input, { toolCallAllowance: 8, turnAllowance: 4 });The effective limit is min(policy bound, max(1, floor(allowance))). The runtime ignores
non-finite allowances. A run allowance never raises the definition’s ceiling.
Hierarchical usage budgets
Section titled “Hierarchical usage budgets”@yielded/agent can apply shared budgets in this order:
global → tenant → agent → thread → run.
Effect.gen(function* () { const root = yield* makeUsageBudgetRoot( UsageBudgetNodeConfig.make({ level: "tenant", id: "acme", limits: UsageBudgetLimits.make({ maxCostMicrousd: 50_000_000 }), }), ); const runNode = yield* root.child( UsageBudgetNodeConfig.make({ level: "run", id: runId, limits: runLimits }), ); return yield* AgentRuntime.run(agent, input, { budget: toRunBudgetHook(runNode) });});Consumption is atomic across the ancestor chain. A rejected increment changes no node and fails
as BudgetExceeded with the scope and bound. The run budget hook checks the model stream and tool
batches. It accounts for observed usage after the fact; it does not reserve a provider’s maximum
charge before I/O.
Delegation budgets
Section titled “Delegation budgets”A requested child tool call allowance is bounded independently by the reservation and definition:
effective allowance = min(normalized request, reserved tool calls, definition maxToolCalls)SubagentPolicy limits child count, concurrency, turns, tool calls, duration, tokens, cost, and
result bytes. The parent reserves the slice before starting the child and releases it once after
settlement. An overrun remains charged and reduces later headroom. The child’s stop policy still
controls its run and final-answer behavior.
Omitted child policy fields inherit the parent’s defaults; explicit child settings override
those defaults within the delegation ceilings. Omitted delegation policy uses one parent-sized
pool shared by child invocations. Inheritance supplies limits, not new budget credit.
An explicit delegation policy derives aggregate caps by multiplying per-child limits by
maxChildren; parentCaps can replace those caps. Every delegation in one parent Run must
use the same aggregate caps. The pool covers delegated work separately from the parent’s own
turn and tool-call counters.
Durable admission checks the recorded pool before reserving another child. Recovery restores the child’s resolved policy and existing reservation. Unknown usage remains charged conservatively, and admission fails closed on inconsistent caps or exhausted concurrency.
Independently funded workers
Section titled “Independently funded workers”A root may request budgetScope: "worker-run" for a background worker only when the native
WorkerBudgetAuthorizer permits its exact source, target, and allocation. Immutable delegation
lineage still controls Tool grants and depth; accounting belongs to the worker’s logical Run.
New Runs reuse the configured allowance, while joined input and every recovery Attempt share the
existing Run’s usage. No cumulative token or cost cap is invented when that dimension is omitted.
Attached descendants remain bounded by the immediate worker Run’s reserved subtree. See the
subagent reference for host configuration.
Request a larger child budget
Section titled “Request a larger child budget”failureMode: "return" converts expected child failures into model-visible result data. Engine
suspension and durability failures stay in the error channel.
toolCallAllowance: { default, fromParameters } grants one child invocation a tool call allowance
under its reservation and definition ceilings. This applies to ephemeral and durable children.
To extend work:
- the child reaches its allowance and returns a partial result with
budgetExhausted; - the parent starts a new delegation with a larger author-owned parameter;
- the runtime clamps the new allowance, and the child continues from forwarded findings.
An extension creates a new child thread. It cannot enlarge a live reservation. Durable establishment records the effective allowance before the child becomes runnable. Each replacement attempt restores that allowance and the committed tool-call usage, including after approval suspension or ownership loss. Recovery never grants a fresh budget. The child’s settlement carries the exhaustion marker for the parent’s result projection.
Programmatic calls and Code Mode
Section titled “Programmatic calls and Code Mode”Code Mode checks and reserves maxToolCalls before each inner handler. Exhaustion becomes that
call’s outcome. Reservations are serialized across concurrent calls and never refunded after
admission. The turn boundary then enforces the combined declared and programmatic count.
CodeExecutionLimits.maxHostCallConcurrency independently bounds inner calls per pass (default 4,
maximum 64); the outer toolConcurrency still bounds simultaneous generated programs.
Sizing guidance
Section titled “Sizing guidance”- Raise
maxTurnswithmaxToolCallsfor research agents. Repeated read cycles often use one of each. - A token quota cannot guarantee a dollar ceiling. Strict spending admission must reserve the maximum request charge before provider I/O, including finalization and cache misses.
- Split independent work into bounded child runs. Grant a larger follow-up only when the partial result warrants it.
Related guides
Section titled “Related guides”- Agent definitions guide covers policy declarations.
- Run and stream guide covers finish reasons and Run options.
- Persistence & durability covers child recovery.