Loop
NAME
LLM::Agent::Loop - the streaming agent loop: tools, retry, fallback, transcript, compaction
SYNOPSIS
use LLM::Agent::Loop;
my $loop = LLM::Agent::Loop.new(
backends => [$primary, $fallback], # tried in order, per round trip
provider => $policy, # any tools-for-llm/execute-tool-calls pair
session => $session, # optional, but this is what makes it resumable
compactor => $compactor, # optional, needs a context-budget
);
# The conversation is history. What is true TODAY ā the date, the cwd,
# the branch, the instruction files ā is a RunContext, rendered into the
# request per run and never into the transcript.
my $run = $loop.run([$user-message], context => $context);
start react whenever $run.events -> $event {
given $event {
when LLM::Agent::Event::Token { print $event.text }
when LLM::Agent::Event::ToolCall { note " -> {$event.name}" }
}
}
my %outcome = await $run.result;
say %outcome<final>;
DESCRIPTION
Everything between "here is a conversation" and "here is the answer": ask the model, stream what it says, run the tools it asks for, feed the results back, and go round again until it stops asking. Around that, the things that make it survive contact with reality ā retry, fallback, inactivity timeouts, cancellation, a transcript that is complete after every line, and compaction when the conversation outgrows the window.
It is a separate class from LLM::Chat::ToolLoop rather than an
extension of it: that loop is streaming-only, has no cancellation, no
retry and no framing between rounds, and bolting those on would change
its behaviour for every existing user. This one keeps a
LLM::Chat::Backend chain, a duck-typed tool provider, and publishes
one Supply of typed events.
What plugs into it
| Attribute | What it is |
|---|
backends | LLM::Chat::Backend chain, tried in order per round trip |
| provider | anything with tools-for-llm + execute-tool-calls (or nothing) |
counter | an LLM::Agent::TokenCount (default: ::Usage) |
| session | an LLM::Agent::Session to write the transcript to |
| compactor | an LLM::Agent::Compactor; needs context-budget |
| context-budget | the window size in tokens |
| request-budget | an LLM::Agent::RequestBudget: what fits, and what it may cost |
| artifact-store | where oversized tool results go (default: beside the transcript) |
| tool-deadline | seconds one tool call may run (default: no deadline) |
idempotency-rules | tool-name patterns ā read-only / idempotent / destructive |
| concurrent-tools | tool-name patterns whose neighbouring calls go down as one batch |
| identical-call-exempt | tool-name patterns the identical-call guard does not count at all |
| steer-source | a thunk answering user messages to inject between rounds |
| completion-bus | an LLM::Agent::CompletionBus: the work a run will not end without |
provider is duck-typed on purpose. An MCP::Client, an
MCP::Client::Registry over several of them, an MCP::Client::Policy
over that, or a hand-rolled object with the two methods ā the loop cannot
tell and does not care. With no provider the loop is a plain streaming
chat loop: no tools are declared, and a model that hallucinates a tool
call is answered as if it had not.
The loop stays entirely policy-unaware. It never imports
MCP::Client::Policy, never asks anything, and never decides what is
allowed. What it offers instead is two shims ā wrap-ask and
log-hook ā that an app wires at construction so that the questions and
the logs a policy or a client produces come out of the same event stream
as everything else. See /Wiring the shims.
One run at a time
run returns a LLM::Agent::Run immediately and does the work behind
it. Calling it again while a run is live dies. That is not a
limitation of the machinery ā most of the state is per-run ā but a
deliberate refusal to have an opinion about scheduling: how many agents
may talk to one backend at once, whether they queue or run in parallel,
what happens to the second one's tokens, are questions an application (or
a future scheduler) answers. A loop that is one run at a time is a loop
something else can queue.
The admission boundary is $run.drained, not $run.result. A cancelled
or deadlined run can have an answer while a detached provider call is still
producing or mutating; admitting another run over it would let callbacks and
side effects cross owners. The state machine therefore finishes the result,
closes every work section, and only then does run accept another call.
The little state the loop keeps across a round trip ā the backend and Response a cancel would have to poke ā is tagged with the run that owns it, and every reader checks the tag. So even an old handle that somehow fires late (a hook captured by an app, a straggler thread) finds a slot that is either its own or empty, and never the next run's.
Two things follow that are worth knowing when you are queueing runs: a
finished run's cancel does nothing at all (see LLM::Agent::Run),
and < $run.drained > ā not result ā is what tells you the previous
run has stopped touching the provider, which on a cancelled or deadlined
tool call is strictly later.
A round, step by step
Completions, first of all. If the loop was given a
completion-bus, everything background work reported since the last round is drained off it and appended as framed user turns. Nothing at all happens here without a bus. See /Background operations: the run that does not end yet.Steering, second. If the loop was given a
steer-source, it is asked whether the user has said anything since the last round, and whatever it answers is appended as ordinary user turns ā before the caps are checked, before a compaction, and before the request is built, so every one of them sees the new turns. Nothing at all happens here without asteer-source. See /Steering: a user turn between rounds.
Completions before steers, and it is not arbitrary: a steer arriving at the same boundary is very often the user reacting to a completion they have just watched arrive, and putting the reaction before the thing reacted to is a conversation that does not read.
Compaction check, at the top. Before spending a request, not after: a round that ends with six tool results is exactly the round that blew the budget, and checking on the way in means the next request is the one that fits. The compactor is asked
needs-compactionā its decision, not this class's, and it answers for every reason there is to make a pass: the conversation is overcompactor.trigger, or it is carrying an epoch's worth of stale tool results that can be elided far more cheaply than the window can be summarized (see LLM::Agent::Compactor's observation aging). Either way the pass runs, the working array is swapped, the session gets a line for what it did ā acompaction, anelision, or both ā andCompactionStarted/CompactionDoneframe it. A pass that reports it could not make the next request fit ends the run ā see /When the context runs out.
The conversation's own count is what the compactor works on, but the
number compared against the trigger includes the rendered run context:
it is real weight on every request, and a trigger that ignored it would
let a conversation sail past the window by the size of an AGENTS.md.
The round trip. Per backend: a preflight against that backend's window (see /What fits, and what it costs), and then
AttemptStartedand a streamed completion, withTokenevents as the text arrives. Retry and fallback happen here, per round trip ā see /Failure.Limits, before anything is committed. If the model asked for tools and a limit has been hit,
LimitReachedis emitted, a system message inToolLoop's exact wording is appended, tools are switched off, and the round starts again so the model can answer with what it has.The assistant turn is committed.
AssistantMessageandTurnCommitted, appended to the working array, written to the session (withreasoningandusageas replay-visible extras).Tools. One
ToolCallper call ā the model asked ā and then each call is dispatched, waited for and settled on its own unless its tool opted into a concurrency group: see /One tool call at a time. Every result is filed under the call'stool_call_idā the model'stool_callsentry is the only authority on that id, and a provider that answers with a different one has its answer corrected and aLogevent emitted saying so. Letting it win would put an id in the transcript that nothing in the conversation asked for, which is a 400 on the next request and a session nobody can resume.Grants. If the provider
can('grants')ā a policy does ā the snapshot is written to the session whenever it changes, so a resumed session does not ask the human again. See /Grants, and when they are written.
No tool calls, or tools switched off: RunCompleted, and the run ends ā
unless there is background work outstanding, in which case the run
parks instead of ending. See /Background operations: the run that does
not end yet.
The limit ordering, and the assistant turn it discards
Limits are checked before the assistant turn is committed, and a limit
therefore discards that turn ā no AssistantMessage, nothing in the
session. This is ToolLoop's behaviour and it is not cosmetic: an
assistant message carrying tool_calls that is never followed by the
matching tool messages is a malformed conversation, and every
OpenAI-compatible provider rejects the next request with a 400. The
choice is between losing one sentence of the model's commentary and
losing the run.
One tool call at a time
A round's tool calls are dispatched sequentially, one at a time, each
through the provider as a batch of one. Per call: a tool-dispatched
envelope, ToolStarted, the provider, and then whichever settle path
the answer (or the absence of one) calls for ā the tool message, the
tool-settled envelope, and ToolResult or ToolAbandoned. Only
then does the next call start.
That is a behaviour change from 0.1.x, and it buys three things a whole-batch dispatch could not have:
Real boundaries. Every call has a moment it started and a moment it settled, both of them durable (LLM::Agent::ToolOperation, LLM::Agent::Session's two envelopes). After a
SIGKILLthe transcript can say which of "never ran", "was running" and "finished, and the settle never landed" applies to each call ā which is the difference between a resume that lies to the model and one that does not.Bounded blast radius. A cancel or a deadline reaches the call it is about. The calls behind it are still
proposed, and a call that was never dispatched can be settled as abandoned: known not to have run, which is the strongest thing this loop is ever able to say.Model order. Side effects now happen in the order the model asked for them.
MCP::Client::Registry's per-server grouping degenerates to singletons, andMCP::Client::Policyno longer decides every call before any of them runs: the permission questions of the third call are asked after the first two have run. Both are deliberate. A model that writes a file and then reads it back is describing a sequence, and a policy that asked about all three up front was asking about a world that no longer existed by the time the third one ran.
What it costs is parallelism ā which is free for every tool whose runtime is a syscall, and not free for one whose runtime is another agent. See /Concurrency groups, and the tools that qualify.
Concurrency groups, and the tools that qualify
concurrent-tools is a list of tool-name patterns ā the same shape
idempotency-rules matches with, an exact name or a trailing-*
prefix glob ā and it is empty by default, which is the whole of the
section above unchanged. Name a tool in it and the loop is allowed to put
that tool's calls down together:
LLM::Agent::Loop.new(
:@backends, :$provider,
concurrent-tools => ['task'], # subagents may run side by side
);
Consecutive runs only, and model order is never rearranged. The batch
is walked in the order the model emitted it and a run of neighbouring
calls whose tools all match becomes one group; the first call that does
not match ends that group and is dispatched on its own. So
task, task, fs_read, task is three groups ā [task, task],
[fs_read], [task] ā and the fs_read still happens after the
first two subagents have answered and before the third one starts.
Grouping only ever merges neighbours, because model order is the
contract the section above buys and gathering the two task calls
around an fs_write into one group would reorder side effects the model
described as a sequence.
A group of one is a single dispatch: same envelopes, same events, same code path. Nothing about a loop with no matching tools in its batch differs from one built without the option at all.
What a group does with its calls is what the whole batch used to get: N
tool-dispatched envelopes and N ToolStarted events before
anything runs, one execute-tool-calls carrying all N calls, and then
one settle per call ā its own tool message, its own tool-settled
envelope, its own ToolResult ā in model order, indexed by position in
the call list.
What a group trades away is the per-call granularity of the three things above, and it is worth being precise about which:
Crash granularity is per group, not per call. A
SIGKILLin the middle of a group of three leaves three dispatched-unsettled operations, and a resume can only say "these three were in flight" ā where an ungrouped batch would have said "this one was running, and those two never started". Coarser, still honest: every one of the three really had been handed to the provider.A cancel or a deadline takes the whole group. The group is one wait; when it is given up on, every call in it settles
outcome-unknownwith the same reason. Calls after the group are untouched and still settle asabandonedā known not to have run.The policy decides the group before any of it runs. A
MCP::Client::Policyin front of a grouped batch asks its permission questions for all N calls up front, which is exactly what per-operation dispatch was introduced to stop. For a delegation tool that is the desired UX (the human answers for the whole fan-out once); for anything that mutates shared state it is the old bug.
So which tools qualify? The argument that carries is task's: a call
whose side effects are confined to its own resources and whose result
is a report rather than a change to the world this conversation is
reasoning about. A subagent runs in its own transcript, spends its own
budget, and hands back prose. Two of them side by side cannot see each
other's half-written state, so "what the world looked like when call
three was decided" is not a question with a wrong answer.
A tool that reads or writes anything the next call in the batch might
touch does not qualify, however tempting the wall clock is. fs_read
looks harmless until the model writes a file and reads it back in one
turn; sh_run is whatever somebody typed. Left out of
concurrent-tools, both keep the ordering guarantee they have always
had.
The tool deadline, and the humans it waits for
tool-deadline (undefined by default) bounds one call. When it passes,
the call is detached ā the bridge has no cancellation, so it carries on
somewhere ā and the operation settles outcome-unknown:
no
ToolResult, because nothing came back;a
ToolAbandonedwith< reason => 'deadline'> and< dispatched => True>;a
toolmessage telling the model, in words, that the call did not return in time and that whether it took effect is unknown. It is notis-error, because a local clock knows nothing about a remote side effect, and a model told thatfs_write"failed" will cheerfully write it again.
No later call in that assistant batch is dispatched: each is settled as
known not to have run, and the run carries on to a fresh model round. The
detached call keeps a work section open, so < $run.drained > waits for
it even though < $run.result > did not.
A concurrency group is bounded the same way and as a whole: the clock
is the longest-running operation in it, and a deadline that passes
detaches the batch and settles every call in it outcome-unknown.
The clock pauses while a human is being asked. At most one group is
ever in flight, so an AskPending raised inside a call belongs to
that group; wrap-ask records the span on every operation in it, and
the deadline is measured against < dispatched-at + ask-seconds >
rather than wall clock. Without that, a 30-second deadline would abandon
every call whose permission prompt sat on somebody's screen for a minute
ā having never let it start. (For an ungrouped call ā every call, until
something opts in ā "that group" and "that call" are the same thing, and
the span is the one this paragraph always described.)
The identical-call guard, and what counts as identical
max-identical-calls counts a call's name plus its canonicalised
arguments: the argument JSON is reparsed and re-rendered with sorted
keys, so < {"path":"a","mode":"r"} > and < {"mode":"r","path":"a"} >
are one call made twice. A model going round in circles rarely re-emits a
byte-identical document; it re-emits the same request.
The count spans the current batch as well as the run's history. Three identical calls in one turn are a loop just as surely as one per turn for three turns, and a guard that only looked at history would let the whole batch through and then run it. Only calls that were dispatched count ā a call abandoned to a cancel leaves the tally where it was, so a resumed run does not find its own limit half spent ā and a call whose outcome is unknown counts, because a call that times out identically for ever must not be able to evade the guard by never coming back.
The number is 30, and it used to be 3. Three could not tell a loop
from patience. A great many perfectly healthy flows make the same call
with the same arguments several times over, on purpose and by design: an
agent contending for a file lease retries the same lock_acquire
until the holder gives it back; an agent that has just written a file
reads it back to check, and reads it back again after the next edit; a
poll is a poll. At 3 the guard fired on the fourth honest repeat, the
loop disabled tools, the agent reported half a job, and whatever was
supervising it started the whole thing again ā which is the failure the
guard exists to prevent, arrived at by the guard. A model that really is
going round in circles will do it thirty times just as happily as four,
and thirty repeats cost some tokens where a false positive costs the
task.
Tools the guard does not count
identical-call-exempt is a list of tool-name patterns ā an exact name
or a trailing-* prefix glob, the same shape concurrent-tools and
idempotency-rules take, and empty by default. A call whose tool
matches is skipped by the identical-call check entirely: it is not
counted, and it cannot trip the limit. The skip happens before the
counting rather than after it, so an exempt call repeated forty times in
one batch does not walk the batch's own tally up to somewhere a later,
non-exempt call would fall off ā and everything else in the batch is
counted exactly as it was.
LLM::Agent::Loop.new(
:@backends, :$provider,
identical-call-exempt => ['lock_acquire', 'lock_release'],
);
What qualifies is a tool whose repetition is a protocol rather than a loop. Taking a lease is the argument's shape: "ask, be told somebody else has it, ask again" is the designed way to use it, the arguments cannot vary (they name the path), and the thing that ends the repetition is the other agent finishing ā nothing the model could think its way out of. The same goes for the release beside it, and for any other tool whose contract is "call me until the answer changes".
Nothing that does work qualifies. fs_edit with the same arguments
twice is either a no-op or a fight with somebody, sh_run is whatever
the model typed, and an exemption on either is the guard turned off for
the calls it was written for. The other limits ā max-tool-rounds and
max-tool-calls ā still bound an exempt tool, so an exemption is not a
licence for an unbounded run; it only says this tool's repeats are not
by themselves the evidence of a loop.
When the context runs out
Compaction is the loop's answer to a conversation that has outgrown its
window, and it is not always enough. One 400KB fs_read in the last
turn is bigger than some windows on its own, and the recent turns are
exactly what a compaction refuses to eat (see LLM::Agent::Compactor).
When the compactor reports that even a hard trim leaves the conversation
over budget, the run ends ā CompactionDone first, so the framing is
complete, then RunFailed with < reason => 'context-exhausted' > on
both the event and the result Map. The alternative is what 0.1.0 did:
send the doomed request anyway, fail, compact again, and repeat until
somebody cancels.
Whatever the compaction did manage is kept. The trimmed conversation is
the one in < %outcome<messages> > and the compaction line is in the
session, so the transcript describes what really happened and a resumed
run starts from the smaller conversation rather than the one that could
not be sent.
Three things reach that terminal, and they are the same fact arriving from three directions: the compactor running out of things to drop (above), a preflight that no backend passes (/The preflight, per attempt and per backend), and a provider that truncated the answer because the prompt had already filled the window (/The length that is really a window). The last of those is the one that costs money to discover, which is why it is never discovered twice for the same conversation.
What fits, and what it costs
An LLM::Agent::RequestBudget is how a deployment tells the loop three things it cannot work out for itself: how big each backend's window is, how much of a tool result is too much, and what this run is allowed to spend. All three are optional, and a loop without one behaves exactly as 0.1.x did.
context-budget on its own is still enough, and now buys something: a
loop given one and no explicit budget synthesizes a default profile of
exactly that window (no safety margin, no completion reserve ā see
LLM::Agent::RequestBudget on why both are zero). A conversation that
does not fit the window then ends the run cleanly instead of being sent
and rejected, which is what context-budget without a compactor used to
do. Both given: they have to agree about the window, or the constructor
says so. A compactor whose budget is above the smallest declared window
is refused for the same kind of reason ā every compaction would converge
on a conversation that backend still cannot take.
The preflight, per attempt and per backend
Before each backend is tried ā before the retry loop, and before any
AttemptStarted ā the loop asks whether the request fits it:
needed = counted messages + estimated tool declarations
+ the run context + margin
it fits when needed + completion-reserve <= context-window
Every term is named separately in the attempts record the skip leaves
behind, including a run context of zero, so the sum in the error message
is one a reader can check against the numbers they configured.
Per backend, and not once per round, because a prompt that fits the primary may not fit the fallback. A chain from a 128k model to an 8k one is exactly the case the check exists for, and a single check against "the" window would either wave the doomed request through or refuse a request the primary would have taken.
A backend that does not fit is skipped: an attempts record saying
what it needed and what it had, one Log event, and the chain moves on.
There is no AttemptStarted/AttemptFailed pair, and that is
deliberate on both halves. Preflight unfitness is arithmetic, so retrying
it would spend max-retries on a sum that cannot come out differently.
And the attempt framing is a contract about transport attempts ā a
Token belongs to the AttemptStarted that opened it ā so a backend
nothing was sent to must not open a token scope for a consumer to
retract. The record it leaves is the only one in attempts that carries
a disposition of its own, because there is no AttemptFailed beside
it to say what the loop did.
When every backend refuses, nothing has been sent and nothing has been
spent, and one move is left: compact to target ā the largest
conversation any of these backends would have taken ā and try the chain
once more. Still refused, or no compactor to ask: RunFailed with
< reason => 'context-exhausted' >, the same terminal the compactor
reaches from the other side.
The 400 that is really a window
Preflight is the defence; this is the seatbelt for the backend nobody
declared. An attempt that fails with status 400 whose error text
contains one of a documented set of phrases ā context length,
context_length, maximum context, too many tokens, prompt is
too long, matched case-insensitively ā is reclassified from abort to
advance, so the chain tries the next backend instead of ending the
run.
It is text sniffing, and it is worth being honest about that: there is no status, header or error class that distinguishes "your prompt is too long" from any other 400, providers word it however they like, and a provider that words it differently gets the old behaviour. What the reclassification buys is that the fallback ā often a model with a bigger window, or a different tokenizer ā gets its turn rather than being skipped because the primary said 400.
The length that is really a window
The expensive way for a window to run out. The provider accepts the
prompt, generates until it reaches the edge of the window, cuts the answer
off and reports finish_reason: 'length' ā which LLM::Chat surfaces as
a quit with an error class of response, which classify-error buckets
as advance, which on a single-backend chain degrades into a
retry-same. The result was a 200k-token request sent three to seven
times, at a full prompt each, to be truncated identically every time.
So a length is intercepted before the retry buckets get it, read off
< $resp.finish-reason > ā structured, not sniffed, because LLM::Chat
stamps the reason before it quits ā and split into the two failures it
really is:
| Verdict | What it means | What the loop does |
|---|---|---|
| near-window | the prompt filled the window | a window refusal: advance, and one compaction for the whole chain |
| completion-cap | a small prompt, a small max_tokens | RunFailed, reason 'completion-truncated' |
A near-window length is the preflight's refusal arriving from the other
side, and it goes exactly where that one goes: it counts as this backend's
window refusal, it is the second advance that does not degrade into a
retry, and when every backend in the chain has refused, the loop compacts
once and tries again ā context-exhausted if that does not help. The
one difference from a preflight refusal is what it cost to find out, which
is why the Log event beside it is a warning rather than an info.
The judgement leans towards near-window deliberately, and the reason
is worth stating plainly: the counter is the thing that let this request
through. A length arriving where the budget thought there was room is
evidence that the count was low, and the provider is the one who actually
tokenized the prompt. So the counted sum has to clear the window by more
than a quarter of itself ā a tokenizer disagrees with a counter by a few
percent, not by 25% ā before the cheerful reading is believed. A truncated
response that carried usage needs no tolerance at all: those are the
provider's own numbers, and prompt + completion reaching the window is
the window saying so in its own words.
The compaction target is not the preflight's "window less the reserve":
aiming a compaction at a number the counter has just been proved wrong
about would ask the compactor to drop nothing, and turn one compaction into
an instant context-exhausted. It aims at the counted size that would
still fit if the counter were as wrong as it is allowed to be ā or as
wrong as the provider's own prompt-tokens says it was ā so the one retry
is always a genuinely smaller request.
A completion-cap length is the other cap: a 12-token conversation whose
answer outgrew max_tokens. Compaction cannot help, a different backend
truncates the same answer just as flatly, and there is nothing to retry ā
so the round ends with RunFailed and < reason =>
'completion-truncated' >, and an error naming the cap that bound and
what to raise. A backend nothing describes lands here too, for the same
reason the preflight waves it through: an undeclared window is not evidence
of a full one, there is no target to compact to, and the terminal names
both possibilities so a deployment can declare a profile if the other one
was true.
The length the provider does not report
The same cut, without the confession. A provider that decodes tool call
arguments under a grammar can reach max_tokens with the JSON closed
ā every brace balanced, every string terminated ā and report a finish
reason of tool_calls rather than length. What arrives is a
syntactically perfect call missing whatever the model had not written
yet: a task brief that stops mid-sentence, arguments railroaded into
whatever key the grammar could still close, and a loop that dispatches it
because nothing about the response looks wrong.
There is one witness left, and it is the provider's own: usage. A
completion billed at the backend's max_tokens was stopped by
max_tokens, whatever the finish reason says. So every successful
streamed turn ā prose as much as tool calls ā is checked against the cap
the backend declares, and a completion of
completion-tokens >= the backend's max_tokens
is treated as the truncation it is: no AttemptSucceeded, and the
failure path from there is exactly the one a reported length takes ā
the same !length-verdict, so a clip on a request near the window
compacts once and a clip on a small one ends the run with
'completion-truncated', and the same refusal to re-send identical
bytes. One Log at warning names the disagreement (the reason the
provider gave, the tokens it billed, the cap they met); there is
deliberately no event class of its own, because the round did not fail in
a new way, it merely failed while claiming not to.
< >= > rather than == because usage that includes reasoning tokens
overshoots the cap on models that bill thinking, and there is no
tolerance in the other direction: a completion that stopped short of
the cap stopped because it was finished. The false positive this admits
is an answer that happens to end on exactly its last permitted token, and
it costs a compaction or a named terminal ā never a silently truncated
brief.
It is best-effort, and the honest path stays primary: a provider that
reports no usage at all leaves nothing to check, and the gate is simply
inert. Asking for usage on a stream (stream_options.include_usage)
would close that gap and is deliberately not requested ā it hangs
OpenRouter in the header phase, which is a worse failure than the one it
would catch. Providers that report usage without being asked (most do)
are covered; the rest keep the finish_reason: 'length' gate above,
which is the one that catches an honest provider every time.
Caps, and what "spent" means
max-cost, max-total-tokens and max-wall-clock on the budget are
checked at the top of every round and between tool operations,
through one hook. Both are points where nothing is in flight: an
operation that has been dispatched always settles first, because a cap is
a reason to stop starting things and never a reason to stop knowing what
the last one did. The calls behind it settle as abandoned with
< reason => 'budget-exhausted' > ā known not to have run ā and the run
ends with RunFailed and < reason => 'budget-exhausted' >, naming
the cap, what was spent and what the limit was.
Cost is only counted when a provider reported one (OpenRouter does,
through usage.cost; most do not), and it rides
AttemptSucceeded.usage and the assistant turn's session extras as
well, so a transcript records what each turn cost.
A run with a budget hands back one extra key in its result Map:
< spent => { total-tokens, wall-clock, cost?, prompt-tokens?,
completion-tokens?, cached-prompt-tokens?, parked-seconds? } >. It is
absent without a budget, because "nobody was counting" and "nothing was
spent" are different answers.
wall-clock is the time the run spent working, which on a loop with
a completion-bus is elapsed time less every second it spent parked
on background work; parked-seconds is that difference, present only
when the run actually parked. See /The caps, and what "wall clock" now
means.
The four optional keys follow the same rule one level down: each appears
only when some attempt in the run reported that number, and carries
the sum over the attempts that did. A backend that reports only a total
leaves the prompt-tokens / completion-tokens split absent rather
than zero, so a caller can tell "this run was 8k in and 2k out" apart
from "this run was 10k, and nobody said which way". cached-prompt-tokens
is a cache-hit SUBSET of prompt-tokens ā never its own addition to
total-tokens ā present only when a provider reported one.
Known limitation, and it is deliberate for 0.2: a resumed run starts its accumulators at zero. The transcript knows what each turn cost; adding those up across processes is a runtime store's job rather than a transcript's, and the loop does not pretend otherwise.
Spend that happened somewhere else
The other half of that limitation is a run that pays for work it did not
do itself: a LLM::Agent::Subagents child is a whole agent run of its
own, with its own loop, its own attempts and its own bill, and none of it
touches the accumulators of the parent that asked for it. A parent
max-cost therefore used to cap the parent's own turns and nothing
else, which for an agent whose entire job is delegating is a cap on the
cheapest part of what it spends.
absorb-spend closes that: it adds an already-settled spend record ā
another run's, in exactly the shape this one hands back ā into this run's
accumulators, so every cap sees it. The composer calls it as each child
settles (see LLM::Agent::Subagents), which makes a parent's budget the
budget of its whole subtree, recursively: a grandchild's spend is in
the child's record by the time the child settles into the parent's.
Two things about it are deliberate:
It is the same money, not a second kind. An absorbed record adds to
costand to the token counts exactly as an attempt of this run would, and the cap check that follows is the ordinary one at the next round or operation boundary. There is no separate "subtree cap" and no separate refusal ā a run that goes over because of what its children spent ends with the sameRunFailedandbudget-exhaustedreason as one that went over on its own.wall-clockis never absorbed. This run's elapsed time is measured from when it started, and children run inside that window (often several at once); adding their seconds to it would count the same minute several times over and tripmax-wall-clockfor a run that has been going thirty seconds.
The artifact path: one huge result, twice
A 400KB fs_read is bigger than some windows on its own, and it does
not go away: it enters the conversation, the transcript and every
subsequent request, and compaction refuses to eat the recent turns it
sits in. So a result over < request-budget.max-observation-size >
characters is excerpted into the conversation and written whole to a
file beside the transcript.
The invariant is one sentence: the excerpt is computed once, on the
settle path, before the tool message is built. So the conversation
content, the session payload, the ToolResult event's content and
the result-digest on the tool-settled envelope are all the same
string, the full bytes are in the artifact file and nowhere else, and a
resumed run replays byte for byte without the artifact existing at
all. Deleting the whole .artifacts directory loses the ability to
read what the tool really said; it cannot break a resume.
The rest of the design ā head-weighted excerpt, the marker line, why the
file is named by basename rather than by path, why retention is
none, and how this differs from the compactor's tool-result-cap ā
is in LLM::Agent::Artifacts. Two things belong here:
A write that fails is shielded. The model still gets an excerpt, it says the full result could not be stored, the operation still settles, and the failure comes out as a
Logevent. A tool call is not a failure because a scratch file could not be opened.A sessionless run has nowhere durable to put anything, so it excerpts and says so in the marker. Nothing is written, and no
artifactmetadata is claimed.
Grants, and when they are written
A provider that can('grants') ā an MCP::Client::Policy does ā is
asked for its snapshot, and a snapshot that has changed is written to
the session as a whole. Two details are load-bearing:
Changed means different, not bigger. The comparison is a digest of the whole snapshot, so a policy that narrows a rule, replaces one, or drops one is written out exactly as one that added a rule is. A loop that only noticed growth would resume tomorrow with permissions the human revoked today.
An "always" answer is written as soon as it exists, rather than when the batch it was asked inside of ends. The gap between those two moments is a tool call ā the slowest and most interruptible part of a round ā and a run that is cancelled or killed inside it would otherwise lose the answer and ask the same question again after the resume.
"As soon as it exists" is doing some work in that sentence, and it is
worth knowing why. wrap-ask runs inside the ask: a policy calls the
loop's shim, and only records the rule once the shim has returned to it.
So the loop cannot write the new grant from inside the shim ā at that
moment there is no new grant. The policy is the only thing that knows
when there is, and so the policy says so: wire < on-grant =>
$loop.grant-hook > and the snapshot is written the moment it changes,
with the tool call the question was about still running.
grant-hook is the whole mechanism. (0.1.x guessed instead: a task that
polled the provider's digest for a couple of seconds after every "always"
answer and wrote whatever it saw. It worked, and it was a poll where a
callback belongs.)
And, unwired, the loss is bounded to one tool call. Whether or not anything calls the hook, the loop compares digests and writes after every operation settles ā so the worst case for a policy with no
on-grantis that a grant answered during a call reaches disk when that call finishes, rather than when the whole batch does. That is what per-operation dispatch buys here.
They are also written on the way out of a run that is ending on a cancelled tool call or on a backend chain that failed, because those are exits that can happen with an answer given and nothing yet on disk.
Every one of these writes is shielded: a session that cannot take the
line produces a Log event rather than an exception. On the exit paths
that is because the run's terminal has already been decided and a grant
that could not be saved must not relabel it; on the per-operation path it
is because the rest of the batch still has to be closed off, and a
transcript with unanswered tool_calls in it is worth more damage than
a missing grant is. (0.1.x let that one failure take the run down. It was
the only unshielded sync, and it was the one in the worst place.)
Failure
Every failure is classified by LLM::Chat::Retry's classify-error ā
the same policy LLM::Data::Inference::Task has run in production ā
into abort, retry-same or advance:
| Bucket | What the loop does |
|---|---|
| abort | AttemptFailed, then RunFailed. 4xx will not heal in eight seconds. |
| retry-same | AttemptFailed with a backoff, sleep, same backend again |
| advance | AttemptFailed, next backend in the chain, no wait |
| advance | ...but with nowhere to go, it waits and tries again instead |
max-retries is the number of attempts per backend (Task's
semantics, not "extra tries"): 3 means one call and two retries before
the chain advances. When the budget runs out, the AttemptFailed says
advance, because disposition reports what the loop does rather
than what the classifier said in the abstract.
backoff-cap (default 30 s) is the ceiling on one wait;
LLM::Chat::Retry's exponential is < min(2 ** (n - 1) + jitter,
cap) >.
Advance, with nowhere to go
advance means "this failure is a property of this backend, so ask
a different one". On the last link of the chain ā very often the
only link ā there is no different one, and the literal reading of the
bucket is "give up". The loop does not take it: it degrades that
advance to a retry-same, waits out the normal backoff, and asks the
same backend again, spending the same per-backend budget a
retry-same would have.
The arithmetic is one-sided. If the retry fails the chain is over exactly as it would have been, one backoff later; if it succeeds, a run that was about to die on a transient hiccup did not. A single-backend config used to end on one flaky connect having made one call, which is what this exists to stop.
The AttemptFailed says retry-same with its backoff, because
that is what the loop is doing. When the budget is spent the next
failure says advance and the chain ends, unchanged.
The advances that do not degrade are the two context-overflow
seatbelts above: a 400 saying the conversation does not fit, and a
provider-reported length on a request at the edge of the window. Both
are deterministic refusals of this conversation by this backend ā
re-sending identical bytes buys an identical refusal and a wait on top of
it, and in the length case a second full prompt's worth of tokens to
learn nothing. Those two still end the backend's turn.
When every backend is spent, the run ends with RunFailed carrying the
full attempts list ā the same < { backend-index, model, error,
raw-text? } > records X::LLM::Chat::Retry::Exhausted would have
carried, and no reason: there is nothing to say about it beyond what
the attempts already say. The failures that do carry a reason are
the three the loop chose rather than suffered:
| reason | Why the loop stopped |
|---|---|
| context-exhausted | it will not fit, and compaction cannot help |
| budget-exhausted | a run cap (cost, tokens, wall clock) was reached |
| completion-truncated | max_tokens cut the answer off, and no smaller conversation changes that |
Mid-stream failure: what a consumer must do
A backend can fail after streaming four hundred tokens. Those tokens were
emitted; a Supply has no undo. The attempt framing is the contract,
and it is documented in full in LLM::Agent::Event ā in short, a
Token belongs to the AttemptStarted that opened it, and an
AttemptFailed retracts every Token since. A consumer that ignores
this renders the model's reply twice.
Nothing a failed attempt streamed reaches the session: only committed messages are written, so a transcript can never contain half a sentence the model was in the middle of when a load balancer dropped the connection.
The timeout is inactivity, not duration
round-trip-timeout (default 120s) measures the gap since the response
last did anything ā < $resp.last-activity-at > ā and not the total
time the request has taken.
This is a deliberate divergence from LLM::Data::Inference::Task, which bounds total duration. Task generates one JSON document per item and a slow one is a stuck one. An agent turn legitimately runs for minutes: a reasoning model thinks, a long file is summarised, a big diff is written. Bounding total time there means killing successful work at the point it has cost the most. What is never legitimate is a stream that stops producing tokens and never closes, which is exactly what a dropped connection behind a proxy looks like, and that is what this catches.
A timeout calls < $backend.cancel($resp) > and is classified as
timeout ā which is the advance bucket, so the next backend gets the
turn immediately rather than after a backoff. On a chain with no next
backend it becomes a wait and another try, like every other advance
with nowhere to go.
This is not the same thing as an HTTP client timeout that expired
while still trying to connect. Those never reach a backend at all,
and LLM::Chat::Backend::OpenAICommon classifies them connection ā
the retry-same bucket ā precisely because they say nothing about the
endpoint beyond "the network was unwell".
Cancellation
< $run.cancel > is cooperative, and the difference between what is
promised and what is not is worth knowing before you build a UI on it.
Promised:
No events after the terminal
RunCancelled.The in-flight stream is cancelled at the backend. (Whether that actually stops the generation upstream is the backend's business ā among the LLM::Chat backends only KoboldCpp really aborts; the others stop reading.)
No further rounds start.
A backoff sleep ends within one 0.25s chunk, not at the end of the backoff.
Not promised:
Interrupting an in-flight tool call. The provider bridge has no cancellation, so the call that is running will finish. The loop stops waiting for it, drains the result and ignores it: no
ToolResultis emitted, and no result reaches the conversation. The run ends promptly; thefs_writethat was already in flight still happened.
Because calls are dispatched one at a time, a cancel arriving mid-batch finds the batch in two halves, and the loop says two different things about them:
the call in flight settles
outcome-unknownāToolAbandonedwith< dispatched => True>, and atoolmessage saying the run was cancelled before it returned and that whether it took effect is unknown;the calls behind it settle
abandonedāToolAbandonedwith< dispatched => False>, and atoolmessage saying they did not run, which is a thing the loop can promise about a call it never dispatched.
Those synthetic messages are not results. They are what keeps the
transcript resumable: an assistant turn carrying tool_calls that
nothing answers is a conversation every OpenAI-compatible provider rejects
with a 400, so a cancelled run would otherwise leave a session that can
never be continued. Neither is is-error: one is unknown and the other
did not happen, and neither is a failure.
Interrupting a blocked permission prompt. An ask is a leaf: the policy holds a non-reentrant lock and is waiting for a human. Cancelling the run cannot dismiss a modal that something else owns. The run ends when the question is answered.
result ends the run; drained ends the work
The detached tool call is exactly where the two Promises on a Run part
company. < $run.result > is kept as soon as the loop has stopped
waiting ā a cancel should not take as long as the fs_write it
interrupted, and neither should a deadline. < $run.drained > is kept
only when that call has actually returned and nothing can emit into the
run any more.
So: result for "the caller has an answer", drained for "it is safe
to tear the provider down". A test that asserts nothing else arrived
should wait for the second one, because "nothing else arrived" is only
worth asserting once nothing else can.
Wiring the shims
The loop never talks to a policy, so the app connects them at
construction. There are four shims ā wrap-ask, log-hook,
grant-hook and progress-hook ā and all of them are no-ops when no
run is live: a late answer, a log or a progress note from a server that
had not noticed the run ended is dropped rather than emitted after the
terminal event.
There is a construction cycle to get round: the loop needs the policy (it
is the provider) and the policy needs the loop (the shim comes from it).
Forward-declare the loop and defer the shim into a closure ā the same
idiom MCP::Client::Policy's own Pod uses for its elicit-hook. What
does not work is < on-ask => $loop.wrap-ask(&ask) > with $loop
still undefined: that calls a method on a type object at construction
time and dies there.
my $loop; # named before it exists
# Permission prompts and server elicitations become AskPending /
# AskAnswered events, and still reach the real asker.
my $policy = MCP::Client::Policy.new(
:$provider,
on-ask => -> %request { $loop.wrap-ask(&ask-the-human).(%request) },
# "Always" answers reach the transcript the moment the policy has
# recorded them, rather than when the tool call ends.
on-grant => -> | { $loop.grant-hook.() },
);
# Server log notifications become Log events, and progress notifications
# become ToolProgress. NOTE the log-level: a modern MCP server sends
# NOTHING without one.
my $mcp = MCP::Client.connect-stdio(
command => 'mcp-filesystem',
on-log => -> %params { $loop.log-hook.(%params) },
on-progress => -> %params { $loop.progress-hook.(%params) },
log-level => 'info',
);
$loop = LLM::Agent::Loop.new(:@backends, provider => $policy);
wrap-ask emits AskPending, calls the real asker (which still
blocks, and still holds the policy's lock), emits AskAnswered, and
returns the answer untouched. An asker that throws is rethrown so the
policy's own handling ā refuse this one call, say why ā is unchanged;
the AskAnswered is still emitted, so an AskPending is never left
open.
It does one thing more, and it is what makes the tool deadline humane: the question is timed against the operation it is about. Dispatch is sequential, so the call in flight when a question is raised is the call the question is about, and the seconds a human takes are subtracted from that call's deadline ā see /The tool deadline, and the humans it waits for.
What it deliberately does not do is write grants. It runs inside the
ask, where the answer it just carried has not been recorded by anything
yet; grant-hook is how a policy says it has. See /Grants, and when
they are written.
Emitting from above the loop
The three shims above turn something the loop is given into an event.
emit-external is the other direction: a layer sitting on top of the
loop hands it an event it built itself, and the loop publishes it onto
the run in flight ā stamped, ordered and mailboxed like the driver's own.
# A subagent composer forwarding a child run's events; see
# LLM::Agent::Subagents, which is the reason this seam exists.
$loop.emit-external(LLM::Agent::Event::Subagent.new(
agent-id => 'reviewer-1', agent-type => 'reviewer',
inner => $child-event.to-hash,
));
It answers False rather than throwing when there is no live run or the
run has finished, which is what makes it safe to call from a thread that
belongs to something with a longer life than the run ā a child that is
still winding down after its parent ended is the ordinary case, not an
error. What it will not do is end a run: a terminal event is refused
loudly, because only _finish can keep the result Promise.
Use emitter-for instead when the emitter outlives the run.
emit-external publishes onto whatever run is live at the moment it
is called, which is right for a hook firing inside the run and wrong
for anything that reports late: a straggler from run A, arriving after
run B has started on the same loop, would be published onto B ā a turn
in B's transcript that never happened. < $loop.emitter-for($run) >
hands back a Callable bound to that one run, which answers False for
ever once it ends rather than following the loop to the next one.
# Captured when the work starts...
my &emit = $loop.emitter-for($loop.live-run);
# ...and used for that work's whole life, wherever it ends up running.
start {
for @late-events -> $event {
last unless &emit($event); # False: my run is over. Stop.
}
}
Steering: a user turn between rounds
A long run is a conversation the user is locked out of: the model works
for ten rounds, and everything the user thinks of in the meantime has to
wait for the run to end. steer-source is the way in. It is a thunk ā
no arguments ā answering a list of Str, and the loop calls it at the
top of every round and appends each answer as an ordinary user message.
# The app's queue, and the app's lock: the UI thread pushes onto it and
# the loop's thread drains it.
my @steers;
my $steer-lock = Lock.new;
my $loop = LLM::Agent::Loop.new(
:@backends, :$provider, :$session,
# DRAINED, not peeked: what this hands back is gone from the queue,
# because the loop has no way of handing it back.
steer-source => {
$steer-lock.protect: { my @taken = @steers; @steers = (); @taken }
},
);
# ...from the UI, while the run is in flight:
$steer-lock.protect: { @steers.push: 'leave the tests alone for now' };
What the placement buys is that a steer is never a surprise:
At a round boundary, never mid-batch. The thunk is called with nothing in flight ā no stream open, no tool call dispatched, no group half settled. A steer can therefore never land between an assistant turn carrying
tool_callsand thetoolmessages that answer them, which is the malformed conversation every provider rejects (see /The limit ordering, and the assistant turn it discards for the other half of the same rule).Before everything that reads the conversation. The caps and the preflight weigh the steer, a compaction can move it,
RoundStarted's token figure counts it, and the request carries it. A steer is not a sidecar on the request ā it is history, and the transcript records it as one more user turn with nothing to mark it out.On the round the limit path restarts, too. When a tool limit switches tools off and the loop goes round again for a final answer, that round asks for steers like any other. It is the same round boundary.
Including the first round, before the model has said anything. The queue is normally empty there ā the run was started from the user's question a moment ago ā and an app with something in it at that point gets what it asked for: a second user turn after the first.
And while the run is parked, once per
park-poll. A run waiting on background work may be waiting for minutes, and a user locked out of it for the duration would be locked out of exactly the situation they most want a word in. A pull that comes back with something ends the park and starts a round, and what it took is carried into that round: the pull is destructive, so a park that dropped it would swallow what the user typed. Nothing about the placement rules above changes ā a parked run has nothing in flight by definition.
Three things are the app's job, and the loop does not help with any of them:
The queue's thread safety. The thunk is only ever called on the driving thread, at a quiescent point, one call at a time ā but whatever it reads is shared with whatever the UI pushes onto, and that is the app's lock (or Channel) to get right.
Coalescing. The loop records exactly what it is handed: three answers are three user messages, in order. An app that would rather send one paragraph joins them itself.
Noticing delivery. There is no
Steeredevent and no acknowledgement, because the thunk already is one: the loop asked, and what the app handed over is on its way. Rendering it is the ordinary< $session.messages >/AssistantMessagepath.
A thunk that throws is shielded: the failure becomes a Log event
(level error, logger llm-agent.loop), the round carries on with no
steers, and the next round asks again. An entry that is not a defined
Str is dropped the same way. A broken queue cannot take a run down.
Background operations: the run that does not end yet
A tool call that answers immediately and does the work afterwards is the difference between an agent that delegates and one that waits. The catch is that the loop's terminal condition ā the model stopped asking for tools ā is exactly wrong for it: the model stopped asking because it was told the answer would arrive later, and a run that ended there would end before the answer it promised.
completion-bus is the fix, and it is the whole of the fix.
LLM::Agent::CompletionBus holds two things: the operations that have
been acknowledged and not yet reported, and the reports themselves. The
loop asks it one question at the terminal and one at every round
boundary.
my $bus = LLM::Agent::CompletionBus.new;
my $loop = LLM::Agent::Loop.new(
:@backends,
provider => $subagents, # given the same bus: see its Pod
completion-bus => $bus,
);
Without a bus none of this exists. No park, no drain, no events, no behaviour change of any kind ā which is what makes the option safe to add to a loop that has always worked one way.
The two moments
At the top of every round, whatever is on the bus is drained and appended as user turns, before the caps, the compaction and the request ā the same placement a steer gets, and for the same reason. Completions go in before steers; see /A round, step by step.
At the no-tool-calls terminal, the bus is asked whether it is
quiet: nothing outstanding and nothing queued, read as one snapshot. Quiet means the run really is over, and it ends exactly as it always did. Not quiet means it parks.
What a park is
A park is the run waiting, with nothing in flight, on a bare timer. It
emits RunParked with the inventory of what it is waiting for, fires
on-park, and then looks at four things every park-poll seconds:
| What it finds | What it does |
|---|
| something queued | resumes; the next round delivers it |
| nothing outstanding either | ends the run: RunCompleted, as the terminal would have |
| the run was cancelled | closes every outstanding op, then RunCancelled |
| a steer | resumes, carrying the steer into the round |
| nothing, for park-idle-timeout | the safety valve ā see below |
Every exit closes the park's span, fires on-unpark, and emits
RunResumed with the reason and how long it waited. Those two hooks are
balanced on every path, including a throw, because a host that lends a
concurrency slot for the duration of a park needs it back.
A bare timer, not a wait. Every poll in this class is written that way
to keep Promise.anyof from accumulating continuations (see the note in
!stream), and a park gets something more out of it: there is no
lost-wakeup race to lose, because there is no wakeup. A completion that
lands between two passes is found by the second one.
The turns a completion arrives as
An injected turn is an ordinary user message with an extraordinary first line. It says, in words, that it is an automated event and not the user speaking, and it names the operation it is about ā because a conversation is compacted, and a turn that only made sense next to the acknowledgement it answers becomes an orphan the first time that acknowledgement is summarised away.
Its content goes through the observation excerpt seam before the
message is built ā the same seam a tool result goes through, the same
< request-budget.max-observation-size >, the same artifact file beside
the transcript. A child that answers with a megabyte cannot land whole in
a user turn; what lands is an excerpt, and the full bytes are in the
artifact the marker names. The framing itself is never excerpted: it is
what makes the turn readable as an event rather than as a person.
The transcript line carries extras ā < injected => <kind> >, the
operation's id, and whatever the producer added (completion-of, the
originating call id) ā and those extras are the durable delivery
marker: a resume can tell a completion that was delivered from one that
is still owed without replaying a thing.
The caps, and what "wall clock" now means
max-wall-clock bounds the time the run spent working, which is
elapsed time less every second it spent parked. A run that delegated
three children and waited twenty minutes for them did not work for twenty
minutes, and a cap that said it did would kill runs for being patient ā
which, for an agent whose whole job is delegating, caps the one thing it
is for. The result Map reports parked-seconds beside wall-clock
whenever a run parked at all, so the two can be added back up.
park-idle-timeout is the real-time bound, and it is the safety
valve rather than a deadline: half an hour, by default, in which not one
thing arrived. Every arrival resets it. When it fires, the run closes
every outstanding operation, ends with RunFailed and
< reason => 'park-idle' >, and names the operations that never
answered ā because "whether they took effect is unknown" is the honest
thing to say about work whose producer went away, and the one thing
somebody reading that failure needs to know is which piece it was.
What a park does not do
It does not increment
max-tool-rounds. A round that was woken by a completion and asks for no tools is not a tool round, and a run that spent its round budget on notifications would run out of them for reasons that have nothing to do with tools.It does not re-enter itself. A wake round is an ordinary round: cancel check, drain, steers, caps, compaction, request. If the model goes quiet again with work still outstanding, it parks again ā a fresh
RunParked, a fresh inventory, a fresh idle clock.It does not produce the work. The loop opens nothing and settles nothing; it reads the bus and drains it. Who acknowledges an operation and who reports it are the composer's business (LLM::Agent::Subagents) or the host's.
Resuming: what "the same message" means
A run given a session that already holds messages has to extend
that transcript rather than rewrite it, so the messages it was handed are
checked against the ones the session replayed, index by index. The check
is a digest (LLM::Agent::Canonical) over everything that survives
the round trip ā role, content, tool calls, tool-call id, sticky,
sysprompt, depth ā and not a comparison of role and text.
That matters because the differences that break a resumed conversation
are exactly the ones prose does not show: an assistant turn that has
grown a tool_calls array, a tool result answering a different call
id, a system prompt that is no longer sticky. A run whose prefix does not
match ends immediately, with a RunFailed naming the index it
disagreed at and nothing written to the session; the fix is always to
start from < $session.messages >.
Runtime context: the half that is not history
< run(@messages, context => $context) > takes an
LLM::Agent::RunContext: today's date, the working directory, the
branch, the project's instruction file ā everything that is true now
rather than everything that happened. It is rendered into the request
and nowhere else:
what goes on the wire [ head, |@conversation, tail ]
what everything else sees @conversation
The wire view is built in exactly one place ā the line that calls
chat-completion-stream ā and is never stored, counted or compared.
@conversation itself never contains the context, which is what keeps
all of the following true and unchanged:
the seed check compares the same messages it always did, so a run may carry a completely different context from the one the transcript was written with (that is the point of the whole mechanism);
the session records the same message lines, in the same order;
the compactor is handed the conversation, never a request;
< RunStarted.message-count >and< %outcome<messages> >count history, not framing.
What is recorded is one run-context envelope per run that had one
(see LLM::Agent::Session), written after the seed check ā a run that is
refused leaves no line ā and shielded, so a transcript that cannot take
an audit record never fails a run that would have worked.
What the context costs, and where that number goes
The rendered head and tail are weighed once through the loop counter for the
ordinary proactive-compaction trigger. Every preflight and provider-length
verdict, however, calls the selected profile counter's count-request
(LLM::Agent::TokenCount) over conversation, context and tool catalogue
together. An Exact counter can therefore render the model's real request
template; an older counter safely composes its conversation and text counts.
The proactive figure travels to the compaction consult as
< needs-compaction(@conversation, :$tokens) >, so the compactor weighs
the trigger against the whole request while still working on the
conversation alone. It is deliberately not passed
to the compactor as a target: Compactor's no-op path reports
exhausted against the target it was given, and a reduced one would turn
a compaction with nothing to drop into a context-exhausted terminal.
For forced compaction, the selected profile's complete request count is split into conversation and non-history weight in that same counter's units. Its usable conversation target travels with the counter that produced it, and the compactor uses the override for every before/after, summary and trim check. Targets in different tokenizers are never compared as raw integers; fallback then re-preflights normally in its own units.
The calibration, across a context change
A calibration is "P prompt tokens for these N messages", and P was billed
for a complete request to one model. Before preflight the loop compares the
selected counter's model, context digest and canonical tool catalogue with
the shape it previously described. A change calls counter.invalidate;
rewritten history is detected by Usage's prefix digest, while an unchanged
prefix extended with new turns keeps the calibration. The next usage-bearing
attempt re-calibrates the selected profile only.
Tool declarations are fetched once
tools-for-llm is called once per run, not once per round. A round
is a network call, and asking an MCP server to re-list its catalogue
before each of them doubles the round trips for a list that changes
approximately never. A server that adds a tool mid-run is not noticed
until the next run.
SEE ALSO
LLM::Agent::Run, LLM::Agent::Event (especially the attempt-framing
contract), LLM::Agent::ToolOperation (what a tool call's states mean),
LLM::Agent::Session, LLM::Agent::Compactor,
LLM::Agent::RunContext (what a run is told about now),
LLM::Agent::RequestBudget (what fits and what it may cost),
LLM::Agent::Artifacts (the excerpt and the file behind it),
LLM::Agent::CompletionBus (the work a parked run is waiting for),
LLM::Agent::Subagents (a provider that spawns child runs, and the
consumer of emit-external), LLM::Chat::Retry,
MCP::Client::Policy.
the bus first, because that is what the run is here for ā and
stateis one atomic snapshot, so "something is deliverable" and "nothing is outstanding and nothing is queued" are read of the same instant. Split into two reads they describe two different worlds, and the operation that settled between them is an answer nobody hears;
the cancel next, and before the steer pull, which is a deliberate departure from the round top's order. Pulling steers is destructive ā the app's queue hands them over ā and a cancelled run that pulled would swallow what the user typed and then end without saying it. Nothing is lost by checking the bus first: reading it is not draining it, and what is on it stays on it for the next run;
the steer pull, which is how a user gets a word in while a fleet of children works. What it takes goes into the caller's stash;
the idle valve last, because it is the only one of the four that is a guess. ) method !park-for-completions( $run, @conversation, @attempts, @stashed, Int:D $round, Str:D $final, --> Hash:D ) { my @inventory = self!bus-inventory($run);