Loop

NAME

LLM::Agent::Loop - the streaming agent loop: tools, retry, fallback, transcript, compaction

SYNOPSIS


use LLM::Agent::Loop;

my $loop = LLM::Agent::Loop.new(
    backends => [$primary, $fallback],   # tried in order, per round trip
    provider => $policy,                 # any tools-for-llm/execute-tool-calls pair
    session  => $session,                # optional, but this is what makes it resumable
    compactor => $compactor,             # optional, needs a context-budget
);

# The conversation is history. What is true TODAY — the date, the cwd,
# the branch, the instruction files — is a RunContext, rendered into the
# request per run and never into the transcript.
my $run = $loop.run([$user-message], context => $context);

start react whenever $run.events -> $event {
    given $event {
        when LLM::Agent::Event::Token    { print $event.text }
        when LLM::Agent::Event::ToolCall { note "  -> {$event.name}" }
    }
}

my %outcome = await $run.result;
say %outcome<final>;

DESCRIPTION

Everything between "here is a conversation" and "here is the answer": ask the model, stream what it says, run the tools it asks for, feed the results back, and go round again until it stops asking. Around that, the things that make it survive contact with reality — retry, fallback, inactivity timeouts, cancellation, a transcript that is complete after every line, and compaction when the conversation outgrows the window.

It is a separate class from LLM::Chat::ToolLoop rather than an extension of it: that loop is streaming-only, has no cancellation, no retry and no framing between rounds, and bolting those on would change its behaviour for every existing user. This one keeps a LLM::Chat::Backend chain, a duck-typed tool provider, and publishes one Supply of typed events.

What plugs into it

Attribute | What it is
backends | LLM::Chat::Backend chain, tried in order per round trip
provider | anything with tools-for-llm + execute-tool-calls (or nothing)
counter | an LLM::Agent::TokenCount (default: ::Usage)
session | an LLM::Agent::Session to write the transcript to
compactor | an LLM::Agent::Compactor; needs context-budget
context-budget | the window size in tokens
request-budget | an LLM::Agent::RequestBudget: what fits, and what it may cost
artifact-store | where oversized tool results go (default: beside the transcript)
tool-deadline | seconds one tool call may run (default: no deadline)
idempotency-rules | tool-name patterns → read-only / idempotent / destructive
concurrent-tools | tool-name patterns whose neighbouring calls go down as one batch
identical-call-exempt | tool-name patterns the identical-call guard does not count at all
steer-source | a thunk answering user messages to inject between rounds
completion-bus | an LLM::Agent::CompletionBus: the work a run will not end without

provider is duck-typed on purpose. An MCP::Client, an MCP::Client::Registry over several of them, an MCP::Client::Policy over that, or a hand-rolled object with the two methods — the loop cannot tell and does not care. With no provider the loop is a plain streaming chat loop: no tools are declared, and a model that hallucinates a tool call is answered as if it had not.

The loop stays entirely policy-unaware. It never imports MCP::Client::Policy, never asks anything, and never decides what is allowed. What it offers instead is two shims — wrap-ask and log-hook — that an app wires at construction so that the questions and the logs a policy or a client produces come out of the same event stream as everything else. See /Wiring the shims.

One run at a time

run returns a LLM::Agent::Run immediately and does the work behind it. Calling it again while a run is live dies. That is not a limitation of the machinery — most of the state is per-run — but a deliberate refusal to have an opinion about scheduling: how many agents may talk to one backend at once, whether they queue or run in parallel, what happens to the second one's tokens, are questions an application (or a future scheduler) answers. A loop that is one run at a time is a loop something else can queue.

The admission boundary is $run.drained, not $run.result. A cancelled or deadlined run can have an answer while a detached provider call is still producing or mutating; admitting another run over it would let callbacks and side effects cross owners. The state machine therefore finishes the result, closes every work section, and only then does run accept another call.

The little state the loop keeps across a round trip — the backend and Response a cancel would have to poke — is tagged with the run that owns it, and every reader checks the tag. So even an old handle that somehow fires late (a hook captured by an app, a straggler thread) finds a slot that is either its own or empty, and never the next run's.

Two things follow that are worth knowing when you are queueing runs: a finished run's cancel does nothing at all (see LLM::Agent::Run), and < $run.drained > — not result — is what tells you the previous run has stopped touching the provider, which on a cancelled or deadlined tool call is strictly later.

A round, step by step

  • Completions, first of all. If the loop was given a completion-bus, everything background work reported since the last round is drained off it and appended as framed user turns. Nothing at all happens here without a bus. See /Background operations: the run that does not end yet.

  • Steering, second. If the loop was given a steer-source, it is asked whether the user has said anything since the last round, and whatever it answers is appended as ordinary user turns — before the caps are checked, before a compaction, and before the request is built, so every one of them sees the new turns. Nothing at all happens here without a steer-source. See /Steering: a user turn between rounds.

Completions before steers, and it is not arbitrary: a steer arriving at the same boundary is very often the user reacting to a completion they have just watched arrive, and putting the reaction before the thing reacted to is a conversation that does not read.

  • Compaction check, at the top. Before spending a request, not after: a round that ends with six tool results is exactly the round that blew the budget, and checking on the way in means the next request is the one that fits. The compactor is asked needs-compaction — its decision, not this class's, and it answers for every reason there is to make a pass: the conversation is over compactor.trigger, or it is carrying an epoch's worth of stale tool results that can be elided far more cheaply than the window can be summarized (see LLM::Agent::Compactor's observation aging). Either way the pass runs, the working array is swapped, the session gets a line for what it did — a compaction, an elision, or both — and CompactionStarted / CompactionDone frame it. A pass that reports it could not make the next request fit ends the run — see /When the context runs out.

The conversation's own count is what the compactor works on, but the number compared against the trigger includes the rendered run context: it is real weight on every request, and a trigger that ignored it would let a conversation sail past the window by the size of an AGENTS.md.

  • The round trip. Per backend: a preflight against that backend's window (see /What fits, and what it costs), and then AttemptStarted and a streamed completion, with Token events as the text arrives. Retry and fallback happen here, per round trip — see /Failure.

  • Limits, before anything is committed. If the model asked for tools and a limit has been hit, LimitReached is emitted, a system message in ToolLoop's exact wording is appended, tools are switched off, and the round starts again so the model can answer with what it has.

  • The assistant turn is committed. AssistantMessage and TurnCommitted, appended to the working array, written to the session (with reasoning and usage as replay-visible extras).

  • Tools. One ToolCall per call — the model asked — and then each call is dispatched, waited for and settled on its own unless its tool opted into a concurrency group: see /One tool call at a time. Every result is filed under the call's tool_call_id — the model's tool_calls entry is the only authority on that id, and a provider that answers with a different one has its answer corrected and a Log event emitted saying so. Letting it win would put an id in the transcript that nothing in the conversation asked for, which is a 400 on the next request and a session nobody can resume.

  • Grants. If the provider can('grants') — a policy does — the snapshot is written to the session whenever it changes, so a resumed session does not ask the human again. See /Grants, and when they are written.

No tool calls, or tools switched off: RunCompleted, and the run ends — unless there is background work outstanding, in which case the run parks instead of ending. See /Background operations: the run that does not end yet.

The limit ordering, and the assistant turn it discards

Limits are checked before the assistant turn is committed, and a limit therefore discards that turn — no AssistantMessage, nothing in the session. This is ToolLoop's behaviour and it is not cosmetic: an assistant message carrying tool_calls that is never followed by the matching tool messages is a malformed conversation, and every OpenAI-compatible provider rejects the next request with a 400. The choice is between losing one sentence of the model's commentary and losing the run.

One tool call at a time

A round's tool calls are dispatched sequentially, one at a time, each through the provider as a batch of one. Per call: a tool-dispatched envelope, ToolStarted, the provider, and then whichever settle path the answer (or the absence of one) calls for — the tool message, the tool-settled envelope, and ToolResult or ToolAbandoned. Only then does the next call start.

That is a behaviour change from 0.1.x, and it buys three things a whole-batch dispatch could not have:

  • Real boundaries. Every call has a moment it started and a moment it settled, both of them durable (LLM::Agent::ToolOperation, LLM::Agent::Session's two envelopes). After a SIGKILL the transcript can say which of "never ran", "was running" and "finished, and the settle never landed" applies to each call — which is the difference between a resume that lies to the model and one that does not.

  • Bounded blast radius. A cancel or a deadline reaches the call it is about. The calls behind it are still proposed, and a call that was never dispatched can be settled as abandoned: known not to have run, which is the strongest thing this loop is ever able to say.

  • Model order. Side effects now happen in the order the model asked for them. MCP::Client::Registry's per-server grouping degenerates to singletons, and MCP::Client::Policy no longer decides every call before any of them runs: the permission questions of the third call are asked after the first two have run. Both are deliberate. A model that writes a file and then reads it back is describing a sequence, and a policy that asked about all three up front was asking about a world that no longer existed by the time the third one ran.

What it costs is parallelism — which is free for every tool whose runtime is a syscall, and not free for one whose runtime is another agent. See /Concurrency groups, and the tools that qualify.

Concurrency groups, and the tools that qualify

concurrent-tools is a list of tool-name patterns — the same shape idempotency-rules matches with, an exact name or a trailing-* prefix glob — and it is empty by default, which is the whole of the section above unchanged. Name a tool in it and the loop is allowed to put that tool's calls down together:


    LLM::Agent::Loop.new(
        :@backends, :$provider,
        concurrent-tools => ['task'],   # subagents may run side by side
    );

Consecutive runs only, and model order is never rearranged. The batch is walked in the order the model emitted it and a run of neighbouring calls whose tools all match becomes one group; the first call that does not match ends that group and is dispatched on its own. So task, task, fs_read, task is three groups — [task, task], [fs_read], [task] — and the fs_read still happens after the first two subagents have answered and before the third one starts. Grouping only ever merges neighbours, because model order is the contract the section above buys and gathering the two task calls around an fs_write into one group would reorder side effects the model described as a sequence.

A group of one is a single dispatch: same envelopes, same events, same code path. Nothing about a loop with no matching tools in its batch differs from one built without the option at all.

What a group does with its calls is what the whole batch used to get: N tool-dispatched envelopes and N ToolStarted events before anything runs, one execute-tool-calls carrying all N calls, and then one settle per call — its own tool message, its own tool-settled envelope, its own ToolResult — in model order, indexed by position in the call list.

What a group trades away is the per-call granularity of the three things above, and it is worth being precise about which:

  • Crash granularity is per group, not per call. A SIGKILL in the middle of a group of three leaves three dispatched-unsettled operations, and a resume can only say "these three were in flight" — where an ungrouped batch would have said "this one was running, and those two never started". Coarser, still honest: every one of the three really had been handed to the provider.

  • A cancel or a deadline takes the whole group. The group is one wait; when it is given up on, every call in it settles outcome-unknown with the same reason. Calls after the group are untouched and still settle as abandoned — known not to have run.

  • The policy decides the group before any of it runs. A MCP::Client::Policy in front of a grouped batch asks its permission questions for all N calls up front, which is exactly what per-operation dispatch was introduced to stop. For a delegation tool that is the desired UX (the human answers for the whole fan-out once); for anything that mutates shared state it is the old bug.

So which tools qualify? The argument that carries is task's: a call whose side effects are confined to its own resources and whose result is a report rather than a change to the world this conversation is reasoning about. A subagent runs in its own transcript, spends its own budget, and hands back prose. Two of them side by side cannot see each other's half-written state, so "what the world looked like when call three was decided" is not a question with a wrong answer.

A tool that reads or writes anything the next call in the batch might touch does not qualify, however tempting the wall clock is. fs_read looks harmless until the model writes a file and reads it back in one turn; sh_run is whatever somebody typed. Left out of concurrent-tools, both keep the ordering guarantee they have always had.

The tool deadline, and the humans it waits for

tool-deadline (undefined by default) bounds one call. When it passes, the call is detached — the bridge has no cancellation, so it carries on somewhere — and the operation settles outcome-unknown:

  • no ToolResult, because nothing came back;

  • a ToolAbandoned with < reason => 'deadline' > and < dispatched => True >;

  • a tool message telling the model, in words, that the call did not return in time and that whether it took effect is unknown. It is not is-error, because a local clock knows nothing about a remote side effect, and a model told that fs_write "failed" will cheerfully write it again.

No later call in that assistant batch is dispatched: each is settled as known not to have run, and the run carries on to a fresh model round. The detached call keeps a work section open, so < $run.drained > waits for it even though < $run.result > did not.

A concurrency group is bounded the same way and as a whole: the clock is the longest-running operation in it, and a deadline that passes detaches the batch and settles every call in it outcome-unknown.

The clock pauses while a human is being asked. At most one group is ever in flight, so an AskPending raised inside a call belongs to that group; wrap-ask records the span on every operation in it, and the deadline is measured against < dispatched-at + ask-seconds > rather than wall clock. Without that, a 30-second deadline would abandon every call whose permission prompt sat on somebody's screen for a minute — having never let it start. (For an ungrouped call — every call, until something opts in — "that group" and "that call" are the same thing, and the span is the one this paragraph always described.)

The identical-call guard, and what counts as identical

max-identical-calls counts a call's name plus its canonicalised arguments: the argument JSON is reparsed and re-rendered with sorted keys, so < {"path":"a","mode":"r"} > and < {"mode":"r","path":"a"} > are one call made twice. A model going round in circles rarely re-emits a byte-identical document; it re-emits the same request.

The count spans the current batch as well as the run's history. Three identical calls in one turn are a loop just as surely as one per turn for three turns, and a guard that only looked at history would let the whole batch through and then run it. Only calls that were dispatched count — a call abandoned to a cancel leaves the tally where it was, so a resumed run does not find its own limit half spent — and a call whose outcome is unknown counts, because a call that times out identically for ever must not be able to evade the guard by never coming back.

The number is 30, and it used to be 3. Three could not tell a loop from patience. A great many perfectly healthy flows make the same call with the same arguments several times over, on purpose and by design: an agent contending for a file lease retries the same lock_acquire until the holder gives it back; an agent that has just written a file reads it back to check, and reads it back again after the next edit; a poll is a poll. At 3 the guard fired on the fourth honest repeat, the loop disabled tools, the agent reported half a job, and whatever was supervising it started the whole thing again — which is the failure the guard exists to prevent, arrived at by the guard. A model that really is going round in circles will do it thirty times just as happily as four, and thirty repeats cost some tokens where a false positive costs the task.

Tools the guard does not count

identical-call-exempt is a list of tool-name patterns — an exact name or a trailing-* prefix glob, the same shape concurrent-tools and idempotency-rules take, and empty by default. A call whose tool matches is skipped by the identical-call check entirely: it is not counted, and it cannot trip the limit. The skip happens before the counting rather than after it, so an exempt call repeated forty times in one batch does not walk the batch's own tally up to somewhere a later, non-exempt call would fall off — and everything else in the batch is counted exactly as it was.


    LLM::Agent::Loop.new(
        :@backends, :$provider,
        identical-call-exempt => ['lock_acquire', 'lock_release'],
    );

What qualifies is a tool whose repetition is a protocol rather than a loop. Taking a lease is the argument's shape: "ask, be told somebody else has it, ask again" is the designed way to use it, the arguments cannot vary (they name the path), and the thing that ends the repetition is the other agent finishing — nothing the model could think its way out of. The same goes for the release beside it, and for any other tool whose contract is "call me until the answer changes".

Nothing that does work qualifies. fs_edit with the same arguments twice is either a no-op or a fight with somebody, sh_run is whatever the model typed, and an exemption on either is the guard turned off for the calls it was written for. The other limits — max-tool-rounds and max-tool-calls — still bound an exempt tool, so an exemption is not a licence for an unbounded run; it only says this tool's repeats are not by themselves the evidence of a loop.

When the context runs out

Compaction is the loop's answer to a conversation that has outgrown its window, and it is not always enough. One 400KB fs_read in the last turn is bigger than some windows on its own, and the recent turns are exactly what a compaction refuses to eat (see LLM::Agent::Compactor).

When the compactor reports that even a hard trim leaves the conversation over budget, the run ends — CompactionDone first, so the framing is complete, then RunFailed with < reason => 'context-exhausted' > on both the event and the result Map. The alternative is what 0.1.0 did: send the doomed request anyway, fail, compact again, and repeat until somebody cancels.

Whatever the compaction did manage is kept. The trimmed conversation is the one in < %outcome<messages> > and the compaction line is in the session, so the transcript describes what really happened and a resumed run starts from the smaller conversation rather than the one that could not be sent.

Three things reach that terminal, and they are the same fact arriving from three directions: the compactor running out of things to drop (above), a preflight that no backend passes (/The preflight, per attempt and per backend), and a provider that truncated the answer because the prompt had already filled the window (/The length that is really a window). The last of those is the one that costs money to discover, which is why it is never discovered twice for the same conversation.

What fits, and what it costs

An LLM::Agent::RequestBudget is how a deployment tells the loop three things it cannot work out for itself: how big each backend's window is, how much of a tool result is too much, and what this run is allowed to spend. All three are optional, and a loop without one behaves exactly as 0.1.x did.

context-budget on its own is still enough, and now buys something: a loop given one and no explicit budget synthesizes a default profile of exactly that window (no safety margin, no completion reserve — see LLM::Agent::RequestBudget on why both are zero). A conversation that does not fit the window then ends the run cleanly instead of being sent and rejected, which is what context-budget without a compactor used to do. Both given: they have to agree about the window, or the constructor says so. A compactor whose budget is above the smallest declared window is refused for the same kind of reason — every compaction would converge on a conversation that backend still cannot take.

The preflight, per attempt and per backend

Before each backend is tried — before the retry loop, and before any AttemptStarted — the loop asks whether the request fits it:


    needed = counted messages + estimated tool declarations
                              + the run context + margin
    it fits when   needed + completion-reserve <= context-window

Every term is named separately in the attempts record the skip leaves behind, including a run context of zero, so the sum in the error message is one a reader can check against the numbers they configured.

Per backend, and not once per round, because a prompt that fits the primary may not fit the fallback. A chain from a 128k model to an 8k one is exactly the case the check exists for, and a single check against "the" window would either wave the doomed request through or refuse a request the primary would have taken.

A backend that does not fit is skipped: an attempts record saying what it needed and what it had, one Log event, and the chain moves on. There is no AttemptStarted/AttemptFailed pair, and that is deliberate on both halves. Preflight unfitness is arithmetic, so retrying it would spend max-retries on a sum that cannot come out differently. And the attempt framing is a contract about transport attempts — a Token belongs to the AttemptStarted that opened it — so a backend nothing was sent to must not open a token scope for a consumer to retract. The record it leaves is the only one in attempts that carries a disposition of its own, because there is no AttemptFailed beside it to say what the loop did.

When every backend refuses, nothing has been sent and nothing has been spent, and one move is left: compact to target — the largest conversation any of these backends would have taken — and try the chain once more. Still refused, or no compactor to ask: RunFailed with < reason => 'context-exhausted' >, the same terminal the compactor reaches from the other side.

The 400 that is really a window

Preflight is the defence; this is the seatbelt for the backend nobody declared. An attempt that fails with status 400 whose error text contains one of a documented set of phrases — context length, context_length, maximum context, too many tokens, prompt is too long, matched case-insensitively — is reclassified from abort to advance, so the chain tries the next backend instead of ending the run.

It is text sniffing, and it is worth being honest about that: there is no status, header or error class that distinguishes "your prompt is too long" from any other 400, providers word it however they like, and a provider that words it differently gets the old behaviour. What the reclassification buys is that the fallback — often a model with a bigger window, or a different tokenizer — gets its turn rather than being skipped because the primary said 400.

The length that is really a window

The expensive way for a window to run out. The provider accepts the prompt, generates until it reaches the edge of the window, cuts the answer off and reports finish_reason: 'length' — which LLM::Chat surfaces as a quit with an error class of response, which classify-error buckets as advance, which on a single-backend chain degrades into a retry-same. The result was a 200k-token request sent three to seven times, at a full prompt each, to be truncated identically every time.

So a length is intercepted before the retry buckets get it, read off < $resp.finish-reason > — structured, not sniffed, because LLM::Chat stamps the reason before it quits — and split into the two failures it really is:

Verdict What it means What the loop does
near-window the prompt filled the window a window refusal: advance, and one compaction for the whole chain
completion-cap a small prompt, a small max_tokens RunFailed, reason 'completion-truncated'

A near-window length is the preflight's refusal arriving from the other side, and it goes exactly where that one goes: it counts as this backend's window refusal, it is the second advance that does not degrade into a retry, and when every backend in the chain has refused, the loop compacts once and tries again — context-exhausted if that does not help. The one difference from a preflight refusal is what it cost to find out, which is why the Log event beside it is a warning rather than an info.

The judgement leans towards near-window deliberately, and the reason is worth stating plainly: the counter is the thing that let this request through. A length arriving where the budget thought there was room is evidence that the count was low, and the provider is the one who actually tokenized the prompt. So the counted sum has to clear the window by more than a quarter of itself — a tokenizer disagrees with a counter by a few percent, not by 25% — before the cheerful reading is believed. A truncated response that carried usage needs no tolerance at all: those are the provider's own numbers, and prompt + completion reaching the window is the window saying so in its own words.

The compaction target is not the preflight's "window less the reserve": aiming a compaction at a number the counter has just been proved wrong about would ask the compactor to drop nothing, and turn one compaction into an instant context-exhausted. It aims at the counted size that would still fit if the counter were as wrong as it is allowed to be — or as wrong as the provider's own prompt-tokens says it was — so the one retry is always a genuinely smaller request.

A completion-cap length is the other cap: a 12-token conversation whose answer outgrew max_tokens. Compaction cannot help, a different backend truncates the same answer just as flatly, and there is nothing to retry — so the round ends with RunFailed and < reason => 'completion-truncated' >, and an error naming the cap that bound and what to raise. A backend nothing describes lands here too, for the same reason the preflight waves it through: an undeclared window is not evidence of a full one, there is no target to compact to, and the terminal names both possibilities so a deployment can declare a profile if the other one was true.

The length the provider does not report

The same cut, without the confession. A provider that decodes tool call arguments under a grammar can reach max_tokens with the JSON closed — every brace balanced, every string terminated — and report a finish reason of tool_calls rather than length. What arrives is a syntactically perfect call missing whatever the model had not written yet: a task brief that stops mid-sentence, arguments railroaded into whatever key the grammar could still close, and a loop that dispatches it because nothing about the response looks wrong.

There is one witness left, and it is the provider's own: usage. A completion billed at the backend's max_tokens was stopped by max_tokens, whatever the finish reason says. So every successful streamed turn — prose as much as tool calls — is checked against the cap the backend declares, and a completion of


    completion-tokens >= the backend's max_tokens

is treated as the truncation it is: no AttemptSucceeded, and the failure path from there is exactly the one a reported length takes — the same !length-verdict, so a clip on a request near the window compacts once and a clip on a small one ends the run with 'completion-truncated', and the same refusal to re-send identical bytes. One Log at warning names the disagreement (the reason the provider gave, the tokens it billed, the cap they met); there is deliberately no event class of its own, because the round did not fail in a new way, it merely failed while claiming not to.

< >= > rather than == because usage that includes reasoning tokens overshoots the cap on models that bill thinking, and there is no tolerance in the other direction: a completion that stopped short of the cap stopped because it was finished. The false positive this admits is an answer that happens to end on exactly its last permitted token, and it costs a compaction or a named terminal — never a silently truncated brief.

It is best-effort, and the honest path stays primary: a provider that reports no usage at all leaves nothing to check, and the gate is simply inert. Asking for usage on a stream (stream_options.include_usage) would close that gap and is deliberately not requested — it hangs OpenRouter in the header phase, which is a worse failure than the one it would catch. Providers that report usage without being asked (most do) are covered; the rest keep the finish_reason: 'length' gate above, which is the one that catches an honest provider every time.

Caps, and what "spent" means

max-cost, max-total-tokens and max-wall-clock on the budget are checked at the top of every round and between tool operations, through one hook. Both are points where nothing is in flight: an operation that has been dispatched always settles first, because a cap is a reason to stop starting things and never a reason to stop knowing what the last one did. The calls behind it settle as abandoned with < reason => 'budget-exhausted' > — known not to have run — and the run ends with RunFailed and < reason => 'budget-exhausted' >, naming the cap, what was spent and what the limit was.

Cost is only counted when a provider reported one (OpenRouter does, through usage.cost; most do not), and it rides AttemptSucceeded.usage and the assistant turn's session extras as well, so a transcript records what each turn cost.

A run with a budget hands back one extra key in its result Map: < spent => { total-tokens, wall-clock, cost?, prompt-tokens?, completion-tokens?, cached-prompt-tokens?, parked-seconds? } >. It is absent without a budget, because "nobody was counting" and "nothing was spent" are different answers.

wall-clock is the time the run spent working, which on a loop with a completion-bus is elapsed time less every second it spent parked on background work; parked-seconds is that difference, present only when the run actually parked. See /The caps, and what "wall clock" now means.

The four optional keys follow the same rule one level down: each appears only when some attempt in the run reported that number, and carries the sum over the attempts that did. A backend that reports only a total leaves the prompt-tokens / completion-tokens split absent rather than zero, so a caller can tell "this run was 8k in and 2k out" apart from "this run was 10k, and nobody said which way". cached-prompt-tokens is a cache-hit SUBSET of prompt-tokens — never its own addition to total-tokens — present only when a provider reported one.

Known limitation, and it is deliberate for 0.2: a resumed run starts its accumulators at zero. The transcript knows what each turn cost; adding those up across processes is a runtime store's job rather than a transcript's, and the loop does not pretend otherwise.

Spend that happened somewhere else

The other half of that limitation is a run that pays for work it did not do itself: a LLM::Agent::Subagents child is a whole agent run of its own, with its own loop, its own attempts and its own bill, and none of it touches the accumulators of the parent that asked for it. A parent max-cost therefore used to cap the parent's own turns and nothing else, which for an agent whose entire job is delegating is a cap on the cheapest part of what it spends.

absorb-spend closes that: it adds an already-settled spend record — another run's, in exactly the shape this one hands back — into this run's accumulators, so every cap sees it. The composer calls it as each child settles (see LLM::Agent::Subagents), which makes a parent's budget the budget of its whole subtree, recursively: a grandchild's spend is in the child's record by the time the child settles into the parent's.

Two things about it are deliberate:

  • It is the same money, not a second kind. An absorbed record adds to cost and to the token counts exactly as an attempt of this run would, and the cap check that follows is the ordinary one at the next round or operation boundary. There is no separate "subtree cap" and no separate refusal — a run that goes over because of what its children spent ends with the same RunFailed and budget-exhausted reason as one that went over on its own.

  • wall-clock is never absorbed. This run's elapsed time is measured from when it started, and children run inside that window (often several at once); adding their seconds to it would count the same minute several times over and trip max-wall-clock for a run that has been going thirty seconds.

The artifact path: one huge result, twice

A 400KB fs_read is bigger than some windows on its own, and it does not go away: it enters the conversation, the transcript and every subsequent request, and compaction refuses to eat the recent turns it sits in. So a result over < request-budget.max-observation-size > characters is excerpted into the conversation and written whole to a file beside the transcript.

The invariant is one sentence: the excerpt is computed once, on the settle path, before the tool message is built. So the conversation content, the session payload, the ToolResult event's content and the result-digest on the tool-settled envelope are all the same string, the full bytes are in the artifact file and nowhere else, and a resumed run replays byte for byte without the artifact existing at all. Deleting the whole .artifacts directory loses the ability to read what the tool really said; it cannot break a resume.

The rest of the design — head-weighted excerpt, the marker line, why the file is named by basename rather than by path, why retention is none, and how this differs from the compactor's tool-result-cap — is in LLM::Agent::Artifacts. Two things belong here:

  • A write that fails is shielded. The model still gets an excerpt, it says the full result could not be stored, the operation still settles, and the failure comes out as a Log event. A tool call is not a failure because a scratch file could not be opened.

  • A sessionless run has nowhere durable to put anything, so it excerpts and says so in the marker. Nothing is written, and no artifact metadata is claimed.

Grants, and when they are written

A provider that can('grants') — an MCP::Client::Policy does — is asked for its snapshot, and a snapshot that has changed is written to the session as a whole. Two details are load-bearing:

  • Changed means different, not bigger. The comparison is a digest of the whole snapshot, so a policy that narrows a rule, replaces one, or drops one is written out exactly as one that added a rule is. A loop that only noticed growth would resume tomorrow with permissions the human revoked today.

  • An "always" answer is written as soon as it exists, rather than when the batch it was asked inside of ends. The gap between those two moments is a tool call — the slowest and most interruptible part of a round — and a run that is cancelled or killed inside it would otherwise lose the answer and ask the same question again after the resume.

"As soon as it exists" is doing some work in that sentence, and it is worth knowing why. wrap-ask runs inside the ask: a policy calls the loop's shim, and only records the rule once the shim has returned to it. So the loop cannot write the new grant from inside the shim — at that moment there is no new grant. The policy is the only thing that knows when there is, and so the policy says so: wire < on-grant => $loop.grant-hook > and the snapshot is written the moment it changes, with the tool call the question was about still running.

grant-hook is the whole mechanism. (0.1.x guessed instead: a task that polled the provider's digest for a couple of seconds after every "always" answer and wrote whatever it saw. It worked, and it was a poll where a callback belongs.)

  • And, unwired, the loss is bounded to one tool call. Whether or not anything calls the hook, the loop compares digests and writes after every operation settles — so the worst case for a policy with no on-grant is that a grant answered during a call reaches disk when that call finishes, rather than when the whole batch does. That is what per-operation dispatch buys here.

They are also written on the way out of a run that is ending on a cancelled tool call or on a backend chain that failed, because those are exits that can happen with an answer given and nothing yet on disk.

Every one of these writes is shielded: a session that cannot take the line produces a Log event rather than an exception. On the exit paths that is because the run's terminal has already been decided and a grant that could not be saved must not relabel it; on the per-operation path it is because the rest of the batch still has to be closed off, and a transcript with unanswered tool_calls in it is worth more damage than a missing grant is. (0.1.x let that one failure take the run down. It was the only unshielded sync, and it was the one in the worst place.)

Failure

Every failure is classified by LLM::Chat::Retry's classify-error — the same policy LLM::Data::Inference::Task has run in production — into abort, retry-same or advance:

Bucket What the loop does
abort AttemptFailed, then RunFailed. 4xx will not heal in eight seconds.
retry-same AttemptFailed with a backoff, sleep, same backend again
advance AttemptFailed, next backend in the chain, no wait
advance ...but with nowhere to go, it waits and tries again instead

max-retries is the number of attempts per backend (Task's semantics, not "extra tries"): 3 means one call and two retries before the chain advances. When the budget runs out, the AttemptFailed says advance, because disposition reports what the loop does rather than what the classifier said in the abstract.

backoff-cap (default 30 s) is the ceiling on one wait; LLM::Chat::Retry's exponential is < min(2 ** (n - 1) + jitter, cap) >.

Advance, with nowhere to go

advance means "this failure is a property of this backend, so ask a different one". On the last link of the chain — very often the only link — there is no different one, and the literal reading of the bucket is "give up". The loop does not take it: it degrades that advance to a retry-same, waits out the normal backoff, and asks the same backend again, spending the same per-backend budget a retry-same would have.

The arithmetic is one-sided. If the retry fails the chain is over exactly as it would have been, one backoff later; if it succeeds, a run that was about to die on a transient hiccup did not. A single-backend config used to end on one flaky connect having made one call, which is what this exists to stop.

The AttemptFailed says retry-same with its backoff, because that is what the loop is doing. When the budget is spent the next failure says advance and the chain ends, unchanged.

The advances that do not degrade are the two context-overflow seatbelts above: a 400 saying the conversation does not fit, and a provider-reported length on a request at the edge of the window. Both are deterministic refusals of this conversation by this backend — re-sending identical bytes buys an identical refusal and a wait on top of it, and in the length case a second full prompt's worth of tokens to learn nothing. Those two still end the backend's turn.

When every backend is spent, the run ends with RunFailed carrying the full attempts list — the same < { backend-index, model, error, raw-text? } > records X::LLM::Chat::Retry::Exhausted would have carried, and no reason: there is nothing to say about it beyond what the attempts already say. The failures that do carry a reason are the three the loop chose rather than suffered:

reason Why the loop stopped
context-exhausted it will not fit, and compaction cannot help
budget-exhausted a run cap (cost, tokens, wall clock) was reached
completion-truncated max_tokens cut the answer off, and no smaller conversation changes that

Mid-stream failure: what a consumer must do

A backend can fail after streaming four hundred tokens. Those tokens were emitted; a Supply has no undo. The attempt framing is the contract, and it is documented in full in LLM::Agent::Event — in short, a Token belongs to the AttemptStarted that opened it, and an AttemptFailed retracts every Token since. A consumer that ignores this renders the model's reply twice.

Nothing a failed attempt streamed reaches the session: only committed messages are written, so a transcript can never contain half a sentence the model was in the middle of when a load balancer dropped the connection.

The timeout is inactivity, not duration

round-trip-timeout (default 120s) measures the gap since the response last did anything — < $resp.last-activity-at > — and not the total time the request has taken.

This is a deliberate divergence from LLM::Data::Inference::Task, which bounds total duration. Task generates one JSON document per item and a slow one is a stuck one. An agent turn legitimately runs for minutes: a reasoning model thinks, a long file is summarised, a big diff is written. Bounding total time there means killing successful work at the point it has cost the most. What is never legitimate is a stream that stops producing tokens and never closes, which is exactly what a dropped connection behind a proxy looks like, and that is what this catches.

A timeout calls < $backend.cancel($resp) > and is classified as timeout — which is the advance bucket, so the next backend gets the turn immediately rather than after a backoff. On a chain with no next backend it becomes a wait and another try, like every other advance with nowhere to go.

This is not the same thing as an HTTP client timeout that expired while still trying to connect. Those never reach a backend at all, and LLM::Chat::Backend::OpenAICommon classifies them connection — the retry-same bucket — precisely because they say nothing about the endpoint beyond "the network was unwell".

Cancellation

< $run.cancel > is cooperative, and the difference between what is promised and what is not is worth knowing before you build a UI on it.

Promised:

  • No events after the terminal RunCancelled.

  • The in-flight stream is cancelled at the backend. (Whether that actually stops the generation upstream is the backend's business — among the LLM::Chat backends only KoboldCpp really aborts; the others stop reading.)

  • No further rounds start.

  • A backoff sleep ends within one 0.25s chunk, not at the end of the backoff.

Not promised:

  • Interrupting an in-flight tool call. The provider bridge has no cancellation, so the call that is running will finish. The loop stops waiting for it, drains the result and ignores it: no ToolResult is emitted, and no result reaches the conversation. The run ends promptly; the fs_write that was already in flight still happened.

Because calls are dispatched one at a time, a cancel arriving mid-batch finds the batch in two halves, and the loop says two different things about them:

  • the call in flight settles outcome-unknown — ToolAbandoned with < dispatched => True >, and a tool message saying the run was cancelled before it returned and that whether it took effect is unknown;

  • the calls behind it settle abandoned — ToolAbandoned with < dispatched => False >, and a tool message saying they did not run, which is a thing the loop can promise about a call it never dispatched.

Those synthetic messages are not results. They are what keeps the transcript resumable: an assistant turn carrying tool_calls that nothing answers is a conversation every OpenAI-compatible provider rejects with a 400, so a cancelled run would otherwise leave a session that can never be continued. Neither is is-error: one is unknown and the other did not happen, and neither is a failure.

  • Interrupting a blocked permission prompt. An ask is a leaf: the policy holds a non-reentrant lock and is waiting for a human. Cancelling the run cannot dismiss a modal that something else owns. The run ends when the question is answered.

result ends the run; drained ends the work

The detached tool call is exactly where the two Promises on a Run part company. < $run.result > is kept as soon as the loop has stopped waiting — a cancel should not take as long as the fs_write it interrupted, and neither should a deadline. < $run.drained > is kept only when that call has actually returned and nothing can emit into the run any more.

So: result for "the caller has an answer", drained for "it is safe to tear the provider down". A test that asserts nothing else arrived should wait for the second one, because "nothing else arrived" is only worth asserting once nothing else can.

Wiring the shims

The loop never talks to a policy, so the app connects them at construction. There are four shims — wrap-ask, log-hook, grant-hook and progress-hook — and all of them are no-ops when no run is live: a late answer, a log or a progress note from a server that had not noticed the run ended is dropped rather than emitted after the terminal event.

There is a construction cycle to get round: the loop needs the policy (it is the provider) and the policy needs the loop (the shim comes from it). Forward-declare the loop and defer the shim into a closure — the same idiom MCP::Client::Policy's own Pod uses for its elicit-hook. What does not work is < on-ask => $loop.wrap-ask(&ask) > with $loop still undefined: that calls a method on a type object at construction time and dies there.


my $loop;                                        # named before it exists

# Permission prompts and server elicitations become AskPending /
# AskAnswered events, and still reach the real asker.
my $policy = MCP::Client::Policy.new(
    :$provider,
    on-ask => -> %request { $loop.wrap-ask(&ask-the-human).(%request) },
    # "Always" answers reach the transcript the moment the policy has
    # recorded them, rather than when the tool call ends.
    on-grant => -> | { $loop.grant-hook.() },
);

# Server log notifications become Log events, and progress notifications
# become ToolProgress. NOTE the log-level: a modern MCP server sends
# NOTHING without one.
my $mcp = MCP::Client.connect-stdio(
    command     => 'mcp-filesystem',
    on-log      => -> %params { $loop.log-hook.(%params) },
    on-progress => -> %params { $loop.progress-hook.(%params) },
    log-level   => 'info',
);

$loop = LLM::Agent::Loop.new(:@backends, provider => $policy);

wrap-ask emits AskPending, calls the real asker (which still blocks, and still holds the policy's lock), emits AskAnswered, and returns the answer untouched. An asker that throws is rethrown so the policy's own handling — refuse this one call, say why — is unchanged; the AskAnswered is still emitted, so an AskPending is never left open.

It does one thing more, and it is what makes the tool deadline humane: the question is timed against the operation it is about. Dispatch is sequential, so the call in flight when a question is raised is the call the question is about, and the seconds a human takes are subtracted from that call's deadline — see /The tool deadline, and the humans it waits for.

What it deliberately does not do is write grants. It runs inside the ask, where the answer it just carried has not been recorded by anything yet; grant-hook is how a policy says it has. See /Grants, and when they are written.

Emitting from above the loop

The three shims above turn something the loop is given into an event. emit-external is the other direction: a layer sitting on top of the loop hands it an event it built itself, and the loop publishes it onto the run in flight — stamped, ordered and mailboxed like the driver's own.


# A subagent composer forwarding a child run's events; see
# LLM::Agent::Subagents, which is the reason this seam exists.
$loop.emit-external(LLM::Agent::Event::Subagent.new(
    agent-id => 'reviewer-1', agent-type => 'reviewer',
    inner    => $child-event.to-hash,
));

It answers False rather than throwing when there is no live run or the run has finished, which is what makes it safe to call from a thread that belongs to something with a longer life than the run — a child that is still winding down after its parent ended is the ordinary case, not an error. What it will not do is end a run: a terminal event is refused loudly, because only _finish can keep the result Promise.

Use emitter-for instead when the emitter outlives the run. emit-external publishes onto whatever run is live at the moment it is called, which is right for a hook firing inside the run and wrong for anything that reports late: a straggler from run A, arriving after run B has started on the same loop, would be published onto B — a turn in B's transcript that never happened. < $loop.emitter-for($run) > hands back a Callable bound to that one run, which answers False for ever once it ends rather than following the loop to the next one.


# Captured when the work starts...
my &emit = $loop.emitter-for($loop.live-run);

# ...and used for that work's whole life, wherever it ends up running.
start {
    for @late-events -> $event {
        last unless &emit($event);   # False: my run is over. Stop.
    }
}

Steering: a user turn between rounds

A long run is a conversation the user is locked out of: the model works for ten rounds, and everything the user thinks of in the meantime has to wait for the run to end. steer-source is the way in. It is a thunk — no arguments — answering a list of Str, and the loop calls it at the top of every round and appends each answer as an ordinary user message.


# The app's queue, and the app's lock: the UI thread pushes onto it and
# the loop's thread drains it.
my @steers;
my $steer-lock = Lock.new;

my $loop = LLM::Agent::Loop.new(
    :@backends, :$provider, :$session,
    # DRAINED, not peeked: what this hands back is gone from the queue,
    # because the loop has no way of handing it back.
    steer-source => {
        $steer-lock.protect: { my @taken = @steers; @steers = (); @taken }
    },
);

# ...from the UI, while the run is in flight:
$steer-lock.protect: { @steers.push: 'leave the tests alone for now' };

What the placement buys is that a steer is never a surprise:

  • At a round boundary, never mid-batch. The thunk is called with nothing in flight — no stream open, no tool call dispatched, no group half settled. A steer can therefore never land between an assistant turn carrying tool_calls and the tool messages that answer them, which is the malformed conversation every provider rejects (see /The limit ordering, and the assistant turn it discards for the other half of the same rule).

  • Before everything that reads the conversation. The caps and the preflight weigh the steer, a compaction can move it, RoundStarted's token figure counts it, and the request carries it. A steer is not a sidecar on the request — it is history, and the transcript records it as one more user turn with nothing to mark it out.

  • On the round the limit path restarts, too. When a tool limit switches tools off and the loop goes round again for a final answer, that round asks for steers like any other. It is the same round boundary.

  • Including the first round, before the model has said anything. The queue is normally empty there — the run was started from the user's question a moment ago — and an app with something in it at that point gets what it asked for: a second user turn after the first.

  • And while the run is parked, once per park-poll. A run waiting on background work may be waiting for minutes, and a user locked out of it for the duration would be locked out of exactly the situation they most want a word in. A pull that comes back with something ends the park and starts a round, and what it took is carried into that round: the pull is destructive, so a park that dropped it would swallow what the user typed. Nothing about the placement rules above changes — a parked run has nothing in flight by definition.

Three things are the app's job, and the loop does not help with any of them:

  • The queue's thread safety. The thunk is only ever called on the driving thread, at a quiescent point, one call at a time — but whatever it reads is shared with whatever the UI pushes onto, and that is the app's lock (or Channel) to get right.

  • Coalescing. The loop records exactly what it is handed: three answers are three user messages, in order. An app that would rather send one paragraph joins them itself.

  • Noticing delivery. There is no Steered event and no acknowledgement, because the thunk already is one: the loop asked, and what the app handed over is on its way. Rendering it is the ordinary < $session.messages > / AssistantMessage path.

A thunk that throws is shielded: the failure becomes a Log event (level error, logger llm-agent.loop), the round carries on with no steers, and the next round asks again. An entry that is not a defined Str is dropped the same way. A broken queue cannot take a run down.

Background operations: the run that does not end yet

A tool call that answers immediately and does the work afterwards is the difference between an agent that delegates and one that waits. The catch is that the loop's terminal condition — the model stopped asking for tools — is exactly wrong for it: the model stopped asking because it was told the answer would arrive later, and a run that ended there would end before the answer it promised.

completion-bus is the fix, and it is the whole of the fix. LLM::Agent::CompletionBus holds two things: the operations that have been acknowledged and not yet reported, and the reports themselves. The loop asks it one question at the terminal and one at every round boundary.


    my $bus  = LLM::Agent::CompletionBus.new;
    my $loop = LLM::Agent::Loop.new(
        :@backends,
        provider       => $subagents,   # given the same bus: see its Pod
        completion-bus => $bus,
    );

Without a bus none of this exists. No park, no drain, no events, no behaviour change of any kind — which is what makes the option safe to add to a loop that has always worked one way.

The two moments

  • At the top of every round, whatever is on the bus is drained and appended as user turns, before the caps, the compaction and the request — the same placement a steer gets, and for the same reason. Completions go in before steers; see /A round, step by step.

  • At the no-tool-calls terminal, the bus is asked whether it is quiet: nothing outstanding and nothing queued, read as one snapshot. Quiet means the run really is over, and it ends exactly as it always did. Not quiet means it parks.

What a park is

A park is the run waiting, with nothing in flight, on a bare timer. It emits RunParked with the inventory of what it is waiting for, fires on-park, and then looks at four things every park-poll seconds:

What it finds | What it does
something queued | resumes; the next round delivers it
nothing outstanding either | ends the run: RunCompleted, as the terminal would have
the run was cancelled | closes every outstanding op, then RunCancelled
a steer | resumes, carrying the steer into the round
nothing, for park-idle-timeout | the safety valve — see below

Every exit closes the park's span, fires on-unpark, and emits RunResumed with the reason and how long it waited. Those two hooks are balanced on every path, including a throw, because a host that lends a concurrency slot for the duration of a park needs it back.

A bare timer, not a wait. Every poll in this class is written that way to keep Promise.anyof from accumulating continuations (see the note in !stream), and a park gets something more out of it: there is no lost-wakeup race to lose, because there is no wakeup. A completion that lands between two passes is found by the second one.

The turns a completion arrives as

An injected turn is an ordinary user message with an extraordinary first line. It says, in words, that it is an automated event and not the user speaking, and it names the operation it is about — because a conversation is compacted, and a turn that only made sense next to the acknowledgement it answers becomes an orphan the first time that acknowledgement is summarised away.

Its content goes through the observation excerpt seam before the message is built — the same seam a tool result goes through, the same < request-budget.max-observation-size >, the same artifact file beside the transcript. A child that answers with a megabyte cannot land whole in a user turn; what lands is an excerpt, and the full bytes are in the artifact the marker names. The framing itself is never excerpted: it is what makes the turn readable as an event rather than as a person.

The transcript line carries extras — < injected => <kind> >, the operation's id, and whatever the producer added (completion-of, the originating call id) — and those extras are the durable delivery marker: a resume can tell a completion that was delivered from one that is still owed without replaying a thing.

The caps, and what "wall clock" now means

max-wall-clock bounds the time the run spent working, which is elapsed time less every second it spent parked. A run that delegated three children and waited twenty minutes for them did not work for twenty minutes, and a cap that said it did would kill runs for being patient — which, for an agent whose whole job is delegating, caps the one thing it is for. The result Map reports parked-seconds beside wall-clock whenever a run parked at all, so the two can be added back up.

park-idle-timeout is the real-time bound, and it is the safety valve rather than a deadline: half an hour, by default, in which not one thing arrived. Every arrival resets it. When it fires, the run closes every outstanding operation, ends with RunFailed and < reason => 'park-idle' >, and names the operations that never answered — because "whether they took effect is unknown" is the honest thing to say about work whose producer went away, and the one thing somebody reading that failure needs to know is which piece it was.

What a park does not do

  • It does not increment max-tool-rounds. A round that was woken by a completion and asks for no tools is not a tool round, and a run that spent its round budget on notifications would run out of them for reasons that have nothing to do with tools.

  • It does not re-enter itself. A wake round is an ordinary round: cancel check, drain, steers, caps, compaction, request. If the model goes quiet again with work still outstanding, it parks again — a fresh RunParked, a fresh inventory, a fresh idle clock.

  • It does not produce the work. The loop opens nothing and settles nothing; it reads the bus and drains it. Who acknowledges an operation and who reports it are the composer's business (LLM::Agent::Subagents) or the host's.

Resuming: what "the same message" means

A run given a session that already holds messages has to extend that transcript rather than rewrite it, so the messages it was handed are checked against the ones the session replayed, index by index. The check is a digest (LLM::Agent::Canonical) over everything that survives the round trip — role, content, tool calls, tool-call id, sticky, sysprompt, depth — and not a comparison of role and text.

That matters because the differences that break a resumed conversation are exactly the ones prose does not show: an assistant turn that has grown a tool_calls array, a tool result answering a different call id, a system prompt that is no longer sticky. A run whose prefix does not match ends immediately, with a RunFailed naming the index it disagreed at and nothing written to the session; the fix is always to start from < $session.messages >.

Runtime context: the half that is not history

< run(@messages, context => $context) > takes an LLM::Agent::RunContext: today's date, the working directory, the branch, the project's instruction file — everything that is true now rather than everything that happened. It is rendered into the request and nowhere else:


    what goes on the wire   [ head, |@conversation, tail ]
    what everything else sees              @conversation

The wire view is built in exactly one place — the line that calls chat-completion-stream — and is never stored, counted or compared. @conversation itself never contains the context, which is what keeps all of the following true and unchanged:

  • the seed check compares the same messages it always did, so a run may carry a completely different context from the one the transcript was written with (that is the point of the whole mechanism);

  • the session records the same message lines, in the same order;

  • the compactor is handed the conversation, never a request;

  • < RunStarted.message-count > and < %outcome<messages> > count history, not framing.

What is recorded is one run-context envelope per run that had one (see LLM::Agent::Session), written after the seed check — a run that is refused leaves no line — and shielded, so a transcript that cannot take an audit record never fails a run that would have worked.

What the context costs, and where that number goes

The rendered head and tail are weighed once through the loop counter for the ordinary proactive-compaction trigger. Every preflight and provider-length verdict, however, calls the selected profile counter's count-request (LLM::Agent::TokenCount) over conversation, context and tool catalogue together. An Exact counter can therefore render the model's real request template; an older counter safely composes its conversation and text counts.

The proactive figure travels to the compaction consult as < needs-compaction(@conversation, :$tokens) >, so the compactor weighs the trigger against the whole request while still working on the conversation alone. It is deliberately not passed to the compactor as a target: Compactor's no-op path reports exhausted against the target it was given, and a reduced one would turn a compaction with nothing to drop into a context-exhausted terminal.

For forced compaction, the selected profile's complete request count is split into conversation and non-history weight in that same counter's units. Its usable conversation target travels with the counter that produced it, and the compactor uses the override for every before/after, summary and trim check. Targets in different tokenizers are never compared as raw integers; fallback then re-preflights normally in its own units.

The calibration, across a context change

A calibration is "P prompt tokens for these N messages", and P was billed for a complete request to one model. Before preflight the loop compares the selected counter's model, context digest and canonical tool catalogue with the shape it previously described. A change calls counter.invalidate; rewritten history is detected by Usage's prefix digest, while an unchanged prefix extended with new turns keeps the calibration. The next usage-bearing attempt re-calibrates the selected profile only.

Tool declarations are fetched once

tools-for-llm is called once per run, not once per round. A round is a network call, and asking an MCP server to re-list its catalogue before each of them doubles the round trips for a list that changes approximately never. A server that adds a tool mid-run is not noticed until the next run.

SEE ALSO

LLM::Agent::Run, LLM::Agent::Event (especially the attempt-framing contract), LLM::Agent::ToolOperation (what a tool call's states mean), LLM::Agent::Session, LLM::Agent::Compactor, LLM::Agent::RunContext (what a run is told about now), LLM::Agent::RequestBudget (what fits and what it may cost), LLM::Agent::Artifacts (the excerpt and the file behind it), LLM::Agent::CompletionBus (the work a parked run is waiting for), LLM::Agent::Subagents (a provider that spawns child runs, and the consumer of emit-external), LLM::Chat::Retry, MCP::Client::Policy.

  • the bus first, because that is what the run is here for — and state is one atomic snapshot, so "something is deliverable" and "nothing is outstanding and nothing is queued" are read of the same instant. Split into two reads they describe two different worlds, and the operation that settled between them is an answer nobody hears;

  • the cancel next, and before the steer pull, which is a deliberate departure from the round top's order. Pulling steers is destructive — the app's queue hands them over — and a cancelled run that pulled would swallow what the user typed and then end without saying it. Nothing is lost by checking the bus first: reading it is not draining it, and what is on it stays on it for the next run;

  • the steer pull, which is how a user gets a word in while a fleet of children works. What it takes goes into the caller's stash;

  • the idle valve last, because it is the only one of the four that is a guess. ) method !park-for-completions( $run, @conversation, @attempts, @stashed, Int:D $round, Str:D $final, --> Hash:D ) { my @inventory = self!bus-inventory($run);

LLM::Agent v0.6.1

a streaming agent loop: tools, retry, fallback, a durable

Authors

  • Matt Doughty

License

Artistic-2.0

Dependencies

Digest::SHA256::Native:ver<1.0.0+>:auth<zef:bduggan>LLM::Chat:ver<0.10.0+>:auth<zef:apogee>MCP::Client:ver<0.5.0+>:auth<zef:apogee>JSONL:ver<0.1.6+>:auth<zef:apogee>JSON::Fast:ver<0.19+>:auth<cpan:TIMOTIMO>UUID::V4:ver<1.0.0+>:auth<zef:masukomi>

Test Dependencies

Provides

  • LLM::Agent
  • LLM::Agent::Artifacts
  • LLM::Agent::Canonical
  • LLM::Agent::Compactor
  • LLM::Agent::CompletionBus
  • LLM::Agent::Event
  • LLM::Agent::Loop
  • LLM::Agent::Prompt
  • LLM::Agent::RequestBudget
  • LLM::Agent::Run
  • LLM::Agent::RunContext
  • LLM::Agent::Session
  • LLM::Agent::Subagents
  • LLM::Agent::TokenCount
  • LLM::Agent::ToolOperation

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite — the markup and publishing tools behind this site.