Subagents

NAME

LLM::Agent::Subagents - a tool provider that spawns child agent runs and answers with what they said

SYNOPSIS


use LLM::Agent::CompletionBus;
use LLM::Agent::Subagents;

my $loop;                       # forward-declared: the composer needs it
my $bus = LLM::Agent::CompletionBus.new;    # optional; see below

my $subagents = LLM::Agent::Subagents.new(
    # Everything the model could already do. The composer publishes this
    # catalogue plus one tool of its own, and forwards every call it does
    # not own to it untouched.
    inner => $policy,

    # What may be spawned. `name` and `description` are required and are
    # what the model is shown; every other key is yours and is handed
    # back to the spawn callback verbatim.
    types => [
        {
            name        => 'reviewer',
            description => 'Reviews a diff and reports problems. Read-only.',
            backends    => @cheap-backends,     # app-owned, passed through
            identity    => 'You are a meticulous code reviewer.',
        },
        {
            name        => 'tester',
            description => 'Runs the test suite and reports what failed.',
            backends    => @backends,
            identity    => 'You run tests and report results.',
        },
    ],

    # The whole of the app's side of the deal: build a child run, hand
    # back something with `.run` and `.session-path`. See THE SPAWN
    # CALLBACK.
    spawn => -> %spec { build-child(%spec) },

    # Deferred, because $loop does not exist yet — the same idiom the
    # policy's on-ask uses. `loop => $loop` here would copy an undefined
    # value and the composer would never find a run to emit into.
    loop    => { $loop },
    session => $session,        # the PARENT's transcript

    # Optional, and it changes everything: with a bus a task call
    # acknowledges and returns, and the answer arrives later as a framed
    # turn. Give the SAME bus to the loop. See BACKGROUND DELEGATION.
    completion-bus => $bus,

    max-live             => 4,
	cancel-grace         => 30,
	drain-grace          => 30,
    max-identical-spawns => 3,

    # The question channel's rails: a parent parked on a child is doing
    # nothing, and a host with a concurrency budget wants to know.
    on-child-park   => -> %info { $slots.suspend(%info<agent-id>) },
    on-child-unpark => -> %info { $slots.resume(%info<agent-id>) },
);

# ... and, from the CHILD's ask handler, on the child's own thread: the
# question goes to the parent model rather than to the human, and the
# answer comes back in the shape an elicitation already speaks.
my %outcome = await $subagents.post-question(
    $agent-id, message => 'Which payments module did you mean?',
);

$loop = LLM::Agent::Loop.new(
    :@backends, provider => $subagents, :$session, completion-bus => $bus,
);

# The parent's stream now carries the children's events too, wrapped.
react whenever $loop.run([$question]).events -> $event {
    given $event {
        when LLM::Agent::Event::Subagent {
            note "[{$event.agent-id}] {$event.inner<kind>}";
        }
        when LLM::Agent::Event::Token { print $event.text }
    }
}

DESCRIPTION

A subagent is a second agent run, started by the model of the first one, with a conversation of its own that the parent never sees. This module is the whole of that mechanism on the LLM::Agent side: a tool provider that stacks over another one, publishes a task tool beside its tools, and answers a task call with the child run's final message.

It publishes two more, task_answer and task_wait, which are the rest of the same protocol: a child that cannot act on its brief asks the agent that wrote it, the task call comes back early carrying the question, and the answer is collected afterwards. See /The question channel; a host that never calls post-question never reaches any of it.

Stacking is the point. MCP::Client::Registry composes several servers into one provider, MCP::Client::Policy composes a provider with a permission model, and this composes a provider with the ability to delegate — all three through the same duck-typed pair (tools-for-llm and execute-tool-calls), so the loop cannot tell them apart and the order they are stacked in is the app's decision:


    Subagents( Policy( Registry( client, client, ... ) ) )
        the model may delegate, and everything it or its children do
        goes through the same permission model

    Policy( Subagents( Registry( ... ) ) )
        delegating is itself a permission question

What it deliberately is not is a scheduler. It does not own a queue, a thread pool, a priority or a notion of what a child costs; it starts a child when the model asks for one, refuses when too many are already running, and waits. An app that wants queueing wraps the spawn callback, which is exactly why that callback is the whole of the seam.

What a task call does, end to end

  • The arguments are read and checked strictly: no key the declaration does not name, an agent-type that is in the table, a prompt that is not empty, an optional label. See /Strict arguments.

  • The two guards run — the identical-spawn digest and the max-live backstop. Either one refusing is an is_error result and not an exception: the model is told, in words, what happened and what to do instead.

  • The spawn callback is called with the spec, and hands back a handle. A subagent-spawned envelope goes to the parent's transcript.

  • The child's event Supply is tapped — this is its only tap — and every event is wrapped in an LLM::Agent::Event::Subagent and published on the parent's stream through < Loop.emit-external >.

  • The child's result is awaited. Its final text becomes the tool result; a failed or cancelled child becomes an is_error result saying so. A subagent-settled envelope records the outcome, an excerpt of what came back, and what the child spent.

  • Unless the child asks something first, in which case the call comes back early with the question and the child carries on — see /The question channel. That is not an error and not the child's answer, and the answer is collected afterwards with task_wait.

  • ...or unless this composer holds a completion-bus, in which case the fourth step does not happen at all: the call acknowledges and returns, and the answer arrives later as a framed turn the loop injects. See /Background delegation, which is the mode most of this Pod's "waits" stop being true in.

THE SPAWN CALLBACK

The one thing this module does not do is build a child. That is deliberate: how a child is configured — which backends, which system prompt, which tools, which transcript, whether it is queued behind others, whether it is even in this process — is entirely the app's, and a composer with an opinion about it would have to grow a dependency on everything an app knows.

So &spawn is called with one Hash and must return a handle:

Spec key What it is
agent-id this child's short unique id ('reviewer-1')
type the type record from types, verbatim, as a Hash
prompt what the model asked for, as it wrote it
label the model's name for this piece of work, or an undefined Str

sub build-child(%spec) {
    my %type = %spec<type>;

    # A transcript of its own. Nothing says a child must have one — an
    # undefined path is fine — but a child that writes into the PARENT's
    # session would interleave two conversations in one file.
    my $path = $sessions-dir.add("{%spec<agent-id>}.jsonl");
    my $child-session = LLM::Agent::Session.create(
        path => $path,
        meta => { agent => %type<name>, label => %spec<label> // '' },
    );

    my $child-loop = LLM::Agent::Loop.new(
        backends => %type<backends>,
        provider => %type<provider>,     # usually NOT this composer:
                                         # see "Children do not delegate"
        session  => $child-session,
    );

    my $context = LLM::Agent::RunContext.new(
        head-sections => [ identity => %type<identity> ],
        facts         => [ date => Date.today.Str, cwd => $*CWD.Str ],
    );

    class { has $.run; has Str $.session-path }.new(
        run => $child-loop.run(
            [LLM::Chat::Conversation::Message.new(
                role => 'user', content => %spec<prompt>,
            )],
            :$context,
        ),
        session-path => $path.Str,
    );
}

The handle is duck-typed, checked with .can, and needs exactly two things:

  • .run — an LLM::Agent::Run, already started. The composer taps it and awaits its result; it never starts anything itself.

  • .session-path — where the child's transcript is, as a Str, for the subagent-spawned envelope. An empty string is fine and means "this child has no transcript"; what is not fine is leaving the parent's reader guessing.

A callback that throws, returns something that is not handle-shaped, or returns a handle whose .run is not a Run is a wiring bug that will happen at three in the morning, so it is not an exception: it is an is_error result naming what was wrong, the child's slot is released, and the parent run carries on.

Children do not delegate, unless you say so

Nothing here stops an app handing the child loop this same composer as its provider — and nothing here would stop the resulting tree from being five levels deep, either. max-live is per composer, so a shared one caps the whole tree; a composer per child caps each level separately and the tree is unbounded. Give a child a provider without a task tool unless recursive delegation is something you want and have bounded.

A composer per delegating child — which is what an app that wants a tree rather than a chain builds — works exactly as the single one does, and the two things that cross levels do so by construction: Event::Subagent nests (a child composer's wrapper becomes the inner of the parent's, and inner is plain data all the way down), and a settled child's spend travels up one level at a time (/A parent pays for its children), so the root's bill is the whole tree's. What does not cross levels is max-live: bound the tree at the seam where the app builds children, not here.

The wrapping contract

A child's events are not merged into the parent's stream: they are wrapped, one LLM::Agent::Event::Subagent per child event, carrying the child's .to-hash as inner. The parent Run stamps the wrapper with its own run-id and seq; the inner hash keeps the child's. See LLM::Agent::Event's Pod for why, and for what a consumer does with it.

Every wrapper also carries call-id: the provider's id for the task call that started this child, constant for its whole life. That is the only thing joining a delegation's tool events to its child events — they are three unrelated ids otherwise — and a UI that draws one card per delegation rather than two needs it.

Three properties follow, and they are the ones worth relying on:

  • The parent's terminal contract is untouched. A Subagent event is never terminal, whatever the child's event was, so a child completing cannot end the parent's Supply.

  • Ordering is the parent's. The wrapper goes through < Loop.emit-external >, which is the same mailbox the loop's own events go through, so a child's event is ordered against the parent's tokens rather than racing them.

  • A late child is dropped, not an error. A child that emits after the parent run has finished — a cancelled parent whose child is still winding down — finds no live run, and emit-external answers False. The "nothing after the terminal" contract wins.

Strict arguments

A task call may carry only the keys its declaration names — agent-type, prompt, label — plus reason. Anything else is an is_error result, checked before agent-type and before prompt, and nothing is started: no child, no slot, no entry in the identical-spawn tally.

That is stricter than this layer is anywhere else, and it is strict because of what the unknown keys turn out to be. A provider generating a turn full of long briefs can hit its completion cap in the middle of one of them and still close the JSON: constrained decoding produces valid syntax at the cap, the finish reason says nothing went wrong, and what was left of the brief is railroaded into keys of the decoder's own invention — a key like 'sh_run. CONTEXT' whose name is a fragment of the sentence that was being written. The call parses. agent-type is there. prompt is there, holding the first hundred characters of a brief that was meant to be three thousand. Ignore the junk and a subagent is started on a sentence and a half, with no way for anyone — child, parent or user — to tell that anything was lost.

So the junk is the diagnosis, not a thing to be dropped. The refusal names the offending keys (flattened and cut, because they can carry newlines and a paragraph of the prompt), names the keys that are allowed, and tells the model what this usually means and what to do: re-emit the whole call, with the complete task in prompt. That is one refused turn against a delegation that would have been quietly wrong.

reason is tolerated whether or not anything is producing it. A stack that has MCP::Client's reasons layer in it adds that parameter to every declaration on the way out and strips it from every call on the way down, so this layer never sees it — but the layer is stacked only when its host turns reasons on, and a model that has been asked for a reason on every call for a whole conversation will write one here too. Refusing a delegation over a habit the harness taught it would be the check doing harm, so the key is allowed by name (a literal, not an import: this module does not depend on that stack).

Nothing else this composer publishes is checked this way, and the task tool is checked this way because it is the one whose arguments are a whole conversation's worth of instructions. A tool whose arguments are a path and a line number cannot be silently gutted by a clip; a brief can.

What the model is told about writing one

The task tool's description carries the doctrine that goes with the check: independent work batched into one turn is how delegations run concurrently, and — since a batch of long briefs is exactly the shape that provokes a clip — a brief is worth its length. It says so: write it in full, and when several long briefs are queued, splitting the spawns across turns is fine. Telling a model to batch without telling it that length is allowed is telling it to compress, and a compressed brief is the same lost context arrived at deliberately.

The guards

Two, and both exist because a model that has discovered delegation will delegate.

The identical-spawn guard counts < (agent-type, prompt) > as a digest, the way LLM::Agent::Loop's identical-call guard counts a call's name and canonicalised arguments: the arguments are reparsed, so key order and JSON whitespace are not part of the identity, and the prompt is trimmed, so a trailing newline is not either. The same spawn more than max-identical-spawns times is refused with an is_error saying so. A spawn that failed to start counts — a model retrying a spawn that cannot work is exactly the loop this is for.

The tally is per parent run by default (identical-spawn-scope), and that default matters: the guard is about a model going round in circles within one turn of work. Three identical delegations across three unrelated runs are three occasions on which somebody asked for the same thing, and a composer that refused the fourth because of what happened an hour ago would get more broken the longer the host stayed up. < identical-spawn-scope => 'composer' > is there for a host that really does mean "this much and no more, ever".

The max-live backstop refuses a spawn while max-live children that can still answer are already running, with a message telling the model to wait for what it started. It is a backstop and not a queue: the refusal is immediate, because a model blocked on an invisible queue looks exactly like a model that has hung.

"Can still answer" is the whole of what changed when wedges became visible: a child whose result is kept and whose drained never comes (drain-grace) is still owned and still cancellable, but it no longer holds a slot — see live-count and owned-count. A composer that let one wedge refuse work for the rest of a host's life would be an accounting detail turning into an outage.

Both are checked and the child's slot is reserved in one critical section, so two spawns arriving at once cannot both take the last slot. That was true before anything arrived at once and it is what makes the next section safe.

Several task calls at once

A batch carrying more than one task call starts all of them together — one thread per call — and reassembles the answers in the caller's order. A task call returns only when its child has settled, so running them one after another would make N delegations take as long as N children in a row however many the host could really run.

What that changes, and what it does not:

  • Order is preserved where order is visible. The results come back indexed by the position of their call, assembled on the calling thread, so the conversation reads exactly as it would have. What is not ordered is which child starts, asks or finishes first — that is the point of the exercise.

  • Every guard still holds. max-live, the identical-spawn tally, the id counter and the slot table are all read and written in one critical section per spawn, so N concurrent calls mint N distinct ids and admit exactly as many children as there was room for. The others are refused in the ordinary words, and the model is told to wait.

  • Every permission question arrives up front. A host that asks a human before a child may start will now be asked N times at once rather than one at a time. That is the visible consequence of concurrency and not a defect — but it is a UX decision, so it is made in the caller: the loop only hands this provider a batch of task calls when the app has named task in LLM::Agent::Loop's concurrent-tools. Left unnamed, calls arrive one at a time and this section describes nothing that happens.

  • Never throws, still. Each call's answer is computed inside its own thread with the same shield the serial code had, and the await that collects it has one of its own.

An app that wants a smaller number of children actually working at once bounds it in the spawn callback — that is what the callback is for (see /THE SPAWN CALLBACK), and a queue there is invisible from here: the task call simply takes longer to answer.

The question channel

A child that is given a brief it cannot act on has three things it can do: guess, give up, or ask. The first is how a delegation goes quietly wrong, the second wastes the run, and the third used to mean asking the human — who did not write the brief, was not asked to review it, and is being interrupted about a conversation they cannot see. The agent that wrote the brief is the one that knows what it meant.

So a host can route a child's question to its parent. This module is the engine half: it parks the question, hands the parent's waiting task call back with it, publishes two tools for answering and collecting, and guarantees that nobody is left blocked whatever happens next. What it deliberately does not do is decide which questions get routed — a child's user_ask about a fact is its parent's business, and a child asking permission to delete something is not (a model must not consent on a human's behalf). That is the host's call, made at the seam where it calls post-question.

The host's side


# in the child's ask handler, on the CHILD's own thread
my %outcome = await $composer.post-question(
    $agent-id, message => $question, schema => %requested-schema,
);

post-question parks the question and answers with a Promise, and the child's thread blocks on it — which is exactly what an elicitation is: a question that suspends the asker. What the Promise is kept with is an elicitation outcome (< { action, content } >), so a host hands it straight back to the server that asked.

Every refusal is a value. An agent-id this composer does not own, one that has already finished, a second question from an agent whose first is still waiting, an empty question: all of them come back as an already-kept < { action => 'cancel' } >. Nothing here throws at a child, and nothing can hand one an accept that nobody wrote. A child that asks into the void carries on, or gives up, in the same way it would if a human had dismissed the prompt.

pending-questions is the snapshot beside it, for a UI drawing "waiting on its parent" against an agent.

One question releases the whole wave

execute-tool-calls answers a batch with one List. That is the provider contract, and it is what decides the shape of everything here: a composer cannot hand back the asking child's call and go on waiting for the others, because there is no way to return half a batch. So a question settles every task call parked in that group at once — a wave:

  • the asking child's call carries the question, in full, with what to do about it;

  • every other parked call says its child is still running and names the tool that collects it;

  • every one of them lists whatever else is still open, so a parent that lost track of a turn has the whole picture in front of it.

The alternative — holding the other calls open — would mean a parent that cannot answer the question until the other children happen to finish, which is the deadlock the channel exists to avoid. Nothing is lost by releasing them: a child whose call came back early is still running, still owned, still cancellable, and still collectable by name.

The same reasoning covers a call that arrives afterwards: while any question is open — delivered or not, from any child — a task or task_wait call comes straight back with an interim result rather than parking. A parent that waits on one child while another is blocked on an answer only it can give is the same deadlock arriving a moment later, and the cost of the blunt rule is a turn spent clearing the table, which the interim result says how to do.

The loop needs to know none of this. An interim result is an ordinary tool result with is_error False whose content starts STATUS: interim; task_answer and task_wait are ordinary tool calls. There is no new event, no new terminal, and no change to the loop at all.

Delivered in full, exactly once

A question is carried to the parent once. The wave that delivers it marks it delivered; every interim result after that mentions it as a line in "questions still waiting" and never repeats the text. A model that is shown the same question three times will answer it three times, or decide it is being ignored and stop — and a long question repeated in every result of every wave is a context window spent on nothing.

Two consequences worth knowing:

  • A question asked while the parent is mid-generation — nothing parked, nobody to release — stays pending and is delivered at the next park, which is whichever task or task_wait call comes next.

  • A question from a child nobody is waiting on is delivered by whichever other waiting call gets there first. That is not an edge case: a child whose call has already come back interim has nobody parked on it, and its next question would otherwise sit pending for ever while interrupting every other wait on every poll.

The two tools

Tool What it does
task_answer answers a parked question: agent-id, and answer / fields / decline
task_wait collects an agent whose task call came back interim

task_answer takes answer (prose) for a question that asked in prose, fields for one that asked for named fields, or < decline => True > with a reason for one that cannot be answered at all. A form is validated against the fields it required: a missing one is an is_error naming it, and the question stays parked — every refusal on this path leaves the child exactly as blocked as it was, because a question dropped on a malformed answer strands the agent that asked it. Answering is not collecting: the child carries on from where it stopped, and its answer arrives through task_wait.

task_wait answers with exactly what the original task call would have: the child's final message, or the reason there is none. It is idempotent — a child that has already been collected answers from a cache, in the same words, however many times a model asks — and it parks on a child that is still going, with the same interim, cancel and unstoppable behaviour the original call had. An id that means nothing here is an is_error that says what that usually means (a restart, or an agent already collected) rather than a bare refusal.

Both are checked against their own declarations by the same strict unknown-key check task uses, and neither is a spawn: they take no slot, mint no id, and cannot trip the identical-spawn guard. A host should, though, name them in LLM::Agent::Loop's identical-call-exempt: collecting the same agent twice is by design, and the loop's identical-call guard cannot tell that from a model going round in circles.

The settle happens once, whoever was waiting

A child can now be waited on more than once, and can also finish with nobody waiting at all. Exactly one of those writes the subagent-settled envelope and absorbs the child's spend:

  • a call that was parked when the child settled records it, and the others answer from the cache. Every way out of a park counts here, not just the terminal one: a call released by a question re-reads the child on its way out and records it if it settled underneath, because the child it left behind is a child the drain that follows will assume somebody else was writing;

  • a child that finishes uncollected records itself as it drains, before its slot is released — otherwise its answer would go with it and the task_wait the parent was told to make would find no such agent.

An interim wave writes no settle envelope and absorbs no spend for a child that is still running: it has not finished, and a transcript that recorded a settle for it would be recording something that did not happen.

Cancelling, and the vows

A parked question is a host thread blocked on a Promise, so every way this composer can stop caring about a child keeps that vow: cancel-children sweeps the whole table (every vow cancel), a child that settles or refuses to stop has its own question swept, and a child that is released takes its question with it.

The parent's terminal — not just its cancellation — is a cascade too. A run that ends with a child still going (the model read the interim result and answered the user instead of answering the child) leaves a child working for a conversation that is over and a host thread blocked on a question nobody will ever see; nothing can reach either of them once the run is finished, so both are ended. When everything settled normally — which is the common case — the hook finds no children and no questions and does nothing at all.

The park rails

on-child-park and on-child-unpark fire when a task or task_wait call starts and stops waiting on a child, with < { agent-id, call-id, tool } > and, on the way out, an outcome of interim, final, unstoppable or error. They are balanced on every path, shielded, and exist for the host's own accounting: a parent parked on a child is doing nothing, which for a host with a concurrency budget is a slot it can lend to somebody else — and the asking child is itself suspended on its question, so the two of them together hold one slot rather than two.

They are also the rails the next thing rides on. Interrupting a wait for a reason other than a question — a steer, a priority change — is the same mechanism with a different predicate, and a host that has wired the callbacks for delegation has already wired them for that.

Background delegation

Everything above describes a task call that waits: the parent's round does not move until the child answers. That is a real cost, and it is the cost that makes delegation not worth doing. A parent that cannot read a file while its fleet works is a parent doing nothing, one slow child wedges a whole turn, and "delegate and carry on" — the entire point of having subagents — is not expressible.

Give this composer a completion-bus and it stops waiting.


my $bus = LLM::Agent::CompletionBus.new;

my $subagents = LLM::Agent::Subagents.new(
    ..., completion-bus => $bus,
);
my $loop = LLM::Agent::Loop.new(
    :@backends, provider => $subagents,
    completion-bus => $bus,          # THE SAME ONE. See below.
);

The same bus, both places. The composer opens operations on it and the loop parks on it; two different buses would mean children tracked on one thing and a run waiting on another, which is a run that ends while its children work — the exact failure the bus exists to prevent.

It cannot be checked at construction (the loop does not exist yet, which is what the deferred loop Callable is for), so it is checked at the first delegation and reported as a Log at error: a composer with a bus whose loop has a different one, or none, says so on every delegation it affects. That is deliberately noisy, because nothing else in the system has a symptom — the child works, the transcript records it, and the conversation simply never hears the answer it was promised.

And with no bus, nothing changes. Not "behaves similarly": a composer without one takes every path it took before background delegation existed, down to the words in every tool description. There is one predicate, and every place the two modes differ asks it.

What a task call does instead

  • The arguments are read, the guards run and the slot is taken — unchanged, all of it.

  • The operation is opened on the bus, before the child is built. Building a child is the app's code and takes as long as it takes; a run that finished in that gap would finish having promised an answer that nothing had started producing yet. A bus at max-outstanding refuses here, and the refusal is a guidance result naming both numbers — the same shape the max-live refusal beside it has, because they are the same situation seen from two tables.

  • The child is spawned and tapped exactly as before, and then the call returns an acknowledgement: a STATUS: started result that says whose it is, where the real answer will appear, and — the sentence that does the work — that finishing the turn is safe. A model will not believe that unless it is told; everything it has been trained on says a turn ends the conversation.

  • The child reports itself through the paths that already existed for a child nobody was waiting on: the drain hook, !record-uncollected, !write-terminal. Nothing new watches a child. The one thing that is new is the last line of !write-terminal — the settle goes onto the bus from inside the once-guard that writes the transcript envelope, so the line in the transcript and the turn the model reads are written by the same winner and cannot disagree about what a child said.

Said once, whichever way round it happened

A child can settle while a task_wait is parked on it, or settle first and be asked for afterwards. Both are ordinary; both must present the answer once.

What happened | Where the answer goes
a task_wait was parked when it settled | that call's result. The bus operation closes COLLECTED and enqueues nothing
it settled with nobody waiting onto the bus, and the loop injects it as a turn
...and a task_wait asked for it after that call's result, and the queued turn is WITHDRAWN

:collected is the whole mechanism, and it withdraws whether or not the operation was still open — which is what makes the third row work, where the operation closed some time ago and the turn is sitting in the queue.

The one case it cannot catch is a model that calls task_wait for a child it has already been shown a turn about: the turn is in the transcript by then, and taking it back is not a thing an append-only transcript does. That is task_wait's documented idempotence — the same words twice, for a model that asked twice — rather than a double delivery, and the background task_wait description exists to make it rare.

A question, without a call to carry it

In blocking mode a question rides home on a parked task call: the wave releases every parked call, and the asking child's carries the text. In background mode there is no parked call to release — the parent is off doing something else, or nothing at all — so the question goes to it as a turn of its own, pushed onto the bus as an untracked deliverable.

Untracked is load-bearing. The child asking a question is stopped, not finished; its operation stays open, the run stays unwilling to end, and the question is news beside it rather than an answer instead of it.

Two things are unchanged, and both of them are the deadlock rules:

  • the question stays in the table until it is ANSWERED. Delivery is not an answer. The pending → delivered transition happens inside the same critical section that parked it, which is what makes "in full, exactly once" true whichever thread gets there;

  • !interrupt-pending stays deliberately blunt. Any open question from any child releases any wait — see its Pod, which is the paragraph to read before touching this. A task_wait on a child whose delivered question is still waiting for an answer must come back interim and not hang, because the parent it is waiting on is the parent that owes the answer.

Where a background run stops

The bus is what a run parks on, and a child that can never answer would park it for ever. Three things stop that, and none of them is a timeout on the work itself:

  • every path that answers the model about a child here and now — a spawn that failed, a child cancelled before it started, one that was asked to stop and would not — closes the operation :collected. The model has been told; a background turn saying it again would read as a second child;

  • a wedged child closes it too. A wedge is a child whose result is kept and whose drained never comes (drain-grace): it already stops counting against max-live, for the reason that it will never answer, and it stops holding a run open for exactly the same one. Its answer exists — the wedge is in the drain, not in the work — so what gets delivered is the real one;

  • and everything else is caught by the loop's own idle valve (park-idle-timeout), which closes what is left and ends the run honestly, naming what never came back.

Cancelling, and the cascade

The first spawn of a batch registers on the parent run's cancellation Promise, so cancelling the parent cancels every child. That is a cascade, not a wait: the parent's own cancel path does not wait for the children, and a child that was mid-tool-call takes as long to wind down as it takes. Every cancelled child's task call settles as an is_error — the batch is never left hanging.

cancel-children is the same thing for a shutdown path, and live-agents is what a UI renders while they run. owned-count and live-count are the two numbers behind that list: everything this composer would cancel, and everything that can still answer.

A parent pays for its children

A child is a whole run of its own, with its own loop and its own bill, and nothing about that reaches the parent's budget by itself: a max-cost on the parent would cap the parent's own turns and be blind to the ten agents it started, which for an agent whose job is delegating is a cap on the cheapest part of what it spends.

So as each child settles, its spent record is handed to the parent loop's absorb-spend (LLM::Agent::Loop) — the same numbers the subagent-settled envelope records, added to the accumulators the parent's own caps are checked against. The parent's budget is thereby the budget of its whole subtree, and recursively so: a composer under a child feeds that child, whose settled record already contains its children's, and so on up.

Three properties worth relying on:

  • It is billed to the run that asked. The absorb is tagged with the parent run captured for this batch, so a child settling after its parent finished — a cancelled parent whose child was still winding down — is dropped rather than charged to whatever run started next.

  • Nothing is refused here. Absorbing does not check a cap and does not end anything; the parent's next round or operation boundary does that, in the ordinary way, with the ordinary budget-exhausted reason. A child that pushes its parent over is refused on the parent's next attempt, not retroactively.

  • A child nobody counted is not billed for zero. A run without a request budget hands back no spent key at all, and that absorbs nothing — "nobody was counting" survives the trip, as it does everywhere else in this class.

A child's life is longer than its call

Three moments, and they are all different:

Moment What it means
child result the task call is ANSWERED; the parent may carry on
child drained the child has stopped PRODUCING; the composer lets go
composer's slot held from admission to drained, not to result

The parent is told as soon as the child has a result, because making a model wait on a call nobody is waiting for is how a run stalls. But the composer keeps the child — its max-live slot, its place in live-agents, and the right to cancel it — until the child's drained Promise is kept.

That gap is not theoretical. A child that abandoned a tool call to a deadline has a result while the abandoned call is still running, still writing files; LLM::Agent::Run is explicit that result and drained diverge exactly there. A composer that let go at result would free the slot to start another child beside the one still working, drop it out of whatever a UI is rendering, and — worst — leave cancel-children with nothing to cancel, so a shutdown would report that everything had stopped while a tool call carried on.

The one thing the slot does not survive is a gap that never closes. Once drain-grace has passed the child is marked wedged: still owned, still in live-agents, still reached by cancel-children — and no longer counted by the max-live admission check, which is live-count. Holding a slot for a child that will never answer would mean one abandoned tool call quietly costing the composer a slot for as long as the host is up.

Where a child's events go, and where they do not

Every event of a child is published through an emitter bound to the run that spawned it, captured before the child exists (LLM::Agent::Loop's emitter-for). Never through a fresh "what is running now?" lookup, because for a child that outlives its parent the answer to that question is the next run:


    run A spawns a child ─┐
    run A is cancelled    │  the child is still winding down
    run A finishes        │
    run B starts          │
                          └─► the child's last events arrive HERE

Published by a lookup, those events land on B: B's transcript grows turns from a conversation it never had, and its seq ordering acquires events with no cause in it. Published through the captured emitter, they are dropped — the emitter answers False once its run is over, for ever. The child's task call still settles (as an is_error when the child was cancelled), because settling belongs to the call and not to the stream.

The same binding covers the two session envelopes and the Log event a failed envelope write produces: everything the composer says about a child belongs to the run that started it.

The window, and why there isn't one

Building a child is somebody else's code — the spawn callback may open a transcript, start a process, or queue behind three other agents — so there is a stretch of time in which a spawn has been admitted and no LLM::Agent::Run exists yet. A cancel arriving in that stretch used to find nothing to cancel and silently do nothing, which stranded the child and hung its task call. It is closed structurally rather than by narrowing:

  • The slot is the cancellation target, not the run. It is taken before the spawn callback is called, and cancel-children writes cancel-requested onto every slot — including the ones with no run yet — under the same lock the spawn path registers its run with.

  • The parent run is captured once per batch, before anything is spawned, and the cascade is registered on it there. .then on an already-kept Promise fires immediately, so "cancelled before the spawn" and "cancelled after it" are one case. (Looking the run up later is what does not work: a cancelled parent has finished by the time its detached tool call gets around to spawning, and Loop.live-run quite correctly answers with nothing for a run that is over.)

  • There are two interception points and they meet in the middle: before the spawn callback is called (the child is never started at all), and in the same critical section that registers the child's run (the child is cancelled the moment there is something to cancel).

So, for a cancel arriving at each point of a child's life:

It arrives What stops the child
before the task call is admitted the pre-spawn check (the captured parent is cancelled): never started
between the guard and the slot same check, one line later: never started
between the slot and the run the flag on the slot; registration reads it and cancels
after the run is registered cancel-children has the run and cancels it
after the result, before drained the slot is still held, so cancel-children still has it
after the child drained nothing to do: it has stopped

There is no ordering in which nothing happens, and every one of them ends with the task call settled — as an is_error for a child that was stopped, and as an ordinary answer for one that had already finished saying it.

A child that is asked to stop and does not is the one case left, and it is a bug in that child rather than a race: a cancelled run keeps its result promptly, so one that has not after thirty seconds is answered without — an is_error saying its outcome is unknown, in the same words the loop uses for a tool call it stopped waiting for. The task call always settles.

What the transcript records

Four envelope types on the parent's session, all through < Session.append-event >, so an older LLM::Agent replays them as unknown types and ignores them:

Type Payload
subagent-spawned agent-id, agent-type, prompt, label?, call-id, child-path
subagent-question agent-id, token, call-id?, message, schema
subagent-answered agent-id, token, call-id, action, content, reason?
subagent-settled agent-id, outcome, result, call-id?, spent?

The two question envelopes are written once each per question — the call-id on a subagent-question is the call that carried it to the parent, which is not the call that started the child, and the one on a subagent-answered is the task_answer that resolved it. message and content are excerpts, for the reason result is.

A question delivered as a background turn has no call-id: nothing carried it, which is what a background delivery is, and inventing one would name a call that never happened.

The call-id on a subagent-settled is a third thing again: the task call that started this child, the same one its spawn envelope records. It is what joins a settle found on disk back to the acknowledgement in the conversation, which is how a result collected while the process was dying is told from one that was lost. Absent for a child whose framing this composer no longer holds — a settle that arrived after the parent run had been replaced — and never invented.

call-id is the provider's id for the task call that started this child, which is what joins the spawn line to the tool-dispatched line a few envelopes above it — and, for a UI replaying a transcript, the tool card to the agent card. A transcript written before the key existed simply has no call-id on its spawn lines, and replays exactly as it always did: nothing here reads the key back, and a reader that wants it treats "absent" as "this pairing is not recorded" rather than as damage.

result is an excerpt (2048 characters) of what the parent model was given, not the child's whole transcript, and child-path is a pointer: replaying the parent needs none of the children's files, and a parent session whose children have been deleted resumes exactly as it would have with them.

Writing either is shielded. A transcript that cannot take an audit record must never be the reason a working tool call fails; the failure becomes a Log event on the run instead.

Wiring: the loop is deferred

The composer needs the loop (to emit into its run) and the loop needs the composer (it is the provider). Forward-declare the loop and hand this class a Callable that returns it:


my $loop;
my $subagents = LLM::Agent::Subagents.new(..., loop => { $loop });
$loop = LLM::Agent::Loop.new(:@backends, provider => $subagents);

< loop => $loop > with $loop still undefined does not work and cannot be made to: the value is copied at construction, and what is copied is Any. set-loop is the same fix for an app that would rather assign than close over:


my $subagents = LLM::Agent::Subagents.new(...);          # no loop yet
my $loop = LLM::Agent::Loop.new(:@backends, provider => $subagents);
$subagents.set-loop($loop);

A composer with no loop still works: it spawns, it waits, it answers. What it cannot do is publish the children's events anywhere, because there is nothing to publish them onto.

Grants, and why there are two classes

The loop synchronises a provider's permission grants to the session only when the provider .can('grants') — so a composer that always had the method would tell the loop to write grants for a stack that has none, and one that never had it would silently break grant persistence for an MCP::Client::Policy underneath it.

So .new returns a LLM::Agent::Subagents::WithGrants — a subclass whose only content is a grants method delegating to the inner provider — when the inner provider has grants, and a plain LLM::Agent::Subagents when it does not. Both are LLM::Agent::Subagents, so nothing that type-checks or dispatches notices; .can('grants') answers honestly either way. Construct through .new, never through .bless.

SEE ALSO

LLM::Agent::Loop (emit-external, the seam this publishes through), LLM::Agent::Event (the Subagent event and its two-layer envelope), LLM::Agent::Run, LLM::Agent::Session (append-event), LLM::Agent::CompletionBus (what a background delegation is tracked on), MCP::Client::Policy and MCP::Client::Registry (the other two composers of the same duck-typed pair).

  • Without one — the default — a task call blocks until its child settles, and everything about this class behaves exactly as it did before background delegation existed. Not approximately: the same waits, the same interim results, the same words in every tool description.

  • With one a task call acknowledges and returns. The child goes on working, the parent model goes on working, and the answer arrives later as a framed turn that LLM::Agent::Loop injects at a round boundary. The run will not end while a child is still owed — that is what the bus is for.

  • it appears the moment a spawn is admitted, before the spawn callback has been called, because that is when the max-live slot is taken. starting is True until there is a child run behind it — those are cancellable exactly like the rest (see cancel-children), they just have nothing to render yet;

  • it survives the task call's answer and disappears only when the child has drained. draining is True in between: the child answered, the parent has been told, and something the child detached — a tool call it abandoned to a deadline — is still running. It is still this composer's to cancel. ) method live-agents(--> List:D) { $!lock.protect: { %!children.values.sort({ $_<seq> }).map({ my Bool $has-run = $_<run> ~~ LLM::Agent::Run:D; %( agent-id => $_<agent-id>, agent-type => $_<agent-type>, label => $_<label>, session-path => $_<session-path>, starting => !$has-run, # Answered, and still producing: its run is done and its # `drained` is not. See cancel-children. draining => $has-run && $_<run>.is-done, wedged => ?($_<wedged> // False), ); }).List; }; }

# in the child's ask handler, on the child's own thread
	    my %outcome = await $composer.post-question(
	        $agent-id, message => $question, schema => %requested-schema,
	    );
	    # -> { action => 'accept', content => { ... } }
	    #    { action => 'decline' } / { action => 'cancel' }

LLM::Agent v0.6.1

a streaming agent loop: tools, retry, fallback, a durable

Authors

  • Matt Doughty

License

Artistic-2.0

Dependencies

Digest::SHA256::Native:ver<1.0.0+>:auth<zef:bduggan>LLM::Chat:ver<0.10.0+>:auth<zef:apogee>MCP::Client:ver<0.5.0+>:auth<zef:apogee>JSONL:ver<0.1.6+>:auth<zef:apogee>JSON::Fast:ver<0.19+>:auth<cpan:TIMOTIMO>UUID::V4:ver<1.0.0+>:auth<zef:masukomi>

Test Dependencies

Provides

  • LLM::Agent
  • LLM::Agent::Artifacts
  • LLM::Agent::Canonical
  • LLM::Agent::Compactor
  • LLM::Agent::CompletionBus
  • LLM::Agent::Event
  • LLM::Agent::Loop
  • LLM::Agent::Prompt
  • LLM::Agent::RequestBudget
  • LLM::Agent::Run
  • LLM::Agent::RunContext
  • LLM::Agent::Session
  • LLM::Agent::Subagents
  • LLM::Agent::TokenCount
  • LLM::Agent::ToolOperation

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite — the markup and publishing tools behind this site.