Run

NAME

LLM::Agent::Run - the handle on one agent run: events, result, cancel

SYNOPSIS


my $run = $loop.run(@messages);   # returns immediately; work happens behind it

# Live: one Supply of typed events.
start react whenever $run.events -> $event {
    print $event.text if $event ~~ LLM::Agent::Event::Token;
}

# Or just wait for the answer. The Promise is KEPT even when the run failed.
my %outcome = await $run.result;
given %outcome<outcome> {
    when 'completed' { say %outcome<final> }
    when 'failed'    { note "gave up: {%outcome<error>}" }
    when 'cancelled' { note 'cancelled' }
}

# From anywhere, at any time, as often as you like:
$run.cancel;

DESCRIPTION

Loop.run hands back one of these and gets out of the way. It is three things bolted together — a Supply to watch, a Promise to await, and a cancel button — plus the guarantees that make those three safe to use together.

The result Promise is kept, never broken

Whatever happens, $run.result is kept with a Map:

Key Value
outcome 'completed', 'failed' or 'cancelled'
final the last assistant text (Str, '' when there is none)
messages the conversation as Message objects, as the run left it
error why it failed (Str; undefined otherwise)
attempts the attempt records behind a failure (List; empty otherwise)
reason the named failure behind it (Str; undefined for everything else)
spent what it cost — only when something was counting; see below

The first six keys are always present, so a caller can index blind and branch on outcome.

spent is the exception, and deliberately: it appears only when the loop had an LLM::Agent::RequestBudget to count against, and carries < { total-tokens, wall-clock, complete, cost?, prompt-tokens?, completion-tokens? } > — the last three only when a provider reported them. complete is true only when every provider attempt that started reported either a total or both token halves. Absent spent means nobody was counting, which is a different answer from "nothing was spent", and collapsing the two into a zero would let a dashboard report a free run for a backend that simply does not price its calls.

The same distinction runs through the optional keys: prompt-tokens and completion-tokens are the sums over the attempts that reported a split, and are absent entirely when none did. A run that only knows its total says so by leaving them out.


my %result = await $run.result;
with %result<spent> -> %spent {
    say "total: %spent<total-tokens> tokens";
    say "  in:  %spent<prompt-tokens>"     if %spent<prompt-tokens>:exists;
    say "  out: %spent<completion-tokens>" if %spent<completion-tokens>:exists;
    say "cost:  \$%spent<cost>"            if %spent<cost>:exists;
}

reason is the one to branch on when a failure needs handling rather than reporting: it carries the same string as the RunFailed event's reason — 'context-exhausted', 'budget-exhausted' and 'completion-truncated' are the three so far — and is undefined for the ordinary failures, whose error is all there is to say.

Kept-not-broken is deliberate. A broken Promise makes a failed run into an exception thrown at whoever happened to be awaiting, which means a caller who wanted to inspect the attempts has to catch first and unwrap second, and a caller who wanted to fire-and-forget gets an unhandled Promise warning at GC time for a failure it had already handled through the event stream. A run that failed is a run that finished; it is data.

The same reasoning applies to the Supply: it is done after the terminal event and never quit.

The events Supply preserves

$run.events comes from a Supplier::Preserving, because a scripted run can finish in under a millisecond — long before a test (or a TUI that is still laying out its widgets) gets around to tapping. With a plain Supplier those events are gone; here they are buffered and the first tap receives the whole run from the beginning.

Caveat, and it bites: the buffer is delivered to the first tap only. A second tap arriving after the first has drained it sees only what is emitted from then on — and on a finished run, that is nothing at all, not even the done. So: tap once, and if two consumers need the stream, tap once and fan out yourself. This is why the test kit's drain-events insists it is the only tap.

Cancellation

cancel is idempotent, safe from any thread, and never throws. It does not itself stop anything: it keeps the cancellation Promise, calls the on-cancel hook the driver registered (once), and returns. The driver notices and winds the run down, ending with a RunCancelled event.

Cancelling a run that has already finished is a total no-op. Not "the outcome stands, but the machinery still fires": the cancellation Promise stays Planned, is-cancelled stays False, and the on-cancel hook is not called. That last one is the point. A hook belongs to a driver, a driver holds handles on the things it was driving, and a finished run's driver has already let go of them — so a hook fired after the end would reach for a stream, a backend or a slot that something else now owns. A run that is over cannot be asked to stop; it has stopped.

The drained Promise

result tells you the run has an answer. drained tells you the run has stopped producing: it is kept once the run is closed and every work section the driver opened has closed — the state machine, the token stream, a detached tool batch, an asker or a log hook still inside an _emit.

The two deliberately diverge on cancellation. A cancelled run's result is kept promptly, because the caller asked it to stop and should not wait on a tool batch nobody can interrupt; its drained stays Planned until that batch finally returns. If you are tearing down a loop, an app, or a test fixture and want to know that nothing is still running behind you, drained is the one to await.

drained attests producer quiescence, not delivery. A subscriber that blocks forever in its whenever block parks the run's delivery task and nothing else: result, cancellation and drained all still resolve.

INTERNAL: THE DRIVING SEAM

Everything below is for whoever drives the Run — normally LLM::Agent::Loop, in tests a hand-rolled driver. It is not part of the API a consumer of a Run uses, which is why it is underscore-prefixed.

A Run is the run's serialisation point: it owns the Supplier, the vow on the result Promise, the terminal-once invariant, and the mailbox that turns emissions from any number of threads into one published order. The driver constructs one, drives it with _emit and finishes it with _finish:


my $run = LLM::Agent::Run.new(
    # Called once, on the cancelling thread, the first time cancel() is
    # invoked. Optional, and shielded: an exception from it cannot break
    # cancellation. Use it to poke a blocking wait awake (aborting an
    # in-flight stream at the backend, say) rather than waiting for the
    # next poll.
    on-cancel => { $backend.cancel($current-response) with $current-response },
);

$run._work-begin;   # the driver's own work section, opened up front

start {
    $run._emit: LLM::Agent::Event::RunStarted.new(
        run-id => $run.id, message-count => @messages.elems,
    );

    # ... rounds, attempts, tools; poll $run.is-cancelled, or wait on
    #     $run.cancellation with a ONE-SHOT Promise.anyof — never one
    #     inside a poll loop; see `cancellation` for why ...

    # The epilogue, in this order: let go of everything this run owns —
    # here the handle the cancel hook above pokes — THEN finish it, which
    # is what lets the next run start, THEN close the driver's work
    # section, which is what keeps `drained` honest.
    $current-response = Nil;
    $run._finish:
        LLM::Agent::Event::RunCompleted.new(final => $text),
        final => $text, messages => @conversation;
    $run._work-done;
}

$run;   # handed to the caller immediately

The invariants the Run enforces so the driver cannot get them wrong:

  • _emit refuses terminal events (that is what _finish is for) and silently drops anything emitted after the run has finished, returning False. A late Log event from a tool thread cannot appear after RunCompleted.

  • _finish is once. The second call returns False and changes nothing — so a driver racing its own cancel path against its own success path cannot keep the result Promise twice (which would throw) or emit two terminals.

  • The outcome key of the result Map is derived from the terminal event class, not passed in, so it can never disagree with the event the consumer just saw.

  • _finish keeps the result Promise before it publishes the terminal event, both inside one critical section. So a consumer that reacts to RunCompleted and immediately reads .result finds it Kept — it never has to await what it can already see happened. (This is the reverse of the 0.1.0 ordering, which emitted first and kept after, leaving a window where the terminal was visible and the result was not.)

Publication: the mailbox

Events reach the Supply from one thread, and it is not the driver's. _emit and _finish stamp the event with run-id and the next seq, push it onto the run's mailbox, and return; a single drain task publishes the mailbox in order. Everything that matters follows from that one critical section:

  • The driver may be as multi-threaded as it likes. The state machine, a token tap, a tool thread, an asker and a server's log hook all emit concurrently, and the order the Supply publishes is the order the lock granted — which is exactly the order seq records.

  • Nothing can follow the terminal. _finish closes the run and enqueues the terminal in the same critical section, so a concurrent _emit either got in before (and is published before) or is refused. The Supply's done is sent only when the mailbox is empty and the run is closed, which cannot be observed before the terminal has been published.

  • _emit returning True means accepted and ordered, not delivered. Delivery is the drain task's job, and it happens after the call returned. Wait on drained (or on the Supply's done) if you need "everything has been published"; is-done is the answer to "the run is over".

  • A subscriber that throws loses its own event and nothing else — the same shielding &!on-cancel gets. A subscriber that blocks parks publication only; the result, cancellation and drained Promises are all resolved off the mailbox.

The lock is a strict leaf: nothing called under it emits into the Supply, calls the on-cancel hook, or calls back into the driver.

Work sections, and drained

_work-begin / _work-done bracket anything that may still emit. _work-begin returns False on a closed run — there is nothing left to attest, and the caller must then not call the paired _work-done. When the last section closes on a closed run, drained is kept.

The driver's own section is the outer one, and its _work-done must be the last thing the driver does, after every handle it holds has been released. That is what makes drained mean "this run is not touching anything any more" rather than "the state machine returned".

A driver that dies must still call _finish with a RunFailed — an abandoned Run leaves its result Promise Planned forever, and awaiting it hangs. LLM::Agent::Loop wraps its state machine accordingly.

SEE ALSO

LLM::Agent::Event

LLM::Agent v0.6.1

a streaming agent loop: tools, retry, fallback, a durable

Authors

  • Matt Doughty

License

Artistic-2.0

Dependencies

Digest::SHA256::Native:ver<1.0.0+>:auth<zef:bduggan>LLM::Chat:ver<0.10.0+>:auth<zef:apogee>MCP::Client:ver<0.5.0+>:auth<zef:apogee>JSONL:ver<0.1.6+>:auth<zef:apogee>JSON::Fast:ver<0.19+>:auth<cpan:TIMOTIMO>UUID::V4:ver<1.0.0+>:auth<zef:masukomi>

Test Dependencies

Provides

  • LLM::Agent
  • LLM::Agent::Artifacts
  • LLM::Agent::Canonical
  • LLM::Agent::Compactor
  • LLM::Agent::CompletionBus
  • LLM::Agent::Event
  • LLM::Agent::Loop
  • LLM::Agent::Prompt
  • LLM::Agent::RequestBudget
  • LLM::Agent::Run
  • LLM::Agent::RunContext
  • LLM::Agent::Session
  • LLM::Agent::Subagents
  • LLM::Agent::TokenCount
  • LLM::Agent::ToolOperation

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite — the markup and publishing tools behind this site.