ToolOperation

NAME

LLM::Agent::ToolOperation - one tool call, from "the model asked" to "here is what we know about what happened"

SYNOPSIS


use LLM::Agent::ToolOperation;

my $op = LLM::Agent::ToolOperation.new(
    run-id      => $run.id,
    round       => 3,
    call-id     => 'call_1',                  # the model's tool_calls id
    tool        => 'fs_write',
    arguments   => canonical-arguments($raw), # the comparison key, not the wire text
    idempotency => 'destructive',
);

# The loop dispatches it, and — with a session — the `tool-dispatched`
# envelope's id becomes the operation's id.
$op.dispatch(op-id => $envelope-id);

# ...and settles it, exactly once, whatever came back.
$op.settle(outcome => 'completed', result-digest => data-digest($content));

say $op.state;             # 'completed'
say $op.duration;          # seconds between dispatch and settle

DESCRIPTION

A tool call is the only part of an agent round that changes the world. Everything else — a request, a stream, a summary — can be repeated, retried or thrown away at no cost; fs_write cannot. So a tool call is the one thing the loop tracks as an operation with a state of its own, rather than as a value that either arrives or does not.

The point of the state machine is a single sentence: an operation whose outcome we do not know must never be recorded as one that failed. A local deadline, a cancelled run, a killed process — none of those say anything about whether the remote side effect happened. failed is a claim; outcome-unknown is the truth.

The states

State What it means
proposed the model asked for it; nothing has been sent anywhere
dispatched the provider has it; it is running
completed it answered, and the answer was not an error
failed it answered, and the answer was an error (or was refused)
outcome-unknown we stopped waiting: it may or may not have taken effect
abandoned it was never dispatched, and is known not to have run

Two of those are worth dwelling on.

abandoned is only reachable from proposed. It is the strongest statement the loop can make about a tool call — "this did not happen" — and it is only true of a call the provider was never handed. The moment an operation is dispatched, the only honest terminals are what the provider said (completed / failed) or outcome-unknown.

outcome-unknown is not an error. A run that trips a tool deadline carries on, and the model is told, in the tool message, that the result is unknown rather than that the tool failed. Downgrading it to failed would let a model "retry" a fs_write that had already landed.

authorized is data, not a state

The original design had an authorized state between proposed and dispatched. It is not here, deliberately.

The loop reaches a policy through the provider contract, which never throws: an MCP::Client::Policy that allows a call simply runs it, and one that denies it answers with an is_error result. There is no moment at which the loop is told "this was authorized" — for an allow-verdict call there is nothing to observe at all. A state the machine cannot observe is a state the machine would be lying about.

What the loop can observe is a question being asked, because that comes back through its own wrap-ask shim. So an ask is recorded as time: ask-seconds (and, while a human is still thinking, an open span). That is what makes working-seconds — and therefore the tool deadline — pause for humans instead of punishing them. A denial arrives as an is_error result and settles as failed.

Idempotency is a classification, not a promise

idempotency is one of read-only, idempotent, destructive or unknown, and it is configuration: LLM::Agent::Loop's idempotency-rules map tool-name patterns onto these classes, exactly as an MCP::Client::Policy rule maps a pattern onto a decision. No MCP server publishes anything of the sort — there is no annotation for it anywhere in the protocol — so guessing from the name is the only thing available, and the default is unknown.

It is recorded so that a repair can word itself honestly ("a destructive operation may have taken effect"), and for nothing else. In particular the loop never retries an operation on the strength of it.

Ownership and durability

Instances are owned by the loop, one per call, for the length of a round. They are not the durable record — the session envelopes are (tool-dispatched / tool-settled; see LLM::Agent::Session) — and after a resume the pending operations come back as plain hashes from < $session.pending-tool-operations > rather than as objects. This class is what the live loop reasons with; the transcript is what survives the process.

Every transition is lock-guarded and win-or-lose: dispatch, settle and abandon return True for the caller that made the transition and False for one that arrived too late, rather than throwing. A deadline firing at the same instant as a result arriving is a race the loop is allowed to lose, and it must lose it quietly: whoever gets there first decides what happened, and the other path checks the return value and does nothing.

SEE ALSO

LLM::Agent::Loop (dispatches and settles them), LLM::Agent::Session (the durable envelopes), LLM::Agent::Event (ToolStarted, ToolProgress, ToolAbandoned).

LLM::Agent v0.6.1

a streaming agent loop: tools, retry, fallback, a durable

Authors

  • Matt Doughty

License

Artistic-2.0

Dependencies

Digest::SHA256::Native:ver<1.0.0+>:auth<zef:bduggan>LLM::Chat:ver<0.10.0+>:auth<zef:apogee>MCP::Client:ver<0.5.0+>:auth<zef:apogee>JSONL:ver<0.1.6+>:auth<zef:apogee>JSON::Fast:ver<0.19+>:auth<cpan:TIMOTIMO>UUID::V4:ver<1.0.0+>:auth<zef:masukomi>

Test Dependencies

Provides

  • LLM::Agent
  • LLM::Agent::Artifacts
  • LLM::Agent::Canonical
  • LLM::Agent::Compactor
  • LLM::Agent::CompletionBus
  • LLM::Agent::Event
  • LLM::Agent::Loop
  • LLM::Agent::Prompt
  • LLM::Agent::RequestBudget
  • LLM::Agent::Run
  • LLM::Agent::RunContext
  • LLM::Agent::Session
  • LLM::Agent::Subagents
  • LLM::Agent::TokenCount
  • LLM::Agent::ToolOperation

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite — the markup and publishing tools behind this site.