Canonical

NAME

LLM::Agent::Canonical - one stable rendering of a message, a conversation or a lump of plain data, and the digest of it

SYNOPSIS


use LLM::Agent::Canonical;

# "Is this the same message the session recorded here?" — all of it, not
# just the prose.
message-digest($mine) eq message-digest($theirs);

# "Is this conversation still the one the provider billed for?"
my $prefix = messages-digest(@conversation);
$counter.record-usage(
    prompt-tokens => 812, message-count => @conversation.elems,
    prefix-digest => $prefix, backend => $backend.model,
);

# "Have the grants changed since the last time I wrote them out?"
data-digest($policy.grants) ne $written;

# "Is this the same tool call the model made two rounds ago?" — key
# order and whitespace are not semantics.
canonical-arguments('{"b":2, "a":1}') eq canonical-arguments('{"a":1,"b":2}');

DESCRIPTION

Several parts of the loop need to answer "is this the same thing as that?" about data that arrives from a model, a provider or a file, where the two copies are equal in every way that matters and different in ways that do not: a JSON object whose keys came back in another order, a Hash that has been round-tripped through a transcript, a conversation prefix rebuilt from a session.

Identity by eqv is no use there (a decoded Hash never equals a hand-built one), and identity by prose — comparing role and content, which is what the loop used to do — is worse than no check at all, because it says "the same" about an assistant turn that has grown three tool calls.

So: one canonical rendering, and a SHA-256 of it.

What canonical means here

canonical-json renders plain data with sorted keys and no whitespace, after normalising it:

  • Associative and Positional values are walked recursively, so nesting is canonicalised all the way down;

  • Str, Int, Rat, Num and Bool are kept as they are;

  • an undefined value becomes JSON null;

  • anything else — a DateTime, an object a provider invented — is stringified rather than thrown at, because these functions sit on the loop's hot path and one that dies takes a run with it.

That last rule is what makes this safe to call on whatever a tool provider hands back. It is not a general-purpose serialiser: it is a comparison key, and two things that render the same string are being treated as the same thing.

Why not Message.get-checksum

LLM::Chat::Conversation::Message has a get-checksum, and it is the wrong tool for this:

  • it hashes to-hash, which omits sticky, sysprompt and depth — the three things that decide whether a compaction is allowed to summarise a message away;

  • it .rakus a Hash, whose iteration order is randomised per process, so the same message can hash differently in two processes;

  • it memoises into an is rw attribute, so a Message that is mutated after its first checksum keeps answering with the old one.

The functions here have none of that state: same input, same digest, every process, every time.

The digests

Function Over
data-digest($value) the canonical JSON of any plain data
message-digest($msg) role, content, tool-calls, tool-call-id, sticky, sysprompt, depth
messages-digest(@msgs) the per-message digests, in order

message-digest covers everything that survives the session round trip (see LLM::Agent::Session's message payload) and nothing that does not — :%extra is deliberately absent, because reasoning and usage are not part of the conversation a model sees.

messages-digest is a digest of digests rather than of the whole concatenation, which keeps it cheap to reason about: two conversations share a digest exactly when they are the same messages in the same order.

SEE ALSO

LLM::Agent::Loop (seeds a session by digest), LLM::Agent::TokenCount (calibrates against one), LLM::Agent::Session.

LLM::Agent v0.6.1

a streaming agent loop: tools, retry, fallback, a durable

Authors

  • Matt Doughty

License

Artistic-2.0

Dependencies

Digest::SHA256::Native:ver<1.0.0+>:auth<zef:bduggan>LLM::Chat:ver<0.10.0+>:auth<zef:apogee>MCP::Client:ver<0.5.0+>:auth<zef:apogee>JSONL:ver<0.1.6+>:auth<zef:apogee>JSON::Fast:ver<0.19+>:auth<cpan:TIMOTIMO>UUID::V4:ver<1.0.0+>:auth<zef:masukomi>

Test Dependencies

Provides

  • LLM::Agent
  • LLM::Agent::Artifacts
  • LLM::Agent::Canonical
  • LLM::Agent::Compactor
  • LLM::Agent::CompletionBus
  • LLM::Agent::Event
  • LLM::Agent::Loop
  • LLM::Agent::Prompt
  • LLM::Agent::RequestBudget
  • LLM::Agent::Run
  • LLM::Agent::RunContext
  • LLM::Agent::Session
  • LLM::Agent::Subagents
  • LLM::Agent::TokenCount
  • LLM::Agent::ToolOperation

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite — the markup and publishing tools behind this site.