Canonical
NAME
LLM::Agent::Canonical - one stable rendering of a message, a conversation or a lump of plain data, and the digest of it
SYNOPSIS
use LLM::Agent::Canonical;
# "Is this the same message the session recorded here?" ā all of it, not
# just the prose.
message-digest($mine) eq message-digest($theirs);
# "Is this conversation still the one the provider billed for?"
my $prefix = messages-digest(@conversation);
$counter.record-usage(
prompt-tokens => 812, message-count => @conversation.elems,
prefix-digest => $prefix, backend => $backend.model,
);
# "Have the grants changed since the last time I wrote them out?"
data-digest($policy.grants) ne $written;
# "Is this the same tool call the model made two rounds ago?" ā key
# order and whitespace are not semantics.
canonical-arguments('{"b":2, "a":1}') eq canonical-arguments('{"a":1,"b":2}');
DESCRIPTION
Several parts of the loop need to answer "is this the same thing as that?" about data that arrives from a model, a provider or a file, where the two copies are equal in every way that matters and different in ways that do not: a JSON object whose keys came back in another order, a Hash that has been round-tripped through a transcript, a conversation prefix rebuilt from a session.
Identity by eqv is no use there (a decoded Hash never equals a
hand-built one), and identity by prose ā comparing role and
content, which is what the loop used to do ā is worse than no check
at all, because it says "the same" about an assistant turn that has
grown three tool calls.
So: one canonical rendering, and a SHA-256 of it.
What canonical means here
canonical-json renders plain data with sorted keys and no
whitespace, after normalising it:
AssociativeandPositionalvalues are walked recursively, so nesting is canonicalised all the way down;Str,Int,Rat,NumandBoolare kept as they are;an undefined value becomes JSON
null;anything else ā a DateTime, an object a provider invented ā is stringified rather than thrown at, because these functions sit on the loop's hot path and one that dies takes a run with it.
That last rule is what makes this safe to call on whatever a tool provider hands back. It is not a general-purpose serialiser: it is a comparison key, and two things that render the same string are being treated as the same thing.
Why not Message.get-checksum
LLM::Chat::Conversation::Message has a get-checksum, and it is the
wrong tool for this:
it hashes
to-hash, which omitssticky,syspromptanddepthā the three things that decide whether a compaction is allowed to summarise a message away;it
.rakus a Hash, whose iteration order is randomised per process, so the same message can hash differently in two processes;it memoises into an
is rwattribute, so a Message that is mutated after its first checksum keeps answering with the old one.
The functions here have none of that state: same input, same digest, every process, every time.
The digests
| Function | Over |
|---|---|
| data-digest($value) | the canonical JSON of any plain data |
| message-digest($msg) | role, content, tool-calls, tool-call-id, sticky, sysprompt, depth |
| messages-digest(@msgs) | the per-message digests, in order |
message-digest covers everything that survives the session round
trip (see LLM::Agent::Session's message payload) and nothing that
does not ā :%extra is deliberately absent, because reasoning and
usage are not part of the conversation a model sees.
messages-digest is a digest of digests rather than of the whole
concatenation, which keeps it cheap to reason about: two conversations
share a digest exactly when they are the same messages in the same order.
SEE ALSO
LLM::Agent::Loop (seeds a session by digest), LLM::Agent::TokenCount (calibrates against one), LLM::Agent::Session.