TokenCount
NAME
LLM::Agent::TokenCount - how big is this conversation, three ways
SYNOPSIS
use LLM::Agent::TokenCount;
# The default: calibrate against what the provider actually billed.
my $counter = LLM::Agent::TokenCount::Usage.new;
$counter.count-messages(@messages); # estimate, before any call
# ... after a successful completion, tell it what the provider charged:
$counter.record-usage(
prompt-tokens => $resp.prompt-tokens,
completion-tokens => $resp.completion-tokens,
message-count => @conversation.elems,
);
$counter.count-messages(@conversation); # now exact for that prefix
# No provider usage to lean on? Pure arithmetic, no dependencies:
my $rough = LLM::Agent::TokenCount::Heuristic.new;
# Have the real tokenizer? Exact, at the cost of loading it:
my $exact = LLM::Agent::TokenCount::Exact.new(
counter => LLM::Chat::TokenCounter.new(:$tokenizer, :$template),
);
DESCRIPTION
Compaction needs a number: is this conversation close enough to the context budget to be worth summarizing? The number does not need to be perfect ā it needs to be available, cheap, and never wildly under. Different deployments can afford different things, so counting is a seam rather than a function.
The role
One required method, and four with an answer already:
count-messages(@messages --Int)> ā required. How many tokens this conversation is worth.record-usage(:$prompt-tokens, :$completion-tokens, :$message-count, :$prefix-digest, :$backend --Nil)> ā a no-op by default. The loop calls it after every successful attempt with what the provider reported, plus the identity of what was billed: the digest of the conversation prefix (LLM::Agent::Canonical'smessages-digest) and the model that charged for it. An implementation that can learn from it does, and the others ignore it.count-text(Str:D $text --Int)> ā how big a lump of text is, weighed as one system message. This is what an LLM::Agent::RunContext is counted with: the context is rendered into the request and is not part of the conversation, so weighing it throughcount-messageswould hand a stateful counter an array that is not the one it is calibrated against. The default answers throughcount-messagesand is right for any counter that is a pure function of its input;Usageoverrides it (see below), and so must any other implementation that keeps state per call.count-request(@messages, :@tools, :$context-head, :$context-tail --Int)> ā the complete request the provider tokenizes. The default keeps older counters source-compatible by composingcount-messageswithcount-textover each defined context half and the compact JSON tool catalogue.Exactdelegates to the wrapped tokenizer'sget-request-countwhen it provides one, retaining the component fallback for older tokenizer objects.Usagetreats calibrated provider prompt tokens as the whole request already and estimates only an appended history tail, so context and tools are never charged twice.invalidate(--Nil)> ā throw away whatever has been learned. A no-op onHeuristicandExact, which learn nothing. Before each backend preflight the loop compares the selected counter's request shape: model, runtime context and tool catalogue. A change invalidates the calibration; rewritten or compacted history is detected independently by the calibrated prefix digest. Appended history deliberately keeps the calibration and is estimated as a delta.
The loop owns exactly one counter instance and hands the same one to its compactor, so the calibration a run accumulates is not thrown away at the moment it matters most.
Which one to use
| Implementation | Needs | Accuracy |
|---|---|---|
| Usage | nothing (calibrates itself) | exact for the billed prefix, estimated beyond |
| Heuristic | nothing | ±20% on English prose, worse on code/CJK |
| Exact | a tokenizer + a template | exact, always |
Usage is the default because it needs no tokenizer, costs nothing per
call, and converges: after one round trip its answer for everything the
model has already seen is the provider's own number, and only the
messages added since are estimated.
Exact is worth it when the tokenizer is loaded anyway ā but note it
re-tokenizes the whole conversation on every call, which on a long chat
inside a per-round compaction check is not free.
How Usage's prefix-plus-delta works
record-usage remembers one record: the message count N and the
prompt tokens P the provider charged for it, the digest of the
messages it charged for, and which model charged. P covers the
whole request whose history prefix is @messages[0 ..^ N] ā including
runtime context, tools and the provider's own template framing, precisely
the parts a heuristic is worst at. The loop therefore invalidates it when
those non-history parts change.
count-messages(@messages) then answers:
| Situation | Answer |
|---|---|
| nothing recorded yet | fallback estimate of the whole conversation |
| @messages is the calibrated prefix | P |
| @messages extends it, unchanged | P plus a fallback estimate of the tail |
| @messages[0 ..^ N] is something else | fallback estimate of the whole conversation |
| @messages.elems < N | fallback estimate of the whole conversation |
The last two rows are the same rule twice, and both are destructive: detecting that the calibration describes something else clears it.
The shrink is compaction ā the working array just went from 40 messages
to 6, so P, charged for 40, describes a conversation that no longer
exists. Keeping it would be a live bug rather than a stale one: the loop
appends messages after compacting, and the moment the array grew back to
40 entries the counter would confidently apply P to a completely
different 40 messages and under-report by the size of everything that was
summarized away.
The digest is the same failure without the size change. One counter is
shared by a loop and its compactor, and nothing stops an application
sharing one across runs; "at least N messages" is not evidence that these
are those N messages. A second conversation of the same length would
otherwise inherit the first one's billed prefix ā a number that has
nothing to do with it. So the first N messages are digested (see
LLM::Agent::Canonical) and compared, and a mismatch falls back exactly
as a shrink does.
A calibration recorded without a digest ā a caller that did not say what was billed ā is taken at its word, because the alternative is to throw away perfectly good information from an application that predates the argument.
The very next successful attempt re-calibrates, so the window in which the estimate is rough is one round.
Which backend billed
record-usage also records the model that reported the usage, and a
report from a different one drops the calibration rather than
leaving it in force. A fallback backend is not the primary with a
different name: it has its own tokenizer, its own template framing and
its own idea of what a system prompt costs, so P from one is not an
estimate of the other. That matters most in the case that would otherwise
be silent ā the fallback answered and reported no usage at all (a local
model, a mock), which leaves nothing to replace the calibration with. An
honest fallback estimate beats the primary's number applied to a
conversation the primary is not serving.
Callers that pass no backend are unaffected: provenance that is not
stated cannot be compared, and nothing is dropped.
completion-tokens is remembered (last-completion-tokens) but plays
no part in counting: those tokens are already inside the assistant message
that got appended, so counting them again would double them.
A record-usage call missing either prompt-tokens or
message-count is ignored for calibration ā a backend that reports no
usage (a local model, a mock) leaves the previous calibration alone
rather than destroying it. Negative values are ignored for the same
reason.
All of Usage's state is behind a Lock, so a run that reports usage from its stream thread while a UI asks for a count from another does not read a half-updated pair.
What Heuristic actually counts
Per message: per-message-overhead plus
ceiling(chars / chars-per-token), where chars is the message content
plus, for a message carrying tool calls, the JSON encoding of those calls
(a tool call is real tokens on the wire, and an assistant turn that asked
for three of them with long arguments is not an empty message).
The defaults ā 4 characters per token, 8 tokens of overhead per message ā are the usual English-prose approximation plus room for the role tags and separators every chat template adds. Tune them per model family if you care; both are constructor arguments.
Deliberately not counted: the tool-call ids and the template's own
preamble. Both are small, fixed, and swamped by per-message-overhead ā
and the whole point of this implementation is that it is arithmetic, not
a model of a tokenizer.
SEE ALSO
LLM::Chat::TokenCounter (what Exact wraps), LLM::Agent::Compactor
(the main consumer of the answer).