TokenCount

NAME

LLM::Agent::TokenCount - how big is this conversation, three ways

SYNOPSIS


use LLM::Agent::TokenCount;

# The default: calibrate against what the provider actually billed.
my $counter = LLM::Agent::TokenCount::Usage.new;

$counter.count-messages(@messages);          # estimate, before any call

# ... after a successful completion, tell it what the provider charged:
$counter.record-usage(
    prompt-tokens     => $resp.prompt-tokens,
    completion-tokens => $resp.completion-tokens,
    message-count     => @conversation.elems,
);

$counter.count-messages(@conversation);      # now exact for that prefix

# No provider usage to lean on? Pure arithmetic, no dependencies:
my $rough = LLM::Agent::TokenCount::Heuristic.new;

# Have the real tokenizer? Exact, at the cost of loading it:
my $exact = LLM::Agent::TokenCount::Exact.new(
    counter => LLM::Chat::TokenCounter.new(:$tokenizer, :$template),
);

DESCRIPTION

Compaction needs a number: is this conversation close enough to the context budget to be worth summarizing? The number does not need to be perfect — it needs to be available, cheap, and never wildly under. Different deployments can afford different things, so counting is a seam rather than a function.

The role

One required method, and four with an answer already:

  • count-messages(@messages -- Int)> — required. How many tokens this conversation is worth.

  • record-usage(:$prompt-tokens, :$completion-tokens, :$message-count, :$prefix-digest, :$backend -- Nil)> — a no-op by default. The loop calls it after every successful attempt with what the provider reported, plus the identity of what was billed: the digest of the conversation prefix (LLM::Agent::Canonical's messages-digest) and the model that charged for it. An implementation that can learn from it does, and the others ignore it.

  • count-text(Str:D $text -- Int)> — how big a lump of text is, weighed as one system message. This is what an LLM::Agent::RunContext is counted with: the context is rendered into the request and is not part of the conversation, so weighing it through count-messages would hand a stateful counter an array that is not the one it is calibrated against. The default answers through count-messages and is right for any counter that is a pure function of its input; Usage overrides it (see below), and so must any other implementation that keeps state per call.

  • count-request(@messages, :@tools, :$context-head, :$context-tail -- Int)> — the complete request the provider tokenizes. The default keeps older counters source-compatible by composing count-messages with count-text over each defined context half and the compact JSON tool catalogue. Exact delegates to the wrapped tokenizer's get-request-count when it provides one, retaining the component fallback for older tokenizer objects. Usage treats calibrated provider prompt tokens as the whole request already and estimates only an appended history tail, so context and tools are never charged twice.

  • invalidate(-- Nil)> — throw away whatever has been learned. A no-op on Heuristic and Exact, which learn nothing. Before each backend preflight the loop compares the selected counter's request shape: model, runtime context and tool catalogue. A change invalidates the calibration; rewritten or compacted history is detected independently by the calibrated prefix digest. Appended history deliberately keeps the calibration and is estimated as a delta.

The loop owns exactly one counter instance and hands the same one to its compactor, so the calibration a run accumulates is not thrown away at the moment it matters most.

Which one to use

Implementation Needs Accuracy
Usage nothing (calibrates itself) exact for the billed prefix, estimated beyond
Heuristic nothing ±20% on English prose, worse on code/CJK
Exact a tokenizer + a template exact, always

Usage is the default because it needs no tokenizer, costs nothing per call, and converges: after one round trip its answer for everything the model has already seen is the provider's own number, and only the messages added since are estimated.

Exact is worth it when the tokenizer is loaded anyway — but note it re-tokenizes the whole conversation on every call, which on a long chat inside a per-round compaction check is not free.

How Usage's prefix-plus-delta works

record-usage remembers one record: the message count N and the prompt tokens P the provider charged for it, the digest of the messages it charged for, and which model charged. P covers the whole request whose history prefix is @messages[0 ..^ N] — including runtime context, tools and the provider's own template framing, precisely the parts a heuristic is worst at. The loop therefore invalidates it when those non-history parts change.

count-messages(@messages) then answers:

Situation Answer
nothing recorded yet fallback estimate of the whole conversation
@messages is the calibrated prefix P
@messages extends it, unchanged P plus a fallback estimate of the tail
@messages[0 ..^ N] is something else fallback estimate of the whole conversation
@messages.elems < N fallback estimate of the whole conversation

The last two rows are the same rule twice, and both are destructive: detecting that the calibration describes something else clears it.

The shrink is compaction — the working array just went from 40 messages to 6, so P, charged for 40, describes a conversation that no longer exists. Keeping it would be a live bug rather than a stale one: the loop appends messages after compacting, and the moment the array grew back to 40 entries the counter would confidently apply P to a completely different 40 messages and under-report by the size of everything that was summarized away.

The digest is the same failure without the size change. One counter is shared by a loop and its compactor, and nothing stops an application sharing one across runs; "at least N messages" is not evidence that these are those N messages. A second conversation of the same length would otherwise inherit the first one's billed prefix — a number that has nothing to do with it. So the first N messages are digested (see LLM::Agent::Canonical) and compared, and a mismatch falls back exactly as a shrink does.

A calibration recorded without a digest — a caller that did not say what was billed — is taken at its word, because the alternative is to throw away perfectly good information from an application that predates the argument.

The very next successful attempt re-calibrates, so the window in which the estimate is rough is one round.

Which backend billed

record-usage also records the model that reported the usage, and a report from a different one drops the calibration rather than leaving it in force. A fallback backend is not the primary with a different name: it has its own tokenizer, its own template framing and its own idea of what a system prompt costs, so P from one is not an estimate of the other. That matters most in the case that would otherwise be silent — the fallback answered and reported no usage at all (a local model, a mock), which leaves nothing to replace the calibration with. An honest fallback estimate beats the primary's number applied to a conversation the primary is not serving.

Callers that pass no backend are unaffected: provenance that is not stated cannot be compared, and nothing is dropped.

completion-tokens is remembered (last-completion-tokens) but plays no part in counting: those tokens are already inside the assistant message that got appended, so counting them again would double them.

A record-usage call missing either prompt-tokens or message-count is ignored for calibration — a backend that reports no usage (a local model, a mock) leaves the previous calibration alone rather than destroying it. Negative values are ignored for the same reason.

All of Usage's state is behind a Lock, so a run that reports usage from its stream thread while a UI asks for a count from another does not read a half-updated pair.

What Heuristic actually counts

Per message: per-message-overhead plus ceiling(chars / chars-per-token), where chars is the message content plus, for a message carrying tool calls, the JSON encoding of those calls (a tool call is real tokens on the wire, and an assistant turn that asked for three of them with long arguments is not an empty message).

The defaults — 4 characters per token, 8 tokens of overhead per message — are the usual English-prose approximation plus room for the role tags and separators every chat template adds. Tune them per model family if you care; both are constructor arguments.

Deliberately not counted: the tool-call ids and the template's own preamble. Both are small, fixed, and swamped by per-message-overhead — and the whole point of this implementation is that it is arithmetic, not a model of a tokenizer.

SEE ALSO

LLM::Chat::TokenCounter (what Exact wraps), LLM::Agent::Compactor (the main consumer of the answer).

LLM::Agent v0.6.1

a streaming agent loop: tools, retry, fallback, a durable

Authors

  • Matt Doughty

License

Artistic-2.0

Dependencies

Digest::SHA256::Native:ver<1.0.0+>:auth<zef:bduggan>LLM::Chat:ver<0.10.0+>:auth<zef:apogee>MCP::Client:ver<0.5.0+>:auth<zef:apogee>JSONL:ver<0.1.6+>:auth<zef:apogee>JSON::Fast:ver<0.19+>:auth<cpan:TIMOTIMO>UUID::V4:ver<1.0.0+>:auth<zef:masukomi>

Test Dependencies

Provides

  • LLM::Agent
  • LLM::Agent::Artifacts
  • LLM::Agent::Canonical
  • LLM::Agent::Compactor
  • LLM::Agent::CompletionBus
  • LLM::Agent::Event
  • LLM::Agent::Loop
  • LLM::Agent::Prompt
  • LLM::Agent::RequestBudget
  • LLM::Agent::Run
  • LLM::Agent::RunContext
  • LLM::Agent::Session
  • LLM::Agent::Subagents
  • LLM::Agent::TokenCount
  • LLM::Agent::ToolOperation

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite — the markup and publishing tools behind this site.