Compactor
NAME
LLM::Agent::Compactor - summarize the middle of a conversation before it outgrows the context window
SYNOPSIS
use LLM::Agent::Compactor;
use LLM::Agent::TokenCount;
my $counter = LLM::Agent::TokenCount::Usage.new;
my $compactor = LLM::Agent::Compactor.new(
backend => @backends[0], # a cheap model is a fine summarizer
counter => $counter, # the SAME instance the loop uses
context-budget => 128_000,
);
# The loop does this for you; this is what it does.
if $compactor.needs-compaction(@conversation) {
my %result = $compactor.compact(@conversation, cancelled => { $run.is-cancelled });
@conversation = %result<messages>.list;
note "compacted {%result<dropped>} messages, "
~ "{%result<tokens-before>} -> {%result<tokens-after>} tokens"
~ (%result<fallback> ?? ' (HARD TRIM ā the summarizer was unreachable)' !! '');
}
DESCRIPTION
A tool-using agent fills a context window fast: one fs_read of a
600-line file is most of a small model's budget, and a round of six of
them is most of a large one's. The choices are to stop, to truncate
blindly, or to trade the middle of the conversation for a summary of it.
This is the third.
The shape is fixed and deliberately boring:
[ system prompt + anything sticky ] kept, always, wherever it was
[ ...the middle... ] replaced by ONE summary message
[ the last keep-recent turns ] kept verbatim
Nothing about that shape is negotiable at runtime, because
LLM::Agent::Session has to be able to reproduce it exactly from the
transcript months later. compact and session replay implement the same
transformation, and the cut-index in the result is what ties them
together.
The cut, and why it moves
The recent window starts keep-recent messages from the end ā and is
then extended backwards while its first message is a tool result,
so an assistant turn that asked for three tools and the three results
that answered it are never separated. Splitting them produces a
conversation that most providers reject outright ("tool_call_id did not
match a preceding tool_calls"), so this is a correctness rule and not a
nicety.
Everything before the cut that is not sticky (sticky, sysprompt
or depth ā Message.is-sticky) is what gets summarized. Sticky
messages stay exactly where they were: an agent's own instructions are
the last thing that should be compressed into "the user wanted some
refactoring".
If there is nothing between the sticky prefix and the recent window, the
compaction is a no-op ā dropped is 0, no backend call is made, and
the conversation comes back unchanged. That happens when keep-recent
is bigger than the conversation, and it means the budget is too small
for the window rather than that anything went wrong.
Observation aging: the cheap first resort
Summarizing costs a request, a model and a wait. Most of what fills an
agent's context does not need any of them: it is old tool results. A
600-line fs_read from twenty rounds ago is bulk the model has already
extracted what it wanted from, and the odds it goes back to re-read the
raw bytes are close to nil ā it would re-run the call instead, which is
cheaper for it and cheaper for everybody.
So with < age-observations => True >, a compaction pass tries elision
before it tries summarization: the content of a fat tool message is
replaced, in place, with a deterministic one-line stub ā
[elided by the harness: 41062-char fs_read result; re-run the call if you need it again]
ā and nothing else about the message changes. Same position, same
role, same tool_call_id. That is the whole reason this is safe:
there is no pair alignment to get wrong, no assistant turn left holding
tool_calls that nothing answers, and no provider with an opinion about
it. Compare dropping the turn, which is all of those risks at once.
The two ways a pass gets entered
A reclaimable-batch epoch.
needs-compactionanswers True ā below the trigger ā as soon as the eligible pool reachesage-reclaim-mincharacters. That pass is elision only: it never summarizes, never calls a backend, and never drops a message. One epoch reclaims one batch, and the pool is empty afterwards, so the next epoch is a long way off.The compaction trigger. A conversation over the trigger is aged first, and then recounted. If that was enough ā the conversation is back under the trigger, or under the
:$targeta caller named ā the pass ends there, with no backend call and no summary message. If it was not, the summarizer runs as it always did, over the stubbed middle (so one enormousfs_readno longer dominates the transcript it is asked to summarize, andtool-result-caprarely has anything left to bite on).
Why epochs, and not "a bit every round"
Providers cache the prompt prefix, and a cache hit is an order of
magnitude cheaper than a miss. Any change to the conversation below the
tail invalidates that prefix, so eliding one result per round would pay a
full cache miss every round to reclaim a few thousand characters. A floor
turns that into one miss per batch: the pass waits until there is
age-reclaim-min worth of dead weight and takes it all at once. Elision
epochs are compaction epochs, and an elision-only epoch is strictly the
cheapest kind ā no summarizer request at all.
What is eligible, and what is protected
A message is eligible only if it is a tool result in the middle
region ā after the sticky prefix, before the recent window, the exact
same cut the summarizer works from ā and then:
its content is longer than
age-min-chars. Stubbing something small saves nothing and costs a cache epoch.at least
age-reclaim-mincharacters of conversation content come after it (the trailing grace). The recent window is positional and one tool-heavy round can push eight messages, so "recent" is measured in content instead: the model has to have demonstrably moved a whole epoch's worth of work past a result before its raw bytes are taken away.it is not already a stub, and it is not the newest result of a call identity that appears more than once.
That last rule is the loop breaker. Elision is a bet, and when it loses
the model re-runs the call ā which is fine once and pathological in a
cycle (stub it, it re-reads, the fresh copy ages out, it re-reads again),
because every cycle pays for a full tool result and a cache epoch. So
call identity ā the tool's name plus its canonicalised arguments, joined
back through tool_call_id to the assistant turn that asked ā is
counted: where the same call appears twice, the newest answer is
protected from then on, permanently, and every older copy of it becomes
eligible regardless of size or grace distance (a stale duplicate of a
result that exists in full later in the conversation is pure dead weight).
A result whose assistant turn is not in the conversation any more ā an
earlier compaction summarized it away ā simply has no identity and gets
the ordinary rules.
What the caller gets back
< %result<elisions> > is a list of < { index, stub } >, where
index is a position in the input array. LLM::Agent::Loop turns
those indices into envelope ids and writes one elision line, exactly as
it turns cut-index into a replaces-through-id; LLM::Agent::Session
replays them by id. The stub text is a pure function of the content length
and the tool name, so two runs of the same conversation produce the same
bytes.
Elision changes the conversation, so the usage counter's calibration digest stops matching and the next billed round recalibrates ā the same path a summarizing compaction has always taken. Nothing new is needed for it.
age-observations defaults to False: with it off, every number and
every byte of this class's behaviour is what it was before elision
existed.
What the summarizer is asked for
One blocking chat-completion against $.backend, with a fixed
instruction: goals, decisions, files and commands, and what is still
unresolved ā in that order, no preamble, no invented facts, exact
identifiers over paraphrase. The transcript it summarizes is the middle
rendered as plain text, with every tool result ā and the arguments of
every tool call ā truncated to tool-result-cap characters. Both
truncations leave a marker saying how many characters went, so the
summarizer reads a transcript it can tell was cut rather than one that
looks whole and stops mid-word.
That cap applies only to the summarizer's input. It is there because a
compaction triggered by one enormous fs_read would otherwise send that
same enormous read to the summarizer and fail for exactly the reason it
was invoked. What the model was told during the real conversation is not
altered.
$.backend is a separate attribute rather than "the loop's first
backend" because summarization is a different job: it is not
latency-sensitive, it does not need tools, and a small cheap model does
it well. Pointing it at @backends[0] is perfectly reasonable, and
pointing it somewhere cheaper is usually better.
A compaction has to prove it compacted
compact does not accept a summary just because the model produced
one. After assembling the new conversation it counts it, and requires
C<< tokens-after < tokens-before >>. A model that answers a request for a
summary with something as long as the transcript has not compacted
anything, and taking it would mean the next round asks for a compaction
again over a conversation that has not moved ā forever.
An elision-only pass proves the same thing by construction: it is
returned only when the stubbed conversation is under the number the pass
was aiming at, and a pass with nothing eligible produces no elisions and
falls straight through to the summarizer. The floor is what keeps the
epoch path from spinning ā a pass spends the pool, and needs-compaction
answers False again until a new batch of dead observations has
accumulated.
A summary that fails that test gets one retry, with the instruction saying in as many words that the previous attempt was too long, and the retry is accepted only if it is both smaller than the conversation and smaller than the first attempt. Otherwise it is the hard trim, which is arithmetic and cannot fail to shrink anything.
(The retry tightens with words rather than with a completion cap on
purpose: LLM::Chat keeps that limit on a backend's Settings, which
this compactor normally shares with the loop that owns it ā turning it
down here would turn it down for the conversation as well, from a thread
the conversation knows nothing about.)
When even a hard trim is not enough
There is a floor: the trim never eats into the recent window, so a
conversation whose sticky prefix plus recent window is itself bigger
than the budget cannot be made to fit. One 400KB fs_read in the last
turn does it.
Rather than hand back a conversation whose next request is guaranteed to
fail, compact says so: < exhausted => True > in the result. It is a
marker, not an exception ā the messages are still the best available
version, and LLM::Agent::Loop still writes them to the session before
ending the run with RunFailed and a reason of context-exhausted.
A no-op compaction over a conversation that is already past the budget
(as opposed to merely past the trigger) is marked the same way, for the
same reason: there was nothing to summarize, so there is nothing left to
try.
When the summarizer fails
It will. The context is large, the call is a single one with no fallback chain, and it happens at the least convenient moment. The policy:
The failure is classified with LLM::Chat::Retry's
classify-error. An abort bucket ā 400, 401, 402, 403, 404 ā means a configuration or account problem that will not heal in eight seconds, so it goes straight to the hard trim without burning a retry.Anything else is retried, up to
max-attemptsin total (default 3), withretry-backoffwaits served throughsleep-with-cancelso a cancelled run does not sit out the last one. There is one backend here, soretry-sameandadvanceare the same action; only the wait distinguishes them, and both get it.A summarizer that succeeds but returns nothing is treated as a failure of the
responseclass. An empty summary is worse than no compaction: it silently deletes the middle of the conversation.Still failing: hard trim. The oldest non-sticky messages are dropped ā pair-aligned, so tool results never outlive the assistant turn that asked for them ā until the conversation fits in
target-ratio Ć context-budget, and the result comes back with< fallback => True>.
The hard trim is the whole point of the design: the loop always makes progress. A compaction that could fail would turn an unreachable summarizer into a run that can never take another turn, and an agent that stops working because a summarizer is down is worse than an agent that forgot what happened an hour ago.
The trim is bounded by the cut: it will not eat into the recent window,
even if that leaves the conversation above target. Trimming the last few
turns removes the very context the next turn needs, and a conversation
whose sticky prefix plus recent window does not fit the budget is a
configuration problem (keep-recent too large, budget too small) that
silently discarding the user's last message would hide rather than fix.
fallback is on the CompactionDone event and in the transcript, so
"the model seems to have forgotten everything" is answerable after the
fact.
Compacting to somebody else's number
compact normally aims at target ā target-ratio of the budget ā
and reports exhausted against context-budget. Both are this class's
own numbers, and both are wrong for one caller: LLM::Agent::Loop's
per-attempt preflight, which knows something this class does not. A
backend's usable room is its window minus the tokens reserved for the
answer, minus a safety margin, minus the tool declarations the
request carries ā and in a fallback chain it is the largest such
number across the backends that are still worth trying.
< compact(@messages, :$target, :$counter) > is how that number gets in.
:$target overrides target and :$counter overrides $.counter
for the one invocation. The override is required when the target came
from a backend whose tokenizer is not the compactor's normal tokenizer:
a number produced in one tokenizer's units must never be proved, trimmed
or reported in another's. The loop's override is also allowed to be a
complete-request adapter: it keeps that backend's runtime context and tools
attached while Compactor supplies each candidate conversation. Four
things follow the pair through:
the hard trim aims at it rather than at
target;a summary is accepted only if it gets under it ā a summary that shrank the conversation but left it too big for the backend that asked is not an answer, and falling through to the trim that would have fitted is;
exhaustedis judged against it rather than againstcontext-budget, so "I did everything and it still will not fit" means what the caller asked rather than what this class assumes.every before/after, summary-acceptance and hard-trim count uses the invocation counter, including the counts returned to the caller.
Without the overrides every one of those is exactly as it was.
The result
compact returns a Map:
| Key | Meaning |
|---|---|
| messages | The new conversation. Hand it to the loop; hand it to the model. |
| summary | The synthetic message's full content, header line included. |
| dropped | How many messages were replaced. |
| cut-index | Index, in the INPUT array, of the last message replaced. |
| tokens-before | The counter's answer before. |
| tokens-after | The counter's answer after. |
| fallback | True when this was a hard trim rather than a summary. |
| exhausted | True when even that could not make the conversation fit. |
| elisions | The observations stubbed in place: < [ { index, stub }, ... ] >. |
elisions is empty unless age-observations is on, and its index
values are positions in the input array ā see
/Observation aging: the cheap first resort. An elision-only pass is
the one result that has < dropped => 0 > and a messages array that
is not the input: nothing was replaced, but some of it weighs less.
summary is the whole content rather than just the model's prose so
that a transcript can store one string and replay it into a byte-identical
message. cut-index is what a caller turns into a
replaces-through-id; it is -1 for a no-op compaction.
The synthetic message is < role => 'user' > (the role every provider
accepts arbitrary text in, including the ones that allow exactly one
system message) and is not sticky, so a later compaction can fold it
into a newer summary. Summaries of summaries are how a very long session
stays finite.
The counter
$.counter should be the same instance the loop holds. It is the
loop's counter that has been calibrated against what the provider
actually billed; a fresh one would fall back to arithmetic at the exact
moment an accurate number matters most.
SEE ALSO
LLM::Agent::TokenCount, LLM::Agent::Session (which replays what this produces), LLM::Chat::Retry (the classification and the backoff).