Compactor

NAME

LLM::Agent::Compactor - summarize the middle of a conversation before it outgrows the context window

SYNOPSIS


use LLM::Agent::Compactor;
use LLM::Agent::TokenCount;

my $counter = LLM::Agent::TokenCount::Usage.new;

my $compactor = LLM::Agent::Compactor.new(
    backend        => @backends[0],   # a cheap model is a fine summarizer
    counter        => $counter,       # the SAME instance the loop uses
    context-budget => 128_000,
);

# The loop does this for you; this is what it does.
if $compactor.needs-compaction(@conversation) {
    my %result = $compactor.compact(@conversation, cancelled => { $run.is-cancelled });

    @conversation = %result<messages>.list;

    note "compacted {%result<dropped>} messages, "
       ~ "{%result<tokens-before>} -> {%result<tokens-after>} tokens"
       ~ (%result<fallback> ?? ' (HARD TRIM — the summarizer was unreachable)' !! '');
}

DESCRIPTION

A tool-using agent fills a context window fast: one fs_read of a 600-line file is most of a small model's budget, and a round of six of them is most of a large one's. The choices are to stop, to truncate blindly, or to trade the middle of the conversation for a summary of it. This is the third.

The shape is fixed and deliberately boring:


    [ system prompt + anything sticky ]   kept, always, wherever it was
    [ ...the middle... ]                  replaced by ONE summary message
    [ the last keep-recent turns ]        kept verbatim

Nothing about that shape is negotiable at runtime, because LLM::Agent::Session has to be able to reproduce it exactly from the transcript months later. compact and session replay implement the same transformation, and the cut-index in the result is what ties them together.

The cut, and why it moves

The recent window starts keep-recent messages from the end — and is then extended backwards while its first message is a tool result, so an assistant turn that asked for three tools and the three results that answered it are never separated. Splitting them produces a conversation that most providers reject outright ("tool_call_id did not match a preceding tool_calls"), so this is a correctness rule and not a nicety.

Everything before the cut that is not sticky (sticky, sysprompt or depth — Message.is-sticky) is what gets summarized. Sticky messages stay exactly where they were: an agent's own instructions are the last thing that should be compressed into "the user wanted some refactoring".

If there is nothing between the sticky prefix and the recent window, the compaction is a no-op — dropped is 0, no backend call is made, and the conversation comes back unchanged. That happens when keep-recent is bigger than the conversation, and it means the budget is too small for the window rather than that anything went wrong.

Observation aging: the cheap first resort

Summarizing costs a request, a model and a wait. Most of what fills an agent's context does not need any of them: it is old tool results. A 600-line fs_read from twenty rounds ago is bulk the model has already extracted what it wanted from, and the odds it goes back to re-read the raw bytes are close to nil — it would re-run the call instead, which is cheaper for it and cheaper for everybody.

So with < age-observations => True >, a compaction pass tries elision before it tries summarization: the content of a fat tool message is replaced, in place, with a deterministic one-line stub —


    [elided by the harness: 41062-char fs_read result; re-run the call if you need it again]

— and nothing else about the message changes. Same position, same role, same tool_call_id. That is the whole reason this is safe: there is no pair alignment to get wrong, no assistant turn left holding tool_calls that nothing answers, and no provider with an opinion about it. Compare dropping the turn, which is all of those risks at once.

The two ways a pass gets entered

  • A reclaimable-batch epoch. needs-compaction answers True — below the trigger — as soon as the eligible pool reaches age-reclaim-min characters. That pass is elision only: it never summarizes, never calls a backend, and never drops a message. One epoch reclaims one batch, and the pool is empty afterwards, so the next epoch is a long way off.

  • The compaction trigger. A conversation over the trigger is aged first, and then recounted. If that was enough — the conversation is back under the trigger, or under the :$target a caller named — the pass ends there, with no backend call and no summary message. If it was not, the summarizer runs as it always did, over the stubbed middle (so one enormous fs_read no longer dominates the transcript it is asked to summarize, and tool-result-cap rarely has anything left to bite on).

Why epochs, and not "a bit every round"

Providers cache the prompt prefix, and a cache hit is an order of magnitude cheaper than a miss. Any change to the conversation below the tail invalidates that prefix, so eliding one result per round would pay a full cache miss every round to reclaim a few thousand characters. A floor turns that into one miss per batch: the pass waits until there is age-reclaim-min worth of dead weight and takes it all at once. Elision epochs are compaction epochs, and an elision-only epoch is strictly the cheapest kind — no summarizer request at all.

What is eligible, and what is protected

A message is eligible only if it is a tool result in the middle region — after the sticky prefix, before the recent window, the exact same cut the summarizer works from — and then:

  • its content is longer than age-min-chars. Stubbing something small saves nothing and costs a cache epoch.

  • at least age-reclaim-min characters of conversation content come after it (the trailing grace). The recent window is positional and one tool-heavy round can push eight messages, so "recent" is measured in content instead: the model has to have demonstrably moved a whole epoch's worth of work past a result before its raw bytes are taken away.

  • it is not already a stub, and it is not the newest result of a call identity that appears more than once.

That last rule is the loop breaker. Elision is a bet, and when it loses the model re-runs the call — which is fine once and pathological in a cycle (stub it, it re-reads, the fresh copy ages out, it re-reads again), because every cycle pays for a full tool result and a cache epoch. So call identity — the tool's name plus its canonicalised arguments, joined back through tool_call_id to the assistant turn that asked — is counted: where the same call appears twice, the newest answer is protected from then on, permanently, and every older copy of it becomes eligible regardless of size or grace distance (a stale duplicate of a result that exists in full later in the conversation is pure dead weight). A result whose assistant turn is not in the conversation any more — an earlier compaction summarized it away — simply has no identity and gets the ordinary rules.

What the caller gets back

< %result<elisions> > is a list of < { index, stub } >, where index is a position in the input array. LLM::Agent::Loop turns those indices into envelope ids and writes one elision line, exactly as it turns cut-index into a replaces-through-id; LLM::Agent::Session replays them by id. The stub text is a pure function of the content length and the tool name, so two runs of the same conversation produce the same bytes.

Elision changes the conversation, so the usage counter's calibration digest stops matching and the next billed round recalibrates — the same path a summarizing compaction has always taken. Nothing new is needed for it.

age-observations defaults to False: with it off, every number and every byte of this class's behaviour is what it was before elision existed.

What the summarizer is asked for

One blocking chat-completion against $.backend, with a fixed instruction: goals, decisions, files and commands, and what is still unresolved — in that order, no preamble, no invented facts, exact identifiers over paraphrase. The transcript it summarizes is the middle rendered as plain text, with every tool result — and the arguments of every tool call — truncated to tool-result-cap characters. Both truncations leave a marker saying how many characters went, so the summarizer reads a transcript it can tell was cut rather than one that looks whole and stops mid-word.

That cap applies only to the summarizer's input. It is there because a compaction triggered by one enormous fs_read would otherwise send that same enormous read to the summarizer and fail for exactly the reason it was invoked. What the model was told during the real conversation is not altered.

$.backend is a separate attribute rather than "the loop's first backend" because summarization is a different job: it is not latency-sensitive, it does not need tools, and a small cheap model does it well. Pointing it at @backends[0] is perfectly reasonable, and pointing it somewhere cheaper is usually better.

A compaction has to prove it compacted

compact does not accept a summary just because the model produced one. After assembling the new conversation it counts it, and requires C<< tokens-after < tokens-before >>. A model that answers a request for a summary with something as long as the transcript has not compacted anything, and taking it would mean the next round asks for a compaction again over a conversation that has not moved — forever.

An elision-only pass proves the same thing by construction: it is returned only when the stubbed conversation is under the number the pass was aiming at, and a pass with nothing eligible produces no elisions and falls straight through to the summarizer. The floor is what keeps the epoch path from spinning — a pass spends the pool, and needs-compaction answers False again until a new batch of dead observations has accumulated.

A summary that fails that test gets one retry, with the instruction saying in as many words that the previous attempt was too long, and the retry is accepted only if it is both smaller than the conversation and smaller than the first attempt. Otherwise it is the hard trim, which is arithmetic and cannot fail to shrink anything.

(The retry tightens with words rather than with a completion cap on purpose: LLM::Chat keeps that limit on a backend's Settings, which this compactor normally shares with the loop that owns it — turning it down here would turn it down for the conversation as well, from a thread the conversation knows nothing about.)

When even a hard trim is not enough

There is a floor: the trim never eats into the recent window, so a conversation whose sticky prefix plus recent window is itself bigger than the budget cannot be made to fit. One 400KB fs_read in the last turn does it.

Rather than hand back a conversation whose next request is guaranteed to fail, compact says so: < exhausted => True > in the result. It is a marker, not an exception — the messages are still the best available version, and LLM::Agent::Loop still writes them to the session before ending the run with RunFailed and a reason of context-exhausted. A no-op compaction over a conversation that is already past the budget (as opposed to merely past the trigger) is marked the same way, for the same reason: there was nothing to summarize, so there is nothing left to try.

When the summarizer fails

It will. The context is large, the call is a single one with no fallback chain, and it happens at the least convenient moment. The policy:

  • The failure is classified with LLM::Chat::Retry's classify-error. An abort bucket — 400, 401, 402, 403, 404 — means a configuration or account problem that will not heal in eight seconds, so it goes straight to the hard trim without burning a retry.

  • Anything else is retried, up to max-attempts in total (default 3), with retry-backoff waits served through sleep-with-cancel so a cancelled run does not sit out the last one. There is one backend here, so retry-same and advance are the same action; only the wait distinguishes them, and both get it.

  • A summarizer that succeeds but returns nothing is treated as a failure of the response class. An empty summary is worse than no compaction: it silently deletes the middle of the conversation.

  • Still failing: hard trim. The oldest non-sticky messages are dropped — pair-aligned, so tool results never outlive the assistant turn that asked for them — until the conversation fits in target-ratio Ɨ context-budget, and the result comes back with < fallback => True >.

The hard trim is the whole point of the design: the loop always makes progress. A compaction that could fail would turn an unreachable summarizer into a run that can never take another turn, and an agent that stops working because a summarizer is down is worse than an agent that forgot what happened an hour ago.

The trim is bounded by the cut: it will not eat into the recent window, even if that leaves the conversation above target. Trimming the last few turns removes the very context the next turn needs, and a conversation whose sticky prefix plus recent window does not fit the budget is a configuration problem (keep-recent too large, budget too small) that silently discarding the user's last message would hide rather than fix.

fallback is on the CompactionDone event and in the transcript, so "the model seems to have forgotten everything" is answerable after the fact.

Compacting to somebody else's number

compact normally aims at target — target-ratio of the budget — and reports exhausted against context-budget. Both are this class's own numbers, and both are wrong for one caller: LLM::Agent::Loop's per-attempt preflight, which knows something this class does not. A backend's usable room is its window minus the tokens reserved for the answer, minus a safety margin, minus the tool declarations the request carries — and in a fallback chain it is the largest such number across the backends that are still worth trying.

< compact(@messages, :$target, :$counter) > is how that number gets in. :$target overrides target and :$counter overrides $.counter for the one invocation. The override is required when the target came from a backend whose tokenizer is not the compactor's normal tokenizer: a number produced in one tokenizer's units must never be proved, trimmed or reported in another's. The loop's override is also allowed to be a complete-request adapter: it keeps that backend's runtime context and tools attached while Compactor supplies each candidate conversation. Four things follow the pair through:

  • the hard trim aims at it rather than at target;

  • a summary is accepted only if it gets under it — a summary that shrank the conversation but left it too big for the backend that asked is not an answer, and falling through to the trim that would have fitted is;

  • exhausted is judged against it rather than against context-budget, so "I did everything and it still will not fit" means what the caller asked rather than what this class assumes.

  • every before/after, summary-acceptance and hard-trim count uses the invocation counter, including the counts returned to the caller.

Without the overrides every one of those is exactly as it was.

The result

compact returns a Map:

Key Meaning
messages The new conversation. Hand it to the loop; hand it to the model.
summary The synthetic message's full content, header line included.
dropped How many messages were replaced.
cut-index Index, in the INPUT array, of the last message replaced.
tokens-before The counter's answer before.
tokens-after The counter's answer after.
fallback True when this was a hard trim rather than a summary.
exhausted True when even that could not make the conversation fit.
elisions The observations stubbed in place: < [ { index, stub }, ... ] >.

elisions is empty unless age-observations is on, and its index values are positions in the input array — see /Observation aging: the cheap first resort. An elision-only pass is the one result that has < dropped => 0 > and a messages array that is not the input: nothing was replaced, but some of it weighs less.

summary is the whole content rather than just the model's prose so that a transcript can store one string and replay it into a byte-identical message. cut-index is what a caller turns into a replaces-through-id; it is -1 for a no-op compaction.

The synthetic message is < role => 'user' > (the role every provider accepts arbitrary text in, including the ones that allow exactly one system message) and is not sticky, so a later compaction can fold it into a newer summary. Summaries of summaries are how a very long session stays finite.

The counter

$.counter should be the same instance the loop holds. It is the loop's counter that has been calibrated against what the provider actually billed; a fresh one would fall back to arithmetic at the exact moment an accurate number matters most.

SEE ALSO

LLM::Agent::TokenCount, LLM::Agent::Session (which replays what this produces), LLM::Chat::Retry (the classification and the backoff).

LLM::Agent v0.6.1

a streaming agent loop: tools, retry, fallback, a durable

Authors

  • Matt Doughty

License

Artistic-2.0

Dependencies

Digest::SHA256::Native:ver<1.0.0+>:auth<zef:bduggan>LLM::Chat:ver<0.10.0+>:auth<zef:apogee>MCP::Client:ver<0.5.0+>:auth<zef:apogee>JSONL:ver<0.1.6+>:auth<zef:apogee>JSON::Fast:ver<0.19+>:auth<cpan:TIMOTIMO>UUID::V4:ver<1.0.0+>:auth<zef:masukomi>

Test Dependencies

Provides

  • LLM::Agent
  • LLM::Agent::Artifacts
  • LLM::Agent::Canonical
  • LLM::Agent::Compactor
  • LLM::Agent::CompletionBus
  • LLM::Agent::Event
  • LLM::Agent::Loop
  • LLM::Agent::Prompt
  • LLM::Agent::RequestBudget
  • LLM::Agent::Run
  • LLM::Agent::RunContext
  • LLM::Agent::Session
  • LLM::Agent::Subagents
  • LLM::Agent::TokenCount
  • LLM::Agent::ToolOperation

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite — the markup and publishing tools behind this site.