RequestBudget
NAME
LLM::Agent::RequestBudget - what fits in this backend, and what this run is allowed to spend
SYNOPSIS
use LLM::Agent::RequestBudget;
my $budget = LLM::Agent::RequestBudget.new(
# One profile per backend, keyed by what that backend calls its model
# ā the same string LLM::Agent::TokenCount::Usage calibrates against.
profiles => {
'moonshotai/kimi-k2' => LLM::Agent::RequestBudget::Profile.new(
context-window => 128_000,
completion-reserve => 8_000,
),
'local' => LLM::Agent::RequestBudget::Profile.new(
context-window => 8_192,
),
},
# For a backend with no profile of its own.
default-profile => LLM::Agent::RequestBudget::Profile.new(
context-window => 32_000,
),
# A tool result bigger than this is excerpted into the conversation
# and spilled whole to an artifact file. See LLM::Agent::Artifacts.
max-observation-size => 16_384,
# Run caps. Undefined (the default) means no cap at all.
max-cost => 2.50,
max-total-tokens => 400_000,
max-wall-clock => 1800,
);
my $loop = LLM::Agent::Loop.new(:@backends, :$provider, request-budget => $budget);
DESCRIPTION
Two questions with one owner: will this request fit the backend it is about to be sent to, and has this run spent more than it was allowed to. They are one object because they are the same kind of fact ā a limit that belongs to the deployment rather than to the conversation ā and because both are read at the same two moments (the top of a round, and the boundary between two tool calls).
Nothing here talks to a backend, counts a token or ends a run. LLM::Agent::Loop does all three; this is the thing it asks.
A profile per backend, keyed by model
%.profiles is keyed by what a backend calls its model ā exactly the
string LLM::Agent::TokenCount's Usage stores its calibration under,
so a profile and a calibration cannot end up describing different things
under the same name. A backend with no profile falls back to
default-profile, and a backend with neither is not preflighted at
all: an unknown window is not a small one, and refusing to send a
request because nobody said how big the window is would be worse than
sending it.
A profile is four numbers:
| Field | What it is |
|---|---|
| context-window | the whole window, in tokens: prompt and completion |
| completion-reserve | tokens kept back for the answer (see below) |
| input-safety-margin | slack for template framing the counter cannot see |
| counter | an LLM::Agent::TokenCount for this model only |
completion-reserve is undefined by default, and is then taken from
the backend at lookup time: < $backend.settings.max_tokens >, or 1024
when there is no answer to be had. That is deliberate ā the number of
tokens a backend will be asked to produce is already configured on the
backend, and carrying it twice creates a way for the two to disagree.
Note what this means in practice: LLM::Chat's Settings defaults
max_tokens to 256, so a deployment that never set it gets a 256-token
reserve, which is honest about what that backend will actually generate.
counter exists for heterogeneous chains. Two models with different
tokenizers do not agree about how big a conversation is, and a chain that
falls back from a 128k model to an 8k one is exactly where the difference
matters; a profile with no counter of its own uses the loop's.
The preflight arithmetic
LLM::Agent::Loop asks, per backend, before it spends an attempt:
request = counter.count-request(
@conversation, :@tools, :$context-head, :$context-tail,
)
needed = request + input-safety-margin
it fits when needed + completion-reserve <= context-window
The conversation, runtime context and tool declarations are all counted
through the selected profile's counter because all three are real tokens
on the wire and tokenizers disagree about JSON and template framing as well
as prose. TokenCount::Exact delegates the complete request to a tokenizer
that supports get-request-count; older counters compose their existing
count-messages and count-text answers. This is repeated for every
fallback backend, in that backend's units. A catalogue also disappears when
the loop switches tools off after a limit, which is precisely the case where
a request that did not fit may suddenly fit.
A forced compaction carries the chosen profile's target and counter as a pair. Targets from different tokenizers are compared only by the fraction of their currently counted conversation they would retain; the chosen target is then used only with the counter that produced it. After compaction every backend is preflighted again in its own units.
Circuit breakers, and what "spent" means
The run limits are circuit breakers checked at safe boundaries, not transactional hard caps. A request or tool group already in flight is allowed to settle before the next check, so reported spend can finish above a limit; the guarantee is that no later operation is started after the breaker trips.
max-cost, max-total-tokens and max-wall-clock are undefined by
default, and undefined means no cap. cap-tripped compares them
against a spend record the loop accumulates, and answers with
< { cap, spent, max } > for the first one that has been reached.
Reached, not exceeded: a cap is a ceiling on what the run may have
spent, so a run that has spent exactly its cap has nothing left and is
stopped. The alternative ā stopping only once the cap is past ā means
every cap is quietly the cap plus one more round trip, which for
max-cost is the round trip somebody set a cap to avoid.
cost is only ever counted when a provider reported one (OpenRouter
does; most do not), and a run whose backends report nothing spends
0 ā so a max-cost against a backend that does not price its calls
never trips. That is the honest behaviour: an unreported cost is not a
zero cost, and inventing a number to compare against would be worse than
having none.
Known limitation, 0.2: a resumed run starts its accumulators at zero. The transcript records what each turn cost, but adding them up across processes is runtime-store territory rather than transcript territory, and the loop deliberately does not go there.
max-observation-size lives here, not on a profile
It is a character count, not a token count, and it is a property of the budget as a whole rather than of any one backend. The reason is timing: the decision to excerpt an enormous tool result is made when that result settles, which is before the loop knows which backend the next round will use ā see LLM::Agent::Artifacts. A per-profile threshold would have to be evaluated against a backend nobody has chosen yet.
The context-budget shorthand
< LLM::Agent::Loop.context-budget > is still the one-number way in,
and a loop given one and no explicit budget synthesizes
for-window($context-budget): a default profile of exactly that window,
with no safety margin and no completion reserve.
Those two zeroes are the point. context-budget is a statement about
the window and nothing else ā it is the same number the compactor is
built with, and the compactor already reserves headroom for the answer
through its ratios (it targets 40% of the budget and triggers at 80%).
Subtracting a margin and a reserve from it as well would double-count
that headroom, and on a small window a 512-token margin is most of the
window. What the shorthand buys is the thing that was missing: a
conversation that does not fit the window at all now ends the run
cleanly instead of being sent and rejected. A deployment that wants a
tight preflight states the numbers per model, because those numbers are
per model.
SEE ALSO
LLM::Agent::Loop (the only consumer), LLM::Agent::Artifacts (what
max-observation-size triggers), LLM::Agent::TokenCount (the
counting seam and the calibration key), LLM::Agent::Compactor (the
other half of "it does not fit").