RequestBudget

NAME

LLM::Agent::RequestBudget - what fits in this backend, and what this run is allowed to spend

SYNOPSIS


use LLM::Agent::RequestBudget;

my $budget = LLM::Agent::RequestBudget.new(
    # One profile per backend, keyed by what that backend calls its model
    # — the same string LLM::Agent::TokenCount::Usage calibrates against.
    profiles => {
        'moonshotai/kimi-k2' => LLM::Agent::RequestBudget::Profile.new(
            context-window     => 128_000,
            completion-reserve => 8_000,
        ),
        'local' => LLM::Agent::RequestBudget::Profile.new(
            context-window => 8_192,
        ),
    },
    # For a backend with no profile of its own.
    default-profile => LLM::Agent::RequestBudget::Profile.new(
        context-window => 32_000,
    ),

    # A tool result bigger than this is excerpted into the conversation
    # and spilled whole to an artifact file. See LLM::Agent::Artifacts.
    max-observation-size => 16_384,

    # Run caps. Undefined (the default) means no cap at all.
    max-cost        => 2.50,
    max-total-tokens => 400_000,
    max-wall-clock  => 1800,
);

my $loop = LLM::Agent::Loop.new(:@backends, :$provider, request-budget => $budget);

DESCRIPTION

Two questions with one owner: will this request fit the backend it is about to be sent to, and has this run spent more than it was allowed to. They are one object because they are the same kind of fact — a limit that belongs to the deployment rather than to the conversation — and because both are read at the same two moments (the top of a round, and the boundary between two tool calls).

Nothing here talks to a backend, counts a token or ends a run. LLM::Agent::Loop does all three; this is the thing it asks.

A profile per backend, keyed by model

%.profiles is keyed by what a backend calls its model — exactly the string LLM::Agent::TokenCount's Usage stores its calibration under, so a profile and a calibration cannot end up describing different things under the same name. A backend with no profile falls back to default-profile, and a backend with neither is not preflighted at all: an unknown window is not a small one, and refusing to send a request because nobody said how big the window is would be worse than sending it.

A profile is four numbers:

Field What it is
context-window the whole window, in tokens: prompt and completion
completion-reserve tokens kept back for the answer (see below)
input-safety-margin slack for template framing the counter cannot see
counter an LLM::Agent::TokenCount for this model only

completion-reserve is undefined by default, and is then taken from the backend at lookup time: < $backend.settings.max_tokens >, or 1024 when there is no answer to be had. That is deliberate — the number of tokens a backend will be asked to produce is already configured on the backend, and carrying it twice creates a way for the two to disagree. Note what this means in practice: LLM::Chat's Settings defaults max_tokens to 256, so a deployment that never set it gets a 256-token reserve, which is honest about what that backend will actually generate.

counter exists for heterogeneous chains. Two models with different tokenizers do not agree about how big a conversation is, and a chain that falls back from a 128k model to an 8k one is exactly where the difference matters; a profile with no counter of its own uses the loop's.

The preflight arithmetic

LLM::Agent::Loop asks, per backend, before it spends an attempt:


	request = counter.count-request(
	              @conversation, :@tools, :$context-head, :$context-tail,
	          )
	needed = request + input-safety-margin

    it fits when   needed + completion-reserve <= context-window

The conversation, runtime context and tool declarations are all counted through the selected profile's counter because all three are real tokens on the wire and tokenizers disagree about JSON and template framing as well as prose. TokenCount::Exact delegates the complete request to a tokenizer that supports get-request-count; older counters compose their existing count-messages and count-text answers. This is repeated for every fallback backend, in that backend's units. A catalogue also disappears when the loop switches tools off after a limit, which is precisely the case where a request that did not fit may suddenly fit.

A forced compaction carries the chosen profile's target and counter as a pair. Targets from different tokenizers are compared only by the fraction of their currently counted conversation they would retain; the chosen target is then used only with the counter that produced it. After compaction every backend is preflighted again in its own units.

Circuit breakers, and what "spent" means

The run limits are circuit breakers checked at safe boundaries, not transactional hard caps. A request or tool group already in flight is allowed to settle before the next check, so reported spend can finish above a limit; the guarantee is that no later operation is started after the breaker trips.

max-cost, max-total-tokens and max-wall-clock are undefined by default, and undefined means no cap. cap-tripped compares them against a spend record the loop accumulates, and answers with < { cap, spent, max } > for the first one that has been reached.

Reached, not exceeded: a cap is a ceiling on what the run may have spent, so a run that has spent exactly its cap has nothing left and is stopped. The alternative — stopping only once the cap is past — means every cap is quietly the cap plus one more round trip, which for max-cost is the round trip somebody set a cap to avoid.

cost is only ever counted when a provider reported one (OpenRouter does; most do not), and a run whose backends report nothing spends 0 — so a max-cost against a backend that does not price its calls never trips. That is the honest behaviour: an unreported cost is not a zero cost, and inventing a number to compare against would be worse than having none.

Known limitation, 0.2: a resumed run starts its accumulators at zero. The transcript records what each turn cost, but adding them up across processes is runtime-store territory rather than transcript territory, and the loop deliberately does not go there.

max-observation-size lives here, not on a profile

It is a character count, not a token count, and it is a property of the budget as a whole rather than of any one backend. The reason is timing: the decision to excerpt an enormous tool result is made when that result settles, which is before the loop knows which backend the next round will use — see LLM::Agent::Artifacts. A per-profile threshold would have to be evaluated against a backend nobody has chosen yet.

The context-budget shorthand

< LLM::Agent::Loop.context-budget > is still the one-number way in, and a loop given one and no explicit budget synthesizes for-window($context-budget): a default profile of exactly that window, with no safety margin and no completion reserve.

Those two zeroes are the point. context-budget is a statement about the window and nothing else — it is the same number the compactor is built with, and the compactor already reserves headroom for the answer through its ratios (it targets 40% of the budget and triggers at 80%). Subtracting a margin and a reserve from it as well would double-count that headroom, and on a small window a 512-token margin is most of the window. What the shorthand buys is the thing that was missing: a conversation that does not fit the window at all now ends the run cleanly instead of being sent and rejected. A deployment that wants a tight preflight states the numbers per model, because those numbers are per model.

SEE ALSO

LLM::Agent::Loop (the only consumer), LLM::Agent::Artifacts (what max-observation-size triggers), LLM::Agent::TokenCount (the counting seam and the calibration key), LLM::Agent::Compactor (the other half of "it does not fit").

LLM::Agent v0.6.1

a streaming agent loop: tools, retry, fallback, a durable

Authors

  • Matt Doughty

License

Artistic-2.0

Dependencies

Digest::SHA256::Native:ver<1.0.0+>:auth<zef:bduggan>LLM::Chat:ver<0.10.0+>:auth<zef:apogee>MCP::Client:ver<0.5.0+>:auth<zef:apogee>JSONL:ver<0.1.6+>:auth<zef:apogee>JSON::Fast:ver<0.19+>:auth<cpan:TIMOTIMO>UUID::V4:ver<1.0.0+>:auth<zef:masukomi>

Test Dependencies

Provides

  • LLM::Agent
  • LLM::Agent::Artifacts
  • LLM::Agent::Canonical
  • LLM::Agent::Compactor
  • LLM::Agent::CompletionBus
  • LLM::Agent::Event
  • LLM::Agent::Loop
  • LLM::Agent::Prompt
  • LLM::Agent::RequestBudget
  • LLM::Agent::Run
  • LLM::Agent::RunContext
  • LLM::Agent::Session
  • LLM::Agent::Subagents
  • LLM::Agent::TokenCount
  • LLM::Agent::ToolOperation

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite — the markup and publishing tools behind this site.