Retry

NAME

LLM::Chat::Retry - the shared retry/fallback policy primitives

SYNOPSIS


use LLM::Chat::Retry;
use LLM::Chat::Retry::Exceptions;

my @attempts;
my Int $retries-left = 2;

for @backends.kv -> $i, $backend {
    loop {
        my $resp = call-somehow($backend);
        last if $resp.is-success;   # advance on success is the caller's job

        @attempts.push: attempt-record(
            backend-index => $i,
            model         => $backend.model,
            error         => "[backend $i] {$resp.err}",
        );

        given classify-error(
            error-class  => $resp.error-class,
            error-status => $resp.error-status,
        ) {
            when 'abort'      { die "aborting: {$resp.err}" }
            when 'retry-same' {
                last unless $retries-left > 0;
                my Num $wait = retry-backoff(3 - $retries-left);
                $retries-left--;
                # Cancel-aware: returns False the moment the hook flips.
                last unless sleep-with-cancel($wait, cancelled => &cancelled);
                next;
            }
            default { last }        # 'advance'
        }
    }
}

X::LLM::Chat::Retry::Exhausted.new(
    :@attempts, summary => @attempts.map({ "  $_<error>" }).join("\n"),
).throw;

DESCRIPTION

Five small, pure, side-effect-light subs that together are the retry policy LLM::Data::Inference::Task has run in production since 0.5: how to classify a failure, how long to wait before retrying, how to sleep without ignoring a cancel, and what shape the attempt and telemetry records have.

They were factored out of that Task verbatim so a second executor (LLM::Agent::Loop, an app's own chain runner) gets the same behaviour without copying it — and, more importantly, so that fixing the policy fixes it in one place. Everything here is deliberately a free function: no state, no I/O beyond sleep, no knowledge of backends, responses or HTTP clients.

What is NOT here, deliberately

The loop itself. Bucket handling is policy shaped by the caller: the 'abort' bucket dies with a message only the caller can compose, the exhaustion hook's payload differs per layer, hook-shielding is the caller's contract with its own users, and the shape of "advance" depends on what a backend even is in that layer. Trying to share the loop would force one caller's error strings and hook semantics onto every other — so the primitives are shared and the loop is not.

telemetry-payload also does not import LLM::Chat::Backend::Response: it probes whatever it is handed with .?, so a caller can pass a real Response, a subclass with provider extras, a test double, or nothing at all.

Error buckets

classify-error maps a structured failure (as recorded on a Response by _set-error-info) to one of three strings:

Bucket Meaning What the caller should do
abort config / account / access error stop the whole chain now
retry-same probably transient retry THIS backend after a backoff
advance model-specific pathology move to the NEXT backend at once

The full rule table, in evaluation order:

Condition Bucket
parser-failed => True (any class) advance
http 400 / 401 / 402 / 403 / 404 abort
http 429 advance
http 500..599 retry-same
http, any other or missing status advance
timeout advance
connection retry-same
response advance
anything else (incl. no class) retry-same

The two classes worth reading closely are 'timeout' and 'connection', because the same Cro exception can produce either and the difference is which phase it expired in — see below.

The reasoning behind the less obvious rows:

  • parser-failed wins over everything. The transport succeeded; what came back was unusable. A different model may well produce parseable output, the same one just proved it does not — so advance, and never spend the transient-retry budget on it.

  • 429 advances rather than retrying. A rate limit is per-account per-provider; the next backend in the chain is usually a different one and answers immediately, which beats sleeping out someone else's quota.

  • 5xx retries the same backend. Against an aggregator like OpenRouter, a 5xx is one upstream provider failing — the retry frequently routes to a different one and succeeds on the same backend config.

  • An unclassifiable error retries the same backend. The conservative reading: an error nobody has classified yet is more often a hiccup than a permanent property of the model.

  • An 'http' class with no status advances. There is no evidence for a retry, and no evidence it will abort either; moving on is the cheap, safe move.

  • A 'timeout' advances, but a failed connect is not one. 'timeout' means the endpoint accepted the connection and then did not answer inside the deadline. That is a statement about that backend, and the next one in the chain will answer sooner than a backoff would end — so, advance.

A connect that never completed is a different animal, and LLM::Chat::Backend::OpenAICommon files it as 'connection' accordingly: nothing was sent, the endpoint said nothing, and all that is known is that the network was unwell for thirty seconds. Retrying the same backend after a backoff is the correct response, and on a one-backend chain it is the only one that is not "give up" — which is exactly what filing it under 'timeout' used to mean there.

Backoff

retry-backoff($n) is min(2 ** ($n - 1) + jitter, $cap) seconds, with $n the 1-based retry number and jitter drawn from [0, 0.5) when the caller does not supply one:

retry-n range (default cap 30)
1 1.0 .. 1.5 s
2 2.0 .. 2.5 s
3 4.0 .. 4.5 s
4 8.0 .. 8.5 s
5 16.0 .. 16.5 s
6+ 30 s (capped)

The jitter exists because these chains are run by batch workers: N processes that hit the same rate limit in the same second would otherwise all retry in the same second, forever. Half a second of spread is enough to decorrelate them without meaningfully changing the wait. Pass :jitter(0e0) for deterministic tests, or a bigger jitter for bigger fleets.

Sleeping without ignoring a cancel

sleep-with-cancel is the reason a cancelled run does not sit out a 16-second backoff. It sleeps in $chunk slices (default 0.25 s), checking &cancelled before every slice, including the first, and returns False the moment the hook says stop — so worst-case latency between a cancel flipping and the caller noticing is one chunk, not one backoff.

It returns a Bool, it does not throw: only the caller knows which exception type its layer promises (X::LLM::Chat::Retry::Cancelled, a subclass, or something else entirely), so it reports and lets the caller decide.


throw-cancelled unless sleep-with-cancel($wait, cancelled => &cancelled-now);

A &cancelled callback that throws propagates: the callback is the caller's own code, and swallowing an exception from it would silently turn a broken cancel hook into "never cancels".

Record shapes

attempt-record builds the entries of X::LLM::Chat::Retry::Exhausted's @.attempts:

Key Type Presence
backend-index Int always
model Str always
error Str always
raw-text Str ONLY when :raw-text was passed (key absent otherwise)

telemetry-payload builds the per-round-trip hook payload:

Key Presence
attempt, backend-index, model-name, latency-ms always
success, stage always
error, error-class, error-status always (value may be undefined)
prompt-tokens, completion-tokens, total-tokens when :$response exposes them, defined
model-used, finish-reason when :$response exposes them, defined
cost, generation-id, provider-name, is-byok when :$response exposes them, defined

Note the asymmetry, which is the pre-existing contract: the error* keys are always present (undefined on success) so a sink can index them unconditionally, while the response-derived keys are presence-gated so a sink can distinguish "the provider reported 0 completion tokens" from "the provider reported nothing". An empty-string error is normalised to an undefined Str for the same reason.

Response probing is duck-typed .? throughout, so a backend whose Response lacks cost / generation-id / provider-name / is-byok (everything that is not OpenRouter) simply omits them, and a test double needs only the accessors it wants to exercise.

CONSUMERS

  • LLM::Data::Inference::Task — the original home of this code; its classify-error method now delegates here, and its exception types subclass LLM::Chat::Retry::Exceptions'.

  • LLM::Agent::Loop — per-round-trip retry/fallback inside the agent loop, including mid-stream failures.

  • Any caller that wants Task's policy without Task's dependencies.

SEE ALSO

LLM::Chat::Retry::Exceptions, LLM::Chat::Backend::Response (the error-class / error-status pair this classifies).

LLM::Chat v0.10.0

Simple framework for LLM inferencing

Authors

  • Matt Doughty

License

Artistic-2.0

Dependencies

Cro::HTTP::Client:ver<0.8.11+>:auth<zef:cro>Digest::SHA256::Native:ver<1.0.0+>:auth<zef:bduggan>JSON::Fast:ver<0.19+>:auth<cpan:TIMOTIMO>Template::Jinja2:ver<0.3.0+>:auth<zef:apogee>Tokenizers:ver<0.3.0+>:auth<zef:apogee>UUID::V4:ver<1.0.0+>:auth<zef:masukomi>

Provides

  • LLM::Chat
  • LLM::Chat::Backend
  • LLM::Chat::Backend::KoboldCpp
  • LLM::Chat::Backend::Mock
  • LLM::Chat::Backend::OpenAICommon
  • LLM::Chat::Backend::OpenRouter
  • LLM::Chat::Backend::Response
  • LLM::Chat::Backend::Response::OpenRouter
  • LLM::Chat::Backend::Response::OpenRouter::Augment
  • LLM::Chat::Backend::Response::OpenRouter::Stream
  • LLM::Chat::Backend::Response::Stream
  • LLM::Chat::Backend::Settings
  • LLM::Chat::Conversation
  • LLM::Chat::Conversation::Message
  • LLM::Chat::Debug
  • LLM::Chat::Retry
  • LLM::Chat::Retry::Exceptions
  • LLM::Chat::Template
  • LLM::Chat::Template::ChatML
  • LLM::Chat::Template::DeepSeekV4
  • LLM::Chat::Template::Gemma2
  • LLM::Chat::Template::Jinja2
  • LLM::Chat::Template::Llama3
  • LLM::Chat::Template::Llama4
  • LLM::Chat::Template::MistralV7
  • LLM::Chat::TokenCounter
  • LLM::Chat::ToolLoop

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite — the markup and publishing tools behind this site.