Retry
NAME
LLM::Chat::Retry - the shared retry/fallback policy primitives
SYNOPSIS
use LLM::Chat::Retry;
use LLM::Chat::Retry::Exceptions;
my @attempts;
my Int $retries-left = 2;
for @backends.kv -> $i, $backend {
loop {
my $resp = call-somehow($backend);
last if $resp.is-success; # advance on success is the caller's job
@attempts.push: attempt-record(
backend-index => $i,
model => $backend.model,
error => "[backend $i] {$resp.err}",
);
given classify-error(
error-class => $resp.error-class,
error-status => $resp.error-status,
) {
when 'abort' { die "aborting: {$resp.err}" }
when 'retry-same' {
last unless $retries-left > 0;
my Num $wait = retry-backoff(3 - $retries-left);
$retries-left--;
# Cancel-aware: returns False the moment the hook flips.
last unless sleep-with-cancel($wait, cancelled => &cancelled);
next;
}
default { last } # 'advance'
}
}
}
X::LLM::Chat::Retry::Exhausted.new(
:@attempts, summary => @attempts.map({ " $_<error>" }).join("\n"),
).throw;
DESCRIPTION
Five small, pure, side-effect-light subs that together are the retry policy LLM::Data::Inference::Task has run in production since 0.5: how to classify a failure, how long to wait before retrying, how to sleep without ignoring a cancel, and what shape the attempt and telemetry records have.
They were factored out of that Task verbatim so a second executor
(LLM::Agent::Loop, an app's own chain runner) gets the same
behaviour without copying it ā and, more importantly, so that fixing
the policy fixes it in one place. Everything here is deliberately a
free function: no state, no I/O beyond sleep, no knowledge of
backends, responses or HTTP clients.
What is NOT here, deliberately
The loop itself. Bucket handling is policy shaped by the caller: the
'abort' bucket dies with a message only the caller can compose, the
exhaustion hook's payload differs per layer, hook-shielding is the
caller's contract with its own users, and the shape of "advance"
depends on what a backend even is in that layer. Trying to share the
loop would force one caller's error strings and hook semantics onto
every other ā so the primitives are shared and the loop is not.
telemetry-payload also does not import
LLM::Chat::Backend::Response: it probes whatever it is handed with
.?, so a caller can pass a real Response, a subclass with provider
extras, a test double, or nothing at all.
Error buckets
classify-error maps a structured failure (as recorded on a Response
by _set-error-info) to one of three strings:
| Bucket | Meaning | What the caller should do |
|---|---|---|
| abort | config / account / access error | stop the whole chain now |
| retry-same | probably transient | retry THIS backend after a backoff |
| advance | model-specific pathology | move to the NEXT backend at once |
The full rule table, in evaluation order:
| Condition | Bucket |
|---|---|
| parser-failed => True (any class) | advance |
| http 400 / 401 / 402 / 403 / 404 | abort |
| http 429 | advance |
| http 500..599 | retry-same |
| http, any other or missing status | advance |
| timeout | advance |
| connection | retry-same |
| response | advance |
| anything else (incl. no class) | retry-same |
The two classes worth reading closely are 'timeout' and
'connection', because the same Cro exception can produce either and
the difference is which phase it expired in ā see below.
The reasoning behind the less obvious rows:
parser-failed wins over everything. The transport succeeded; what came back was unusable. A different model may well produce parseable output, the same one just proved it does not ā so advance, and never spend the transient-retry budget on it.
429 advances rather than retrying. A rate limit is per-account per-provider; the next backend in the chain is usually a different one and answers immediately, which beats sleeping out someone else's quota.
5xx retries the same backend. Against an aggregator like OpenRouter, a 5xx is one upstream provider failing ā the retry frequently routes to a different one and succeeds on the same backend config.
An unclassifiable error retries the same backend. The conservative reading: an error nobody has classified yet is more often a hiccup than a permanent property of the model.
An
'http'class with no status advances. There is no evidence for a retry, and no evidence it will abort either; moving on is the cheap, safe move.A
'timeout'advances, but a failed connect is not one.'timeout'means the endpoint accepted the connection and then did not answer inside the deadline. That is a statement about that backend, and the next one in the chain will answer sooner than a backoff would end ā so, advance.
A connect that never completed is a different animal, and
LLM::Chat::Backend::OpenAICommon files it as 'connection'
accordingly: nothing was sent, the endpoint said nothing, and all that
is known is that the network was unwell for thirty seconds. Retrying
the same backend after a backoff is the correct response, and on a
one-backend chain it is the only one that is not "give up" ā which
is exactly what filing it under 'timeout' used to mean there.
Backoff
retry-backoff($n) is min(2 ** ($n - 1) + jitter, $cap) seconds,
with $n the 1-based retry number and jitter drawn from
[0, 0.5) when the caller does not supply one:
| retry-n | range (default cap 30) |
|---|---|
| 1 | 1.0 .. 1.5 s |
| 2 | 2.0 .. 2.5 s |
| 3 | 4.0 .. 4.5 s |
| 4 | 8.0 .. 8.5 s |
| 5 | 16.0 .. 16.5 s |
| 6+ | 30 s (capped) |
The jitter exists because these chains are run by batch workers: N
processes that hit the same rate limit in the same second would
otherwise all retry in the same second, forever. Half a second of
spread is enough to decorrelate them without meaningfully changing the
wait. Pass :jitter(0e0) for deterministic tests, or a bigger jitter
for bigger fleets.
Sleeping without ignoring a cancel
sleep-with-cancel is the reason a cancelled run does not sit out a
16-second backoff. It sleeps in $chunk slices (default 0.25 s),
checking &cancelled before every slice, including the first, and
returns False the moment the hook says stop ā so worst-case latency
between a cancel flipping and the caller noticing is one chunk, not one
backoff.
It returns a Bool, it does not throw: only the caller knows which
exception type its layer promises (X::LLM::Chat::Retry::Cancelled,
a subclass, or something else entirely), so it reports and lets the
caller decide.
throw-cancelled unless sleep-with-cancel($wait, cancelled => &cancelled-now);
A &cancelled callback that throws propagates: the callback is the
caller's own code, and swallowing an exception from it would silently
turn a broken cancel hook into "never cancels".
Record shapes
attempt-record builds the entries of X::LLM::Chat::Retry::Exhausted's
@.attempts:
| Key | Type | Presence |
|---|---|---|
| backend-index | Int | always |
| model | Str | always |
| error | Str | always |
| raw-text | Str | ONLY when :raw-text was passed (key absent otherwise) |
telemetry-payload builds the per-round-trip hook payload:
| Key | Presence |
|---|---|
| attempt, backend-index, model-name, latency-ms | always |
| success, stage | always |
| error, error-class, error-status | always (value may be undefined) |
| prompt-tokens, completion-tokens, total-tokens | when :$response exposes them, defined |
| model-used, finish-reason | when :$response exposes them, defined |
| cost, generation-id, provider-name, is-byok | when :$response exposes them, defined |
Note the asymmetry, which is the pre-existing contract: the error*
keys are always present (undefined on success) so a sink can index them
unconditionally, while the response-derived keys are presence-gated
so a sink can distinguish "the provider reported 0 completion tokens"
from "the provider reported nothing". An empty-string error is
normalised to an undefined Str for the same reason.
Response probing is duck-typed .? throughout, so a backend whose
Response lacks cost / generation-id / provider-name /
is-byok (everything that is not OpenRouter) simply omits them, and a
test double needs only the accessors it wants to exercise.
CONSUMERS
LLM::Data::Inference::Task ā the original home of this code; its
classify-errormethod now delegates here, and its exception types subclass LLM::Chat::Retry::Exceptions'.LLM::Agent::Loopā per-round-trip retry/fallback inside the agent loop, including mid-stream failures.Any caller that wants Task's policy without Task's dependencies.
SEE ALSO
LLM::Chat::Retry::Exceptions, LLM::Chat::Backend::Response
(the error-class / error-status pair this classifies).