Exceptions

NAME

LLM::Chat::Retry::Exceptions - shared typed failures for retry/fallback loops

SYNOPSIS


use LLM::Chat::Retry;
use LLM::Chat::Retry::Exceptions;

my @attempts;
my $bucket = classify-error(:error-class<http>, :error-status(500));
@attempts.push: attempt-record(
    :backend-index(0), :model<gpt-4o-mini>, :error('http 500: upstream'),
);

# ... chain exhausted ...
X::LLM::Chat::Retry::Exhausted.new(
    attempts => @attempts,
    summary  => "all 1 backend(s) exhausted.\n  {@attempts[0]<error>}",
).throw;

Catching, with the subclasses first — a when chain takes the first matching arm, and Truncated / TimedOut are Exhausteds:


CATCH {
    when X::LLM::Chat::Retry::Cancelled {
        # The caller asked for this — clean up quietly.
    }
    when X::LLM::Chat::Retry::Truncated {
        note "every attempt hit the token budget:\n{.message}";
    }
    when X::LLM::Chat::Retry::TimedOut {
        note "the chain ran out of time:\n{.message}";
    }
    when X::LLM::Chat::Retry::Exhausted {
        for .attempts.grep(*.<raw-text>.defined) -> %a {
            note "model %a<model> said:\n%a<raw-text>";
        }
    }
}

DESCRIPTION

The four exception types every retry/fallback loop in this ecosystem throws, factored out of LLM::Data::Inference::Task so that LLM::Chat-level consumers (LLM::Agent::Loop, bespoke chains, an app's own executor) can throw and catch the same types instead of each inventing a private hierarchy that nothing downstream can match on.

They are deliberately dependency-free: this module imports nothing but core Raku, so an exception can travel through any layer without dragging a backend, a Response, or an HTTP client behind it.

The four types

  • X::LLM::Chat::Retry::Exhausted — every backend in the chain was tried without producing a usable result. Carries @.attempts (one record per recorded failure, in order) and a required $.summary, which is also its .message. The summary is the human-readable aggregate; the records are the machine-readable detail, including the raw model output for parser failures, so a caller can show the user what the model actually produced instead of only "parser: unexpected ']'".

  • X::LLM::Chat::Retry::Truncated — a subclass of Exhausted, thrown in its place when at least one attempt in the exhausted chain was cut off by the completion budget (the provider reported finish_reason 'length') rather than producing something unusable. Being a subclass is the whole point: existing Exhausted handlers — dead-letter routing, CLI error reporting, exhaustion hooks — keep catching it unchanged, while callers who care can add a narrower arm saying "this one is fixable: raise the budget". It adds no attributes; the lever belongs in the summary, and truncated attempt records carry their partial output as raw-text.

  • X::LLM::Chat::Retry::TimedOut — likewise a subclass of Exhausted, thrown in its place when at least one attempt died on a response deadline (a caller-side timeout firing mid-poll, or a backend reporting error class 'timeout'). Same rationale: exhaustion by deadline IS a chain failure and must keep flowing to every existing Exhausted handler, while callers who want to distinguish "fixable by raising the deadline" from "the model can't do this" can match it before the Exhausted arm.

  • X::LLM::Chat::Retry::Cancelled — the caller withdrew: an is-cancelled-style hook reported that nobody wants the result any more, before a round-trip, between retries, or mid-call. Carries the same @.attempts so anything already produced stays inspectable.

Why Cancelled is NOT an Exhausted

Everything else in this file is a failure of the chain; a cancel is the caller's own intent. A handler that routes exhaustion to a dead-letter queue, counts model failures, or opens an incident must not have a user pressing Ctrl-C silently folded into those numbers — so Cancelled descends straight from Exception and has to be caught on purpose. That also means a bare when X::LLM::Chat::Retry::Exhausted arm can never swallow it by accident.

Precedence: Truncated beats TimedOut

A chain that saw both a truncation and a deadline miss throws Truncated. It is the more specific diagnosis, and — see the advice contract below — it is the deterministic one: it declares item-retryable False, which tells an orchestrator to stop re-running the item rather than buying back retries that the truncation guarantees will fail. Downgrading to TimedOut would do exactly that. Throwers that observed both pathologies should name both levers in the summary; they describe different fixes and suppressing one hides half the diagnosis.

Attempt record shape

@.attempts holds one Hash per recorded failure, in the order they happened. LLM::Chat::Retry's attempt-record builds them:


%(
    backend-index => 0,          # 0-based position in the chain
    model         => 'gpt-4o',   # model name, or 'unknown'
    error         => '[backend 0 gpt-4o] http 500: upstream on fire',
    raw-text      => '{"a": 1',  # OPTIONAL — key ABSENT when there is
                                 # no body worth keeping (network
                                 # failures have none); present for
                                 # parser failures and truncations
)

The raw-text key is absent, not undefined, when it does not apply, so .grep(*.<raw-text>.defined) and %a<raw-text>:exists both work and neither reports a phantom empty body.

THE item-retryable ADVICE CONTRACT

(This is the canonical home of the contract; it moved here from LLM::Data::Inference::Exceptions when the retry policy was factored out, and the types there are now subclasses of these.)

An orchestration layer that re-runs failed work (LLM::Data::Pipeline's item engine, a job queue, a supervisor loop) has no way of knowing which of these deaths a blind re-roll could plausibly fix. Some of them are deterministic: re-issuing the identical request produces the identical death, so the retries are pure latency and spend on the way to the dead-letter queue.

An exception may therefore advise the layer above it by exposing:


method item-retryable(--> Bool:D) { False }

The contract is duck-typed in both directions, deliberately:

  • Absent means retryable. A consumer probes with .? and defaults to True, so every exception that has never heard of the contract — including every exception in this distribution before 0.8.0 — keeps its full attempt budget.

  • No type dependency either way. The consumer never imports a retry type (it calls a method by name); this distribution never imports the orchestrator. Any exception from any library can opt in.

  • It is advice, not control flow. The layer above decides what to do with it. Nothing in this distribution reads the method.

Only Truncated declares it, as False: a completion cut off by max_tokens will be cut off at exactly the same place on the next blind re-roll, because max_tokens is per-backend Settings state that a re-run cannot change. TimedOut deliberately does not declare it — a deadline miss is genuinely transient (a slow provider, a queued request, an upstream hiccup), so it keeps the full item budget. Plain Exhausted abstains for the same reason: the base exhaustion carries no evidence that a re-run is pointless.

A consumer reads it like this:


my Bool $retry = $exception.?item-retryable // True;

SUBCLASSING

Distributions that want their own names in the type tree — for backwards compatibility, or just for a message that says which layer died — subclass these and inherit the whole contract:


class X::My::Chain::Exhausted is X::LLM::Chat::Retry::Exhausted { }
class X::My::Chain::Truncated
    is X::My::Chain::Exhausted
    is X::LLM::Chat::Retry::Truncated { }   # keeps item-retryable

The diamond is intentional and resolves cleanly: X::My::Chain::Truncated smartmatches both X::My::Chain::Exhausted and X::LLM::Chat::Retry::Truncated, so old handlers and new ones both fire. One gotcha when overriding methods in such a subclass: use the public accessor @.attempts, never the private @!attempts — the attribute lives in the parent class and is not visible to the child.

Why these are plain global classes, not is exported ones

There is no unit module here and no is export trait, so use-ing this file simply makes X::LLM::Chat::Retry::* resolvable — the same way core's X::AdHoc is. That is deliberate, and it is not the style LLM::Data::Inference::Exceptions uses.

is export on a nested-name class exports it under its leaf name as well: a module doing class X::Foo::Exhausted is export puts a bare Exhausted into every importer's lexical scope. Two such modules therefore cannot be imported into the same scope at all — "Cannot import the following symbols … because they already exist: 'Exhausted', 'Truncated', …" — which would make it impossible to write the very handler this module exists to enable, one that matches a subclass hierarchy against the shared one. Declaring the classes globally instead costs nothing (they are reachable by their real names, which are also their .^names) and keeps this module compatible with every other exception module in the ecosystem, present and future.

LLM::Chat v0.10.0

Simple framework for LLM inferencing

Authors

  • Matt Doughty

License

Artistic-2.0

Dependencies

Cro::HTTP::Client:ver<0.8.11+>:auth<zef:cro>Digest::SHA256::Native:ver<1.0.0+>:auth<zef:bduggan>JSON::Fast:ver<0.19+>:auth<cpan:TIMOTIMO>Template::Jinja2:ver<0.3.0+>:auth<zef:apogee>Tokenizers:ver<0.3.0+>:auth<zef:apogee>UUID::V4:ver<1.0.0+>:auth<zef:masukomi>

Provides

  • LLM::Chat
  • LLM::Chat::Backend
  • LLM::Chat::Backend::KoboldCpp
  • LLM::Chat::Backend::Mock
  • LLM::Chat::Backend::OpenAICommon
  • LLM::Chat::Backend::OpenRouter
  • LLM::Chat::Backend::Response
  • LLM::Chat::Backend::Response::OpenRouter
  • LLM::Chat::Backend::Response::OpenRouter::Augment
  • LLM::Chat::Backend::Response::OpenRouter::Stream
  • LLM::Chat::Backend::Response::Stream
  • LLM::Chat::Backend::Settings
  • LLM::Chat::Conversation
  • LLM::Chat::Conversation::Message
  • LLM::Chat::Debug
  • LLM::Chat::Retry
  • LLM::Chat::Retry::Exceptions
  • LLM::Chat::Template
  • LLM::Chat::Template::ChatML
  • LLM::Chat::Template::DeepSeekV4
  • LLM::Chat::Template::Gemma2
  • LLM::Chat::Template::Jinja2
  • LLM::Chat::Template::Llama3
  • LLM::Chat::Template::Llama4
  • LLM::Chat::Template::MistralV7
  • LLM::Chat::TokenCounter
  • LLM::Chat::ToolLoop

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite — the markup and publishing tools behind this site.