Exceptions

NAME

LLM::Data::Inference::Exceptions - typed failures thrown by the inference chain

DESCRIPTION

X::LLM::Data::Inference::Exhausted is thrown by LLM::Data::Inference::Task when every backend in the chain has been tried without producing a usable result. Its message is the same human-readable summary the module has always died with, so string-based callers keep working — but the exception additionally carries the per-attempt failure records, including the raw model output for parser failures, so callers can show the user exactly what the model produced instead of only "parser: unexpected ']'".

X::LLM::Data::Inference::Truncated is a subclass of Exhausted, thrown in its place when at least one attempt in the exhausted chain died because the model ran out of completion budget (the provider reported finish_reason 'length') rather than because it produced something unusable. Being a subclass is the whole point: every existing Exhausted handler — DLQ routing, CLI error reporting, :on-exhausted hooks — keeps catching it unchanged, while callers that care can add a narrower arm to say "this one is fixable: raise the budget" instead of "the model can't do this". It carries no extra attributes; the lever is named in the summary, and the truncated attempt records carry their partial output as raw-text. Only a Task with truncation-policy 'fail' can throw it — see LLM::Data::Inference::Task.

X::LLM::Data::Inference::TimedOut is likewise a subclass of Exhausted, thrown in its place when at least one attempt in the exhausted chain died on a response deadline — the Task's own $.timeout firing mid-poll, or a backend reporting error class 'timeout'. Same rationale as Truncated: exhaustion by deadline IS a chain failure and must keep flowing to every existing Exhausted handler untouched, while callers who want to distinguish "fixable by raising the deadline" from "the model can't do this" can match it before the Exhausted arm. The lever is named in the summary, which quotes the deadline that was hit.

When a chain saw both a truncation and a timeout, Truncated wins: it is the more specific diagnosis and, unlike a timeout, it is deterministic — see the advice contract below.

X::LLM::Data::Inference::Cancelled is thrown by the same Task when its is-cancelled hook reports that the caller no longer wants the result — before a round-trip, between retries, or mid-call (the in-flight response is aborted first). It carries the same per-attempt records as Exhausted so anything already produced stays inspectable. Catch it separately from Exhausted: cancellation is the caller's own intent, not a failure of the chain.

EXAMPLES


use LLM::Data::Inference::Task;
use LLM::Data::Inference::Exceptions;

my $task = LLM::Data::Inference::Task.new(
    :@backends, :$user-prompt, :&parser, parse-retries => 2,
    is-cancelled => -> { $job.cancelled },
);

my $result = do {
    CATCH {
        when X::LLM::Data::Inference::Cancelled {
            # The caller asked for this — clean up quietly.
            %();
        }
        # Truncated and TimedOut ARE Exhausteds, so they must come
        # FIRST — a `when` chain takes the first matching arm, and the
        # Exhausted arm below would otherwise swallow them.
        when X::LLM::Data::Inference::Truncated {
            note "every attempt hit the token budget:\n{.message}";
            %();
        }
        when X::LLM::Data::Inference::TimedOut {
            note "the chain ran out of time:\n{.message}";
            %();
        }
        when X::LLM::Data::Inference::Exhausted {
            for .attempts.grep(*.<raw-text>.defined) -> %a {
                note "model %a<model> said:\n%a<raw-text>";
            }
            %();
        }
    }
    $task.execute;
};

Callers that do not care about the distinction need no change at all: a bare when X::LLM::Data::Inference::Exhausted still catches a Truncated or a TimedOut, and :on-exhausted fires identically for all three.

RELATIONSHIP TO LLM::Chat::Retry::Exceptions

Since 0.9.0 these four types are thin subclasses of the shared retry exceptions in LLM::Chat::Retry::Exceptions, which is where the retry/fallback policy this distribution pioneered now lives so a second executor can share it:


X::LLM::Chat::Retry::Exhausted
ā”œā”€ā”€ X::LLM::Data::Inference::Exhausted
│   ā”œā”€ā”€ X::LLM::Data::Inference::Truncated  (also X::LLM::Chat::Retry::Truncated)
│   └── X::LLM::Data::Inference::TimedOut   (also X::LLM::Chat::Retry::TimedOut)
ā”œā”€ā”€ X::LLM::Chat::Retry::Truncated
└── X::LLM::Chat::Retry::TimedOut

X::LLM::Chat::Retry::Cancelled
└── X::LLM::Data::Inference::Cancelled

The consequence is that both hierarchies match, so nothing has to choose:


my $ex = ... ;  # thrown by a Task
$ex ~~ X::LLM::Data::Inference::Exhausted;   # True — every existing handler
$ex ~~ X::LLM::Chat::Retry::Exhausted;       # True — a generic retry handler

Attributes (@.attempts, $.summary), .message strings, the attempt-record shape and the item-retryable advice are all unchanged; the diamond exists purely so a caller that handles chains from several executors can write one when arm against the shared types, while a caller that only knows this distribution never notices.

Two notes for anyone subclassing further: match the more specific type first in a when chain, and use the public accessor @.attempts rather than @!attempts when overriding a method — the attribute now lives in the LLM::Chat parent and is not visible to a subclass's private-attribute syntax.

THE item-retryable ADVICE CONTRACT

The contract's canonical documentation moved to LLM::Chat::Retry::Exceptions along with the types themselves. In brief: an exception may advise an orchestration layer that re-runs failed work (LLM::Data::Pipeline's item engine, a job queue, a supervisor loop) that a blind re-roll cannot possibly help, by exposing


method item-retryable(--> Bool:D) { False }

which the layer reads as < $ex.?item-retryable // True > — absent means retryable, so every exception that has never heard of the contract keeps its full attempt budget, and neither side imports the other's types.

In this distribution only Truncated declares it, as False: a completion cut off by max_tokens will be cut off at exactly the same place on the next blind re-roll, because max_tokens is per-backend Settings state that a re-run cannot change. TimedOut deliberately does not — a deadline miss is genuinely transient (a slow provider, a queued request, an upstream hiccup), so it keeps the full item budget. Plain Exhausted abstains for the same reason, and nothing in this distribution reads the method: it is advice, not control flow.

LLM::Data::Inference v0.10.0

Structured LLM task layer with retry, JSON parsing, and query-based routing

Authors

  • Matt Doughty

License

Artistic-2.0

Dependencies

LLM::Chat:ver<0.10.0+>:auth<zef:apogee>Roaring::Tags:ver<0.2.3+>:auth<zef:apogee>CRoaring:ver<0.2.3+>:auth<zef:apogee>JSON::Fast:ver<0.19>:auth<cpan:TIMOTIMO>

Test Dependencies

Provides

  • LLM::Data::Inference
  • LLM::Data::Inference::Exceptions
  • LLM::Data::Inference::JSONTask
  • LLM::Data::Inference::PromptBuilder
  • LLM::Data::Inference::Router
  • LLM::Data::Inference::Task

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite — the markup and publishing tools behind this site.