Exceptions
NAME
LLM::Data::Inference::Exceptions - typed failures thrown by the inference chain
DESCRIPTION
X::LLM::Data::Inference::Exhausted is thrown by
LLM::Data::Inference::Task when every backend in the chain has been
tried without producing a usable result. Its message is the same
human-readable summary the module has always died with, so string-based
callers keep working ā but the exception additionally carries the
per-attempt failure records, including the raw model output for
parser failures, so callers can show the user exactly what the model
produced instead of only "parser: unexpected ']'".
X::LLM::Data::Inference::Truncated is a subclass of
Exhausted, thrown in its place when at least one attempt in the
exhausted chain died because the model ran out of completion budget
(the provider reported finish_reason 'length') rather than
because it produced something unusable. Being a subclass is the whole
point: every existing Exhausted handler ā DLQ routing, CLI error
reporting, :on-exhausted hooks ā keeps catching it unchanged, while
callers that care can add a narrower arm to say "this one is fixable:
raise the budget" instead of "the model can't do this". It carries no
extra attributes; the lever is named in the summary, and the
truncated attempt records carry their partial output as raw-text.
Only a Task with truncation-policy 'fail' can throw it ā see
LLM::Data::Inference::Task.
X::LLM::Data::Inference::TimedOut is likewise a subclass of
Exhausted, thrown in its place when at least one attempt in the
exhausted chain died on a response deadline ā the Task's own
$.timeout firing mid-poll, or a backend reporting error class
'timeout'. Same rationale as Truncated: exhaustion by deadline
IS a chain failure and must keep flowing to every existing
Exhausted handler untouched, while callers who want to distinguish
"fixable by raising the deadline" from "the model can't do this" can
match it before the Exhausted arm. The lever is named in the
summary, which quotes the deadline that was hit.
When a chain saw both a truncation and a timeout, Truncated
wins: it is the more specific diagnosis and, unlike a timeout, it is
deterministic ā see the advice contract below.
X::LLM::Data::Inference::Cancelled is thrown by the same Task when
its is-cancelled hook reports that the caller no longer wants the
result ā before a round-trip, between retries, or mid-call (the
in-flight response is aborted first). It carries the same per-attempt
records as Exhausted so anything already produced stays
inspectable. Catch it separately from Exhausted: cancellation is
the caller's own intent, not a failure of the chain.
EXAMPLES
use LLM::Data::Inference::Task;
use LLM::Data::Inference::Exceptions;
my $task = LLM::Data::Inference::Task.new(
:@backends, :$user-prompt, :&parser, parse-retries => 2,
is-cancelled => -> { $job.cancelled },
);
my $result = do {
CATCH {
when X::LLM::Data::Inference::Cancelled {
# The caller asked for this ā clean up quietly.
%();
}
# Truncated and TimedOut ARE Exhausteds, so they must come
# FIRST ā a `when` chain takes the first matching arm, and the
# Exhausted arm below would otherwise swallow them.
when X::LLM::Data::Inference::Truncated {
note "every attempt hit the token budget:\n{.message}";
%();
}
when X::LLM::Data::Inference::TimedOut {
note "the chain ran out of time:\n{.message}";
%();
}
when X::LLM::Data::Inference::Exhausted {
for .attempts.grep(*.<raw-text>.defined) -> %a {
note "model %a<model> said:\n%a<raw-text>";
}
%();
}
}
$task.execute;
};
Callers that do not care about the distinction need no change at all:
a bare when X::LLM::Data::Inference::Exhausted still catches a
Truncated or a TimedOut, and :on-exhausted fires identically
for all three.
RELATIONSHIP TO LLM::Chat::Retry::Exceptions
Since 0.9.0 these four types are thin subclasses of the shared retry exceptions in LLM::Chat::Retry::Exceptions, which is where the retry/fallback policy this distribution pioneered now lives so a second executor can share it:
X::LLM::Chat::Retry::Exhausted
āāā X::LLM::Data::Inference::Exhausted
ā āāā X::LLM::Data::Inference::Truncated (also X::LLM::Chat::Retry::Truncated)
ā āāā X::LLM::Data::Inference::TimedOut (also X::LLM::Chat::Retry::TimedOut)
āāā X::LLM::Chat::Retry::Truncated
āāā X::LLM::Chat::Retry::TimedOut
X::LLM::Chat::Retry::Cancelled
āāā X::LLM::Data::Inference::Cancelled
The consequence is that both hierarchies match, so nothing has to choose:
my $ex = ... ; # thrown by a Task
$ex ~~ X::LLM::Data::Inference::Exhausted; # True ā every existing handler
$ex ~~ X::LLM::Chat::Retry::Exhausted; # True ā a generic retry handler
Attributes (@.attempts, $.summary), .message strings, the
attempt-record shape and the item-retryable advice are all
unchanged; the diamond exists purely so a caller that handles chains
from several executors can write one when arm against the shared
types, while a caller that only knows this distribution never notices.
Two notes for anyone subclassing further: match the more specific type
first in a when chain, and use the public accessor @.attempts
rather than @!attempts when overriding a method ā the attribute now
lives in the LLM::Chat parent and is not visible to a subclass's
private-attribute syntax.
THE item-retryable ADVICE CONTRACT
The contract's canonical documentation moved to
LLM::Chat::Retry::Exceptions along with the types themselves. In
brief: an exception may advise an orchestration layer that re-runs
failed work (LLM::Data::Pipeline's item engine, a job queue, a
supervisor loop) that a blind re-roll cannot possibly help, by
exposing
method item-retryable(--> Bool:D) { False }
which the layer reads as < $ex.?item-retryable // True > ā absent
means retryable, so every exception that has never heard of the
contract keeps its full attempt budget, and neither side imports the
other's types.
In this distribution only Truncated declares it, as False: a
completion cut off by max_tokens will be cut off at exactly the
same place on the next blind re-roll, because max_tokens is
per-backend Settings state that a re-run cannot change. TimedOut
deliberately does not ā a deadline miss is genuinely transient (a
slow provider, a queued request, an upstream hiccup), so it keeps the
full item budget. Plain Exhausted abstains for the same reason, and
nothing in this distribution reads the method: it is advice, not
control flow.