Task
NAME
LLM::Data::Inference::Task - Chat-completion with model-chain fallback
SYNOPSIS
use LLM::Data::Inference::Task;
use LLM::Chat::Backend; # or a concrete Backend subclass
# Single backend (legacy shape ā unchanged)
my $task = LLM::Data::Inference::Task.new(
backend => $backend,
user-prompt => 'Tell me a joke',
);
my $text = $task.execute;
# Model-chain fallback ā try backends in order, falling through on
# model-specific failures (timeout / 4xx-non-auth / empty body /
# parse failure), retry the current backend once on transient errors
# (connection drop / 5xx), abort immediately on config/account errors
# (400 / 401 / 402 / 403 / 404).
my $task2 = LLM::Data::Inference::Task.new(
backends => [$primary, $secondary, $fallback],
user-prompt => 'Write a scene',
);
my $text2 = $task2.execute;
DESCRIPTION
One-shot LLM invocation with a robust retry + fallback policy.
Callers supply either a single :$backend (legacy, backward-compat)
or an ordered :@backends chain; the Task iterates the chain with
a three-bucket error classifier:
Error buckets
abort ā HTTP 400/401/402/403/404. These are config, account, or access errors; iterating the chain is wasted effort because the same error repeats. Re-raises immediately with context.
retry-same ā connection drop, 5xx, or an unclassifiable error. Likely transient (a specific OpenRouter upstream provider failed; a retry often routes to a different one). The current backend gets one additional attempt, then the Task advances to the next backend if it still fails.
advance ā timeout, 429 rate-limit, empty-body response, finish-reason quit (length / content_filter), or parser validation failure. These are model-specific: the current model ran into a pathology (reasoning loop, sanitisation, malformed JSON). Skip to the next backend.
The "finish-reason quit" case above is streaming-only. A streamed
'length' / 'content_filter' makes the backend quit the supply
with error class 'response', which lands in this bucket. The
blocking chat-completion path the Task actually uses does not
fail on either: a completion cut off by max_tokens comes back as an
ordinary HTTP 200 success carrying a partial body, with the reason
only visible as $response.finish-reason. See /Truncated output
for how that is intercepted.
Per-backend attempt budget
Each backend gets up to $.max-retries HTTP round-trips (default 3):
one initial attempt plus max-retries - 1 retries reserved for
retry-same-class errors. An advance-class error on any attempt
short-circuits the budget and moves straight to the next backend.
The chain is exhausted when every backend has been tried.
Retries use exponential backoff with jitter (2^(retry-1) + [0,0.5)
seconds, capped at 30 s) to avoid thundering-herd on concurrent
workers.
Parser failures
When :&parser is set, the parser is invoked on the raw text. A thrown
exception inside &parser is treated as an advance-class failure
(different model may produce parseable output). By default the Task
does NOT retry the same backend on parser failure ā the previous
per-backend retry loop produced identical malformed output enough times
in practice (batch pipelines, deterministic-ish sampling) that burning
budget on same-model parse retries was net-negative.
Interactive callers on stochastic samplers can opt into same-backend
re-rolls with :parse-retries(N): each parser failure consumes one
re-roll (a fresh completion from the same backend, no backoff) before
the advance rule kicks in. The raw model output of every failed parse
is preserved on the thrown X::LLM::Data::Inference::Exhausted
(see LLM::Data::Inference::Exceptions), so callers can surface
what the model actually said instead of only the parse error.
Retry feedback
A parse re-roll is blind by default: the same prompt goes back to the same model, and nothing tells it what was wrong with the answer it just gave. That is fine for malformed JSON on a stochastic sampler ā a fresh roll usually closes its brackets ā but useless for a semantic rejection. A validator that says "quote 2 does not appear verbatim in the passage" will keep saying it, three times, and the Task will dead-letter an item the model could have fixed on the second try.
:retry-feedback (opt-in, default False) makes the re-roll
informed: after a parser or validator failure, the next attempt against
the same backend carries one extra user turn naming the rejection.
my $task = LLM::Data::Inference::JSONTask.new(
:backends($primary, $fallback),
:$user-prompt,
:validator("es-appear-verbatim),
:parse-retries(2),
:retry-feedback, # re-rolls now say WHY the last one lost
);
The appended turn reads:
Your previous response was rejected: <the parser's message>. Respond
again with the same JSON contract, corrected.
Three properties are worth being explicit about, because each one is a decision rather than an accident:
Replace, never accumulate. Each attempt derives its message list from the pristine
@messages, so re-roll 2 carries re-roll 2's rejection and not re-roll 1's. The alternative ā appending to a growing transcript ā would spend an ever-larger prompt on progressively staler complaints, and models reliably fixate on the first item in such a list.Parse failures only. The feedback slot is filled in the parser-failure branch and nowhere else. A truncation is a budget fact, not a mistake the model can correct by trying harder (see /Truncated output); a network retry-same never produced output to critique; and the slot lives per backend, so a backend advance starts the next model on the pristine prompt ā its predecessor's mistakes are not evidence about anything it did. (The slot is sticky until the backend changes: if a parse failure is followed by a transient network error, the retry after that error still carries the pending feedback, because the rejection it describes is still the last thing the model actually said.)
No echo of the rejected output. The feedback names the error, not the text that produced it. The model has its own last turn in context on every provider worth using, so echoing it back doubles the prompt to restate what the model already knows; and validator diagnostics in this ecosystem are written to identify the failing item precisely ("beat 3, quote 2: ..."), which is the part the model cannot infer. If a caller wants the raw text in the loop, it is on every attempt record already ā write a parser that folds it into its own message.
The canned wording names a JSON contract, which makes it correct for
JSONTask and for hand-rolled JSON parsers, and wrong for prose
Tasks. That is deliberate: prose Tasks simply do not opt in. (Callers
who want feedback with different wording can raise it from inside the
parser instead, since the parser's message is what gets quoted.)
Telemetry is unchanged ā &on-call-complete carries no message-count
field, so the extra turn shows up only as a larger prompt-tokens
figure on the retry row. Budget roughly a hundred extra prompt tokens
per informed re-roll.
Truncated output
A blocking completion that runs out of completion budget is not an
error at the transport level: HTTP 200, a well-formed body, a partial
answer, and finish_reason 'length'. Left alone, the Task would
hand that half-finished text to &.parser, watch it fail, and ā with
:parse-retries set ā re-roll the identical doomed request until
the budget was spent, because a request that overran once will overrun
again at the same max_tokens.
$.truncation-policy decides what happens instead:
'fail'ā a'length'finish reason is recorded as a failure and the Task advances to the next backend immediately, without consuming a parse re-roll.max_tokensis per-backendSettingsstate, so the next backend in the chain can legitimately have a bigger budget and succeed where this one could not ā but a same-backend re-roll cannot. If the whole chain truncates, the Task throwsX::LLM::Data::Inference::Truncated(a subclass ofExhausted, so existing handlers are unaffected) and the summary names the levers.'accept'ā historical behaviour. The partial text is the result: returned as-is when there is no parser, handed to the parser when there is.''(the default) ā derive the policy:'fail'when&.parseris set,'accept'when it is not. A truncated payload headed for a parser can never parse, so failing fast only removes a guaranteed-waste path; truncated prose, on the other hand, is frequently still the product (continuations, style passes), so parser-less Tasks keep their historical behaviour bit-for-bit. Set the attribute explicitly to override in either direction.
LLM::Data::Inference::JSONTask always installs a parser, so it
defaults to 'fail'.
The empty-body rule is checked after truncation, deliberately: a
reasoning model that spends its entire budget inside
< <think>ā¦</think> > returns 200 with an empty body and
finish_reason 'length'. That is budget exhaustion, and reporting
it as "empty response body" sends the reader hunting the wrong bug.
# Prose task that wants the partial text anyway (the default already
# does this, since no parser is set ā spelled out for clarity).
my $prose = LLM::Data::Inference::Task.new(
:$backend, :$user-prompt, truncation-policy => 'accept',
);
# JSON task on a chain whose second backend has a bigger budget.
my $json = LLM::Data::Inference::JSONTask.new(
backends => [$small-budget, $big-budget], :$user-prompt,
parse-retries => 2, # never burned on a truncation
);
Typed exhaustion and levers
When the chain runs out, the Task throws one of three types ā all of
them X::LLM::Data::Inference::Exhausted or a subclass of it, so a
handler that never heard of the subclasses is unaffected:
Truncatedā at least one attempt was cut off bymax_tokens(only reachable undertruncation-policy'fail'; see above).TimedOutā at least one attempt died on a response deadline:$.timeoutfiring while the Task polled the in-flight response, or a backend reporting error class'timeout'itself.Exhaustedā everything else.
Precedence: Truncated wins over TimedOut when a chain saw
both. It is the more specific diagnosis, and it is the deterministic
one ā it declares item-retryable False (see
LLM::Data::Inference::Exceptions), which tells an orchestrator to
stop re-running the item rather than buying back retries a truncation
guarantees will fail. TimedOut makes no such claim: a deadline miss
is transient, so it keeps its full item budget.
The summary (and therefore .message) gains one lever line per
pathology observed, appended after the per-attempt error list. A chain
that both truncated and timed out gets both lines ā they name
different fixes:
LLM::Data::Inference::Task: all 2 backend(s) exhausted.
[backend 0 small] truncated: output was cut off by the max_tokens ā¦
[backend 1 slow] timeout: response timed out after 300s
At least one attempt was cut off by the completion budget: raise ā¦
At least one attempt hit the 300-second response deadline before ā¦
Summaries of chains that hit neither are byte-identical to what they have always been.
Telemetry
Every HTTP round-trip fires &on-call-complete (if provided), with
the attempt number, model name, backend index (position in the chain),
latency, success flag, usage data, and error details. Parser failures
do not fire separately ā they're classified from the preceding
telemetry row.
&on-exhausted (if provided) fires exactly once, on the execute
thread, immediately before the chain throws
X::LLM::Data::Inference::Exhausted or either of its subclasses
(see /Typed exhaustion and levers) ā never on success, and never
on the Cancelled path. The payload is stage => 'exhausted',
attempts (the same @attempt-records array the thrown exception
carries ā including raw-text for parse failures), summary (the
same aggregate string as the exception's .message), and
backends (the chain length). Like &on-call-complete, a hook
that throws is shielded ā it can never suppress or alter the
Exhausted throw that always follows.
my $task = LLM::Data::Inference::Task.new(
backends => [$primary, $fallback],
user-prompt => 'q',
on-exhausted => -> %info {
$dlq.append: %(
attempts => %info<attempts>,
summary => %info<summary>,
);
},
);
Cooperative cancellation
:&is-cancelled is polled at every decision point: before each
round-trip, inside the in-flight completion poll (the pending response
is aborted via Backend.cancel, exactly like the timeout path),
between backoff sleeps, and before parse re-rolls / backend advances.
Aborting through Backend.cancel rather than Response.cancel is
what makes the cancel reach the server: against KoboldCpp it fires
POST /api/extra/abort alongside the local stream close, so a
cancelled Task stops the generation instead of only looking away from
it. Backends with no upstream abort endpoint degrade to a local close.
When it returns True the Task throws
X::LLM::Data::Inference::Cancelled (see
LLM::Data::Inference::Exceptions) carrying whatever attempt records
already exist. Worst-case latency between the hook flipping and the
throw is one poll interval (~10 ms in-call, ~250 ms mid-backoff) ā
a cancelled Task never burns a full retry chain.
my $task = LLM::Data::Inference::Task.new(
backend => $backend,
user-prompt => 'Reconcile this exchange as JSON',
parser => &json-parser,
parse-retries => 2,
is-cancelled => -> { $job.cancelled },
);
BACKWARD COMPATIBILITY
Callers passing :$backend only see no behaviour change beyond the
new retry policy. The $.backend accessor remains available and
always points at the first backend in the chain; legacy consumers that
read it (e.g. to label log messages with the model) keep working.