Task

NAME

LLM::Data::Inference::Task - Chat-completion with model-chain fallback

SYNOPSIS


use LLM::Data::Inference::Task;
use LLM::Chat::Backend;  # or a concrete Backend subclass

# Single backend (legacy shape — unchanged)
my $task = LLM::Data::Inference::Task.new(
    backend     => $backend,
    user-prompt => 'Tell me a joke',
);
my $text = $task.execute;

# Model-chain fallback — try backends in order, falling through on
# model-specific failures (timeout / 4xx-non-auth / empty body /
# parse failure), retry the current backend once on transient errors
# (connection drop / 5xx), abort immediately on config/account errors
# (400 / 401 / 402 / 403 / 404).
my $task2 = LLM::Data::Inference::Task.new(
    backends => [$primary, $secondary, $fallback],
    user-prompt => 'Write a scene',
);
my $text2 = $task2.execute;

DESCRIPTION

One-shot LLM invocation with a robust retry + fallback policy. Callers supply either a single :$backend (legacy, backward-compat) or an ordered :@backends chain; the Task iterates the chain with a three-bucket error classifier:

Error buckets

  • abort — HTTP 400/401/402/403/404. These are config, account, or access errors; iterating the chain is wasted effort because the same error repeats. Re-raises immediately with context.

  • retry-same — connection drop, 5xx, or an unclassifiable error. Likely transient (a specific OpenRouter upstream provider failed; a retry often routes to a different one). The current backend gets one additional attempt, then the Task advances to the next backend if it still fails.

  • advance — timeout, 429 rate-limit, empty-body response, finish-reason quit (length / content_filter), or parser validation failure. These are model-specific: the current model ran into a pathology (reasoning loop, sanitisation, malformed JSON). Skip to the next backend.

The "finish-reason quit" case above is streaming-only. A streamed 'length' / 'content_filter' makes the backend quit the supply with error class 'response', which lands in this bucket. The blocking chat-completion path the Task actually uses does not fail on either: a completion cut off by max_tokens comes back as an ordinary HTTP 200 success carrying a partial body, with the reason only visible as $response.finish-reason. See /Truncated output for how that is intercepted.

Per-backend attempt budget

Each backend gets up to $.max-retries HTTP round-trips (default 3): one initial attempt plus max-retries - 1 retries reserved for retry-same-class errors. An advance-class error on any attempt short-circuits the budget and moves straight to the next backend. The chain is exhausted when every backend has been tried.

Retries use exponential backoff with jitter (2^(retry-1) + [0,0.5) seconds, capped at 30 s) to avoid thundering-herd on concurrent workers.

Parser failures

When :&parser is set, the parser is invoked on the raw text. A thrown exception inside &parser is treated as an advance-class failure (different model may produce parseable output). By default the Task does NOT retry the same backend on parser failure — the previous per-backend retry loop produced identical malformed output enough times in practice (batch pipelines, deterministic-ish sampling) that burning budget on same-model parse retries was net-negative.

Interactive callers on stochastic samplers can opt into same-backend re-rolls with :parse-retries(N): each parser failure consumes one re-roll (a fresh completion from the same backend, no backoff) before the advance rule kicks in. The raw model output of every failed parse is preserved on the thrown X::LLM::Data::Inference::Exhausted (see LLM::Data::Inference::Exceptions), so callers can surface what the model actually said instead of only the parse error.

Retry feedback

A parse re-roll is blind by default: the same prompt goes back to the same model, and nothing tells it what was wrong with the answer it just gave. That is fine for malformed JSON on a stochastic sampler — a fresh roll usually closes its brackets — but useless for a semantic rejection. A validator that says "quote 2 does not appear verbatim in the passage" will keep saying it, three times, and the Task will dead-letter an item the model could have fixed on the second try.

:retry-feedback (opt-in, default False) makes the re-roll informed: after a parser or validator failure, the next attempt against the same backend carries one extra user turn naming the rejection.


my $task = LLM::Data::Inference::JSONTask.new(
    :backends($primary, $fallback),
    :$user-prompt,
    :validator(&quotes-appear-verbatim),
    :parse-retries(2),
    :retry-feedback,       # re-rolls now say WHY the last one lost
);

The appended turn reads:


Your previous response was rejected: <the parser's message>. Respond
again with the same JSON contract, corrected.

Three properties are worth being explicit about, because each one is a decision rather than an accident:

  • Replace, never accumulate. Each attempt derives its message list from the pristine @messages, so re-roll 2 carries re-roll 2's rejection and not re-roll 1's. The alternative — appending to a growing transcript — would spend an ever-larger prompt on progressively staler complaints, and models reliably fixate on the first item in such a list.

  • Parse failures only. The feedback slot is filled in the parser-failure branch and nowhere else. A truncation is a budget fact, not a mistake the model can correct by trying harder (see /Truncated output); a network retry-same never produced output to critique; and the slot lives per backend, so a backend advance starts the next model on the pristine prompt — its predecessor's mistakes are not evidence about anything it did. (The slot is sticky until the backend changes: if a parse failure is followed by a transient network error, the retry after that error still carries the pending feedback, because the rejection it describes is still the last thing the model actually said.)

  • No echo of the rejected output. The feedback names the error, not the text that produced it. The model has its own last turn in context on every provider worth using, so echoing it back doubles the prompt to restate what the model already knows; and validator diagnostics in this ecosystem are written to identify the failing item precisely ("beat 3, quote 2: ..."), which is the part the model cannot infer. If a caller wants the raw text in the loop, it is on every attempt record already — write a parser that folds it into its own message.

The canned wording names a JSON contract, which makes it correct for JSONTask and for hand-rolled JSON parsers, and wrong for prose Tasks. That is deliberate: prose Tasks simply do not opt in. (Callers who want feedback with different wording can raise it from inside the parser instead, since the parser's message is what gets quoted.)

Telemetry is unchanged — &on-call-complete carries no message-count field, so the extra turn shows up only as a larger prompt-tokens figure on the retry row. Budget roughly a hundred extra prompt tokens per informed re-roll.

Truncated output

A blocking completion that runs out of completion budget is not an error at the transport level: HTTP 200, a well-formed body, a partial answer, and finish_reason 'length'. Left alone, the Task would hand that half-finished text to &.parser, watch it fail, and — with :parse-retries set — re-roll the identical doomed request until the budget was spent, because a request that overran once will overrun again at the same max_tokens.

$.truncation-policy decides what happens instead:

  • 'fail' — a 'length' finish reason is recorded as a failure and the Task advances to the next backend immediately, without consuming a parse re-roll. max_tokens is per-backend Settings state, so the next backend in the chain can legitimately have a bigger budget and succeed where this one could not — but a same-backend re-roll cannot. If the whole chain truncates, the Task throws X::LLM::Data::Inference::Truncated (a subclass of Exhausted, so existing handlers are unaffected) and the summary names the levers.

  • 'accept' — historical behaviour. The partial text is the result: returned as-is when there is no parser, handed to the parser when there is.

  • '' (the default) — derive the policy: 'fail' when &.parser is set, 'accept' when it is not. A truncated payload headed for a parser can never parse, so failing fast only removes a guaranteed-waste path; truncated prose, on the other hand, is frequently still the product (continuations, style passes), so parser-less Tasks keep their historical behaviour bit-for-bit. Set the attribute explicitly to override in either direction.

LLM::Data::Inference::JSONTask always installs a parser, so it defaults to 'fail'.

The empty-body rule is checked after truncation, deliberately: a reasoning model that spends its entire budget inside < <think>…</think> > returns 200 with an empty body and finish_reason 'length'. That is budget exhaustion, and reporting it as "empty response body" sends the reader hunting the wrong bug.


# Prose task that wants the partial text anyway (the default already
# does this, since no parser is set — spelled out for clarity).
my $prose = LLM::Data::Inference::Task.new(
    :$backend, :$user-prompt, truncation-policy => 'accept',
);

# JSON task on a chain whose second backend has a bigger budget.
my $json = LLM::Data::Inference::JSONTask.new(
    backends => [$small-budget, $big-budget], :$user-prompt,
    parse-retries => 2,   # never burned on a truncation
);

Typed exhaustion and levers

When the chain runs out, the Task throws one of three types — all of them X::LLM::Data::Inference::Exhausted or a subclass of it, so a handler that never heard of the subclasses is unaffected:

  • Truncated — at least one attempt was cut off by max_tokens (only reachable under truncation-policy 'fail'; see above).

  • TimedOut — at least one attempt died on a response deadline: $.timeout firing while the Task polled the in-flight response, or a backend reporting error class 'timeout' itself.

  • Exhausted — everything else.

Precedence: Truncated wins over TimedOut when a chain saw both. It is the more specific diagnosis, and it is the deterministic one — it declares item-retryable False (see LLM::Data::Inference::Exceptions), which tells an orchestrator to stop re-running the item rather than buying back retries a truncation guarantees will fail. TimedOut makes no such claim: a deadline miss is transient, so it keeps its full item budget.

The summary (and therefore .message) gains one lever line per pathology observed, appended after the per-attempt error list. A chain that both truncated and timed out gets both lines — they name different fixes:


LLM::Data::Inference::Task: all 2 backend(s) exhausted.
  [backend 0 small] truncated: output was cut off by the max_tokens …
  [backend 1 slow] timeout: response timed out after 300s
  At least one attempt was cut off by the completion budget: raise …
  At least one attempt hit the 300-second response deadline before …

Summaries of chains that hit neither are byte-identical to what they have always been.

Telemetry

Every HTTP round-trip fires &on-call-complete (if provided), with the attempt number, model name, backend index (position in the chain), latency, success flag, usage data, and error details. Parser failures do not fire separately — they're classified from the preceding telemetry row.

&on-exhausted (if provided) fires exactly once, on the execute thread, immediately before the chain throws X::LLM::Data::Inference::Exhausted or either of its subclasses (see /Typed exhaustion and levers) — never on success, and never on the Cancelled path. The payload is stage => 'exhausted', attempts (the same @attempt-records array the thrown exception carries — including raw-text for parse failures), summary (the same aggregate string as the exception's .message), and backends (the chain length). Like &on-call-complete, a hook that throws is shielded — it can never suppress or alter the Exhausted throw that always follows.


my $task = LLM::Data::Inference::Task.new(
    backends      => [$primary, $fallback],
    user-prompt   => 'q',
    on-exhausted  => -> %info {
        $dlq.append: %(
            attempts => %info<attempts>,
            summary  => %info<summary>,
        );
    },
);

Cooperative cancellation

:&is-cancelled is polled at every decision point: before each round-trip, inside the in-flight completion poll (the pending response is aborted via Backend.cancel, exactly like the timeout path), between backoff sleeps, and before parse re-rolls / backend advances.

Aborting through Backend.cancel rather than Response.cancel is what makes the cancel reach the server: against KoboldCpp it fires POST /api/extra/abort alongside the local stream close, so a cancelled Task stops the generation instead of only looking away from it. Backends with no upstream abort endpoint degrade to a local close. When it returns True the Task throws X::LLM::Data::Inference::Cancelled (see LLM::Data::Inference::Exceptions) carrying whatever attempt records already exist. Worst-case latency between the hook flipping and the throw is one poll interval (~10 ms in-call, ~250 ms mid-backoff) — a cancelled Task never burns a full retry chain.


my $task = LLM::Data::Inference::Task.new(
    backend      => $backend,
    user-prompt  => 'Reconcile this exchange as JSON',
    parser       => &json-parser,
    parse-retries => 2,
    is-cancelled => -> { $job.cancelled },
);

BACKWARD COMPATIBILITY

Callers passing :$backend only see no behaviour change beyond the new retry policy. The $.backend accessor remains available and always points at the first backend in the chain; legacy consumers that read it (e.g. to label log messages with the model) keep working.

LLM::Data::Inference v0.10.0

Structured LLM task layer with retry, JSON parsing, and query-based routing

Authors

  • Matt Doughty

License

Artistic-2.0

Dependencies

LLM::Chat:ver<0.10.0+>:auth<zef:apogee>Roaring::Tags:ver<0.2.3+>:auth<zef:apogee>CRoaring:ver<0.2.3+>:auth<zef:apogee>JSON::Fast:ver<0.19>:auth<cpan:TIMOTIMO>

Test Dependencies

Provides

  • LLM::Data::Inference
  • LLM::Data::Inference::Exceptions
  • LLM::Data::Inference::JSONTask
  • LLM::Data::Inference::PromptBuilder
  • LLM::Data::Inference::Router
  • LLM::Data::Inference::Task

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite — the markup and publishing tools behind this site.