Mock
NAME
LLM::Chat::Backend::Mock - Canned-response backend for tests
SYNOPSIS
use LLM::Chat::Backend::Mock;
use LLM::Chat::Backend::Settings;
# Single canned response
my $mock = LLM::Chat::Backend::Mock.new(
settings => LLM::Chat::Backend::Settings.new,
responses => ['hello from the mock backend'],
);
# A sequence ā each call pops the next one
my $mock2 = LLM::Chat::Backend::Mock.new(
settings => LLM::Chat::Backend::Settings.new,
responses => [
'first answer',
'second answer',
'third answer',
],
);
# Streaming: tokens arrive split by whitespace (default).
# Customise with :token-splitter if you want per-character streaming or
# a specific token boundary.
my $stream = $mock2.chat-completion-stream(@messages);
react {
whenever $stream.supply -> $token {
print $token;
}
whenever $stream.supply.done { say "done" }
}
DESCRIPTION
Non-network backend for use in tests. Returns pre-configured responses
in order ā one response per chat-completion/text-completion call.
When the queue is exhausted, subsequent calls either repeat the last
response (default) or fail (:fail-on-empty).
For streaming calls, the response is split into tokens and emitted
through the normal Supplier path. By default each whitespace-
separated word becomes a token; override :token-splitter for
finer-grained control (per-character streaming, specific tokenisation,
etc.).
Not safe for production use. Doesn't talk to any real model; any
parameter in Settings is ignored (temperature, top_p, stop, etc.).
Only the raw text in responses gets returned.
ATTRIBUTES
@.responsesā array of Str, one per expected call.$.fail-on-emptyā when True, calling beyond the queue fails instead of repeating the last response. Default False.$.stream-delayā seconds between tokens in streaming mode. Default 0 (emit as fast as possible).&.token-splitterā sub that takes a Str and returns a list of tokens. Default splits on whitespace and preserves spacing by re-appending a space after each token except the last.&.error-producerā optional(Int $call-index --Hash)> callback. When it returns a defined hash, that call is scripted to fail. Used to exercise the Task fallback policy. See /SIMULATING FAILURES.&.finish-reason-producerā optional(Int $call-index --Str)> callback. A defined Str return is stamped on the Response as itsfinish-reasonwhile the call still succeeds ā the way a real blocking completion reports'length'. See /SIMULATING TRUNCATION.$.call-indexā monotonically-increasing per-backend call count, bumped on every completion call whether it succeeded or failed. Distinct from the internal response cursor ā tests can read this to assert call counts without reasoning about the response queue.$.response-classā class minted for non-streaming completion calls. Defaults to LLM::Chat::Backend::Response; pass LLM::Chat::Backend::Response::OpenRouter to exercise the OR-flavoured Response shape.$.stream-response-classā streaming counterpart to$.response-class. Pair with LLM::Chat::Backend::Response::OpenRouter::Stream.%.fake-usageā per-call usage payload. OAI-spec keys (prompt,completion,total,model) flow into._set-usage; OR-specific keys (cost,generation-id,provider-name,is-byok) flow into._set-or-usagewhen the configured response class consumes the OR Augment role.
SIMULATING FAILURES
Pass &.error-producer to script per-call failures. Every
chat-completion invocation calls it with the 0-based call index.
Returning a defined hash fails the call; returning Nil (or an
undefined value) lets the call proceed normally through the
@.responses queue.
# Fail the first two calls with retryable errors, succeed on the third
my $mock = LLM::Chat::Backend::Mock.new(
settings => LLM::Chat::Backend::Settings.new,
responses => ['finally'],
error-producer => -> $i {
when $i == 0 { { class => 'connection', message => 'ECONNRESET' } }
when $i == 1 { { class => 'http', status => 503, message => 'bad gateway' } }
default { Nil }
},
);
Recognised class values match the Response error classification:
'http' (with numeric status), 'timeout', 'connection',
'response' (malformed / empty body / finish-reason failure), and
'unknown'. Failed calls DO NOT consume a slot from @.responses,
so the queue addresses only the successful calls regardless of where
failures fire.
SIMULATING TRUNCATION
A real blocking completion that runs out of completion budget does
not fail: it returns HTTP 200 with a perfectly well-formed body, a
partial content, and finish_reason 'length'. Callers only
learn the output was cut off by reading $response.finish-reason.
&.finish-reason-producer reproduces exactly that shape.
The callback receives the 0-based per-backend call index (the same
index &.error-producer sees) and returns either a Str ā stamped on
the Response via _set-finish-reason before the emission task starts,
so it is readable the instant the caller has the Response ā or an
undefined value, which leaves finish-reason unset (the historical
Mock shape).
# First call comes back truncated, second one is clean.
my $mock = LLM::Chat::Backend::Mock.new(
settings => LLM::Chat::Backend::Settings.new,
responses => ['{"beats": [{"first": 0,', '{"beats": []}'],
finish-reason-producer => -> $i { $i == 0 ?? 'length' !! Str },
);
my $first = $mock.chat-completion(@messages);
await $first.supply; # succeeds ā 200 with a partial body
is $first.finish-reason, 'length'; # ... but the body was cut off
Two deliberate limits:
The truncated call still succeeds.
is-successstays True and.msgholds whatever@.responsessupplied ā that is the point, since it is what a real backend does. Use&.error-producerwhen you want the call to fail outright.&.error-producerwins. When both callbacks fire on the same index the call is scripted to fail and no finish reason is stamped, because a failed call never produced a body to have a finish reason for.
The knob is wired into the non-streaming paths only
(chat-completion, and text-completion through its delegation).
The streaming paths already surface finish reasons the way the real
backends do ā a 'length' chunk quits the stream with error class
'response' ā so scripting one there would contradict the transport
being modelled. Tests that need a truncated stream should exercise
LLM::Chat::Backend::OpenAICommon against a canned SSE body instead
(see t/19-length-finish-stream.rakutest).
RECORDING CALLS
Every completion call is appended to @.recorded-calls as a hash with:
kindā 'chat-completion', 'chat-completion-stream', 'text-completion', or 'text-completion-stream'messagesā the@messagesarray that was passed intoolsā the@toolsarray (empty if none were passed)continuationā for text-completion only, the flag valueresponseā the Str that was returnedatā Instant when the call was made
Tests can assert on what reached the backend rather than just what came back. Typical pattern:
my $mock = LLM::Chat::Backend::Mock.new(
settings => LLM::Chat::Backend::Settings.new,
responses => ['ok'],
);
# ... code under test calls $mock.chat-completion-stream(@msgs) ...
is $mock.recorded-calls.elems, 1, 'one call';
is $mock.recorded-calls[0]<kind>, 'chat-completion-stream';
is $mock.recorded-calls[0]<messages>[0].role, 'system', 'first message is system prompt';
is $mock.recorded-calls[0]<messages>[*-1].content, 'hello', 'last message is user turn';
Use clear-recorded-calls to reset the log between phases of a test.