Mock

NAME

LLM::Chat::Backend::Mock - Canned-response backend for tests

SYNOPSIS


use LLM::Chat::Backend::Mock;
use LLM::Chat::Backend::Settings;

# Single canned response
my $mock = LLM::Chat::Backend::Mock.new(
    settings => LLM::Chat::Backend::Settings.new,
    responses => ['hello from the mock backend'],
);

# A sequence — each call pops the next one
my $mock2 = LLM::Chat::Backend::Mock.new(
    settings  => LLM::Chat::Backend::Settings.new,
    responses => [
        'first answer',
        'second answer',
        'third answer',
    ],
);

# Streaming: tokens arrive split by whitespace (default).
# Customise with :token-splitter if you want per-character streaming or
# a specific token boundary.
my $stream = $mock2.chat-completion-stream(@messages);
react {
    whenever $stream.supply -> $token {
        print $token;
    }
    whenever $stream.supply.done { say "done" }
}

DESCRIPTION

Non-network backend for use in tests. Returns pre-configured responses in order — one response per chat-completion/text-completion call. When the queue is exhausted, subsequent calls either repeat the last response (default) or fail (:fail-on-empty).

For streaming calls, the response is split into tokens and emitted through the normal Supplier path. By default each whitespace- separated word becomes a token; override :token-splitter for finer-grained control (per-character streaming, specific tokenisation, etc.).

Not safe for production use. Doesn't talk to any real model; any parameter in Settings is ignored (temperature, top_p, stop, etc.). Only the raw text in responses gets returned.

ATTRIBUTES

  • @.responses — array of Str, one per expected call.

  • $.fail-on-empty — when True, calling beyond the queue fails instead of repeating the last response. Default False.

  • $.stream-delay — seconds between tokens in streaming mode. Default 0 (emit as fast as possible).

  • &.token-splitter — sub that takes a Str and returns a list of tokens. Default splits on whitespace and preserves spacing by re-appending a space after each token except the last.

  • &.error-producer — optional (Int $call-index -- Hash)> callback. When it returns a defined hash, that call is scripted to fail. Used to exercise the Task fallback policy. See /SIMULATING FAILURES.

  • &.finish-reason-producer — optional (Int $call-index -- Str)> callback. A defined Str return is stamped on the Response as its finish-reason while the call still succeeds — the way a real blocking completion reports 'length'. See /SIMULATING TRUNCATION.

  • $.call-index — monotonically-increasing per-backend call count, bumped on every completion call whether it succeeded or failed. Distinct from the internal response cursor — tests can read this to assert call counts without reasoning about the response queue.

  • $.response-class — class minted for non-streaming completion calls. Defaults to LLM::Chat::Backend::Response; pass LLM::Chat::Backend::Response::OpenRouter to exercise the OR-flavoured Response shape.

  • $.stream-response-class — streaming counterpart to $.response-class. Pair with LLM::Chat::Backend::Response::OpenRouter::Stream.

  • %.fake-usage — per-call usage payload. OAI-spec keys (prompt, completion, total, model) flow into ._set-usage; OR-specific keys (cost, generation-id, provider-name, is-byok) flow into ._set-or-usage when the configured response class consumes the OR Augment role.

SIMULATING FAILURES

Pass &.error-producer to script per-call failures. Every chat-completion invocation calls it with the 0-based call index. Returning a defined hash fails the call; returning Nil (or an undefined value) lets the call proceed normally through the @.responses queue.


# Fail the first two calls with retryable errors, succeed on the third
my $mock = LLM::Chat::Backend::Mock.new(
    settings  => LLM::Chat::Backend::Settings.new,
    responses => ['finally'],
    error-producer => -> $i {
        when $i == 0 { { class => 'connection', message => 'ECONNRESET' } }
        when $i == 1 { { class => 'http', status => 503, message => 'bad gateway' } }
        default      { Nil }
    },
);

Recognised class values match the Response error classification: 'http' (with numeric status), 'timeout', 'connection', 'response' (malformed / empty body / finish-reason failure), and 'unknown'. Failed calls DO NOT consume a slot from @.responses, so the queue addresses only the successful calls regardless of where failures fire.

SIMULATING TRUNCATION

A real blocking completion that runs out of completion budget does not fail: it returns HTTP 200 with a perfectly well-formed body, a partial content, and finish_reason 'length'. Callers only learn the output was cut off by reading $response.finish-reason. &.finish-reason-producer reproduces exactly that shape.

The callback receives the 0-based per-backend call index (the same index &.error-producer sees) and returns either a Str — stamped on the Response via _set-finish-reason before the emission task starts, so it is readable the instant the caller has the Response — or an undefined value, which leaves finish-reason unset (the historical Mock shape).


# First call comes back truncated, second one is clean.
my $mock = LLM::Chat::Backend::Mock.new(
    settings  => LLM::Chat::Backend::Settings.new,
    responses => ['{"beats": [{"first": 0,', '{"beats": []}'],
    finish-reason-producer => -> $i { $i == 0 ?? 'length' !! Str },
);

my $first = $mock.chat-completion(@messages);
await $first.supply;                 # succeeds — 200 with a partial body
is $first.finish-reason, 'length';   # ... but the body was cut off

Two deliberate limits:

  • The truncated call still succeeds. is-success stays True and .msg holds whatever @.responses supplied — that is the point, since it is what a real backend does. Use &.error-producer when you want the call to fail outright.

  • &.error-producer wins. When both callbacks fire on the same index the call is scripted to fail and no finish reason is stamped, because a failed call never produced a body to have a finish reason for.

The knob is wired into the non-streaming paths only (chat-completion, and text-completion through its delegation). The streaming paths already surface finish reasons the way the real backends do — a 'length' chunk quits the stream with error class 'response' — so scripting one there would contradict the transport being modelled. Tests that need a truncated stream should exercise LLM::Chat::Backend::OpenAICommon against a canned SSE body instead (see t/19-length-finish-stream.rakutest).

RECORDING CALLS

Every completion call is appended to @.recorded-calls as a hash with:

  • kind — 'chat-completion', 'chat-completion-stream', 'text-completion', or 'text-completion-stream'

  • messages — the @messages array that was passed in

  • tools — the @tools array (empty if none were passed)

  • continuation — for text-completion only, the flag value

  • response — the Str that was returned

  • at — Instant when the call was made

Tests can assert on what reached the backend rather than just what came back. Typical pattern:


my $mock = LLM::Chat::Backend::Mock.new(
    settings  => LLM::Chat::Backend::Settings.new,
    responses => ['ok'],
);

# ... code under test calls $mock.chat-completion-stream(@msgs) ...

is $mock.recorded-calls.elems, 1, 'one call';
is $mock.recorded-calls[0]<kind>, 'chat-completion-stream';
is $mock.recorded-calls[0]<messages>[0].role, 'system', 'first message is system prompt';
is $mock.recorded-calls[0]<messages>[*-1].content, 'hello', 'last message is user turn';

Use clear-recorded-calls to reset the log between phases of a test.

LLM::Chat v0.10.0

Simple framework for LLM inferencing

Authors

  • Matt Doughty

License

Artistic-2.0

Dependencies

Cro::HTTP::Client:ver<0.8.11+>:auth<zef:cro>Digest::SHA256::Native:ver<1.0.0+>:auth<zef:bduggan>JSON::Fast:ver<0.19+>:auth<cpan:TIMOTIMO>Template::Jinja2:ver<0.3.0+>:auth<zef:apogee>Tokenizers:ver<0.3.0+>:auth<zef:apogee>UUID::V4:ver<1.0.0+>:auth<zef:masukomi>

Provides

  • LLM::Chat
  • LLM::Chat::Backend
  • LLM::Chat::Backend::KoboldCpp
  • LLM::Chat::Backend::Mock
  • LLM::Chat::Backend::OpenAICommon
  • LLM::Chat::Backend::OpenRouter
  • LLM::Chat::Backend::Response
  • LLM::Chat::Backend::Response::OpenRouter
  • LLM::Chat::Backend::Response::OpenRouter::Augment
  • LLM::Chat::Backend::Response::OpenRouter::Stream
  • LLM::Chat::Backend::Response::Stream
  • LLM::Chat::Backend::Settings
  • LLM::Chat::Conversation
  • LLM::Chat::Conversation::Message
  • LLM::Chat::Debug
  • LLM::Chat::Retry
  • LLM::Chat::Retry::Exceptions
  • LLM::Chat::Template
  • LLM::Chat::Template::ChatML
  • LLM::Chat::Template::DeepSeekV4
  • LLM::Chat::Template::Gemma2
  • LLM::Chat::Template::Jinja2
  • LLM::Chat::Template::Llama3
  • LLM::Chat::Template::Llama4
  • LLM::Chat::Template::MistralV7
  • LLM::Chat::TokenCounter
  • LLM::Chat::ToolLoop

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite — the markup and publishing tools behind this site.