OpenAICommon
NAME
LLM::Chat::Backend::OpenAICommon - Generic OpenAI-compatible chat backend
SYNOPSIS
use LLM::Chat::Backend::OpenAICommon;
use LLM::Chat::Backend::Settings;
# Any OpenAI-compatible endpoint (Together, Groq, Fireworks, vLLM,
# llama.cpp's OAI server, ...). For OpenRouter specifically, prefer
# the L<LLM::Chat::Backend::OpenRouter> subclass β it adds the
# attribution headers, body extras, and cost/generation-id lifts.
my $backend = LLM::Chat::Backend::OpenAICommon.new(
api_url => 'https://api.together.xyz/v1',
api_key => %*ENV<TOGETHER_API_KEY>,
model => 'meta-llama/Llama-3.3-70B-Instruct-Turbo',
settings => LLM::Chat::Backend::Settings.new(:max_tokens(4096)),
);
my $resp = $backend.chat-completion(@messages);
react {
whenever $resp.supply -> $tok { print $tok }
whenever $resp.supply.done { say "done β {$resp.prompt-tokens} in / {$resp.completion-tokens} out" }
}
DESCRIPTION
OpenAI-compatible HTTP client. Uses Cro::HTTP::Client against
/chat/completions and /completions endpoints, lifts the OAI-spec
usage block onto the returned Response, and classifies failures
into the categorical error-class shape the LLM::Data::Inference
fallback policy expects.
EXTENSION
Provider-specific subclasses extend this class to add wire fields,
headers, or Response shapes that aren't part of the OAI spec. To
keep that path clean, the following internal methods are
underscore-prefixed (rather than !-private) so subclasses can
override them:
_get-api-settingsβ body params for every request._get-api-headersβ HTTP headers for every request._finalize-request-bodyβ last look at the assembled body, after messages / prompt / tools / stream are attached and before it goes on the wire._lift-usageβ extract usage / model from a response body or stream chunk into the Response.make-responseβ factory for non-streaming Response objects (used bychat-completion/text-completion).make-stream-responseβ factory for streaming Response objects._on-blocking-completeβ fires after a non-streaming completion's body has been parsed and lifted, before$response.done. Subclasses use it to attach post-call metadata (e.g. OpenRouter's/generationcost lookup)._on-stream-completeβ same hook for the streaming path; fires after[DONE], before$response.done.
LLM::Chat::Backend::OpenRouter overrides all of these.
_blob-text, _body-text, _decode-json-body and
_stream-decoder are also underscore-prefixed, but as shared
internals rather than override points: they exist so subclasses
making their own HTTP calls (such as OpenRouter's /generation
lookup) read response bodies the same way the completion paths do. See
their declarations for why LLM::Chat decodes response bytes itself
instead of using Cro's body-text / await .body, and why a
streamed body needs an incremental decoder rather than a
.decode per chunk.
!classify-exception stays private β it's pure Raku/Cro mapping
with no provider-specific behaviour.
Request-body finalization
_get-api-settings can only describe how to generate β model,
samplers, response format β because it runs before the payload is
attached. Anything that has to look at the conversation itself needs
_finalize-request-body, which every completion path calls with the
finished body just before the POST:
class MyBackend is LLM::Chat::Backend::OpenAICommon {
# Some gateways bill a "priority" tier per request and want the
# flag next to the payload rather than in the sampler block.
method _finalize-request-body(%settings --> Hash) {
my %out = %settings;
%out<priority> = 'fast' if (%out<messages> // []).elems > 20;
%out;
}
}
Three rules make an override safe:
Match the signature exactly.
method _finalize-request-body(%settings --> Hash), character for character. Raku treats a divergent signature as a new method rather than an override, and the parent's no-op keeps answering the call with no error anywhere.Return the body. The caller replaces its body with your return value, so an override that mutates in place and returns something else (the last statement of an
if, say) throws the request away.Build, don't reach back.
%settings<messages>holds hashes minted byLLM::Chat::Conversation::Message.to-hashfor this one request. Rewriting them by constructing new hashes is fine; treating them as a handle on the caller's conversation is not.
Both chat paths and both text paths call it, so an override that only
makes sense for one shape must say so itself β check for
%settings<messages>:exists (chat) or %settings<prompt>:exists
(text) and return the body unchanged otherwise.
LLM::Chat::Backend::OpenRouter does exactly that for its
cache_control breakpoints.
STREAM TERMINATION CONTRACT
A streamed generation can end in exactly three ways, and both
chat-completion-stream and text-completion-stream route every
one of them through the same private !settle-stream, so the
Response a consumer inspects means the same thing whichever method
produced it and whichever provider was on the other end.
1. The provider says it is finished
Either the [DONE] sentinel arrives, or a finish_reason of
stop / tool_calls does and the body then closes. This is the
success path: _on-stream-complete fires first, then the supply
completes.
.is-done True
.is-success True
.finish-reason 'stop' / 'tool_calls' / undefined ([DONE] with no
reason chunk β some providers send none)
.error-class undefined
2. The provider names a terminal failure
A finish_reason of length or content_filter, or one this
library does not recognise. The supply is quit where the reason is
read; !settle-stream sees an already-terminal Response and does
nothing.
.is-done True
.is-success False
.finish-reason 'length' / 'content_filter' / whatever arrived
.error-class 'response'
.err "Hit max tokens" / "Blocked by content filter" /
"Unknown finish reason: ..."
Note that length is a failure here. A generation stopped by the
completion budget is a truncated one, and the caller decides whether
a truncated answer is worth keeping β it never arrives labelled as a
complete one. The partial text is still on the Response
(.latest / .msg).
3. The body just stops
The connection closes with neither [DONE] nor a finish reason: a
dropped connection, a killed upstream worker, a proxy that gave up
mid-stream. There is no way to tell a reply that ended early from one
that ended, so this fails.
.is-done True
.is-success False
.finish-reason undefined
.error-class 'response'
.err 'Stream closed without [DONE] or a finish reason'
The important half of this case is that the Response settles at
all. It used to be left unterminated, and every consumer waiting on
.is-done waited out its own timeout instead β minutes, per
occurrence, for a stream that had been over the whole time.
Dropped frames
A data: line that does not parse as JSON is skipped rather than
fatal β providers interleave error fragments and truncated frames
into otherwise healthy streams β but each one bumps
Response.dropped-frames, and that counter changes how the stream
is allowed to end:
Prose only. The stream terminates normally, per the three cases above, and
.dropped-framesis left for the consumer to read. A hole in prose costs a few words and is visible to whoever reads the reply.Tool calls assembled. The stream fails with error-class
response, whatever else was going on β including a perfectly well-formed[DONE]. Tool-callargumentsare streamed as JSON fragments that are concatenated blind, so a hole in them produces either a parse failure or, far worse, valid JSON that says something the model never asked for. That is corruption no downstream consumer can detect, so it never reachesis-success.
Subclasses
Neither LLM::Chat::Backend::OpenRouter nor
LLM::Chat::Backend::KoboldCpp overrides the streaming methods β
they extend settings, headers, Response shapes and the completion
hooks β so this contract is theirs too. A subclass that did override
a stream loop would have to settle the Response itself; the
straightforward way is to keep the react block's structure and
call !settle-stream the same three places this class does (the
[DONE] arm, a stop-family arm that wants to close early, and
once on the way out of the react block).