OpenAICommon

NAME

LLM::Chat::Backend::OpenAICommon - Generic OpenAI-compatible chat backend

SYNOPSIS


use LLM::Chat::Backend::OpenAICommon;
use LLM::Chat::Backend::Settings;

# Any OpenAI-compatible endpoint (Together, Groq, Fireworks, vLLM,
# llama.cpp's OAI server, ...). For OpenRouter specifically, prefer
# the L<LLM::Chat::Backend::OpenRouter> subclass β€” it adds the
# attribution headers, body extras, and cost/generation-id lifts.
my $backend = LLM::Chat::Backend::OpenAICommon.new(
    api_url  => 'https://api.together.xyz/v1',
    api_key  => %*ENV<TOGETHER_API_KEY>,
    model    => 'meta-llama/Llama-3.3-70B-Instruct-Turbo',
    settings => LLM::Chat::Backend::Settings.new(:max_tokens(4096)),
);

my $resp = $backend.chat-completion(@messages);
react {
    whenever $resp.supply -> $tok { print $tok }
    whenever $resp.supply.done    { say "done β€” {$resp.prompt-tokens} in / {$resp.completion-tokens} out" }
}

DESCRIPTION

OpenAI-compatible HTTP client. Uses Cro::HTTP::Client against /chat/completions and /completions endpoints, lifts the OAI-spec usage block onto the returned Response, and classifies failures into the categorical error-class shape the LLM::Data::Inference fallback policy expects.

EXTENSION

Provider-specific subclasses extend this class to add wire fields, headers, or Response shapes that aren't part of the OAI spec. To keep that path clean, the following internal methods are underscore-prefixed (rather than !-private) so subclasses can override them:

  • _get-api-settings β€” body params for every request.

  • _get-api-headers β€” HTTP headers for every request.

  • _finalize-request-body β€” last look at the assembled body, after messages / prompt / tools / stream are attached and before it goes on the wire.

  • _lift-usage β€” extract usage / model from a response body or stream chunk into the Response.

  • make-response β€” factory for non-streaming Response objects (used by chat-completion / text-completion).

  • make-stream-response β€” factory for streaming Response objects.

  • _on-blocking-complete β€” fires after a non-streaming completion's body has been parsed and lifted, before $response.done. Subclasses use it to attach post-call metadata (e.g. OpenRouter's /generation cost lookup).

  • _on-stream-complete β€” same hook for the streaming path; fires after [DONE], before $response.done.

LLM::Chat::Backend::OpenRouter overrides all of these.

_blob-text, _body-text, _decode-json-body and _stream-decoder are also underscore-prefixed, but as shared internals rather than override points: they exist so subclasses making their own HTTP calls (such as OpenRouter's /generation lookup) read response bodies the same way the completion paths do. See their declarations for why LLM::Chat decodes response bytes itself instead of using Cro's body-text / await .body, and why a streamed body needs an incremental decoder rather than a .decode per chunk.

!classify-exception stays private β€” it's pure Raku/Cro mapping with no provider-specific behaviour.

Request-body finalization

_get-api-settings can only describe how to generate β€” model, samplers, response format β€” because it runs before the payload is attached. Anything that has to look at the conversation itself needs _finalize-request-body, which every completion path calls with the finished body just before the POST:


class MyBackend is LLM::Chat::Backend::OpenAICommon {
    # Some gateways bill a "priority" tier per request and want the
    # flag next to the payload rather than in the sampler block.
    method _finalize-request-body(%settings --> Hash) {
        my %out = %settings;
        %out<priority> = 'fast' if (%out<messages> // []).elems > 20;
        %out;
    }
}

Three rules make an override safe:

  • Match the signature exactly. method _finalize-request-body(%settings --> Hash), character for character. Raku treats a divergent signature as a new method rather than an override, and the parent's no-op keeps answering the call with no error anywhere.

  • Return the body. The caller replaces its body with your return value, so an override that mutates in place and returns something else (the last statement of an if, say) throws the request away.

  • Build, don't reach back. %settings<messages> holds hashes minted by LLM::Chat::Conversation::Message.to-hash for this one request. Rewriting them by constructing new hashes is fine; treating them as a handle on the caller's conversation is not.

Both chat paths and both text paths call it, so an override that only makes sense for one shape must say so itself β€” check for %settings<messages>:exists (chat) or %settings<prompt>:exists (text) and return the body unchanged otherwise. LLM::Chat::Backend::OpenRouter does exactly that for its cache_control breakpoints.

STREAM TERMINATION CONTRACT

A streamed generation can end in exactly three ways, and both chat-completion-stream and text-completion-stream route every one of them through the same private !settle-stream, so the Response a consumer inspects means the same thing whichever method produced it and whichever provider was on the other end.

1. The provider says it is finished

Either the [DONE] sentinel arrives, or a finish_reason of stop / tool_calls does and the body then closes. This is the success path: _on-stream-complete fires first, then the supply completes.

.is-done      True
    .is-success   True
    .finish-reason  'stop' / 'tool_calls' / undefined ([DONE] with no
                    reason chunk β€” some providers send none)
    .error-class    undefined

2. The provider names a terminal failure

A finish_reason of length or content_filter, or one this library does not recognise. The supply is quit where the reason is read; !settle-stream sees an already-terminal Response and does nothing.

.is-done      True
    .is-success   False
    .finish-reason  'length' / 'content_filter' / whatever arrived
    .error-class    'response'
    .err            "Hit max tokens" / "Blocked by content filter" /
                    "Unknown finish reason: ..."

Note that length is a failure here. A generation stopped by the completion budget is a truncated one, and the caller decides whether a truncated answer is worth keeping β€” it never arrives labelled as a complete one. The partial text is still on the Response (.latest / .msg).

3. The body just stops

The connection closes with neither [DONE] nor a finish reason: a dropped connection, a killed upstream worker, a proxy that gave up mid-stream. There is no way to tell a reply that ended early from one that ended, so this fails.

.is-done      True
    .is-success   False
    .finish-reason  undefined
    .error-class    'response'
    .err            'Stream closed without [DONE] or a finish reason'

The important half of this case is that the Response settles at all. It used to be left unterminated, and every consumer waiting on .is-done waited out its own timeout instead β€” minutes, per occurrence, for a stream that had been over the whole time.

Dropped frames

A data: line that does not parse as JSON is skipped rather than fatal β€” providers interleave error fragments and truncated frames into otherwise healthy streams β€” but each one bumps Response.dropped-frames, and that counter changes how the stream is allowed to end:

  • Prose only. The stream terminates normally, per the three cases above, and .dropped-frames is left for the consumer to read. A hole in prose costs a few words and is visible to whoever reads the reply.

  • Tool calls assembled. The stream fails with error-class response, whatever else was going on β€” including a perfectly well-formed [DONE]. Tool-call arguments are streamed as JSON fragments that are concatenated blind, so a hole in them produces either a parse failure or, far worse, valid JSON that says something the model never asked for. That is corruption no downstream consumer can detect, so it never reaches is-success.

Subclasses

Neither LLM::Chat::Backend::OpenRouter nor LLM::Chat::Backend::KoboldCpp overrides the streaming methods β€” they extend settings, headers, Response shapes and the completion hooks β€” so this contract is theirs too. A subclass that did override a stream loop would have to settle the Response itself; the straightforward way is to keep the react block's structure and call !settle-stream the same three places this class does (the [DONE] arm, a stop-family arm that wants to close early, and once on the way out of the react block).

LLM::Chat v0.10.0

Simple framework for LLM inferencing

Authors

  • Matt Doughty

License

Artistic-2.0

Dependencies

Cro::HTTP::Client:ver<0.8.11+>:auth<zef:cro>Digest::SHA256::Native:ver<1.0.0+>:auth<zef:bduggan>JSON::Fast:ver<0.19+>:auth<cpan:TIMOTIMO>Template::Jinja2:ver<0.3.0+>:auth<zef:apogee>Tokenizers:ver<0.3.0+>:auth<zef:apogee>UUID::V4:ver<1.0.0+>:auth<zef:masukomi>

Provides

  • LLM::Chat
  • LLM::Chat::Backend
  • LLM::Chat::Backend::KoboldCpp
  • LLM::Chat::Backend::Mock
  • LLM::Chat::Backend::OpenAICommon
  • LLM::Chat::Backend::OpenRouter
  • LLM::Chat::Backend::Response
  • LLM::Chat::Backend::Response::OpenRouter
  • LLM::Chat::Backend::Response::OpenRouter::Augment
  • LLM::Chat::Backend::Response::OpenRouter::Stream
  • LLM::Chat::Backend::Response::Stream
  • LLM::Chat::Backend::Settings
  • LLM::Chat::Conversation
  • LLM::Chat::Conversation::Message
  • LLM::Chat::Debug
  • LLM::Chat::Retry
  • LLM::Chat::Retry::Exceptions
  • LLM::Chat::Template
  • LLM::Chat::Template::ChatML
  • LLM::Chat::Template::DeepSeekV4
  • LLM::Chat::Template::Gemma2
  • LLM::Chat::Template::Jinja2
  • LLM::Chat::Template::Llama3
  • LLM::Chat::Template::Llama4
  • LLM::Chat::Template::MistralV7
  • LLM::Chat::TokenCounter
  • LLM::Chat::ToolLoop

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite β€” the markup and publishing tools behind this site.