OpenRouter
NAME
LLM::Chat::Backend::OpenRouter - OpenRouter-specific chat backend
SYNOPSIS
use LLM::Chat::Backend::OpenRouter;
use LLM::Chat::Backend::Settings;
my $backend = LLM::Chat::Backend::OpenRouter.new(
api_key => %*ENV<OPENROUTER_API_KEY>,
model => 'anthropic/claude-opus-4-7',
settings => LLM::Chat::Backend::Settings.new(:max_tokens(8192)),
# Optional attribution headers β let your app appear on the
# OpenRouter rankings page and in users' generation logs.
http-referer => 'https://example.com/my-app',
x-title => 'My App',
);
my $resp = $backend.chat-completion(@messages);
react {
whenever $resp.supply -> $tok { print $tok }
whenever $resp.supply.done {
say "cost: \${$resp.cost // 0}";
say "served by: {$resp.provider-name // 'unknown'}";
say "lookup: /generation?id={$resp.generation-id}" if $resp.generation-id.defined;
}
}
DESCRIPTION
Subclass of LLM::Chat::Backend::OpenAICommon that wires in OpenRouter-specific behaviour without touching the generic OpenAI-compatible code path:
Sends
include_reasoning: True+top_kon every request, mirroring the body shape SillyTavern sends to OpenRouter. Addsreasoning: { effort: ... }whensettings.reasoning_effortis set, again following SillyTavern's shape. Does not sendusage: { include: true }orstream_options: { include_usage: true }β those caused intermittent header-phase hangs against some upstream providers (DeepSeek-V3.2 mostly fine, others ~80 % fail).Marks up to three
cache_controlbreakpoints on outgoing chat bodies (controlled by$.cache-breakpoints, on by default) so providers that cache explicitly can reuse the prompt prefix between rounds β see Prompt caching below.Adds
HTTP-Referer/X-Titleattribution headers when:http-referer/:x-titleare configured. (See https://openrouter.ai/docs/api-reference/overview#headers.)Returns LLM::Chat::Backend::Response::OpenRouter / LLM::Chat::Backend::Response::OpenRouter::Stream so callers can read
.cost,.generation-id,.provider-name, and.is-byok.Lifts those fields off the response body / stream chunks via
_lift-usage(callscallsamefirst to handle OAI-spec).After a completion finishes β streaming or blocking β fires a one-shot GET against
/generation?id=...to populate.cost/.provider-namefrom OpenRouter's metadata endpoint. Replaces the inlineusage.costwe used to ask for viausage: { include: true }; lookup is best-effort, so a failure leaves.costNil rather than erroring the call. Wired in via the_on-stream-complete/_on-blocking-completehooks that the parent class fires before$response.done, so consumers see populated cost metadata by the time the response is observably done on either path.
Everything else β request shape, error classification, fallback /
retry interaction, streaming mechanics, cancel β is inherited
unchanged from OpenAICommon.
Prompt caching
A long-running conversation re-sends its whole history every round,
and the input bill is the bulk of what that costs. Providers will
serve a repeated prefix out of cache at a fraction of the price β
some (OpenAI, DeepSeek, Moonshot, Z.AI, Grok) do it implicitly, and
some (Anthropic, Qwen, Gemini) only when the request says where the
reusable prefix ends, with a cache_control marker on a
content-parts block.
This backend places those markers itself, on the head message and on
the last two user / assistant turns, so the marker written one
round ago still covers everything the next request shares with it:
# On by default β nothing to configure for the common case.
my $backend = LLM::Chat::Backend::OpenRouter.new(
api_key => %*ENV<OPENROUTER_API_KEY>,
model => 'anthropic/claude-opus-4-7',
);
# Off: messages go on the wire as plain `content => 'text'` strings,
# byte for byte what earlier releases sent.
my $plain = LLM::Chat::Backend::OpenRouter.new(
api_key => %*ENV<OPENROUTER_API_KEY>,
model => 'anthropic/claude-opus-4-7',
cache-breakpoints => False,
);
Implicit-caching routes ignore the markers, so the flag can stay on
across a mixed model fleet. Conversation state is untouched either
way β the rewrite happens on the serialized request body, so
Message.get-checksum is unaffected by where a breakpoint landed.
Read the payoff back off the Response: .cached-prompt-tokens is
the cached slice of .prompt-tokens when the provider reports one.
with $resp.cached-prompt-tokens -> $cached {
say "cache: $cached / {$resp.prompt-tokens} prompt tokens";
}
ATTRIBUTES
$.api_urlβ base URL. Defaults tohttps://openrouter.ai/api/v1.$.api_keyβ bearer token (an OpenRouter inference key).$.modelβ model id (e.g.anthropic/claude-opus-4-7).$.http-refererβ optional. Sent as theHTTP-Refererheader.$.x-titleβ optional. Sent as theX-Titleheader.$.cache-breakpointsβBool, defaultTrue. Whether chat bodies carrycache_controlbreakpoints. Text-completion bodies are never annotated.
RESPONSE FIELDS
The Response objects returned by this backend are typed
LLM::Chat::Backend::Response::OpenRouter (or its ::Stream
subclass) and carry, in addition to the inherited OAI-spec fields:
$.costβ USD spent (Num), fromusage.cost.$.generation-idβ OpenRouter'sgen-XXXXid, suitable for the/generationendpoint.$.provider-nameβ provider that actually served the request (e.g.Anthropic).$.is-byokβ True when the call used the user's BYOK keys.
All four are presence-gated β read with .cost.defined etc.
The inherited $.cached-prompt-tokens is filled in on this backend
from either end of the call: from the completion body's
usage.prompt_tokens_details.cached_tokens when the route reports
it inline, and otherwise from the /generation poll that also
carries cost. The body's number always wins β the poll only fills a
field the body left unreported, which is what makes the streaming
path (no inline usage frame under this backend's request shape)
report cache hits at all.