OpenRouter

NAME

LLM::Chat::Backend::OpenRouter - OpenRouter-specific chat backend

SYNOPSIS


use LLM::Chat::Backend::OpenRouter;
use LLM::Chat::Backend::Settings;

my $backend = LLM::Chat::Backend::OpenRouter.new(
    api_key  => %*ENV<OPENROUTER_API_KEY>,
    model    => 'anthropic/claude-opus-4-7',
    settings => LLM::Chat::Backend::Settings.new(:max_tokens(8192)),

    # Optional attribution headers β€” let your app appear on the
    # OpenRouter rankings page and in users' generation logs.
    http-referer => 'https://example.com/my-app',
    x-title      => 'My App',
);

my $resp = $backend.chat-completion(@messages);
react {
    whenever $resp.supply -> $tok { print $tok }
    whenever $resp.supply.done {
        say "cost: \${$resp.cost // 0}";
        say "served by: {$resp.provider-name // 'unknown'}";
        say "lookup: /generation?id={$resp.generation-id}" if $resp.generation-id.defined;
    }
}

DESCRIPTION

Subclass of LLM::Chat::Backend::OpenAICommon that wires in OpenRouter-specific behaviour without touching the generic OpenAI-compatible code path:

  • Sends include_reasoning: True + top_k on every request, mirroring the body shape SillyTavern sends to OpenRouter. Adds reasoning: { effort: ... } when settings.reasoning_effort is set, again following SillyTavern's shape. Does not send usage: { include: true } or stream_options: { include_usage: true } β€” those caused intermittent header-phase hangs against some upstream providers (DeepSeek-V3.2 mostly fine, others ~80 % fail).

  • Marks up to three cache_control breakpoints on outgoing chat bodies (controlled by $.cache-breakpoints, on by default) so providers that cache explicitly can reuse the prompt prefix between rounds β€” see Prompt caching below.

  • Adds HTTP-Referer / X-Title attribution headers when :http-referer / :x-title are configured. (See https://openrouter.ai/docs/api-reference/overview#headers.)

  • Returns LLM::Chat::Backend::Response::OpenRouter / LLM::Chat::Backend::Response::OpenRouter::Stream so callers can read .cost, .generation-id, .provider-name, and .is-byok.

  • Lifts those fields off the response body / stream chunks via _lift-usage (calls callsame first to handle OAI-spec).

  • After a completion finishes β€” streaming or blocking β€” fires a one-shot GET against /generation?id=... to populate .cost / .provider-name from OpenRouter's metadata endpoint. Replaces the inline usage.cost we used to ask for via usage: { include: true }; lookup is best-effort, so a failure leaves .cost Nil rather than erroring the call. Wired in via the _on-stream-complete / _on-blocking-complete hooks that the parent class fires before $response.done, so consumers see populated cost metadata by the time the response is observably done on either path.

Everything else β€” request shape, error classification, fallback / retry interaction, streaming mechanics, cancel β€” is inherited unchanged from OpenAICommon.

Prompt caching

A long-running conversation re-sends its whole history every round, and the input bill is the bulk of what that costs. Providers will serve a repeated prefix out of cache at a fraction of the price β€” some (OpenAI, DeepSeek, Moonshot, Z.AI, Grok) do it implicitly, and some (Anthropic, Qwen, Gemini) only when the request says where the reusable prefix ends, with a cache_control marker on a content-parts block.

This backend places those markers itself, on the head message and on the last two user / assistant turns, so the marker written one round ago still covers everything the next request shares with it:


# On by default β€” nothing to configure for the common case.
my $backend = LLM::Chat::Backend::OpenRouter.new(
    api_key => %*ENV<OPENROUTER_API_KEY>,
    model   => 'anthropic/claude-opus-4-7',
);

# Off: messages go on the wire as plain `content => 'text'` strings,
# byte for byte what earlier releases sent.
my $plain = LLM::Chat::Backend::OpenRouter.new(
    api_key           => %*ENV<OPENROUTER_API_KEY>,
    model             => 'anthropic/claude-opus-4-7',
    cache-breakpoints => False,
);

Implicit-caching routes ignore the markers, so the flag can stay on across a mixed model fleet. Conversation state is untouched either way β€” the rewrite happens on the serialized request body, so Message.get-checksum is unaffected by where a breakpoint landed.

Read the payoff back off the Response: .cached-prompt-tokens is the cached slice of .prompt-tokens when the provider reports one.


with $resp.cached-prompt-tokens -> $cached {
    say "cache: $cached / {$resp.prompt-tokens} prompt tokens";
}

ATTRIBUTES

  • $.api_url β€” base URL. Defaults to https://openrouter.ai/api/v1.

  • $.api_key β€” bearer token (an OpenRouter inference key).

  • $.model β€” model id (e.g. anthropic/claude-opus-4-7).

  • $.http-referer β€” optional. Sent as the HTTP-Referer header.

  • $.x-title β€” optional. Sent as the X-Title header.

  • $.cache-breakpoints β€” Bool, default True. Whether chat bodies carry cache_control breakpoints. Text-completion bodies are never annotated.

RESPONSE FIELDS

The Response objects returned by this backend are typed LLM::Chat::Backend::Response::OpenRouter (or its ::Stream subclass) and carry, in addition to the inherited OAI-spec fields:

  • $.cost β€” USD spent (Num), from usage.cost.

  • $.generation-id β€” OpenRouter's gen-XXXX id, suitable for the /generation endpoint.

  • $.provider-name β€” provider that actually served the request (e.g. Anthropic).

  • $.is-byok β€” True when the call used the user's BYOK keys.

All four are presence-gated β€” read with .cost.defined etc.

The inherited $.cached-prompt-tokens is filled in on this backend from either end of the call: from the completion body's usage.prompt_tokens_details.cached_tokens when the route reports it inline, and otherwise from the /generation poll that also carries cost. The body's number always wins β€” the poll only fills a field the body left unreported, which is what makes the streaming path (no inline usage frame under this backend's request shape) report cache hits at all.

LLM::Chat v0.10.0

Simple framework for LLM inferencing

Authors

  • Matt Doughty

License

Artistic-2.0

Dependencies

Cro::HTTP::Client:ver<0.8.11+>:auth<zef:cro>Digest::SHA256::Native:ver<1.0.0+>:auth<zef:bduggan>JSON::Fast:ver<0.19+>:auth<cpan:TIMOTIMO>Template::Jinja2:ver<0.3.0+>:auth<zef:apogee>Tokenizers:ver<0.3.0+>:auth<zef:apogee>UUID::V4:ver<1.0.0+>:auth<zef:masukomi>

Provides

  • LLM::Chat
  • LLM::Chat::Backend
  • LLM::Chat::Backend::KoboldCpp
  • LLM::Chat::Backend::Mock
  • LLM::Chat::Backend::OpenAICommon
  • LLM::Chat::Backend::OpenRouter
  • LLM::Chat::Backend::Response
  • LLM::Chat::Backend::Response::OpenRouter
  • LLM::Chat::Backend::Response::OpenRouter::Augment
  • LLM::Chat::Backend::Response::OpenRouter::Stream
  • LLM::Chat::Backend::Response::Stream
  • LLM::Chat::Backend::Settings
  • LLM::Chat::Conversation
  • LLM::Chat::Conversation::Message
  • LLM::Chat::Debug
  • LLM::Chat::Retry
  • LLM::Chat::Retry::Exceptions
  • LLM::Chat::Template
  • LLM::Chat::Template::ChatML
  • LLM::Chat::Template::DeepSeekV4
  • LLM::Chat::Template::Gemma2
  • LLM::Chat::Template::Jinja2
  • LLM::Chat::Template::Llama3
  • LLM::Chat::Template::Llama4
  • LLM::Chat::Template::MistralV7
  • LLM::Chat::TokenCounter
  • LLM::Chat::ToolLoop

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite β€” the markup and publishing tools behind this site.