Artifacts

NAME

LLM::Agent::Artifacts - when one tool result is bigger than the conversation can afford

SYNOPSIS


use LLM::Agent::Artifacts;

# What the loop does for you, when a result is over
# `request-budget.max-observation-size`:

my $store = LLM::Agent::Artifacts::Store.new(
    dir => artifact-dir-for($session.path),   # <transcript-stem>.artifacts/
);

my %stored = $store.store($op.op-id, $huge);  # { file, digest, bytes, chars }

my Str $seen = excerpt(
    $huge, 16_384,
    marker => stored-marker(elided-chars($huge.chars, 16_384), %stored),
);

# $seen — not $huge — is what the model sees, what the transcript holds,
# and what the ToolResult event carries. The full bytes are in the file
# and nowhere else.

DESCRIPTION

One fs_read of a 400KB file is bigger than some context windows on its own, and — this is the part that hurts — it is immortal: it goes into the conversation, into the transcript, and into every request for the rest of the run, and compaction refuses to eat the recent turns it sits in. The answer is to put an excerpt in the conversation and the whole thing in a file beside the transcript.

The invariant

One string, four places. The excerpt is computed once, on the settle path, before the tool message is built — so the conversation content, the session payload, the ToolResult event's content and the digest on the tool-settled envelope are all the same string. The full bytes live in the artifact file and nowhere else.

Everything follows from that:

  • a resumed run replays the excerpt byte for byte, and the seed check (LLM::Agent::Canonical's digests) passes;

  • replay never needs the artifact to exist. Delete the whole .artifacts directory and the transcript still loads, still resumes, and still produces exactly the same next request. What is lost is the ability to go and read what the tool really said, which is a different kind of loss from a session that will not open;

  • a crash between the write and the settle envelope leaves an orphan file, which nothing reads and nothing trips over. The write happens first on purpose: an envelope that named a file which does not exist would be a lie in the durable record, and an unreferenced file is not.

Characters, not bytes

The threshold is .chars — graphemes — because everything downstream is Str-shaped and Raku's substr is grapheme-safe: an excerpt can never be cut through the middle of a character, or between a base character and its combining accent, or between the CR and the LF of a Windows line ending (Raku makes those one grapheme). A byte-oriented cut would have to worry about all three.

The byte count is recorded in the metadata, because that is the number that matches what ls and shasum say about the file.

The excerpt

Head-weighted, three quarters to one quarter, with one marker line between them:


    <the first 3/4 of the budget>
    [... elided 391,204 chars (402,110 bytes total); full result: 019903...txt sha256 4f2c... ...]
    <the last 1/4 of the budget>

Head-weighted because the top of a file, a directory listing or a diff is where the identifying information is, and the tail is kept at all because the end is where errors and totals are.

The marker names the artifact by basename, never by path. Transcripts get moved, copied out of a container, or read on a different machine; a path recorded inside the conversation would be wrong the first time any of that happened, and the file is always in the .artifacts directory beside the transcript that names it.

A run with no session has nowhere to put the full bytes, so it says so — the marker records the size and the digest and states that the full result was not stored. The model is told the same thing either way: this is an excerpt, and here is how much is missing.

Retention: none

Nothing here deletes anything, ever. A long-lived agent directory grows one file per oversized tool result, and cleaning it up — by age, by size, by which transcripts still exist — is a decision an application makes, not one a library makes silently about files somebody may want.

Not the same thing as the compactor's cap

LLM::Agent::Compactor's tool-result-cap truncates tool results in the text it sends the summarizer, and changes nothing about the conversation. This changes what the conversation contains in the first place. They are different layers and both are worth having: this one stops a huge result entering the transcript, that one stops a compaction choking on the ones that already did.

Server-side truncation (an MCP tool that returns its own "output truncated" note) is upstream of all of this and unparseable in general, so nothing here tries: the size is measured, never inferred from the payload.

SEE ALSO

LLM::Agent::RequestBudget (where max-observation-size lives), LLM::Agent::Loop (the settle path that calls this), LLM::Agent::Session (the artifact key on a tool-settled line).

LLM::Agent v0.6.1

a streaming agent loop: tools, retry, fallback, a durable

Authors

  • Matt Doughty

License

Artistic-2.0

Dependencies

Digest::SHA256::Native:ver<1.0.0+>:auth<zef:bduggan>LLM::Chat:ver<0.10.0+>:auth<zef:apogee>MCP::Client:ver<0.5.0+>:auth<zef:apogee>JSONL:ver<0.1.6+>:auth<zef:apogee>JSON::Fast:ver<0.19+>:auth<cpan:TIMOTIMO>UUID::V4:ver<1.0.0+>:auth<zef:masukomi>

Test Dependencies

Provides

  • LLM::Agent
  • LLM::Agent::Artifacts
  • LLM::Agent::Canonical
  • LLM::Agent::Compactor
  • LLM::Agent::CompletionBus
  • LLM::Agent::Event
  • LLM::Agent::Loop
  • LLM::Agent::Prompt
  • LLM::Agent::RequestBudget
  • LLM::Agent::Run
  • LLM::Agent::RunContext
  • LLM::Agent::Session
  • LLM::Agent::Subagents
  • LLM::Agent::TokenCount
  • LLM::Agent::ToolOperation

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite — the markup and publishing tools behind this site.