Fetcher
NAME
MCP::Server::Tool::Web::Fetcher - one guarded HTTP GET, redirects, budgets and character sets included
SYNOPSIS
use MCP::Server::Tool::Web::Fetcher;
use MCP::Server::Tool::Web::Guard;
my $fetcher = MCP::Server::Tool::Web::Fetcher.new(
guard => MCP::Server::Tool::Web::Guard.new,
max-bytes => 2 * 1024 * 1024,
);
my $page = $fetcher.fetch('https://docs.example.com/guide');
say $page.status; # 200
say $page.final-url; # https://docs.example.com/guide/ (after a 301)
say $page.redirect-chain; # [https://docs.example.com/guide https://โฆ/guide/]
say $page.content-type; # text/html
say $page.charset; # utf-8 โ the encoding that actually worked
say $page.byte-count; # 48219
say $page.text.lines.head; # <!DOCTYPE html>
say $page.elapsed; # 0.31
A 404 is a result, not an exception. Only the request never happening is:
my $missing = $fetcher.fetch('https://example.com/nope');
say $missing.status; # 404
say $missing.text.chars; # the error page is still the body
{
$fetcher.fetch('http://169.254.169.254/latest/meta-data/');
CATCH {
when X::MCP::Server::Tool::Web::Blocked { say .message } # metadata endpoint
when X::MCP::Server::Tool::Web::BadUrl { say .message }
when X::MCP::Server::Tool::Web::Deadline { say .message }
when X::MCP::Server::Tool::Web::Transport { say .message }
}
}
One budget, shared by every page of a crawl:
my $budget = $fetcher.new-budget; # this Fetcher's caps, one call's worth
for @urls -> $url {
last if $budget.expired;
my $page = $fetcher.fetch($url, :$budget);
say "$url: {$page.byte-count} bytes"
~ ($page.truncated ?? " (cut: {$page.truncated-reason})" !! '');
}
DESCRIPTION
Everything between "here is a URL" and "here is the text of that page", with the pack's rules applied at every step: the guard vets the URL and, through MCP::Server::Tool::Web::Transport, the address behind it; the budget bounds the time and the bytes; redirects are followed by hand so that every hop is checked again; and the body is decoded with the encoding the bytes are actually in rather than the one the fetcher would prefer.
Redirects are followed here, not by Cro
Cro::HTTP::Client can follow redirects itself, and this class asks it not
to (follow => False). Every hop goes back through
Guard.check-url and, on connect, through the address check: a
Location: file:///etc/passwd, a Location: http://10.0.0.1/ or a redirect
that quietly drops from https to http is refused on the hop that produced it,
naming the rule. Cro's own loop would have skipped all of that from hop two
onwards.
The other reason is bookkeeping the tool surface needs: the chain, in order, and the final URL โ which is what a crawl's same-origin test and a relative link's base have to be resolved against.
Loops are detected on (origin, path, query), so a site that alternates between two URLs is refused as a loop rather than as "too many redirects": the first is a fact about the site that no setting will fix, the second is a budget.
The body is read against a ledger
The response body is streamed and counted. When the budget's byte allowance
runs out mid-body the download is cancelled and what arrived is kept; when the
deadline passes mid-body, likewise. Either way the result carries
truncated => True and a truncated-reason of 'max-bytes' or
'deadline', and a body that simply ended is never marked truncated,
however close to a limit it came.
Why head-only, and not head-and-tail
The shell tool keeps the head and the tail of a long output, because the
interesting part of a build log is usually at the end. This one keeps only the
head, deliberately: the tail of a truncated HTML document is closing markup,
the extractor needs a parseable prefix rather than two fragments spliced
together, and web_grep's line numbers have to mean something. Please do not
"improve" this into head+tail.
The cancel that looks exactly like a timeout
Cancelling a body makes Cro quit the byte stream with
X::Cro::HTTP::Client::Timeout(phase => 'body') โ the very same
exception, of the very same type, that a real body timeout raises (a real body
timeout is a cancel, internally). There is no way to tell them apart from
the exception, so a $stopped flag records that the cancel was ours. Without
it, a genuinely stalled server would be reported as a tidy truncation.
Cro's own body timeout is switched off (body => Inf) for the same
reason: two things racing to cancel the same stream makes
truncated-reason unattributable. The ledger and the deadline in
MCP::Server::Tool::Web::Budget end the body; Cro does not.
Character sets
Bodies are decoded here, never by body-text โ Cro's own encoding guess
falls back to latin-1 in a way that overwrites a successful UTF-8 decode, so
cafรฉ arrives as cafรยฉ. The order is:
the
charsetparameter of theContent-Typeheader;failing that, for HTML, a
<meta charset>in the first 1 KB (sniffed from a latin-1 decode of those bytes, which cannot throw);failing that, UTF-8 for JSON (RFC 8259 ยง8.1 leaves no other option);
and a byte-order mark beats all three, because it is evidence rather than a claim.
Whatever is chosen, the decode is attempted with fallbacks โ chosen, then
UTF-8, then latin-1, which cannot fail โ and charset on the result reports
the encoding that actually worked, not the one the server claimed.
When the body was truncated, an incomplete UTF-8 sequence at the cut is trimmed first (at most three bytes): without that, one severed character makes the whole document fail to decode as UTF-8 and fall back to latin-1, turning a clean truncation into a page of mojibake.
Binary bodies โ images, PDFs, archives โ are not decoded at all: text is
empty and bytes is intact, so the tool layer can say what it is and how big
it is. MCP::Server::Tool::Web::Extract owns that message.
robots.txt
The Fetcher holds a MCP::Server::Tool::Web::Robots policy, Off by
default, and consults it before every hop โ a redirect onto a path the
origin disallows is still a request to that path. web_fetch and depth-zero
web_grep keep the default; crawling swaps in Respect. See that module
for how the two point at each other without a dependency cycle.