Fetcher

NAME

MCP::Server::Tool::Web::Fetcher - one guarded HTTP GET, redirects, budgets and character sets included

SYNOPSIS


use MCP::Server::Tool::Web::Fetcher;
use MCP::Server::Tool::Web::Guard;

my $fetcher = MCP::Server::Tool::Web::Fetcher.new(
    guard     => MCP::Server::Tool::Web::Guard.new,
    max-bytes => 2 * 1024 * 1024,
);

my $page = $fetcher.fetch('https://docs.example.com/guide');

say $page.status;             # 200
say $page.final-url;          # https://docs.example.com/guide/  (after a 301)
say $page.redirect-chain;     # [https://docs.example.com/guide https://โ€ฆ/guide/]
say $page.content-type;       # text/html
say $page.charset;            # utf-8      โ€” the encoding that actually worked
say $page.byte-count;         # 48219
say $page.text.lines.head;    # <!DOCTYPE html>
say $page.elapsed;            # 0.31

A 404 is a result, not an exception. Only the request never happening is:


my $missing = $fetcher.fetch('https://example.com/nope');
say $missing.status;          # 404
say $missing.text.chars;      # the error page is still the body

{
    $fetcher.fetch('http://169.254.169.254/latest/meta-data/');
    CATCH {
        when X::MCP::Server::Tool::Web::Blocked   { say .message }  # metadata endpoint
        when X::MCP::Server::Tool::Web::BadUrl    { say .message }
        when X::MCP::Server::Tool::Web::Deadline  { say .message }
        when X::MCP::Server::Tool::Web::Transport { say .message }
    }
}

One budget, shared by every page of a crawl:


my $budget = $fetcher.new-budget;         # this Fetcher's caps, one call's worth
for @urls -> $url {
    last if $budget.expired;
    my $page = $fetcher.fetch($url, :$budget);
    say "$url: {$page.byte-count} bytes"
        ~ ($page.truncated ?? " (cut: {$page.truncated-reason})" !! '');
}

DESCRIPTION

Everything between "here is a URL" and "here is the text of that page", with the pack's rules applied at every step: the guard vets the URL and, through MCP::Server::Tool::Web::Transport, the address behind it; the budget bounds the time and the bytes; redirects are followed by hand so that every hop is checked again; and the body is decoded with the encoding the bytes are actually in rather than the one the fetcher would prefer.

Redirects are followed here, not by Cro

Cro::HTTP::Client can follow redirects itself, and this class asks it not to (follow => False). Every hop goes back through Guard.check-url and, on connect, through the address check: a Location: file:///etc/passwd, a Location: http://10.0.0.1/ or a redirect that quietly drops from https to http is refused on the hop that produced it, naming the rule. Cro's own loop would have skipped all of that from hop two onwards.

The other reason is bookkeeping the tool surface needs: the chain, in order, and the final URL โ€” which is what a crawl's same-origin test and a relative link's base have to be resolved against.

Loops are detected on (origin, path, query), so a site that alternates between two URLs is refused as a loop rather than as "too many redirects": the first is a fact about the site that no setting will fix, the second is a budget.

The body is read against a ledger

The response body is streamed and counted. When the budget's byte allowance runs out mid-body the download is cancelled and what arrived is kept; when the deadline passes mid-body, likewise. Either way the result carries truncated => True and a truncated-reason of 'max-bytes' or 'deadline', and a body that simply ended is never marked truncated, however close to a limit it came.

Why head-only, and not head-and-tail

The shell tool keeps the head and the tail of a long output, because the interesting part of a build log is usually at the end. This one keeps only the head, deliberately: the tail of a truncated HTML document is closing markup, the extractor needs a parseable prefix rather than two fragments spliced together, and web_grep's line numbers have to mean something. Please do not "improve" this into head+tail.

The cancel that looks exactly like a timeout

Cancelling a body makes Cro quit the byte stream with X::Cro::HTTP::Client::Timeout(phase => 'body') โ€” the very same exception, of the very same type, that a real body timeout raises (a real body timeout is a cancel, internally). There is no way to tell them apart from the exception, so a $stopped flag records that the cancel was ours. Without it, a genuinely stalled server would be reported as a tidy truncation.

Cro's own body timeout is switched off (body => Inf) for the same reason: two things racing to cancel the same stream makes truncated-reason unattributable. The ledger and the deadline in MCP::Server::Tool::Web::Budget end the body; Cro does not.

Character sets

Bodies are decoded here, never by body-text โ€” Cro's own encoding guess falls back to latin-1 in a way that overwrites a successful UTF-8 decode, so cafรฉ arrives as cafรƒยฉ. The order is:

  • the charset parameter of the Content-Type header;

  • failing that, for HTML, a <meta charset> in the first 1 KB (sniffed from a latin-1 decode of those bytes, which cannot throw);

  • failing that, UTF-8 for JSON (RFC 8259 ยง8.1 leaves no other option);

  • and a byte-order mark beats all three, because it is evidence rather than a claim.

Whatever is chosen, the decode is attempted with fallbacks โ€” chosen, then UTF-8, then latin-1, which cannot fail โ€” and charset on the result reports the encoding that actually worked, not the one the server claimed.

When the body was truncated, an incomplete UTF-8 sequence at the cut is trimmed first (at most three bytes): without that, one severed character makes the whole document fail to decode as UTF-8 and fall back to latin-1, turning a clean truncation into a page of mojibake.

Binary bodies โ€” images, PDFs, archives โ€” are not decoded at all: text is empty and bytes is intact, so the tool layer can say what it is and how big it is. MCP::Server::Tool::Web::Extract owns that message.

robots.txt

The Fetcher holds a MCP::Server::Tool::Web::Robots policy, Off by default, and consults it before every hop โ€” a redirect onto a path the origin disallows is still a request to that path. web_fetch and depth-zero web_grep keep the default; crawling swaps in Respect. See that module for how the two point at each other without a dependency cycle.

MCP::Server::Tool::Web v0.1.1

web search, fetch, crawl and grep for MCP::Server

Authors

  • Matt Doughty

License

Artistic-2.0

Dependencies

MCP::Server:auth<zef:apogee>:ver<0.6.0+>Cro::HTTP:auth<zef:cro>:ver<0.8.11+>Cro::Core:auth<zef:cro>:ver<0.8.10+>IO::Socket::Async::SSL:auth<zef:raku-community-modules>:ver<0.8.2+>JSON::Fast:ver<0.19>:auth<cpan:TIMOTIMO>

Test Dependencies

Provides

  • MCP::Server::Tool::Web
  • MCP::Server::Tool::Web::Addr
  • MCP::Server::Tool::Web::Budget
  • MCP::Server::Tool::Web::Crawl
  • MCP::Server::Tool::Web::Extract
  • MCP::Server::Tool::Web::Fetcher
  • MCP::Server::Tool::Web::Guard
  • MCP::Server::Tool::Web::Provider::Brave
  • MCP::Server::Tool::Web::Robots
  • MCP::Server::Tool::Web::SearchProvider
  • MCP::Server::Tool::Web::Transport
  • MCP::Server::Tool::Web::Url
  • MCP::Server::Tool::Web::X

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite โ€” the markup and publishing tools behind this site.