Crawl
NAME
MCP::Server::Tool::Web::Crawl - one breadth-first walk of a site, with a callback per page
SYNOPSIS
use MCP::Server::Tool::Web::Crawl;
use MCP::Server::Tool::Web::Fetcher;
use MCP::Server::Tool::Web::Guard;
my $fetcher = MCP::Server::Tool::Web::Fetcher.new(
guard => MCP::Server::Tool::Web::Guard.new,
);
my $crawl = MCP::Server::Tool::Web::Crawl.new(
fetcher => $fetcher,
max-depth => 1, # 0 = the starting page and nothing else
max-pages => 20, # counting the starting page
delay => 0.25, # seconds between requests, single-threaded
);
my @index;
my $report = $crawl.run('https://docs.example.com/', on-page => -> $page {
@index.push("{$page.final-url} {$page.status} {$page.byte-count} bytes");
True; # False would stop the crawl here
});
say $report.fetched; # 7
say $report.origin; # https://docs.example.com
say $report.stopped; # max-pages | deadline | max-bytes | callback | Str
say $report.failures; # [{url => ā¦, reason => ā¦}, ā¦]
say $report.robots-skipped; # [https://docs.example.com/private/, ā¦]
Stopping early is the callback's job ā this is how web_grep spends a global
result budget without the cap logic being written twice:
my Int $matched = 0;
my $report = $crawl.run($seed, on-page => -> $page {
$matched += count-matches($page);
$matched < 50; # once 50 lines have matched, nothing more is fetched
});
say $report.stopped; # callback ā and the pages after it were never requested
DESCRIPTION
One traversal, two consumers: web_crawl builds an index out of the callback
and web_grep greps each page as it arrives and stops the walk when it has
found enough. Writing the cap, dedup and same-origin logic twice is how the two
tools would drift apart, so they share this class and differ only in what their
callback does.
What the walk does
Breadth first. Every page at depth n is fetched before any page at depth n+1, and within a level the order is the order the pages themselves list their links in ā the site's own idea of what matters most, and deterministic for a fixed site.
Same origin only. Origin means scheme, host and effective port, and it is taken from the seed's final URL: a seed that redirects to
www.example.comcrawlswww.example.com, which is what the redirect was for. Links elsewhere are not followed and are not failures.Every page once. Both the URL a page was requested by and the URL it finally came from are remembered, so a page reachable by two paths ā or by a redirect from one of them ā is fetched once.
Links come from the raw HTML, resolved against the document's own
<base href>or the final URL, http(s) only, fragments dropped, andrel="nofollow"honoured. Only successful HTML responses are read for links: the link list of a 404 page is the site's error template.max-pages counts fetched pages, including the seed.
One budget for the whole walk. Time and bytes are shared, so a crawl cannot cost twenty times a single fetch. Running out of either stops the walk with a
stoppedreason rather than an exception.
Failures are results, except the first one
A page that cannot be fetched ā refused address, dead connection, redirect
loop ā is recorded in failures and the walk carries on: one broken link in a
documentation site is not a reason to report nothing. The seed is the
exception: if the page the caller actually named cannot be fetched, the
exception travels, because "0 pages, 1 failure" is a worse answer than the
refusal that explains why.
Pages refused by robots.txt are counted separately in robots-skipped: they
are a policy decision rather than a fault, and the tool layer says so with a
notice naming the setting that would change it.
Politeness
The walk is single-threaded and sleeps delay seconds between requests. When
the fetcher's robots.txt policy asks for a longer Crawl-delay for a URL,
that one is used instead ā robots.txt may slow the crawl down, never speed it
up. Tests set delay => 0.