Crawl

NAME

MCP::Server::Tool::Web::Crawl - one breadth-first walk of a site, with a callback per page

SYNOPSIS


use MCP::Server::Tool::Web::Crawl;
use MCP::Server::Tool::Web::Fetcher;
use MCP::Server::Tool::Web::Guard;

my $fetcher = MCP::Server::Tool::Web::Fetcher.new(
    guard => MCP::Server::Tool::Web::Guard.new,
);

my $crawl = MCP::Server::Tool::Web::Crawl.new(
    fetcher   => $fetcher,
    max-depth => 1,       # 0 = the starting page and nothing else
    max-pages => 20,      # counting the starting page
    delay     => 0.25,    # seconds between requests, single-threaded
);

my @index;
my $report = $crawl.run('https://docs.example.com/', on-page => -> $page {
    @index.push("{$page.final-url} {$page.status} {$page.byte-count} bytes");
    True;             # False would stop the crawl here
});

say $report.fetched;          # 7
say $report.origin;           # https://docs.example.com
say $report.stopped;          # max-pages | deadline | max-bytes | callback | Str
say $report.failures;         # [{url => …, reason => …}, …]
say $report.robots-skipped;   # [https://docs.example.com/private/, …]

Stopping early is the callback's job — this is how web_grep spends a global result budget without the cap logic being written twice:


my Int $matched = 0;
my $report = $crawl.run($seed, on-page => -> $page {
    $matched += count-matches($page);
    $matched < 50;        # once 50 lines have matched, nothing more is fetched
});
say $report.stopped;      # callback — and the pages after it were never requested

DESCRIPTION

One traversal, two consumers: web_crawl builds an index out of the callback and web_grep greps each page as it arrives and stops the walk when it has found enough. Writing the cap, dedup and same-origin logic twice is how the two tools would drift apart, so they share this class and differ only in what their callback does.

What the walk does

  • Breadth first. Every page at depth n is fetched before any page at depth n+1, and within a level the order is the order the pages themselves list their links in — the site's own idea of what matters most, and deterministic for a fixed site.

  • Same origin only. Origin means scheme, host and effective port, and it is taken from the seed's final URL: a seed that redirects to www.example.com crawls www.example.com, which is what the redirect was for. Links elsewhere are not followed and are not failures.

  • Every page once. Both the URL a page was requested by and the URL it finally came from are remembered, so a page reachable by two paths — or by a redirect from one of them — is fetched once.

  • Links come from the raw HTML, resolved against the document's own <base href> or the final URL, http(s) only, fragments dropped, and rel="nofollow" honoured. Only successful HTML responses are read for links: the link list of a 404 page is the site's error template.

  • max-pages counts fetched pages, including the seed.

  • One budget for the whole walk. Time and bytes are shared, so a crawl cannot cost twenty times a single fetch. Running out of either stops the walk with a stopped reason rather than an exception.

Failures are results, except the first one

A page that cannot be fetched — refused address, dead connection, redirect loop — is recorded in failures and the walk carries on: one broken link in a documentation site is not a reason to report nothing. The seed is the exception: if the page the caller actually named cannot be fetched, the exception travels, because "0 pages, 1 failure" is a worse answer than the refusal that explains why.

Pages refused by robots.txt are counted separately in robots-skipped: they are a policy decision rather than a fault, and the tool layer says so with a notice naming the setting that would change it.

Politeness

The walk is single-threaded and sleeps delay seconds between requests. When the fetcher's robots.txt policy asks for a longer Crawl-delay for a URL, that one is used instead — robots.txt may slow the crawl down, never speed it up. Tests set delay => 0.

MCP::Server::Tool::Web v0.1.1

web search, fetch, crawl and grep for MCP::Server

Authors

  • Matt Doughty

License

Artistic-2.0

Dependencies

MCP::Server:auth<zef:apogee>:ver<0.6.0+>Cro::HTTP:auth<zef:cro>:ver<0.8.11+>Cro::Core:auth<zef:cro>:ver<0.8.10+>IO::Socket::Async::SSL:auth<zef:raku-community-modules>:ver<0.8.2+>JSON::Fast:ver<0.19>:auth<cpan:TIMOTIMO>

Test Dependencies

Provides

  • MCP::Server::Tool::Web
  • MCP::Server::Tool::Web::Addr
  • MCP::Server::Tool::Web::Budget
  • MCP::Server::Tool::Web::Crawl
  • MCP::Server::Tool::Web::Extract
  • MCP::Server::Tool::Web::Fetcher
  • MCP::Server::Tool::Web::Guard
  • MCP::Server::Tool::Web::Provider::Brave
  • MCP::Server::Tool::Web::Robots
  • MCP::Server::Tool::Web::SearchProvider
  • MCP::Server::Tool::Web::Transport
  • MCP::Server::Tool::Web::Url
  • MCP::Server::Tool::Web::X

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite — the markup and publishing tools behind this site.