Robots

NAME

MCP::Server::Tool::Web::Robots - robots.txt policy: the seam, the "do not consult it" implementation, and the RFC 9309 one

SYNOPSIS


use MCP::Server::Tool::Web::Robots;
use MCP::Server::Tool::Web::Url;

my $url = MCP::Server::Tool::Web::Url.parse('https://example.com/private/x');

# What a single-page fetch uses: robots.txt is never consulted.
my $off = MCP::Server::Tool::Web::Robots::Off.new;
say $off.allow($url, agent => 'MCP-Server-Tool-Web/0.1.0');   # True
say $off.describe;    # robots.txt is not consulted (single-page fetches)

# What a crawl uses. The one thing it cannot do for itself is fetch, so the
# fetch is handed in โ€” which is also how tests give it bodies directly.
my $respect = MCP::Server::Tool::Web::Robots::Respect.new(
    fetch-robots => -> Str $origin {
        "User-agent: *\nDisallow: /private/\nCrawl-delay: 2\n";
    },
);

say $respect.allow($url, agent => 'MCP-Server-Tool-Web/0.1.0');        # False
say $respect.disallow-rule($url, agent => 'MCP-Server-Tool-Web/0.1.0');# /private/
say $respect.crawl-delay($url, agent => 'MCP-Server-Tool-Web/0.1.0');  # 2

Wiring it to a real fetcher is a two-step, because the robots policy and the fetcher that feeds it point at each other. The closure is not called until the first crawl, so binding the fetcher after the fact is safe:


my $fetcher;
my $robots = MCP::Server::Tool::Web::Robots::Respect.new(
    fetch-robots => -> Str $origin { $fetcher.robots-text($origin) },
);
$fetcher = MCP::Server::Tool::Web::Fetcher.new(:$guard, :$transport, :$robots);

DESCRIPTION

Per the pack's ruling, web_crawl (and web_grep with a depth above zero) respects robots.txt; web_fetch and depth-zero grep do not. A single page a person explicitly asked for is not crawling, and no robots.txt has ever meant "this URL may not be read once, deliberately, by a human's agent". Traversal is a different act, and that one is asked for.

So there are two implementations of one role, and the difference between "we crawl politely" and "we fetch one page" is which object the Fetcher holds.

The role

  • allow(Url:D $url, Str:D :$agent!) โ€” may we request this URL?

  • disallow-rule(Url:D $url, Str:D :$agent!) โ€” the Disallow pattern that refused it, for the refusal message, or the Str type object.

  • crawl-delay(Url:D $url, Str:D :$agent!) โ€” seconds this origin asks to be left between requests. A crawl takes the larger of this and its own configured delay: robots.txt may slow us down, never speed us up.

  • describe() โ€” one line for notices and logs.

What Respect implements (RFC 9309)

  • Groups are User-agent lines followed by their rules; consecutive User-agent lines share one group.

  • The group whose product token matches ours wins; failing that, the * group; failing that, everything is allowed. Matching is case-insensitive, and a full user-agent string is reduced to its product token (the part before the /) before comparing.

  • Allow and Disallow patterns support * (any run of characters) and a trailing $ (end of path). The longest matching pattern decides, and Allow wins a tie โ€” so Disallow: /docs/ plus Allow: /docs/public/ reads the way its author meant.

  • An empty Disallow: is not a rule at all (it is the idiomatic way to say "everything is permitted"), and neither is an empty Allow:.

  • Crawl-delay is read, and only ever raises a caller's delay.

  • Anything else โ€” Sitemap, unknown fields, junk โ€” is ignored rather than treated as an error.

Failure is permission

A missing robots.txt, a 4xx, a timeout, a body that does not parse, or a fetch that throws all mean allowed. That is RFC 9309's own reading ("unavailable" means unrestricted), and it is the only choice that does not turn a flaky origin into a silently empty crawl. The one thing that would be worse than crawling a page we should not have is reporting "no pages found" for a site that was simply slow to answer.

The cache

One entry per origin, holding the parsed groups and an expiry. Lookups drop expired entries as they pass them, and an insert that would take the cache over cache-size evicts the least recently used entry. There are no timers and no threads: a cache that needed cleaning up would be a shutdown obligation the pack does not otherwise have.

The fetch happens outside the cache lock, so two crawls starting at once may both fetch the same robots.txt. That is one wasted request rather than a lock held across the network, which is the right trade.

E

<gt> [...], rules => #| [%(allow, pattern)], crawl-delay)>. Never throws โ€” a file it cannot make #| sense of yields the groups it could, which may be none, and none means #| "allowed". my sub parse-robots(Str:D $text --> List:D) { my @groups; my @agents; my @rules; my $delay; my Bool $in-agents = False;

MCP::Server::Tool::Web v0.1.1

web search, fetch, crawl and grep for MCP::Server

Authors

  • Matt Doughty

License

Artistic-2.0

Dependencies

MCP::Server:auth<zef:apogee>:ver<0.6.0+>Cro::HTTP:auth<zef:cro>:ver<0.8.11+>Cro::Core:auth<zef:cro>:ver<0.8.10+>IO::Socket::Async::SSL:auth<zef:raku-community-modules>:ver<0.8.2+>JSON::Fast:ver<0.19>:auth<cpan:TIMOTIMO>

Test Dependencies

Provides

  • MCP::Server::Tool::Web
  • MCP::Server::Tool::Web::Addr
  • MCP::Server::Tool::Web::Budget
  • MCP::Server::Tool::Web::Crawl
  • MCP::Server::Tool::Web::Extract
  • MCP::Server::Tool::Web::Fetcher
  • MCP::Server::Tool::Web::Guard
  • MCP::Server::Tool::Web::Provider::Brave
  • MCP::Server::Tool::Web::Robots
  • MCP::Server::Tool::Web::SearchProvider
  • MCP::Server::Tool::Web::Transport
  • MCP::Server::Tool::Web::Url
  • MCP::Server::Tool::Web::X

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite โ€” the markup and publishing tools behind this site.