Robots
NAME
MCP::Server::Tool::Web::Robots - robots.txt policy: the seam, the "do not consult it" implementation, and the RFC 9309 one
SYNOPSIS
use MCP::Server::Tool::Web::Robots;
use MCP::Server::Tool::Web::Url;
my $url = MCP::Server::Tool::Web::Url.parse('https://example.com/private/x');
# What a single-page fetch uses: robots.txt is never consulted.
my $off = MCP::Server::Tool::Web::Robots::Off.new;
say $off.allow($url, agent => 'MCP-Server-Tool-Web/0.1.0'); # True
say $off.describe; # robots.txt is not consulted (single-page fetches)
# What a crawl uses. The one thing it cannot do for itself is fetch, so the
# fetch is handed in โ which is also how tests give it bodies directly.
my $respect = MCP::Server::Tool::Web::Robots::Respect.new(
fetch-robots => -> Str $origin {
"User-agent: *\nDisallow: /private/\nCrawl-delay: 2\n";
},
);
say $respect.allow($url, agent => 'MCP-Server-Tool-Web/0.1.0'); # False
say $respect.disallow-rule($url, agent => 'MCP-Server-Tool-Web/0.1.0');# /private/
say $respect.crawl-delay($url, agent => 'MCP-Server-Tool-Web/0.1.0'); # 2
Wiring it to a real fetcher is a two-step, because the robots policy and the fetcher that feeds it point at each other. The closure is not called until the first crawl, so binding the fetcher after the fact is safe:
my $fetcher;
my $robots = MCP::Server::Tool::Web::Robots::Respect.new(
fetch-robots => -> Str $origin { $fetcher.robots-text($origin) },
);
$fetcher = MCP::Server::Tool::Web::Fetcher.new(:$guard, :$transport, :$robots);
DESCRIPTION
Per the pack's ruling, web_crawl (and web_grep with a depth above zero)
respects robots.txt; web_fetch and depth-zero grep do not. A single page a
person explicitly asked for is not crawling, and no robots.txt has ever meant
"this URL may not be read once, deliberately, by a human's agent". Traversal
is a different act, and that one is asked for.
So there are two implementations of one role, and the difference between "we crawl politely" and "we fetch one page" is which object the Fetcher holds.
The role
allow(Url:D $url, Str:D :$agent!)โ may we request this URL?disallow-rule(Url:D $url, Str:D :$agent!)โ theDisallowpattern that refused it, for the refusal message, or theStrtype object.crawl-delay(Url:D $url, Str:D :$agent!)โ seconds this origin asks to be left between requests. A crawl takes the larger of this and its own configured delay: robots.txt may slow us down, never speed us up.describe()โ one line for notices and logs.
What Respect implements (RFC 9309)
Groups are
User-agentlines followed by their rules; consecutiveUser-agentlines share one group.The group whose product token matches ours wins; failing that, the
*group; failing that, everything is allowed. Matching is case-insensitive, and a full user-agent string is reduced to its product token (the part before the/) before comparing.AllowandDisallowpatterns support*(any run of characters) and a trailing$(end of path). The longest matching pattern decides, andAllowwins a tie โ soDisallow: /docs/plusAllow: /docs/public/reads the way its author meant.An empty
Disallow:is not a rule at all (it is the idiomatic way to say "everything is permitted"), and neither is an emptyAllow:.Crawl-delayis read, and only ever raises a caller's delay.Anything else โ
Sitemap, unknown fields, junk โ is ignored rather than treated as an error.
Failure is permission
A missing robots.txt, a 4xx, a timeout, a body that does not parse, or a fetch that throws all mean allowed. That is RFC 9309's own reading ("unavailable" means unrestricted), and it is the only choice that does not turn a flaky origin into a silently empty crawl. The one thing that would be worse than crawling a page we should not have is reporting "no pages found" for a site that was simply slow to answer.
The cache
One entry per origin, holding the parsed groups and an expiry. Lookups drop
expired entries as they pass them, and an insert that would take the cache
over cache-size evicts the least recently used entry. There are no timers
and no threads: a cache that needed cleaning up would be a shutdown obligation
the pack does not otherwise have.
The fetch happens outside the cache lock, so two crawls starting at once may both fetch the same robots.txt. That is one wasted request rather than a lock held across the network, which is the right trade.
E
<gt> [...], rules => #| [%(allow, pattern)], crawl-delay)>. Never throws โ a file it cannot make #| sense of yields the groups it could, which may be none, and none means #| "allowed". my sub parse-robots(Str:D $text --> List:D) { my @groups; my @agents; my @rules; my $delay; my Bool $in-agents = False;