Web
NAME
MCP::Server::Tool::Web - web search, fetch, crawl and grep for MCP::Server
SYNOPSIS
use MCP::Server;
use MCP::Server::Tool::Web;
my $server = MCP::Server.new(name => 'research');
$server.plug: MCP::Server::Tool::Web.new; # web_search, web_fetch,
# web_crawl, web_grep
$server.run;
Configured rather than constructed ā the shape a host's JSON config takes:
my $server = MCP::Server.new(
name => 'research',
tools => [
'Web' => {
provider => 'brave',
'api-key-env' => 'BRAVE_API_KEY',
'max-bytes' => 4 * 1024 * 1024,
'total-timeout' => 90,
'crawl-delay' => 0.5,
'allow-hosts' => <wiki.corp docs.corp>, # exact names only
},
],
);
What the tools answer with:
# web_search ā JSON, so the model can pick a URL out of it
# {"count":10,"provider":"brave","query":"raku grammars",
# "results":[{"snippet":"ā¦","title":"Grammars","url":"https://docs.raku.org/ā¦"}]}
# web_fetch ā plain text, because a page is text
# # Grammars
# Retrieved from https://docs.raku.org/language/grammars
#
# Grammars are a powerful tool used to destructure text ā¦
# web_crawl ā an index, not the pages
# Crawled 3 pages from https://docs.example.com/ (depth 1, same origin).
#
# https://docs.example.com/ 200 4213 bytes Documentation
# https://docs.example.com/install 200 2288 bytes Installing
# web_grep ā only the lines that matched, exactly as fs_grep reports them
# https://docs.example.com/install:42: zef install Foo::Bar
#
# [stopped at 50 matching lines; raise max-results to see more]
DESCRIPTION
Four tools, one prefix (web), and one rule that runs under all of them: a
URL the model supplied is fetched only after the address of the socket that was
actually opened has been checked (see MCP::Server::Tool::Web::Guard and
MCP::Server::Tool::Web::Transport), so a hostname that resolves onto this
machine's own network is refused whatever it is spelled like.
The four tools
web_search(query, count?) ā the search engine, behind a seam. Answers JSON so the model can pick a URL out of it. Zero results is an empty list and a notice, not an error.
web_fetch(url) ā one page, as text: title, the final URL when a redirect moved it, then the page with its markup, scripts and navigation removed. Exactly one parameter, deliberately: a windowing tool would invite the model to page through a document it should have grepped.
web_crawl(url, depth?, max-pages?) ā an index of a site's pages, never their bodies. See below.
web_grep(pattern, url, regex?, context?, max-results?, depth?, max-pages?) ā
fs_grepfor the web, output format and notices included. The cheap tool, and the one to reach for first.
Why web_crawl returns an index
Concatenating twenty pages into one tool result is a way to spend a context
window without answering anything: what the model actually receives is page one,
half of page two, and a truncation marker. So a crawl reports what is there ā
URL, status, size, title, one line each ā and the corpus stays reachable
through web_grep (with a depth) and web_fetch.
robots.txt
Per the pack's ruling: web_crawl, and web_grep with a depth above
zero, respect robots.txt; web_fetch and depth-zero web_grep never consult
it. A page a person asked for by name is not crawling, and no robots.txt has
ever meant "this URL may not be read once, deliberately, on a human's behalf" ā
traversal is a different act, and that one is asked for. Setting
respect-robots to False turns it off for the traversing tools too.
Mechanically that is two fetchers over one transport (so they share the vetted,
already-established connections): one holding
MCP::Server::Tool::Web::Robots::Off, one holding Respect. The decision
is per call rather than per pack, which is why it cannot simply be a setting on
one fetcher.
The search provider is not guard-subject
The endpoint web_search calls is the operator's, named in configuration ā
not the model's. A SearXNG on the operator's own LAN is a legitimate thing to
point this at, so the search provider talks to its API directly rather than
through the SSRF guard. The guard exists for URLs an attacker (or a confused
model) chose; a config key is neither.
The API key is resolved at call time, so a pack with no BRAVE_API_KEY in
its environment still serves web_fetch, web_crawl and web_grep
perfectly well, and web_search answers with a message naming the variable
and the two config keys that would fix it.
Configuration
Every public attribute below is a JSON config key (from-config validates
against exactly this set, so a typo names the valid keys instead of doing
nothing quietly).
provider, api-key-env, api-key, search-country, search-lang, safesearch, timeout ā the search provider.
max-bytes, total-timeout, connect-timeout, headers-timeout ā what one tool call may spend.
max-bytesandtotal-timeoutare the whole call's: twenty pages of a crawl share them, so a crawl cannot cost twenty times a fetch.crawl-delay, respect-robots, user-agent ā how the traversing tools behave.
allow-private, allow-loopback, allow-hosts, allow-cidrs, proxy ā the SSRF floor's escape hatches, all off by default. See MCP::Server::Tool::Web::Guard:
allow-hostsis exact matching, because suffix matching is where allow-list CVEs come from.
Two more settings exist and are deliberately not config keys, because a
Callable and a live object cannot come out of JSON ā they are is built
private attributes, so .new can pass them and from-config cannot:
search-provider ā anything doing MCP::Server::Tool::Web::SearchProvider, used instead of building the named
provider. This is how a test dispatchesweb_searchwithout a network, and how a host wires up an engine this pack has never heard of.fetcher ā a prepared MCP::Server::Tool::Web::Fetcher. Given one, the pack uses it for single-page work and clones it (with the robots policy swapped in) for traversal, so the transport ā and its cache of vetted connections ā is shared.
connect-tcp ā the one closure that touches the network,
($host, $port --> Promise[IO::Socket::Async]), handed straight to MCP::Server::Tool::Web::Transport. Test rigs replace it so that hostnames stay real end to end while the sockets go to a local origin; replacing it does not weaken the floor, since whatever socket comes back still has its peer address checked.
AUTHOR
Matt Doughty
COPYRIGHT AND LICENSE
Copyright 2026 Matt Doughty
This library is free software; you can redistribute it and/or modify it under the Artistic License 2.0.