Web

NAME

MCP::Server::Tool::Web - web search, fetch, crawl and grep for MCP::Server

SYNOPSIS


use MCP::Server;
use MCP::Server::Tool::Web;

my $server = MCP::Server.new(name => 'research');
$server.plug: MCP::Server::Tool::Web.new;      # web_search, web_fetch,
                                               # web_crawl, web_grep
$server.run;

Configured rather than constructed — the shape a host's JSON config takes:


my $server = MCP::Server.new(
    name  => 'research',
    tools => [
        'Web' => {
            provider       => 'brave',
            'api-key-env'  => 'BRAVE_API_KEY',
            'max-bytes'    => 4 * 1024 * 1024,
            'total-timeout' => 90,
            'crawl-delay'  => 0.5,
            'allow-hosts'  => <wiki.corp docs.corp>,   # exact names only
        },
    ],
);

What the tools answer with:


# web_search — JSON, so the model can pick a URL out of it
# {"count":10,"provider":"brave","query":"raku grammars",
#  "results":[{"snippet":"…","title":"Grammars","url":"https://docs.raku.org/…"}]}

# web_fetch — plain text, because a page is text
# # Grammars
# Retrieved from https://docs.raku.org/language/grammars
#
# Grammars are a powerful tool used to destructure text …

# web_crawl — an index, not the pages
# Crawled 3 pages from https://docs.example.com/ (depth 1, same origin).
#
# https://docs.example.com/  200  4213 bytes  Documentation
# https://docs.example.com/install  200  2288 bytes  Installing

# web_grep — only the lines that matched, exactly as fs_grep reports them
# https://docs.example.com/install:42:  zef install Foo::Bar
#
# [stopped at 50 matching lines; raise max-results to see more]

DESCRIPTION

Four tools, one prefix (web), and one rule that runs under all of them: a URL the model supplied is fetched only after the address of the socket that was actually opened has been checked (see MCP::Server::Tool::Web::Guard and MCP::Server::Tool::Web::Transport), so a hostname that resolves onto this machine's own network is refused whatever it is spelled like.

The four tools

  • web_search(query, count?) — the search engine, behind a seam. Answers JSON so the model can pick a URL out of it. Zero results is an empty list and a notice, not an error.

  • web_fetch(url) — one page, as text: title, the final URL when a redirect moved it, then the page with its markup, scripts and navigation removed. Exactly one parameter, deliberately: a windowing tool would invite the model to page through a document it should have grepped.

  • web_crawl(url, depth?, max-pages?) — an index of a site's pages, never their bodies. See below.

  • web_grep(pattern, url, regex?, context?, max-results?, depth?, max-pages?) — fs_grep for the web, output format and notices included. The cheap tool, and the one to reach for first.

Why web_crawl returns an index

Concatenating twenty pages into one tool result is a way to spend a context window without answering anything: what the model actually receives is page one, half of page two, and a truncation marker. So a crawl reports what is there — URL, status, size, title, one line each — and the corpus stays reachable through web_grep (with a depth) and web_fetch.

robots.txt

Per the pack's ruling: web_crawl, and web_grep with a depth above zero, respect robots.txt; web_fetch and depth-zero web_grep never consult it. A page a person asked for by name is not crawling, and no robots.txt has ever meant "this URL may not be read once, deliberately, on a human's behalf" — traversal is a different act, and that one is asked for. Setting respect-robots to False turns it off for the traversing tools too.

Mechanically that is two fetchers over one transport (so they share the vetted, already-established connections): one holding MCP::Server::Tool::Web::Robots::Off, one holding Respect. The decision is per call rather than per pack, which is why it cannot simply be a setting on one fetcher.

The search provider is not guard-subject

The endpoint web_search calls is the operator's, named in configuration — not the model's. A SearXNG on the operator's own LAN is a legitimate thing to point this at, so the search provider talks to its API directly rather than through the SSRF guard. The guard exists for URLs an attacker (or a confused model) chose; a config key is neither.

The API key is resolved at call time, so a pack with no BRAVE_API_KEY in its environment still serves web_fetch, web_crawl and web_grep perfectly well, and web_search answers with a message naming the variable and the two config keys that would fix it.

Configuration

Every public attribute below is a JSON config key (from-config validates against exactly this set, so a typo names the valid keys instead of doing nothing quietly).

  • provider, api-key-env, api-key, search-country, search-lang, safesearch, timeout — the search provider.

  • max-bytes, total-timeout, connect-timeout, headers-timeout — what one tool call may spend. max-bytes and total-timeout are the whole call's: twenty pages of a crawl share them, so a crawl cannot cost twenty times a fetch.

  • crawl-delay, respect-robots, user-agent — how the traversing tools behave.

  • allow-private, allow-loopback, allow-hosts, allow-cidrs, proxy — the SSRF floor's escape hatches, all off by default. See MCP::Server::Tool::Web::Guard: allow-hosts is exact matching, because suffix matching is where allow-list CVEs come from.

Two more settings exist and are deliberately not config keys, because a Callable and a live object cannot come out of JSON — they are is built private attributes, so .new can pass them and from-config cannot:

  • search-provider — anything doing MCP::Server::Tool::Web::SearchProvider, used instead of building the named provider. This is how a test dispatches web_search without a network, and how a host wires up an engine this pack has never heard of.

  • fetcher — a prepared MCP::Server::Tool::Web::Fetcher. Given one, the pack uses it for single-page work and clones it (with the robots policy swapped in) for traversal, so the transport — and its cache of vetted connections — is shared.

  • connect-tcp — the one closure that touches the network, ($host, $port --> Promise[IO::Socket::Async]), handed straight to MCP::Server::Tool::Web::Transport. Test rigs replace it so that hostnames stay real end to end while the sockets go to a local origin; replacing it does not weaken the floor, since whatever socket comes back still has its peer address checked.

AUTHOR

Matt Doughty

COPYRIGHT AND LICENSE

Copyright 2026 Matt Doughty

This library is free software; you can redistribute it and/or modify it under the Artistic License 2.0.

MCP::Server::Tool::Web v0.1.1

web search, fetch, crawl and grep for MCP::Server

Authors

  • Matt Doughty

License

Artistic-2.0

Dependencies

MCP::Server:auth<zef:apogee>:ver<0.6.0+>Cro::HTTP:auth<zef:cro>:ver<0.8.11+>Cro::Core:auth<zef:cro>:ver<0.8.10+>IO::Socket::Async::SSL:auth<zef:raku-community-modules>:ver<0.8.2+>JSON::Fast:ver<0.19>:auth<cpan:TIMOTIMO>

Test Dependencies

Provides

  • MCP::Server::Tool::Web
  • MCP::Server::Tool::Web::Addr
  • MCP::Server::Tool::Web::Budget
  • MCP::Server::Tool::Web::Crawl
  • MCP::Server::Tool::Web::Extract
  • MCP::Server::Tool::Web::Fetcher
  • MCP::Server::Tool::Web::Guard
  • MCP::Server::Tool::Web::Provider::Brave
  • MCP::Server::Tool::Web::Robots
  • MCP::Server::Tool::Web::SearchProvider
  • MCP::Server::Tool::Web::Transport
  • MCP::Server::Tool::Web::Url
  • MCP::Server::Tool::Web::X

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite — the markup and publishing tools behind this site.