Url

NAME

MCP::Server::Tool::Web::Url - strict, normalising URL parsing for the web tools

SYNOPSIS


use MCP::Server::Tool::Web::Url;

my $url = MCP::Server::Tool::Web::Url.parse('HTTP://Example.COM:80/A/b?q=1#frag');
say $url.Str;            # http://example.com/A/b?q=1  β€” host folded, default
                         # port dropped, fragment gone, path untouched
say $url.origin;         # http://example.com
say $url.host;           # example.com
say $url.port;           # 80  β€” the effective port, never undefined
say $url.target;         # /A/b?q=1  β€” what goes on the wire

# Redirects: RFC 3986 reference resolution, then the whole parse again. A
# Location header cannot smuggle in a scheme or a host shape that
# MCP::Server::Tool::Web::Url.parse would have refused.
my $next = $url.resolve('/other?x=2');
say $next.Str;                       # http://example.com/other?x=2
say $url.same-origin($next);         # True

# Same origin means scheme, host and *effective* port β€” so the default port
# spelled out is still the same origin.
say MCP::Server::Tool::Web::Url.parse('https://example.com/a')
    .same-origin(MCP::Server::Tool::Web::Url.parse('https://example.com:443/b'));  # True

# Everything below is refused, with a message that says which rule refused it.
for <
    file:///etc/passwd
    javascript:alert(1)
    http://user:[email protected]/
    http://2130706433/
    http://0x7f000001/
    http://127.1/
    http://017700000001/
    http://example.com:99999/
    http://exΓ€mple.com/
> -> $bad {
    my $why = (try MCP::Server::Tool::Web::Url.parse($bad)) // $!.message;
    say $why;
}

DESCRIPTION

Cro::Uri does RFC 3986 β€” it parses the grammar and resolves references, and it is right to be liberal about it, because RFC 3986 describes every URI anyone has ever written. This class is the opposite: it accepts the small, boring subset of URLs a web tool should ever act on, and refuses everything else by name, before a socket is opened.

Everything parse returns is normalised, so two spellings of one page compare equal:

  • scheme and host lowercased;

  • a single trailing root dot removed from the host (example.com. and example.com are the same name in DNS, and treating them as two origins would split a crawl in half);

  • an IPv6 literal rewritten in its canonical RFC 5952 form, so [0:0:0:0:0:0:0:1] and [::1] are one origin;

  • the default port for the scheme dropped from origin and Str, but always available as port;

  • the fragment dropped β€” it is never sent to a server, and keeping it would make two requests for one page look like two pages.

The path is not touched beyond that: /a/../b is requested as written, because the server, not this class, decides what a dot segment means to it. Reference resolution is the one place dot segments are removed, because RFC 3986 Β§5 says to β€” so parse('http://x/a/../b') keeps the path and parse('http://x/').resolve('/a/../b') does not. Compare fetched pages by their final URL if you need one canonical spelling per page.

What is refused, and why

Each of these throws X::MCP::Server::Tool::Web::BadUrl, whose message names the rule. No allow-list rescues any of them: these are shape rules, decided before anything is resolved, and allow-private / allow-hosts / allow-cidrs / allow-loopback are about addresses, not shapes.

  • Non-absolute, or a scheme other than http/https. file: would read the disk, javascript: and data: mean nothing to a fetcher, and a relative URL has no host to check.

  • Credentials in the URL. http://user:pw@host/ exists mostly to make a host look like a path to a human ("http://[email protected]"). Refused outright rather than stripped, so nobody has to wonder which half was used.

  • Numerically obfuscated hosts. 2130706433, 0x7f000001, 017700000001 and 127.1 are all 127.0.0.1 β€” legal for getaddrinfo, unreadable for a human reviewing a log. The connector floor would catch them anyway (it checks the address of the socket it actually holds), but a URL nobody can read is refused as a matter of clarity, not just safety.

  • Percent-encoding in the host. http://127.0.0.%31/ decodes to 127.0.0.1. There is no legitimate use.

  • Non-ASCII hosts. Cro percent-encodes them rather than punycoding, and a percent-encoded host does not resolve; the refusal asks for the A-label (xn--…) form instead, which is what DNS wants anyway.

  • Ports outside 1..65535, IPv6 zone identifiers, empty or over-long host labels, and characters that are not valid in a host name.

Resolving redirects

resolve is Cro::Uri.add β€” the RFC 3986 Β§5 reference resolution algorithm, which is fiddly enough to be worth not reimplementing β€” followed by the whole of parse again. That second half is the point: a Location: file:///etc/passwd or Location: http://10.0.0.1/ is refused on the hop that produced it, with the same message it would have got as a starting URL.

MCP::Server::Tool::Web v0.1.1

web search, fetch, crawl and grep for MCP::Server

Authors

  • Matt Doughty

License

Artistic-2.0

Dependencies

MCP::Server:auth<zef:apogee>:ver<0.6.0+>Cro::HTTP:auth<zef:cro>:ver<0.8.11+>Cro::Core:auth<zef:cro>:ver<0.8.10+>IO::Socket::Async::SSL:auth<zef:raku-community-modules>:ver<0.8.2+>JSON::Fast:ver<0.19>:auth<cpan:TIMOTIMO>

Test Dependencies

Provides

  • MCP::Server::Tool::Web
  • MCP::Server::Tool::Web::Addr
  • MCP::Server::Tool::Web::Budget
  • MCP::Server::Tool::Web::Crawl
  • MCP::Server::Tool::Web::Extract
  • MCP::Server::Tool::Web::Fetcher
  • MCP::Server::Tool::Web::Guard
  • MCP::Server::Tool::Web::Provider::Brave
  • MCP::Server::Tool::Web::Robots
  • MCP::Server::Tool::Web::SearchProvider
  • MCP::Server::Tool::Web::Transport
  • MCP::Server::Tool::Web::Url
  • MCP::Server::Tool::Web::X

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite β€” the markup and publishing tools behind this site.