Url
NAME
MCP::Server::Tool::Web::Url - strict, normalising URL parsing for the web tools
SYNOPSIS
use MCP::Server::Tool::Web::Url;
my $url = MCP::Server::Tool::Web::Url.parse('HTTP://Example.COM:80/A/b?q=1#frag');
say $url.Str; # http://example.com/A/b?q=1 β host folded, default
# port dropped, fragment gone, path untouched
say $url.origin; # http://example.com
say $url.host; # example.com
say $url.port; # 80 β the effective port, never undefined
say $url.target; # /A/b?q=1 β what goes on the wire
# Redirects: RFC 3986 reference resolution, then the whole parse again. A
# Location header cannot smuggle in a scheme or a host shape that
# MCP::Server::Tool::Web::Url.parse would have refused.
my $next = $url.resolve('/other?x=2');
say $next.Str; # http://example.com/other?x=2
say $url.same-origin($next); # True
# Same origin means scheme, host and *effective* port β so the default port
# spelled out is still the same origin.
say MCP::Server::Tool::Web::Url.parse('https://example.com/a')
.same-origin(MCP::Server::Tool::Web::Url.parse('https://example.com:443/b')); # True
# Everything below is refused, with a message that says which rule refused it.
for <
file:///etc/passwd
javascript:alert(1)
http://user:[email protected]/
http://2130706433/
http://0x7f000001/
http://127.1/
http://017700000001/
http://example.com:99999/
http://exΓ€mple.com/
> -> $bad {
my $why = (try MCP::Server::Tool::Web::Url.parse($bad)) // $!.message;
say $why;
}
DESCRIPTION
Cro::Uri does RFC 3986 β it parses the grammar and resolves references,
and it is right to be liberal about it, because RFC 3986 describes every URI
anyone has ever written. This class is the opposite: it accepts the small,
boring subset of URLs a web tool should ever act on, and refuses everything
else by name, before a socket is opened.
Everything parse returns is normalised, so two spellings of one page
compare equal:
scheme and host lowercased;
a single trailing root dot removed from the host (
example.com.andexample.comare the same name in DNS, and treating them as two origins would split a crawl in half);an IPv6 literal rewritten in its canonical RFC 5952 form, so
[0:0:0:0:0:0:0:1]and[::1]are one origin;the default port for the scheme dropped from
originandStr, but always available asport;the fragment dropped β it is never sent to a server, and keeping it would make two requests for one page look like two pages.
The path is not touched beyond that: /a/../b is requested as written,
because the server, not this class, decides what a dot segment means to it.
Reference resolution is the one place dot segments are removed, because RFC
3986 Β§5 says to β so parse('http://x/a/../b') keeps the path and
parse('http://x/').resolve('/a/../b') does not. Compare fetched pages by
their final URL if you need one canonical spelling per page.
What is refused, and why
Each of these throws X::MCP::Server::Tool::Web::BadUrl, whose message
names the rule. No allow-list rescues any of them: these are shape rules,
decided before anything is resolved, and allow-private / allow-hosts /
allow-cidrs / allow-loopback are about addresses, not shapes.
Non-absolute, or a scheme other than http/https.
file:would read the disk,javascript:anddata:mean nothing to a fetcher, and a relative URL has no host to check.Credentials in the URL.
http://user:pw@host/exists mostly to make a host look like a path to a human ("http://[email protected]"). Refused outright rather than stripped, so nobody has to wonder which half was used.Numerically obfuscated hosts.
2130706433,0x7f000001,017700000001and127.1are all127.0.0.1β legal forgetaddrinfo, unreadable for a human reviewing a log. The connector floor would catch them anyway (it checks the address of the socket it actually holds), but a URL nobody can read is refused as a matter of clarity, not just safety.Percent-encoding in the host.
http://127.0.0.%31/decodes to127.0.0.1. There is no legitimate use.Non-ASCII hosts. Cro percent-encodes them rather than punycoding, and a percent-encoded host does not resolve; the refusal asks for the A-label (
xn--β¦) form instead, which is what DNS wants anyway.Ports outside 1..65535, IPv6 zone identifiers, empty or over-long host labels, and characters that are not valid in a host name.
Resolving redirects
resolve is Cro::Uri.add β the RFC 3986 Β§5 reference resolution algorithm,
which is fiddly enough to be worth not reimplementing β followed by the whole
of parse again. That second half is the point: a Location: file:///etc/passwd
or Location: http://10.0.0.1/ is refused on the hop that produced it, with
the same message it would have got as a starting URL.