Extract
NAME
MCP::Server::Tool::Web::Extract - turning a fetched page into text a model can read
SYNOPSIS
use MCP::Server::Tool::Web::Extract;
my $html = q:to/HTML/;
<!doctype html>
<html>
<head><title>Raku & regexes</title><script>var x = "</div>";</script></head>
<body>
<nav><a href="/">Home</a></nav>
<main>
<h1>Grammars</h1>
<p>See the <a href="/docs/grammar">grammar docs</a> for more.</p>
<pre><code>rule TOP { <thing>+ }</code></pre>
</main>
</body>
</html>
HTML
say extract-title($html);
# Raku & regexes
say html-to-text($html, base => 'https://example.test/page');
# # Grammars
#
# See the [grammar docs](https://example.test/docs/grammar) for more.
#
# rule TOP { <thing>+ }
.say for extract-links($html, base => 'https://example.test/page');
# {nofollow => False, url => https://example.test/}
# {nofollow => False, url => https://example.test/docs/grammar}
say text-for('{"b":1,"a":2}', content-type => 'application/json');
# {
# "a": 2,
# "b": 1
# }
DESCRIPTION
Everything in this module is a pure function of its arguments: no I/O, no
network, no module state, no clock, no locale. The same string in gives the
same string out, byte for byte, on every platform — web_grep numbers the
lines of this output and quotes them back with line numbers, so a rendering
that shifted by a blank line between two runs would make every citation a lie.
The determinism is tested, not hoped for.
The parser is a single hand-rolled character scanner with a small state
machine. It is deliberately not a chain of regexes: real pages carry
unterminated comments, tag soup, minified scripts with < </div> > inside a
JavaScript string, and megabytes of markup, and a regex chain over that is one
pathological input away from taking a minute. Every scan here advances
monotonically through the document, so a 1MB page is linear work (a test pins
the wall clock).
What is thrown away
Comments (including an unterminated C<< <!-- >> that runs to the end of the
document), < <!DOCTYPE …> >, < <![CDATA[…]]> > and processing
instructions never reach the renderer. Neither does the content of
< <script> >, < <style> >, < <svg> >, < <template> >,
< <iframe> >, < <noscript> > or < <head> >.
Those elements are skipped by scanning for their literal closing tag rather than by parsing what is inside them, which is the only way to survive real scripts:
say html-to-text('<script>var a = "</div>";</script><p>Kept</p>');
# Kept
< <head> > is the one exception to "skip by literal close tag": it is
dropped a step later, from the token stream, so that a < <base href> >
inside it is still available for resolving links, and so that a document with
no < </head> > loses its head only as far as its < <body> > rather than
losing everything. < <title> > and < <textarea> > keep their text as one
verbatim chunk, so < <title>a < b</title> > reads correctly.
Which part of the page is rendered
In order, first hit wins, and a candidate that turns out to hold no text at all is passed over:
the first
< <main> >;otherwise every
< <article> >, concatenated in document order;otherwise
< <body> >with< <nav> >,< <header> >,< <footer> >and< <aside> >removed;otherwise the whole document.
The furniture removal only happens in the < <body> > case: a page whose
author put everything inside a < <main> > has already said where the content
is, and a document with no < <body> > at all is usually a fragment in which
every element is content.
How it renders
Output is markdown-flavoured, because that is the shape a language model reads best and the shape it will quote back:
| Element | Rendered as |
|---|
< <h1> >…< <h6> > | # … ###### then the heading text |
< <p> > | text with a blank line around it |
< <div> > | text with a line break around it |
< <li> > | - item, indented two spaces per level of nesting |
< <pre> > | verbatim: line breaks and indentation survive |
< <a href> > | [text](absolute-url) |
< <img alt> > | [image: alt] |
< <br> > | a line break |
< <td> >, < <th> > | cells joined with | , one row per line |
Anchors are resolved against the document's < <base href> > if it has one,
otherwise against the :base you pass (normally the fetch's final URL). A
link that cannot be made absolute, that is not http or https, or that
points at the page being rendered is written as its bare text — a
[Skip to content](#main) line teaches a model nothing and costs it tokens.
An anchor with no text at all disappears entirely, as does an < <img> > with
no alt; both are decoration.
Inside < <pre> > collapsing is suspended and links are left as plain text,
because code samples are the payload and [foo](https://…) in the middle of
one is corruption rather than annotation.
Entities
decode-entities handles the named entities that actually turn up in prose
(about eighty of them), plus numeric — and hex — forms with
full codepoint validation. Deliberate choices:
becomes an ordinary space (U+0020), and so do , , and a literal U+00A0 in the source. A model that greps for"foo bar"has to match text a designer wrote with a non-breaking space.Zero-width and directional marks (
­,‍,‎, …) become nothing at all: invisible characters that break a literal search are worse than useless in extracted text.Surrogates,
�and anything above U+10FFFF are dropped — they cannot be represented, and a replacement character would only be noise a grep would have to learn to ignore.€..Ÿare read as Windows-1252, exactly as the HTML standard requires; that is where real pages' curly quotes live.An unknown name is left verbatim:
&foo;stays&foo;, and a bare&stays a bare&. Guessing is howAT&Tturns into something that is notAT&T.
A named entity must carry its semicolon. The legacy semicolon-less forms are
ambiguous (¬it; is famously two different things) and this is a reading
aid, not a browser.
Whitespace, and why it is nailed down
Outside < <pre> >: runs of whitespace collapse to one space, every line is
trimmed, and a run of blank lines becomes a single blank line. Leading and
trailing blank lines are removed and the result does not end with a newline.
That is the whole contract web_grep depends on. Line n of this output is
line n of the same output tomorrow, on macOS, Linux and Windows alike; input
line endings are normalised to \n before anything else happens, so a page
served with CRLF and the same page served with LF extract identically.
Links for the crawler
extract-links works on the raw HTML rather than on the extracted region:
a crawler wants the site's navigation, which is exactly what html-to-text
throws away. It reports each link once, in document order, with whether it was
marked rel="nofollow":
my @links = extract-links(
'<a href="/a">A</a> <a href="/a#x">again</a> <a rel="nofollow" href="/b">B</a>',
base => 'https://example.test/',
);
# ({url => https://example.test/a, nofollow => False},
# {url => https://example.test/b, nofollow => True })
Fragments are stripped before deduplication (/a and /a#x are one page),
only http and https survive, and a URL that appears both with and without
nofollow is reported as followable — the restrictive marking has to be
unanimous to count.
Dispatching on content type
text-for is the seam the tool layer calls with whatever the server said the
body was:
| Content type | Result |
|---|
text/html, application/xhtml+xml | extracted as HTML |
text/* | verbatim |
application/json, …+json | re-printed with :pretty :sorted-keys |
application/xml, …+xml | verbatim |
| absent | sniffed from the content |
image/*, application/pdf, … | refused, naming the type |
JSON is pretty-printed because a 200KB API response on one line is one grep hit that contains everything and locates nothing; sorted keys are what make two fetches of the same document produce the same line numbers. A body that claims to be JSON and is not comes back verbatim with a bracketed notice appended — appended, not prepended, so the line numbers of the content itself do not shift because of the diagnosis.
A binary type dies rather than returning bytes-as-text. The message names
the type and the size, because "I cannot read a PDF" is something a model can
act on and 40KB of mojibake is not.
EXAMPLES
Extracting the readable part of a documentation page and numbering it the way
web_grep does:
my $text = html-to-text($fetched.text, base => $fetched.final-url);
for $text.lines.kv -> $index, $line {
say "{$fetched.final-url}:{$index + 1}:$line" if $line.contains('Signature');
}
Feeding a crawler's frontier, honouring nofollow:
my @next = extract-links($body, base => $final-url)
.grep({ !.<nofollow> })
.map({ .<url> });
Reading a JSON API response and an HTML page through one call:
my $body = text-for(
$response.text,
content-type => $response.content-type,
url => $response.final-url,
byte-count => $response.byte-count,
);
Decoding entities on their own — useful for a title or a meta description
pulled out of markup by something else:
say decode-entities('café & crème — 100 %');
# café & crème — 100 %
say decode-entities('AT&T &unknown; &#xZZ;');
# AT&T &unknown; &#xZZ;
SEE ALSO
MCP::Server::Tool::Web — the tool pack whose
web_fetch, web_crawl and web_grep tools consume this module.