Extract

NAME

MCP::Server::Tool::Web::Extract - turning a fetched page into text a model can read

SYNOPSIS

use MCP::Server::Tool::Web::Extract;

my $html = q:to/HTML/;
	<!doctype html>
	<html>
	<head><title>Raku &amp; regexes</title><script>var x = "</div>";</script></head>
	<body>
	  <nav><a href="/">Home</a></nav>
	  <main>
	    <h1>Grammars</h1>
	    <p>See the <a href="/docs/grammar">grammar docs</a> for&nbsp;more.</p>
	    <pre><code>rule TOP { &lt;thing&gt;+ }</code></pre>
	  </main>
	</body>
	</html>
	HTML

say extract-title($html);
# Raku & regexes

say html-to-text($html, base => 'https://example.test/page');
# # Grammars
#
# See the [grammar docs](https://example.test/docs/grammar) for more.
#
# rule TOP { <thing>+ }

.say for extract-links($html, base => 'https://example.test/page');
# {nofollow => False, url => https://example.test/}
# {nofollow => False, url => https://example.test/docs/grammar}

say text-for('{"b":1,"a":2}', content-type => 'application/json');
# {
#   "a": 2,
#   "b": 1
# }

DESCRIPTION

Everything in this module is a pure function of its arguments: no I/O, no network, no module state, no clock, no locale. The same string in gives the same string out, byte for byte, on every platform — web_grep numbers the lines of this output and quotes them back with line numbers, so a rendering that shifted by a blank line between two runs would make every citation a lie. The determinism is tested, not hoped for.

The parser is a single hand-rolled character scanner with a small state machine. It is deliberately not a chain of regexes: real pages carry unterminated comments, tag soup, minified scripts with < </div> > inside a JavaScript string, and megabytes of markup, and a regex chain over that is one pathological input away from taking a minute. Every scan here advances monotonically through the document, so a 1MB page is linear work (a test pins the wall clock).

What is thrown away

Comments (including an unterminated C<< <!-- >> that runs to the end of the document), < <!DOCTYPE …> >, < <![CDATA[…]]> > and processing instructions never reach the renderer. Neither does the content of < <script> >, < <style> >, < <svg> >, < <template> >, < <iframe> >, < <noscript> > or < <head> >.

Those elements are skipped by scanning for their literal closing tag rather than by parsing what is inside them, which is the only way to survive real scripts:

say html-to-text('<script>var a = "</div>";</script><p>Kept</p>');
# Kept

< <head> > is the one exception to "skip by literal close tag": it is dropped a step later, from the token stream, so that a < <base href> > inside it is still available for resolving links, and so that a document with no < </head> > loses its head only as far as its < <body> > rather than losing everything. < <title> > and < <textarea> > keep their text as one verbatim chunk, so < <title>a < b</title> > reads correctly.

Which part of the page is rendered

In order, first hit wins, and a candidate that turns out to hold no text at all is passed over:

  • the first < <main> >;

  • otherwise every < <article> >, concatenated in document order;

  • otherwise < <body> > with < <nav> >, < <header> >, < <footer> > and < <aside> > removed;

  • otherwise the whole document.

The furniture removal only happens in the < <body> > case: a page whose author put everything inside a < <main> > has already said where the content is, and a document with no < <body> > at all is usually a fragment in which every element is content.

How it renders

Output is markdown-flavoured, because that is the shape a language model reads best and the shape it will quote back:

Element | Rendered as
< <h1> >…< <h6> > | # … ###### then the heading text
< <p> > | text with a blank line around it
< <div> > | text with a line break around it
< <li> > | - item, indented two spaces per level of nesting
< <pre> > | verbatim: line breaks and indentation survive
< <a href> > | [text](absolute-url)
< <img alt> > | [image: alt]
< <br> > | a line break
< <td> >, < <th> > | cells joined with | , one row per line

Anchors are resolved against the document's < <base href> > if it has one, otherwise against the :base you pass (normally the fetch's final URL). A link that cannot be made absolute, that is not http or https, or that points at the page being rendered is written as its bare text — a [Skip to content](#main) line teaches a model nothing and costs it tokens. An anchor with no text at all disappears entirely, as does an < <img> > with no alt; both are decoration.

Inside < <pre> > collapsing is suspended and links are left as plain text, because code samples are the payload and [foo](https://…) in the middle of one is corruption rather than annotation.

Entities

decode-entities handles the named entities that actually turn up in prose (about eighty of them), plus numeric &#8212; and hex &#x2014; forms with full codepoint validation. Deliberate choices:

  • &nbsp; becomes an ordinary space (U+0020), and so do &ensp;, &emsp;, &thinsp; and a literal U+00A0 in the source. A model that greps for "foo bar" has to match text a designer wrote with a non-breaking space.

  • Zero-width and directional marks (&shy;, &zwj;, &lrm;, …) become nothing at all: invisible characters that break a literal search are worse than useless in extracted text.

  • Surrogates, &#0; and anything above U+10FFFF are dropped — they cannot be represented, and a replacement character would only be noise a grep would have to learn to ignore.

  • &#128;..&#159; are read as Windows-1252, exactly as the HTML standard requires; that is where real pages' curly quotes live.

  • An unknown name is left verbatim: &foo; stays &foo;, and a bare & stays a bare &. Guessing is how AT&T turns into something that is not AT&T.

A named entity must carry its semicolon. The legacy semicolon-less forms are ambiguous (&notit; is famously two different things) and this is a reading aid, not a browser.

Whitespace, and why it is nailed down

Outside < <pre> >: runs of whitespace collapse to one space, every line is trimmed, and a run of blank lines becomes a single blank line. Leading and trailing blank lines are removed and the result does not end with a newline.

That is the whole contract web_grep depends on. Line n of this output is line n of the same output tomorrow, on macOS, Linux and Windows alike; input line endings are normalised to \n before anything else happens, so a page served with CRLF and the same page served with LF extract identically.

extract-links works on the raw HTML rather than on the extracted region: a crawler wants the site's navigation, which is exactly what html-to-text throws away. It reports each link once, in document order, with whether it was marked rel="nofollow":

my @links = extract-links(
	'<a href="/a">A</a> <a href="/a#x">again</a> <a rel="nofollow" href="/b">B</a>',
	base => 'https://example.test/',
);
# ({url => https://example.test/a, nofollow => False},
#  {url => https://example.test/b, nofollow => True })

Fragments are stripped before deduplication (/a and /a#x are one page), only http and https survive, and a URL that appears both with and without nofollow is reported as followable — the restrictive marking has to be unanimous to count.

Dispatching on content type

text-for is the seam the tool layer calls with whatever the server said the body was:

Content type | Result
text/html, application/xhtml+xml | extracted as HTML
text/* | verbatim
application/json, …+json | re-printed with :pretty :sorted-keys
application/xml, …+xml | verbatim
absent | sniffed from the content
image/*, application/pdf, … | refused, naming the type

JSON is pretty-printed because a 200KB API response on one line is one grep hit that contains everything and locates nothing; sorted keys are what make two fetches of the same document produce the same line numbers. A body that claims to be JSON and is not comes back verbatim with a bracketed notice appended — appended, not prepended, so the line numbers of the content itself do not shift because of the diagnosis.

A binary type dies rather than returning bytes-as-text. The message names the type and the size, because "I cannot read a PDF" is something a model can act on and 40KB of mojibake is not.

EXAMPLES

Extracting the readable part of a documentation page and numbering it the way web_grep does:

my $text = html-to-text($fetched.text, base => $fetched.final-url);
for $text.lines.kv -> $index, $line {
	say "{$fetched.final-url}:{$index + 1}:$line" if $line.contains('Signature');
}

Feeding a crawler's frontier, honouring nofollow:

my @next = extract-links($body, base => $final-url)
	.grep({ !.<nofollow> })
	.map({ .<url> });

Reading a JSON API response and an HTML page through one call:

my $body = text-for(
	$response.text,
	content-type => $response.content-type,
	url          => $response.final-url,
	byte-count   => $response.byte-count,
);

Decoding entities on their own — useful for a title or a meta description pulled out of markup by something else:

say decode-entities('caf&eacute; &amp; cr&#232;me &mdash; 100&nbsp;%');
# café & crème — 100 %
say decode-entities('AT&T &unknown; &#xZZ;');
# AT&T &unknown; &#xZZ;

SEE ALSO

MCP::Server::Tool::Web — the tool pack whose web_fetch, web_crawl and web_grep tools consume this module.

MCP::Server::Tool::Web v0.1.1

web search, fetch, crawl and grep for MCP::Server

Authors

  • Matt Doughty

License

Artistic-2.0

Dependencies

MCP::Server:auth<zef:apogee>:ver<0.6.0+>Cro::HTTP:auth<zef:cro>:ver<0.8.11+>Cro::Core:auth<zef:cro>:ver<0.8.10+>IO::Socket::Async::SSL:auth<zef:raku-community-modules>:ver<0.8.2+>JSON::Fast:ver<0.19>:auth<cpan:TIMOTIMO>

Test Dependencies

Provides

  • MCP::Server::Tool::Web
  • MCP::Server::Tool::Web::Addr
  • MCP::Server::Tool::Web::Budget
  • MCP::Server::Tool::Web::Crawl
  • MCP::Server::Tool::Web::Extract
  • MCP::Server::Tool::Web::Fetcher
  • MCP::Server::Tool::Web::Guard
  • MCP::Server::Tool::Web::Provider::Brave
  • MCP::Server::Tool::Web::Robots
  • MCP::Server::Tool::Web::SearchProvider
  • MCP::Server::Tool::Web::Transport
  • MCP::Server::Tool::Web::Url
  • MCP::Server::Tool::Web::X

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite — the markup and publishing tools behind this site.