Lex
NAME
Markdown::Lex - a total, streaming-tolerant Markdown lexer: typed blocks and flat styled spans, no renderer attached
SYNOPSIS
use Markdown::Lex;
for parse("# Title\n\nHello **world** and `code`.\n") -> $block {
given $block {
when Markdown::Lex::Heading {
say "H{$block.level}: ", $block.inlines».text.join;
}
when Markdown::Lex::Para {
for $block.inlines -> $span {
print bold($span.text) if $span.bold;
print italic($span.text) if $span.italic;
print plain($span.text) unless $span.bold || $span.italic;
}
}
}
}
DESCRIPTION
A lexer for the Markdown a language model actually emits, aimed at a terminal renderer that has to draw the text while it is still arriving.
Two properties drive every design decision in here:
It is total.
parsenever throws. Not onStr:U, not on the empty string, not on a document that is one unterminated code fence, not on ten thousand asterisks in a row. Whatever comes in, aListof blocks comes out.It is prefix-stable. Every prefix of a document is itself a valid document. A consumer that re-lexes the accumulated text on every token batch never sees an error and never sees text vanish — a half-typed
**bois literal text until the closing**arrives, at which point it becomes bold.
Everything else — colour, wrapping, indentation, hyperlink escape sequences,
whether code gets a background — belongs to the renderer. This module has no
opinion and no dependencies.
THE OUTPUT
parse returns an immutable List of block objects. Every block is a
Markdown::Lex::Block, is immutable, and has a useful .gist:
| Class | Attributes |
|---|---|
| Markdown::Lex::Heading | Int:D $.level (1..6), @.inlines |
| Markdown::Lex::Para | @.inlines |
| Markdown::Lex::CodeFence | Str $.lang, @.lines, Bool:D $.closed |
| Markdown::Lex::Bullet | Int:D $.depth (0-based), Str:D $.marker, @.inlines |
| Markdown::Lex::Quote | @.inlines |
| Markdown::Lex::Rule | (none) |
| Markdown::Lex::Table | @.alignments, @.header, @.rows, .columns |
The classes are plain global names. use Markdown::Lex; exports exactly one
symbol — the parse sub — and the classes are reachable fully qualified,
which is what you want in a given/when anyway.
A table's cells are Markdown::Lex::Cell, which is not a Block — it is a
holder for one cell's @.inlines and nothing else. See #TABLES.
Spans are flat, not a tree
@.inlines is a flat list of Markdown::Lex::Span, never a tree:
class Markdown::Lex::Span {
has Str:D $.text is required;
has Bool:D $.bold = False;
has Bool:D $.italic = False;
has Bool:D $.code = False;
has Str $.link; # undefined unless this span is a link
}
Nesting is expressed by combining flags, because that is exactly what a
terminal can draw. **bold with *both* inside** lexes to three spans:
my @spans = parse("**bold with *both* inside**").head.inlines;
@spans[0]; # Span("bold with " :b)
@spans[1]; # Span("both" :bi)
@spans[2]; # Span(" inside" :b)
There is therefore no nesting depth limit, and no recursion: twenty levels of
emphasis saturate at bold + italic and cost nothing. Adjacent runs with
identical flags are merged, and a span with empty text is never emitted, so a
renderer can loop over @.inlines without defensive checks.
The span vocabulary has exactly one other member, and you only ever meet it if
you asked for it: Markdown::Lex::Break, a zero-text Span subclass emitted
under parse's :hard-breaks option to mark a newline the author meant. A
renderer that has never heard of it prints .text and draws nothing, which is
why it is a Span and not a class of its own. See #HARD LINE BREAKS.
STREAMING
The prefix guarantee, as a consumer uses it:
my $accumulated = '';
react whenever $token-supply -> $chunk {
$accumulated ~= $chunk;
render(parse($accumulated)); # always safe, always sensible
}
Mid-stream constructs degrade gracefully rather than disappearing:
| Partial text | What you get while it is partial |
|---|---|
**bo | one literal span "**bo" |
``` then code lines | a CodeFence with :!closed that grows line by line |
[label](htt | one literal span "[label](htt" |
`unfinished | one literal span "`unfinished" |
A pipe table takes two lines before it is a table at all, so what it degrades to is an honest paragraph rather than a placeholder:
parse("| a | b |"); # one Para: a header row on its own is text
parse("| a | b |\n|---"); # one Para: one delimiter cell, two header cells
parse("| a | b |\n|-|-"); # a Table, two columns wide, with no rows yet
CodeFence.closed is the one place the distinction is surfaced deliberately:
a renderer can draw an unclosed fence with a "still writing" affordance, and it
will close itself when the final ``` arrives.
WHAT IT LEXES
Blocks
ATX headings — one to six
#followed by a space (or end of line, for an empty heading). A trailing run of#is stripped:"## Title ##"isTitle.Fenced code — three or more backticks or tildes. The first word of the info string becomes
$.lang(undefined when absent). The fence closes on a line of the same character, at least as long, with nothing else on it. An unclosed fence at end of input is still aCodeFence, with:!closed. The opening fence's indentation is stripped from the content lines; everything else about them is verbatim, Markdown syntax included.Bullets —
-,*or+, or an ordinal (1.,2),17.), followed by a space.$.markeris the marker exactly as written, so a renderer can keep the author's numbering.$.depthis the leading indentation divided by two, floored: zero, two, four... spaces are depths 0, 1, 2, and a tab counts as however many columns it takes to reach the next four-column stop. A following non-blank, non-structural line continues the bullet rather than starting a paragraph.Quotes — a leading
< >>, with or without the space after it. Consecutive quote lines merge into oneQuote.Rules — three or more
-,*or_alone on a line, spaces allowed between them. This wins over the bullet reading, so- - -is a rule, as CommonMark says.Tables — a row of pipe-separated cells with a delimiter row under it. Both lines have to carry a pipe, and the two have to agree on how many columns there are; without that the header row is ordinary paragraph text. #TABLES has the whole of it.
Paragraphs — everything else. Consecutive non-blank, non-structural lines merge into one
Paraand the line breaks between them become single spaces. A blank line ends whatever block is open.
Inline
parse('`code`').head.inlines; # Span("code" :c)
parse('**bold**').head.inlines; # Span("bold" :b)
parse('*italic*').head.inlines; # Span("italic" :i)
parse('_italic_').head.inlines; # Span("italic" :i)
parse('__bold__').head.inlines; # Span("bold" :b)
parse('***both***').head.inlines; # Span("both" :bi)
parse('[docs](http://x)').head.inlines; # Span("docs" ->http://x)
Code spans take a run of N backticks and close on the next run of exactly
N, so ``a `b` c`` is one code span containing a `b` c. The content is
verbatim: no emphasis, no links, no escapes. One leading and one trailing space
are dropped if both are present, which is how you write a code span containing
a backtick.
Emphasis follows CommonMark's delimiter-run rules closely enough that the cases which bite in practice come out right:
parse('*a **b** c*').head.inlines; # "a " :i / "b" :bi / " c" :i
parse('**a*').head.inlines; # "*" / "a" :i (leftover is literal)
parse('a * b * c').head.inlines; # one literal span (space-flanked)
parse('snake_case_name').head.inlines;# one literal span (intraword _)
parse('foo*bar*baz').head.inlines; # "foo" / "bar" :i / "baz"
Links are [text](destination). The destination may be wrapped in angle
brackets, and a title after it is discarded. The bracket text is not lexed
for styles — the span's .text is exactly what you wrote between the brackets
— but it does inherit the emphasis it sits inside, so a link inside **...**
comes back bold. [](url) renders the URL as its own label.
An unmatched delimiter is always literal text. Nothing is ever dropped.
TABLES
A GFM pipe table is a header row, a delimiter row underneath it that gives every column its alignment, and any number of body rows:
my $table = parse(q:to/MD/).head;
| Package | Version | Status |
|:--------|:-------:|--------:|
| Selkie | 0.13.0 | shipped |
| Sadna | 0.2.0 | soon |
MD
$table.columns; # 3
$table.alignments; # ("left", "center", "right")
$table.header[0].inlines».text.join; # "Package"
$table.rows.elems; # 2
$table.rows[1][2].inlines».text.join; # "soon"
$table.gist; # "[table 3x2] Package | Version | Status"
The three attributes are rectangular by construction: @.alignments,
@.header and every row in @.rows are all .columns long, always, so a
renderer can zip a row against the alignments without counting or guarding
anything.
@.alignments holds plain lowercase strings — 'left', 'center' or
'right' — rather than an enum, so that a when 'center' in a renderer
needs no import and a debug dump reads as itself. A delimiter cell with no
colons is 'left', because that is what every renderer draws it as.
What makes a table
Two lines, checked together:
The header row is any line with an unescaped pipe in it that is not already something else. Structural markers win, so
< - a | b >is a bullet and< > a | b> is a quote, whatever sits under them.The delimiter row is the line directly below, split into cells the same way. Every cell must be a run of one or more
-with an optional leading and/or trailing:, and there must be exactly as many of them as the header had. It must contain a pipe of its own, which is what keeps---a thematic break and-a bullet instead of turning every dash under a piped line into a one-column table.
parse("a | b\n---|---"); # a Table: 2 header cells, 2 delimiter cells
parse("a | b\n---"); # a Para and a Rule: one delimiter cell, not two
parse("| a | b |\n|- - -|-|"); # one Para: spaces are not allowed in a delimiter
parse("| a |\n|-|"); # a Table, one column wide
parse("a\n-|-"); # one Para of both lines: the header has no pipe
Alignment comes from the colons: :--- and --- are 'left', :---: is
'center', ---: is 'right'. A single - is a perfectly good delimiter
cell, so |-|-| is the shortest two-column table you can write.
A table interrupts an open paragraph, as GFM says it does — the paragraph is closed above it and the table starts:
parse("intro\n| a | b |\n|-|-|").map(*.gist);
# [para] intro
# [table 2x0] a | b
Cells
Cells are trimmed, and the outer pipes are the table's punctuation rather than
empty cells: < | a | b | >, < a | b > and < | a | b > are all the same
two cells. An empty cell that no outer pipe created survives, which is how you
write a trailing empty one:
parse("| a | |\n|-|-|").head.header».inlines».elems; # (1, 0)
Each cell's @.inlines is lexed exactly the way a paragraph's is — same
Markdown::Lex::Span, same flags, same merging — so emphasis, code spans and
links inside a cell all work:
my $row = parse("| a | b |\n|-|-|\n| **hi** | [d](u) |").head.rows[0];
$row[0].inlines.head.gist; # Span("hi" :b)
$row[1].inlines.head.gist; # Span("d" ->u)
The order matters and is GFM's: cells are cut first, then each cell is lexed.
A pipe inside a code span is therefore still a cell boundary —
< | `a|b` | > is two cells holding `a and b` — which looks wrong for
about a second and is the only rule that can be implemented without lexing the
whole row twice. To put a pipe inside a cell, escape it:
parse("| a \\| b |\n|-|").head.header[0].inlines.head.text; # "a | b"
\| is the only backslash escape in this module (see #DIVERGENCES AND
NON-GOALS), and it exists because there is otherwise no way at all to get a
pipe into a cell. It is recognised anywhere on a line, so a paragraph line whose
only pipes are escaped is not a header candidate at all.
Ragged rows
Rows are normalised to the header's width, which is GFM's rule: cells past the last column are dropped, and a row that stops early is padded with empty cells.
my @rows = parse("| a | b |\n|-|-|\n| 1 | 2 | 3 |\n| 4 |").head.rows;
@rows[0]».inlines».elems; # (1, 1) -- the 3 went nowhere
@rows[1]».inlines».elems; # (1, 0) -- the missing cell is empty
This is the one place in the module where text goes in and does not come out,
and it is a deliberate divergence from "nothing is ever dropped": the
alternative — a row wider than its own table — pushes the problem onto every
renderer instead. A consumer that must not lose a keystroke should treat a
Table as advisory and keep the source text, the way it already has to for the
:!closed case.
Where a table ends
A table takes rows until one of these, whichever comes first:
a blank line,
a line with no unescaped pipe in it,
a line that opens a block of its own — a heading, a fence, a rule, a bullet or a quote — even if it does have a pipe,
the end of the input.
parse("| a |\n|-|\n| 1 |\nprose").map(*.gist);
# [table 1x1] a
# [para] prose
That second rule is a deliberate divergence: cmark-gfm reads a pipeless line
under a table as a one-cell row and pads the rest. Ending the table is the
better answer for text that is still arriving, where the line after a table is
almost always the next paragraph and swallowing it into the table makes the
whole screen jump.
A table while it is still arriving
The two-line rule falls out of prefix-stability rather than fighting it. Each
stage below is what parse returns for that prefix, and no character of the
document is off the screen at any of them:
parse("| Name"); # [para] | Name
parse("| Name | Age |"); # [para] | Name | Age |
parse("| Name | Age |\n|-"); # [para] | Name | Age | |-
parse("| Name | Age |\n|-|-"); # [table 2x0] Name | Age
parse("| Name | Age |\n|-|-\n| Bob | 4 |"); # [table 2x1] Name | Age
parse("| Name | Age |\n|-|-\n| Bob | 4 |\n| Sue |");
# [table 2x2] Name | Age
A renderer that draws .header and .rows as they are gets a table that
grows a row at a time and never redraws a cell it has already drawn.
HARD LINE BREAKS
parse takes one option:
parse($text, :hard-breaks);
Off — the default, and byte-for-byte the behaviour described everywhere else in this document — the newlines inside a paragraph, bullet or quote are soft: the lines are joined with single spaces and lexed as one run of text, because that is what Markdown means by a line break.
On, every one of those newlines becomes a Markdown::Lex::Break span sitting
between the two lines' spans:
parse("one\ntwo").head.inlines; # Span("one two")
parse("one\ntwo", :hard-breaks).head.inlines; # Span("one") Break Span("two")
It is for rendering text a human typed, where the newlines are where they
are on purpose — a chat message, a commit message, a form field — and joining
them into a wall of prose is visibly wrong. It changes nothing else: blank
lines still separate blocks, headings and rules are single lines anyway, a
CodeFence was already line-faithful, and table cells never contain a
Break because a row is a line.
A Break has no text, no flags and no link, and it is the only span in the
module whose .text is the empty string. Draw it as a newline; ignore it and
you get the old wall of prose back, minus the spaces.
Emphasis does not cross a hard break
Under :hard-breaks each line is lexed on its own, so a delimiter cannot reach
across a newline to find its partner:
parse("**bold across\nthe break**").head.inlines;
# Span("bold across the break" :b) -- one bold span
parse("**bold across\nthe break**", :hard-breaks).head.inlines;
# Span("**bold across") Break Span("the break**") -- literal, both halves
This diverges from GFM, where emphasis does span a soft break. It is the honest reading of what the option asks for — if a newline is a line break, the line is the unit — and it is the cheap one: the alternative is lexing the joined text and then mapping span offsets back onto lines, which is a second model to keep correct for a case (emphasis opened on one line and closed on the next) that human-typed text almost never contains. Both behaviours are pinned by the test suite, so the divergence cannot drift.
Prefix-stability is unaffected, and for the same reason: the option changes only how one block's already-collected lines are lexed.
DIVERGENCES AND NON-GOALS
Deliberately not implemented. Each of these degrades to literal text, which is the only failure mode this module has:
Setext headings (
Titleover=====).---after a paragraph is a rule here, not an H2.Indented (four-space) code blocks. There is consequently no four-space cliff: structural markers are recognised at any indentation, and indentation means depth for bullets and nothing anywhere else.
Images, raw HTML, autolinks, footnotes, reference links and entity references.
lexes as a literal!followed by a link.Hard line breaks written as two trailing spaces, or as a trailing backslash. Trailing whitespace is trimmed and that is the end of it — a deliberate newline is asked for with
:hard-breaks(#HARD LINE BREAKS), which is honest about being a mode rather than pretending a renderer can see invisible characters in a diff.Backslash escapes, with exactly one exception:
\|inside a table row.\*is a backslash and an asterisk, and the asterisk still opens emphasis.Nested block structure. A quote inside a quote, a list inside a quote, or a fence inside a bullet are all lexed at the top level;
Bullet.depthis the only nesting a consumer gets, and a second< >> stays as literal text.Lazy continuation into a quote. A non-quote line ends the quote and starts a paragraph. (Bullets do take continuation lines, because a soft-wrapped bullet is common and losing the association is visibly wrong.)
Input normalisation is not a divergence but is worth stating: \r\n and lone
\r both become \n on entry. These blocks describe something on a screen
and are never written back to disk, so there is nothing for the normalisation to
corrupt.
AUTHOR
Matt Doughty <[email protected]>
COPYRIGHT AND LICENSE
Copyright 2026 Matt Doughty
This library is free software; you can redistribute it and/or modify it under the Artistic License 2.0.
for @($table.rows) -> @row {
for @row Z @($table.alignments) -> ($cell, $align) {
print pad($cell.inlines».text.join, $align);
}
}
my @blocks = parse("- one\n- two\n\n> quoted **hard**\n");
@blocks.map(*.gist).join("\n");
# [bullet d0 -] one
# [bullet d0 -] two
# [quote] quoted hard
parse("one\ntwo").head.inlines; # Span("one two")
parse("one\ntwo", :hard-breaks).head.inlines; # Span("one") Break Span("two")