Lex

NAME

Markdown::Lex - a total, streaming-tolerant Markdown lexer: typed blocks and flat styled spans, no renderer attached

SYNOPSIS


use Markdown::Lex;

for parse("# Title\n\nHello **world** and `code`.\n") -> $block {
    given $block {
        when Markdown::Lex::Heading {
            say "H{$block.level}: ", $block.inlines».text.join;
        }
        when Markdown::Lex::Para {
            for $block.inlines -> $span {
                print bold($span.text)   if $span.bold;
                print italic($span.text) if $span.italic;
                print plain($span.text)  unless $span.bold || $span.italic;
            }
        }
    }
}

DESCRIPTION

A lexer for the Markdown a language model actually emits, aimed at a terminal renderer that has to draw the text while it is still arriving.

Two properties drive every design decision in here:

  • It is total. parse never throws. Not on Str:U, not on the empty string, not on a document that is one unterminated code fence, not on ten thousand asterisks in a row. Whatever comes in, a List of blocks comes out.

  • It is prefix-stable. Every prefix of a document is itself a valid document. A consumer that re-lexes the accumulated text on every token batch never sees an error and never sees text vanish — a half-typed **bo is literal text until the closing ** arrives, at which point it becomes bold.

Everything else — colour, wrapping, indentation, hyperlink escape sequences, whether code gets a background — belongs to the renderer. This module has no opinion and no dependencies.

THE OUTPUT

parse returns an immutable List of block objects. Every block is a Markdown::Lex::Block, is immutable, and has a useful .gist:

Class Attributes
Markdown::Lex::Heading Int:D $.level (1..6), @.inlines
Markdown::Lex::Para @.inlines
Markdown::Lex::CodeFence Str $.lang, @.lines, Bool:D $.closed
Markdown::Lex::Bullet Int:D $.depth (0-based), Str:D $.marker, @.inlines
Markdown::Lex::Quote @.inlines
Markdown::Lex::Rule (none)
Markdown::Lex::Table @.alignments, @.header, @.rows, .columns

The classes are plain global names. use Markdown::Lex; exports exactly one symbol — the parse sub — and the classes are reachable fully qualified, which is what you want in a given/when anyway.

A table's cells are Markdown::Lex::Cell, which is not a Block — it is a holder for one cell's @.inlines and nothing else. See #TABLES.

Spans are flat, not a tree

@.inlines is a flat list of Markdown::Lex::Span, never a tree:


class Markdown::Lex::Span {
    has Str:D  $.text is required;
    has Bool:D $.bold   = False;
    has Bool:D $.italic = False;
    has Bool:D $.code   = False;
    has Str    $.link;              # undefined unless this span is a link
}

Nesting is expressed by combining flags, because that is exactly what a terminal can draw. **bold with *both* inside** lexes to three spans:


my @spans = parse("**bold with *both* inside**").head.inlines;

@spans[0];   # Span("bold with " :b)
@spans[1];   # Span("both" :bi)
@spans[2];   # Span(" inside" :b)

There is therefore no nesting depth limit, and no recursion: twenty levels of emphasis saturate at bold + italic and cost nothing. Adjacent runs with identical flags are merged, and a span with empty text is never emitted, so a renderer can loop over @.inlines without defensive checks.

The span vocabulary has exactly one other member, and you only ever meet it if you asked for it: Markdown::Lex::Break, a zero-text Span subclass emitted under parse's :hard-breaks option to mark a newline the author meant. A renderer that has never heard of it prints .text and draws nothing, which is why it is a Span and not a class of its own. See #HARD LINE BREAKS.

STREAMING

The prefix guarantee, as a consumer uses it:


my $accumulated = '';

react whenever $token-supply -> $chunk {
    $accumulated ~= $chunk;
    render(parse($accumulated));       # always safe, always sensible
}

Mid-stream constructs degrade gracefully rather than disappearing:

Partial text What you get while it is partial
**bo one literal span "**bo"
``` then code lines a CodeFence with :!closed that grows line by line
[label](htt one literal span "[label](htt"
`unfinished one literal span "`unfinished"

A pipe table takes two lines before it is a table at all, so what it degrades to is an honest paragraph rather than a placeholder:


parse("| a | b |");             # one Para: a header row on its own is text
parse("| a | b |\n|---");       # one Para: one delimiter cell, two header cells
parse("| a | b |\n|-|-");       # a Table, two columns wide, with no rows yet

CodeFence.closed is the one place the distinction is surfaced deliberately: a renderer can draw an unclosed fence with a "still writing" affordance, and it will close itself when the final ``` arrives.

WHAT IT LEXES

Blocks

  • ATX headings — one to six # followed by a space (or end of line, for an empty heading). A trailing run of # is stripped: "## Title ##" is Title.

  • Fenced code — three or more backticks or tildes. The first word of the info string becomes $.lang (undefined when absent). The fence closes on a line of the same character, at least as long, with nothing else on it. An unclosed fence at end of input is still a CodeFence, with :!closed. The opening fence's indentation is stripped from the content lines; everything else about them is verbatim, Markdown syntax included.

  • Bullets — -, * or +, or an ordinal (1., 2), 17.), followed by a space. $.marker is the marker exactly as written, so a renderer can keep the author's numbering. $.depth is the leading indentation divided by two, floored: zero, two, four... spaces are depths 0, 1, 2, and a tab counts as however many columns it takes to reach the next four-column stop. A following non-blank, non-structural line continues the bullet rather than starting a paragraph.

  • Quotes — a leading < > >, with or without the space after it. Consecutive quote lines merge into one Quote.

  • Rules — three or more -, * or _ alone on a line, spaces allowed between them. This wins over the bullet reading, so - - - is a rule, as CommonMark says.

  • Tables — a row of pipe-separated cells with a delimiter row under it. Both lines have to carry a pipe, and the two have to agree on how many columns there are; without that the header row is ordinary paragraph text. #TABLES has the whole of it.

  • Paragraphs — everything else. Consecutive non-blank, non-structural lines merge into one Para and the line breaks between them become single spaces. A blank line ends whatever block is open.

Inline


parse('`code`').head.inlines;             # Span("code" :c)
parse('**bold**').head.inlines;           # Span("bold" :b)
parse('*italic*').head.inlines;           # Span("italic" :i)
parse('_italic_').head.inlines;           # Span("italic" :i)
parse('__bold__').head.inlines;           # Span("bold" :b)
parse('***both***').head.inlines;         # Span("both" :bi)
parse('[docs](http://x)').head.inlines;   # Span("docs" ->http://x)

Code spans take a run of N backticks and close on the next run of exactly N, so ``a `b` c`` is one code span containing a `b` c. The content is verbatim: no emphasis, no links, no escapes. One leading and one trailing space are dropped if both are present, which is how you write a code span containing a backtick.

Emphasis follows CommonMark's delimiter-run rules closely enough that the cases which bite in practice come out right:


parse('*a **b** c*').head.inlines;    # "a " :i / "b" :bi / " c" :i
parse('**a*').head.inlines;           # "*" / "a" :i        (leftover is literal)
parse('a * b * c').head.inlines;      # one literal span    (space-flanked)
parse('snake_case_name').head.inlines;# one literal span    (intraword _)
parse('foo*bar*baz').head.inlines;    # "foo" / "bar" :i / "baz"

Links are [text](destination). The destination may be wrapped in angle brackets, and a title after it is discarded. The bracket text is not lexed for styles — the span's .text is exactly what you wrote between the brackets — but it does inherit the emphasis it sits inside, so a link inside **...** comes back bold. [](url) renders the URL as its own label.

An unmatched delimiter is always literal text. Nothing is ever dropped.

TABLES

A GFM pipe table is a header row, a delimiter row underneath it that gives every column its alignment, and any number of body rows:


my $table = parse(q:to/MD/).head;
    | Package | Version | Status  |
    |:--------|:-------:|--------:|
    | Selkie  | 0.13.0  | shipped |
    | Sadna   | 0.2.0   | soon    |
    MD

$table.columns;                          # 3
$table.alignments;                       # ("left", "center", "right")
$table.header[0].inlines».text.join;     # "Package"
$table.rows.elems;                       # 2
$table.rows[1][2].inlines».text.join;    # "soon"
$table.gist;                     # "[table 3x2] Package | Version | Status"

The three attributes are rectangular by construction: @.alignments, @.header and every row in @.rows are all .columns long, always, so a renderer can zip a row against the alignments without counting or guarding anything.

@.alignments holds plain lowercase strings — 'left', 'center' or 'right' — rather than an enum, so that a when 'center' in a renderer needs no import and a debug dump reads as itself. A delimiter cell with no colons is 'left', because that is what every renderer draws it as.

What makes a table

Two lines, checked together:

  • The header row is any line with an unescaped pipe in it that is not already something else. Structural markers win, so < - a | b > is a bullet and < > a | b > is a quote, whatever sits under them.

  • The delimiter row is the line directly below, split into cells the same way. Every cell must be a run of one or more - with an optional leading and/or trailing :, and there must be exactly as many of them as the header had. It must contain a pipe of its own, which is what keeps --- a thematic break and - a bullet instead of turning every dash under a piped line into a one-column table.


parse("a | b\n---|---");       # a Table: 2 header cells, 2 delimiter cells
parse("a | b\n---");           # a Para and a Rule: one delimiter cell, not two
parse("| a | b |\n|- - -|-|"); # one Para: spaces are not allowed in a delimiter
parse("| a |\n|-|");           # a Table, one column wide
parse("a\n-|-");               # one Para of both lines: the header has no pipe

Alignment comes from the colons: :--- and --- are 'left', :---: is 'center', ---: is 'right'. A single - is a perfectly good delimiter cell, so |-|-| is the shortest two-column table you can write.

A table interrupts an open paragraph, as GFM says it does — the paragraph is closed above it and the table starts:


parse("intro\n| a | b |\n|-|-|").map(*.gist);
# [para] intro
# [table 2x0] a | b

Cells

Cells are trimmed, and the outer pipes are the table's punctuation rather than empty cells: < | a | b | >, < a | b > and < | a | b > are all the same two cells. An empty cell that no outer pipe created survives, which is how you write a trailing empty one:


parse("| a | |\n|-|-|").head.header».inlines».elems;   # (1, 0)

Each cell's @.inlines is lexed exactly the way a paragraph's is — same Markdown::Lex::Span, same flags, same merging — so emphasis, code spans and links inside a cell all work:


my $row = parse("| a | b |\n|-|-|\n| **hi** | [d](u) |").head.rows[0];
$row[0].inlines.head.gist;    # Span("hi" :b)
$row[1].inlines.head.gist;    # Span("d" ->u)

The order matters and is GFM's: cells are cut first, then each cell is lexed. A pipe inside a code span is therefore still a cell boundary — < | `a|b` | > is two cells holding `a and b` — which looks wrong for about a second and is the only rule that can be implemented without lexing the whole row twice. To put a pipe inside a cell, escape it:


parse("| a \\| b |\n|-|").head.header[0].inlines.head.text;   # "a | b"

\| is the only backslash escape in this module (see #DIVERGENCES AND NON-GOALS), and it exists because there is otherwise no way at all to get a pipe into a cell. It is recognised anywhere on a line, so a paragraph line whose only pipes are escaped is not a header candidate at all.

Ragged rows

Rows are normalised to the header's width, which is GFM's rule: cells past the last column are dropped, and a row that stops early is padded with empty cells.


my @rows = parse("| a | b |\n|-|-|\n| 1 | 2 | 3 |\n| 4 |").head.rows;
@rows[0]».inlines».elems;    # (1, 1)   -- the 3 went nowhere
@rows[1]».inlines».elems;    # (1, 0)   -- the missing cell is empty

This is the one place in the module where text goes in and does not come out, and it is a deliberate divergence from "nothing is ever dropped": the alternative — a row wider than its own table — pushes the problem onto every renderer instead. A consumer that must not lose a keystroke should treat a Table as advisory and keep the source text, the way it already has to for the :!closed case.

Where a table ends

A table takes rows until one of these, whichever comes first:

  • a blank line,

  • a line with no unescaped pipe in it,

  • a line that opens a block of its own — a heading, a fence, a rule, a bullet or a quote — even if it does have a pipe,

  • the end of the input.


parse("| a |\n|-|\n| 1 |\nprose").map(*.gist);
# [table 1x1] a
# [para] prose

That second rule is a deliberate divergence: cmark-gfm reads a pipeless line under a table as a one-cell row and pads the rest. Ending the table is the better answer for text that is still arriving, where the line after a table is almost always the next paragraph and swallowing it into the table makes the whole screen jump.

A table while it is still arriving

The two-line rule falls out of prefix-stability rather than fighting it. Each stage below is what parse returns for that prefix, and no character of the document is off the screen at any of them:


parse("| Name");                            # [para] | Name
parse("| Name | Age |");                    # [para] | Name | Age |
parse("| Name | Age |\n|-");                # [para] | Name | Age | |-
parse("| Name | Age |\n|-|-");              # [table 2x0] Name | Age
parse("| Name | Age |\n|-|-\n| Bob | 4 |"); # [table 2x1] Name | Age
parse("| Name | Age |\n|-|-\n| Bob | 4 |\n| Sue |");
                                            # [table 2x2] Name | Age

A renderer that draws .header and .rows as they are gets a table that grows a row at a time and never redraws a cell it has already drawn.

HARD LINE BREAKS

parse takes one option:


parse($text, :hard-breaks);

Off — the default, and byte-for-byte the behaviour described everywhere else in this document — the newlines inside a paragraph, bullet or quote are soft: the lines are joined with single spaces and lexed as one run of text, because that is what Markdown means by a line break.

On, every one of those newlines becomes a Markdown::Lex::Break span sitting between the two lines' spans:


parse("one\ntwo").head.inlines;                # Span("one two")
parse("one\ntwo", :hard-breaks).head.inlines;  # Span("one") Break Span("two")

It is for rendering text a human typed, where the newlines are where they are on purpose — a chat message, a commit message, a form field — and joining them into a wall of prose is visibly wrong. It changes nothing else: blank lines still separate blocks, headings and rules are single lines anyway, a CodeFence was already line-faithful, and table cells never contain a Break because a row is a line.

A Break has no text, no flags and no link, and it is the only span in the module whose .text is the empty string. Draw it as a newline; ignore it and you get the old wall of prose back, minus the spaces.

Emphasis does not cross a hard break

Under :hard-breaks each line is lexed on its own, so a delimiter cannot reach across a newline to find its partner:


parse("**bold across\nthe break**").head.inlines;
# Span("bold across the break" :b)                 -- one bold span

parse("**bold across\nthe break**", :hard-breaks).head.inlines;
# Span("**bold across") Break Span("the break**")  -- literal, both halves

This diverges from GFM, where emphasis does span a soft break. It is the honest reading of what the option asks for — if a newline is a line break, the line is the unit — and it is the cheap one: the alternative is lexing the joined text and then mapping span offsets back onto lines, which is a second model to keep correct for a case (emphasis opened on one line and closed on the next) that human-typed text almost never contains. Both behaviours are pinned by the test suite, so the divergence cannot drift.

Prefix-stability is unaffected, and for the same reason: the option changes only how one block's already-collected lines are lexed.

DIVERGENCES AND NON-GOALS

Deliberately not implemented. Each of these degrades to literal text, which is the only failure mode this module has:

  • Setext headings (Title over =====). --- after a paragraph is a rule here, not an H2.

  • Indented (four-space) code blocks. There is consequently no four-space cliff: structural markers are recognised at any indentation, and indentation means depth for bullets and nothing anywhere else.

  • Images, raw HTML, autolinks, footnotes, reference links and entity references. ![alt](url) lexes as a literal ! followed by a link.

  • Hard line breaks written as two trailing spaces, or as a trailing backslash. Trailing whitespace is trimmed and that is the end of it — a deliberate newline is asked for with :hard-breaks (#HARD LINE BREAKS), which is honest about being a mode rather than pretending a renderer can see invisible characters in a diff.

  • Backslash escapes, with exactly one exception: \| inside a table row. \* is a backslash and an asterisk, and the asterisk still opens emphasis.

  • Nested block structure. A quote inside a quote, a list inside a quote, or a fence inside a bullet are all lexed at the top level; Bullet.depth is the only nesting a consumer gets, and a second < > > stays as literal text.

  • Lazy continuation into a quote. A non-quote line ends the quote and starts a paragraph. (Bullets do take continuation lines, because a soft-wrapped bullet is common and losing the association is visibly wrong.)

Input normalisation is not a divergence but is worth stating: \r\n and lone \r both become \n on entry. These blocks describe something on a screen and are never written back to disk, so there is nothing for the normalisation to corrupt.

AUTHOR

Matt Doughty <[email protected]>

COPYRIGHT AND LICENSE

Copyright 2026 Matt Doughty

This library is free software; you can redistribute it and/or modify it under the Artistic License 2.0.

for @($table.rows) -> @row {
        for @row Z @($table.alignments) -> ($cell, $align) {
            print pad($cell.inlines».text.join, $align);
        }
    }
my @blocks = parse("- one\n- two\n\n> quoted **hard**\n");
    @blocks.map(*.gist).join("\n");
    # [bullet d0 -] one
    # [bullet d0 -] two
    # [quote] quoted hard
parse("one\ntwo").head.inlines;                 # Span("one two")
    parse("one\ntwo", :hard-breaks).head.inlines;   # Span("one") Break Span("two")

Markdown::Lex v0.2.0

a total, streaming-tolerant Markdown lexer: typed blocks and

Authors

  • Matt Doughty

License

Artistic-2.0

Dependencies

Test Dependencies

Provides

  • Markdown::Lex

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite — the markup and publishing tools behind this site.