pdf-tag-dump
SYNOPSIS
pdf-dom-dump.raku [options] file.pdf
Options: --password password for an encrypted PDF --max-depth=n maximum tag-depth to descend --select=XPath nodes to be included --omit=tag-name nodes to be excluded --root=tag-name define outer root tag --roles translate role-map to tags, where possible --class-names set 'class' attribute, don't expand classes --/fields disable field values --marks descend into marked content --artifacts descend into artifacts --debug add debugging to output --/atts omit attributes in tags --/strict suppress warnings --quiet avoid printing messages --dtd=link External DtD to use --/valid omit external DtD declaration --xsl=link xsl (stage1) stylesheet to use --css=link css (stage2) stylesheet to use --/style omit stylesheet declarations
DESCRIPTION
Dumps structure elements from a tagged PDF.
Produces tagged output in an XML format.
Only some PDF files contain tagged PDF. pdf-info can be used to check this:
% pdf-info my-doc.pdf | grep Tagged:
Tagged: yes
DEPENDENCIES
This script requires the freetype6 native library and the PDF::Font::Loader Raku module to be installed on your system.