pdf-tag-dump

SYNOPSIS

pdf-dom-dump.raku [options] file.pdf

Options: --password password for an encrypted PDF --max-depth=n maximum tag-depth to descend --select=XPath nodes to be included --omit=tag-name nodes to be excluded --root=tag-name define outer root tag --roles translate role-map to tags, where possible --class-names set 'class' attribute, don't expand classes --/fields disable field values --marks descend into marked content --artifacts descend into artifacts --debug add debugging to output --/atts omit attributes in tags --/strict suppress warnings --quiet avoid printing messages --dtd=link External DtD to use --/valid omit external DtD declaration --xsl=link xsl (stage1) stylesheet to use --css=link css (stage2) stylesheet to use --/style omit stylesheet declarations

DESCRIPTION

Dumps structure elements from a tagged PDF.

Produces tagged output in an XML format.

Only some PDF files contain tagged PDF. pdf-info can be used to check this:

% pdf-info my-doc.pdf | grep Tagged:
    Tagged:     yes

DEPENDENCIES

This script requires the freetype6 native library and the PDF::Font::Loader Raku module to be installed on your system.

PDF::Tags::Reader v0.0.18

Tagged PDF reader

Authors

  • David Warring

License

Artistic-2.0

Dependencies

PDF::Content:ver<0.9.10+>PDF::Tags:ver<0.2.7+>PDF::Font::Loader:ver<0.6.13+>Method::Also

Test Dependencies

Provides

  • PDF::Tags::Reader
  • PDF::Tags::Reader::TextDecoder

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite — the markup and publishing tools behind this site.