Native

NAME

TreeSitter::Native - parse and navigate source code with tree-sitter

SYNOPSIS

use TreeSitter::Native;
my $parser = Parser.new('python');
    my $tree   = $parser.parse(q:to/PY/);
        def greet(name):
            return f"hello {name}"
        PY
my $root = $tree.root-node;
    say $root.type;                        # module
    say $root.sexp;                        # (module (function_definition …
my $def = $root.named-child(0);
    say $def.child-by-field-name('name').text;   # greet
    say $def.start-point.gist;                   # 0:0
$tree.dispose;
    $parser.dispose;

Or, letting scope do the freeing:

{
        my $parser = Parser.new('go');
        my $tree   = $parser.parse($source);
        say $tree.root-node.walk.grep(*.type eq 'function_declaration').elems;
    }   # both collected, both freed, in any order

DESCRIPTION

Four classes, in the order you meet them.

Parser holds a grammar and turns source into trees. Tree owns a parse result and the bytes it describes. Node is an immutable handle on one node of a tree β€” a value object, so two handles on the same node compare equal and hash the same. Cursor walks a tree without allocating a Node per step.

Parser, Tree and Node are exported. Cursor and Point are not, because Raku's core already has a Cursor and importing a second one over it would be rude; reach them through Node.cursor and Node.start-point, or by their full names TreeSitter::Native::Cursor and TreeSitter::Native::Point.

Byte offsets, always

Every offset in this distribution β€” start-byte, end-byte, point columns, the byte ranges you can restrict a query to β€” is a byte offset into the UTF-8 encoding of the source. Not a character index, and emphatically not a grapheme index.

This matters more than it sounds. In Raku "\r\n" is one grapheme, so a Windows-line-ended file's .chars and its byte length disagree by one per line; a file with any non-ASCII identifier disagrees by more. Slicing source text with .substr and a tree-sitter offset produces quiet nonsense.

Parser.parse encodes its argument to UTF-8 exactly once and hands the resulting Blob to the Tree, which keeps it. Node.text slices that buffer and decodes the slice. Do your own slicing the same way:

my $bytes = $tree.bytes;
    my $slice = $bytes.subbuf($node.start-byte,
                              $node.end-byte - $node.start-byte);
    say $slice.decode('utf-8');     # exactly $node.text

Lifetimes

Parser, Tree and Cursor each own a C allocation. Each has an idempotent .dispose and a DESTROY that calls it, so you may free them explicitly or leave them to the garbage collector, and calling .dispose twice β€” or after DESTROY already ran β€” is harmless.

A Tree outlives the Parser that made it; disposing the parser does not touch trees it produced. A Node keeps its Tree alive by holding a reference to it, so a node is never left pointing into freed memory by ordinary garbage collection.

The one thing you can still do wrong is dispose a Tree explicitly and then use a Node from it. That is checked: every node accessor asks its tree whether it is still alive first, and throws rather than reading freed memory.

TreeSitter::Native v0.1.0

parse and navigate source code with tree-sitter

Authors

  • Matt Doughty

License

Artistic-2.0

Dependencies

NativeCall

Test Dependencies

Provides

  • TreeSitter::Native
  • TreeSitter::Native::FFI
  • TreeSitter::Native::Languages
  • TreeSitter::Native::Query
  • TreeSitter::Native::Types

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite β€” the markup and publishing tools behind this site.