Native
NAME
TreeSitter::Native - parse and navigate source code with tree-sitter
SYNOPSIS
use TreeSitter::Native;
my $parser = Parser.new('python');
my $tree = $parser.parse(q:to/PY/);
def greet(name):
return f"hello {name}"
PY
my $root = $tree.root-node;
say $root.type; # module
say $root.sexp; # (module (function_definition β¦
my $def = $root.named-child(0);
say $def.child-by-field-name('name').text; # greet
say $def.start-point.gist; # 0:0
$tree.dispose;
$parser.dispose;
Or, letting scope do the freeing:
{
my $parser = Parser.new('go');
my $tree = $parser.parse($source);
say $tree.root-node.walk.grep(*.type eq 'function_declaration').elems;
} # both collected, both freed, in any order
DESCRIPTION
Four classes, in the order you meet them.
Parser holds a grammar and turns source into trees. Tree owns a
parse result and the bytes it describes. Node is an immutable handle
on one node of a tree β a value object, so two handles on the same node
compare equal and hash the same. Cursor walks a tree without
allocating a Node per step.
Parser, Tree and Node are exported. Cursor and Point are
not, because Raku's core already has a Cursor and importing a second
one over it would be rude; reach them through Node.cursor and
Node.start-point, or by their full names
TreeSitter::Native::Cursor and TreeSitter::Native::Point.
Byte offsets, always
Every offset in this distribution β start-byte, end-byte, point
columns, the byte ranges you can restrict a query to β is a byte offset
into the UTF-8 encoding of the source. Not a character index, and
emphatically not a grapheme index.
This matters more than it sounds. In Raku "\r\n" is one grapheme,
so a Windows-line-ended file's .chars and its byte length disagree by
one per line; a file with any non-ASCII identifier disagrees by more.
Slicing source text with .substr and a tree-sitter offset produces
quiet nonsense.
Parser.parse encodes its argument to UTF-8 exactly once and hands the
resulting Blob to the Tree, which keeps it. Node.text slices
that buffer and decodes the slice. Do your own slicing the same way:
my $bytes = $tree.bytes;
my $slice = $bytes.subbuf($node.start-byte,
$node.end-byte - $node.start-byte);
say $slice.decode('utf-8'); # exactly $node.text
Lifetimes
Parser, Tree and Cursor each own a C allocation. Each has an
idempotent .dispose and a DESTROY that calls it, so you may free
them explicitly or leave them to the garbage collector, and calling
.dispose twice β or after DESTROY already ran β is harmless.
A Tree outlives the Parser that made it; disposing the parser does
not touch trees it produced. A Node keeps its Tree alive by
holding a reference to it, so a node is never left pointing into freed
memory by ordinary garbage collection.
The one thing you can still do wrong is dispose a Tree explicitly and
then use a Node from it. That is checked: every node accessor asks
its tree whether it is still alive first, and throws rather than reading
freed memory.