Query
NAME
TreeSitter::Native::Query - tree-sitter queries, with predicates applied
SYNOPSIS
use TreeSitter::Native;
use TreeSitter::Native::Query;
my $parser = Parser.new('python');
my $tree = $parser.parse($source);
my $q = Query.new('python', tags-query('python'));
for $q.matches($tree.root-node) -> $match {
with $match.captures<name> -> $name {
say $name.text, ' @ ', $name.start-point;
}
}
# Or, ignoring which pattern each capture came from:
for $q.captures($tree.root-node) -> (:key($name), :value($node)) {
say "$name: {$node.text}";
}
DESCRIPTION
A query is one or more S-expression patterns, compiled against a
grammar, matched against a subtree. This is the same query language the
tree-sitter CLI and every editor integration uses, and the .scm
files each grammar ships (TreeSitter::Native::Languages) are written
in it.
Predicates are evaluated here
tree-sitter's C library parses predicates ā the (#eq? @a "x")
clauses inside a pattern ā and then deliberately does nothing with them,
because what they mean depends on how the binding can read node text.
Bindings that skip that step return matches their queries plainly say
should be excluded. The bundled javascript tags.scm is exactly such a
query: without predicate evaluation it reports every constructor as a
method definition and every require(ā¦) as a function call.
So this binding evaluates them. Six are supported:
| Predicate | Meaning |
|---|---|
| #eq? @a "text" | every node captured as @a has that text |
| #eq? @a @b | @a and @b capture the same text, pairwise |
| #not-eq? ⦠| the negation of either form |
| #match? @a "regex" | every node captured as @a matches |
| #not-match? @a "regex" | no node captured as @a matches |
| #any-of? @a "x" "y" | every @a node's text is one of the list |
| #not-any-of? @a "x" "y" | no @a node's text is |
A capture that a match does not contain has no nodes, and a predicate over no nodes is vacuously true ā the same rule tree-sitter's own Rust binding uses.
Text comparison for #eq? and #any-of? is done on bytes, so it
is exact regardless of whether the source decodes as UTF-8.
Directives are exposed, not applied
Clauses ending in ! rather than ? ā #set!, #strip!,
#select-adjacent! ā are directives: they do not filter matches,
they tell the consuming application to do something with one. There is
no agreed set of them and no agreed semantics, so this binding parses
them, hands them to you, and applies none:
for $q.pattern-directives(0) -> $d {
say $d.name, ' ', $d.args.map(*.value).join(' ');
}
# strip! doc ^[\s\*/]+|^[\s\*/]$
# select-adjacent! doc definition.method
In the bundled tags.scm files, every directive shapes only the
@doc capture ā which comment block belongs to which definition, and
how much leading * to strip off it. Nothing else is affected by their
absence.
Unknown predicates fail closed
A query naming a predicate that is not in the table above is rejected at
construction with an X::TreeSitter::Query of kind
QUERY_ERROR_PREDICATE. Silently ignoring it would mean quietly
returning matches the query says to exclude, which is the bug this whole
module exists to avoid.
If you would rather have the matches anyway ā you are exploring someone
else's query, or the predicate is one your own code will apply ā pass
:lenient:
my $q = Query.new('go', $scm, :lenient);
# (#foo? @x "y") is now parsed, exposed via .pattern-predicates,
# and ignored when filtering.
:lenient covers unrecognised names only. A #eq? with three
arguments is a malformed query whichever way you look at it, and is
rejected either way.
One vendored query needs :lenient, and exactly one:
highlights-query('javascript') uses #is-not? @x local, an
editor-ecosystem property predicate rather than a text filter. All
twenty-five other vendored query files ā every tags.scm, every other
highlights.scm, and the locals, injections and javascript's two
extra highlight queries ā compile without it.
# The one that needs it:
my $q = Query.new('javascript', highlights-query('javascript'),
:lenient);
Regular expressions
#match? patterns are written in the query file in Perl/PCRE-ish
syntax. They are translated, at query construction, into the equivalent
Raku regex, and anything the translator does not recognise is a
construction-time error rather than a silently different match.
Translated: literals, ., ^, $, \A, \z, \Z, \b,
\B, the \d \D \w \W \s \S \n \r \t \f \e escapes, \xHH and
\x{HHHH}, character classes including ranges and negation, groups
(capturing, (?:ā¦), < (?<name>ā¦) >), lookaround ((?=ā¦),
(?!ā¦), C<< (?<=ā¦) >>, C<< (?<!ā¦) >>), the (?i) and (?i:ā¦)
case-insensitivity
flags, alternation, * + ? with their lazy *? +? ?? forms, and
{n} / {n,} / {n,m}.
Rejected: backreferences, possessive quantifiers, POSIX bracket classes
([:alpha:]), and any other (?ā¦) construct.
Two deliberate approximations, both documented so they cannot surprise
you quietly: $ means end-of-string here, where Perl also lets it
match before a final newline; and alternation is translated to Raku's
||, which tries alternatives left to right exactly as Perl does
(Raku's bare | would prefer the longest match instead). Since
#match? only ever asks whether a match exists, neither changes an
answer for any pattern that is not itself ambiguous.