Rules
NAME
MCP::Client::Policy::Rules - the permission engine: plain data in, a decision out
SYNOPSIS
use MCP::Client::Policy::Rules;
my @rules =
{ tool => 'fs_read', decision => 'allow' },
{ tool => 'fs_write', decision => 'allow', under => '/srv/scratch' },
{ tool => 'fs_*', decision => 'deny', under => '/srv/secrets' },
;
evaluate(@rules, 'fs_write', { path => '/srv/scratch/notes.md' },
roots => { fs => '/srv' });
# { decision => 'allow', reason => 'rule', rule => {...},
# paths => ('/srv/scratch/notes.md',),
# suggestion => { tool => 'fs_write', under => '/srv/scratch' } }
evaluate(@rules, 'fs_write', { path => '../../etc/passwd' },
roots => { fs => '/srv' });
# { decision => 'ask', reason => 'unevaluable-path', ... }
DESCRIPTION
The half of MCP::Client::Policy that has no state, no locks, no provider and
no user: a list of rules, the name of a tool, the arguments a model asked to
call it with, and out comes one of allow, deny or ask together with
everything a permission prompt needs in order to explain itself. Keeping it
separate is what makes the interesting part testable ā every path in here is
reachable from a single function call with plain data.
A rule
A rule is JSON-safe plain data with two required members and a handful of optional narrowers:
{ tool => 'fs_read', decision => 'allow' } # bare
{ tool => 'fs_*', decision => 'deny', under => '/etc' } # narrowed by path
{ tool => 'sh_run', decision => 'ask', command => 'git', # narrowed by command
args => ['push'], args-any => ['--force', '-f'],
note => 'force pushes rewrite history', severity => 'danger' }
{ tool => 'web_*', decision => 'allow', host => '*.raku.org' } # narrowed by website
toolā an exact tool name, or a prefix glob: a name ending in*matches every name that starts with what comes before it.*on its own matches everything. A*anywhere else is refused, because a rule that silently matches nothing is worse than one that will not load.decisionāallow,denyorask.underā a directory. The rule only has an opinion about calls whose location arguments are inside it.commandā a program basename, exact or a trailing-*glob (never a path separator: it matches the basename, so a/could not fire). The rule only speaks about calls whose program is this one./usr/bin/git,gitandC:\bin\git.EXEall match acommandofgit.argsā a list of tokens the call's argv must begin with, in order. Needs acommand.args-anyā a list of tokens, at least one of which must appear anywhere in the argv (--forceor-f, order no matter). Needs acommand. Combined withargs, both must hold.hostā a website. Either an exact host (docs.raku.org) or a suffix glob of the one shape*.plus a host (*.raku.org). The rule only speaks about calls whoseurlargument names that site.noteā free text saying why the rule exists; a permission prompt shows it, so the human is told what they are being asked about.severityādanger, the one value there is. A prompt renders these loudly.
Anything else in a rule is refused: dir where under was meant would
otherwise turn a narrow rule into a blanket one, which is exactly the mistake a
permission system must not make quietly. New members only ever refuse more, so
an older engine rejects a rule it does not understand loudly, rather than
under-enforcing it ā the safe direction for a permission file to break in.
Narrowing by command, and failing closed
A rule with a command is measured against the program a call runs and its
argv ā the command and args arguments, or whatever command-params
renames them to. The tri-state and the fail-closed asymmetry are the same as
for paths: an allow fires only when the command is one it named; a deny or
ask fires on a match or on anything it cannot read. A call whose argv the
engine cannot see to the end ā because a token is not a literal string ā is
unknown past that point, so an args prefix that would need to read into
the dark, or an args-any token that is not already visible, keeps a deny or
ask firing and stops an allow.
command and under on one rule are ANDed: both must hold. A rule that
allows git only inside a scratch directory means exactly that.
Narrowing by host
A rule with a host is measured against the website a call names in its
url argument ā the one URL-bearing parameter name there is, the naming a
web_search/web_fetch/web_crawl/web_grep pack follows. Without it,
saying yes once to web_fetch would be saying yes to every site there will
ever be, which is not what a human clicking "always allow" on a documentation
page means.
A host is an exact host, or a suffix glob of exactly one shape: *.
followed by a host. Nothing else glob-ish is accepted ā no ?, no * in the
middle, no bare *, no empty string ā because a pattern that matches nothing
is a permission rule that never fires, and one that matches more than its author
read is worse.
Matching is on the ASCII-lowercased host and on a . boundary, never on the
string:
host => 'docs.raku.org' # docs.raku.org, DOCS.RAKU.ORG -- and nothing else
host => '*.raku.org' # docs.raku.org, a.b.raku.org
# NOT raku.org (the glob is a subdomain glob)
# NOT evilraku.org (the '.' boundary is the point)
The host comes out of a deliberately small, strict URL parse: an absolute
http or https URL, its authority taken up to the first /, ? or
#, brackets stripped from an IPv6 literal, any port dropped. Everything else
ā a relative reference, a data: or file: URL, a percent-encoded or
backslash-bearing authority, a value that is not a string ā is unparseable,
and so is a URL carrying userinfo: https://[email protected]#@good.com is a
shape whose host two readers disagree about, and a permission engine may not be
the reader that guesses.
The tri-state and the fail-closed asymmetry are the path predicate's exactly: a
call whose url parses and matches is yes, one that parses and does not is
no, and a call with no url at all or one that cannot be parsed is
unknown. So an allow narrowed by host fires only on a URL it could read
and did recognise, while a deny or ask fires on that or on anything it
could not read. host ANDs with under, command and check as those AND
with each other.
Two things follow, both deliberate:
urlis not one of the location arguments (see below). A URL is never treated as a filesystem path:path-underwould readhttps://evil.com/as a relative path with ahttps:segment and answer a question nobody asked;an
undernarrower on aweb_*rule can never fire. Such a call names no location arguments, so the path predicate isunknown, and unknown fails closed ā the allow stays silent, the deny always speaks. Narrow a web tool byhost, not byunder.
Where paths come from
The location arguments of a call are, by convention, the ones named path,
from and to ā the naming the MCP::Server::Tool::FileSystem pack
follows for precisely this reason. path-params overrides the convention per
tool for anything that does not:
path-params => { 'git_*' => ['repo',], 'sql_query' => [] }
Relative arguments are made absolute against roots, whose keys are
tool-name prefixes (not globs ā the longest matching prefix wins, and the
empty string is a catch-all):
roots => { fs => '/srv/docs', '' => '/tmp/agent' }
A rule's under is absolutized the same way, so a rule may be written
relative to the root it is about.
Containment is lexical, and tri-state
path-under compares two paths segment by segment, and never touches the
filesystem. Two reasons, both load-bearing:
the paths belong to the server's filesystem, which may not be this machine's at all ā resolving
/srv/docshere would answer a question nobody asked;symlink truth is the server pack's job.
MCP::Server::Tool::FileSystemresolves every argument inside its sandbox root before touching it; a client-side.resolvewould be a second, weaker opinion about the same question.
Segment-wise, never string-prefix: /tmp/root2 does not start with
/tmp/root in any sense a permission system may act on, however much the two
strings look alike.
The answer has three values, not two. unknown is returned when the question
cannot honestly be answered lexically:
a
..segment anywhere (the only honest resolution is on the server);a null byte, or a backslash ā on a Windows server
\is a separator and on a POSIX one it is an ordinary character, and guessing wrong either hides an escape or invents one. Write rules and roots with/, which Windows accepts too;an absolute path measured against a relative directory, or the reverse: no root was configured, so there is nothing to measure from;
an argument that is missing, or is not a string at all.
What unknown does
Fail closed, in both directions:
an allow rule with a predicate fires only when the call has location arguments and every one of them is
yes;a deny or ask rule with a predicate fires on any
yesorunknownā an argument that cannot be ruled out is not ruled out;every rule is evaluated, and the strongest decision wins: deny > ask > allow. Order-independence is the point: rule sets get serialized, merged and re-ordered, and a permission system whose answer depends on which half of the file was read first is not one;
no rule matched at all is
ask, not allow. A tool nobody wrote a rule about is a tool nobody has consented to.
EXAMPLES
The two-argument case, where fail-closed earns its keep. fs_move declares
both from and to, so a rule that allows moves inside a scratch directory
must be satisfied about both ends:
my @rules = { tool => 'fs_move', decision => 'allow', under => '/srv/scratch' },;
evaluate(@rules, 'fs_move',
{ from => '/srv/scratch/a', to => '/srv/scratch/b' })<decision>; # allow
evaluate(@rules, 'fs_move',
{ from => '/srv/scratch/a', to => '/srv/live/b' })<decision>; # ask
evaluate(@rules, 'fs_move', { from => '/srv/scratch/a' })<decision>; # ask
The suggestion is what a permission prompt offers as its "always allow" ā the deepest directory that contains everything this call touches, never wider than the tool's configured root:
evaluate((), 'fs_read', { path => 'notes/today.md' },
roots => { fs => '/srv/docs' })<suggestion>;
# { tool => 'fs_read', under => '/srv/docs/notes' }
A call that names a website rather than a directory is offered the site it named, so that "always allow" means this site and not the web:
evaluate((), 'web_fetch', { url => 'https://Docs.Raku.org/lang.html' })<suggestion>;
# { tool => 'web_fetch', host => 'docs.raku.org' }
evaluate((), 'web_fetch', { url => 'not a url' })<suggestion>;
# { tool => 'web_fetch' } -- nothing honest to scope it to
And the host predicate, failing closed in both directions:
my @rules =
{ tool => 'web_fetch', decision => 'allow', host => '*.raku.org' },
{ tool => 'web_fetch', decision => 'deny', host => 'evil.com' },
;
evaluate(@rules, 'web_fetch', { url => 'https://docs.raku.org/x' })<decision>; # allow
evaluate(@rules, 'web_fetch', { url => 'https://raku.org/x' })<decision>; # ask
evaluate(@rules, 'web_fetch', { url => 'https://evil.com/x' })<decision>; # deny
evaluate(@rules, 'web_fetch', { url => 'javascript:x' })<decision>; # deny