Leases

NAME

MCP::Client::Leases - advisory file leases for concurrent agents over one workspace

SYNOPSIS

use MCP::Client::Leases;
use MCP::Client::Leases::Table;
use MCP::Client::Policy;

# One table for the workspace, made by whatever owns the agents.
my $table = MCP::Client::Leases::Table.new;

# One composer per agent, under the policy: permission is resolved before a
# lease is ever consumed.
sub provider-for(Str:D $agent-id) {
	MCP::Client::Policy.new(
		provider => MCP::Client::Leases.new(
			inner      => $registry,
			table      => $table,
			agent-id   => $agent-id,
			roots      => { fs => '/srv/work', lock => '/srv/work' },
			concurrent => { $engine.live-agents },      # > 1 turns strict mode on
		),
		rules => MCP::Client::Policy.default-rules,
		:&on-ask,
	);
}

# The model's side of it:
#   lock_acquire { "paths": ["src/app.raku"], "ttl": 600 }
#   fs_edit      { "path": "src/app.raku", ... }
#   lock_release { "paths": ["src/app.raku"] }

DESCRIPTION

Two agents editing one checkout will, eventually, edit one file. Not at the same instant — the calls are serialized by the server — but across the window that matters: read the file, think about it, write it back. The second agent's read happened before the first agent's write, and its write throws that away. Nothing in the permission layer notices, because both calls were things their agent was perfectly entitled to do.

MCP::Client::Leases is the smallest thing that closes that window: an in-process, advisory, bounded-wait lock table with two tools in front of it. An agent about to work on part of the tree says so; another agent trying to change the same part is told who has it and for how long. It is a provider composer, satisfying the same duck type MCP::Client::Registry and MCP::Client::Policy do — tools-for-llm plus execute-tool-calls — so it stacks in wherever a provider goes.

What a lease is, and is not

A lease is advisory. It is not a filesystem lock, it does not survive the process, and nothing outside this table honours it. It is a claim one agent makes and the other agents' tool layers respect, which is enough precisely because every agent in the pack goes through this layer.

The single-call race is already handled elsewhere and needs no lease: fs_edit's exact-match old-string is optimistic concurrency control — the edit lands only if the text it expected is still there, so a call that lost the race fails loudly instead of clobbering. What a lease protects is the multi-step window fs_edit cannot see: read, decide, write; or the four edits that only make sense applied together; or the rename that has to follow the change to the file it renames.

Reads are unrestricted in v1. There are no read-locks, and a lease does not stop anybody reading what it covers — which means a non-locking reader can still read a file halfway through somebody else's multi-step change and act on what it saw. If that matters for a particular flow, have the reader take the lease too: lock_acquire then read then lock_release gives a reader the same window a writer gets, at the cost of contending for it.

Acquisition waits, but never for ever

lock_acquire tries at once, and if the path is taken it waits — up to wait seconds, polling the table about once a second — before it gives up and returns the refusal. wait defaults to 90 seconds and is capped at the table's default-ttl (300 out of the box).

This is a change of mind, and the economics are the whole of the argument. The original design refused immediately, on the grounds that a bounded thing that never waits cannot deadlock. It cannot — but it does not stop anybody waiting either; it moves the waiting into the model, which is the most expensive place in the system to put it. A refused agent re-reads its instructions, thinks about what to do instead, and calls lock_acquire again — and every one of those laps re-sends the entire conversation to a language model. Two agents sharing one file can burn hundreds of thousands of prompt tokens taking turns at a lock that was free after four seconds. Polling a hash inside the process costs nothing at all.

What keeps it safe is that the wait is bounded twice over. A wait is at most wait seconds, and every lease expires on its own at its TTL. More importantly, an agent that already holds any lease is never allowed to park on another: contention fails immediately and teaches it to release, deduplicate, and atomically acquire its complete working set. The Table also registers each real waiter and rejects a cycle of any length as defence in depth. Registration, blocker discovery, cycle detection and grant are one locked transition; no thread sleeps while the Table lock is held.

wait: 0 is the original behaviour exactly: attempt once, return the refusal, never sleep. Nothing else about the refusal changed — same wording, same holder-and-held-for shape — because a model that has learned to read it should not have to learn again just because it arrived a minute later.

There is no queue and no fairness. A lease that comes free is taken by whoever's poll lands first, which may not be whoever has waited longest, and a third agent that never waited at all can win it by asking at the right microsecond. A fair queue is a much bigger object — it has to survive holders that die, waiters that are cancelled, and a released lock nobody is left to take — and none of it is needed for "two agents want one file". The refusal at the deadline says who is holding it, which is enough to work with.

The refusal is still the interesting message, and it is still written to be acted on: it names the holder and how long they have held it, so the model can work on something else and come back, or say who has the file. The tool description teaches that, and teaches that the call may take a while.

Stale-write fence and provider pins

With MCP::Server::Tool::FileSystem, successful read, stat and list results carry a hidden sha256-v1 revision. The layer caches it and adds the reserved _expected-revisions argument to the next fs_write, fs_edit, fs_mkdir, fs_move or fs_delete. The filesystem provider compares and mutates under one lock. A changed target is refused before any side effect; success advances the cache and failure invalidates it. A target never observed by this agent is expected to be absent, so a blind overwrite fails closed.

Before forwarding a mutation, the exact covering lease generation is pinned. Expiry and explicit release become pending until the provider returns, so the same path cannot be granted to another agent while the original write is still landing. Table.status reports leases, waiters and pins together.

Waiting is a suspension, and hosts want to know

A batch's lock calls settle first (see /Never throws), so an acquire that waits delays the rest of its own batch. That is correct and not a regression: the calls behind it are the edit the lease was taken to protect, and running them first is precisely the race the lease exists to close.

It does mean an agent can be parked inside a tool call for a minute or more, which a host with a fleet of them very much wants to know about — an agent waiting on a lock is not working, and if the host caps concurrency by counting working agents, a waiter should not be counted. Hence two optional callbacks:

  • on-wait-begin — called once, when a wait really begins. An immediate grant never fires it, and neither does wait: 0.

  • on-wait-end — called once, on every exit from the wait: the grant, the deadline, a cancel, or an exception on the way out.

Neither takes an argument: a composer belongs to one agent and carries its agent-id, so the host already knows who it is about. Both are shielded — a hook that throws is swallowed, because a broken bit of host bookkeeping must not turn an acquire that was granted into an error the model has to interpret.

Cancelling a wait

The optional cancelled thunk is asked, once per poll, whether anybody is still waiting for this call's answer. When it answers true the wait stops at that tick and the standard refusal comes back with a sentence noting that the wait was cut short.

The question is deliberately that one rather than "has somebody pressed cancel", because the two come apart at the moment that matters. A host whose run was cancelled typically settles it and stops waiting for the tool call while the call is still in flight — the call is detached, its answer has nowhere to go — so a thunk that only reports "cancelled" goes false again the instant the run finishes, and the wait it should have ended carries on for its full deadline holding a thread (and possibly taking a lease nobody will release). "Is there still a live, uncancelled run behind this call?" is the shape that closes both cases.

It fails open: a thunk that throws is read as "not cancelled" and the wait runs its course. That is the opposite of the concurrent thunk two sections down, deliberately: over-waiting costs seconds of wall clock, while a falsely cancelled acquire hands back a refusal for a lock that was there for the taking — and the agent then edits a file it does not hold, or gives up on work it could have done. When the two failure modes are "slow" and "wrong", the broken signal should choose slow.

Strict when concurrent

By default this layer only enforces other agents' leases: a call that would change a path somebody else has claimed comes back as an is_error naming them, and everything else goes through untouched. A solo agent therefore never sees the lease layer at all, which is the point — locking discipline is pure overhead when there is nobody to contend with.

When the concurrent thunk says more than one agent is live, the layer switches to strict: a mutating call must be covered by a lease this agent holds, or it is refused with an error that teaches the fix ("another agent is live; lock_acquire first"). Strictness relaxes on its own when the last sibling drains.

The thunk fails closed. If it throws, or answers with something that cannot be read as a count, the layer assumes concurrency and goes strict. This is the opposite of the usual "broken signal → feature inactive" reflex, and deliberately so: the failure mode of over-strictness is an agent that asks for locks it did not need, and the failure mode of under-strictness is two agents silently overwriting each other's work. One of those is visible in the transcript and the other is not.

Where in the stack

Stack it under the policy — Policy(Leases(Registry)):

  • permission resolves before a lease is consumed, so a call the human refuses never takes or checks a lock, and the lease layer's error is never the reason a permission prompt did not appear;

  • it is also under any escalation or retry layer, so a re-run of a refused call cannot walk around the lease by being asked a second time;

  • and it means the lease layer never has to know a policy exists. It has no ask surface, calls nothing that could open one, and holds no lock while calling anybody else's code.

That placement is also why this class does not delegate interactive, grants, elicit-hook or anything else policy-shaped to its inner provider, and why it needs no .can('grants') subclass the way LLM::Agent::Subagents grew one. A composer that sits over a policy has to keep the policy's own surface reachable through it, or a host that introspects the provider loses the human seam; a composer that sits under one has the policy on the outside already, where the host is looking. If you invert the stack you take that on yourself.

Rule names, roots keys and path-params patterns are read at this layer's position, exactly as a policy's are: under a registry a call arrives with the prefix already stripped, over one it has not been. Getting it wrong is quiet — no path is ever located, so nothing is ever enforced.

The published tools

Two, published without any registry prefix, the way LLM::Agent::Subagents publishes task:

  • lock_acquire(paths, ttl?, wait?) — claim paths, all or nothing, waiting out a contended one for up to wait seconds;

  • lock_release(paths?) — give them back; bare, everything you hold.

An inner provider that publishes either name has its declaration dropped from the catalogue rather than published twice — this composer routes on those names, so publishing somebody else's version of them would be a lie. That holds even with < publish-tools => False >, the solo-only configuration where the two tools are hidden from the model entirely.

Never throws

execute-tool-calls keeps the provider contract exactly: a refusal, a malformed call, a lock that could not be taken, an inner provider that died — each is one is_error result in the caller's own order, and the calls around it are unaffected. tools-for-llm keeps the deliberate asymmetry: an inner provider that cannot list its tools throws, because publishing a silently shorter catalogue leaves a model wondering where a capability went.

Within a mixed batch the lock calls settle first, in order, before any other call is judged. So lock_acquire and the edit it protects may travel in one batch, which is how a model that has just been told to lock first will write it.

Results from the inner provider are passed through as they came, with one exception: when the layer could not judge a path and is not being strict, the result gets a leases-warning key saying so (only if it was an object — a result of some other shape is never reshaped to carry a warning).

EXAMPLES

The solo case, where the layer costs nothing. No concurrent thunk means never strict, so an unlocked write goes straight through and only another holder's lease can stop anything:

my $leases = MCP::Client::Leases.new(
	inner => $registry, table => $table, agent-id => 'main',
	roots => { fs => '/srv/work', lock => '/srv/work' },
);

Note the lock entry in roots. Roots are matched by tool-name prefix, and the published tools are called lock_acquire/lock_release, so a table that only knows about fs leaves the lock tools with no root to measure a relative path from — and 'src/app.raku' locked as a relative path will not be recognised as the same location as /srv/work/src/app.raku written by fs_write. Either add a lock entry pointing at the same directory the fs tools use, or give the table a '' catch-all.

Enforcement follows the tools that change things, which is configurable because not every server calls them what the MCP::Server::Tool::FileSystem pack does. Patterns are the same trailing-* globs a rule's tool is:

my $leases = MCP::Client::Leases.new(
	inner => $inner, table => $table, agent-id => 'writer-1',
	enforced-tools => <fs_write fs_edit fs_mkdir fs_move fs_delete patch_*>,
	path-params    => { patch_apply => ['target',] },
);

A fleet that counts working agents wants the suspension hooks, and a fleet with a cancel button wants the thunk. All three are plain closures over what the host already has:

my $leases = MCP::Client::Leases.new(
	inner => $registry, table => $table, agent-id => $agent-id,
	roots => %( fs => $root, lock => $root ),
	concurrent => &live-agents,
	# An agent parked on a lock is not working: give its queue slot back
	# for the duration, exactly as one parked on a question does.
	on-wait-begin => { $subagents.suspend-agent($agent-id) },
	on-wait-end   => { $subagents.resume-agent($agent-id) },
	# And stop waiting the moment this agent's run is cancelled.
	cancelled     => { $run.defined && $run.is-cancelled },
);

Releasing when an agent stops is not this layer's job — it has no idea when that happened. The engine that owns the agents does it, off the run's drained (never its result: a detached write may still be landing when the answer is already in):

$run.drained.then({ $table.release-holder($agent-id) });
$engine.on-shutdown({ $table.release-all });

SEE ALSO

MCP::Client::Policy — the permission layer this one stacks under, and the source of the roots/path-params conventions it locates paths with.

MCP::Client::Leases::Table — the lease book itself, and the plain-data API a status surface reads.

  • granted — the answer of the attempt that won, returned as-is, so a grant here is byte-identical to a grant that never waited;

  • deadline — the LAST attempt's refusal, which names whoever is holding it now rather than whoever was holding it a minute ago;

  • cancelled — the last refusal with cut-short set, at the first poll after the thunk went true;

  • no longer a conflict — a refusal for some other reason (a table that has started answering differently) ends the wait too; waiting is for contention and nothing else.

MCP::Client v0.5.0

talk to an MCP server, in either protocol era

Authors

  • Matt Doughty

License

Artistic-2.0

Dependencies

MCP::Server:ver<0.6.0+>:auth<zef:apogee>JSON::Fast:ver<0.19+>:auth<cpan:TIMOTIMO>Cro::HTTP:ver<0.8.11+>:auth<zef:cro>:api<0>MIME::Base64:ver<1.2.5+>:auth<zef:raku-community-modules>

Test Dependencies

Provides

  • MCP::Client
  • MCP::Client::Cache
  • MCP::Client::Correlator
  • MCP::Client::Exceptions
  • MCP::Client::Leases
  • MCP::Client::Leases::Table
  • MCP::Client::Policy
  • MCP::Client::Policy::Commands
  • MCP::Client::Policy::Floor
  • MCP::Client::Policy::Grants
  • MCP::Client::Policy::Rules
  • MCP::Client::Protocol
  • MCP::Client::Reasons
  • MCP::Client::Registry
  • MCP::Client::SSE
  • MCP::Client::Transport
  • MCP::Client::Transport::HTTP
  • MCP::Client::Transport::Stdio
  • MCP::Client::UnknownKeys

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite — the markup and publishing tools behind this site.