← all writing
Canopy field notes

The Agentic Mesh: why coordination is the real multiplier

Parallel agents do not become a team by sharing a repository. They need identity, visible ownership, retained messages, and stalls that cannot disappear.

Running more agents is easy. Getting useful parallel work is not.

Open four terminals, start four coding agents, and the machine looks busy. But without a coordination layer, each session sees only its own prompt and its own slice of the repository. Two agents can edit the same file. A request can be injected into the wrong terminal. A useful decision can vanish in scrollback. A blocked agent can sit quietly for an hour while every dashboard still looks green.

The multiplier is not the number of agents. It is the quality of the system between them.

Identity before orchestration

Every coordination feature depends on one question: who is acting?

Canopy gives each terminal its own bridge credential. When a terminal sends an agent message or places a claim, the sender identity comes from that credential, not from a name supplied in the request body. The credential identifies the Canopy terminal; it does not pretend to authenticate a model, provider account, or human.

That narrow guarantee matters. It means one terminal cannot simply name another terminal as the owner of a claim and then release it. It also lets Canopy fail closed: if the destination is a known shell, a run terminal, or a terminal whose role is still unknown, direct agent messaging is refused instead of typing into it and hoping.

credential-derived identity
terminal 39checkout-flowcredential / 9d38:39
signs
mesh requestfrom_pty = 39body cannot override sender
routes
terminal 47api-contractrole = agent
The terminal credential names the sender; the message body never gets that authority.

Messages need a trail

A terminal write is not a coordination record. It is just bytes.

Canopy’s mesh keeps a bounded local history of recent agent-to-agent messages. Retained messages get IDs that can be used for retrieval and replies. A message can carry a multiline body, local file references, a pointer to the message it answers, and a queryable {kind, id} tag such as a task, PR, or attempt.

Those tags are useful metadata, not verified foreign keys. Reply pointers are not a full threaded-conversation engine. A successful terminal write is not a read receipt. The distinction is important because honest systems say exactly what they observed.

The store is intentionally bounded: up to 500 recent messages, with seven-day retention. It survives app restarts when local persistence succeeds, without turning the IDE into an unbounded archive.

a retained exchange
m41 · checkout-flow

The webhook contract is stable. I left the generated client untouched.

ref / attempt / run_8e2
m42 · replying to m41

Confirmed. I am taking the handler and its tests.

submitted / terminal 47
Messages carry a stable local trail: body, reply pointer, work reference, and submission state.

Shared checkouts need visible ownership

Git is excellent at reconciling commits. It does not stop two live agents from rewriting the same working-tree file at the same time.

Canopy uses advisory claims on files and directories. A claim says, “this terminal is working here.” If another agent asks for an overlapping path, the request is refused and the conflict becomes visible: who asked, which paths overlapped, and when.

Claims are not filesystem locks. They do not enforce line ranges or block an editor write. They are coordination signals, keyed to the terminal owner’s credential and released when that terminal exits. That makes collision risk visible before it becomes a mystery diff.

shared checkout / visible ownership
src/payments/checkout-flow3 files
src/api/webhook.tsapi-contract1 file
overlap refusedmobile-polish requested payment.tsheld by checkout-flow
Claims do not lock the filesystem. They make the collision visible before the edits overlap.

Silence is a state, not success

Agent systems often optimize the happy path and ignore quiet failure.

Canopy watches lifecycle evidence from agent hooks, terminal activity, and the running process. Today, two watchdog rules turn specific silence into attention:

  • W1: a session that has remained unexpectedly quiet for five minutes becomes a persistent alert.
  • W2: an already-detected human block that has gone unseen for two minutes is escalated.

The watchdog makes the stall impossible to miss, classifies the failure from the evidence around it, and starts the cheapest safe recovery rung. A human-required question remains a human-required question; recovery never routes around it.

Recovery needs portable state

Identity, retained messages, claims, and stall detection make route recovery safe enough to use.

Canopy keeps the durable goal, acceptance criteria, baseline, evidence, and attempt history in a TaskEnvelope. Failure classification separates transient, route, task, human-required, and safety-stop conditions. Fleet readiness excludes routes that are unavailable, unhealthy, or outside policy before the scorecard ranks the rest.

The recovery ladder is explicit and bounded: retry the same route when evidence says the failure was transient; choose another eligible profile or model tier when policy allows; reseed a replacement attempt from portable task state; escalate to a human when switching routes will not solve the task.

That system only works because the substrate is honest. A route switch without a durable task identity loses the job. A resurrection without a safe baseline can absorb unrelated edits. A “verified” result based on the agent’s own confidence is not verification.

bounded recovery ladder
01detect

heartbeat, lifecycle, provider, and terminal evidence

→
02classify

transient, route, task, human-required, safety stop

→
03recover

retry, eligible profile, model tier, cross-CLI reseed, human

A replacement attempt inherits the task, safe baseline, completed milestones, and required verification.

A mesh is a set of rules

The mesh graph connects lead agents, bounded domains, tasks, attempts, claims, messages, replies, evidence, and recovery edges. PR follow-ups return to the recorded conversation when it remains resumable or start a replacement attempt with the PR context.

That is already more powerful than four isolated terminals because coordination is inspectable.

The point is not to make the system look autonomous. The point is to make parallel work safer, legible, and recoverable: each agent leads a bounded domain, every job attaches to durable work, and recovery follows evidence instead of guesswork.

Explore the mesh that ships today, see the Engineer workspace, or read the Build mode architecture.