Development workflow
How work actually flows through a fleet of agents.
Running several agents at once is a workflow problem before it is a tooling problem. This is the loop that works in practice; nothing here is enforced by the product, it is how to use it.
The loop
Leader -> Worker tasking (the ticket is claimed first; the worker owns the checkout)
Worker -> Leader design (what it will change, where, the failing test)
Leader -> Worker approval (or revision turns until approved)
Worker -> Leader completion (branch + SHA, gates run, what was NOT done)
Leader -> Reviewer review kickoff (a FRESH agent, read-only, on the diff)
Reviewer -> Leader findings (ranked, most severe first, a concrete failure case each)
Leader -> Worker relay (the reviewer never talks to the worker)
Worker -> ship build, deploy, verify by effect
Why it is shaped that way
- One agent per checkout. Two agents editing one working tree steal each other's uncommitted work. Give each a checkout, or hand the checkout over explicitly.
- The reviewer is a fresh agent, every time. Never the author reviewing itself, and never the one who approved the design — it shares the blind spot. The best review outcome is usually a deletion from the plan.
- Size the model to the task. A mechanical sweep does not need your most expensive model; a security change does.
- Verify by effect. "The code is deployed" is not a result. Measure the thing the change was supposed to do, on the live system, and record what you measured.
- The handoff sentence is the one nobody reviews. Make it name every behaviour it claims, and say plainly what was skipped.
What each message carries
Tasking: the ticket, the checkout and branch, the gate to run, what "done" looks like as an observable effect, and the constraints that actually bite. Design: files touched, the failing test first, what it deliberately does not do. Completion: branch and SHA, gates with their real exit codes, anything skipped.