Delivery lifecycle
The fleet is a reconciler. Each tick reads durable state, performs the work that is currently safe,
writes the result, and reports why anything could not advance. fleet drive repeats that process
until the fleet is quiet or reaches its configured tick limit — or, with --loop, until it is
stopped.
The wave is the unit that moves
A wave carries several stories and advances as one. Understanding that is the difference between reading a stalled fleet correctly and guessing at it, because a wave holds at whichever state its slowest or weakest story leaves it in.
┌───────────── one wave, `max_wave_size` stories ───────────┐
dispatched ──▶ awaiting-handoffs ──▶ handoffs-ready ──▶ reviewing ──▶ integrating ──▶ gated ──▶ applied
│ │ │ │ │
│ every story one judge per rebase onto the whole
│ needs a turn story, up to 7 the branch post-merge
│ that completed at once tip, 1 at tree runs
│ │ │ a time the gate
└── a turn that fails and has attempts left comes back here ────┘
Three properties follow, and each one explains something an operator sees:
Every story must land a completed turn before any of them move. The wave sits
at awaiting-handoffs until the last one arrives, so a single slow story holds
four finished ones with it.
Review is per story, not per wave. Each handoff gets its own independent judge verdict, and up to seven run concurrently. There is no second, joint review of the wave as a whole.
Integration is per wave, and all-or-nothing. A wave whose stories were reviewed separately still integrates together, so one rejection strands every acceptance beside it. The attempt budget is the wave's too: one story's rejection spends an attempt the others were relying on.
Nothing is destroyed when that happens. Every attempt lives on its own
impl/<wave>-worker-<n> branch and stays reachable after the wave is retired,
so accepted work is re-dispatchable in a fresh wave rather than lost. But it does
have to go round again, which is why wave size is a real trade: a larger wave
amortises review and integration over more stories, and enlarges the unit of
loss when one of them fails.
[fleet] max_wave_size = 1 removes the trade entirely, and is what this
repository runs. At one story per wave a handoff is a wave: there is no
sibling to wait for, one rejection strands nothing, and the three properties
above collapse into "each story reviews, integrates and lands on its own". Read
the section above as what a larger wave costs, not as what you are running,
unless you raised the number deliberately.
Where the queue actually forms
At size 1 the coupling is not between stories in a wave, it is between waves at the integrate and apply lanes — both capped at one, because both move the shared branch tip and two rebases cannot hold it at once. So a project running seven concurrent turns funnels them through a one-wide integration:
turns 7 concurrent ─┐
reviews 7 concurrent ─┤
integrations 1 concurrent ─┴─▶ the serial section
applies 1 concurrent
That is the constraint to reason about when work is accepted and not landing. It is a correctness bound rather than an arbitrary limit, and widening it means changing what holds the tip, not raising a number.
What did change (2026-08-12) is how often the serial section is offered
work. The integration lane used to read the wave list once per tick, so a
review that closed mid-tick waited a whole tick to integrate and another to
apply — and a tick lasts as long as its longest running turn, so that tail
measured 20 to 70 minutes. The progression pass now offers integration
and apply whatever became eligible while the tick's turns and reviews still
run, and once more after they drain: a passing verdict reaches applied in
the tick that heard it. Still one at a time — the pass changes how often the
lanes are offered work, never how many run at once.
What triggers each step
Nothing is event-driven. A single loop — the tick, every few seconds inside
autodev daemon start — walks every wave and runs four lanes, each with its own
concurrency ceiling. Two of them are offered work continuously within the tick:
admission keeps asking what became dispatchable while turns run, and the
progression pass keeps asking what became integrable or appliable — so the
lanes below describe ceilings, not phases:
turns 7 concurrent a dispatched wave with a spawned worker
reviews 7 concurrent a handoff waiting on a verdict
integrations 1 concurrent a reviewed wave, rebasing onto the branch tip
applies 1 concurrent a gated-green wave, or one that owes the board
Integration and apply are capped at one because both move the shared branch tip, and two waves cannot hold it at once.
Every tick prints all four lanes whether or not they did anything, and an idle lane names the condition it was waiting for:
turns peak 0 of 7 — idle: no wave is dispatched or awaiting handoffs with a spawned worker
reviews peak 0 of 7 — idle: no wave is handoffs waiting on a verdict
integrations peak 0 of 1 — idle: no wave is integrating
applies peak 0 of 1 — idle: no wave is gated green or owing the board
That is the answer to "why has nothing integrated yet", in the log, at the moment it did not happen — rather than a state an operator has to reconstruct afterwards.
1. Make the story dispatchable
Before dispatch, the board checks that a story can be represented and acted on. The story needs a usable goal and acceptance criteria, a state recognized by its workflow, and satisfied dependency requirements. Unresolved clarification markers and other contract defects keep work out of the fleet with a reason.
Readiness is part of the delivery contract, not a label an agent is expected to interpret.
2. Dispatch a pinned wave
The scheduler selects eligible work subject to dependencies, live claims, configured capacity, and
resource preflight. Only typed facts withhold a story: the fleet's own
delivery record, the board's status, a live claim, an unsatisfied dependency, a
named ceiling. What a commit message says — a Refs: trailer naming the
story on the base branch — refuses nothing: it files one standing task
("close the story or record a correction") and the story dispatches meanwhile.
A missed landing costs one worker turn; an invented one used to cost the story
for ever. It creates a wave with:
- the exact commit behind the base ref;
- a snapshot of each selected story's execution contract;
- an attempt budget; and
- the areas claimed by those stories.
Each story is assigned a worker in its own git worktree. That isolation lets worker turns overlap without sharing a mutable checkout.
3. Run and observe the worker turn
The selected bridge launches the agent harness with the story contract and repository validation command. While the turn runs, autodev streams its transcript and records the process it owns, so a killed or orphaned turn does not disappear.
When the harness exits, autodev compares git before and after the turn. The resulting receipt names
the commits that actually appeared, the measured wall time, and either per-model usage reported by
the harness or an explicit unknown value. A successful-sounding transcript is not a commit.
4. Build a verified handoff
For a candidate commit, the host derives the write set from git and evaluates failing-first evidence in throwaway verification worktrees:
- place the candidate's test on the pinned base without carrying the fix with it;
- run the configured validation and require a meaningful red result;
- run validation at the candidate and require green; and
- compare observed writes with the claims held by other live waves.
Step 2 has a third answer besides red and green. If every file that carried a test onto the base is a file the base commit does not contain — a story that creates a new crate, for instance — then the base could hold no opinion at all: there was no red to be had, and a green there proves nothing. autodev reads that from git, before the run, rather than from anything the worker said about it, and treats it as an absence of a failing-first proof rather than a failed one. The work is still accepted on the strength of the same tests passing at the candidate that creates them, and the acceptance record says which ground it stands on. The older case — tests placed in a compilation unit the base does have, where nothing compiles — is still a refusal.
If those checks hold, autodev records a handoff containing the commit, write set, evidence, and any worker note. Otherwise it refuses the handoff with actionable defects and can reopen the story for a bounded retry.
A retry is briefed with everything the fleet already paid to learn: every prior attempt's refusals, judge rejections and worker notes — across every wave the story has run in, not only the current one. A wave's attempt budget is the wave's; a story's history is the story's, and a cancelled wave does not erase it. The brief is byte-bounded and drops whole oldest attempts first, saying how much it dropped.
5. Review independently
The judge receives the story contract and candidate diff in a separate review turn. If the diff cannot be read, review holds rather than asking the judge to decide blind. The judge must finish its own turn and produce a parseable verdict; a killed process cannot leave behind a usable verdict. A refusal records its reason and returns eligible work to another attempt while the budget permits.
“Independent” describes the turn's role and context: the reviewer is not the worker turn whose work it evaluates. Worker and judge roles can be assigned to different configured bridges.
A rejection is not discarded: it becomes part of the story's retry brief (step 4), so the next attempt — in this wave or a later one — starts from the criticism instead of rediscovering it.
6. Integrate on the current tip
Accepted handoffs are integrated in a throwaway integration worktree. autodev first attempts a normal merge. If that conflicts, it tries to replay the handoff's commits onto the current tip before declaring a real contest. The resulting tree runs the repository gate.
When a multi-story batch turns the gate red, candidates can be attributed and excluded individually; one bad handoff does not have to discard every clean candidate in the wave.
Integration is intentionally serial even when worker and review turns overlap. Every integration must include the previously accepted tip, or two apparently green branches could omit each other.
7. Gate the tree that will remain
A green integration branch is merged into the local base branch — in the fleet's own checkout under
worktree_root, not in yours. autodev then runs the repository gate again on that post-merge tree.
This second gate matters because the base may have moved since integration. A red or un-runnable gate
restores the pre-merge tree and records why the change did not remain.
Only when that gate holds does autodev advance the base ref, and it does so as a compare-and-swap against the tip it started from. Where the base branch stands is a repository fact, independent of which checkout has it or what state that checkout is in, so an operator's uncommitted work does not hold a delivery. After the ref moves, the fleet tries to bring your checkout forward without touching that work, updating only the files the landing changed; when a local edit collides it refuses rather than overwriting, and says which command reconciles it when you are ready.
Only after the gate holds does autodev write delivery records. It then advances the planning board. If the board update fails after the commit landed — including when your checkout could not be moved, because the board commit is written there — the wave enters a typed planning-pending state and retries the owed update instead of delivering the same work twice.
The intended result is a chain from intent to outcome in which every boundary has an artifact. See Trust and evidence for the rules behind that chain and Current limitations for gaps in the current release.