Troubleshooting
Start with the system's own facts:
autodev board check
autodev fleet status
autodev fleet doctor
autodev fleet disk
autodev retro
Add --output json to preserve typed fields when collecting a report.
For a registered workspace, also inspect the central owner:
autodev daemon status
autodev workspace list
Initialization is refused or keeps existing files
autodev init requires an existing Git repository because board mutations commit and planning
lifetimes come from Git history. Run git init first when necessary. Init is convergent: it creates
missing planning directories, a commented .autodev/fleet.toml, and runtime-state ignore rules, but
keeps existing files byte-for-byte. Use --no-commit when you want to commit the scaffold yourself.
A new story cannot become ready
Run:
autodev board get APP-42
autodev board explain APP-42 --to ready
autodev board check
Common blocking defects are an empty goal, missing or vague acceptance criteria, [NEEDS CLARIFICATION] markers, duplicate criterion numbers, unknown status, unresolved dependencies, or dangling epic/design/glossary references. board create intentionally leaves clarification markers when goal or criteria are absent.
A story is ready but is not dispatched
Read the withholding reason from fleet tick or fleet status. Typical reasons include:
- the fleet is paused;
- dependencies are not done;
- another live or parked wave holds the story;
- declared areas overlap live work;
- the open-wave or concurrency limit is reached;
- free space is below
disk_floor_gb; - the board contract has a blocking defect;
- the base branch gate or fleet configuration is invalid.
Do not “fix” this by editing SQLite. Resolve the named reason.
Fleet configuration is refused
The fleet.toml parser rejects unknown keys and reports all detected configuration defects.
workflows.toml denies unknown keys too, at file, workflow, and edge level, and names the offending
key. Check that:
- every bridge
kindis one ofcodex,claude-code, orcommand; - a command bridge has a non-empty
argv; - every named role points to a declared bridge;
- a configured worker has a non-empty
validation.argv; - role harnesses enforce confinement, or the driver explicitly accepted the exception.
Run autodev fleet tick after editing; all fleet readers use the same validated configuration path, so an invalid file can also prevent status, disk, or usage commands from opening it.
Validation says Bubblewrap is unavailable
On Linux, install bwrap and make sure it is on the driver's PATH. Containers may have the binary but lack kernel/capability support to start its sandbox.
AUTODEV_ALLOW_UNCONFINED_VALIDATION=1 bypasses filesystem/network confinement and leaves only the restricted environment. That is a risk acceptance, not a general installation fix.
A harness cannot authenticate or find a tool
Worker environments are cleared except for PATH, HOME, and explicitly configured values. Confirm the harness is authenticated under the same HOME as the driver. Add only necessary names under env.pass_through or explicit values under env.set.
Avoid passing secrets merely to make a broad environment work. Harness and tool output may be stored in transcripts.
For kind = "command", also confirm the CLI accepts the prompt as an argument: autodev substitutes
{prompt} in argv and closes stdin. A generic harness that needs writable session state outside
the worktree can fail under confinement = "fleet"; prefer its ephemeral mode or a reviewed wrapper
over accepting an unconfined role. See Connect coding-agent CLIs.
The integrated gate is red
The fleet runs the configured gate against the integrated candidate and retries one red run to distinguish a flaky flip from a repeatable failure. Inspect the wave events and transcript, reproduce the exact argument array in a clean checkout, and correct the repository or story work.
If validation.gate is absent, the gate is validation.argv. A command that passes only in a developer's ambient environment may fail because fleet validation is deliberately default-deny and network-isolated.
A registered workspace is not using the daemon
Run autodev daemon status and autodev workspace list. The accessor uses the daemon only when the
recorded endpoint answers and the canonical repository root is registered. A registered workspace
whose endpoint is dead reports the daemon-owned store as unreachable; it does not fall back to a
checkout-local .autodev/fleet.sqlite and silently present different history.
Restore and verify the daemon before running a store-backed command. Git-only board inspection may still read planning files directly, but checks or transitions that require runtime evidence must observe the daemon-owned history.
If a legacy store is reported as unimported, carry it across explicitly:
autodev workspace import .
Do not copy the SQLite file into the daemon directory by hand. Import verifies the record counts and records what happened to the legacy path.
The browser UI is unavailable or read-only
One daemon process owns two listeners. The machine API on 7788 does not serve HTML; the embedded
operator client is on 7777. Start the daemon once, verify it, then open the client URL:
autodev daemon start --single-user
autodev daemon status
scripts/open-autodev-brave.sh
--single-user grants act on loopback. Without it or an act token, the browser is read-only by
design. If an act token was configured, present it to the browser and check the role indicator
before retrying a mutation. Both listeners refuse unsafe assumptions about exposure; do not work
around the loopback boundary with a public reverse proxy.
If Brave stays CPU-bound, inspect the dedicated process tree before killing a child:
ps -eo pid,ppid,pcpu,args | grep '[b]rave.*\(gpu-process\|remote-debugging\)'
curl --fail --silent http://127.0.0.1:9333/json/list
A gpu-process command line containing --use-gl=disabled means software
rasterization can be consuming CPU despite the process name. Killing that
child is temporary because Brave recreates it. The repository launcher avoids
the shared-profile access pattern, does not disable GPU acceleration, uses an
isolated profile, and binds the debugging endpoint to loopback. Override its
port with AUTODEV_BRAVE_DEBUG_PORT when 9333 is already occupied.
Your checkout is behind the branch the fleet landed on
A dirty working tree no longer blocks anything (A-151). Apply stages its merge in a fleet-owned
checkout under worktree_root, runs the post-merge gate there, and advances the base ref; your
checkout is a consumer.
After a landing the fleet tries to bring your checkout forward without touching your work, updating only the files the landing changed and keeping every uncommitted edit that does not collide with one. When an edit does collide it refuses rather than overwriting, and the tick says so with the command that reconciles it when you are ready:
git stash && git reset --hard <landed-sha> && git stash pop
Until you run it, your uncommitted work is exactly where you left it and the delivery is already on the ref. One consequence worth knowing: the board transition for that delivery is owed rather than attempted while the checkout is behind, because the hook that writes it commits in your checkout. The next tick retries it once the tree can move.
A wave is stuck or malformed
autodev fleet status
autodev fleet doctor
autodev retro
Pause new dispatch, then park a wave while investigating:
autodev fleet park wave-12 --reason "investigating malformed state"
Parking keeps its stories held. A running tick no longer holds the repository write lock across its turns, so parking does not need the drive stopped first; it waits only on the short sections a delivery or an apply takes.
Use fleet unpark when the cause is corrected, or fleet retire <wave> --reason <text> to terminate
it and release its stories. There is no fleet cancel command.
Unpark can answer in two ways. A wave re-earns the areas it claims before it re-enters the loop, so if another live wave now holds one of them, the wave stays parked and the answer names the area and the holding wave. Nothing is written in that case. Wait for the holding wave to close, or retire one of them. Note that this deferral is reported only in human output; see Current limitations.
Disk dispatch is withheld
Run autodev fleet disk and inspect the filesystem containing worktree_root. Terminal and crash-left worktrees are reclaimed by later ticks where possible, but resident worktrees can each hold a build cache. Reduce concurrency/open-wave limits, move the root to an appropriately sized filesystem, or raise the floor only when measurement supports it.
Never lower the floor just to silence the refusal: a build that runs the disk out mid-gate cannot be distinguished from a real red gate by wishful configuration.
A second drive or board write reports a lock
Only one drive may own a repository. Board and fleet mutations wait a bounded time for the repository
write lock — 120s for board, 10s for fleet, or whatever --lock-wait-ms sets. A tick does not
take that lock, so the holder you collide with is normally another operator verb, or the short
section a delivery or apply holds. Read-only diagnosis never takes it at all.
Check whether the named holder process is alive and whether another drive is intentionally running. A restarted drive can recover a stale owner whose process is gone. Avoid manually deleting lock files as a first response because a live holder and a stale record require different actions.
For a registered workspace, exclusion is owned by the daemon rather than competing SQLite openers.
A busy or unavailable daemon should still be diagnosed through daemon status, the typed refusal,
and the workspace registry—not by deleting repository lock markers.
Metrics show unknown
unknown means the necessary fact was not recorded. Examples include a story with no committed history, a harness with no usage report, an unpriced model/rate, or no measured delivery rate for a forecast. It is not equivalent to zero and should remain distinct in exported dashboards.
The drive was interrupted
Do not immediately start editing worktrees. Restart with fleet status, fleet doctor, and one
fleet tick. The reconciler recovers ordinary dead workers and stale resources, but missing process
groups and refused kills can still require operator intervention. Use park, unpark, or retire
with a reason so the recovery remains on the record.