Skip to main content

Troubleshooting

Start with the system's own facts:

autodev board check
autodev fleet status
autodev fleet doctor
autodev fleet disk
autodev retro

Add --output json to preserve typed fields when collecting a report.

For a registered workspace, also inspect the central owner:

autodev daemon status
autodev workspace list

Initialization is refused or keeps existing files​

autodev init requires an existing Git repository because board mutations commit and planning lifetimes come from Git history. Run git init first when necessary. Init is convergent: it creates missing planning directories, a commented .autodev/fleet.toml, and runtime-state ignore rules, but keeps existing files byte-for-byte. Use --no-commit when you want to commit the scaffold yourself.

A new story cannot become ready​

Run:

autodev board get APP-42
autodev board explain APP-42 --to ready
autodev board check

Common blocking defects are an empty goal, missing or vague acceptance criteria, [NEEDS CLARIFICATION] markers, duplicate criterion numbers, unknown status, unresolved dependencies, or dangling epic/design/glossary references. board create intentionally leaves clarification markers when goal or criteria are absent.

A story is ready but is not dispatched​

Read the withholding reason from fleet tick or fleet status. Typical reasons include:

  • the fleet is paused;
  • dependencies are not done;
  • another live or parked wave holds the story;
  • declared areas overlap live work;
  • the open-wave or concurrency limit is reached;
  • free space is below disk_floor_gb;
  • the board contract has a blocking defect;
  • the base branch gate or fleet configuration is invalid.

Do not “fix” this by editing SQLite. Resolve the named reason.

Fleet configuration is refused​

The fleet.toml parser rejects unknown keys and reports all detected configuration defects. workflows.toml denies unknown keys too, at file, workflow, and edge level, and names the offending key. Check that:

  • every bridge kind is one of codex, claude-code, or command;
  • a command bridge has a non-empty argv;
  • every named role points to a declared bridge;
  • a configured worker has a non-empty validation.argv;
  • role harnesses enforce confinement, or the driver explicitly accepted the exception.

Run autodev fleet tick after editing; all fleet readers use the same validated configuration path, so an invalid file can also prevent status, disk, or usage commands from opening it.

Validation says Bubblewrap is unavailable​

On Linux, install bwrap and make sure it is on the driver's PATH. Containers may have the binary but lack kernel/capability support to start its sandbox.

AUTODEV_ALLOW_UNCONFINED_VALIDATION=1 bypasses filesystem/network confinement and leaves only the restricted environment. That is a risk acceptance, not a general installation fix.

A harness cannot authenticate or find a tool​

Worker environments are cleared except for PATH, HOME, and explicitly configured values. Confirm the harness is authenticated under the same HOME as the driver. Add only necessary names under env.pass_through or explicit values under env.set.

Avoid passing secrets merely to make a broad environment work. Harness and tool output may be stored in transcripts.

For kind = "command", also confirm the CLI accepts the prompt as an argument: autodev substitutes {prompt} in argv and closes stdin. A generic harness that needs writable session state outside the worktree can fail under confinement = "fleet"; prefer its ephemeral mode or a reviewed wrapper over accepting an unconfined role. See Connect coding-agent CLIs.

The integrated gate is red​

The fleet runs the configured gate against the integrated candidate and retries one red run to distinguish a flaky flip from a repeatable failure. Inspect the wave events and transcript, reproduce the exact argument array in a clean checkout, and correct the repository or story work.

If validation.gate is absent, the gate is validation.argv. A command that passes only in a developer's ambient environment may fail because fleet validation is deliberately default-deny and network-isolated.

A registered workspace is not using the daemon​

Run autodev daemon status and autodev workspace list. The accessor uses the daemon only when the recorded endpoint answers and the canonical repository root is registered. A registered workspace whose endpoint is dead reports the daemon-owned store as unreachable; it does not fall back to a checkout-local .autodev/fleet.sqlite and silently present different history.

Restore and verify the daemon before running a store-backed command. Git-only board inspection may still read planning files directly, but checks or transitions that require runtime evidence must observe the daemon-owned history.

If a legacy store is reported as unimported, carry it across explicitly:

autodev workspace import .

Do not copy the SQLite file into the daemon directory by hand. Import verifies the record counts and records what happened to the legacy path.

The browser UI is unavailable or read-only​

One daemon process owns two listeners. The machine API on 7788 does not serve HTML; the embedded operator client is on 7777. Start the daemon once, verify it, then open the client URL:

autodev daemon start --single-user
autodev daemon status
scripts/open-autodev-brave.sh

--single-user grants act on loopback. Without it or an act token, the browser is read-only by design. If an act token was configured, present it to the browser and check the role indicator before retrying a mutation. Both listeners refuse unsafe assumptions about exposure; do not work around the loopback boundary with a public reverse proxy.

If Brave stays CPU-bound, inspect the dedicated process tree before killing a child:

ps -eo pid,ppid,pcpu,args | grep '[b]rave.*\(gpu-process\|remote-debugging\)'
curl --fail --silent http://127.0.0.1:9333/json/list

A gpu-process command line containing --use-gl=disabled means software rasterization can be consuming CPU despite the process name. Killing that child is temporary because Brave recreates it. The repository launcher avoids the shared-profile access pattern, does not disable GPU acceleration, uses an isolated profile, and binds the debugging endpoint to loopback. Override its port with AUTODEV_BRAVE_DEBUG_PORT when 9333 is already occupied.

Your checkout is behind the branch the fleet landed on​

A dirty working tree no longer blocks anything (A-151). Apply stages its merge in a fleet-owned checkout under worktree_root, runs the post-merge gate there, and advances the base ref; your checkout is a consumer.

After a landing the fleet tries to bring your checkout forward without touching your work, updating only the files the landing changed and keeping every uncommitted edit that does not collide with one. When an edit does collide it refuses rather than overwriting, and the tick says so with the command that reconciles it when you are ready:

git stash && git reset --hard <landed-sha> && git stash pop

Until you run it, your uncommitted work is exactly where you left it and the delivery is already on the ref. One consequence worth knowing: the board transition for that delivery is owed rather than attempted while the checkout is behind, because the hook that writes it commits in your checkout. The next tick retries it once the tree can move.

A wave is stuck or malformed​

autodev fleet status
autodev fleet doctor
autodev retro

Pause new dispatch, then park a wave while investigating:

autodev fleet park wave-12 --reason "investigating malformed state"

Parking keeps its stories held. A running tick no longer holds the repository write lock across its turns, so parking does not need the drive stopped first; it waits only on the short sections a delivery or an apply takes.

Use fleet unpark when the cause is corrected, or fleet retire <wave> --reason <text> to terminate it and release its stories. There is no fleet cancel command.

Unpark can answer in two ways. A wave re-earns the areas it claims before it re-enters the loop, so if another live wave now holds one of them, the wave stays parked and the answer names the area and the holding wave. Nothing is written in that case. Wait for the holding wave to close, or retire one of them. Note that this deferral is reported only in human output; see Current limitations.

Disk dispatch is withheld​

Run autodev fleet disk and inspect the filesystem containing worktree_root. Terminal and crash-left worktrees are reclaimed by later ticks where possible, but resident worktrees can each hold a build cache. Reduce concurrency/open-wave limits, move the root to an appropriately sized filesystem, or raise the floor only when measurement supports it.

Never lower the floor just to silence the refusal: a build that runs the disk out mid-gate cannot be distinguished from a real red gate by wishful configuration.

A second drive or board write reports a lock​

Only one drive may own a repository. Board and fleet mutations wait a bounded time for the repository write lock — 120s for board, 10s for fleet, or whatever --lock-wait-ms sets. A tick does not take that lock, so the holder you collide with is normally another operator verb, or the short section a delivery or apply holds. Read-only diagnosis never takes it at all.

Check whether the named holder process is alive and whether another drive is intentionally running. A restarted drive can recover a stale owner whose process is gone. Avoid manually deleting lock files as a first response because a live holder and a stale record require different actions.

For a registered workspace, exclusion is owned by the daemon rather than competing SQLite openers. A busy or unavailable daemon should still be diagnosed through daemon status, the typed refusal, and the workspace registry—not by deleting repository lock markers.

Metrics show unknown​

unknown means the necessary fact was not recorded. Examples include a story with no committed history, a harness with no usage report, an unpriced model/rate, or no measured delivery rate for a forecast. It is not equivalent to zero and should remain distinct in exported dashboards.

The drive was interrupted​

Do not immediately start editing worktrees. Restart with fleet status, fleet doctor, and one fleet tick. The reconciler recovers ordinary dead workers and stale resources, but missing process groups and refused kills can still require operator intervention. Use park, unpark, or retire with a reason so the recovery remains on the record.