herdr hands a freshly created pane/agent back before it's actually ready,
and rejects the very next call with a range of different transient errors
("not an available shell", "not an active named agent", "target ... not
found") depending on timing. String-matching each wording as it turned up
live proved unwinnable across three live redeploy-and-test rounds, so
StartAgent and Prompt now retry any error for up to 15s (bounded by
wall-clock time, not attempt count) rather than pattern-matching herdr's
error text.
Confirmed live against workpc: a fresh lease (wD:p1) now reaches a real
attached claude session instead of failing before the agent starts.
Live testing also exposed a second, separate defect (B13, documented in
AUDIT.md, not fixed here): agent.start can return success while never
actually starting an agent when two leases land close together, with no
error for a retry to catch. Left three test panes on workpc untouched
(wD:p1, wE:p1, wF:p1) pending manual cleanup, per the standing rule against
destructive herdr calls without asking first.
Also folds in the already-flattened AUDIT.md/progress.md merge that was
staged ahead of this session's changes.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1rkJ2hBMybnJctPbcy4tT
39 KiB
Orchestra — spec conformance audit & remediation status
Audited 2026-07-27 against orchestra-spec (1).md at commit 325c684.
Method: read every non-test file in internal/ and cmd/, traced each spec
section to its call site, and checked whether the live path (main.go →
router → coordinator → adapter) actually reaches it. This file is the single
merged record — a prior split between AUDIT.md (audit + plan) and
progress.md (chronological session log) has been flattened here; treat any
older chronological git history of either file as session notes, not ground
truth.
Ground truth rule for future sessions: this repo has a documented history of code that looks wired but isn't (packages with tests that pass in isolation while the live call path silently no-ops). Before trusting any "done"/"fixed" claim below, check the actual call site, not just this document.
Current verdict (as of 2026-07-28)
The substrate (Layer 1) is solid. Layers 2–4 were originally shaped-but-not- wired; most of the blocking defects below are now closed, verified by reading the current code (not just by trusting this file) and by targeted tests. Two things are not yet proven:
- A real end-to-end unattended run. The first live attempt (2026-07-28)
surfaced B12 (fixed and confirmed live, see below), and re-verifying
that fix surfaced B13 (open) —
agent.startcan return success while silently never starting an agent, under back-to-back leases. One task (wD:p1) did make it all the way to a real attachedclaudesession, so the lease path is provably reachable, but B13 means it's not yet reliable enough to call proven — occupancy/rotation/handoff/completion still haven't been exercised end-to-end against a session that's guaranteed to actually exist. - Cross-machine (federation) correctness. Deliberately deferred — see "The federation fork" below.
go build ./..., go vet ./..., and go test ./... all pass.
| Spec layer | State |
|---|---|
| L1 substrate (§3, §4) | Built and correct in the main path. |
| L2 harness (§5) | Occupancy, rotation, and completion all wired and reachable; B12's readiness race is fixed and confirmed live, but B13 (agent.start silently no-op'ing under back-to-back leases) still blocks reliable live verification of what's downstream. |
| L3 continuity (§6) | Handoff schema, pickup validation, scratch branches, TASK.md are all wired into the live path (Phase 4 complete). |
| L4 surfaces (§7) | Brief/standup/delivery real; quota has a post-hoc producer only (no live push feed yet); authz bypass closed. |
Blocking defect currently open
B13 — agent.start silently no-ops under back-to-back leases (found live 2026-07-28)
Discovered while live-verifying the B12 fix below. B12 itself is confirmed fixed (see that entry), but re-testing it exposed a second, deeper defect that B12's retry logic cannot address.
Three fresh tasks were leased in quick succession against workpc-claude
(06FTAKSJZTB73FZQE3QT7XQ1J0/pane wD:p1, 06FTAMFKVEPGPKR0H86C08ZZNG/pane
wE:p1, 06FTANAYZPA55BABR2J4MWQ050/pane wF:p1). All three got a
worktree.create success and an agent.start call that returned no
error. Checked live via pane.get/pane.list (and independently
confirmed by the user cd'ing into each worktree by hand):
wD:p1(the first, given time alone to settle): a realclaudeprocess did eventually attach —agent: claude,agent_status: idle,revision: 2. This is consistent with B12's model:agent.startacknowledges the request but the actual attach is asynchronous, sometimes taking well over a minute.wE:p1andwF:p1(started moments afterwD:p1, while it was presumably still initializing): still no agent attached after 2+ minutes of polling —agent_status: unknown, noagentfield,revisionnever incremented past its creation value. Empty shells, confirmed both viapane.getand by the user directlycding into the worktree and finding no session running.
So agent.start's "success" is not just delayed for these two — it never
happened at all, and agent.start never returned an error to tell Orchestra
that. One plausible cause: something about launching a second/third claude
CLI process while a prior one on the same host/account is still mid-startup
(auth handshake, config lock, etc.) silently drops the later request(s)
rather than queuing or erroring them. Not confirmed — would need a
controlled single-at-a-time repro to isolate from herdr's own internals,
which requires spinning up more real panes and wasn't done this pass (see
"three orphaned panes" below).
Consequence: unlike B12 (which at least produced an observable
TaskBlocked), this failure mode leaves the task leased indefinitely against
a pane that will never produce a session, with no error surfaced anywhere —
worse than B12 was, because nothing currently distinguishes "still booting,
give it more time" from "silently dead, will never start." No fix attempted
yet; a client-side retry (B12's approach) cannot fix this because the
agent.start call that should have started the process already returned
success.
Side effect of this investigation — three live orphaned panes on workpc,
left untouched on purpose: wD:p1 (has a real but abandoned claude
session, task now stuck in blocked from before B12's fix), wE:p1 and
wF:p1 (empty shells, no agent ever attached). Not cleaned up — same
standing rule as the pre-existing wA:p1 stuck task: don't call
pane.close/pane.release_agent against live panes without asking first.
B12 — StartAgent races the pane's shell readiness (found live 2026-07-28) — closed (this specific race; see B13 for a second issue it exposed)
Discovered during the first real end-to-end run against a freshly rebuilt/
redeployed binary (the running service had been on stale pre-audit commit
325c684 the whole time).
Created a fresh task (06FTAH4MCRCAJYY7V75Z4V2YGG, project test-e2e,
capability opencode). It reached TaskLeased onto workpc-opencode
(worktree created for real at /tmp/test-e2e-worktrees/<task-id> on workpc,
confirmed via a live pane.list — pane wB:p1, correct cwd), then
445ms later hit TaskBlocked:
lease: herdr protocol error: agent target pane wB:p1 is not an available shell
internal/herdr/herdr.go:199 (StartAgent) calls agent.start on the pane
Worktree just recorded, with no wait or retry after worktree.create
returns. herdr hands the pane back before its shell has actually finished
initializing, and agent.start refuses it as not-yet-a-shell. None of B1–B11
exercised this — every prior fix assumed a session already existed
(occupancy, rotation, handoff); nothing in the remediation plan tested the
very first Lease call against a real herdr pane end-to-end. Phase 0's
probing used params:{}/missing-field tricks, never a real
worktree.create → agent.start sequence timed against actual pane boot.
Reproduced a second time, 2026-07-28, same session: a second fresh task
(06FTAHW8GZP0S6DCHEQ6PJXWB0) leased onto a brand-new pane (wC:p1) and hit
the identical error on the same sub-second timeline. Two for two — a
systematic race, not a flake.
Consequence: every fresh lease that hits this race goes straight to
TaskBlocked, and the router never retries a blocked task
(internal/router/router.go has no reference to StateBlocked at all,
confirmed by grep) — so it sits there permanently until a human intervenes
via the approval surface. This is now the single blocking defect for proving
the rest of the remediation plan against a real live task.
Fix, iterated live across three rounds, 2026-07-28:
- First pass: bounded-retry
agent.start(10×, 500ms apart) specifically on the"not an available shell"string. Redeployed and re-tested against a real fresh lease — this cleared that exact error, but the very next call in the sequence (Lease's post-agent.startPrompt) then hit a different transient error,"agent ... is not an active named agent"— same underlying readiness race, one call later, worded differently. - Second pass: added the identical string-matched retry to
Promptfor that specific message. Redeployed, re-tested — this time hit a third distinct wording,"agent target ... not found", on the same call. - Given three different error strings for the same race across three live
attempts, string-matching was abandoned as unwinnable (herdr's wording for
"not ready yet" isn't fixed enough to enumerate). Final shape: both
StartAgent'sagent.startcall andPrompt'sagent.promptcall now retry on any error, bounded by wall-clock time (15s,bootRetryDelayof 500ms) rather than attempt count or specific wording — there is no other legitimate reason a call against a pane/worktree orchestra itself just created would fail immediately.StartAgentstill special-cases"already"as a non-error (idempotent re-lease).
Confirmed this closes the original race: task 06FTAKSJZTB73FZQE3QT7XQ1J0
(pane wD:p1) leased successfully, and polling pane.get live shows a real
agent: claude / agent_status: idle attach — independently confirmed by
the user cd-ing into that worktree by hand. go build/go vet/go test ./... all pass throughout.
Not fully clean, however: re-testing this fix by firing two more tasks
shortly after the first one surfaced a second, separate defect — agent.start
returning success while never actually starting an agent, with no error for
the retry logic to even see. That's B13, above — this entry only covers
the specific shell/agent-readiness race B12 was originally filed for, which
is fixed; B13 is a distinct root cause the same investigation exposed.
Closed defects
Each entry is the flattened final state — design + what's verified — not the session-by-session narration. Grouped by original defect ID.
B1 — Occupancy measured from the wrong value (§5.2.1) — closed
CLIAdapter.Occupancy was calling a.Usage(s.PaneID) (a herdr pane id)
against readers (ClaudeUsage/CodexUsage/OpenCodeUsage) that want a
filesystem path to session state — every call failed and was swallowed by a
bare continue. Fixed: herdr.Session.SessionFile is resolved at lease time
per harness (Claude: newest-mtime transcript under Claude Code's own project
dir; Codex: existing CodexActiveUsage sqlite discovery; opencode: refuses
loudly rather than guess, since it needs a live session id not resolvable
from the worktree alone). A missing/unreadable session file is now a hard
error surfaced via SessionHealth.Occupancy/OccupancyError on
GET /v1/tasks/{id}/health, never a silent zero.
Still open: live verification against a real Claude Code session at a known context fill — not possible from this sandbox, blocked further by B13 now that fresh leases don't reliably survive.
B2 — Adapter lookup keyed wrong in 3 of 4 call sites (§5.3, §5.4) — closed
AdapterFactory.Herdrs is keyed by herdr instance id (homesrv-claude);
Reconcile/expire/rotate were looking it up by session.Harness (the
harness kind, claude) and silently no-op'ing (continue) on every miss —
orphaned panes never killed, expired leases never killed their pane,
rotation never began. Fixed with a single Coordinator.adapterFor(session)
routed through at all four call sites. Regression test registers an adapter
under a herdr-id key distinct from the harness kind and asserts rotation
still fires.
B3 — Nothing emitted TaskCompleted (§4, §5.2) — closed
Only producer used to be a human calling POST /v1/tasks/{id}/complete.
Added POST /v1/harness/complete: a Claude Code Stop hook
(deploy/hooks/orchestra-stop.sh) fires every turn boundary but only
reports completion once the agent has written a .orchestra-report.md
marker at the worktree root (an ordinary turn boundary is a no-op). The
server builds the receipt itself from herdr.ClaudeUsage against the
transcript (never trusts a self-reported number) and uploads the report
body to CAS for report_ref. Gated by optional ORCHESTRA_HARNESS_TOKEN;
event appended with Surface: system set directly in Go, consistent with
B8. Extended to Codex/opencode via an optional harness field in the
request body ("codex" → CodexUsage, "opencode" → OpenCodeUsage).
Still open: no automated test for the HTTP handler (cmd/orchestra/main.go
has zero handler test coverage of any kind, pre-existing gap — this follows
the existing pattern rather than introducing a one-off harness).
B4 — Router counted rotation as a retry (§5.3, §5.4) — closed
Every TaskReleased (including rotation, which carries a valid
handoff_ref) advanced MaxAttempts, and lease also double-counted — a task
healthy enough to rotate twice was killed. Now only a release without a
handoff_ref (expiry/crash) advances the counter.
B5 — Herdr protocol methods unverified/invented — closed
Confirmed live against 192.168.1.105:9245 (raw JSON-RPC probing, no local
herdr CLI available — schema reconstructed from Rust serde's
"unknown variant"/"missing field" errors). pane.release, pane.kill,
pane.rotation_signal, pane.status do not exist. Real replacements:
pane.close({pane_id}) for kill (drop-in); pane.release_agent({pane_id, source, agent}) for release — structurally different, does not return
a handoff_ref (herdr never writes handoffs; the agent does, §6.1).
RotationSignal interface/method/call site deleted outright (no replacement
exists; herdr has no concept of Orchestra rotation).
CLIAdapter.Release now: reads the agent-authored .orchestra-handoff.json,
validates with continuity.Decode, cross-checks its anchor SHA against the
worktree's real HeadSHA, re-verifies every Anchor.Dirty file hash,
scratch-commits any dirty files (continuity.ScratchCommit, idempotent
across multiple rotations) and rewrites the handoff's anchor to the new
scratch commit before uploading to CAS via continuity.Save (minting the
real handoff_ref), and only then calls pane.release_agent — sequenced
last so a herdr-side error can't strand an uploaded handoff. Any failure
(missing file, invalid schema, anchor mismatch, stale dirty hash, herdr
error) is a refusal, which rotate already treats as "retry next tick."
CLIAdapter.Lease's bootstrap prompt also fixed to pass time.Minute (not
wait=0) for inline-wait, matching Bootstrap and closing the
send-into-a-half-rendered-prompt race §5.1 requires guarding against.
Covered by internal/herdr/adapter_test.go against a real git worktree and a
fake in-process herdr TCP listener: upload-and-release on a valid handoff,
refusal with no handoff file, refusal on anchor mismatch, refusal on stale
dirty-file hash, refusal with no CAS configured.
Protocol version confirmed live as a bare JSON number (17), not the string
config.jsonc declares — CheckProtocol's raw-bytes fallback happens to
compare correctly today; don't "clean up" that code without re-checking
this, or it may start doing a real numeric-vs-string comparison and break.
Still open: nothing makes the agent actually write
.orchestra-handoff.json unprompted (closed separately, see B6/Phase 4 item
2 below) — that's the piece this entry originally deferred.
B6 — Layer 3 (continuity) was entirely dead code (§6) — closed
Nothing wrote TASK.md, so pickup validation had nothing to check and never
ran. Fixed in full:
GitWorktrees.Createnow writes and commits an immutableTASK.md(continuity.RenderTaskFile) into every fresh worktree;continuity.TaskFileHashreads it back forw.TaskFileSHA.Coordinator.Startrunscontinuity.ValidatePickup(anchor SHA + dirty-file hashes + TASK.md hash) before bootstrapping a successor onto ahandoff_ref— failure kills the session and emitsTaskBlockedinstead of trusting an unvalidated ref.ScratchCommitwired intoRelease(see B5) so pickup collapses to a single HEAD compare instead of per-file rehashing.CLIAdapter.Bootstrap's prompt rewritten to point the agent atgit log/the scratch branch (trust already established by the plane's ownValidatePickup) instead of vague "read the handoff" prose, and deliberately does not claim aGET /v1/artifacts/<ref>endpoint since none exists (/v1/artifactsis POST-only — an earlier draft invented this and was corrected before landing).continuity.MarkdownChanges(zero callers, zero tests despite being listed as implemented in an earlier snapshot) was deleted rather than half-wired. §6.3's actual requirement ("notice to agents whose task is adjacent") was rebuilt independently:continuity.ConventionsHash(root)hashesAGENTS.md/CLAUDE.md/VOCAB.md;Coordinator.checkConventionsruns everyMonitortick, recomputes the project base repo's hash for every leased session, and on drift calls the optionalherdr.ConventionsNotifiercapability (in-pane prompt) once per drift.- Handoff production, the item that most of B6 hinged on:
Coordinator.rotatechecks the adapter's optionalherdr.HandoffRequestercapability; if.orchestra-handoff.jsonis missing, it prompts the agent once (CLIAdapter.RequestHandoff, naming the exact §6.1 JSON shape and explicitly telling the agent not to fabricate SHA/hashes) and skipsReleasethat tick, retrying every subsequent tick — mirrors the.orchestra-report.md/B3 convention exactly.Session.HandoffRequestedavoids re-prompting every tick.
Covered by TestGitWorktreesCommitsTaskFile, TestStartBlocksOnInvalidPickup,
TestReleaseScratchCommitsDirtyFilesBeforeUpload,
TestReleaseRefusesOnStaleDirtyFile, TestConventionsDriftNotifiesActiveSession,
TestRotationRequestsHandoffBeforeReleasing (internal/herdr, internal/orchestrator).
Caveat still open: TASK.md hashing is best-effort/untested for the
herdr-hosted (WorktreeCreator) worktree path specifically — no adapter or
test exercises that path with a real handoff_ref, so pickup validation
there runs with an empty taskFileSHA (anchor + dirty-file hashes still
checked). Same cross-host caveat as the federation fork, below.
B7 — Quota projection had no producer (§7.2) — closed (post-hoc only)
POST /v1/harness/complete now appends a QuotaReported event
(consumed from the same usage.Numerator() used for the receipt), so the
router's 5h/weekly filter and the brief's quota_consumed stop evaluating
against a permanent zero.
Design gap recorded, not yet built — a live push producer, since a
harness that never completes a task cleanly (e.g. the stuck wA pane, see
below) currently under-counts its consumption forever:
- Claude: the statusline stdin JSON already carries real
.rate_limits.five_hour.used_percentage/.seven_day.used_percentage— confirmed by reading~/.claude/statusline.shon this machine. Prefer this overclaude.ai/api/.../usage, which needs browser session cookies, not an API key. - Codex: every
token_countevent in the rollout carries arate_limitsobject withused_percent/window_minutes/resets_at— confirmed directly from a real rollout file. Nocodex usagesubcommand; the rollout tail (orcodex app-server, same events live) is the only transport. - opencode: no first-class quota surface.
opencode statsis historical only. Zen's free-tier daily quota arrives asx-ratelimit-*response headers, not a queryable endpoint — would need an opencode plugin to intercept. - Structural consequence if this is ever built:
sumSince's additive model is wrong for a used-percentage feed (need a secondAvailabilityreading the latest report per harness, not summing); staticquota_limit_5h/weeklybecome unnecessary for Claude/Codex once real fractions are reported;QuotaWindowLimits{FiveHour,Weekly}is too narrow for Codex's self-describingwindow_minutesor opencode's daily Zen window.
B8 — Surface: system unauthenticated bypass (§7.1) — closed
X-Orchestra-Surface: system was reachable from any HTTP request header in
both authz.HTTP and main.go's surface closure — since no deployment
sets ORCHESTRA_SYSTEM_TOKEN, this was an unauthenticated full-control
bypass reachable from any LAN caller. Both call sites now downgrade system
to web before doing anything else with it. System is only constructible
in-process (router leases/failures, coordinator releases/blocks, standup
advisory/apply) via the bus-level authz.AuthorizeEvent check now enforced
inside store.Append itself.
S1 — duplicate JSON struct tags — closed
Brief.From/To and GitSync.Branch/Head/Status each shared one JSON
tag (Go only honors the first json:"..." tag on a combined field
declaration), making go vet ./... fail and the brief's git state
unparseable. Fixed; go vet ./... passes clean.
S2 + S3 — brief git state and completion receipts — closed
Brief.Git used to read from ORCHESTRA_DATA (the event-log dir, never a
git checkout — always "git unavailable") and completions were counted but
never surfaced with proof. operations.Brief.Git is now map[string]GitSync
keyed by project ID, built from registry.Project.Repo (falling back to a
single "default" entry off ORCHESTRA_REPO), with Ahead/Behind vs
upstream added. Brief.Receipts []CompletionReceipt pulls report_ref/
receipt straight out of each TaskCompleted event's existing payload.
S4 — notification goroutine died on first send error — closed
delivery.Fanout.Run used to return on the first sender error,
permanently killing notifications for the rest of the process's lifetime.
Failed sends now go through an OnError hook instead of aborting the loop.
Cursor is persisted (SaveCursor → a delivery-cursor file next to
ORCHESTRA_DATA), so a restart resumes from the last delivered event
instead of re-notifying the entire log from seq 0.
S5 — colliding event IDs on lease — closed
Store.Lease/ExpireLeases set Event.ID to the task id, so every lease of
a task produced colliding event IDs. Now uses domain.NewID().
S6 — ingest dedup returned the wrong event with 200 — closed
Dedup path returned nil (success) without appending; the HTTP handler then
returned an unrelated event with 201. Added domain.ErrDuplicate and
Store.TaskBySource; POST /v1/tasks now returns the existing task with
200 on a duplicate. Every other Append caller (Gitea poll/webhook, JSONL
ingest) now treats ErrDuplicate as expected.
S7 — TaskAmended dropped most fields — closed
Only title was ever applied from an amendment payload; due/
description/inherent_priority were accepted, logged, and silently
dropped by the projection. Also added the missing Task.Description field
entirely (it didn't exist, so TaskCreated dropped it too). Both
TaskCreated and TaskAmended now populate all four fields.
S8 — no compensating-event mechanism (§3.1) — closed
Added TaskCorrected: payload requires corrects (the id of the event it
repairs) plus at least one change (state, or the existing amend-style
fields). Store.Append rejects a corrects that doesn't name a real prior
event on the same task; apply applies the changes and clears Lease like
every other terminal-state branch. Covered by a mistaken TaskFailed
reverted to queued, an unknown-corrects rejection, and both events
surviving snapshot+replay.
S9 — duplicate lease-expiry tickers — closed
Coordinator.Monitor's 30s ticker and main.go's own 1s reclaim ticker both
called Store.ExpireLeases independently. Previously filed as "harmless —
CAS rejects the loser," but the real effect was worse: the coordinator's
expire() is the only place that kills the herdr session/pane for an
expired lease, and the 1s ticker running 30x more often almost always won
the race, leaving the coordinator's own call with nothing left to expire —
silently orphaning herdr panes past TTL. Fixed: main.go's ticker skips
ExpireLeases entirely when a coordinator is configured (deferring reclaim
to the coordinator's loop), keeping only AssignPending as periodic retry.
S10 — federation worker registration had no admission control — closed
Any caller could self-declare a worker id and self-choose a token, and a
second caller could silently re-register an existing worker id with a
different token, hijacking its identity/capacity. Added
Registry.AdmitToken (pre-shared secret via
ORCHESTRA_FEDERATION_ADMIT_TOKEN, checked against the registration
request's bearer header); same-ID re-registration now requires presenting
the existing worker's own token (legitimate restarts still succeed;
different-token same-id is rejected as a hijack).
S11 — soft threshold, milestone, thrash, agent-initiated ROTATE (§5.3) — closed
All four of §5.3's rotation triggers now land:
- Soft threshold:
Coordinator.Soft(default 0.55,ORCHESTRA_OCCUPANCY_SOFT) — bothrotate()andTurnDecisionrequest a handoff once occupancy crossesSoft, beforeHardforces one;TurnDecisionreturnsprepare_handoffadvisory (task stays leased, no turn boundary required). - Agent-initiated ROTATE:
handoffReason(worktree)reads a written.orchestra-handoff.json; ifreason == "manual", bothrotate()andTurnDecisionskip occupancy and the turn-boundary probe entirely and release immediately — the agent's own handoff is the boundary signal. Shared release-and-certify tail extracted intoCoordinator.finishReleaseso this path gets the same anchor-safety guarantee as the threshold path. - Milestone + thrash: needed a transcript/tool-call introspection source
this repo didn't have — built new,
internal/herdr/activity.go:ToolCall{Name, Kind, Key, Success, IsTest}normalizes one tool/function call across harnesses.ClaudeActivitypairstool_use/tool_resultblocks bytool_use_idfrom the transcript already opened for occupancy (an unresolved tool_use — mid-turn — is dropped, not reported).CodexActivity— verified against a real live rollout, 2026-07-28, after the original guessedfunction_call/function_call_outputshape turned out not to exist anywhere in~/.codex/sessions. Real shape: file edits arrive asevent_msg/patch_apply_end(haschanges+ top-levelsuccessdirectly, no pairing needed); shell commands arrive as a freeformresponse_item/custom_tool_callnamed"exec"whoseinputis a JS snippet —codexExecCommandregex-extracts the first embeddedcmd:"..."; failure is signaled by the literal output prefix"Script error:"(confirmed for both a JS syntax error and anapply_patchverification failure on this machine's real transcripts). Manually re-ran the rewritten parser against a real multi-hundred-line rollout end-to-end and spot-checked the output by eye, in addition toTestCodexActivityParsesRealRolloutShape/TestCodexActivityMarksScriptErrorAsFailure.OpenCodeActivityrefuses outright — opencode's on-disk message storage is only confirmed to carry aggregate token counts, not per-tool-call records; fabricating a parser against an unconfirmed shape would repeat the exact mistake this audit exists to catch.DetectThrash(calls, ThrashConfig)implements all three rules: N consecutive failed test runs (regex-matched build/test commands), the same file edited M times (edit tools only, explicitly excludingRead— caught by an initially-failing test), and an identical tool call repeated K times back-to-back (excluding test re-runs and file reads — also caught by an initially-failing test). Defaults 3/5/4, overridable viaCoordinator.Thrash.DetectMilestone(calls)— deliberately narrow: only "the most recent call was a successfulgit commit." Fuzzier definitions (passing test suite, finished subtask) were not guessed at.CLIAdapter.Activityresolves the session file the same wayOccupancydoes;ActivityReaderis an optional Face-B capability like the others.- New
ReasonedHandoffRequester/CLIAdapter.RequestHandoffReasonnames why (thrash's dead ends, milestone reasoning) instead of the generic threshold framing. rotate()/TurnDecisiongeneralize "reason=manual bypasses occupancy" tomanual/milestone/thrashalike:checkActivityTriggersruns before falling through to occupancy/soft/hard (thrash takes priority over milestone), and a hit callsrequestReasonedHandoff— request only, never release, same as the soft-threshold path.
Covered by internal/herdr/activity_test.go and
internal/orchestrator/rotation_test.go's
TestActivityTriggersRequestReasonedHandoffWithoutReleasing,
TestTurnDecision subtests for soft-threshold and manual-reason, and
TestRotationRequestsHandoffBeforeReleasing.
Known remaining limitation: DetectMilestone's single-command
extraction only sees the first cmd:"..." in a chained Codex exec script,
so a git commit issued later in the same script isn't recognized as the
"last call" even though it executed last. Opencode still has no verified
per-tool-call source at all.
Also closed this pass (not separately numbered in the original audit)
- Codex/opencode Stop-hook-equivalent scripts. Neither harness has a
native Stop hook, so the server-side
/v1/harness/turnand/v1/harness/complete(harness dispatch) existed with no caller for either. Addeddeploy/hooks/orchestra-codex-poll.shandorchestra-opencode-poll.sh— background poll loops (default 60s,ORCHESTRA_POLL_INTERVAL) that find the newest session-state file (codex: newestrollout-*.jsonl; opencode: newest file under~/.local/share/opencode/storage/message/), POST to/completewhen.orchestra-report.mdexists, otherwise POST to/turnand log (not act on)refuse/rotate_now— advisory-only, since there's no real turn boundary to refuse at from outside either harness process without deeper app-server/SSE integration. - Bus-level authorization.
authz.AuthorizeEventenforced insidestore.Appenditself (the single choke point every event passes through), not just at HTTP handlers. Event schema bumped to v2, requiring every event to declare aSurface; schema v1 events on disk still replay (tolerant reader). - Dual quota windows.
router.QuotaAvailabilitytracks a 5-hour rolling window and a 7-day weekly window independently per harness, applying the conservative 80% rule to each separately. - Turn-boundary detection made observable. An adapter that fails to
answer
TurnBoundarynow blocks that tick's release rather than treating the failure as "safe to proceed," and incrementsMonitorHealth.TurnBoundaryDegraded. - Cross-machine lease correctness has a primitive-level test.
TestCrossMachineLeaseAnchorAndQuotaArePerHostexercises a lease claimed through the federation worker HTTP API, validates the anchor against that worker's own local checkout, and asserts quota is accounted per-host — proof for the primitives that exist today; does not yet run against two real physical machines (see federation fork, below). - Fuzz coverage (
internal/domain/fuzz_test.go,FuzzValidatePayload/FuzzValidateEvent) for all event types including malformed nested payloads — no panic, always a typed error. - Multi-repo Gitea ingestion.
provider.Giteagained aProjectfield and namespacedSourceName();provider.MultiGiteadispatches bytask.Source;LoadGiteaConfigsloads a JSON array of per-project sources. Legacy single-repo env vars still work unchanged. Previously zero tests existed for the Gitea provider at all (an earlier progress.md claim of coverage was inaccurate) —internal/provider/gitea_test.goadded. - Per-project repos.
registry.Projectgained optionalrepo/worktree_root;main.gobuilds aPerProjectGitWorktrees, falling back to the global default for any project that omits these fields. - Rotation's
TaskReleasedpayload validity.Coordinator.rotateused to build{"handoff_ref","reason"}, omitting theanchor_shathe spec anddomain.ValidatePayloadrequire wheneverhandoff_refis present —store.Appendwould reject it, the error was discarded (if c.Store.Append(e) == nil), and rotation silently never happened, no visible failure. Fixed:herdr.HeadSHA(worktree)populatesanchor_shafrom the real worktree HEAD before appending; if the anchor can't be read, rotation now skips that tick instead of emitting a payload guaranteed to fail validation. The federation worker release endpoint (/v1/federation/workers/{id}/release) had the identical gap and now requires/forwards a 40-hex-charanchor_sha, rejecting with 400 otherwise.
The federation fork — decided 2026-07-27, still the deferred half
Two incompatible federation designs coexist in the tree.
Design A — "drive the remote socket" (currently deployed,
clients/herdr-bridge.go): homesrv calls worktree.create/agent.start/
etc. directly on workpc's herdr over TCP as if it were local — meaning
anchor validation (git rev-parse HEAD) executes on the wrong machine
relative to the actual checkout. Two outcomes if the coordinator's
session.Worktree path happens to also exist on homesrv (likely, since
every project shares the same directory layout): either HeadSHA errors and
rotation silently skips forever, or — the dangerous case — it returns
homesrv's HEAD for an unrelated checkout, passing validation while
certifying a commit the agent never touched. Same class of bug applies to
per-host quota accounting and to cleanupCompleted's git worktree remove,
which runs on homesrv for a worktree that lives on workpc.
Design B — "workers pull tasks" (/v1/federation/*): fully built
server-side (registration, heartbeat/TTL offline detection, event-cursor
polling/ack, lease claim), zero clients — no worker binary exists anywhere
in this repo. This is what the spec actually describes (§2.1: "everything
crossing a machine boundary is git + a validated artifact, never live state
over the wire"), but every endpoint is currently unreachable in the real
deployment. Latent, never-surfaced defects in this unused half: offline
detection only runs inside Snapshot(), called solely from a GET endpoint —
nothing ticks it on its own, so OnOffline (the hook that releases leases
held by a vanished worker) only fires if a human hits that endpoint; the
registry is in-memory with no persistence, so a restart forgets all workers
and cursors.
Decision, unchanged: keep Design A through Phase 5 (single-host concerns
— occupancy, Face B, rotation, continuity — are provable on homesrv alone
with workpc's herdr as just another pane host); commit to Design B in Phase
6. Two guardrails were meant to land immediately so Design A can't corrupt
state in the meantime: (1) refuse to rotate a lease held by a non-local
herdr rather than validate against the wrong checkout, since the protocol
schema can't run rev-parse where the checkout is; (2) same treatment for
cleanupCompleted's worktree removal. Status of these two guardrails is
unverified in this pass — re-check internal/orchestrator before assuming
they landed; they are not confirmed closed above the way B1–B11/S1–S11 are.
If the worker binary is ever abandoned, delete /v1/federation/* and record
the deviation — leaving both designs in place unmarked is explicitly not
acceptable; the repo would read as though the spec-conformant path is
implemented when nothing can reach it.
Live deployment facts (as of 2026-07-28)
- Runs as
orchestra.serviceon homesrv via this machine's own systemd.homesrvhas no local herdr running (connection refused on 9245) — onlyworkpc's herdr (192.168.1.105:9245) is live and reachable.main.goonly logs herdr connection failures at startup, never successes, so "no log line" for a herdr does not mean it's down. - Stuck task, deliberately left untouched: workspace
wA, task06FT6CKD9Y98AZRX6X8K3QXFZG, opencode harness, panewA:p1. From Orchestra's own point of view it is no longer "stuck" — it now readsstate: "failed"(retries exhausted,MaxAttemptshit before B4's fix landed). But the underlying herdr pane is still live andagent_status: "blocked", confirmed via a directagent.getprobe — the router gave up and moved on, but nothing ever released or killed the actual agent, confirming the orphaned-pane prediction from B2/B5. Left untouched on purpose: user was asked and chose to leave it rather than have it cleaned up mid-audit. Do not callpane.close/pane.release_agentagainst it without asking again first. - First live end-to-end task run (2026-07-28, after rebuilding/redeploying —
the running service had been on stale pre-audit commit
325c684) is what surfaced B12 above, since fixed and confirmed live; re-verifying it surfaced B13, still open. - Three more live orphaned panes from this pass, left untouched:
wD:p1(task06FTAKSJZTB73FZQE3QT7XQ1J0, worktree/tmp/test-e2e-worktrees/06FTAKSJZTB73FZQE3QT7XQ1J0— has a real attachedclaudesession but the task itself is stuckblockedfrom before B12's fix landed),wE:p1andwF:p1(tasks06FTAMFKVEPGPKR0H86C08ZZNG/06FTANAYZPA55BABR2J4MWQ050— empty shells, no agent ever attached, see B13). Same rule aswA:p1: don'tpane.close/pane.release_agentany of these without asking first.
Known open gaps (as of 2026-07-28)
- B13 (above) —
agent.startcan silently no-op under back-to-back leases, with no error surfaced; blocks reliable live verification of everything downstream (occupancy, rotation, handoff, completion) against a real task, even though B12's readiness race is now fixed. - Cross-machine lease correctness proven at the primitive level only — needs an actual homesrv/workpc pair over the real mesh; this repo cannot exercise that by itself.
- Turn-boundary Face B degrades to occupancy-only for adapters that
don't implement it, by design — degradation is observable
(
MonitorHealth.TurnBoundaryDegraded) and blocks-on-failure, but whether Claude/Codex/opencode's native hooks are wired in a live deployment is a deployment-config fact, not provable from source alone. - Quota has no live push producer — only the post-hoc
/v1/harness/completeproducer exists; see B7's "design consequences" for the (unbuilt) Claude statusline / Codex rollout-tail feeds. - Federation guardrails' landed-status is unverified — see federation fork section above.
- opencode has no verified per-tool-call activity source and no first-class quota surface (Zen headers only, would need a plugin) — both named, not attempted, in S11/B7.
DetectMilestone's single-command extraction can miss a Codex commit issued later in a chained exec script.
Everything else audited (provider layer, continuity schema/validation, router matching, delivery, federation registration primitives) was spot-checked against the code and its tests and matched described behavior.