6 Commits

Author SHA1 Message Date
kami 99b209ba10 Record run 18: an implement successor inherits verified phase progress
The rung the last two runs missed. A five-phase task with ten named checks
kept the implement phase open long enough to rotate inside it. The successor
picked up in implement and its launch context carried the whole sealed plan,
phase-1 and phase-2 as verified, phase-3 as the first unfinished phase, and
the current human authority.

Unasked-for bonus: the progress block renders SHA staleness itself, naming
the tree each phase was verified against and the tree it is now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CVbaKucEYBjMqVeUgJUsc1
2026-08-29 17:48:17 +04:00
kami ab3258833d Record run 17: F62 proven live, from a production trigger
The rig was not needed. An implement to review phase change asked for the
handoff through the ordinary path, and the agent ignored it, which is the
exact shape F62 was written for. Every assertion held: a causal release
rather than an idle expiry, four renewals during the bounded wait, one
resend carrying the original reason, the timeout class at 10m5s, no stale
transaction, and a successor that leased normally a minute later.

Also recorded: a five-phase plan does not lengthen the implement phase, and
the state-file lever cannot be driven with systemctl restart alone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CVbaKucEYBjMqVeUgJUsc1
2026-08-29 17:38:15 +04:00
kami c587f2cc8d Bound the wait for a handoff nobody answers
F62. The rotation is agent-driven: the worker asks, and the agent must write
its handoff. When the agent never does, renewals stopped on the ordinary
progress gate, the lease expired, and the task lost an attempt with nothing on
record saying a handoff had ever been requested. Run 16 showed only "agent
status idle and pane unchanged", 34 times.

The request is now stamped, and the wait around it is bounded. While Orchestra
is explicitly waiting the lease renews, because a quiet pane is the answer the
agent was told to give. The request is re-sent once after four minutes, with
the reason it was first asked with. At ten minutes the worker nacks with
failure class handoff_unanswered, and the coordinator releases the task naming
that cause instead of letting the lease die as generic idleness.

The class is known to DebtClassForFailureClass, so a harness that ignores
handoff requests accumulates as its own debt item rather than hiding inside
lease_expired.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CVbaKucEYBjMqVeUgJUsc1
2026-08-29 16:21:35 +04:00
kami 41658aea5b Carry the resume order and the tag's meaning in the handoff
The next session should not rebuild either from the commits. A defect
found in experimental work is not a reason to reopen settled
architecture.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CVbaKucEYBjMqVeUgJUsc1
2026-08-29 15:51:42 +04:00
kami 765bf2afc6 Hand off with the release path settled and the UI as a truth detector
F57 to F60 close the release-transaction family, F18 closes the single
last_error slot, the debt ledger exists as a read-only projection, and
the operator console is rebuilt on ethos.

The part worth acting on is what the UI could not do honestly: no
web-facing human-decision write path, no keystroke forwarding, context
occupancy trapped in herdr, project configuration unserved, and no
federated handoff request. F61 and F62 are recorded and unbuilt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CVbaKucEYBjMqVeUgJUsc1
2026-08-29 14:59:05 +04:00
kami cbd6b11c49 Record run 16: the rotation rung, half proven
A successor inherits the whole sealed plan. Phase progress is withheld
from a review successor on purpose, so the rung still needs an
implement-phase successor. F62: a requested handoff nobody answers is
invisible, and its only consequence is an expiry indistinguishable from
an idle one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CVbaKucEYBjMqVeUgJUsc1
2026-08-29 14:14:37 +04:00
7 changed files with 787 additions and 4 deletions
+233
View File
@@ -2527,3 +2527,236 @@ transaction and no session behind, on either worker.
Both workers hold no sessions and no release transactions. Deployed pair is
still `3c7cf95` on both halves.
## Run 16, 2026-08-29: the rotation rung, half proven
Task `06G4SWEVP71FYYKAV5FV0ZK5ZG` on `34f3c28`, a three-phase plan
(`78d25730140c`) with automated checks inside policy.
### Proven: a successor inherits the whole sealed plan
The launch context of the session that picked the task up after its rotation
carries the complete plan: overview, current state, desired end state,
non-goals, approach, all three phases with their files, changes and
verification, testing strategy, risks, migration, and ten research citations.
### Not a defect: that successor was not given phase progress
`renderPlanProgress` runs only for `WorkPhaseImplement`
(`internal/agentctx/agentctx.go:321`). This successor picked up in review, and
the code says why: an independent review must reconstruct the change from the
diff rather than inherit the implementer's account. Progress is withheld there
deliberately.
So the rung still needs an **implement-phase** successor. The records exist and
are durable, projected by `Task.PlanPhases()`.
### Why the mid-implement rotation missed
A trivial task spends about four minutes in implement, and all three phases
verified inside two and a half of them, every one against the same tree
`3da86dbf8ad5`. The implementer does the whole change first, then verifies each
phase. The suspend, edit and restart that sets `handoff_requested` takes about
a minute, and it landed after the implementer had already asked for review.
Two mechanics worth keeping:
- **The web handoff action cannot drive this.** `RequestHandoff` needs a local
coordinator and every session here is worker-owned, so it answers 503. That
is the Design A guardrail working, and a fourth UI-exposed gap.
- **A state-file edit does not survive a running worker.** It holds sessions in
memory and writes them back. Suspend, edit, then restart; resuming lets the
old copy win. This cost one attempt to learn.
### F62: a requested handoff that is never answered is invisible
The rotation is agent-driven: the worker asks, and the agent must write
`HANDOFF.md` at a turn boundary. The review agent never did. Renewals stopped,
the lease expired at 10:12:05, and the task requeued having lost an attempt.
Worker health recorded exactly one thing, 34 times:
```text
x34 lease 06G4SWEVP71FYYKAV5FV0ZK5ZG not renewed: agent status idle and pane unchanged
```
Nothing says a handoff was requested and left unanswered. There is no timeout,
no retry, and no observation. The only visible consequence is an expiry that
looks identical to an ordinary idle one. Diagnosing it needed the worker state
file.
The ring earned its keep again: one distinct message with a count of 34, rather
than 34 overwrites of one slot.
**Fixed, not yet proven live.** The request is now stamped
(`herdr.Session.HandoffRequestedAt`) and the wait is bounded by
`watchHandoff` in the worker:
- the lease renews while Orchestra is explicitly waiting, because a quiet pane
is the answer the agent was asked for — the renewal gate's new case, bounded
by `handoffAnswerTimeout`;
- the request is re-sent once at `handoffRetryAfter` (4 minutes), with the same
reason it was first asked with;
- at 10 minutes the worker nacks with failure class `handoff_unanswered`, and
the coordinator emits `TaskReleased reason=handoff_unanswered` rather than
letting the lease die as generic idleness.
`DebtClassForFailureClass` knows the class, so a harness that repeatedly
ignores handoff requests now accumulates in the debt ledger instead of hiding
inside `lease_expired`.
Runtime proof still owed: force a request the agent will not answer, and read
the release event rather than the worker state file.
## Run 17, 2026-08-29: F62 proven live on `c587f2c`
Task `06G4V20T528ZTER7KZBVGNAYXC`, a five-phase plan
(`3b16b77054b9`) on the deployed pair, coordinator and worker both at
`c587f2c`.
### The trigger was production, not a rig
The rig planned for this run was a state-file lever. It was never needed. The
`implement -> review` phase change asked for a handoff through the ordinary
path (`rotateForPhase`, reason `phase_changed`), and the agent ignored it
because this task's brief instructed it to from phase 4 onward. Nothing touched
tmux, no process was stopped, and the pane stayed open for the whole wait.
```text
12:48:09 implement -> review, handoff requested, reason=phase_changed
12:50:09 TaskLeaseRenewed agent idle, pane capture unchanged
12:52:15 pane input confirmed the single resend, 4m06s after the request
12:52:39 TaskLeaseRenewed
12:55:14 TaskLeaseRenewed
12:57:49 TaskLeaseRenewed
12:58:14 TaskReleased reason=handoff_unanswered failure_class=handoff_unanswered
last_error="handoff requested (phase_changed) and unanswered for 10m5s"
12:59:14 TaskLeased successor leases normally
13:02:20 TaskSubmitted
```
Every assertion the operator set for this rung held:
- the release does not look like an ordinary `lease_expired`: it carries
`reason=handoff_unanswered` and `lifecycle_phase=handoff_unanswered`;
- the idle-renewal rule did not kill the wait: four renewals on an idle pane
with no capture movement, each one the gate would have refused before this
commit;
- the resend preserved the reason the request was first made with;
- exactly one resend, not one per tick;
- the timeout produced the class at `10m5s`, one tick past the bound;
- nothing stale was left: `releases: {}`, no session, no lease, and zero tmux
panes matching the id;
- the next attempt leased 60s later and reached `TaskSubmitted`.
The debt projection classified it with no new code beyond the class mapping:
```text
v1:operational:handoff_unanswered:workpc-claude:lease operational important
"handoff requested (phase_changed) and unanswered for 10m5s"
```
**Inference, not observation.** The resent prompt's text was not captured. That
it carried the `phase_changed` wording rests on the code path and the timing,
not on a pane capture. A future rig should publish a capture across the resend.
### The implement-successor rung missed again, and why
A five-phase plan does not lengthen the implement phase. This one ran 12:45:18
to 12:48:09, about 2m50s, with verifications 20 to 40 seconds apart:
```text
12:45:47 phase-1 12:46:07 phase-2 12:46:32 phase-3
12:47:12 phase-4 12:47:52 phase-5 12:48:09 implement -> review
```
The brief's explicit one-phase-at-a-time requirement was followed, and it
still took under three minutes. Phase count is not the lever; volume of
mechanical work is.
`awaiting_manual_verification` does not hold the phase open either. All five
phases carried that status and the agent moved on regardless.
### The state-file lever cannot be driven with `restart` alone
The documented procedure is stop, edit, start. The unattended sudo rule grants
`install` and `systemctl restart orchestra-worker`, with no `stop` or `start`.
With `restart` only, the running worker's five-second tick save landed between
the edit and the restart, and the edit was gone. Widening the sudoers line is
the fix.
### One cosmetic note
`Session.HandoffRequestedAt` serializes as `0001-01-01T00:00:00Z` when unset,
because `omitempty` does not omit a zero `time.Time`. The worker parses it back
to a zero value and `watchHandoff` starts the clock, so behaviour is correct.
It is the same marshalling trap the operator console hit.
## Run 18, 2026-08-29: the implement-successor rung, proven
Task `06G4VF5HZW7Q4JBM3TTY7W1Y64` on `c587f2c`, a five-phase plan whose phases
are ten named checks rather than one refactor. Volume of mechanical work is
what keeps the implement phase open; phase count does not.
### The rotation
```text
13:43:10 implement session launched
13:43:44 phase-1 verified
13:44:04 phase-2 verified
13:44:35 handoff requested, reason=milestone, delivered by watchHandoff's resend
13:45:25 TaskLeaseRenewed
13:46:51 TaskReleased handoff_ref 0b4fae6f705a anchor 8ea3c7d1ae8f
13:46:51 TaskLeased successor, same harness
13:46:55 TaskPickupValidated
13:47:06 TaskLaunchAcknowledged, still in implement
```
The lever was the worker state file, backdated past `handoffRetryAfter` so the
request went out on the next tick. Under `c587f2c` that lever now produces a
real prompt: the flag alone used to sit there unasked, which is how run 16
ended in a silent expiry.
### What the successor was given
From `.orchestra/launch.md`, written at 13:47:05:
```text
## Current phase
implement: Implement the accepted plan below. Verify as you go.
## Verified git state
- head: 8ea3c7d1ae8fd960897344694645117adbd8ee82
- uncommitted changes: false
## Current human decisions
None recorded. Work from the goal and acceptance above.
## Plan progress
Orchestra established this by running the plan's own verification. You cannot
write it.
- phase-1 (...): automated checks passed at d3acd4989a20, stale because the
tree is now at 8ea3c7d1ae8f, waiting for the human to confirm the manual steps
- phase-2 (...): automated checks passed at d3acd4989a20, stale because the
tree is now at 8ea3c7d1ae8f, waiting for the human to confirm the manual steps
- phase-3 (Checks 7 to 8): not started
- phase-4 (...): not started
- phase-5 (...): not started
```
Everything the rung asked for is there: the complete accepted plan with all
five phases, their files, changes and verification; the phases already
verified; the first unfinished phase; and the current human authority. The
launch context also carries the accepted research and its dead ends.
**SHA staleness renders itself.** Neither phase is reported as simply passed.
Each says the checks passed at `d3acd4989a20` and are stale because the tree
has moved to `8ea3c7d1ae8f`. That is half of the manual-verification rung
observed without being asked for.
### The lever still loses a race, sometimes
`systemctl restart` alone leaves a five-second window in which the running
worker's tick save can clobber the edit. It clobbered run 17's attempt and
survived run 18's. Stop, edit, start is the reliable sequence, and it needs
`systemctl stop` and `start` in the unattended sudo rule.
+349
View File
@@ -0,0 +1,349 @@
# Handoff: the release path is settled, and the UI became a truth detector
Written 2026-08-29, 14:20 local (10:20 UTC). Read with `BURNIN.md` (the run
ledger, current through run 16), `DEBT-DESIGN.md`, `PLAN-SPEC-DESIGN.md`,
`AUDIT.md` and `CLAUDE.md`. The previous handoff is
`HANDOFF-2026-08-28-plan-v1.md`.
Everything below was observed live unless it says otherwise.
## The headline
**The expired-release defect is fixed and proven both ways.** The whole family
around it is closed. F57 through F60 settle what happens to a release
transaction in every case. That includes the ones that used to need an operator
with a text editor.
**Three new things exist that did not before.** A bounded observation ring on
worker health, closing F18. A read-only debt ledger projected from the event
log. An operator console rebuilt on the ethos design system.
**The UI turned out to be a truth detector.** Nine screens were built against
real endpoints. They found four places where Orchestra has no capability to
support the intended interface. That list is the most valuable output of the
session.
```text
orchestra-plan-v1 plan machinery proven
orchestra-f18-baseline d6ee10f, the bounded observation ring
↓ 6 commits
34f3c28 deployed now: release path settled, debt ledger, new UI
```
## Deployed state
| Half | Revision |
|---|---|
| Coordinator, homesrv container | `34f3c28` |
| Worker, workpc systemd | `34f3c28` |
```text
commit 34f3c2888fc7d45d190d93e1ea42501b6cd3e474
coordinator sha256 1d32d83ec859b36e12473fab01bbfc3c769b97b4f6ae4db2841db5161e0eda19
worker sha256 0e3877321dea8a1eeda51ccc6f3ead1a14aa5f5cae4d95b704f61c364248ec65
```
`cbd6b11` is one commit above and is documentation only. Three commits are
unpushed.
**Worker installs no longer need a human.** The operator installed the
`/etc/sudoers.d` line, so `sudo -n install …` and
`sudo -n systemctl restart orchestra-worker` both work unattended. Verify the
running revision from the journal, never the installed file.
## The defects fixed, and how each was found
Not one came from reading code. Every one came from a live run failing.
| Id | Commit | What |
|---|---|---|
| F57 | `6565b9f` | An expired lease could never commit the anchor it had already pushed. The worker sent an epoch the expiry replay had deleted, and the coordinator refused any handoff without a live lease. The epoch now belongs to the transaction, `TaskReleased` retains the ending epoch, and `lateHandoffAccepted` lets exactly that owner commit while the task is queued and unleased. |
| F58 | `03663f4` | A superseded transaction retried a permanent 409 forever, holding the pane and pinning `ActiveTask`. Run 10's task did it for seven hours. `TaskLeased` now abandons a transaction whose id the lease does not carry. |
| F59 | `8e37989` | F58 fires on `TaskLeased`, and a failed task is never leased again. `TaskFailed` now drops the transaction too. |
| F60 | `3c7cf95` | The general rule the other two were reaching for. Terminal is failed or completed. Blocked keeps the transaction, because a reopen can still commit it, so `TaskBlocked` now retains the ending epoch as well. A refusal parks the commit for 30s backing off to 5 minutes, and any event about the task un-parks it. A transport failure is not an answer and retries at once. |
| F18 | `d6ee10f` | The single `last_error` slot. Worker health now carries up to sixteen distinct observations with repeat counts and first/last times, collapsing by message rather than by position. |
### The rig that proved F57, and the guard
```text
19:00:10.742 transaction opens at prepared, anchor pushing
19:00:10.727 TaskReleased v15 reason=lease_expired surface=tui
19:00:11.662 TaskReleased v16 the late commit, accepted 935ms after the lease died
19:00:11.665 TaskLeased v17 successor picks up the handoff
19:00:15.524 TaskPickupValidated v18
```
Race guard, next boundary: force the expiry, then lease the task to a probe
harness before the push finishes. The late commit is refused, no handoff is
written, and the successor's lease stands.
**The rig technique matters more than the rig.** Suspending the worker cannot
produce this ordering. The event replay runs at the top of every tick and
discards the transaction. The ordering exists only inside one call:
transaction opened, anchor pushing, commit not yet sent. So poll the worker
state file at 2ms and fire `POST /v1/tasks/<id>/release` the instant a
transaction appears at `prepared`.
**A named probe harness owns a lease without starting an agent.**
`race-guard-probe` never picks anything up and expires on the normal TTL.
## Corrections to the previous handoff
**The operator lifecycle actions do not lose a version race.** On a leased
task, `block`, `release` and `attention` are refused by
`internal/store/store.go:885-901` when the payload omits `harness_id` and
`lease_epoch`. Ten attempts in 550ms all failed that way. Send both fencing
fields and they succeed on the first try.
**The OpenCode Zen free tier is not blocked.** The selected model was.
## The debt ledger
`DEBT-DESIGN.md` answers nine design questions and carries four amendments the
operator made. Slice one is built, deployed and run.
**Slice one writes nothing.** A read-only projection over the existing log,
plus a pure eligibility function and `GET /v1/debt`. It folded 881 events and
produced 15 candidates and 3 gaps.
```text
v1:operational:lease_expired:workpc-opencode:lease r=41 tasks=4
v1:operational:lease_expired:workpc-claude:lease r=29 tasks=13
v1:operational:lease_failure:-:lease r=20 tasks=16
v1:correctness:handoff_validation:-:lease r=5 tasks=5
```
The opencode failure shape is the top item, found mechanically. Run 14 reached
the same conclusion by hand from a pane capture.
**It reported what it cannot see, which was the point.** The 409 release loop
does not appear. That evidence lived in the F18 worker ring, and no event
carries it. Manual interventions are a non-durable gap for the same reason.
**The first run exposed four defects in the model**, all recorded at the end of
`DEBT-DESIGN.md`:
- the component part is too coarse for lease evidence
- `harness` is often empty on block reasons
- path normalization mangled a mismatch reference
- recurrence alone is the wrong sort order
Layout, so the read model does not end up in the command layer:
```text
internal/domain/debt.go types, signature, classification
internal/store/debt_projection.go the fold, and the gap report
internal/operations/debt.go CheckDebtEligibility
```
## The operator console
Nine screens on the ethos system, signal violet `#8F7AE5`, routing fork motif.
Each screen was built by its own agent against a foundation with one author.
The shell, tokens and primitives could not drift into nine dialects.
**Render before signing off.** Three bugs existed that no computed value would
have caught. All three came from looking at a screenshot:
- The previous stylesheet fought every shared class name and leaked properties
the new rules never mention, which is how `position: fixed` survived on
`.topbar`. It is now scoped under `.legacy` and reaches only the login route.
That also stops its green accent and its `backdrop-filter` from reaching the
console.
- Go marshals a zero `time.Time` as `0001-01-01T00:00:00Z` and `omitempty` does
not omit a struct, so absent timestamps arrived populated-looking and
rendered as `739855d ago`. Stripped once in `client.ts`, with a test.
- Long machine ids overflowed their cards and painted under the next one.
Chromium is installed at `/usr/bin/chromium`. To see a screen without a live
session, write a throwaway harness that stubs `window.fetch` and renders
`<Console>` inside a `MemoryRouter`, served by vite on a spare port. Note that
`npx` and `./node_modules/.bin/*` do not work on this filesystem: call
`node ./node_modules/vite/bin/vite.js` directly.
## What the UI proved Orchestra cannot do
This is the part worth acting on. Each screen refused to fake something, and
the refusals name real capability gaps.
| Gap | Evidence |
|---|---|
| **No web-facing human-decision write path** | `Steer / Correct` is disabled. `internal/ui/ui.go`'s action switch has grant/deny approval, resubmit, handoff, release, block and complete, and nothing writes `HumanDecisionRecorded`. The spec makes steering the primary action of the task detail screen. |
| **No keystroke forwarding** | `Take control` is disabled. Only resubmit and approval grant/deny reach a live pane. |
| **Context occupancy is trapped in herdr** | Three screens independently hit it. No projection carries it. |
| **Project configuration is not served** | Repo, remote, quality gate and verification policy live only in `config.jsonc`. The projects screen can show none of it. |
| **The web cannot request a handoff for a federated task** | `RequestHandoff` needs a local coordinator and answers 503. That is the Design A guardrail working. |
The operator's direction on these. Treat first-class direct human input as the
highest-value backend feature. Build it as `POST /v1/tasks/<id>/decisions`,
using the same durable decision semantics as Gitea comments, so Gitea, CLI and
web converge on one `HumanDecisionRecorded`. Keep take-control disabled,
because arbitrary pane input bypasses the durable authority model. Expose
occupancy through a session health projection rather than teaching the web
server about herdr. Add a read-only effective project configuration endpoint,
which F61 will also need.
## Where the plan-machinery ladder stands
Proven in run 14: the worker executes the **sealed plan's** commands rather
than the request's, and every `PlanPhaseVerified` binds `plan_ref`, `phase_id`,
`at_sha`, `evidence_ref`, `lease_epoch` and `harness_id`.
Proven in run 16: a successor inherits the **whole sealed plan**, all phases
with their files, changes, verification, and the research citations.
**Still unproven, and the next runtime item:**
```text
mid-implement rotation
→ successor picks up in implement
→ launch context states which phases are already verified
manual verification
→ SHA goes stale
→ re-verification
plan mismatch
→ human decision
→ real replan, old plan retained, replacement launched
```
Two things make the first one hard, and both are now known:
- **A trivial task spends about four minutes in implement**, and verifies every
phase against one tree near the end. The implementer writes the whole change
first, then verifies each phase in turn. Use a task whose implement phase
genuinely runs long.
- **A state-file edit does not survive a running worker.** It holds sessions in
memory and writes them back. Suspend, edit, then `sudo systemctl restart`.
Resuming lets the old copy win. Setting `handoff_requested` on the session is
the production rotation lever.
Progress is withheld from a **review** successor on purpose
(`internal/agentctx/agentctx.go:321`), because an independent review must
reconstruct the change from the diff. That absence is not the defect.
## F61 and F62, recorded and not built
**F61: the planner learns the verification policy by refusal.** The brief says
a policy exists, not what is in it. Every plan therefore pays one refused round
trip. The recovery loop works: run 14's planner consumed the refusal and
resealed 31 seconds later. It is not a one-liner, because `agentctx.Input.Policy` is filled
from the worker's `SafeOperations` while the verification policy is
coordinator-side. The reason to promote it later is local models, which may
propose forbidden commands repeatedly because they cannot infer the allowed
substitute.
**F62: a requested handoff nobody answers is invisible.** The rotation is
agent-driven. In run 16 the agent never wrote `HANDOFF.md`, renewals stopped,
the lease expired, and the task lost an attempt. Worker health recorded only
`agent status idle and pane unchanged`, 34 times. There is no timeout, no
retry, and no observation saying a handoff was requested and left unanswered.
The expiry is indistinguishable from an ordinary idle one.
## Live state
Three pull requests are open and unreviewed:
```text
06G4M6HF1Z3EREX1X3NEKSHP24 pulls/18
06G4M8WHGQ4P3GQMPEEH0RJRHM pulls/19
06G4SWEVP71FYYKAV5FV0ZK5ZG pulls/20
```
Both workers are online on `34f3c28` with no release transactions and no
sessions. 29 blocked `test-e2e` tasks are burn-in debris. Three `correx` tasks
are queued and unschedulable, because `correx` has no entry in the
coordinator's `config.jsonc`.
**opencode now runs.** The model was the problem, and `hy3-free` works. It then
stops on a permission prompt, because `~/.config/opencode/opencode.jsonc` sets
`"bash": "ask"`. Orchestra can answer that exact dialog, but only when an
operator queues `grant_approval`, so unattended work stalls on the first
command. Set `"bash": "allow"` for unattended runs. That file also has no
top-level `model` key, so OpenCode picks whatever sits at the top of
`~/.local/state/opencode/model.json`, which any manual pick silently changes.
**The opencode adapter is now a debt item, not a curiosity.** It cannot resolve
a session file. Activity therefore reads `unknown`, and the worker cannot tell
finished from never-started. The debt projection surfaced it independently as
the top recurring operational item.
## Things that will bite
- **Background python tasks get killed here.** Three watchers died before doing
anything. Foreground polling and the `Monitor` tool both work.
- **`/v1/events` is one line of JSON.** A `grep` for two substrings matches
across unrelated tasks. Parse it.
- **`npx` and `./node_modules/.bin/*` fail on this filesystem.** Call node
directly.
- **`rm` and `cp` are interactive.** Use `/bin/rm -f` and `install`.
- **Secrets are guarded.** Expand a token inside the container in one remote
command:
`T=$(docker exec orchestra-api printenv ORCHESTRA_TUI_TOKEN); curl -s -H "X-Orchestra-Surface: tui" -H "Authorization: Bearer $T" ...`
- **Rebuild the coordinator from a detached worktree**, and split the worktree
add, the docker build and the compose up into separate commands.
- **`deploy/build.sh` builds both halves from one commit with one stamp.** Use
it.
- **Geist is not on disk.** Both stacks fall back to system faces, and the
ethos threat model rules out the font CDN.
## Resume in this order
Set by the operator at the session boundary. Do not rebuild it from the commits.
1. **F62 first.** Make a requested-but-unanswered handoff visible and bounded.
Preserve the lease while Orchestra is explicitly waiting, retry the confirmed
request, then emit a causal timeout. Today it ends as generic idleness.
2. **Repeat the implement-successor rung** with a deliberately longer task. Get
at least one `PlanPhaseVerified`, force the handoff while still in
`implement`, and prove the successor sees the complete accepted plan, the
verified previous phase, the first unfinished phase, and the current human
authority.
3. **Finish the remaining ladder.** Manual verification, SHA staleness and
reverification, human-decision mismatch, then a real replan with old-plan
provenance.
4. **Use the capability table above as backend work discovery.** Do not add fake
controls. Each disabled action is concrete evidence of a missing capability.
5. **Continue the debt slices independently.** Durable worker observations and
manual-intervention events are the next evidence gaps. Not automatic
maintenance yet.
## What the tag means
```text
orchestra-release-v1 -> 34f3c28
deployed, settled lifecycle plus the truth-detector UI baseline
later HEADs
experimental plan, debt and runtime work that must earn their own
release proof
```
Use that distinction. A defect found in experimental work is not a reason to
reopen settled architecture.
## The roadmap the operator set
```text
A. runtime correctness finish the plan-machinery live proof
B. evidence and debt durable worker observations, manual-intervention
events, then rerun the projection
C. operator surface first-class HumanDecision write API,
effective project-config read API,
session and context health projection
D. adapter fix opencode activity and session resolution
E. UI wire capabilities as backend support becomes real
```
The framing that ties them together is the thing to keep. The UI says what
Orchestra cannot expose or control. The debt ledger says which of those
shortcomings repeatedly costs something. The burn-in says which runtime
semantics are reliable. Those three decide what gets built next.
One design question to settle before B's slice two writes any code. The worker
ring is bounded and lossy by construction. Ingesting it durably means deciding
whether the coordinator stores every observation as an event, or only
transitions. Storing every heartbeat's ring would write the same 41-count
observation hundreds of times.
+110
View File
@@ -0,0 +1,110 @@
package main
import (
"context"
"encoding/json"
"net/http"
"net/http/httptest"
"path/filepath"
"strings"
"testing"
"time"
"orchestra/internal/domain"
"orchestra/internal/federation"
"orchestra/internal/herdr"
)
// F62. Run 16: the agent was asked to hand off, never wrote HANDOFF.md,
// renewals stopped, and the lease died as ordinary idleness. Waiting is now
// bounded: re-ask once, then give the task up with a class that says why.
func TestUnansweredHandoffIsRetriedThenGivenUp(t *testing.T) {
var nack map[string]any
w, backend, _, done := phaseWorker(t, func(rw http.ResponseWriter, r *http.Request) {
if strings.HasSuffix(r.URL.Path, "/nack") {
_ = json.NewDecoder(r.Body).Decode(&nack)
}
rw.Write([]byte(`{}`))
})
defer done()
ctx := context.Background()
requested := func(ago time.Duration) herdr.Session {
s := w.sessions["task"]
s.HandoffRequested, s.HandoffReason = true, "phase_changed"
s.HandoffRequestedAt = time.Now().UTC().Add(-ago)
w.sessions["task"] = s
return s
}
// Still inside the answering window: nothing said, nothing given up.
if s, gaveUp := w.watchHandoff(ctx, "task", requested(time.Minute)); gaveUp || s.HandoffRetried {
t.Fatalf("gave up while still waiting: gaveUp=%v session=%+v", gaveUp, s)
}
if len(backend.prompts) != 0 {
t.Fatalf("re-asked too early: %q", backend.prompts)
}
// Past the retry point: asked again, exactly once.
s, gaveUp := w.watchHandoff(ctx, "task", requested(handoffRetryAfter+time.Minute))
if gaveUp || !s.HandoffRetried || len(backend.prompts) != 1 {
t.Fatalf("retry: gaveUp=%v retried=%v prompts=%q", gaveUp, s.HandoffRetried, backend.prompts)
}
if _, gaveUp = w.watchHandoff(ctx, "task", s); gaveUp || len(backend.prompts) != 1 {
t.Fatalf("re-asked every tick: %q", backend.prompts)
}
// Past the bound: a causal reclaim, and no lease left to renew.
s = requested(handoffAnswerTimeout + time.Second)
s.HandoffRetried = true
w.sessions["task"] = s
if _, gaveUp = w.watchHandoff(ctx, "task", s); !gaveUp {
t.Fatal("an unanswered handoff waited forever")
}
if nack["failure_class"] != "handoff_unanswered" {
t.Fatalf("nack = %+v", nack)
}
if detail, _ := nack["last_error"].(string); !strings.Contains(detail, "phase_changed") {
t.Fatalf("the reclaim does not name the request: %q", detail)
}
if _, held := w.leases["task"]; held {
t.Fatal("the given-up task kept its lease")
}
}
// The lease must survive the wait it was asked to make: an idle pane is the
// answer Orchestra requested, not evidence of an agent that stopped working.
func TestWaitingForAHandoffKeepsTheLease(t *testing.T) {
renewals := 0
api := httptest.NewServer(http.HandlerFunc(func(rw http.ResponseWriter, r *http.Request) {
renewals++
rw.Write([]byte(`{}`))
}))
defer api.Close()
backend := &recordingBackend{status: "idle", progress: "same screen"}
w := &worker{
api: federation.Client{BaseURL: api.URL, WorkerID: "h", Token: "t"},
backend: backend,
harness: "claude",
sessions: map[string]herdr.Session{"task": {PaneID: "pane", HandoffRequested: true, HandoffRequestedAt: time.Now().UTC()}},
leases: map[string]lease{"task": {Epoch: "e", Version: 1, Until: time.Now(), ProgressSHA: domain.Hash([]byte("same screen"))}},
quarantined: map[string]bool{},
statePath: filepath.Join(t.TempDir(), "state.json"),
}
w.renewLeases(context.Background())
if renewals != 1 {
t.Fatalf("a lease waiting on a requested handoff renewed %d times, want 1", renewals)
}
// Past the bound the exemption stops: watchHandoff has given the task up
// by then, and nothing keeps an unanswered request alive.
s := w.sessions["task"]
s.HandoffRequestedAt = time.Now().UTC().Add(-handoffAnswerTimeout - time.Second)
w.sessions["task"] = s
l := w.leases["task"]
l.Until = time.Now()
w.leases["task"] = l
w.renewLeases(context.Background())
if renewals != 1 {
t.Fatalf("the exemption outlived its bound: renewals=%d", renewals)
}
}
+80 -3
View File
@@ -174,6 +174,7 @@ func releaseBackoff(attempts int) time.Duration {
}
return d
}
type projectConfig struct {
Repo string `json:"repo"`
Root string `json:"worktree_root"`
@@ -673,6 +674,13 @@ func (w *worker) releaseReady(ctx context.Context) {
w.advanceRelease(ctx, id, s)
continue
}
if s.HandoffRequested {
next, gaveUp := w.watchHandoff(ctx, id, s)
if gaveUp {
continue
}
s = next
}
w.rotationTick(ctx, id, s)
}
}
@@ -772,7 +780,7 @@ func (w *worker) rotationTick(ctx context.Context, id string, s herdr.Session) {
w.recordError(fmt.Errorf("rotation %s threshold prompt: %w", id, err))
return
}
s.HandoffRequested, s.HandoffReason = true, d.Reason
s.HandoffRequested, s.HandoffReason, s.HandoffRequestedAt = true, d.Reason, time.Now().UTC()
w.sessions[id] = s
_ = w.save()
}
@@ -1166,6 +1174,70 @@ func (w *worker) submit(ctx context.Context, id string, s herdr.Session, e compl
// the directory also holds for a worktree that has no inner .gitignore.
const stageExclude = ":!.orchestra"
// F62: a requested handoff nobody answers was invisible. Renewals stopped,
// the lease expired, and the task lost an attempt with nothing on record
// saying a handoff had ever been asked for — worker health showed only "agent
// status idle and pane unchanged", 34 times in run 16.
const (
handoffRetryAfter = 4 * time.Minute
handoffAnswerTimeout = 10 * time.Minute
)
// watchHandoff bounds the wait for an agent's handoff answer: re-send the
// request once, then give the task up with a class that names the cause. It
// returns the session to keep using and whether the task was given up.
func (w *worker) watchHandoff(ctx context.Context, id string, s herdr.Session) (herdr.Session, bool) {
if s.HandoffRequestedAt.IsZero() {
// A session persisted before the stamp existed, or requested by a path
// that does not set it. Start the clock now rather than time out a
// request retroactively.
s.HandoffRequestedAt = time.Now().UTC()
w.sessions[id] = s
_ = w.save()
return s, false
}
waited := time.Since(s.HandoffRequestedAt)
if waited < handoffRetryAfter {
return s, false
}
if waited < handoffAnswerTimeout {
if s.HandoffRetried || w.executionBackend() == nil {
return s, false
}
a := herdr.CLIAdapter{Backend: w.executionBackend(), Harness: w.harness}
var err error
if s.HandoffReason != "" {
err = a.RequestHandoffReason(ctx, s, s.HandoffReason, nil)
} else {
err = a.RequestHandoff(ctx, s)
}
if err != nil {
w.recordError(fmt.Errorf("handoff %s re-request: %w", id, err))
return s, false
}
s.HandoffRetried = true
w.sessions[id] = s
_ = w.save()
w.recordError(fmt.Errorf("handoff %s (%s) unanswered for %s: request re-sent", id, s.HandoffReason, waited.Round(time.Second)))
return s, false
}
l, ok := w.leases[id]
if !ok {
return s, false
}
detail := fmt.Sprintf("handoff requested (%s) and unanswered for %s", s.HandoffReason, waited.Round(time.Second))
if err := w.api.NackStart(ctx, id, l.Epoch, l.Version, "handoff_unanswered", detail, w.sessionEvidence(ctx, id, s)); err != nil {
w.recordError(fmt.Errorf("handoff timeout %s: %w", id, err))
return s, false
}
// The coordinator answers with TaskReleased; its replay quarantines the
// pane. Drop the lease here so nothing renews it in the meantime.
delete(w.leases, id)
_ = w.save()
w.recordError(errors.New(detail))
return s, true
}
func (w *worker) renewLeases(ctx context.Context) {
if w.executionBackend() == nil {
return
@@ -1199,6 +1271,11 @@ func (w *worker) renewLeases(ctx context.Context) {
case l.ProgressSHA == "":
// First renewal has no baseline to compare against. Record one and
// allow this renewal; the next one must show real movement.
case s.HandoffRequested && !s.HandoffRequestedAt.IsZero() && time.Since(s.HandoffRequestedAt) < handoffAnswerTimeout:
// Orchestra told this agent to stop and write its handoff. A quiet
// pane is the answer it was asked for, so the lease is held while
// the wait is explicitly bounded (F62). Past the bound the case
// stops matching and watchHandoff has already given the task up.
default:
w.recordError(fmt.Errorf("lease %s not renewed: agent status %s and pane unchanged since the last renewal", taskID, status))
continue
@@ -1932,7 +2009,7 @@ func (w *worker) federatedTurn(ctx context.Context, id string, a herdr.Adapter,
w.recordError(fmt.Errorf("reconcile failure handoff %s: %w", id, err))
return
}
s.HandoffRequested, s.HandoffReason = true, "reconcile_failure"
s.HandoffRequested, s.HandoffReason, s.HandoffRequestedAt = true, "reconcile_failure", time.Now().UTC()
w.sessions[id] = s
_ = w.save()
return
@@ -2009,7 +2086,7 @@ func (w *worker) rotateForPhase(ctx context.Context, id string, a herdr.Adapter,
w.recordError(fmt.Errorf("phase rotation %s: %w", id, err))
return
}
s.HandoffRequested, s.HandoffReason = true, "phase_changed"
s.HandoffRequested, s.HandoffReason, s.HandoffRequestedAt = true, "phase_changed", time.Now().UTC()
w.sessions[id] = s
_ = w.save()
log.Printf("phase changed for %s: session rotating", id)
+6
View File
@@ -1589,6 +1589,12 @@ func main() {
case "invalid_handoff":
typ = "TaskBlocked"
p, _ = json.Marshal(map[string]any{"blocker": b.LastError, "block_reason": string(domain.BlockReasonHandoffValidation), "harness_id": parts[3], "lease_epoch": b.LeaseEpoch, "expected_version": t.Version, "lifecycle_phase": "launch_nacked", "last_error": b.LastError, "session_evidence": b.SessionEvidence})
case "handoff_unanswered":
// F62: not a launch failure. The agent was asked to hand off
// and never did, so the reclaim says exactly that instead of
// arriving as an ordinary idle expiry.
typ = "TaskReleased"
p, _ = json.Marshal(map[string]any{"reason": "handoff_unanswered", "failure_class": b.FailureClass, "harness_id": parts[3], "lease_epoch": b.LeaseEpoch, "expected_version": t.Version, "lifecycle_phase": "handoff_unanswered", "last_error": b.LastError, "session_evidence": b.SessionEvidence})
case "launch_uncertain":
typ = "TaskNeedsAttention"
p, _ = json.Marshal(map[string]any{"blocker": b.LastError, "block_reason": string(domain.BlockReasonLeaseFailure), "harness_id": parts[3], "lease_epoch": b.LeaseEpoch, "expected_version": t.Version, "lifecycle_phase": "launch_uncertain", "last_error": b.LastError, "session_evidence": b.SessionEvidence})
+1 -1
View File
@@ -251,7 +251,7 @@ func DebtClassForBlockReason(r BlockReason) (DebtClass, bool) {
// classes a worker actually emits are listed; an unknown one is not guessed at.
func DebtClassForFailureClass(f string) (DebtClass, bool) {
switch f {
case "retry_limit", "launch_failed", "launch_transient", "launch_uncertain", "prompt_not_submitted", "lease_expired":
case "retry_limit", "launch_failed", "launch_transient", "launch_uncertain", "prompt_not_submitted", "lease_expired", "handoff_unanswered":
return DebtOperational, true
case "invalid_handoff":
return DebtCorrectness, true
+8
View File
@@ -174,6 +174,14 @@ type Session struct {
// its §6.1 handoff (HandoffFile) — avoids re-sending the same prompt
// every tick while Release keeps waiting for the file to appear.
HandoffRequested bool `json:"handoff_requested,omitempty"`
// HandoffRequestedAt stamps that prompt. A request nobody answers used to
// end as an ordinary idle expiry, indistinguishable from an agent that
// never started (F62); the stamp is what makes the wait bounded and the
// giving-up causal.
HandoffRequestedAt time.Time `json:"handoff_requested_at,omitempty"`
// HandoffRetried records that the request was re-sent once, so a session
// waiting on an answer is not re-prompted every tick.
HandoffRetried bool `json:"handoff_retried,omitempty"`
// HandoffReason is selected by the coordinator when it asks for the
// semantic report. The checkout worker, rather than the harness, copies
// it into the canonical handoff it seals at release time.