Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011xsXyr5J1RACo71YeKG3Pu
7.6 KiB
Handoff — run 4 failed, four fixes landed, one deployment step outstanding
Written 2026-08-27, 14:00 UTC. Read with BURNIN.md (the ledger, current
through run 4 and F28), HANDOFF-2026-08-27-burnin-2.md (the session before
this one), AUDIT.md and CLAUDE.md.
Everything below was observed live unless it says otherwise.
Evidence standard, unchanged
event log → lifecycle truth
worker journal → worker-local observations and confirmer receipts
pane capture → harness evidence only
code-path argument → supporting evidence, not live proof
This session broke that rule twice and both are corrected in BURNIN.md. It
reported F23 as broken after grepping only internal/operations, and it
proposed KillMode=mixed for F28 without checking what systemd actually does
after the main process exits. Check the whole call graph, and check the manual.
Deployed state
| Half | Revision | Evidence |
|---|---|---|
| Coordinator, homesrv container | fbaaf79 |
docker logs orchestra-api, 13:30:33 UTC, dirty false, /readyz 200 |
| Worker, workpc systemd | fbaaf79 |
worker registrations, both online |
Branch webui-and-audit-reconciliation, HEAD daa5d20. Worker sha256
28bfe635e5311d631ad12cfdb826d80fbfc5e5c78eb12cec2f5a369670ef0964, staged at
~/orchestra-deploy/orchestra-worker.fbaaf79.
daa5d20 is systemd unit files and ledger only. No rebuild is needed for it.
The one thing that is not done
F28's unit split is written but not installed. Until it is, every worker restart destroys every agent pane, and no clean conformance run is possible.
sudo install -m 0644 \
/home/kami/orchestra-deploy/orchestra-tmux.service \
/etc/systemd/system/orchestra-tmux.service
sudo systemctl daemon-reload
sudo systemctl enable --now orchestra-tmux
The installed orchestra-worker.service also needs the two ordering lines from
deploy/orchestra-worker.service, then daemon-reload.
The running worker still owns the current tmux server. The split takes hold
only once that server is gone and orchestra-tmux owns the next one.
User must match between the units, because the socket is /tmp/tmux-$UID.
Installed worker runs as kami; deploy/orchestra-worker.service still says
orchestra, which is pre-existing drift. The staged tmux unit says kami.
What run 4 established
Task 06G46P6KE25Y04VVF7VRZMHZ78, issue kami/test-e2e#4. Failed conformance.
Cause: F25. Its front half is the cleanest baseline so far.
| seq | at (UTC) | event |
|---|---|---|
| 370 | 13:11:29 | TaskCreated v1, 7 acceptance criteria |
| 371 | 13:11:30 | TaskLeased v2 |
| 372 | 13:11:33 | TaskLaunchAcknowledged v3 |
Issue filed to confirmed launch: 4 seconds, no operator lifecycle intervention.
One worker-journal line at 13:32:16, after the fbaaf79 restart, live-proved
four things at once:
deliver phase refusal 06G46P6KE25Y04VVF7VRZMHZ78: tmux send-keys ... -l -- Orchestra refused your phase request: phase request refused: task 06G46P6KE25Y04VVF7VRZMHZ78 may only move to "research", not "implement"
...: no server running on /tmp/tmux-1000/orchestra: exit status 1
F25 (the boundary ran on claude at all), F21's reading half, F26's premise (the
coordinator naming the one valid target), F27 (409 classified, routed to
sendPrompt). The request file survived the failed send, the correct branch.
The receipt is still unproven, because F28 had already destroyed the pane.
Ledger
F6 closed
F14 closed
F15 closed by detection, transport fix pending live proof
F16 closed, both branches live-proven (trigger was self-inflicted, see F28)
F17 fixed, isolated live proof
F18 open observability, non-blocking
F19 fixed
F20 fixed, live receipt still pending
F21 reading and artifact halves live-proven; accepted transition pending
F22 fixed, live proof pending
F23 closed as already implemented; was undeliverable on claude until F25
F24 open correctness, dormant on current topology
F25 fixed, live-proven
F26 fixed, live-proven
F27 fixed, live-proven
F28 fix written, NOT INSTALLED. Blocks every clean run.
Next session, in order
- Install the F28 units. Nothing else is worth doing first.
- Run the F28 proof, both directions:
start task, obtain live pane P, record identity
systemctl restart orchestra-worker
→ pane P still exists with identical agent identity
→ worker reconciles lease + pane P
→ no TaskReleased or new lease caused merely by the restart
→ agent continues under the same lease epoch
then, separately:
restart orchestra-tmux
→ pane disappears
→ F16 refuses renewal
→ lease expires and requeues
The second half restores the meaning of F16's missing-pane branch: real execution-runtime loss rather than deployment killing its own child.
- File a fresh issue for run 5, the next clean conformance run. Run 4 is closed; do not shepherd it. The run 5 chain to prove:
frame → phase-request.json → WorkPhaseChanged → phase_changed rotation
research → research.json sealed → request → rotate
plan → accepted research visible
→ post ONE trusted issue comment while plan is live, before it asks to leave
→ HumanDecisionRecorded → RemoteTurn → DecisionNotice → confirmer receipt, once
plan.json sealed → request → rotate
implement launch carries: decision above accepted plan above accepted research
Withhold the correction comment until plan is actually live. Posting during frame or research still tests F20 and F23, but it stops proving that a mid-plan correction outranks accepted research and the current trajectory.
Suggested comment, from the operator:
for --json, use snake_case keys. keep the default text output byte-for-byte unchanged.
- Classify separately, as the operator asked: a
WorkPhaseChangedwith no request artifact, and a request artifact with no transition, are different defects from a failed rotation.
Things that will bite
- Another session still owns 16 uncommitted paths, including
AUDIT.md,deploy/build.shand theweb/frontend. Commit by path. Nevergit add -A. - Run 4's worktree still holds
{"from":"frame","to":"implement"}at/tmp/test-e2e-worktrees/06G46P6KE25Y04VVF7VRZMHZ78/.orchestra/. Harmless, and useful if that task is ever retried diagnostically. w.recordErrorkeeps one slot (cmd/orchestra-worker/main.go:70). It overwrites, so a causal sequence cannot be read from worker health. That is F18, still open, and it cost this session real diagnosis time twice.- Gitea writes without reading the token: run curl on homesrv and expand the
value there.
ssh kami@192.168.1.104 'TOKEN=$(docker exec orchestra-api printenv ORCHESTRA_GITEA_TOKEN); curl -H "Authorization: token $TOKEN" ...'Basehttps://gitea.kvmx.ru, ownerkami, repotest-e2e. - Do not use
docker compose build. Build from a detached worktree, asdeploy/build.shdoes and asHANDOFF-2026-08-26-burnin.mdshows. - Installing the worker and the units needs root, so it is always the operator's step. Hand them the command with the checksum.
- This machine is workpc,
hostnameisbugmachine. homesrv iskami@192.168.1.104with/usr/bin/ssh, not thesshon PATH. /tmpis tmpfs with a 10-day sweep. Check/tmp/test-e2eand/tmp/test-e2e-worktreesexist before every run.- Worker capacity is 1 per harness. A leased task on
workpc-claudeblocks the next one, and a Gitea task carries no capability, so the router may hand it toworkpc-opencodeinstead. Wait for a free claude slot before filing a run's issue, or the conformance evidence lands on the wrong harness. - Three queued
corrextasks will never lease. No worker declares that project.