Files
orchestra/HANDOFF-2026-08-27-burnin-3.md
T

7.6 KiB

Handoff — run 4 failed, four fixes landed, one deployment step outstanding

Written 2026-08-27, 14:00 UTC. Read with BURNIN.md (the ledger, current through run 4 and F28), HANDOFF-2026-08-27-burnin-2.md (the session before this one), AUDIT.md and CLAUDE.md.

Everything below was observed live unless it says otherwise.

Evidence standard, unchanged

event log       → lifecycle truth
worker journal  → worker-local observations and confirmer receipts
pane capture    → harness evidence only
code-path argument → supporting evidence, not live proof

This session broke that rule twice and both are corrected in BURNIN.md. It reported F23 as broken after grepping only internal/operations, and it proposed KillMode=mixed for F28 without checking what systemd actually does after the main process exits. Check the whole call graph, and check the manual.

Deployed state

Half Revision Evidence
Coordinator, homesrv container fbaaf79 docker logs orchestra-api, 13:30:33 UTC, dirty false, /readyz 200
Worker, workpc systemd fbaaf79 worker registrations, both online

Branch webui-and-audit-reconciliation, HEAD daa5d20. Worker sha256 28bfe635e5311d631ad12cfdb826d80fbfc5e5c78eb12cec2f5a369670ef0964, staged at ~/orchestra-deploy/orchestra-worker.fbaaf79.

daa5d20 is systemd unit files and ledger only. No rebuild is needed for it.

The one thing that is not done

F28's unit split is written but not installed. Until it is, every worker restart destroys every agent pane, and no clean conformance run is possible.

sudo install -m 0644 \
  /home/kami/orchestra-deploy/orchestra-tmux.service \
  /etc/systemd/system/orchestra-tmux.service
sudo systemctl daemon-reload
sudo systemctl enable --now orchestra-tmux

The installed orchestra-worker.service also needs the two ordering lines from deploy/orchestra-worker.service, then daemon-reload.

The running worker still owns the current tmux server. The split takes hold only once that server is gone and orchestra-tmux owns the next one.

User must match between the units, because the socket is /tmp/tmux-$UID. Installed worker runs as kami; deploy/orchestra-worker.service still says orchestra, which is pre-existing drift. The staged tmux unit says kami.

What run 4 established

Task 06G46P6KE25Y04VVF7VRZMHZ78, issue kami/test-e2e#4. Failed conformance. Cause: F25. Its front half is the cleanest baseline so far.

seq at (UTC) event
370 13:11:29 TaskCreated v1, 7 acceptance criteria
371 13:11:30 TaskLeased v2
372 13:11:33 TaskLaunchAcknowledged v3

Issue filed to confirmed launch: 4 seconds, no operator lifecycle intervention.

One worker-journal line at 13:32:16, after the fbaaf79 restart, live-proved four things at once:

deliver phase refusal 06G46P6KE25Y04VVF7VRZMHZ78: tmux send-keys ... -l -- Orchestra refused your phase request: phase request refused: task 06G46P6KE25Y04VVF7VRZMHZ78 may only move to "research", not "implement"
...: no server running on /tmp/tmux-1000/orchestra: exit status 1

F25 (the boundary ran on claude at all), F21's reading half, F26's premise (the coordinator naming the one valid target), F27 (409 classified, routed to sendPrompt). The request file survived the failed send, the correct branch.

The receipt is still unproven, because F28 had already destroyed the pane.

Ledger

F6   closed
F14  closed
F15  closed by detection, transport fix pending live proof
F16  closed, both branches live-proven (trigger was self-inflicted, see F28)
F17  fixed, isolated live proof
F18  open observability, non-blocking
F19  fixed
F20  fixed, live receipt still pending
F21  reading and artifact halves live-proven; accepted transition pending
F22  fixed, live proof pending
F23  closed as already implemented; was undeliverable on claude until F25
F24  open correctness, dormant on current topology
F25  fixed, live-proven
F26  fixed, live-proven
F27  fixed, live-proven
F28  fix written, NOT INSTALLED. Blocks every clean run.

Next session, in order

  1. Install the F28 units. Nothing else is worth doing first.
  2. Run the F28 proof, both directions:
start task, obtain live pane P, record identity
systemctl restart orchestra-worker
→ pane P still exists with identical agent identity
→ worker reconciles lease + pane P
→ no TaskReleased or new lease caused merely by the restart
→ agent continues under the same lease epoch

then, separately:
restart orchestra-tmux
→ pane disappears
→ F16 refuses renewal
→ lease expires and requeues

The second half restores the meaning of F16's missing-pane branch: real execution-runtime loss rather than deployment killing its own child.

  1. File a fresh issue for run 5, the next clean conformance run. Run 4 is closed; do not shepherd it. The run 5 chain to prove:
frame → phase-request.json → WorkPhaseChanged → phase_changed rotation
research → research.json sealed → request → rotate
plan → accepted research visible
  → post ONE trusted issue comment while plan is live, before it asks to leave
  → HumanDecisionRecorded → RemoteTurn → DecisionNotice → confirmer receipt, once
plan.json sealed → request → rotate
implement launch carries: decision above accepted plan above accepted research

Withhold the correction comment until plan is actually live. Posting during frame or research still tests F20 and F23, but it stops proving that a mid-plan correction outranks accepted research and the current trajectory.

Suggested comment, from the operator:

for --json, use snake_case keys. keep the default text output byte-for-byte unchanged.
  1. Classify separately, as the operator asked: a WorkPhaseChanged with no request artifact, and a request artifact with no transition, are different defects from a failed rotation.

Things that will bite

  • Another session still owns 16 uncommitted paths, including AUDIT.md, deploy/build.sh and the web/ frontend. Commit by path. Never git add -A.
  • Run 4's worktree still holds {"from":"frame","to":"implement"} at /tmp/test-e2e-worktrees/06G46P6KE25Y04VVF7VRZMHZ78/.orchestra/. Harmless, and useful if that task is ever retried diagnostically.
  • w.recordError keeps one slot (cmd/orchestra-worker/main.go:70). It overwrites, so a causal sequence cannot be read from worker health. That is F18, still open, and it cost this session real diagnosis time twice.
  • Gitea writes without reading the token: run curl on homesrv and expand the value there. ssh kami@192.168.1.104 'TOKEN=$(docker exec orchestra-api printenv ORCHESTRA_GITEA_TOKEN); curl -H "Authorization: token $TOKEN" ...' Base https://gitea.kvmx.ru, owner kami, repo test-e2e.
  • Do not use docker compose build. Build from a detached worktree, as deploy/build.sh does and as HANDOFF-2026-08-26-burnin.md shows.
  • Installing the worker and the units needs root, so it is always the operator's step. Hand them the command with the checksum.
  • This machine is workpc, hostname is bugmachine. homesrv is kami@192.168.1.104 with /usr/bin/ssh, not the ssh on PATH.
  • /tmp is tmpfs with a 10-day sweep. Check /tmp/test-e2e and /tmp/test-e2e-worktrees exist before every run.
  • Worker capacity is 1 per harness. A leased task on workpc-claude blocks the next one, and a Gitea task carries no capability, so the router may hand it to workpc-opencode instead. Wait for a free claude slot before filing a run's issue, or the conformance evidence lands on the wrong harness.
  • Three queued correx tasks will never lease. No worker declares that project.