Hand off with F28 uninstalled and run 5 unstarted

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011xsXyr5J1RACo71YeKG3Pu
This commit is contained in:
2026-08-27 17:53:42 +04:00
parent daa5d20d9b
commit fc63dda3b9
+184
View File
@@ -0,0 +1,184 @@
# Handoff — run 4 failed, four fixes landed, one deployment step outstanding
Written 2026-08-27, 14:00 UTC. Read with `BURNIN.md` (the ledger, current
through run 4 and F28), `HANDOFF-2026-08-27-burnin-2.md` (the session before
this one), `AUDIT.md` and `CLAUDE.md`.
Everything below was observed live unless it says otherwise.
## Evidence standard, unchanged
```
event log → lifecycle truth
worker journal → worker-local observations and confirmer receipts
pane capture → harness evidence only
code-path argument → supporting evidence, not live proof
```
This session broke that rule twice and both are corrected in `BURNIN.md`. It
reported F23 as broken after grepping only `internal/operations`, and it
proposed `KillMode=mixed` for F28 without checking what systemd actually does
after the main process exits. Check the whole call graph, and check the manual.
## Deployed state
| Half | Revision | Evidence |
|---|---|---|
| Coordinator, homesrv container | `fbaaf79` | `docker logs orchestra-api`, 13:30:33 UTC, `dirty false`, `/readyz` 200 |
| Worker, workpc systemd | `fbaaf79` | worker registrations, both online |
Branch `webui-and-audit-reconciliation`, HEAD `daa5d20`. Worker sha256
`28bfe635e5311d631ad12cfdb826d80fbfc5e5c78eb12cec2f5a369670ef0964`, staged at
`~/orchestra-deploy/orchestra-worker.fbaaf79`.
`daa5d20` is systemd unit files and ledger only. No rebuild is needed for it.
## The one thing that is not done
**F28's unit split is written but not installed.** Until it is, every worker
restart destroys every agent pane, and no clean conformance run is possible.
```bash
sudo install -m 0644 \
/home/kami/orchestra-deploy/orchestra-tmux.service \
/etc/systemd/system/orchestra-tmux.service
sudo systemctl daemon-reload
sudo systemctl enable --now orchestra-tmux
```
The installed `orchestra-worker.service` also needs the two ordering lines from
`deploy/orchestra-worker.service`, then `daemon-reload`.
The running worker still owns the current tmux server. The split takes hold
only once that server is gone and `orchestra-tmux` owns the next one.
`User` must match between the units, because the socket is `/tmp/tmux-$UID`.
Installed worker runs as `kami`; `deploy/orchestra-worker.service` still says
`orchestra`, which is pre-existing drift. The staged tmux unit says `kami`.
## What run 4 established
Task `06G46P6KE25Y04VVF7VRZMHZ78`, issue `kami/test-e2e#4`. Failed conformance.
Cause: F25. Its front half is the cleanest baseline so far.
| seq | at (UTC) | event |
|---|---|---|
| 370 | 13:11:29 | `TaskCreated` v1, 7 acceptance criteria |
| 371 | 13:11:30 | `TaskLeased` v2 |
| 372 | 13:11:33 | `TaskLaunchAcknowledged` v3 |
Issue filed to confirmed launch: 4 seconds, no operator lifecycle intervention.
One worker-journal line at 13:32:16, after the `fbaaf79` restart, live-proved
four things at once:
```
deliver phase refusal 06G46P6KE25Y04VVF7VRZMHZ78: tmux send-keys ... -l -- Orchestra refused your phase request: phase request refused: task 06G46P6KE25Y04VVF7VRZMHZ78 may only move to "research", not "implement"
...: no server running on /tmp/tmux-1000/orchestra: exit status 1
```
F25 (the boundary ran on claude at all), F21's reading half, F26's premise (the
coordinator naming the one valid target), F27 (409 classified, routed to
`sendPrompt`). The request file survived the failed send, the correct branch.
The receipt is still unproven, because F28 had already destroyed the pane.
## Ledger
```
F6 closed
F14 closed
F15 closed by detection, transport fix pending live proof
F16 closed, both branches live-proven (trigger was self-inflicted, see F28)
F17 fixed, isolated live proof
F18 open observability, non-blocking
F19 fixed
F20 fixed, live receipt still pending
F21 reading and artifact halves live-proven; accepted transition pending
F22 fixed, live proof pending
F23 closed as already implemented; was undeliverable on claude until F25
F24 open correctness, dormant on current topology
F25 fixed, live-proven
F26 fixed, live-proven
F27 fixed, live-proven
F28 fix written, NOT INSTALLED. Blocks every clean run.
```
## Next session, in order
1. Install the F28 units. Nothing else is worth doing first.
2. Run the F28 proof, both directions:
```text
start task, obtain live pane P, record identity
systemctl restart orchestra-worker
→ pane P still exists with identical agent identity
→ worker reconciles lease + pane P
→ no TaskReleased or new lease caused merely by the restart
→ agent continues under the same lease epoch
then, separately:
restart orchestra-tmux
→ pane disappears
→ F16 refuses renewal
→ lease expires and requeues
```
The second half restores the meaning of F16's missing-pane branch: real
execution-runtime loss rather than deployment killing its own child.
3. File a fresh issue for **run 5**, the next clean conformance run. Run 4 is
closed; do not shepherd it. The run 5 chain to prove:
```text
frame → phase-request.json → WorkPhaseChanged → phase_changed rotation
research → research.json sealed → request → rotate
plan → accepted research visible
→ post ONE trusted issue comment while plan is live, before it asks to leave
→ HumanDecisionRecorded → RemoteTurn → DecisionNotice → confirmer receipt, once
plan.json sealed → request → rotate
implement launch carries: decision above accepted plan above accepted research
```
Withhold the correction comment until plan is actually live. Posting during
frame or research still tests F20 and F23, but it stops proving that a mid-plan
correction outranks accepted research and the current trajectory.
Suggested comment, from the operator:
```text
for --json, use snake_case keys. keep the default text output byte-for-byte unchanged.
```
4. Classify separately, as the operator asked: a `WorkPhaseChanged` with no
request artifact, and a request artifact with no transition, are different
defects from a failed rotation.
## Things that will bite
- **Another session still owns 16 uncommitted paths**, including `AUDIT.md`,
`deploy/build.sh` and the `web/` frontend. Commit by path. Never `git add -A`.
- **Run 4's worktree still holds `{"from":"frame","to":"implement"}`** at
`/tmp/test-e2e-worktrees/06G46P6KE25Y04VVF7VRZMHZ78/.orchestra/`. Harmless,
and useful if that task is ever retried diagnostically.
- **`w.recordError` keeps one slot** (`cmd/orchestra-worker/main.go:70`). It
overwrites, so a causal sequence cannot be read from worker health. That is
F18, still open, and it cost this session real diagnosis time twice.
- **Gitea writes without reading the token**: run curl on homesrv and expand the
value there.
`ssh kami@192.168.1.104 'TOKEN=$(docker exec orchestra-api printenv ORCHESTRA_GITEA_TOKEN); curl -H "Authorization: token $TOKEN" ...'`
Base `https://gitea.kvmx.ru`, owner `kami`, repo `test-e2e`.
- **Do not use `docker compose build`.** Build from a detached worktree, as
`deploy/build.sh` does and as `HANDOFF-2026-08-26-burnin.md` shows.
- **Installing the worker and the units needs root**, so it is always the
operator's step. Hand them the command with the checksum.
- **This machine is workpc**, `hostname` is `bugmachine`. homesrv is
`kami@192.168.1.104` with `/usr/bin/ssh`, not the `ssh` on PATH.
- **`/tmp` is tmpfs with a 10-day sweep.** Check `/tmp/test-e2e` and
`/tmp/test-e2e-worktrees` exist before every run.
- **Worker capacity is 1 per harness.** A leased task on `workpc-claude` blocks
the next one, and a Gitea task carries no capability, so the router may hand
it to `workpc-opencode` instead. Wait for a free claude slot before filing a
run's issue, or the conformance evidence lands on the wrong harness.
- **Three queued `correx` tasks will never lease.** No worker declares that
project.