Files
orchestra/BURNIN.md
T
kami daa5d20d9b Own the tmux execution runtime as its own service
F28. The worker spawns the tmux server on its first command, so the server and
every agent pane sit in the worker unit's cgroup. Restarting the worker
destroyed the sessions it was restarting to manage, and F16's missing-pane
branch has been firing on deployment rather than on real execution loss.

KillMode is not the fix. Under mixed systemd still SIGKILLs the cgroup
remainder once the main process exits, and process only encodes accidental
orphaning. The runtime becomes its own service instead.

The worker gains After= and Wants= on it, ordering only: a worker that finds
the runtime missing must report that rather than be stopped by it. The unit
holds an idle session so the server outlives its last agent pane.

User must match between the units, since the socket lives under /tmp/tmux-$UID.
The installed worker on workpc runs as kami while this file still says
orchestra; the staged copy is set to kami to match.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011xsXyr5J1RACo71YeKG3Pu
2026-08-27 17:48:28 +04:00

1189 lines
50 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Burn-in: live conformance, not more architecture
Written 2026-08-26. The workflow is feature-complete enough to exercise. What
remains is empirical: run the same real task shape through each harness and
classify what breaks. Isolated tests are insufficient by design here, and only
the live owner path establishes conformance.
Do not add workflow features while this is running. Findings decide the next
implementation work.
## Burn-in build identity
`77a2b323fabcf080d7542061ae2d7b3c34eef5b7`
Both halves must report exactly this revision before a task is created. Neither
needs a credential now: the coordinator prints it in `docker logs orchestra-api`
and the worker in `journalctl -u orchestra-worker`. The same object is at
`GET /v1/admin/diagnostics` and `GET /v1/federation/workers` behind the operator
login.
Build both with `deploy/build.sh <outdir>`, which refuses a dirty tree. Later
documentation-only commits do not change this identity, so a rebuild either
passes this revision explicitly or accepts the new one and redeploys both
halves. Never one half.
## Deployment state, 2026-08-26 18:35
| Half | State |
|---|---|
| Coordinator (homesrv) | **Deployed at 6f9300b.** Rebuilt with `--build-arg BUILD_REVISION`, recreated, `/readyz` ready, self-reporting the revision in its log |
| Worker (workpc) | **Staged, not installed.** `~/orchestra-deploy/orchestra-worker` sha256 `2b1c43071eaf6b5c15b110de39f204038a9629897e0ce3dacf45be96a6e6529e`. Installed binary is still the 2026-07-30 build |
Two operator steps remain, both needing root:
```sh
sudo install -m 0755 /home/kami/orchestra-deploy/orchestra-worker /usr/local/bin/orchestra-worker
sudo systemctl restart orchestra-worker
journalctl -u orchestra-worker -n 5 --no-pager # must print revision 6f9300b...
```
A third blocker, found while preparing the pane check:
**herdr is not running on workpc.** `herdr status server` reports `not running`,
the socket at `/home/kami/.config/herdr/herdr.sock` refuses connections, and its
log stops at 2026-07-30. `workpc-opencode` therefore cannot start a pane, even
with the worker installed. The worker logs `serving harness workpc-opencode
(opencode) on herdr backend` at startup without touching the socket, so that
line is not evidence the backend is reachable. herdr is an interactive terminal
workspace manager: running `herdr` in a terminal launches or attaches to the
persistent session and starts the server. Confirm with `herdr status server`
before creating a task.
Then scrub the pane environment before letting an agent run.
### Where the pane environment actually comes from
The two workpc harnesses inherit different environments, so one scrub does not
cover both.
- **`workpc-opencode` (herdr backend).** The pane is created by the herdr
daemon, which is a separate long-running process. It inherits herdr's
environment, not the worker's. Scrubbing `worker.env` does nothing here.
Whatever environment herdr is started with is what every opencode pane gets.
- **`workpc-claude` (tmux backend).** `TmuxBackend.StartAgent` runs `tmux
new-session` through `exec.CommandContext` with no `Env` set
(`internal/herdr/tmux.go`). If the tmux server is not already up, the worker
starts it, and that server inherits the worker's full environment. Every
claude pane then inherits it too.
This exposes a conflict the scrub alone cannot resolve. The worker legitimately
needs `ORCHESTRA_WORKER_TOKEN` (or `ORCHESTRA_WORKER_TOKEN_<ID>`) and
`ORCHESTRA_FEDERATION_ADMIT_TOKEN`, and deleting them breaks the worker. Keeping
them means the tmux path hands them to the agent. Closing it needs either a
filtered `cmd.Env` when the backend spawns a pane, or a tmux server started
separately with a clean environment. Not on flow 1's path, because flow 1 is
opencode only. Close it before step 6.
## Baseline: one harness, one path
`workpc-opencode` only. Codex is a coverage gap, not a prerequisite, and no
worker runs on homesrv. Prove one real path first:
```
real vikunja task
→ workpc-opencode
→ frame/research/plan
→ trajectory gate if configured
→ implement
→ independent review
→ task pr
→ human merge
→ completed
```
Then raise difficulty in this order:
```
1. opencode boring success
2. opencode mid-session human correction
3. opencode forced rotation
4. opencode pr rejection -> fix -> merge
5. opencode reconcile outage/recovery
6. representative cases on claude
7. add the codex adapter/config
8. cross-harness rotation: opencode -> claude/codex
```
## Herdr transports, so nobody misdiagnoses again
No entry in the live `config.jsonc` sets `backend` or `address`, so all six
resolve to `<machine>:9245` (`registry.defaultHerdrPort`). Both ports are
closed. That says nothing about the workpc harnesses: they are worker-owned,
and the coordinator logs `worker-owned on workpc; coordinator probe skipped`
for each. The worker reaches `workpc-claude` over a tmux socket and
`workpc-opencode` over `/home/kami/.config/herdr/herdr.sock`. Do not diagnose
harness availability from a TCP probe of 9245.
## Evidence to record per run
One row per run. The launch instruction is now dumped to
`<worktree>/.orchestra/launch.md` (`herdr.LaunchContextFile`) at every launch,
local and federated, so the context is auditable without reading pane
scrollback.
```
harness (claude | codex | opencode)
task id
lease epoch (every epoch, if the task rotates)
session ids (pane ids per epoch)
human decision ids
launch context (sha256 of .orchestra/launch.md, per launch)
handoff ref
git sha (at every boundary: launch, each handoff, review, submit, merge)
review ref
submission ref
pr id
completion receipt
```
## The five flows
1. **Boring success.** `task → research → plan → implement → review → pr → merge`.
2. **Mid-session correction.** Agent is doing A, the human says B, the next
verified boundary delivers B, and A becomes history rather than a competing
instruction.
3. **Rotation.** Agent A hands off, agent B resumes. Switch harnesses across the
boundary if both are up.
4. **Human rejection.** Submit sha A, comment on the pull request, task reopens,
fix at sha B, a fresh review, the same pull request, merge.
5. **Failure path.** Human source goes down, the reconcile streak escalates, the
session hands off, the successor lease is refused, the source recovers, and a
corrected successor starts.
## What to inspect by hand
For every launch, read `.orchestra/launch.md` and ask only:
```
does this agent know
- what the task actually wants?
- what was most recently decided?
- what phase it is in?
- what earlier material is merely historical?
- what to do next?
```
Then compare that against what the model did.
## Classify before fixing
Every failure gets exactly one label before any code is written:
```
authority bug the wrong thing outranked the right thing
context-selection bug the renderer showed or hid the wrong material
lifecycle bug state, lease, phase or event handling is wrong
adapter/harness bug the pane, occupancy, boundary or prompt path is wrong
model-following failure the context was right and the model ignored it
operator-policy gap credentials, deployment or configuration
```
This matters because the cheap response to every failure is another prompt
rule, and prompt rules accumulated to compensate for lifecycle or adapter bugs
are how the enforcement boundary rots. A model-following failure is the only
class a prompt change should ever answer.
## Pane credential cleanup (operator, in parallel)
The agent pane inherits the worker's environment. On workpc the worker is
`orchestra-worker.service`, `User=kami`, `EnvironmentFile=/etc/orchestra/worker.env`
(mode 0600, root-owned, deliberately unreadable here). Remove from that file, and
from anything else the pane inherits:
```
gitea mutation token
vikunja mutation credentials
orchestra operator / system token
```
Retain only `ORCHESTRA_AGENT_TOKEN` and credentials a specific task genuinely
needs. If an agent needs a forge operation, it asks Orchestra to perform it.
Then verify from inside a real pane:
```bash
git push # must fail unless Orchestra supplied the auth
curl .../v1/tasks/<id>/submission # must be 403 for the agent surface
tea pr create # must have no usable mutation credential
```
Note what is *not* enforced by code: nothing in this repo scrubs the pane
environment, and a fake `git` earlier in `$PATH` is not enforcement because
`/usr/bin/git` bypasses it. Credential isolation is the enforcement.
Confinement of network, filesystem and destructive commands belongs to whatever
launches the process (herdr, a container, systemd, bubblewrap), not to
Orchestra, which supplies identity and policy inputs.
## Exit criterion
Roughly 10 to 20 real tasks, with the five flows covered on each harness that
is actually up. Then read the classified failures and decide the next
implementation unit from them.
## Live findings, 2026-08-26 21:35
Both halves are paired at `6f9300b`, and herdr is up (0.7.5, protocol 17). What
remains before flow 1, and what the in-pane credential checks will really show.
### Flow 1 has no repository to run in
`workpc-opencode` declares exactly one project, `test-e2e`, pointing at
`/tmp/test-e2e` and `/tmp/test-e2e-worktrees`. Neither path exists on workpc
after today's reboot, and `/var/lib/orchestra/repos` does not exist there at
all. Nothing can be leased to this harness until it has a project backed by a
durable local repository.
The only configured task source is Gitea `kami/correx` at
`https://gitea.kvmx.ru` (`ORCHESTRA_GITEA_URL/OWNER/REPO`), so a real task means
a `correx` issue. `correx` already has `machine_affinity: ["homesrv","workpc"]`
in the coordinator's `config.jsonc`. What is missing is a workpc-local entry:
```jsonc
// /etc/orchestra/worker-projects.json, on workpc
{
"correx": {
"repo": "/home/kami/orchestra/repos/correx.git",
"worktree_root": "/home/kami/orchestra/worktrees/correx",
"remote": "origin"
}
}
```
Not `/tmp`. The burn-in outlives a reboot.
### There is no vikunja ingest
This repo has no Vikunja provider. `/readyz` reports `gitea` and `jsonl`. A
"real vikunja task" cannot enter Orchestra today: it arrives as a Gitea issue or
through the JSONL watcher.
### `git push` from a pane will succeed, and no Orchestra setting stops it
Corrected 21:45. The mechanism is SSH, not HTTPS. HTTPS has no stored
credential on workpc, where `git push` over `https://` fails with `could not
read Username`. Pushing works over `ssh://git@gitea.kvmx.ru:2222` using kami's
default RSA identity, which Gitea lists as the key named `workpc`. It carries no
passphrase, so no agent is needed.
That identity is what lets the *worker* push, and panes run as the same user
with the same home directory, so an agent inherits it. `worker.env` is
irrelevant to this. Real isolation needs panes under a different unix user.
Expect `git push --dry-run` to succeed from a pane, and record it as an
operator-policy gap rather than a code defect.
The worker inherits ambient Git credentials by design: `git()` in
`cmd/orchestra-worker/main.go` runs `exec.CommandContext` with no environment of
its own. There is no separate credential path for the worker to hold something
the pane does not.
### The agent surface is currently unauthenticated
`ORCHESTRA_AGENT_TOKEN` is unset on the coordinator, and an unset surface token
means the middleware performs no check for that surface. Probed live:
```
POST /v1/tasks/<id>/decision-request -H 'X-Orchestra-Surface: agent' -> 404 (reached the handler)
POST /v1/tasks/<id>/phase -H 'X-Orchestra-Surface: agent' -> 403 (refused)
```
The capability boundary holds. Authentication does not. Set
`ORCHESTRA_AGENT_TOKEN` in the coordinator `.env` before the burn-in, and note
the same is true of `ORCHESTRA_MCP_TOKEN` and `ORCHESTRA_MAVEN_TOKEN`, both
unset.
### Two more unset settings that change burn-in behaviour
- `ORCHESTRA_FEDERATION_ADMIT_TOKEN` is unset, so any caller may register a new
worker identity. Existing identities stay protected, because registering an
existing id with a different token is refused.
- `ORCHESTRA_REVIEW_ACTORS` is unset, which `human.Trust` reads as "anyone not
explicitly ignored". Flow 4 will accept a task-moving comment from any Gitea
actor. Set it to `kami`.
## Flow 1 setup, 2026-08-26 21:45
Target: `kami/test-e2e` on Gitea, a throwaway repo, rather than `correx`.
Done:
- **Ingest switched.** `ORCHESTRA_GITEA_REPO=test-e2e` in the coordinator `.env`
(previous file kept as `.env.pre-burnin-20260826`), container recreated, still
reporting `6f9300b`. The project id must equal the repo name, because
`main.go` sets `Project: os.Getenv("ORCHESTRA_GITEA_REPO")`. `test-e2e`
already exists in `config.jsonc` with `machine_affinity: ["workpc"]`.
- **Durable repo on workpc.** Bare clone at
`/home/kami/orchestra/repos/test-e2e.git`, worktree root
`/home/kami/orchestra/worktrees/test-e2e`. Origin points at
`ssh://git@gitea.kvmx.ru:2222/kami/test-e2e.git`, and `git push --dry-run`
reports `Everything up-to-date`. The old `/tmp` paths did not survive the
reboot and must not come back.
Remaining, needs root:
```jsonc
// /etc/orchestra/worker-projects.json
{
"test-e2e": {
"repo": "/home/kami/orchestra/repos/test-e2e.git",
"worktree_root": "/home/kami/orchestra/worktrees/test-e2e",
"remote": "origin"
}
}
```
Then `sudo systemctl restart orchestra-worker`.
Store state, checked before the first run: 30 tasks, all `test-e2e`, with 23
blocked, 6 completed and 1 failed. All are July leftovers and none holds a
lease. The coordinator has run `ResumeAnsweredBlockers` every second for hours
without resuming any of them, so they are inert rather than merely quiet. The
July stuck task `06FT6CKD9Y98AZRX6X8K3QXFZG` is now `failed`.
A stale branch `orchestra/scratch/oc-06ftgkjadcd2hwjn2zwjen90q4` exists on the
remote from an earlier run. Harmless, but it is not from this burn-in.
## Run 1: task 06G3YR34117MAYT6KEAC9RJHD0, 2026-08-26
Evidence ledger:
```
harness workpc-opencode (herdr backend)
task 06G3YR34117MAYT6KEAC9RJHD0
source gitea:test-e2e/1 "Add a --version flag to the healthcheck script"
attempts 3 leases: two nacked, third launched
pane wN:p1, agent oc-06g3yr34117mayt6keac9rjhd0
agent session ses_fc096de31ffelZeMngQEyBTxfh
worktree /tmp/test-e2e-worktrees/06G3YR34117MAYT6KEAC9RJHD0
branch orchestra/06G3YR34117MAYT6KEAC9RJHD0
head at launch 5e6c4783d4213a934dc486160b888b6014bc2d03
launch context .orchestra/launch.md, 2138 bytes
phase frame
decisions none
```
### What worked
- **The launch context is right.** Authority order, the ambiguity ladder, the
phase brief with `frame: ... Do not change code`, `Current human decisions:
None recorded`, and a verified git state with worktree, branch and head.
- **The agent respected the phase.** `agent_status: done`, no code changes, only
Orchestra's own scratch directory in the tree.
- **A bad launch nacked cleanly.** The first attempt failed on the base checkout
and the task went back to `queued` with `lifecycle_phase: launch_nacked` and
`attempt: 1`. No orphan pane, no stuck lease.
### Fixed during the run
- **F1, authority bug, `2753a8d`.** Every federated launch died with `effective
intent: federation: 401 Unauthorized: unauthorized surface`.
`GET /v1/tasks/<id>/intent` exists for workers, and the authz worker-path
exemption never included it. Twenty test packages passed throughout.
- **F2, authority bug, `09e572f`.** The rendered goal was the issue title alone
and acceptance read `Not stated.`, because `Gitea.event` parsed the issue body
and dropped it from the `TaskCreated` payload. The store already read
`description`. Every Gitea-sourced task so far ran on its title.
- **F3, self-inflicted, `a0209a2`.** The launch dump left `.orchestra/`
untracked in a worktree whose own context said `uncommitted changes: false`.
It would have polluted the gate, the review diff and the agent's `git status`.
### Open, classified
- **F4, adapter/harness.** `rotation ... activity degraded: activity unknown:
adapter: resolve session file: adapter: harness "opencode" has no
session-file resolver`. Thrash and activity triggers are permanently degraded
on opencode. Observable in the worker's `last_error`, non-fatal.
- **F5, lifecycle.** The router never leased this task. Every gate checks out by
hand: worker `online: true`, `herdr_status: reachable`, fresh `checked_at`,
`supported_projects: ["test-e2e"]`, quota empty, no active leases, capability
empty. A direct `POST /v1/tasks/<id>/lease` succeeded instantly. The refusal
is upstream of `Store.Lease` and silent by construction
(`internal/router/router.go:197`). All three runs were leased by hand.
- **F6, observability.** Worker health reports `active_task: null` while the
store shows the task leased and herdr shows a live opencode agent in the
task's worktree.
- **F7, operator policy.** `ORCHESTRA_TUI_TOKEN` is unset, and an unset surface
token means no authentication for that surface. That is how the manual leases
above were issued, unauthenticated, from another machine. The TUI surface is
`FullControl`, so this is the whole control plane, not just the two agent
request endpoints. Set the token.
- **F8, correctness.** `human.Reconciler.Reconcile` iterates every configured
source for every task, so the three queued `correx` tasks have their external
ids looked up in `kami/test-e2e`. A source should only reconcile the tasks
that came from it. Likely why those three never lease.
## Fix pass before run 2, 2026-08-26 23:50
Burn-in identity is now `77a2b32`. The coordinator is deployed at it. The worker
is staged at it, sha256 `2a850f2d1102d390...`, still needing root to install.
### Closed
- **F7, security.** `authz.RequireCredentials` refuses startup when a
full-control surface has no token, rather than logging it.
`ORCHESTRA_TUI_TOKEN` is set in the coordinator `.env`. Verified live: an
unauthenticated `POST` on the TUI surface now returns 401. Web is exempt
because `Sessions` makes its login mandatory. `ORCHESTRA_MCP_TOKEN`,
`ORCHESTRA_MAVEN_TOKEN` and `ORCHESTRA_AGENT_TOKEN` remain unset, so those
surfaces are still unauthenticated for reads and for their three request
endpoints. Bounded by capability, worth closing, not startup-fatal.
- **F5, lifecycle.** Every eligibility gate now records a `router.Rejection`,
exposed at `GET /v1/router/health` and reset per pass. No gate was weakened.
The live output immediately explained the three stuck `correx` tasks:
`worker has not declared project correx`. A queued task in retry backoff was
skipped before the candidate loop and recorded nothing at all, which is the
shape that hid the original case; it now reports
`retry backoff until <time>`.
- **F8, correctness.** Reconciliation is bound to `task.Source`, the
`provider:project` identity the ingest stamped. A source that cannot prove it
owns the task is skipped, and a task with no matching source reconciles to
nothing and still launches. The integration fixture had encoded the bug: it
ingested from `jsonl` and reconciled from `gitea`.
- **F11, outage.** The API entered a restart loop exiting with `invalid event:
until_ns required`. `ValidateEvent` compared `until_ns` against `time.Now()`
for `TaskLeased` and `TaskLeaseRenewed`, so a lease event that was valid when
written failed validation once it expired. `store.Open` replays the tail after
the snapshot and `log.Fatal`s on the first invalid event, so the coordinator
refused its own history. Validation of a durable event is now
time-independent. Latent since the field existed: it needed a renewal in the
post-snapshot tail plus a restart after that renewal expired.
### Also found
- **F9, lifecycle.** An operator cannot release a leased task through
`POST /v1/tasks/<id>/release` without knowing its `harness_id` and
`lease_epoch`, because `Store.Append` fences lifecycle events on a leased
task and the endpoint passes the request body through unchanged. Correct that
only the owner releases, but there is no operator escape hatch.
- **F10, hygiene.** A coordinator-side release under a live worker leaves the
worker renewing a lease it no longer holds. Combined with F11 that produced a
per-second invalid-event log line.
### Shared checkout
Another session is working on the auth and frontend layers in this same
checkout. Two consequences, both handled: `deploy/build.sh` and the container
image now build in a detached worktree of the revision they stamp, so no
uncommitted work is compiled into a stamped binary and no in-progress edit
blocks a deploy. Commits from this session name their paths rather than using
`git add -A`. The first commit, `7f12c7f`, predates that discipline and swept in
whatever was uncommitted at the time, including `web/src` and `internal/authn`.
### Run 1 is closed as a diagnostic
Task `06G3YR34117MAYT6KEAC9RJHD0` was released with reason `abandoned diagnostic
run`, supplying the lease fence by hand per F9. It sits at `attempt: 3` and the
router will fail it on its next pass. It is not a conformance run: it needed
three manual leases and its authority was built before the issue-body fix.
### Run 2 preconditions
```
worker installed at 77a2b32 and restarted
coordinator at 77a2b32 done
F7 closed done
F5 exposed done
F8 fixed done
fresh issue, no manual leasing
```
## Runs 2 and 3, 2026-08-27
Written from live evidence. Neither run was a conformance pass, and both were
worth more than one: run 2 found three blocking bugs behind each other, and
run 3 found two more plus the reason six isolated probes disagreed with
production.
`AUDIT.md` is uncommitted and owned by another session, so the burn-in ledger
lives here.
### Fixed, with the revision each landed in
- **F12**, lifecycle, `0d67af9`. `Store.QuotaSince` reported an empty window as
unknown, `QuotaAvailability` fails closed on unknown, and every herdr in
`config.jsonc` declares a quota limit. The only producer of a receipt is a
completed lease, so nothing could ever be leased. An empty window is now
observable zero. A receipt that declares its own consumption unknown still
fails closed. Verified live: run 2 leased three seconds after the fixed
coordinator started.
- **F13**, observability, `0d67af9`. `federatedAvailability` restated a quota
refusal as worker health, so router health said `stale heartbeat` against a
heartbeat one second old. Gates name themselves through
`router.ReasonedAvailability`. Verified live.
- **F14**, authority, `a5d361b`. Gitea ingest never set `acceptance`. One
recognized heading, bullets and checkboxes until the next heading, order
preserved, section removed from the description. Verified live on run 3: five
ordered items, prose above the heading kept as the description.
- **F15**, adapter, `1fd82f8`. Two bugs. `TmuxBackend.Prompt` wrote the whole
instruction with `send-keys -l`, Claude Code coalesced it into a paste, and
the Enter was absorbed. And `TaskLaunchAcknowledged` meant "Prompt returned
nil", not "the harness accepted it". Launch transport became a backend
property, and `ConfirmLaunch` began polling for proof. The detection works.
The fix did not: see F17.
- **F16**, lifecycle, `1888d42`. `renewLeases` renewed whenever a session
existed and `PaneCapture` succeeded, so a pane that opened and never started
renewed forever. This is why the July task stayed orphaned: F15 explains why
nothing started, F16 why the lease never let go. Renewal now needs the agent
busy, or the pane capture to differ from the hash recorded at the previous
renewal. `lease.ProgressSHA` carries that hash.
- **F17**, adapter, `f54fb00`. `pendingInput` scanned every line beginning with
the prompt marker, but queued and already-accepted input renders with the
same prefix. Only the editor owning the pane cursor is unsubmitted, so
`inputState` reads `#{cursor_y}` and captures screen rows without `-J`, which
would invalidate the row index. On that footing `ConfirmLaunch` became an
active submit protocol: resend Enter while the live editor still holds
exactly what was submitted, at most three times, no closer than two poll
intervals, then observe until the deadline. Queued input confirms rather than
fails. Evidence records `confirmation`, `submit_attempts` and both timestamps.
- **F19**, lifecycle, `edff021`. A task that reached the router's
`MaxAttempts` was permanently terminal. `TaskReleased` only increments
`Attempt`, `TaskCorrected` could not touch it, and no HTTP route emitted a
correction. `POST /v1/tasks/{id}/retry` requires the task to be failed,
unleased, and failed with `reason: retry_limit`, then appends one
`TaskCorrected` naming that failure with `state: queued` and `attempt: 0`.
Identity, goal, acceptance, decisions, phase and artifact refs all survive,
and the original failures stay in the log. `operation_id` is required and
makes it idempotent. Deliberately not a generic correction endpoint.
### The finding that mattered most
Run 3 failed three times with `prompt_not_submitted`, and six isolated probes
of the same code path could not reproduce it: fresh session, untrusted
directory with the trust dialog, `launch.md` present before startup, 100ms
readiness polling, `env -i` with only the worker's systemd variables, and the
real `StartAgent`/`Prompt`/`ConfirmLaunch` against a byte-identical git
worktree. All six submitted on the first Enter.
F17 turned that disagreement into data instead of a theory. The first live
launch after it deployed:
```
launch 06G44JZB80MZBEY97196EZN8EC confirmed: confirmation=editor_cleared
submit_attempts=2 first_submit_at=2026-08-27T09:12:00.801855937Z
confirmed_at=2026-08-27T09:12:01.315587139Z
```
Production needs the second Enter. Isolated probes need one. The submit is not
deterministic, which is exactly what F15's `ConfirmLaunch` comment asserted it
was.
### Still open
- **F18**, observability. A queued launch used to be misread as unsubmitted.
F17 fixes the predicate, but nothing else that parses pane text distinguishes
the active editor from history, so the same class of error can recur wherever
`PaneCapture` output is matched.
- **F9** is unchanged by F19. An operator still cannot release a lease someone
else owns without supplying that owner's `harness_id` and `lease_epoch`,
which means reading them out of the event log first. F19 recovers a terminal
task, which is a different escape hatch.
- **F6**, observability. Worker health reported `active_task: null` while the
store showed the task leased. Run 2 showed this was partly honest, because
the worker really had no working agent. Recheck now that launches confirm.
- **F4**, adapter quality. `harness "opencode" has no session-file resolver`,
so activity and thrash triggers are degraded on opencode.
- **F10**, hygiene. A coordinator-side release under a live worker leaves the
worker renewing a lease it no longer holds.
- Unset gated tokens: `ORCHESTRA_MCP_TOKEN`, `ORCHESTRA_MAVEN_TOKEN`,
`ORCHESTRA_AGENT_TOKEN`.
- `gofmt -l` fails on `internal/provider/provider.go`,
`internal/router/router.go` and `internal/webui/webui.go` at `edff021`. None
were touched by these runs. `internal/router/router.go` picked it up in
`0d67af9`. `go vet` passes, which is why nobody noticed.
### Operator actions taken by hand, and why
- Task `06G3ZCZWJ3QHF992ZMDGSJ0PYG`, run 2, was blocked with
`block_reason: operator_block` rather than released. A released task returns
to `queued`, and the router leases `queued` tasks, so releasing it would have
put it in competition with run 3.
- The first block attempt returned `task version conflict`.
`Store.validateTransition` fences every lifecycle event on a leased task, so
the payload needs `harness_id` and `lease_epoch` even though the HTTP handler
does not ask for them. That is F9 in practice.
- Task `06G44JZB80MZBEY97196EZN8EC`, run 3, was recovered with F19's retry
rather than by filing a fourth issue. That keeps `kami/test-e2e#3` served by
its own task and proves terminal recovery in the same run.
### Deployment boundary
Both halves report `edff021`, built from a detached worktree of that revision.
Worker sha256 `68ac265455acb0e007bfb0e898a91b317cd7289d7f2bf5da6df1224e1d6706a5`
at `/usr/local/bin/orchestra-worker`. Coordinator image built 13:08:30 +0400.
## Run 3: autonomous expiry and relaunch, 2026-08-27 11:02 UTC
Task `06G44JZB80MZBEY97196EZN8EC`, issue `kami/test-e2e#3`. Observed live, no
operator action in the window. The stranded epoch was
`06G44ZBZX4YH6PRN7Y4ZH8GG3W`.
| seq | at (UTC) | event | epoch |
|---|---|---|---|
| 358 | 10:32:09 | `TaskLeaseRenewed` v15 | `06G44ZBZX4YH6PRN7Y4ZH8GG3W` |
| 359 | 11:02:16 | `TaskReleased` v16, `reason: lease_expired` | `06G44ZBZX4YH6PRN7Y4ZH8GG3W` |
| 360 | 11:03:17 | `TaskLeased` v17 | `06G45RVKAW20H1D22NHBRYKMN8` |
| 361 | 11:03:22 | `TaskLaunchAcknowledged` v18, `lifecycle_phase: started` | `06G45RVKAW20H1D22NHBRYKMN8` |
Expiry to acknowledged launch: 66 seconds, autonomous. Every event from seq 353
onward carries `surface: system`. The last `tui` event is the `TaskCorrected` at
09:11:55, before the window.
Launch receipt, worker journal, 15:03:22 +04:
```
launch 06G44JZB80MZBEY97196EZN8EC confirmed: confirmation=editor_cleared submit_attempts=2 first_submit_at=2026-08-27T11:03:22.203859355Z confirmed_at=2026-08-27T11:03:22.717657436Z
```
### F16 is partial, not closed
The hard predicate passed: no renewal carried the stranded epoch after
10:45:26. The run did not exercise the progress comparison. Worker health
records why the renewal was skipped:
```
"last_error": "validate lease 06G44JZB80MZBEY97196EZN8EC: tmux capture-pane -p -J -S -200 -t =orchestra-06g44jzb80mzbey97196ezn8ec-e53607ea:1.0: no server running on /tmp/tmux-1000/orchestra: exit status 1",
"error_at": "2026-08-27T11:02:16.213318366Z"
```
That is the `paneProgress` error path at `cmd/orchestra-worker/main.go:973`, not
the `default` branch at line 992. The tmux server was gone, so no pane text
could be hashed.
- **Proven:** dead or missing pane, no renewal, lease expires, autonomous
re-lease and relaunch.
- **Unproven:** live pane with an idle agent and unchanged output, renewal
refused.
### Epoch reset, stated precisely
The new epoch and its acknowledged launch establish the reset structurally.
`leases` in `/var/lib/orchestra/worker-state/workpc-claude.json` holds only
`{epoch, version, until}`, with `progress_sha` absent under `omitempty`. No
runtime value of `ProgressSHA`, `UsageBaseline` or `PickupAcknowledged` was
observed.
### F6 closed
Worker health reports `active_task_id: 06G44JZB80MZBEY97196EZN8EC` with
`active_pane_id` set, against a leased task.
### F18 sharpened, non-blocking
`recordError` (`cmd/orchestra-worker/main.go:70`) keeps one slot, so the earlier
renewal decision near 10:52 was overwritten and cannot be read. One
`last_error` slot cannot preserve a causal sequence. Later replacement: a
bounded recent-error ring or event-backed observations, roughly the last 8
`{at, class, message}` entries. Do not land this during a live run.
### Ledger after this checkpoint
```
F6 closed
F14 closed
F15 closed by detection, transport fix pending live proof
F16 partial: missing-pane branch proven, static-live-pane branch unproven
F17 fixed, isolated live proof
F18 open observability
F19 fixed
F20 awaiting first post-frame confirmed input
```
### Next decisive checkpoint, F20
```
frame completes
→ Orchestra emits the next input
→ worker journal must contain: input to <pane> confirmed: confirmation=<kind> submit_attempts=<n>
→ only then may the agent continue
```
Receipt appears and the model proceeds: let the run continue through phase
transitions. Input lands with no receipt: F20 fails. No input attempted after
`frame`: a different lifecycle or phase-advance bug.
## Run 3 conformance result, 2026-08-27 12:00 UTC
```text
run 3 conformance result: failed
cause: no autonomous phase-advance path
```
### F16 closed
```text
missing-pane branch: proven
live-pane idle + unchanged output branch: proven
external editor input ignored as progress: proven
```
Second branch evidence, worker health at 11:53:46 with the pane alive and
`herdr_status: reachable`:
```
"last_error": "lease 06G44JZB80MZBEY97196EZN8EC not renewed: agent status idle and pane unchanged since the last renewal",
"error_at": "2026-08-27T11:53:46.195553513Z"
```
That is the `default` branch at `cmd/orchestra-worker/main.go:992`.
### F21, lifecycle: the phase brief promises a protocol that does not exist
`.orchestra/launch.md:62` tells the agent `Orchestra decides when this phase
ends. Ask for a phase change, do not declare one.` The agent complied and
printed `Nothing blocks. Ready for a phase change to implement.`
```text
phase brief: "ask for a phase change"
agent: asks
worker: has no representation of that request
coordinator: AdvanceWorkPhase exists, but nothing invokes it autonomously
```
`operations.AdvanceWorkPhase` is reachable only from the HTTP handler at
`cmd/orchestra/main.go:883` and from `internal/operations/review.go:101`. The
worker's only post-frame send is a decision notice at
`cmd/orchestra-worker/main.go:1639`, gated on `len(answer.Decisions) == 0`
returning early.
The fix is a bounded agent intent, not pane-text matching:
```go
type PhaseAdvanceRequest struct {
From WorkPhase
To WorkPhase
}
```
```text
verified turn boundary
→ obtain bounded phase intent from harness
→ validate requested transition
→ submit to coordinator
→ operations.AdvanceWorkPhase
→ if transition requires sealed artifact:
collect/seal artifact first
→ rotate/start successor if phase policy requires fresh context
```
`frame` needs no artifact. `research` and `plan` take the same path but must
seal their artifact before `AdvanceWorkPhase` accepts them.
### F22, lifecycle: nothing reacts to `WorkPhaseChanged` at runtime
`t.WorkPhase` is read only when a launch context is built, at
`cmd/orchestra-worker/main.go:369` and
`internal/orchestrator/orchestrator.go:1218`. No code path rotates, relaunches
or notifies a live session when the phase changes. A manual advance therefore
takes effect only at the next launch, and leaves the current idle pane idle.
### F23, lifecycle: a decision cannot exist before a submission
`answer.Decisions` is the only input `sendPrompt` ever carries post-frame.
Decisions come from `operations.ReflectSubmission`, and
`human.PullRequestState.FeedbackAfter` returns input only strictly after a
submission. A task still in `frame` has no submission, so no decision can be
recorded for it.
Consequence: **F20 cannot be exercised on this task in its current state.** A
manual phase advance does not produce a post-frame input, because nothing
reacts to the phase change and no decision can exist yet.
### Observation kept separate from F20
` go ahead and implement it` appeared in the editor again on attempt 2, cursor
at `cursor_x=2`, never submitted. The agent transcript at
`~/.claude/projects/-tmp-test-e2e-worktrees-06G44JZB80MZBEY97196EZN8EC/97b30e4d-90a7-49bc-a9b8-7d9e450831e4.jsonl`
holds exactly one user message, the launch prompt at 11:55:52.236Z. No worker
send and no receipt exist for it. Origin unknown, external to Orchestra. F20
does not absorb it.
### Ledger after this checkpoint
```
F6 closed
F14 closed
F15 closed by detection, transport fix pending live proof
F16 closed, both branches live-proven
F17 fixed, isolated live proof
F18 open observability
F19 fixed
F20 blocked, not merely awaiting: see F23
F21 open lifecycle, no autonomous phase-advance path
F22 open lifecycle, no runtime reaction to WorkPhaseChanged
F23 open lifecycle, no decision path before a submission
```
## F21, F22, F23 implemented, 2026-08-27
Run 3 was not shepherded further. Attempt 3 was left unspent: F16 is closed on
both branches, and another idle phase would have proven nothing new.
### F23 was already implemented, and the earlier entry was wrong
`human.Reconciler.Reconcile` imports issue comments from `task.Source` and
records `HumanDecisionRecorded` with no submission involved. It is wired at
two points in `cmd/orchestra/main.go`: `Store.PreLease`, and
`Coordinator.ReconcileHumanInput`, which `RemoteTurn` calls at every verified
turn boundary. `provider.GiteaComments.FetchAfter` reads the task's own issue.
The earlier F23 entry traced `EventHumanDecisionRecorded` through
`internal/operations` only and concluded decisions required a submission. That
was a scoping error in the search, not a gap in the code. `ReflectSubmission`
is the pull-request path and stays narrow; it was never the general decision
source.
Consequence for the record: **F20 had a legitimate post-launch send path
throughout run 3.** A comment on `kami/test-e2e#3` would have produced a
decision, a `DecisionNotice` at the next boundary, and a receipt.
F23 is closed as already-implemented, with tests added for the boundary it
turns on.
### F21, the phase-request protocol
The agent asks with a bounded file. Prose is not a protocol, so nothing
matches on pane text.
```
.orchestra/phase-request.json {"from": "frame", "to": "research"}
.orchestra/research.json sealed before leaving research
.orchestra/plan.json sealed before leaving plan
```
At a verified turn boundary the worker validates what it can see locally: the
phase the agent believes it is in, the legality of the transition, and the
presence and decode of the artifact the phase must seal. It then calls
`POST /v1/federation/phase` with the lease epoch and a derived operation id.
The coordinator calls the existing `operations.AdvanceWorkPhase`.
- `internal/operations/workphase.go`: `RequestWorkPhase`, `ErrPhaseRequest`.
Fences on lease epoch, refuses a stale phase belief, refuses any target but
the project's next phase, idempotent per operation id.
- `internal/federation/client.go`: `Client.AdvancePhase`.
- `cmd/orchestra/main.go`: the `/v1/federation/phase` route.
- `cmd/orchestra-worker/main.go`: `requestPhase`, `phaseArtifact`,
`phaseRequestFile`.
- `internal/agentctx/agentctx.go`: `phaseRequestBrief` renders the protocol
under the sentence that used to promise it.
The operation id is derived, not random: `phase:<task>:<epoch>:<from>:<to>`. A
redelivery after a lost response carries the same id, so the coordinator
returns the first event instead of advancing twice.
### F22, a phase change ends that cognitive session
`herdr.Session` now carries `Phase`, the phase the session was launched to
run. At a turn boundary a session whose task has moved on is rotated with
reason `phase_changed`, which is added to the closed handoff-reason set in
`internal/herdr/adapter.go` and given its own wording in
`RequestHandoffReason`.
One comparison covers both cases: a change this worker requested, and one an
operator made through `POST /v1/tasks/{id}/phase`. Both leave the same
evidence, a session built for a phase that is no longer current.
### F20 hole found and closed while implementing F22
`CLIAdapter.prompt` called `backend.Prompt` and returned. Every rotation and
handoff prompt therefore went out unconfirmed, which is precisely the failure
F20 exists to catch. `RequestHandoffReason` is one of its callers, so the
phase rotation would have inherited it.
Fixed at the shared call site rather than per caller: `prompt` now routes
through `InputConfirmer` and logs the same `input to <pane> confirmed:`
receipt the worker logs.
### Tests
`go build ./...`, `go vet ./...` and `go test ./...` all pass.
- `internal/operations/phase_request_test.go`: full path with each artifact
sealed, skipped phase refused, stale phase belief refused, stale epoch
refused, missing operation id refused, redelivery idempotent, operation id
recorded.
- `cmd/orchestra-worker/phase_test.go`: request accepted then session rotates,
external phase change rotates, unsealed artifact refused locally, malformed
artifact refused locally, sealed artifact travels with the request, refused
request retried, decision notice stays undelivered until confirmed.
- `internal/integration/phase_protocol_test.go`: a pre-submission issue
comment steers a live leased session and is not re-sent once delivered;
the phase-request path seals and fences at the coordinator.
- `internal/human/pullrequest_window_test.go`: pull-request feedback ignores
anything not strictly after the submission, and applies trust.
- `internal/agentctx/agentctx_test.go`: the brief names the request file, the
request shape, and the artifact to seal.
### Ledger
```
F6 closed
F14 closed
F15 closed by detection, transport fix pending live proof
F16 closed, both branches live-proven
F17 fixed, isolated live proof
F18 open observability, non-blocking
F19 fixed
F20 fixed, adapter hole closed; live proof pending run 4
F21 fixed, tests only
F22 fixed, tests only
F23 closed, already implemented; tests added
```
Run 4 starts from here. Both halves must be rebuilt and redeployed before it
begins, and the deployment boundary recorded as usual.
## Deployment boundary for run 4, 2026-08-27
Revision `1f5bf7e`, both halves, built from a detached worktree of that commit.
| Half | Evidence |
|---|---|
| Coordinator, homesrv container | `docker logs orchestra-api`, `orchestra revision 1f5bf7e66e0c6dc8dc1db794ab91193c6c38e118 built 2026-08-27T16:35:37+04:00 dirty false`, `/readyz` 200 |
| Worker, workpc systemd | pending operator install |
Worker sha256 `c188ec78f7b11bd0f3b72da2c3544c5c9477b4505816b30bd0f243f6fa8032da`,
staged at `~/orchestra-deploy/orchestra-worker.1f5bf7e`.
Both worker registrations must report `1f5bf7e` before run 4 starts.
### F24, correctness, dormant on the current topology
`CLIAdapter.LeasePrompt` (`internal/herdr/adapter.go:161`) calls
`backend.Prompt` and returns without confirming. It is the last Orchestra-owned
editor write outside the delivery guarantee.
Dormant, not fixed: the coordinator-local harness path cannot execute here. No
local herdr is eligible, and `Coordinator.adapterFor` refuses a session owned by
a non-local herdr. Run 4 exercises the federated worker path, whose launch
delivery is already confirmed at `cmd/orchestra-worker/main.go:454`.
The eventual fix:
```text
CLIAdapter.LeasePrompt
→ backend.Prompt
→ ConfirmInput
→ only then acknowledge launch
```
The comment in place there argues that waiting would turn a long first turn
into a false lease failure. It is obsolete: `ConfirmInput` proves submission,
not completion of the first turn.
Required before claiming coordinator-local harness conformance. Not required
before this federated burn-in, and changing the revision now would add churn
run 4 cannot observe.
### The y/n approval paths stay separate
`internal/herdr/adapter.go:703` and `cmd/orchestra-worker/main.go:1086` write to
a y/n dialog, not an editor. They need a capture-revision-aware confirmation
protocol of their own. The editor confirmer would report nonsense on them, so
they are deliberately outside F20 and stay that way.
### Run 3
Left running and untouched. Whatever it does from here is additional diagnostic
evidence. Run 4 does not wait on it.
## Run 4: failed conformance, 2026-08-27
Task `06G46P6KE25Y04VVF7VRZMHZ78`, issue `kami/test-e2e#4`, revision `1f5bf7e`.
```text
run 4: failed conformance
cause:
f25, claude rotationTick bypassed federatedTurn
evidence:
phase-request.json existed
no WorkPhaseChanged followed
no boundary error was emitted
additional defects:
f26 phase brief advertised invalid shortcut
f27 phase refusal was not delivered to agent
```
### What run 4 did prove
The front half of the chain is clean, with no operator lifecycle intervention.
| seq | at (UTC) | event |
|---|---|---|
| 370 | 13:11:29 | `TaskCreated` v1, 7 acceptance criteria |
| 371 | 13:11:30 | `TaskLeased` v2, epoch `06G46P6Q4EAPF7DFDBFBTA9RJC` |
| 372 | 13:11:33 | `TaskLaunchAcknowledged` v3 |
Issue to confirmed launch: 4 seconds. F14 passes, against run 1 where ingest
set no acceptance at all. F15 and F17 pass again with
`confirmation=editor_cleared submit_attempts=1`.
The agent wrote its request 25 seconds after launch:
```
/tmp/test-e2e-worktrees/06G46P6KE25Y04VVF7VRZMHZ78/.orchestra/phase-request.json
{"from": "frame", "to": "implement"}
```
The artifact half of F21 works. The reading half never ran.
### F25, lifecycle: the turn boundary was unreachable on claude
`rotationTick` returned early for this harness before reaching the boundary,
and `federatedTurn` has exactly one call site, below that return.
```go
if w.harness == "claude" {
if err := w.advanceClaudeContextReset(ctx, id, s); err != nil { ... }
return
}
```
Introduced in `7f12c7f`. Consequences on the harness both burn-in runs used:
no phase request could ever be read, and **every human decision recorded
against a live claude session went undelivered**.
This retracts an earlier claim in this ledger. F23's decisions do reach
`RemoteTurn`, but on claude the worker never asked, so run 3 had no
post-launch send path either. The statement that a comment on `#3` would have
produced a receipt was wrong.
Fixed: claude runs the common boundary after its context reset. It still skips
the occupancy state machine, because it owns its own rollover. A turn boundary
is not a rotation.
### F26, authority: the brief advertised an invalid shortcut
The deployed brief rendered `Legal values for "to" from here: research,
implement`, listing every domain-legal move. The project's path allows only
`research`. The agent took the shortcut, reasonably.
Fixed: the brief names one target and states that a wrong one comes back with
the right one. The project's path is Orchestra's to know.
### F27, lifecycle: a refusal never reached the agent
A refused request only reached `w.recordError`. The agent would rewrite the
same rejected file at every boundary with nothing telling it why, which is the
silent-loop shape this codebase keeps producing.
Fixed: `federation.StatusError` makes a 409 classifiable. A refusal is
delivered through `sendPrompt`, so it travels under the F20 guarantee, and the
request file is cleared. A transport failure keeps the file and tells the agent
nothing, because it is not an answer.
### Ledger
```
F6 closed
F14 closed
F15 closed by detection, transport fix pending live proof
F16 closed, both branches live-proven
F17 fixed, isolated live proof
F18 open observability, non-blocking
F19 fixed
F20 fixed, live proof still pending
F21 artifact half live-proven, reading half pending
F22 fixed, live proof pending
F23 closed as already implemented, undeliverable on claude until F25
F24 open correctness, dormant on current topology
F25 fixed, tests only
F26 fixed, tests only
F27 fixed, tests only
```
### F28, deployment: the execution runtime shares the worker's cgroup
Worker restart destroys agent sessions because the tmux execution runtime
shares the worker service cgroup. tmux must be independently lifecycle-managed.
The worker spawns the server implicitly on its first tmux command, so the
server and every pane land in `orchestra-worker.service`:
```text
orchestra-worker.service
├── orchestra-worker
└── tmux server
└── claude panes
```
`KillMode` does not fix this. Under `mixed` systemd still sends the final
SIGKILL to whatever remains in the cgroup once the main process exits, and
`process` only encodes accidental orphaning, which systemd itself discourages.
An earlier draft of this entry proposed `mixed` and was wrong.
The split:
```text
orchestra-worker.service
└── orchestra-worker
orchestra-tmux.service
└── tmux -L orchestra
└── claude sessions
```
`deploy/orchestra-tmux.service` is new. `deploy/orchestra-worker.service` gains
`After=`/`Wants=` on it, ordering only: a worker that finds the runtime missing
must report that rather than be stopped by it. The unit carries an idle
`orchestra-runtime` session so the server outlives its last agent pane.
Probed live on a throwaway socket: `tmux -L <s> new-session -d` leaves the
server running under its own pid after the client exits, so `Type=forking`
resolves a main pid. `systemd-analyze verify` passes.
`User` must match between the two units, because the socket lives under
`/tmp/tmux-$UID`. The installed worker unit on workpc runs as `kami` while
`deploy/orchestra-worker.service` still says `orchestra`. The staged copy at
`~/orchestra-deploy/orchestra-tmux.service` is set to `kami` to match reality.
### What F28 retroactively explains
- Run 3's `no server running` at 11:02:16 followed the 10:45:26 worker restart.
- The same at 13:06:22.
- Run 4's refusal could not be delivered at 13:32:16, seconds after the
13:31:51 restart.
- F16's missing-pane branch has been firing on deployment, not on real
execution-runtime loss. Its proof stands as written, but the trigger was
self-inflicted.
### F28 proof, to run after the split
```text
1. start task and obtain live pane P
2. record pane/session identity
3. systemctl restart orchestra-worker
4. assert pane P still exists with identical agent identity
5. new worker starts
6. worker reconciles lease + pane P
7. no TaskReleased / new lease caused merely by worker restart
8. agent continues under the same lease epoch
```
Then separately, to restore the meaning of F16's missing-pane branch:
```text
restart orchestra-tmux
→ pane disappears
→ F16 refuses renewal
→ lease expires and requeues
```
### Run 4 diagnostic continuation, partial
The chain ran and failed only at the last step, at 13:32:16:
```
deliver phase refusal 06G46P6KE25Y04VVF7VRZMHZ78: tmux send-keys ... -l -- Orchestra refused your phase request: phase request refused: task 06G46P6KE25Y04VVF7VRZMHZ78 may only move to "research", not "implement"
...: no server running on /tmp/tmux-1000/orchestra: exit status 1
```
Live-proven by that one line: F25, the boundary ran on claude at all; F21's
reading half; F26's premise, the coordinator naming the one valid target; F27,
the 409 classified and routed to `sendPrompt`. The request file survived the
failed send, which is the correct branch.
Not proven: the receipt itself, because F28 had already destroyed the pane.