Files
orchestra/BURNIN.md
T

293 lines
12 KiB
Markdown

# Burn-in: live conformance, not more architecture
Written 2026-08-26. The workflow is feature-complete enough to exercise. What
remains is empirical: run the same real task shape through each harness and
classify what breaks. Isolated tests are insufficient by design here, and only
the live owner path establishes conformance.
Do not add workflow features while this is running. Findings decide the next
implementation work.
## Burn-in build identity
`6f9300b549362c4c5788f8845b56aaff9672d993`
Both halves must report exactly this revision before a task is created. Neither
needs a credential now: the coordinator prints it in `docker logs orchestra-api`
and the worker in `journalctl -u orchestra-worker`. The same object is at
`GET /v1/admin/diagnostics` and `GET /v1/federation/workers` behind the operator
login.
Build both with `deploy/build.sh <outdir>`, which refuses a dirty tree. Later
documentation-only commits do not change this identity, so a rebuild either
passes this revision explicitly or accepts the new one and redeploys both
halves. Never one half.
## Deployment state, 2026-08-26 18:35
| Half | State |
|---|---|
| Coordinator (homesrv) | **Deployed at 6f9300b.** Rebuilt with `--build-arg BUILD_REVISION`, recreated, `/readyz` ready, self-reporting the revision in its log |
| Worker (workpc) | **Staged, not installed.** `~/orchestra-deploy/orchestra-worker` sha256 `2b1c43071eaf6b5c15b110de39f204038a9629897e0ce3dacf45be96a6e6529e`. Installed binary is still the 2026-07-30 build |
Two operator steps remain, both needing root:
```sh
sudo install -m 0755 /home/kami/orchestra-deploy/orchestra-worker /usr/local/bin/orchestra-worker
sudo systemctl restart orchestra-worker
journalctl -u orchestra-worker -n 5 --no-pager # must print revision 6f9300b...
```
A third blocker, found while preparing the pane check:
**herdr is not running on workpc.** `herdr status server` reports `not running`,
the socket at `/home/kami/.config/herdr/herdr.sock` refuses connections, and its
log stops at 2026-07-30. `workpc-opencode` therefore cannot start a pane, even
with the worker installed. The worker logs `serving harness workpc-opencode
(opencode) on herdr backend` at startup without touching the socket, so that
line is not evidence the backend is reachable. herdr is an interactive terminal
workspace manager: running `herdr` in a terminal launches or attaches to the
persistent session and starts the server. Confirm with `herdr status server`
before creating a task.
Then scrub the pane environment before letting an agent run.
### Where the pane environment actually comes from
The two workpc harnesses inherit different environments, so one scrub does not
cover both.
- **`workpc-opencode` (herdr backend).** The pane is created by the herdr
daemon, which is a separate long-running process. It inherits herdr's
environment, not the worker's. Scrubbing `worker.env` does nothing here.
Whatever environment herdr is started with is what every opencode pane gets.
- **`workpc-claude` (tmux backend).** `TmuxBackend.StartAgent` runs `tmux
new-session` through `exec.CommandContext` with no `Env` set
(`internal/herdr/tmux.go`). If the tmux server is not already up, the worker
starts it, and that server inherits the worker's full environment. Every
claude pane then inherits it too.
This exposes a conflict the scrub alone cannot resolve. The worker legitimately
needs `ORCHESTRA_WORKER_TOKEN` (or `ORCHESTRA_WORKER_TOKEN_<ID>`) and
`ORCHESTRA_FEDERATION_ADMIT_TOKEN`, and deleting them breaks the worker. Keeping
them means the tmux path hands them to the agent. Closing it needs either a
filtered `cmd.Env` when the backend spawns a pane, or a tmux server started
separately with a clean environment. Not on flow 1's path, because flow 1 is
opencode only. Close it before step 6.
## Baseline: one harness, one path
`workpc-opencode` only. Codex is a coverage gap, not a prerequisite, and no
worker runs on homesrv. Prove one real path first:
```
real vikunja task
→ workpc-opencode
→ frame/research/plan
→ trajectory gate if configured
→ implement
→ independent review
→ task pr
→ human merge
→ completed
```
Then raise difficulty in this order:
```
1. opencode boring success
2. opencode mid-session human correction
3. opencode forced rotation
4. opencode pr rejection -> fix -> merge
5. opencode reconcile outage/recovery
6. representative cases on claude
7. add the codex adapter/config
8. cross-harness rotation: opencode -> claude/codex
```
## Herdr transports, so nobody misdiagnoses again
No entry in the live `config.jsonc` sets `backend` or `address`, so all six
resolve to `<machine>:9245` (`registry.defaultHerdrPort`). Both ports are
closed. That says nothing about the workpc harnesses: they are worker-owned,
and the coordinator logs `worker-owned on workpc; coordinator probe skipped`
for each. The worker reaches `workpc-claude` over a tmux socket and
`workpc-opencode` over `/home/kami/.config/herdr/herdr.sock`. Do not diagnose
harness availability from a TCP probe of 9245.
## Evidence to record per run
One row per run. The launch instruction is now dumped to
`<worktree>/.orchestra/launch.md` (`herdr.LaunchContextFile`) at every launch,
local and federated, so the context is auditable without reading pane
scrollback.
```
harness (claude | codex | opencode)
task id
lease epoch (every epoch, if the task rotates)
session ids (pane ids per epoch)
human decision ids
launch context (sha256 of .orchestra/launch.md, per launch)
handoff ref
git sha (at every boundary: launch, each handoff, review, submit, merge)
review ref
submission ref
pr id
completion receipt
```
## The five flows
1. **Boring success.** `task → research → plan → implement → review → pr → merge`.
2. **Mid-session correction.** Agent is doing A, the human says B, the next
verified boundary delivers B, and A becomes history rather than a competing
instruction.
3. **Rotation.** Agent A hands off, agent B resumes. Switch harnesses across the
boundary if both are up.
4. **Human rejection.** Submit sha A, comment on the pull request, task reopens,
fix at sha B, a fresh review, the same pull request, merge.
5. **Failure path.** Human source goes down, the reconcile streak escalates, the
session hands off, the successor lease is refused, the source recovers, and a
corrected successor starts.
## What to inspect by hand
For every launch, read `.orchestra/launch.md` and ask only:
```
does this agent know
- what the task actually wants?
- what was most recently decided?
- what phase it is in?
- what earlier material is merely historical?
- what to do next?
```
Then compare that against what the model did.
## Classify before fixing
Every failure gets exactly one label before any code is written:
```
authority bug the wrong thing outranked the right thing
context-selection bug the renderer showed or hid the wrong material
lifecycle bug state, lease, phase or event handling is wrong
adapter/harness bug the pane, occupancy, boundary or prompt path is wrong
model-following failure the context was right and the model ignored it
operator-policy gap credentials, deployment or configuration
```
This matters because the cheap response to every failure is another prompt
rule, and prompt rules accumulated to compensate for lifecycle or adapter bugs
are how the enforcement boundary rots. A model-following failure is the only
class a prompt change should ever answer.
## Pane credential cleanup (operator, in parallel)
The agent pane inherits the worker's environment. On workpc the worker is
`orchestra-worker.service`, `User=kami`, `EnvironmentFile=/etc/orchestra/worker.env`
(mode 0600, root-owned, deliberately unreadable here). Remove from that file, and
from anything else the pane inherits:
```
gitea mutation token
vikunja mutation credentials
orchestra operator / system token
```
Retain only `ORCHESTRA_AGENT_TOKEN` and credentials a specific task genuinely
needs. If an agent needs a forge operation, it asks Orchestra to perform it.
Then verify from inside a real pane:
```bash
git push # must fail unless Orchestra supplied the auth
curl .../v1/tasks/<id>/submission # must be 403 for the agent surface
tea pr create # must have no usable mutation credential
```
Note what is *not* enforced by code: nothing in this repo scrubs the pane
environment, and a fake `git` earlier in `$PATH` is not enforcement because
`/usr/bin/git` bypasses it. Credential isolation is the enforcement.
Confinement of network, filesystem and destructive commands belongs to whatever
launches the process (herdr, a container, systemd, bubblewrap), not to
Orchestra, which supplies identity and policy inputs.
## Exit criterion
Roughly 10 to 20 real tasks, with the five flows covered on each harness that
is actually up. Then read the classified failures and decide the next
implementation unit from them.
## Live findings, 2026-08-26 21:35
Both halves are paired at `6f9300b`, and herdr is up (0.7.5, protocol 17). What
remains before flow 1, and what the in-pane credential checks will really show.
### Flow 1 has no repository to run in
`workpc-opencode` declares exactly one project, `test-e2e`, pointing at
`/tmp/test-e2e` and `/tmp/test-e2e-worktrees`. Neither path exists on workpc
after today's reboot, and `/var/lib/orchestra/repos` does not exist there at
all. Nothing can be leased to this harness until it has a project backed by a
durable local repository.
The only configured task source is Gitea `kami/correx` at
`https://gitea.kvmx.ru` (`ORCHESTRA_GITEA_URL/OWNER/REPO`), so a real task means
a `correx` issue. `correx` already has `machine_affinity: ["homesrv","workpc"]`
in the coordinator's `config.jsonc`. What is missing is a workpc-local entry:
```jsonc
// /etc/orchestra/worker-projects.json, on workpc
{
"correx": {
"repo": "/home/kami/orchestra/repos/correx.git",
"worktree_root": "/home/kami/orchestra/worktrees/correx",
"remote": "origin"
}
}
```
Not `/tmp`. The burn-in outlives a reboot.
### There is no vikunja ingest
This repo has no Vikunja provider. `/readyz` reports `gitea` and `jsonl`. A
"real vikunja task" cannot enter Orchestra today: it arrives as a Gitea issue or
through the JSONL watcher.
### `git push` from a pane will succeed, and no Orchestra setting stops it
Panes run as `kami`, the same user as the worker, and `kami` has
`credential.helper store` set globally. An agent in a herdr pane therefore
inherits a usable Gitea credential over HTTPS. `worker.env` is irrelevant to
this. Real isolation needs panes under a different unix user, or a credential
store the pane cannot read. Expect `git push --dry-run` to succeed and record
that as an operator-policy gap, not a code defect.
### The agent surface is currently unauthenticated
`ORCHESTRA_AGENT_TOKEN` is unset on the coordinator, and an unset surface token
means the middleware performs no check for that surface. Probed live:
```
POST /v1/tasks/<id>/decision-request -H 'X-Orchestra-Surface: agent' -> 404 (reached the handler)
POST /v1/tasks/<id>/phase -H 'X-Orchestra-Surface: agent' -> 403 (refused)
```
The capability boundary holds. Authentication does not. Set
`ORCHESTRA_AGENT_TOKEN` in the coordinator `.env` before the burn-in, and note
the same is true of `ORCHESTRA_MCP_TOKEN` and `ORCHESTRA_MAVEN_TOKEN`, both
unset.
### Two more unset settings that change burn-in behaviour
- `ORCHESTRA_FEDERATION_ADMIT_TOKEN` is unset, so any caller may register a new
worker identity. Existing identities stay protected, because registering an
existing id with a different token is refused.
- `ORCHESTRA_REVIEW_ACTORS` is unset, which `human.Trust` reads as "anyone not
explicitly ignored". Flow 4 will accept a task-moving comment from any Gitea
actor. Set it to `kami`.