Files
orchestra/BURNIN.md
T
2026-08-26 21:42:04 +04:00

14 KiB

Burn-in: live conformance, not more architecture

Written 2026-08-26. The workflow is feature-complete enough to exercise. What remains is empirical: run the same real task shape through each harness and classify what breaks. Isolated tests are insufficient by design here, and only the live owner path establishes conformance.

Do not add workflow features while this is running. Findings decide the next implementation work.

Burn-in build identity

6f9300b549362c4c5788f8845b56aaff9672d993

Both halves must report exactly this revision before a task is created. Neither needs a credential now: the coordinator prints it in docker logs orchestra-api and the worker in journalctl -u orchestra-worker. The same object is at GET /v1/admin/diagnostics and GET /v1/federation/workers behind the operator login.

Build both with deploy/build.sh <outdir>, which refuses a dirty tree. Later documentation-only commits do not change this identity, so a rebuild either passes this revision explicitly or accepts the new one and redeploys both halves. Never one half.

Deployment state, 2026-08-26 18:35

Half State
Coordinator (homesrv) Deployed at 6f9300b. Rebuilt with --build-arg BUILD_REVISION, recreated, /readyz ready, self-reporting the revision in its log
Worker (workpc) Staged, not installed. ~/orchestra-deploy/orchestra-worker sha256 2b1c43071eaf6b5c15b110de39f204038a9629897e0ce3dacf45be96a6e6529e. Installed binary is still the 2026-07-30 build

Two operator steps remain, both needing root:

sudo install -m 0755 /home/kami/orchestra-deploy/orchestra-worker /usr/local/bin/orchestra-worker
sudo systemctl restart orchestra-worker
journalctl -u orchestra-worker -n 5 --no-pager   # must print revision 6f9300b...

A third blocker, found while preparing the pane check:

herdr is not running on workpc. herdr status server reports not running, the socket at /home/kami/.config/herdr/herdr.sock refuses connections, and its log stops at 2026-07-30. workpc-opencode therefore cannot start a pane, even with the worker installed. The worker logs serving harness workpc-opencode (opencode) on herdr backend at startup without touching the socket, so that line is not evidence the backend is reachable. herdr is an interactive terminal workspace manager: running herdr in a terminal launches or attaches to the persistent session and starts the server. Confirm with herdr status server before creating a task.

Then scrub the pane environment before letting an agent run.

Where the pane environment actually comes from

The two workpc harnesses inherit different environments, so one scrub does not cover both.

  • workpc-opencode (herdr backend). The pane is created by the herdr daemon, which is a separate long-running process. It inherits herdr's environment, not the worker's. Scrubbing worker.env does nothing here. Whatever environment herdr is started with is what every opencode pane gets.
  • workpc-claude (tmux backend). TmuxBackend.StartAgent runs tmux new-session through exec.CommandContext with no Env set (internal/herdr/tmux.go). If the tmux server is not already up, the worker starts it, and that server inherits the worker's full environment. Every claude pane then inherits it too.

This exposes a conflict the scrub alone cannot resolve. The worker legitimately needs ORCHESTRA_WORKER_TOKEN (or ORCHESTRA_WORKER_TOKEN_<ID>) and ORCHESTRA_FEDERATION_ADMIT_TOKEN, and deleting them breaks the worker. Keeping them means the tmux path hands them to the agent. Closing it needs either a filtered cmd.Env when the backend spawns a pane, or a tmux server started separately with a clean environment. Not on flow 1's path, because flow 1 is opencode only. Close it before step 6.

Baseline: one harness, one path

workpc-opencode only. Codex is a coverage gap, not a prerequisite, and no worker runs on homesrv. Prove one real path first:

real vikunja task
→ workpc-opencode
→ frame/research/plan
→ trajectory gate if configured
→ implement
→ independent review
→ task pr
→ human merge
→ completed

Then raise difficulty in this order:

1. opencode boring success
2. opencode mid-session human correction
3. opencode forced rotation
4. opencode pr rejection -> fix -> merge
5. opencode reconcile outage/recovery
6. representative cases on claude
7. add the codex adapter/config
8. cross-harness rotation: opencode -> claude/codex

Herdr transports, so nobody misdiagnoses again

No entry in the live config.jsonc sets backend or address, so all six resolve to <machine>:9245 (registry.defaultHerdrPort). Both ports are closed. That says nothing about the workpc harnesses: they are worker-owned, and the coordinator logs worker-owned on workpc; coordinator probe skipped for each. The worker reaches workpc-claude over a tmux socket and workpc-opencode over /home/kami/.config/herdr/herdr.sock. Do not diagnose harness availability from a TCP probe of 9245.

Evidence to record per run

One row per run. The launch instruction is now dumped to <worktree>/.orchestra/launch.md (herdr.LaunchContextFile) at every launch, local and federated, so the context is auditable without reading pane scrollback.

harness            (claude | codex | opencode)
task id
lease epoch        (every epoch, if the task rotates)
session ids        (pane ids per epoch)
human decision ids
launch context     (sha256 of .orchestra/launch.md, per launch)
handoff ref
git sha            (at every boundary: launch, each handoff, review, submit, merge)
review ref
submission ref
pr id
completion receipt

The five flows

  1. Boring success. task → research → plan → implement → review → pr → merge.
  2. Mid-session correction. Agent is doing A, the human says B, the next verified boundary delivers B, and A becomes history rather than a competing instruction.
  3. Rotation. Agent A hands off, agent B resumes. Switch harnesses across the boundary if both are up.
  4. Human rejection. Submit sha A, comment on the pull request, task reopens, fix at sha B, a fresh review, the same pull request, merge.
  5. Failure path. Human source goes down, the reconcile streak escalates, the session hands off, the successor lease is refused, the source recovers, and a corrected successor starts.

What to inspect by hand

For every launch, read .orchestra/launch.md and ask only:

does this agent know
- what the task actually wants?
- what was most recently decided?
- what phase it is in?
- what earlier material is merely historical?
- what to do next?

Then compare that against what the model did.

Classify before fixing

Every failure gets exactly one label before any code is written:

authority bug            the wrong thing outranked the right thing
context-selection bug    the renderer showed or hid the wrong material
lifecycle bug            state, lease, phase or event handling is wrong
adapter/harness bug       the pane, occupancy, boundary or prompt path is wrong
model-following failure  the context was right and the model ignored it
operator-policy gap      credentials, deployment or configuration

This matters because the cheap response to every failure is another prompt rule, and prompt rules accumulated to compensate for lifecycle or adapter bugs are how the enforcement boundary rots. A model-following failure is the only class a prompt change should ever answer.

Pane credential cleanup (operator, in parallel)

The agent pane inherits the worker's environment. On workpc the worker is orchestra-worker.service, User=kami, EnvironmentFile=/etc/orchestra/worker.env (mode 0600, root-owned, deliberately unreadable here). Remove from that file, and from anything else the pane inherits:

gitea mutation token
vikunja mutation credentials
orchestra operator / system token

Retain only ORCHESTRA_AGENT_TOKEN and credentials a specific task genuinely needs. If an agent needs a forge operation, it asks Orchestra to perform it.

Then verify from inside a real pane:

git push                 # must fail unless Orchestra supplied the auth
curl .../v1/tasks/<id>/submission   # must be 403 for the agent surface
tea pr create            # must have no usable mutation credential

Note what is not enforced by code: nothing in this repo scrubs the pane environment, and a fake git earlier in $PATH is not enforcement because /usr/bin/git bypasses it. Credential isolation is the enforcement. Confinement of network, filesystem and destructive commands belongs to whatever launches the process (herdr, a container, systemd, bubblewrap), not to Orchestra, which supplies identity and policy inputs.

Exit criterion

Roughly 10 to 20 real tasks, with the five flows covered on each harness that is actually up. Then read the classified failures and decide the next implementation unit from them.

Live findings, 2026-08-26 21:35

Both halves are paired at 6f9300b, and herdr is up (0.7.5, protocol 17). What remains before flow 1, and what the in-pane credential checks will really show.

Flow 1 has no repository to run in

workpc-opencode declares exactly one project, test-e2e, pointing at /tmp/test-e2e and /tmp/test-e2e-worktrees. Neither path exists on workpc after today's reboot, and /var/lib/orchestra/repos does not exist there at all. Nothing can be leased to this harness until it has a project backed by a durable local repository.

The only configured task source is Gitea kami/correx at https://gitea.kvmx.ru (ORCHESTRA_GITEA_URL/OWNER/REPO), so a real task means a correx issue. correx already has machine_affinity: ["homesrv","workpc"] in the coordinator's config.jsonc. What is missing is a workpc-local entry:

// /etc/orchestra/worker-projects.json, on workpc
{
  "correx": {
    "repo": "/home/kami/orchestra/repos/correx.git",
    "worktree_root": "/home/kami/orchestra/worktrees/correx",
    "remote": "origin"
  }
}

Not /tmp. The burn-in outlives a reboot.

There is no vikunja ingest

This repo has no Vikunja provider. /readyz reports gitea and jsonl. A "real vikunja task" cannot enter Orchestra today: it arrives as a Gitea issue or through the JSONL watcher.

git push from a pane will succeed, and no Orchestra setting stops it

Corrected 21:45. The mechanism is SSH, not HTTPS. HTTPS has no stored credential on workpc, where git push over https:// fails with could not read Username. Pushing works over ssh://git@gitea.kvmx.ru:2222 using kami's default RSA identity, which Gitea lists as the key named workpc. It carries no passphrase, so no agent is needed.

That identity is what lets the worker push, and panes run as the same user with the same home directory, so an agent inherits it. worker.env is irrelevant to this. Real isolation needs panes under a different unix user. Expect git push --dry-run to succeed from a pane, and record it as an operator-policy gap rather than a code defect.

The worker inherits ambient Git credentials by design: git() in cmd/orchestra-worker/main.go runs exec.CommandContext with no environment of its own. There is no separate credential path for the worker to hold something the pane does not.

The agent surface is currently unauthenticated

ORCHESTRA_AGENT_TOKEN is unset on the coordinator, and an unset surface token means the middleware performs no check for that surface. Probed live:

POST /v1/tasks/<id>/decision-request  -H 'X-Orchestra-Surface: agent'  -> 404 (reached the handler)
POST /v1/tasks/<id>/phase             -H 'X-Orchestra-Surface: agent'  -> 403 (refused)

The capability boundary holds. Authentication does not. Set ORCHESTRA_AGENT_TOKEN in the coordinator .env before the burn-in, and note the same is true of ORCHESTRA_MCP_TOKEN and ORCHESTRA_MAVEN_TOKEN, both unset.

Two more unset settings that change burn-in behaviour

  • ORCHESTRA_FEDERATION_ADMIT_TOKEN is unset, so any caller may register a new worker identity. Existing identities stay protected, because registering an existing id with a different token is refused.
  • ORCHESTRA_REVIEW_ACTORS is unset, which human.Trust reads as "anyone not explicitly ignored". Flow 4 will accept a task-moving comment from any Gitea actor. Set it to kami.

Flow 1 setup, 2026-08-26 21:45

Target: kami/test-e2e on Gitea, a throwaway repo, rather than correx.

Done:

  • Ingest switched. ORCHESTRA_GITEA_REPO=test-e2e in the coordinator .env (previous file kept as .env.pre-burnin-20260826), container recreated, still reporting 6f9300b. The project id must equal the repo name, because main.go sets Project: os.Getenv("ORCHESTRA_GITEA_REPO"). test-e2e already exists in config.jsonc with machine_affinity: ["workpc"].
  • Durable repo on workpc. Bare clone at /home/kami/orchestra/repos/test-e2e.git, worktree root /home/kami/orchestra/worktrees/test-e2e. Origin points at ssh://git@gitea.kvmx.ru:2222/kami/test-e2e.git, and git push --dry-run reports Everything up-to-date. The old /tmp paths did not survive the reboot and must not come back.

Remaining, needs root:

// /etc/orchestra/worker-projects.json
{
  "test-e2e": {
    "repo": "/home/kami/orchestra/repos/test-e2e.git",
    "worktree_root": "/home/kami/orchestra/worktrees/test-e2e",
    "remote": "origin"
  }
}

Then sudo systemctl restart orchestra-worker.

Store state, checked before the first run: 30 tasks, all test-e2e, with 23 blocked, 6 completed and 1 failed. All are July leftovers and none holds a lease. The coordinator has run ResumeAnsweredBlockers every second for hours without resuming any of them, so they are inert rather than merely quiet. The July stuck task 06FT6CKD9Y98AZRX6X8K3QXFZG is now failed.

A stale branch orchestra/scratch/oc-06ftgkjadcd2hwjn2zwjen90q4 exists on the remote from an earlier run. Harmless, but it is not from this burn-in.