diff --git a/AUDIT.md b/AUDIT.md index 80e25fe..4881e9c 100644 --- a/AUDIT.md +++ b/AUDIT.md @@ -333,3 +333,30 @@ Probed live state, and the burn-in is blocked on deployment, not on code: that is *not* evidence about the workpc harnesses, which use a tmux socket and a unix socket. Federated reachability defers to worker heartbeat. Do not diagnose harness availability from a TCP probe of 9245. + +## Burn-in deployment, 2026-08-26 18:35 + +Burn-in build identity: `86b67d9fbff459fe5ae3415cc59594e26cc665f3`. The v3 stack +is committed (`7f12c7f`), followed by startup revision logging (`86b67d9`). + +- **Coordinator deployed.** Rebuilt on homesrv with `--build-arg + BUILD_REVISION/BUILD_TIME/BUILD_DIRTY`, container recreated, `/readyz` ready. + It now logs `orchestra revision 86b67d9... dirty false` at startup. +- **Worker staged, not installed.** `~/orchestra-deploy/orchestra-worker`, + sha256 `1400358...`. `install` and `systemctl restart` need root, which this + sandbox does not have, so the running worker is still the 2026-07-30 build. + Until it is installed the pair is mismatched and no task should be created. +- **Observability fixed before proceeding**, per the requirement that deployed + identity be evidence. Revision was previously visible only behind the operator + login, and the worker never logged its own build at all. Both now print it at + startup, so `docker logs orchestra-api` and `journalctl -u orchestra-worker` + are sufficient. +- **`deploy/build.sh`** stamps both binaries from one commit and refuses a dirty + tree, so a burn-in run cannot pair a new coordinator with an old worker. +- The deployed coordinator confirms the transport split directly: + `herdr workpc-opencode is worker-owned on workpc; coordinator probe skipped`, + while the three `homesrv-*` herdrs report `dial tcp 192.168.1.104:9245: + connect: connection refused`. + +Not done, and both need root: the `/etc/orchestra/worker.env` scrub (mode 0600, +root-owned) and its in-pane verification. No agent should run before that. diff --git a/BURNIN.md b/BURNIN.md index 9bac248..1e9502e 100644 --- a/BURNIN.md +++ b/BURNIN.md @@ -8,39 +8,74 @@ the live owner path establishes conformance. Do not add workflow features while this is running. Findings decide the next implementation work. -## Deployment state, probed 2026-08-26 +## Burn-in build identity -Burn-in cannot start yet. Two blockers, one config gap. +`86b67d9fbff459fe5ae3415cc59594e26cc665f3` -| Fact | Evidence | +Both halves must report exactly this revision before a task is created. Neither +needs a credential now: the coordinator prints it in `docker logs orchestra-api` +and the worker in `journalctl -u orchestra-worker`. The same object is at +`GET /v1/admin/diagnostics` and `GET /v1/federation/workers` behind the operator +login. + +Build both with `deploy/build.sh `, which refuses a dirty tree. + +## Deployment state, 2026-08-26 18:35 + +| Half | State | |---|---| -| The API is up on homesrv | `GET 192.168.1.104:9145/readyz` → `ready:true`, store/router/gitea/jsonl all ready | -| The workpc worker is up and serving two harnesses | `orchestra-worker` pid 741, restarted 2026-08-26 11:39, no errors since | -| **The deployed worker predates every v3 unit** | `go version -m /usr/local/bin/orchestra-worker` → `vcs.revision=97a9c65`, `vcs.time=2026-07-30` | -| **The v3 work is uncommitted** | Both sessions' units are working-tree changes on `webui-and-audit-reconciliation` | -| Codex has no harness entry at all | `/etc/orchestra/harnesses.json` declares only `workpc-claude` (tmux) and `workpc-opencode` (herdr) | -| No homesrv worker is running | Only the API container runs there; the three `homesrv-*` herdrs have no worker | +| Coordinator (homesrv) | **Deployed at 86b67d9.** Rebuilt with `--build-arg BUILD_REVISION`, recreated, `/readyz` ready, self-reporting the revision in its log | +| Worker (workpc) | **Staged, not installed.** `~/orchestra-deploy/orchestra-worker` sha256 `140035811bf22f6c2ae25f0130c80095eea272fe8fed7987312ccee6f418e2a7`. Installed binary is still the 2026-07-30 build | -Consequences: +Two operator steps remain, both needing root: -1. **Commit, then rebuild and redeploy both artifacts.** The API container - (`docker compose -f compose.yaml -f compose.live.yaml up -d --build` in - `/mnt/server/home/kami/docker-apps/orchestra-web-ui`) and the worker binary - (staged at `workpc:~/orchestra-deploy/orchestra-worker`, then installed to - `/usr/local/bin`, then `systemctl restart orchestra-worker`). Confirm the API - revision at `GET /v1/admin/diagnostics` and the worker's at - `go version -m`. Neither currently has phases, review, submission, the agent - surface, or the reconcile escalation. -2. **Codex cannot be burned in until it has a harnesses.json entry.** Three of - four target harnesses are unreachable today: codex is unconfigured, and both - `homesrv-claude`/`homesrv-codex`/`homesrv-opencode` have no worker. -3. The registry is misleading about herdr transports. No entry in the live - `config.jsonc` sets `backend` or `address`, so all six resolve to - `:9245` (`registry.defaultHerdrPort`), and both ports are closed. - The workpc worker actually uses a tmux socket and a unix socket - (`/home/kami/.config/herdr/herdr.sock`). Federated reachability defers to - worker heartbeat, so this does not block leasing, but a TCP probe of 9245 is - not evidence about any of these harnesses. Do not diagnose from it. +```sh +sudo install -m 0755 /home/kami/orchestra-deploy/orchestra-worker /usr/local/bin/orchestra-worker +sudo systemctl restart orchestra-worker +journalctl -u orchestra-worker -n 5 --no-pager # must print revision 86b67d9... +``` + +Then scrub the pane environment before letting an agent run. + +## Baseline: one harness, one path + +`workpc-opencode` only. Codex is a coverage gap, not a prerequisite, and no +worker runs on homesrv. Prove one real path first: + +``` +real vikunja task +→ workpc-opencode +→ frame/research/plan +→ trajectory gate if configured +→ implement +→ independent review +→ task pr +→ human merge +→ completed +``` + +Then raise difficulty in this order: + +``` +1. opencode boring success +2. opencode mid-session human correction +3. opencode forced rotation +4. opencode pr rejection -> fix -> merge +5. opencode reconcile outage/recovery +6. representative cases on claude +7. add the codex adapter/config +8. cross-harness rotation: opencode -> claude/codex +``` + +## Herdr transports, so nobody misdiagnoses again + +No entry in the live `config.jsonc` sets `backend` or `address`, so all six +resolve to `:9245` (`registry.defaultHerdrPort`). Both ports are +closed. That says nothing about the workpc harnesses: they are worker-owned, +and the coordinator logs `worker-owned on workpc; coordinator probe skipped` +for each. The worker reaches `workpc-claude` over a tmux socket and +`workpc-opencode` over `/home/kami/.config/herdr/herdr.sock`. Do not diagnose +harness availability from a TCP probe of 9245. ## Evidence to record per run