Record the burn-in build identity and deployment state

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-08-26 18:35:47 +04:00
parent 86b67d9fbf
commit 6f9300b549
2 changed files with 90 additions and 28 deletions
+27
View File
@@ -333,3 +333,30 @@ Probed live state, and the burn-in is blocked on deployment, not on code:
that is *not* evidence about the workpc harnesses, which use a tmux socket and
a unix socket. Federated reachability defers to worker heartbeat. Do not
diagnose harness availability from a TCP probe of 9245.
## Burn-in deployment, 2026-08-26 18:35
Burn-in build identity: `86b67d9fbff459fe5ae3415cc59594e26cc665f3`. The v3 stack
is committed (`7f12c7f`), followed by startup revision logging (`86b67d9`).
- **Coordinator deployed.** Rebuilt on homesrv with `--build-arg
BUILD_REVISION/BUILD_TIME/BUILD_DIRTY`, container recreated, `/readyz` ready.
It now logs `orchestra revision 86b67d9... dirty false` at startup.
- **Worker staged, not installed.** `~/orchestra-deploy/orchestra-worker`,
sha256 `1400358...`. `install` and `systemctl restart` need root, which this
sandbox does not have, so the running worker is still the 2026-07-30 build.
Until it is installed the pair is mismatched and no task should be created.
- **Observability fixed before proceeding**, per the requirement that deployed
identity be evidence. Revision was previously visible only behind the operator
login, and the worker never logged its own build at all. Both now print it at
startup, so `docker logs orchestra-api` and `journalctl -u orchestra-worker`
are sufficient.
- **`deploy/build.sh`** stamps both binaries from one commit and refuses a dirty
tree, so a burn-in run cannot pair a new coordinator with an old worker.
- The deployed coordinator confirms the transport split directly:
`herdr workpc-opencode is worker-owned on workpc; coordinator probe skipped`,
while the three `homesrv-*` herdrs report `dial tcp 192.168.1.104:9245:
connect: connection refused`.
Not done, and both need root: the `/etc/orchestra/worker.env` scrub (mode 0600,
root-owned) and its in-pane verification. No agent should run before that.
+63 -28
View File
@@ -8,39 +8,74 @@ the live owner path establishes conformance.
Do not add workflow features while this is running. Findings decide the next
implementation work.
## Deployment state, probed 2026-08-26
## Burn-in build identity
Burn-in cannot start yet. Two blockers, one config gap.
`86b67d9fbff459fe5ae3415cc59594e26cc665f3`
| Fact | Evidence |
Both halves must report exactly this revision before a task is created. Neither
needs a credential now: the coordinator prints it in `docker logs orchestra-api`
and the worker in `journalctl -u orchestra-worker`. The same object is at
`GET /v1/admin/diagnostics` and `GET /v1/federation/workers` behind the operator
login.
Build both with `deploy/build.sh <outdir>`, which refuses a dirty tree.
## Deployment state, 2026-08-26 18:35
| Half | State |
|---|---|
| The API is up on homesrv | `GET 192.168.1.104:9145/readyz` `ready:true`, store/router/gitea/jsonl all ready |
| The workpc worker is up and serving two harnesses | `orchestra-worker` pid 741, restarted 2026-08-26 11:39, no errors since |
| **The deployed worker predates every v3 unit** | `go version -m /usr/local/bin/orchestra-worker``vcs.revision=97a9c65`, `vcs.time=2026-07-30` |
| **The v3 work is uncommitted** | Both sessions' units are working-tree changes on `webui-and-audit-reconciliation` |
| Codex has no harness entry at all | `/etc/orchestra/harnesses.json` declares only `workpc-claude` (tmux) and `workpc-opencode` (herdr) |
| No homesrv worker is running | Only the API container runs there; the three `homesrv-*` herdrs have no worker |
| Coordinator (homesrv) | **Deployed at 86b67d9.** Rebuilt with `--build-arg BUILD_REVISION`, recreated, `/readyz` ready, self-reporting the revision in its log |
| Worker (workpc) | **Staged, not installed.** `~/orchestra-deploy/orchestra-worker` sha256 `140035811bf22f6c2ae25f0130c80095eea272fe8fed7987312ccee6f418e2a7`. Installed binary is still the 2026-07-30 build |
Consequences:
Two operator steps remain, both needing root:
1. **Commit, then rebuild and redeploy both artifacts.** The API container
(`docker compose -f compose.yaml -f compose.live.yaml up -d --build` in
`/mnt/server/home/kami/docker-apps/orchestra-web-ui`) and the worker binary
(staged at `workpc:~/orchestra-deploy/orchestra-worker`, then installed to
`/usr/local/bin`, then `systemctl restart orchestra-worker`). Confirm the API
revision at `GET /v1/admin/diagnostics` and the worker's at
`go version -m`. Neither currently has phases, review, submission, the agent
surface, or the reconcile escalation.
2. **Codex cannot be burned in until it has a harnesses.json entry.** Three of
four target harnesses are unreachable today: codex is unconfigured, and both
`homesrv-claude`/`homesrv-codex`/`homesrv-opencode` have no worker.
3. The registry is misleading about herdr transports. No entry in the live
`config.jsonc` sets `backend` or `address`, so all six resolve to
`<machine>:9245` (`registry.defaultHerdrPort`), and both ports are closed.
The workpc worker actually uses a tmux socket and a unix socket
(`/home/kami/.config/herdr/herdr.sock`). Federated reachability defers to
worker heartbeat, so this does not block leasing, but a TCP probe of 9245 is
not evidence about any of these harnesses. Do not diagnose from it.
```sh
sudo install -m 0755 /home/kami/orchestra-deploy/orchestra-worker /usr/local/bin/orchestra-worker
sudo systemctl restart orchestra-worker
journalctl -u orchestra-worker -n 5 --no-pager # must print revision 86b67d9...
```
Then scrub the pane environment before letting an agent run.
## Baseline: one harness, one path
`workpc-opencode` only. Codex is a coverage gap, not a prerequisite, and no
worker runs on homesrv. Prove one real path first:
```
real vikunja task
→ workpc-opencode
→ frame/research/plan
→ trajectory gate if configured
→ implement
→ independent review
→ task pr
→ human merge
→ completed
```
Then raise difficulty in this order:
```
1. opencode boring success
2. opencode mid-session human correction
3. opencode forced rotation
4. opencode pr rejection -> fix -> merge
5. opencode reconcile outage/recovery
6. representative cases on claude
7. add the codex adapter/config
8. cross-harness rotation: opencode -> claude/codex
```
## Herdr transports, so nobody misdiagnoses again
No entry in the live `config.jsonc` sets `backend` or `address`, so all six
resolve to `<machine>:9245` (`registry.defaultHerdrPort`). Both ports are
closed. That says nothing about the workpc harnesses: they are worker-owned,
and the coordinator logs `worker-owned on workpc; coordinator probe skipped`
for each. The worker reaches `workpc-claude` over a tmux socket and
`workpc-opencode` over `/home/kami/.config/herdr/herdr.sock`. Do not diagnose
harness availability from a TCP probe of 9245.
## Evidence to record per run