Files
orchestra/CLAUDE.md
T
kami bdc0d4d5be Say that a coordinator deploy leaves the console behind
The ethos console sat undeployed for hours while two coordinator deploys
went out, because both used --no-deps orchestra-api and the web image is
built separately. Checking the commit does not catch it; checking the
served bundle does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CVbaKucEYBjMqVeUgJUsc1
2026-08-29 20:02:29 +04:00

163 lines
9.8 KiB
Markdown

# Orchestra
A Go implementation of `orchestra-spec (1).md` — an unattended multi-agent
task orchestrator that leases coding tasks to CLI harnesses (Claude Code,
Codex, opencode) running inside `herdr`-managed panes, rotates them across
context-window limits, and hands off work via a git-anchored continuity
protocol.
Layout: `internal/{domain,store,provider,registry,router,herdr,orchestrator,
continuity,federation,delivery,authz,operations,admin}` + `cmd/orchestra/main.go`.
## Ground truth over documentation
This repo has a documented history of code that *looks* wired but isn't —
packages with tests that pass in isolation while the live call path silently
no-ops (bare `continue` on error, discarded return values). See `AUDIT.md`
for the full audit; it is also the running log (there is no separate log
file). **Before trusting a claim in `AUDIT.md` that something "works"
or "is fixed," check the actual call site** — the file is written by past
sessions of this same assistant and has previously overstated completion.
The single most reliable way to verify herdr-adapter code is right: don't
read `internal/herdr/adapter.go` and assume the method names are real. Ping
the live herdr instance and check.
## herdr protocol — verified against a live instance, 2026-07-27
- herdr speaks JSON-RPC over a raw TCP (or unix-socket) connection — **not
HTTP**. `internal/herdr/herdr.go`'s `Client.Call` is the only correct way
to talk to it; a bare `curl` to the port returns nothing.
- Request shape: `{"id":"<n>","method":"<name>","params":<object>}`. Herdr's
Rust JSON-RPC decoder requires `params` to be present and rejects a bare
`null` — always send `{}` for parameterless calls (the client does this
automatically).
- Full real method list is committed at `deploy/herdr-schema.json`, captured
live from `192.168.1.105:9245` (the `workpc` herdr) since no local `herdr`
CLI is available in this sandbox — the schema was reconstructed by sending
an unknown method name and reading the `unknown variant ... expected one
of ...` error, then probing each method of interest with `params:{}` /
`params:{pane_id:"nonexistent"}` to read Rust serde's `missing field
<x>` errors for its param shape.
- **Confirmed invented (do not use, they don't exist):** `pane.release`,
`pane.kill`, `pane.rotation_signal`, `pane.status`. If you see these
anywhere, it's a bug, not a valid call.
- **Real replacements:** `pane.close({pane_id})` for kill;
`pane.release_agent({pane_id, source, agent})` for release (structurally
different — does *not* return a `handoff_ref`, see below). No replacement
exists for `rotation_signal` — herdr has no concept of Orchestra rotation.
- **Architectural point that's easy to get wrong:** herdr never produces a
handoff. The agent writes the handoff artifact (§6.1 of the spec); herdr's
role in "release" is only to drop its own claim on the pane/agent binding.
Any adapter code that expects herdr to hand back a `handoff_ref` is wrong
by construction, independent of whether the method name is right.
- Protocol version is returned as a **JSON number** (`17`), not a string,
even though `config.jsonc` declares `"protocol": "17"` as a string.
`CheckProtocol`'s raw-bytes fallback happens to make this compare correctly
today — don't "clean up" that code without checking this note first, or it
might start doing a real numeric-vs-string comparison and break.
## Deployment topology (as of 2026-07-29)
- **Runs under Docker Compose, not systemd.** `docker compose -f compose.yaml
-f compose.live.yaml` in `/home/kami/docker-apps/orchestra-web-ui`, building
both images from this repo: `orchestra-api` (bound `0.0.0.0:9145`) and
`orchestra-web-ui` (nginx proxy, `127.0.0.1:19145`). Logs are
`docker logs orchestra-api`. **Deploying a code change means rebuilding the
compose images** (`up -d --build`) — the running image can silently predate
recent commits, so compare its build time against `git log`.
- **A coordinator deploy does not deploy the console.** They are two images
built from the same repo, and the usual `up -d --no-deps orchestra-api`
leaves `orchestra-web-ui` on whatever it was. This has bitten twice: the
ethos console landed in the repo on 2026-08-29 02:44 and was still serving a
2026-07-30 image hours later, through two coordinator deploys. Rebuild it
explicitly from the same clean worktree (`docker build` in `web/`, then
`up -d --no-deps orchestra-web-ui`), and check the served bundle rather than
the commit: `curl -s http://127.0.0.1:19145/assets/<css> | grep 8F7AE5`.
- `orchestra.service` was the **previous** deployment; the unit file was
deleted from `deploy/` on 2026-07-31 along with `redeploy.sh` (which
`sudo install`ed to `/usr/local/bin` and restarted it). If a stale copy is
still installed on a host, it must stay stopped — it binds the same port and
data dir as the container. (`sudo` isn't available in this sandbox, so
`systemctl disable` needs the operator.) `orchestra-worker.service` is a
*different*, still-current unit — don't delete it by association.
- Config: env vars come from **`.env` in the compose directory**, loaded via
`env_file:` in `compose.yaml` — that is where the Gitea/ntfy/web tokens and
the bcrypt operator hash live. The only thing bind-mounted into
`/etc/orchestra/` is a single file, `config.jsonc`
(projects/machines/herdrs registry), via `compose.override.yaml`. There is no
container entrypoint script — `Dockerfile.api` execs `/app/orchestra`
directly, and `ORCHESTRA_DATA`/`ORCHESTRA_PORT` come from the image `ENV`
plus compose. Neither the deployed `config.jsonc` nor `.env` is the repo's
`deploy/config.example.jsonc`.
- The `0.0.0.0` bind on 9145 is **intentional**: ufw restricts the port to one
other LAN machine. Don't report it as an exposure.
- Two machines in the registry: `homesrv` (192.168.1.104) and `workpc`
(192.168.1.105), each nominally running 3 herdrs (claude/codex/opencode).
As of 2026-07-29 **all six are unreachable** — workpc refuses on 9245-9247,
homesrv times out (filtered). Nothing can be leased until one is brought up.
- `main.go` logs herdr connection *failures* at startup and **never logs a
success**, so a herdr with no log line is *up*, not down — absence of a line
is evidence in the opposite direction, and an operator cannot confirm a
herdr came back by tailing the logs. Always probe reachability directly.
- There was a real, live, stuck task as of 2026-07-27: workspace `wA`, task
id `06FT6CKD9Y98AZRX6X8K3QXFZG`, opencode harness, pane `wA:p1`,
`agent_status: "blocked"`. Likely stuck because rotation/release could
never reach it (B2/B5). Check whether it's still stuck before assuming
fixes here have taken effect operationally — code fixes don't retroactively
unstick an already-orphaned pane; that needs a manual kill/restart once the
release path is trustworthy.
## Federation — Design B is the live design (as of 2026-07-30)
**Design B** ("workers pull tasks", `/v1/federation/*`) is the live design and
has a real client: `cmd/orchestra-worker/main.go` (~1,131 lines, with tests in
`cmd/orchestra-worker/main_test.go`) is the deployed worker — the workpc
OpenCode worker runs it. Build new cross-machine work on Design B.
The **Design A guardrail has landed**: `Coordinator.adapterFor`
(`internal/orchestrator/orchestrator.go`, the `LocalHerdr` check) refuses to
resolve an adapter for a session owned by a non-local herdr, returning
`session %s is owned by non-local herdr %s` instead of validating a git anchor
(`git rev-parse HEAD`) against the wrong machine's checkout. Rotation/cleanup
therefore no longer act on remote leases.
**Design A is gone (deleted 2026-07-31).** `clients/herdr-bridge.go` ("drive
the remote socket": homesrv calling `worktree.create`/`agent.start` directly on
workpc's herdr over TCP as if it were local) was deleted along with the whole
`clients/` directory — the operator confirmed it is undeployed now that workers
carry cross-machine work, which superseded the 2026-07-27 "keep through Phase
5" decision. There is no bridge to preserve; do not reintroduce coordinator-side
calls to a remote herdr socket.
**Completion is worker-owned.** `orchestra-worker` watches for an
`.orchestra/done` marker in the worktree, confirms via `AgentStatus` that the
agent is no longer busy (a marker alone is intent, not proof), then finalizes
and posts through `/v1/federation/*` with both the lease epoch and the expected
version. The old harness-hook path — `.orchestra-report.md` plus
`POST /v1/harness/complete` — is **deleted**: the endpoint, its handler, and the
`deploy/hooks/` scripts are all gone, because an unaffiliated hook has no
durable worker identity or fencing epoch. `/v1/harness/turn` remains for
turn-boundary decisions.
## Working conventions
- **workpc worker deployment target:** copy the built worker binary to
`workpc:~/orchestra-deploy/orchestra-worker` (that is,
`/home/kami/orchestra-deploy/orchestra-worker`), not directly to
`/usr/local/bin`. The workpc deployment process installs from this staging
path. Verify the remote checksum and Go build revision before restart.
- `go build ./...`, `go vet ./...`, and `go test ./...` must all pass — `go
vet` was broken for a while (duplicate JSON struct tags) and nobody
noticed because only `build`/`test` were being checked. Always run all
three.
- Silent `continue`-on-error is the recurring bug pattern in this codebase
(adapter lookups, rotation, expiry). When touching `internal/orchestrator`
or `internal/herdr`, prefer a recorded/observable failure
(`MonitorHealth` fields) over a bare `continue` — that's literally what
turned B1/B2 invisible for as long as they were.
- Don't invoke destructive herdr calls (`pane.close`, `pane.release_agent`)
against a real pane from an investigative/audit session without asking
first — there is live operator state on the other end (see the stuck-task
note above).