95a96d87a5
The deployment moved from orchestra.service to Docker Compose, but both CLAUDE.md and AUDIT.md still described the systemd unit. A stopped unit plus a stale /usr/local/bin/orchestra reads exactly like a dead service, so document where the live deployment actually is and that deploying now means rebuilding the compose images. Also record what a live check found: the running image predates the B18 auth commit, all six herdrs are unreachable (the sole remaining blocker on live proof for B13-B17), GET /v1/tasks has no authz gate but sits behind a ufw-restricted bind, and ntfy still 403s. Sharpen the herdr-logging note: successes are never logged, so a herdr with no log line is up, not down, and the logs can't confirm one coming back. Probe the ports directly. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01535A3Y8RtkAi8wYuWhtkEd
131 lines
7.7 KiB
Markdown
131 lines
7.7 KiB
Markdown
# Orchestra
|
|
|
|
A Go implementation of `orchestra-spec (1).md` — an unattended multi-agent
|
|
task orchestrator that leases coding tasks to CLI harnesses (Claude Code,
|
|
Codex, opencode) running inside `herdr`-managed panes, rotates them across
|
|
context-window limits, and hands off work via a git-anchored continuity
|
|
protocol.
|
|
|
|
Layout: `internal/{domain,store,provider,registry,router,herdr,orchestrator,
|
|
continuity,federation,delivery,authz,operations,admin}` + `cmd/orchestra/main.go`.
|
|
|
|
## Ground truth over documentation
|
|
|
|
This repo has a documented history of code that *looks* wired but isn't —
|
|
packages with tests that pass in isolation while the live call path silently
|
|
no-ops (bare `continue` on error, discarded return values). See `AUDIT.md`
|
|
for the full audit and `progress.md` for a running log. **Before trusting a
|
|
claim in progress.md that something "works" or "is fixed," check the actual
|
|
call site** — the file is written by past sessions of this same assistant and
|
|
has previously overstated completion.
|
|
|
|
The single most reliable way to verify herdr-adapter code is right: don't
|
|
read `internal/herdr/adapter.go` and assume the method names are real. Ping
|
|
the live herdr instance and check.
|
|
|
|
## herdr protocol — verified against a live instance, 2026-07-27
|
|
|
|
- herdr speaks JSON-RPC over a raw TCP (or unix-socket) connection — **not
|
|
HTTP**. `internal/herdr/herdr.go`'s `Client.Call` is the only correct way
|
|
to talk to it; a bare `curl` to the port returns nothing.
|
|
- Request shape: `{"id":"<n>","method":"<name>","params":<object>}`. Herdr's
|
|
Rust JSON-RPC decoder requires `params` to be present and rejects a bare
|
|
`null` — always send `{}` for parameterless calls (the client does this
|
|
automatically).
|
|
- Full real method list is committed at `deploy/herdr-schema.json`, captured
|
|
live from `192.168.1.105:9245` (the `workpc` herdr) since no local `herdr`
|
|
CLI is available in this sandbox — the schema was reconstructed by sending
|
|
an unknown method name and reading the `unknown variant ... expected one
|
|
of ...` error, then probing each method of interest with `params:{}` /
|
|
`params:{pane_id:"nonexistent"}` to read Rust serde's `missing field
|
|
<x>` errors for its param shape.
|
|
- **Confirmed invented (do not use, they don't exist):** `pane.release`,
|
|
`pane.kill`, `pane.rotation_signal`, `pane.status`. If you see these
|
|
anywhere, it's a bug, not a valid call.
|
|
- **Real replacements:** `pane.close({pane_id})` for kill;
|
|
`pane.release_agent({pane_id, source, agent})` for release (structurally
|
|
different — does *not* return a `handoff_ref`, see below). No replacement
|
|
exists for `rotation_signal` — herdr has no concept of Orchestra rotation.
|
|
- **Architectural point that's easy to get wrong:** herdr never produces a
|
|
handoff. The agent writes the handoff artifact (§6.1 of the spec); herdr's
|
|
role in "release" is only to drop its own claim on the pane/agent binding.
|
|
Any adapter code that expects herdr to hand back a `handoff_ref` is wrong
|
|
by construction, independent of whether the method name is right.
|
|
- Protocol version is returned as a **JSON number** (`17`), not a string,
|
|
even though `config.jsonc` declares `"protocol": "17"` as a string.
|
|
`CheckProtocol`'s raw-bytes fallback happens to make this compare correctly
|
|
today — don't "clean up" that code without checking this note first, or it
|
|
might start doing a real numeric-vs-string comparison and break.
|
|
|
|
## Deployment topology (as of 2026-07-29)
|
|
|
|
- **Runs under Docker Compose, not systemd.** `docker compose -f compose.yaml
|
|
-f compose.live.yaml` in `/home/kami/docker-apps/orchestra-web-ui`, building
|
|
both images from this repo: `orchestra-api` (bound `0.0.0.0:9145`) and
|
|
`orchestra-web-ui` (nginx proxy, `127.0.0.1:19145`). Logs are
|
|
`docker logs orchestra-api`. **Deploying a code change means rebuilding the
|
|
compose images** (`up -d --build`) — the running image can silently predate
|
|
recent commits, so compare its build time against `git log`.
|
|
- `orchestra.service` is the **previous** deployment and is retired —
|
|
`compose.live.yaml` requires it stopped, since both bind the same port and
|
|
data dir. Do not start it. (`sudo` isn't available in this sandbox, so
|
|
`systemctl disable` needs the operator.)
|
|
- Config: `/etc/orchestra/orchestra.env` (env vars) + `/etc/orchestra/
|
|
config.jsonc` (projects/machines/herdrs registry) — the deployed copies, not
|
|
the repo's `.orchestra-config/` or `deploy/config.example.jsonc`. The
|
|
container bind-mounts `/etc/orchestra:ro` and its entrypoint sources the env
|
|
file, so Docker never copies the secrets.
|
|
- The `0.0.0.0` bind on 9145 is **intentional**: ufw restricts the port to one
|
|
other LAN machine. Don't report it as an exposure.
|
|
- Two machines in the registry: `homesrv` (192.168.1.104) and `workpc`
|
|
(192.168.1.105), each nominally running 3 herdrs (claude/codex/opencode).
|
|
As of 2026-07-29 **all six are unreachable** — workpc refuses on 9245-9247,
|
|
homesrv times out (filtered). Nothing can be leased until one is brought up.
|
|
- `main.go` logs herdr connection *failures* at startup and **never logs a
|
|
success**, so a herdr with no log line is *up*, not down — absence of a line
|
|
is evidence in the opposite direction, and an operator cannot confirm a
|
|
herdr came back by tailing the logs. Always probe reachability directly.
|
|
- There was a real, live, stuck task as of 2026-07-27: workspace `wA`, task
|
|
id `06FT6CKD9Y98AZRX6X8K3QXFZG`, opencode harness, pane `wA:p1`,
|
|
`agent_status: "blocked"`. Likely stuck because rotation/release could
|
|
never reach it (B2/B5). Check whether it's still stuck before assuming
|
|
fixes here have taken effect operationally — code fixes don't retroactively
|
|
unstick an already-orphaned pane; that needs a manual kill/restart once the
|
|
release path is trustworthy.
|
|
|
|
## The federation fork — read before touching anything cross-machine
|
|
|
|
Two incompatible designs coexist. **Design A** ("drive the remote socket",
|
|
currently deployed via `clients/herdr-bridge.go`) has homesrv call
|
|
`worktree.create`/`agent.start` etc. directly on workpc's herdr over TCP as
|
|
if it were local — meaning anchor validation (`git rev-parse HEAD`) run by
|
|
the coordinator executes on the *wrong machine* relative to the actual
|
|
checkout. **Design B** ("workers pull tasks", `/v1/federation/*`) is fully
|
|
built server-side but has zero clients — no worker binary exists. Decision
|
|
(AUDIT.md, 2026-07-27): keep Design A through Phase 5, commit to Design B in
|
|
Phase 6, with two guardrails landed immediately (refuse to rotate/cleanup a
|
|
lease held by a non-local herdr rather than validate against the wrong
|
|
checkout). Don't build on top of Design A's cross-machine calls without
|
|
reading that section first.
|
|
|
|
## Working conventions
|
|
|
|
- **workpc worker deployment target:** copy the built worker binary to
|
|
`workpc:~/orchestra-deploy/orchestra-worker` (that is,
|
|
`/home/kami/orchestra-deploy/orchestra-worker`), not directly to
|
|
`/usr/local/bin`. The workpc deployment process installs from this staging
|
|
path. Verify the remote checksum and Go build revision before restart.
|
|
- `go build ./...`, `go vet ./...`, and `go test ./...` must all pass — `go
|
|
vet` was broken for a while (duplicate JSON struct tags) and nobody
|
|
noticed because only `build`/`test` were being checked. Always run all
|
|
three.
|
|
- Silent `continue`-on-error is the recurring bug pattern in this codebase
|
|
(adapter lookups, rotation, expiry). When touching `internal/orchestrator`
|
|
or `internal/herdr`, prefer a recorded/observable failure
|
|
(`MonitorHealth` fields) over a bare `continue` — that's literally what
|
|
turned B1/B2 invisible for as long as they were.
|
|
- Don't invoke destructive herdr calls (`pane.close`, `pane.release_agent`)
|
|
against a real pane from an investigative/audit session without asking
|
|
first — there is live operator state on the other end (see the stuck-task
|
|
note above).
|