Files
orchestra/AGENTS.md
T
kami 118ac9fbcb Commit the concurrent session's pending web and docs work
Not my work. These nine files sat uncommitted in the shared checkout
while another session worked on them, and the UI redesign that follows
rewrites web/src/main.tsx and web/src/style.css. Committing first means
that work is recoverable rather than overwritten.

Contents, by inspection rather than by authorship: whitespace
normalisation and edits across main.tsx, 568 added lines of style.css,
client and client test changes, the orchestra-user line in build.sh, and
docs updates to AGENTS.md, AUDIT.md, DEPLOYMENT.md and the env example.

Committed at the operator's explicit instruction.
2026-08-29 02:17:32 +04:00

8.7 KiB

Orchestra

A Go implementation of orchestra-spec (1).md — an unattended multi-agent task orchestrator that leases coding tasks to CLI harnesses (Claude Code, Codex, opencode) running inside herdr-managed panes, rotates them across context-window limits, and hands off work via a git-anchored continuity protocol.

Layout: internal/{domain,store,provider,registry,router,herdr,orchestrator, continuity,federation,delivery,authz,operations,admin} + cmd/orchestra/main.go.

Ground truth over documentation

This repo has a documented history of code that looks wired but isn't — packages with tests that pass in isolation while the live call path silently no-ops (bare continue on error, discarded return values). See AUDIT.md for the full audit; it is also the running log (there is no separate log file). Before trusting a claim in AUDIT.md that something "works" or "is fixed," check the actual call site — the file is written by past sessions of this same assistant and has previously overstated completion.

The single most reliable way to verify herdr-adapter code is right: don't read internal/herdr/adapter.go and assume the method names are real. Ping the live herdr instance and check.

herdr protocol — verified against a live instance, 2026-07-27

  • herdr speaks JSON-RPC over a raw TCP (or unix-socket) connection — not HTTP. internal/herdr/herdr.go's Client.Call is the only correct way to talk to it; a bare curl to the port returns nothing.
  • Request shape: {"id":"<n>","method":"<name>","params":<object>}. Herdr's Rust JSON-RPC decoder requires params to be present and rejects a bare null — always send {} for parameterless calls (the client does this automatically).
  • Full real method list is committed at deploy/herdr-schema.json, captured live from 192.168.1.105:9245 (the workpc herdr) since no local herdr CLI is available in this sandbox — the schema was reconstructed by sending an unknown method name and reading the unknown variant ... expected one of ... error, then probing each method of interest with params:{} / params:{pane_id:"nonexistent"} to read Rust serde's missing field <x> errors for its param shape.
  • Confirmed invented (do not use, they don't exist): pane.release, pane.kill, pane.rotation_signal, pane.status. If you see these anywhere, it's a bug, not a valid call.
  • Real replacements: pane.close({pane_id}) for kill; pane.release_agent({pane_id, source, agent}) for release (structurally different — does not return a handoff_ref, see below). No replacement exists for rotation_signal — herdr has no concept of Orchestra rotation.
  • Architectural point that's easy to get wrong: herdr never produces a handoff. The agent writes the handoff artifact (§6.1 of the spec); herdr's role in "release" is only to drop its own claim on the pane/agent binding. Any adapter code that expects herdr to hand back a handoff_ref is wrong by construction, independent of whether the method name is right.
  • Protocol version is returned as a JSON number (17), not a string, even though config.jsonc declares "protocol": "17" as a string. CheckProtocol's raw-bytes fallback happens to make this compare correctly today — don't "clean up" that code without checking this note first, or it might start doing a real numeric-vs-string comparison and break.

Deployment topology (as of 2026-07-31)

  • Runs under Docker Compose, not systemd. docker compose -f compose.yaml -f compose.live.yaml in /home/kami/docker-apps/orchestra-web-ui, building both images from this repo: orchestra-api (bound 0.0.0.0:9145, which is intentional — ufw restricts the port to one other LAN machine) and orchestra-web-ui (nginx proxy, 127.0.0.1:19145). Logs are docker logs orchestra-api. Deploying a code change means rebuilding the compose images (up -d --build) — the running image can silently predate recent commits, so compare its build time against git log.
  • orchestra.service was the previous deployment; its unit file and redeploy.sh were deleted from deploy/ on 2026-07-31. A stale installed copy must stay stopped — it binds the same port and data dir as the container. orchestra-worker.service is a different, still-current unit.
  • Config: env vars come from .env in the compose directory via env_file:; only config.jsonc is bind-mounted into /etc/orchestra/. There is no container entrypoint script — Dockerfile.api execs /app/orchestra directly. Neither deployed file is the repo's deploy/config.example.jsonc.
  • Browser operator accounts live in $ORCHESTRA_DATA/auth.db. Create or reset one with orchestra-user set -data /data -username NAME while the API is stopped, or use the authenticated Settings screen. The old ORCHESTRA_WEB_USERNAME/ORCHESTRA_WEB_PASSWORD_HASH pair is accepted only for a one-time import into an empty database and should then be removed.
  • Two machines in the registry: homesrv (192.168.1.104) and workpc (192.168.1.105), each nominally running 3 herdrs (claude/codex/opencode). In practice homesrv has no local herdr running (connection refused on 9245) — only workpc's herdr is live and reachable. main.go only logs herdr connection failures at startup, never successes, so "no log line" for a herdr does not mean it's down — check reachability directly.
  • There was a real, live, stuck task as of 2026-07-27: workspace wA, task id 06FT6CKD9Y98AZRX6X8K3QXFZG, opencode harness, pane wA:p1, agent_status: "blocked". Likely stuck because rotation/release could never reach it (B2/B5). Check whether it's still stuck before assuming fixes here have taken effect operationally — code fixes don't retroactively unstick an already-orphaned pane; that needs a manual kill/restart once the release path is trustworthy.

Federation — Design B is the live design (as of 2026-07-30)

Design B ("workers pull tasks", /v1/federation/*) is the live design and has a real client: cmd/orchestra-worker/main.go (~1,131 lines, with tests in cmd/orchestra-worker/main_test.go) is the deployed worker — the workpc OpenCode worker runs it. Build new cross-machine work on Design B.

The Design A guardrail has landed: Coordinator.adapterFor (internal/orchestrator/orchestrator.go, the LocalHerdr check) refuses to resolve an adapter for a session owned by a non-local herdr, returning session %s is owned by non-local herdr %s instead of validating a git anchor (git rev-parse HEAD) against the wrong machine's checkout. Rotation/cleanup therefore no longer act on remote leases.

Design A is gone (deleted 2026-07-31). clients/herdr-bridge.go ("drive the remote socket": homesrv calling worktree.create/agent.start directly on workpc's herdr over TCP as if it were local) was deleted along with the whole clients/ directory — the operator confirmed it is undeployed now that workers carry cross-machine work, which superseded the 2026-07-27 "keep through Phase 5" decision. There is no bridge to preserve; do not reintroduce coordinator-side calls to a remote herdr socket.

Completion is worker-owned. orchestra-worker watches for an .orchestra/done marker, confirms via AgentStatus that the agent is no longer busy, then posts through /v1/federation/* with the lease epoch and expected version. The old harness-hook path — .orchestra-report.md plus POST /v1/harness/complete — is deleted, endpoint, handler, and deploy/hooks/ scripts alike. /v1/harness/turn remains for turn-boundary decisions.

Working conventions

  • workpc worker deployment target: copy the built worker binary to workpc:~/orchestra-deploy/orchestra-worker (that is, /home/kami/orchestra-deploy/orchestra-worker), not directly to /usr/local/bin. The workpc deployment process installs from this staging path. Verify the remote checksum and Go build revision before restart.
  • go build ./..., go vet ./..., and go test ./... must all pass — go vet was broken for a while (duplicate JSON struct tags) and nobody noticed because only build/test were being checked. Always run all three.
  • Silent continue-on-error is the recurring bug pattern in this codebase (adapter lookups, rotation, expiry). When touching internal/orchestrator or internal/herdr, prefer a recorded/observable failure (MonitorHealth fields) over a bare continue — that's literally what turned B1/B2 invisible for as long as they were.
  • Don't invoke destructive herdr calls (pane.close, pane.release_agent) against a real pane from an investigative/audit session without asking first — there is live operator state on the other end (see the stuck-task note above).