The tag orchestra-conformance-v1 marks the full original lifecycle, live-proven by run 6. Both deployed halves sit there deliberately. The four plan-machinery commits above it pass build, vet and test, and have never run against a live agent. The handoff records what they need before they can: a verification policy on test-e2e, both halves rebuilt together, and F18 fixed first, since the burn-in will generate the retry loops it exists to make visible. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CVbaKucEYBjMqVeUgJUsc1
13 KiB
Handoff: the lifecycle is proven and frozen, the plan machinery is written and unproven
Written 2026-08-28, 12:22 local (08:22 UTC). Read with BURNIN.md (the ledger,
current through F46 and run 6), PLAN-SPEC-DESIGN.md (the design this
implements), HANDOFF-2026-08-28-burnin-5.md (the session before this one),
AUDIT.md and CLAUDE.md.
Everything below was observed live unless it says otherwise.
The headline
Two things happened, in this order, deliberately.
The original Orchestra lifecycle is proven end to end and tagged. Run 6
carried issue kami/test-e2e#7 from creation to TaskCompleted. The chain
covered a real pull request, a trusted human rejection, a resubmission onto the
same pull request, and a merge. No manual lifecycle intervention.
The plan machinery is written, tested, and deployed nowhere. Four commits
sit above the tag. go build, go vet and go test ./... pass across 21
packages, and not one line of it has run against a live agent.
orchestra-conformance-v1 full lifecycle proven, both halves deployed here
↓ four commits
HEAD plan.md + verified phase progress
+ policy-bounded verification + typed mismatch
Deployed state
| Half | Revision | Evidence |
|---|---|---|
| Worker, workpc systemd | 063a3ab |
PID 3840261, sha256 51321b93…, verified from /proc/<pid>/exe |
| Coordinator, homesrv container | 4af9880 |
docker logs orchestra-api prints the revision at startup |
Both are the tag. 063a3ab contains 4af9880: the two halves are one linear
history, and the later commits are worker-only.
Verify the running worker, not the installed file. This bit two sessions
ago, twice in one night. sudo install lands the new binary while the old
process keeps serving from the unlinked inode. Always
sha256sum /proc/$(systemctl show -p MainPID --value orchestra-worker)/exe.
sudo needs the operator in this sandbox. The worker deploy staging path is
~/orchestra-deploy/orchestra-worker, and it currently holds the same
51321b93… that is running.
What run 6 proved
Task 06G4E6F69AKP2PA00S28D7ASBC, issue kami/test-e2e#7, workpc-claude.
06:41:14 TaskCreated issue #7 polled
06:41:37 frame -> research
06:42:57 research -> plan
06:45:12 plan -> implement
06:46:27 implement -> review
06:47:48 ReviewRecorded blocking 0, sha 41dee802
06:47:49 TaskSubmitted PR 8 opened
06:55:11 TaskChangesRequested trusted comment by kami
06:55:11 review -> implement
06:57:47 implement -> review
07:00:09 ReviewRecorded blocking 0, sha 0d1585df
07:00:09 TaskSubmitted PR 8 reused
07:02:07 merge by kami
07:02:11 TaskCompleted receipt binds merge f48e363d
Gate sha, review sha, submitted sha and the pull request head are one commit,
0d1585df. Each was checked against its artifact rather than its log line. The
full acceptance table is in BURNIN.md under "Run 6".
The premise that was wrong
The previous handoff said submission, human review and merge needed
configuration rather than code. The event log held zero ReviewRecorded and
zero TaskSubmitted across its whole history. Nothing had ever called either
endpoint. Four defects, all live-proven and fixed:
- F43: the reviewed change had no path to the human. The worker completed
directly, and review, submission and the publisher had no live caller.
e8d04d7 - F44: the Gitea source ingested Orchestra's own pull request as a new task,
one per cycle, forever.
6ccc755 - F45: the submission reflection loop was dead code, guarded on a map filled
600 lines later. A merge could never complete a task.
4af9880 - F46: stale
review.jsonsurvives a changes-requested round trip.063a3ab
Plus c11bf0e, delivering a refusal to a reviewer whose findings file is
missing rather than retrying every five seconds into worker health.
The lesson worth carrying: a documented endpoint is not a live call path. Grep for the caller before believing a feature exists. This is the third time this repo has produced that failure.
What the four new commits do
822f086, research findings are citable
Finding gains ID and Confidence (fact|inference|assumption). The brief
at agentctx.go had advertised both for months while the struct dropped them
silently on every seal.
DecodeStoredResearch reads what is already in the CAS and backfills both,
labelling a backfilled finding inference. DecodeResearch stays strict for
new seals.
57c028f, plan.md
Markdown, 128 KiB, no per-line cap, same seal path and same PlanRef. The old
artifact capped every entry at 500 single-line characters, which makes a
specification impossible.
The parser enforces required sections, contiguous phase numbering from 1,
Files/Changes/Verification per phase, and at least one check per phase.
run: lines are JSON argv arrays. Headings inside fenced blocks are content,
so a plan can show markdown without parsing its own example.
Citations resolve on the coordinator at seal time, the only party holding
ResearchRef.
The property this exists for: the plan renders byte for byte into the
implement launch. A rotated successor receives the same complete document.
TestAcceptedPlanRendersVerbatim and TestRotatedSuccessorReceivesTheWholePlan
guard it.
DecodeStoredPlan renders a legacy JSON plan into the same type, carrying zero
phases. The implement brief then says "legacy accepted plan, phase progress
unavailable" and the task continues under old semantics.
a221502, progress Orchestra establishes
agent may request: ready_for_verification
agent may not assert: verified, awaiting_manual_verification, failed, skipped
The worker resolves commands from the coordinator, never from the request. They
run as argv through exec with Dir set to the worktree, the quality gate's
existing envelope, with no shell.
registry.Project.Verification matches argv positionally. An absent policy
refuses everything: a plan command is agent-authored and must not inherit the
operator-authored gate's reach.
PlanPhaseRecord binds plan_ref and at_sha. Progress earned under plan A
cannot survive into plan B. A verified phase whose commit has moved renders as
stale rather than as a claim about the current tree.
Manual steps hold a phase at awaiting_manual_verification. The sign-off is an
ordinary human decision whose subject is
plan_phase_verification:<plan_ref>:<phase_id>, so a later "looks good" on an
unrelated thread cannot satisfy it.
c76112a, typed mismatch
The report carries an observation and never a replacement plan.
requested_action is advisory.
The worker fills plan_ref and at_sha from what it can verify, so a stale
report is refused rather than replayed.
The distinction that matters in the reducer:
mismatch recorded != plan superseded
A plan stops being accepted only when a replacement is sealed. An abandoned
replan leaves the accepted plan and its progress intact. On a real re-seal the
old ref moves to PlanHistory and its verification stays in the log as
provenance.
human_decision never reopens. It blocks with a packet, and a human answer can
resolve the contradiction without resealing anything.
CanReopenPhase is separate from CanTransitionPhase. Every path validating an
agent request uses the latter, so phase-request.json still refuses a backward
move.
What the plan machinery needs before it can run
Configuration, in two places. The dual-config trap from run 6 applies here too. It is worse for verification, because the worker is the half that executes.
The coordinator's config.jsonc on homesrv
(~/docker-apps/orchestra-web-ui/config.jsonc) needs a verification block on
test-e2e:
"verification": { "allowed": [["bash", "-n", "*"], ["bash", "scripts/orchestra_e2e_healthcheck.sh"]] }
deploy/config.example.jsonc documents the matcher rules. Absent policy refuses
every plan command, so without this a plan seals and then no phase can ever
verify.
Operational invariant worth recording: the worker-side project policy is
authoritative for execution, because the worker runs the verification. The
coordinator's copy is what PlanPhaseCommands authorises against. Today the
policy lives on registry.Project, which is the coordinator's copy only. The
worker's /etc/orchestra/worker-projects.json carries quality_gate and no
verification field. It does not need one: the worker asks the coordinator which
commands to run. Do not "fix" that by duplicating the policy.
Both halves need rebuilding and redeploying together. The four commits touch
internal/domain, internal/store, internal/agentctx, internal/operations,
internal/registry, internal/federation, internal/workphase, and both
binaries. A split deployment would have a worker calling
/v1/federation/workers/<id>/plan-phase against a coordinator that returns 404.
The burn-in this needs
The proof ladder from PLAN-SPEC-DESIGN.md section 7, unchanged:
- Plan seals three detailed phases. Implement launch contains the plan
verbatim. Grep the rendered
launch.mdfor the phase-three verification command. It is there or it is not. - Phase 1 requests verification. Worker runs the exact plan command. Durable state says phase 1 verified.
- Rotation. Successor receives the complete original plan, phase 1 verified, phase 2 explicitly current.
- Human correction lands above the plan.
- Phase 2 hits a contradiction. Mismatch recorded, no improvisation, old plan retained until the revised one seals, fresh launch carries the revised plan in full.
Steps 1 and 3 are the rungs that matter for smaller local models.
Two more worth adding, from what this session found:
- A plan whose commands fall outside project policy. The refusal must reach the planner at seal time, not the implementer at verification time.
- A legacy plan. The implement context must say phase progress is unavailable and the task must still complete.
Things that will bite
- Verify the running worker binary, not the installed one.
/proc/<pid>/exe. - A documented endpoint is not a live call path. Grep for the caller. F43, F45, and the two before them were all this.
- Gitea returns 422 for a review on your own pull request. Orchestra opens
the pull request as the same forge user the operator reviews as. That makes
REQUEST_CHANGESunreachable in this deployment. The comment path works and is what run 6 exercised. A separate bot account forORCHESTRA_GITEA_TOKENwould restore it. quality_gatelives in two files.config.jsoncon the coordinator and/etc/orchestra/worker-projects.jsonon the worker.finalizereads the worker copy. Verification policy is coordinator-only, deliberately.- Task
06G4E83E4KRXM8DS90M2648MGMis blocked, deliberately. It is the task F44 spawned from Orchestra's own pull request. Its worktree and branch are cleaned up. The task is left blocked onoperator_requiredbecause it is not real work. Fail it or leave it. rmandcpare interactive in this shell. A plainrmprompts, appears to succeed, and leaves the file. Use/bin/rm -fandinstall.- Another session owns 13 uncommitted paths, including
AUDIT.md,deploy/build.sh,web/andAGENTS.md. Commit by path. Nevergit add -Awithout a pathspec. - Secrets are guarded. Reading
.envorprintenvof a token through the shell is blocked. The working recipe is expanding the token inside the container in a single remote command. /tmpis tmpfs with a 10-day sweep. Check/tmp/test-e2eand/tmp/test-e2e-worktreesexist before every run. Eight worktrees are there now, most from completed tasks.- Rebuild the coordinator from a detached worktree, never
docker compose build. The recipe is inHANDOFF-2026-08-28-burnin-5.mdand was used four times today without trouble.
What I would do next, in order
- Fix F18. Three more instances landed today: every refusal in the new plan
machinery records into the same single
last_errorslot. The bounded observation ring is overdue. The plan burn-in will generate exactly the kind of five-second retry loop it exists to make visible. - Configure
test-e2everification policy in the coordinator config. - Rebuild and redeploy both halves at HEAD, together.
- Run the seven-step ladder above. Freeze the result as
orchestra-plan-v1if it passes. - Only then consider the Gitea bot account and the duplicated
quality_gateconfiguration. Both are real, neither blocks anything.
Do step 1 first. Steps 3 and 4 will produce failures whose only trace is otherwise the journal.