diff --git a/BURNIN.md b/BURNIN.md index bef04eb..1eee101 100644 --- a/BURNIN.md +++ b/BURNIN.md @@ -1093,3 +1093,96 @@ F25 fixed, tests only F26 fixed, tests only F27 fixed, tests only ``` + +### F28, deployment: the execution runtime shares the worker's cgroup + +Worker restart destroys agent sessions because the tmux execution runtime +shares the worker service cgroup. tmux must be independently lifecycle-managed. + +The worker spawns the server implicitly on its first tmux command, so the +server and every pane land in `orchestra-worker.service`: + +```text +orchestra-worker.service +├── orchestra-worker +└── tmux server + └── claude panes +``` + +`KillMode` does not fix this. Under `mixed` systemd still sends the final +SIGKILL to whatever remains in the cgroup once the main process exits, and +`process` only encodes accidental orphaning, which systemd itself discourages. +An earlier draft of this entry proposed `mixed` and was wrong. + +The split: + +```text +orchestra-worker.service +└── orchestra-worker + +orchestra-tmux.service +└── tmux -L orchestra + └── claude sessions +``` + +`deploy/orchestra-tmux.service` is new. `deploy/orchestra-worker.service` gains +`After=`/`Wants=` on it, ordering only: a worker that finds the runtime missing +must report that rather than be stopped by it. The unit carries an idle +`orchestra-runtime` session so the server outlives its last agent pane. + +Probed live on a throwaway socket: `tmux -L new-session -d` leaves the +server running under its own pid after the client exits, so `Type=forking` +resolves a main pid. `systemd-analyze verify` passes. + +`User` must match between the two units, because the socket lives under +`/tmp/tmux-$UID`. The installed worker unit on workpc runs as `kami` while +`deploy/orchestra-worker.service` still says `orchestra`. The staged copy at +`~/orchestra-deploy/orchestra-tmux.service` is set to `kami` to match reality. + +### What F28 retroactively explains + +- Run 3's `no server running` at 11:02:16 followed the 10:45:26 worker restart. +- The same at 13:06:22. +- Run 4's refusal could not be delivered at 13:32:16, seconds after the + 13:31:51 restart. +- F16's missing-pane branch has been firing on deployment, not on real + execution-runtime loss. Its proof stands as written, but the trigger was + self-inflicted. + +### F28 proof, to run after the split + +```text +1. start task and obtain live pane P +2. record pane/session identity +3. systemctl restart orchestra-worker +4. assert pane P still exists with identical agent identity +5. new worker starts +6. worker reconciles lease + pane P +7. no TaskReleased / new lease caused merely by worker restart +8. agent continues under the same lease epoch +``` + +Then separately, to restore the meaning of F16's missing-pane branch: + +```text +restart orchestra-tmux +→ pane disappears +→ F16 refuses renewal +→ lease expires and requeues +``` + +### Run 4 diagnostic continuation, partial + +The chain ran and failed only at the last step, at 13:32:16: + +``` +deliver phase refusal 06G46P6KE25Y04VVF7VRZMHZ78: tmux send-keys ... -l -- Orchestra refused your phase request: phase request refused: task 06G46P6KE25Y04VVF7VRZMHZ78 may only move to "research", not "implement" +...: no server running on /tmp/tmux-1000/orchestra: exit status 1 +``` + +Live-proven by that one line: F25, the boundary ran on claude at all; F21's +reading half; F26's premise, the coordinator naming the one valid target; F27, +the 409 classified and routed to `sendPrompt`. The request file survived the +failed send, which is the correct branch. + +Not proven: the receipt itself, because F28 had already destroyed the pane. diff --git a/deploy/orchestra-tmux.service b/deploy/orchestra-tmux.service new file mode 100644 index 0000000..5b2afd8 --- /dev/null +++ b/deploy/orchestra-tmux.service @@ -0,0 +1,42 @@ +# The tmux execution runtime, owned separately from the worker. +# +# The worker used to spawn this server implicitly on its first tmux command, +# which put the server and every agent pane inside the worker unit's cgroup. +# Restarting the worker then destroyed the sessions it was restarting to +# manage: a deploy killed the run it was deploying for, and F16's missing-pane +# branch fired on deployment rather than on real execution-runtime loss. +# +# KillMode does not fix that. Under `mixed` systemd still sends the final +# SIGKILL to whatever remains in the cgroup, and `process` only encodes +# accidental orphaning. The runtime has to be a service of its own. +# +# Lifecycle this gives you: +# systemctl restart orchestra-worker -> panes survive, worker reconciles +# systemctl stop orchestra-tmux -> panes die, deliberately +# +# User MUST match orchestra-worker.service. The socket lives under +# /tmp/tmux-$UID, so a mismatch leaves the worker talking to a different +# server, or to none. +[Unit] +Description=Orchestra tmux execution runtime +After=network.target + +[Service] +Type=forking +User=orchestra +# The socket name is the worker's tmux_socket for this harness (-L, not -S). +# Keep the two in step; the worker does not create the server any more. +# +# The idle session exists so the server outlives its last agent pane. Without +# it the runtime exits whenever a task completes, and the next launch spawns a +# fresh server back inside whichever cgroup asked for it. +ExecStart=/usr/bin/tmux -L orchestra new-session -d -s orchestra-runtime +ExecStop=/usr/bin/tmux -L orchestra kill-server +# The runtime coming back empty is honest: the panes are gone either way, and +# F16 observes that as real execution loss. Refusing to restart would instead +# block every later launch until an operator noticed. +Restart=on-failure +RestartSec=5 + +[Install] +WantedBy=multi-user.target diff --git a/deploy/orchestra-worker.service b/deploy/orchestra-worker.service index a137adc..7862390 100644 --- a/deploy/orchestra-worker.service +++ b/deploy/orchestra-worker.service @@ -2,6 +2,11 @@ Description=Orchestra federation worker After=network-online.target Wants=network-online.target +# Ordering only, deliberately not Requires. The execution runtime owns the +# agent panes and must outlive a worker restart, and a worker that finds it +# missing has to report that rather than be stopped by it. +After=orchestra-tmux.service +Wants=orchestra-tmux.service [Service] Type=simple