# Session post-mortem: 954da1a9 (freestyle web-ui run) Date: 2026-08-11 Branch: `master` HEAD at analysis: `700f59ef` Session: `954da1a9-56cb-48da-a7dc-3ef472c42bf6`, 11:52:14 to 13:31:10 UTC Method: full event-log replay from `~/.config/correx/correx.db` plus CAS artifact reads (`scripts/artread.py`). No inference was re-run. Every claim below cites a session sequence number, an artifact hash, or a source line. ## Executive summary The run did not fail at the end. It failed at minute 7. It then passed every gate for 92 minutes and died on a build gate that was never its job. | | | |---|---| | Wall clock | 98.9 min | | Inference | 73.8 min, 274 rounds, p50 12.9 s, max 190 s | | Blocked on approvals | 19.0 min, 94 pauses, 100% approved, 0 steering | | Tool calls | 258, of which **110 were waste** (43%) | | Files written | 25 `FileWritten` events, **9 distinct files** | | UI views delivered | 0 of 8 requested | `postcss.config.js` was written 5 times. `QueryClientProvider.tsx` was written 8 times. Eight findings. Seven are open. One (§1) was closed by #699 after the run ended. Vikunja: #705, #706, #707, #708, #709, #710, #711, #712. ## Timeline | seq | time | event | |---|---|---| | 2 | 11:52 | `InitialIntent` — eight-view web UI, React/Vite/Tailwind/Ethos/TanStack | | 171 | 11:57 | discovery artifact: `brief.scope` holds all 8 items, `ready: true`, no questions | | 236 | 11:59 | dod artifact: **4 criteria, all `part: "Project Foundation"`** | | 249-256 | 12:02 | execution plan locked; plan-compile, plan-lint (score 0) and grounding all PASS | | 609 | 12:14 | first retry — Tailwind v4 PostCSS plugin moved, `npm run build` exit 1 | | 951 | 12:28 | `WorkspaceVerificationObserved` PROJECT `npm --prefix frontend run build` **passed** | | 1510 | 12:41 | retry — `verbatimModuleSyntax` type-only import | | 1697 | 12:49 | retry — `final_review` PROJECT gate runs **`./gradlew assemble`**, exit 1 | | 1994 | 13:00 | reviewer correctly names `~/.gradle/init.d/offline.gradle` as the cause | | 2306 | 13:12 | reviewer **retracts** the correct diagnosis: "was not found" | | 2374 | 13:14 | `FailureTicketOpened` `stage_loop_break`, 6 identical failures, routes to recovery | | 2584 | 13:23 | `RefinementIteration` recovery→final_review, back into the same wall | | 2723-2731 | 13:30 | 4 retries in 40 s on `No provider satisfies capabilities [ToolCalling]` | | 2733 | 13:31 | `WorkflowFailed` | ## 1. Scope collapse at the analyst (closed by #699) Discovery settled 8 scope items (artifact `8053f0d7`): session driver, sessions list, workflow listing, event log viewer, artifacts/timeline, ideas board, profiles, configuration. The analyst emitted 4 criteria (artifact `5d46f805`), every one `part: "Project Foundation"`. Its summary reads *"Initialize the Correx web-ui project with the required tech stack."* The architect planned against the shrunken DoD (`13d08806`): `scaffold_frontend`, `install_styling_and_ethos`, `configure_tanstack_query`, `final_review`. Its `goal` field never mentions a view. Nothing downstream sees discovery again. Implementer stages receive `neededArtifact: dod` and nothing else, so the loss is total and silent from seq 236 onward. All three plan gates passed, because all three grade structure and the structure was sound. Closed by `ScopeCoverage` (#699, commit `700f59ef`, 19:56 the same day). Stated ceiling holds: it catches a dropped index, not a weak criterion. ## 2. The build gate ran the wrong toolchain (#705, open) [`SessionOrchestratorGates2.kt:156-178`](../../core/kernel/src/main/kotlin/com/correx/core/kernel/orchestration/SessionOrchestratorGates2.kt#L156-L178): - `sessionProducedBuildTarget(sessionId)` is **session-scoped** and turns the gate on. - `stageProducedToolchain(sessionId, stageId)` is **stage-scoped** and returns null when the stage wrote no files. - `commandFor(profileCommands, null)` then falls back to the flat `build` alias. `final_review` is a reviewer. Its tools are `[file_read]` and the plan set `build_expectation: "none"`. It wrote nothing, so the toolchain lookup returned null and the gate ran `./gradlew assemble` against a run whose every write was `frontend/**`. The correct command, `node.build` = `npm --prefix frontend run build`, had already passed at seq 951, 1286 and 1683. The failure it produced was environmental: ``` > Could not resolve io.ktor:ktor-client-websockets:3.0.3. > No cached version ... available for offline mode. ``` That is `~/.gradle/init.d/offline.gradle` on workpc, outside the workspace jail, predating the session. Cost: 4 retries plus a recovery detour, seq 1694 to 2733. That is **32.5 min, 33% of wall clock**, and the run's death. Fix: when the stage produced no toolchain, fall back to the session's toolchain (the same manifest `sessionProducedBuildTarget` already reads) before the flat alias. Two lookups deciding one command must share a scope. ## 3. The harness destroyed a correct diagnosis (#707, open) Only 4 of 274 inference rounds returned any text at all (§5). Three of them are this finding. - **seq 1994**: *"The build failed because the Gradle environment is in offline mode... The `org.gradle.offline=true` flag is present in `~/.gradle/init.d/offline.gradle`."* Correct root cause, reached in about ten minutes. - **seq 2306**: *"Since `~/.gradle/init.d/offline.gradle` was not found, I will check for global Gradle properties..."* The model **retracted the correct answer**, because at seq 1793 `file_read ~/.gradle/gradle.properties` returned `[REFERENCE_EXISTS] ... does not exist. Do NOT keep retrying`. The tool layer does not expand `~`, and the path is outside the jail regardless. The guard reported a true thing as false. - **seq 2715**: *"the DoD criteria do not require a successful backend build... the frontend build passed... I cannot fix the backend dependencies, it's out of scope. However, since the PROJECT build gate is a hard gate, I must reject it."* Right three times, and wrong-footed each time by the harness. The structural gap: every verdict the reviewer can emit (`approved`, `changes_requested`, `rejected`) routes back into the agent. There is no way to say *this failure is not attributable to this run*. `FailureTicketOpened` fired correctly at seq 2374 after 6 identical failures, routed to `recovery`, and recovery routed straight back (`RefinementIteration`, seq 2584). The retry feedback closed with *"Do not re-discover unrelated files before it builds."* The harness forbade the only move that would have worked. ## 4. Ten-turn amnesia inside a mostly empty window (#706, open) 176 `ContextTruncated` events, every one L2, `entriesDropped` climbing to 40+ and pinning. | stage | truncations | max dropped | |---|---|---| | discovery | 14 | 28 | | scaffold_frontend | 40 | 40 | | install_styling_and_ethos | 20 | 40 | | configure_tanstack_query | 21 | 16 | | final_review | 62 | 42 | | recovery | 16 | 32 | [`DefaultContextPackBuilder.kt:380`](../../core/context/src/main/kotlin/com/correx/core/context/builder/DefaultContextPackBuilder.kt#L380) uses `CompressionStrategy.Conversation()`. The default is `keepLast = 10` ([`CompressionStrategy.kt:7`](../../core/context/src/main/kotlin/com/correx/core/context/compression/CompressionStrategy.kt#L7)). L2 holds **5 tool-call/result pairs**. The assembled pack at round 31 of `scaffold_frontend` (seq 599) confirms it. It carries 4 L0 entries, 2 L1, 3 L3, and 10 L2, exactly the last ten. **The window is not full.** Those same packs run 1.7k to 13k tokens, median 5.3k. The cap is entry count, not tokens, so the model is starved at 5k while a local model's context sits mostly idle. Consequence, measured: **76 byte-identical repeat calls**, 29% of all tool invocations. | repeats | stage | call | |---|---|---| | 9 | final_review | `list_dir {"path":"frontend"}` | | 7 | scaffold_frontend | `list_dir {"path":"frontend/"}` | | 6 | final_review | `file_read {"path":"gradle.properties"}` | | 5 | scaffold_frontend | `file_read {"path":"frontend/package.json"}` | | 4 | final_review | `shell ./gradlew assemble --no-configuration-cache` | The `decisionJournal` is the only memory surviving eviction. Its complete content at the end of the run: ``` - Goal: Write a web-ui for Correx. ... - scaffold_frontend: 1 retry, resolved — stage completed - configure_tanstack_query: 1 retry, resolved — stage completed - final_review: 3 retries, resolved — stage completed ``` Bookkeeping, not knowledge. Preferred fix is not a larger `keepLast` (it costs local-inference latency and still forgets). Add a deterministic **action ledger**. One line per tool call for the stage: tool, target, exit, short result hash. Never evicted. 258 calls at roughly 12 tokens is about 3k tokens for a whole run. That is cheaper than the duplicates it removes, and it carries the "already tried, same failure" signal no window size provides. Related: the plan set `kind: "process_result"` on every stage, which silently disables the `file_written` manifest injection ([`SessionOrchestratorContext.kt:458`](../../core/kernel/src/main/kotlin/com/correx/core/kernel/orchestration/SessionOrchestratorContext.kt#L458) builds it only for `file_written` slots). A feature built to prevent this thrash was switched off by one planner field. ## 5. There was no chain of thought **270 of 274 inference rounds returned an empty response body** (blake3 `af1349b9...`, the empty string). Per stage: discovery 25/25, analyst 9/9, architect 1/1, scaffold_frontend 62/62, install_styling 31/31, configure_tanstack 32/33, final_review 83/86, recovery 27/27. Every assistant message in the assembled prompts is a bare function call with no content. The system prompt instructs *"make one coherent step per response."* The loop reduces to this. Context in, one tool call out, no reasoning, no plan, no self-check, with a 10-turn memory (§4) at temperature 1.0 (§6). The 29% repeat rate is what that combination produces. Surfacing reasoning is already tracked as #298. This finding is that none was produced to surface. ## 6. Temperature 1.0 on structured arguments (#710, open) [`Main.kt:435-451`](../../apps/server/src/main/kotlin/com/correx/apps/server/Main.kt#L435-L451) sends `[sampling] temperature = 1.0, top_p = 0.95, top_k = 64` on **every** stage inference. Chat-quality settings applied to fields where exactly one string is correct. | seq | emitted | intended | |---|---|---| | 277 | `create_vite@latest` | `create vite@latest` | | 465 | `npm_prefix=frontend` | `npm --prefix frontend` | | 1302 | `npm_install_package` | `npm install ` | | 2060 | `./gradlew_` | `./gradlew` | | 2708 | `gradlew` | `./gradlew` | Space-to-underscore substitution on an argv field is sampling noise. correx already sends a per-request `GenerationConfig`, so the override is per-call. Keep 1.0 for artifact and review prose. Go near-greedy when the response is constrained to tool calls. #303 repairs these after the fact. Not emitting them is cheaper. **Same task, second half: files are copied through the token stream.** There is no `file_copy` tool. `ethos-icons.svg` moved by a `file_read` of 6124 characters, then a `file_write` of 6845 characters of tool-call arguments. A `READ_BEFORE_WRITE` rejection round came first (seq 540), when the model tried to write before reading. About 6 rounds and 13k tokens for two static assets, at temperature 1.0, with silent corruption possible. The project profile says *"COPY design-system/ethos.tokens.css and ethos-icons.svg into the project and import them, do NOT read their contents."* The toolset offers no way to obey that instruction. ## 7. The plan contradicts itself, and the kernel sides against the model (#711, open) The architect emits a stage's `prompt` and its `writes` manifest independently. They disagreed twice, and the kernel enforces `writes`. - `install_styling_and_ethos` prompt step 4: *"Copy ... into `frontend/src/assets/ethos/`."* Manifest: `[tailwind.config.js, postcss.config.js, src/index.css]`. Result at seq 1086: `[PATH_OUTSIDE_MANIFEST]`. - `configure_tanstack_query` prompt step 2: *"Create a `QueryClient` instance and a wrapper component."* Manifest: `[frontend/src/lib/query-client.ts]`. Blocked at seq 1321. `QueryClientProvider.tsx` was then written 8 times. Both fields are already parsed by the plan-compile gate, so this is a string check over paths the prompt itself spells out. ## 8. The stage prompt never updates (#709, open) Verified against the assembled prompt for round 31 of `scaffold_frontend` (artifact `47ebf7b0`). The system message still read: > Use the `shell` tool to initialize a new React + TypeScript project using Vite in the > `frontend/` directory. Run `npm create vite@latest frontend -- --template react-ts` ... The project had existed for 30 rounds. The instruction is a fixed string from the execution plan, and nothing recomputes it against what exists. Combined with §4, the model's strongest signal every round is an order to redo step one. The plan's stage prompts are literal command transcripts written from training memory, and both were wrong: - `npx tailwindcss init -p`. `npx` is **blocked by tool policy**, rejected at seq 329 and 988. The policy message helpfully names the alternative. A plan-compile lint rejecting a plan that names a forbidden executable is roughly ten lines. - `npm install -D tailwindcss postcss autoprefixer` plus `tailwind.config.js` is the **Tailwind v3** procedure. `npm install tailwindcss` now resolves v4, where the PostCSS plugin moved to `@tailwindcss/postcss`. That is the build failure at seq 609 verbatim, and 22 rounds of thrash followed it. ## 9. Ninety-four approvals that carried no information (#712, open) Every `ApprovalDecisionResolved` in the run is `APPROVED` with `reason: null` and `userSteering: null`. Pause durations: p50 13.6 s, p90 19.0 s, max 20.3 s. Total **19.0 min, 19% of wall clock**. By tool: `shell` T2 x63, `file_write` T2 x13, `file_edit` T3 x10, remainder task and delete calls. Two problems. The gate is priced per tool call, so it charges 94 interrupts for one operator intention. The attention also lands in the wrong place. All 94 chances to intervene were spent on "may I run `npm install`". The decision that determined the outcome (seq 236, §1) had no operator checkpoint at all. Auto-approving reads and writes inside the stage's declared `writes` manifest costs nothing in safety, because `PATH_OUTSIDE_MANIFEST` already enforces that boundary. It returns most of the 19 minutes. ## Waste breakdown 258 tool calls. 110 wasted, counting each call once (repeat, failed, or policy-rejected). | stage | waste / total | | |---|---|---| | discovery | 2 / 24 | 8% | | analyst | 1 / 8 | 12% | | scaffold_frontend | 26 / 60 | 43% | | install_styling_and_ethos | 11 / 30 | 37% | | configure_tanstack_query | 9 / 31 | 29% | | **final_review** | **55 / 79** | **70%** | | recovery | 6 / 26 | 23% | Context composition across all 279 assemblies: `toolResult` 47.9%, `projectProfile` 8.4%, `assistantToolCall` 8.1%, `decisionJournal` 6.8%, `retryFeedback` 5.4%. By layer: L2 56.0%, L0 18.0%, L1 13.2%, L3 12.9%. ## Ranked remediation inference, from this run's own timings: 1. **Baseline every gate command at session start (#708).** One build at t=0. A command already red before the agent touched anything reports `pre-existing` and never fails a stage. Recovers `final_review` plus most of `recovery`. **~40 min.** 2. **Fix the toolchain scope mismatch (#705).** One line. Prevents the same 40 minutes by an independent route. Do both. 3. **Deterministic action ledger (#706).** ~16 min of duplicate calls, plus the loops they feed. Raising `keepLast` alone does not fix it. 4. **Auto-approve inside the declared manifest (#712).** **~17 min.** 5. **Split the sampling config (#710).** Near-greedy for tool-call rounds. Items 1, 3 and 4 account for roughly 70 of the 99 minutes. ## The structural point Every gate in this run was local. Each asked whether a stage did its stage correctly. None asked whether the run was still building the requested thing, until `final_review` at minute 90 graded against the already-collapsed DoD. The path from intent to work is five lossy model summarisations: intent, discovery, DoD, plan, stage prompt. Only the last link was ever checked. #699 now checks the DoD-to-discovery link, which is the one that broke here. The general shape remains. The cheap version needs no model. The intent names eight views. After three of four stages, `FileWritten` holds nine files and none is a view. That comparison is a set difference over recorded events and costs nothing. It would have fired at minute 25, not failed silently at minute 99. The deeper item, for #168: this run produced a genuinely valuable artifact and discarded it. A local model correctly diagnosed a Gradle offline-mode misconfiguration in ten minutes from build output alone. That knowledge existed at seq 1994 and was gone by seq 2306. A guard reporting a real file as missing erased it. The event log still has it. Nothing reads it back. ## Verification and scope - fact: every number above is derived from `events` rows for session `954da1a9` and from CAS artifacts resolved through `~/.config/correx/artifacts/index.sqlite`. No inference re-run. - fact: `ScopeCoverage` (#699) landed at 19:56 on 2026-08-11, after this session ended at 13:31. It was not active during the run. - inference: the `./gradlew assemble` failure is attributed to `~/.gradle/init.d/offline.gradle` on workpc. The build output names offline mode and an uncached `io.ktor:ktor-client-websockets`. The file itself was not read during this analysis. - unknown: whether the 94 approvals were resolved by a human operator or by an auto-approver. The tight 13-20 s clustering suggests a poll interval. No event records the decider. - This post-mortem implemented no fixes. It filed Vikunja #705 through #712.