The run died on a build gate running the wrong toolchain and burned 43% of its tool calls on repeats, rejections and failures. Five fixes, each traceable to a measured cost in docs/audits/2026-08-11-session-954da1a9-postmortem.md. #705 build gate toolchain scope. The gate is armed session-scoped but read the toolchain stage-scoped, so a reviewer stage that wrote nothing fell back to the flat `build` alias and ran ./gradlew assemble on an all-frontend session. It now falls back to the session's own manifest first. 32.5 min, 33% of that run. #706 action ledger. L2 keeps ten conversation entries, so a stage past round five has no memory of what it tried; 29% of tool calls were byte-identical repeats. One pinned line per call (tool, target, outcome) with repeats collapsed to a count, ~3k tokens for a whole run. #709 plan-compile lint for blocked runners. Plans prescribed `npx tailwindcss init -p`, rejected by the shell denylist at every attempt. Rejected at compile time now, where the architect can still rewrite the step. #710 near-greedy sampling on tool-call rounds. temperature 1.0 on argv emitted `./gradlew_`, `npm_prefix=frontend`, `create_vite@latest`. Prose rounds keep the operator's sampling. #712 auto-approve manifest-contained writes. 94 approvals, all APPROVED, no steering, 19 min. A write inside the declared manifest already proved its containment by getting past ManifestContainmentRule. DENY mode still denies. Tests: core:kernel 133, infrastructure:workflow 99, testing:integration 176, testing:deterministic 79, all green. detekt clean (no new findings). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013dVqqci5H5b3s6xzv6Lojq
18 KiB
Session post-mortem: 954da1a9 (freestyle web-ui run)
Date: 2026-08-11
Branch: master
HEAD at analysis: 700f59ef
Session: 954da1a9-56cb-48da-a7dc-3ef472c42bf6, 11:52:14 to 13:31:10 UTC
Method: full event-log replay from ~/.config/correx/correx.db plus CAS artifact reads
(scripts/artread.py). No inference was re-run. Every claim below cites a session sequence
number, an artifact hash, or a source line.
Executive summary
The run did not fail at the end. It failed at minute 7. It then passed every gate for 92 minutes and died on a build gate that was never its job.
| Wall clock | 98.9 min |
| Inference | 73.8 min, 274 rounds, p50 12.9 s, max 190 s |
| Blocked on approvals | 19.0 min, 94 pauses, 100% approved, 0 steering |
| Tool calls | 258, of which 110 were waste (43%) |
| Files written | 25 FileWritten events, 9 distinct files |
| UI views delivered | 0 of 8 requested |
postcss.config.js was written 5 times. QueryClientProvider.tsx was written 8 times.
Eight findings. Seven are open. One (§1) was closed by #699 after the run ended. Vikunja: #705, #706, #707, #708, #709, #710, #711, #712.
Timeline
| seq | time | event |
|---|---|---|
| 2 | 11:52 | InitialIntent — eight-view web UI, React/Vite/Tailwind/Ethos/TanStack |
| 171 | 11:57 | discovery artifact: brief.scope holds all 8 items, ready: true, no questions |
| 236 | 11:59 | dod artifact: 4 criteria, all part: "Project Foundation" |
| 249-256 | 12:02 | execution plan locked; plan-compile, plan-lint (score 0) and grounding all PASS |
| 609 | 12:14 | first retry — Tailwind v4 PostCSS plugin moved, npm run build exit 1 |
| 951 | 12:28 | WorkspaceVerificationObserved PROJECT npm --prefix frontend run build passed |
| 1510 | 12:41 | retry — verbatimModuleSyntax type-only import |
| 1697 | 12:49 | retry — final_review PROJECT gate runs ./gradlew assemble, exit 1 |
| 1994 | 13:00 | reviewer correctly names ~/.gradle/init.d/offline.gradle as the cause |
| 2306 | 13:12 | reviewer retracts the correct diagnosis: "was not found" |
| 2374 | 13:14 | FailureTicketOpened stage_loop_break, 6 identical failures, routes to recovery |
| 2584 | 13:23 | RefinementIteration recovery→final_review, back into the same wall |
| 2723-2731 | 13:30 | 4 retries in 40 s on No provider satisfies capabilities [ToolCalling] |
| 2733 | 13:31 | WorkflowFailed |
1. Scope collapse at the analyst (closed by #699)
Discovery settled 8 scope items (artifact 8053f0d7): session driver, sessions list,
workflow listing, event log viewer, artifacts/timeline, ideas board, profiles, configuration.
The analyst emitted 4 criteria (artifact 5d46f805), every one part: "Project Foundation".
Its summary reads "Initialize the Correx web-ui project with the required tech stack." The
architect planned against the shrunken DoD (13d08806): scaffold_frontend,
install_styling_and_ethos, configure_tanstack_query, final_review. Its goal field
never mentions a view.
Nothing downstream sees discovery again. Implementer stages receive neededArtifact: dod and
nothing else, so the loss is total and silent from seq 236 onward.
All three plan gates passed, because all three grade structure and the structure was sound.
Closed by ScopeCoverage (#699, commit 700f59ef, 19:56 the same day). Stated ceiling holds:
it catches a dropped index, not a weak criterion.
2. The build gate ran the wrong toolchain (#705, open)
SessionOrchestratorGates2.kt:156-178:
sessionProducedBuildTarget(sessionId)is session-scoped and turns the gate on.stageProducedToolchain(sessionId, stageId)is stage-scoped and returns null when the stage wrote no files.commandFor(profileCommands, null)then falls back to the flatbuildalias.
final_review is a reviewer. Its tools are [file_read] and the plan set
build_expectation: "none". It wrote nothing, so the toolchain lookup returned null and the
gate ran ./gradlew assemble against a run whose every write was frontend/**. The correct
command, node.build = npm --prefix frontend run build, had already passed at seq 951, 1286
and 1683.
The failure it produced was environmental:
> Could not resolve io.ktor:ktor-client-websockets:3.0.3.
> No cached version ... available for offline mode.
That is ~/.gradle/init.d/offline.gradle on workpc, outside the workspace jail, predating the
session.
Cost: 4 retries plus a recovery detour, seq 1694 to 2733. That is 32.5 min, 33% of wall clock, and the run's death.
Fix: when the stage produced no toolchain, fall back to the session's toolchain (the same
manifest sessionProducedBuildTarget already reads) before the flat alias. Two lookups
deciding one command must share a scope.
3. The harness destroyed a correct diagnosis (#707, open)
Only 4 of 274 inference rounds returned any text at all (§5). Three of them are this finding.
- seq 1994: "The build failed because the Gradle environment is in offline mode... The
org.gradle.offline=trueflag is present in~/.gradle/init.d/offline.gradle." Correct root cause, reached in about ten minutes. - seq 2306: "Since
~/.gradle/init.d/offline.gradlewas not found, I will check for global Gradle properties..." The model retracted the correct answer, because at seq 1793file_read ~/.gradle/gradle.propertiesreturned[REFERENCE_EXISTS] ... does not exist. Do NOT keep retrying. The tool layer does not expand~, and the path is outside the jail regardless. The guard reported a true thing as false. - seq 2715: "the DoD criteria do not require a successful backend build... the frontend build passed... I cannot fix the backend dependencies, it's out of scope. However, since the PROJECT build gate is a hard gate, I must reject it."
Right three times, and wrong-footed each time by the harness.
The structural gap: every verdict the reviewer can emit (approved, changes_requested,
rejected) routes back into the agent. There is no way to say this failure is not
attributable to this run. FailureTicketOpened fired correctly at seq 2374 after 6 identical
failures, routed to recovery, and recovery routed straight back (RefinementIteration,
seq 2584).
The retry feedback closed with "Do not re-discover unrelated files before it builds." The harness forbade the only move that would have worked.
4. Ten-turn amnesia inside a mostly empty window (#706, open)
176 ContextTruncated events, every one L2, entriesDropped climbing to 40+ and pinning.
| stage | truncations | max dropped |
|---|---|---|
| discovery | 14 | 28 |
| scaffold_frontend | 40 | 40 |
| install_styling_and_ethos | 20 | 40 |
| configure_tanstack_query | 21 | 16 |
| final_review | 62 | 42 |
| recovery | 16 | 32 |
DefaultContextPackBuilder.kt:380
uses CompressionStrategy.Conversation(). The default is keepLast = 10
(CompressionStrategy.kt:7).
L2 holds 5 tool-call/result pairs. The assembled pack at round 31 of scaffold_frontend
(seq 599) confirms it. It carries 4 L0 entries, 2 L1, 3 L3, and 10 L2, exactly the last ten.
The window is not full. Those same packs run 1.7k to 13k tokens, median 5.3k. The cap is entry count, not tokens, so the model is starved at 5k while a local model's context sits mostly idle.
Consequence, measured: 76 byte-identical repeat calls, 29% of all tool invocations.
| repeats | stage | call |
|---|---|---|
| 9 | final_review | list_dir {"path":"frontend"} |
| 7 | scaffold_frontend | list_dir {"path":"frontend/"} |
| 6 | final_review | file_read {"path":"gradle.properties"} |
| 5 | scaffold_frontend | file_read {"path":"frontend/package.json"} |
| 4 | final_review | shell ./gradlew assemble --no-configuration-cache |
The decisionJournal is the only memory surviving eviction. Its complete content at the end of
the run:
- Goal: Write a web-ui for Correx. ...
- scaffold_frontend: 1 retry, resolved — stage completed
- configure_tanstack_query: 1 retry, resolved — stage completed
- final_review: 3 retries, resolved — stage completed
Bookkeeping, not knowledge.
Preferred fix is not a larger keepLast (it costs local-inference latency and still forgets).
Add a deterministic action ledger. One line per tool call for the stage: tool, target,
exit, short result hash. Never evicted. 258 calls at roughly 12 tokens is about 3k tokens for
a whole run. That is cheaper than the duplicates it removes, and it carries the "already
tried, same failure" signal no window size provides.
Related: the plan set kind: "process_result" on every stage, which silently disables the
file_written manifest injection
(SessionOrchestratorContext.kt:458
builds it only for file_written slots). A feature built to prevent this thrash was switched
off by one planner field.
5. There was no chain of thought
270 of 274 inference rounds returned an empty response body (blake3
af1349b9..., the empty string). Per stage: discovery 25/25, analyst 9/9, architect 1/1,
scaffold_frontend 62/62, install_styling 31/31, configure_tanstack 32/33, final_review 83/86,
recovery 27/27.
Every assistant message in the assembled prompts is a bare function call with no content. The system prompt instructs "make one coherent step per response."
The loop reduces to this. Context in, one tool call out, no reasoning, no plan, no self-check, with a 10-turn memory (§4) at temperature 1.0 (§6). The 29% repeat rate is what that combination produces. Surfacing reasoning is already tracked as #298. This finding is that none was produced to surface.
6. Temperature 1.0 on structured arguments (#710, open)
Main.kt:435-451
sends [sampling] temperature = 1.0, top_p = 0.95, top_k = 64 on every stage inference.
Chat-quality settings applied to fields where exactly one string is correct.
| seq | emitted | intended |
|---|---|---|
| 277 | create_vite@latest |
create vite@latest |
| 465 | npm_prefix=frontend |
npm --prefix frontend |
| 1302 | npm_install_package |
npm install <package> |
| 2060 | ./gradlew_ |
./gradlew |
| 2708 | gradlew |
./gradlew |
Space-to-underscore substitution on an argv field is sampling noise. correx already sends a
per-request GenerationConfig, so the override is per-call. Keep 1.0 for artifact and review
prose. Go near-greedy when the response is constrained to tool calls. #303 repairs these after
the fact. Not emitting them is cheaper.
Same task, second half: files are copied through the token stream. There is no file_copy
tool. ethos-icons.svg moved by a file_read of 6124 characters, then a file_write of 6845
characters of tool-call arguments. A READ_BEFORE_WRITE rejection round came first (seq 540),
when the model tried to write before reading. About 6 rounds and 13k tokens for two static
assets, at temperature 1.0, with silent corruption possible. The project profile says "COPY
design-system/ethos.tokens.css and ethos-icons.svg into the project and import them, do NOT
read their contents." The toolset offers no way to obey that instruction.
7. The plan contradicts itself, and the kernel sides against the model (#711, open)
The architect emits a stage's prompt and its writes manifest independently. They disagreed
twice, and the kernel enforces writes.
install_styling_and_ethosprompt step 4: "Copy ... intofrontend/src/assets/ethos/." Manifest:[tailwind.config.js, postcss.config.js, src/index.css]. Result at seq 1086:[PATH_OUTSIDE_MANIFEST].configure_tanstack_queryprompt step 2: "Create aQueryClientinstance and a wrapper component." Manifest:[frontend/src/lib/query-client.ts]. Blocked at seq 1321.QueryClientProvider.tsxwas then written 8 times.
Both fields are already parsed by the plan-compile gate, so this is a string check over paths the prompt itself spells out.
8. The stage prompt never updates (#709, open)
Verified against the assembled prompt for round 31 of scaffold_frontend (artifact
47ebf7b0). The system message still read:
Use the
shelltool to initialize a new React + TypeScript project using Vite in thefrontend/directory. Runnpm create vite@latest frontend -- --template react-ts...
The project had existed for 30 rounds. The instruction is a fixed string from the execution plan, and nothing recomputes it against what exists. Combined with §4, the model's strongest signal every round is an order to redo step one.
The plan's stage prompts are literal command transcripts written from training memory, and both were wrong:
npx tailwindcss init -p.npxis blocked by tool policy, rejected at seq 329 and 988. The policy message helpfully names the alternative. A plan-compile lint rejecting a plan that names a forbidden executable is roughly ten lines.npm install -D tailwindcss postcss autoprefixerplustailwind.config.jsis the Tailwind v3 procedure.npm install tailwindcssnow resolves v4, where the PostCSS plugin moved to@tailwindcss/postcss. That is the build failure at seq 609 verbatim, and 22 rounds of thrash followed it.
9. Ninety-four approvals that carried no information (#712, open)
Every ApprovalDecisionResolved in the run is APPROVED with reason: null and
userSteering: null. Pause durations: p50 13.6 s, p90 19.0 s, max 20.3 s. Total 19.0 min,
19% of wall clock.
By tool: shell T2 x63, file_write T2 x13, file_edit T3 x10, remainder task and delete
calls.
Two problems. The gate is priced per tool call, so it charges 94 interrupts for one operator
intention. The attention also lands in the wrong place. All 94 chances to intervene were
spent on "may I run npm install". The decision that determined the outcome (seq 236, §1)
had no operator checkpoint at all.
Auto-approving reads and writes inside the stage's declared writes manifest costs nothing in
safety, because PATH_OUTSIDE_MANIFEST already enforces that boundary. It returns most of the
19 minutes.
Waste breakdown
258 tool calls. 110 wasted, counting each call once (repeat, failed, or policy-rejected).
| stage | waste / total | |
|---|---|---|
| discovery | 2 / 24 | 8% |
| analyst | 1 / 8 | 12% |
| scaffold_frontend | 26 / 60 | 43% |
| install_styling_and_ethos | 11 / 30 | 37% |
| configure_tanstack_query | 9 / 31 | 29% |
| final_review | 55 / 79 | 70% |
| recovery | 6 / 26 | 23% |
Context composition across all 279 assemblies: toolResult 47.9%, projectProfile 8.4%,
assistantToolCall 8.1%, decisionJournal 6.8%, retryFeedback 5.4%. By layer: L2 56.0%,
L0 18.0%, L1 13.2%, L3 12.9%.
Ranked remediation
inference, from this run's own timings:
- Baseline every gate command at session start (#708). One build at t=0. A command already
red before the agent touched anything reports
pre-existingand never fails a stage. Recoversfinal_reviewplus most ofrecovery. ~40 min. - Fix the toolchain scope mismatch (#705). One line. Prevents the same 40 minutes by an independent route. Do both.
- Deterministic action ledger (#706). ~16 min of duplicate calls, plus the loops they feed.
Raising
keepLastalone does not fix it. - Auto-approve inside the declared manifest (#712). ~17 min.
- Split the sampling config (#710). Near-greedy for tool-call rounds.
Items 1, 3 and 4 account for roughly 70 of the 99 minutes.
The structural point
Every gate in this run was local. Each asked whether a stage did its stage correctly. None
asked whether the run was still building the requested thing, until final_review at minute 90
graded against the already-collapsed DoD.
The path from intent to work is five lossy model summarisations: intent, discovery, DoD, plan, stage prompt. Only the last link was ever checked. #699 now checks the DoD-to-discovery link, which is the one that broke here. The general shape remains.
The cheap version needs no model. The intent names eight views. After three of four stages,
FileWritten holds nine files and none is a view. That comparison is a set difference over
recorded events and costs nothing. It would have fired at minute 25, not failed silently at
minute 99.
The deeper item, for #168: this run produced a genuinely valuable artifact and discarded it. A local model correctly diagnosed a Gradle offline-mode misconfiguration in ten minutes from build output alone. That knowledge existed at seq 1994 and was gone by seq 2306. A guard reporting a real file as missing erased it. The event log still has it. Nothing reads it back.
Verification and scope
- fact: every number above is derived from
eventsrows for session954da1a9and from CAS artifacts resolved through~/.config/correx/artifacts/index.sqlite. No inference re-run. - fact:
ScopeCoverage(#699) landed at 19:56 on 2026-08-11, after this session ended at 13:31. It was not active during the run. - inference: the
./gradlew assemblefailure is attributed to~/.gradle/init.d/offline.gradleon workpc. The build output names offline mode and an uncachedio.ktor:ktor-client-websockets. The file itself was not read during this analysis. - unknown: whether the 94 approvals were resolved by a human operator or by an auto-approver. The tight 13-20 s clustering suggests a poll interval. No event records the decider.
- This post-mortem implemented no fixes. It filed Vikunja #705 through #712.