Files
correx/docs/audits/2026-08-11-session-954da1a9-postmortem.md
T
claude d892587420 fix(kernel,workflow): five fixes from the 954da1a9 post-mortem (#705,#706,#709,#710,#712)
The run died on a build gate running the wrong toolchain and burned 43% of its
tool calls on repeats, rejections and failures. Five fixes, each traceable to a
measured cost in docs/audits/2026-08-11-session-954da1a9-postmortem.md.

#705 build gate toolchain scope. The gate is armed session-scoped but read the
toolchain stage-scoped, so a reviewer stage that wrote nothing fell back to the
flat `build` alias and ran ./gradlew assemble on an all-frontend session. It now
falls back to the session's own manifest first. 32.5 min, 33% of that run.

#706 action ledger. L2 keeps ten conversation entries, so a stage past round
five has no memory of what it tried; 29% of tool calls were byte-identical
repeats. One pinned line per call (tool, target, outcome) with repeats collapsed
to a count, ~3k tokens for a whole run.

#709 plan-compile lint for blocked runners. Plans prescribed `npx tailwindcss
init -p`, rejected by the shell denylist at every attempt. Rejected at compile
time now, where the architect can still rewrite the step.

#710 near-greedy sampling on tool-call rounds. temperature 1.0 on argv emitted
`./gradlew_`, `npm_prefix=frontend`, `create_vite@latest`. Prose rounds keep the
operator's sampling.

#712 auto-approve manifest-contained writes. 94 approvals, all APPROVED, no
steering, 19 min. A write inside the declared manifest already proved its
containment by getting past ManifestContainmentRule. DENY mode still denies.

Tests: core:kernel 133, infrastructure:workflow 99, testing:integration 176,
testing:deterministic 79, all green. detekt clean (no new findings).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013dVqqci5H5b3s6xzv6Lojq
2026-08-11 22:15:49 +04:00

18 KiB

Session post-mortem: 954da1a9 (freestyle web-ui run)

Date: 2026-08-11 Branch: master HEAD at analysis: 700f59ef Session: 954da1a9-56cb-48da-a7dc-3ef472c42bf6, 11:52:14 to 13:31:10 UTC Method: full event-log replay from ~/.config/correx/correx.db plus CAS artifact reads (scripts/artread.py). No inference was re-run. Every claim below cites a session sequence number, an artifact hash, or a source line.

Executive summary

The run did not fail at the end. It failed at minute 7. It then passed every gate for 92 minutes and died on a build gate that was never its job.

Wall clock 98.9 min
Inference 73.8 min, 274 rounds, p50 12.9 s, max 190 s
Blocked on approvals 19.0 min, 94 pauses, 100% approved, 0 steering
Tool calls 258, of which 110 were waste (43%)
Files written 25 FileWritten events, 9 distinct files
UI views delivered 0 of 8 requested

postcss.config.js was written 5 times. QueryClientProvider.tsx was written 8 times.

Eight findings. Seven are open. One (§1) was closed by #699 after the run ended. Vikunja: #705, #706, #707, #708, #709, #710, #711, #712.

Timeline

seq time event
2 11:52 InitialIntent — eight-view web UI, React/Vite/Tailwind/Ethos/TanStack
171 11:57 discovery artifact: brief.scope holds all 8 items, ready: true, no questions
236 11:59 dod artifact: 4 criteria, all part: "Project Foundation"
249-256 12:02 execution plan locked; plan-compile, plan-lint (score 0) and grounding all PASS
609 12:14 first retry — Tailwind v4 PostCSS plugin moved, npm run build exit 1
951 12:28 WorkspaceVerificationObserved PROJECT npm --prefix frontend run build passed
1510 12:41 retry — verbatimModuleSyntax type-only import
1697 12:49 retry — final_review PROJECT gate runs ./gradlew assemble, exit 1
1994 13:00 reviewer correctly names ~/.gradle/init.d/offline.gradle as the cause
2306 13:12 reviewer retracts the correct diagnosis: "was not found"
2374 13:14 FailureTicketOpened stage_loop_break, 6 identical failures, routes to recovery
2584 13:23 RefinementIteration recovery→final_review, back into the same wall
2723-2731 13:30 4 retries in 40 s on No provider satisfies capabilities [ToolCalling]
2733 13:31 WorkflowFailed

1. Scope collapse at the analyst (closed by #699)

Discovery settled 8 scope items (artifact 8053f0d7): session driver, sessions list, workflow listing, event log viewer, artifacts/timeline, ideas board, profiles, configuration.

The analyst emitted 4 criteria (artifact 5d46f805), every one part: "Project Foundation". Its summary reads "Initialize the Correx web-ui project with the required tech stack." The architect planned against the shrunken DoD (13d08806): scaffold_frontend, install_styling_and_ethos, configure_tanstack_query, final_review. Its goal field never mentions a view.

Nothing downstream sees discovery again. Implementer stages receive neededArtifact: dod and nothing else, so the loss is total and silent from seq 236 onward.

All three plan gates passed, because all three grade structure and the structure was sound.

Closed by ScopeCoverage (#699, commit 700f59ef, 19:56 the same day). Stated ceiling holds: it catches a dropped index, not a weak criterion.

2. The build gate ran the wrong toolchain (#705, open)

SessionOrchestratorGates2.kt:156-178:

  • sessionProducedBuildTarget(sessionId) is session-scoped and turns the gate on.
  • stageProducedToolchain(sessionId, stageId) is stage-scoped and returns null when the stage wrote no files.
  • commandFor(profileCommands, null) then falls back to the flat build alias.

final_review is a reviewer. Its tools are [file_read] and the plan set build_expectation: "none". It wrote nothing, so the toolchain lookup returned null and the gate ran ./gradlew assemble against a run whose every write was frontend/**. The correct command, node.build = npm --prefix frontend run build, had already passed at seq 951, 1286 and 1683.

The failure it produced was environmental:

> Could not resolve io.ktor:ktor-client-websockets:3.0.3.
  > No cached version ... available for offline mode.

That is ~/.gradle/init.d/offline.gradle on workpc, outside the workspace jail, predating the session.

Cost: 4 retries plus a recovery detour, seq 1694 to 2733. That is 32.5 min, 33% of wall clock, and the run's death.

Fix: when the stage produced no toolchain, fall back to the session's toolchain (the same manifest sessionProducedBuildTarget already reads) before the flat alias. Two lookups deciding one command must share a scope.

3. The harness destroyed a correct diagnosis (#707, open)

Only 4 of 274 inference rounds returned any text at all (§5). Three of them are this finding.

  • seq 1994: "The build failed because the Gradle environment is in offline mode... The org.gradle.offline=true flag is present in ~/.gradle/init.d/offline.gradle." Correct root cause, reached in about ten minutes.
  • seq 2306: "Since ~/.gradle/init.d/offline.gradle was not found, I will check for global Gradle properties..." The model retracted the correct answer, because at seq 1793 file_read ~/.gradle/gradle.properties returned [REFERENCE_EXISTS] ... does not exist. Do NOT keep retrying. The tool layer does not expand ~, and the path is outside the jail regardless. The guard reported a true thing as false.
  • seq 2715: "the DoD criteria do not require a successful backend build... the frontend build passed... I cannot fix the backend dependencies, it's out of scope. However, since the PROJECT build gate is a hard gate, I must reject it."

Right three times, and wrong-footed each time by the harness.

The structural gap: every verdict the reviewer can emit (approved, changes_requested, rejected) routes back into the agent. There is no way to say this failure is not attributable to this run. FailureTicketOpened fired correctly at seq 2374 after 6 identical failures, routed to recovery, and recovery routed straight back (RefinementIteration, seq 2584).

The retry feedback closed with "Do not re-discover unrelated files before it builds." The harness forbade the only move that would have worked.

4. Ten-turn amnesia inside a mostly empty window (#706, open)

176 ContextTruncated events, every one L2, entriesDropped climbing to 40+ and pinning.

stage truncations max dropped
discovery 14 28
scaffold_frontend 40 40
install_styling_and_ethos 20 40
configure_tanstack_query 21 16
final_review 62 42
recovery 16 32

DefaultContextPackBuilder.kt:380 uses CompressionStrategy.Conversation(). The default is keepLast = 10 (CompressionStrategy.kt:7). L2 holds 5 tool-call/result pairs. The assembled pack at round 31 of scaffold_frontend (seq 599) confirms it. It carries 4 L0 entries, 2 L1, 3 L3, and 10 L2, exactly the last ten.

The window is not full. Those same packs run 1.7k to 13k tokens, median 5.3k. The cap is entry count, not tokens, so the model is starved at 5k while a local model's context sits mostly idle.

Consequence, measured: 76 byte-identical repeat calls, 29% of all tool invocations.

repeats stage call
9 final_review list_dir {"path":"frontend"}
7 scaffold_frontend list_dir {"path":"frontend/"}
6 final_review file_read {"path":"gradle.properties"}
5 scaffold_frontend file_read {"path":"frontend/package.json"}
4 final_review shell ./gradlew assemble --no-configuration-cache

The decisionJournal is the only memory surviving eviction. Its complete content at the end of the run:

- Goal: Write a web-ui for Correx. ...
- scaffold_frontend: 1 retry, resolved — stage completed
- configure_tanstack_query: 1 retry, resolved — stage completed
- final_review: 3 retries, resolved — stage completed

Bookkeeping, not knowledge.

Preferred fix is not a larger keepLast (it costs local-inference latency and still forgets). Add a deterministic action ledger. One line per tool call for the stage: tool, target, exit, short result hash. Never evicted. 258 calls at roughly 12 tokens is about 3k tokens for a whole run. That is cheaper than the duplicates it removes, and it carries the "already tried, same failure" signal no window size provides.

Related: the plan set kind: "process_result" on every stage, which silently disables the file_written manifest injection (SessionOrchestratorContext.kt:458 builds it only for file_written slots). A feature built to prevent this thrash was switched off by one planner field.

5. There was no chain of thought

270 of 274 inference rounds returned an empty response body (blake3 af1349b9..., the empty string). Per stage: discovery 25/25, analyst 9/9, architect 1/1, scaffold_frontend 62/62, install_styling 31/31, configure_tanstack 32/33, final_review 83/86, recovery 27/27.

Every assistant message in the assembled prompts is a bare function call with no content. The system prompt instructs "make one coherent step per response."

The loop reduces to this. Context in, one tool call out, no reasoning, no plan, no self-check, with a 10-turn memory (§4) at temperature 1.0 (§6). The 29% repeat rate is what that combination produces. Surfacing reasoning is already tracked as #298. This finding is that none was produced to surface.

6. Temperature 1.0 on structured arguments (#710, open)

Main.kt:435-451 sends [sampling] temperature = 1.0, top_p = 0.95, top_k = 64 on every stage inference. Chat-quality settings applied to fields where exactly one string is correct.

seq emitted intended
277 create_vite@latest create vite@latest
465 npm_prefix=frontend npm --prefix frontend
1302 npm_install_package npm install <package>
2060 ./gradlew_ ./gradlew
2708 gradlew ./gradlew

Space-to-underscore substitution on an argv field is sampling noise. correx already sends a per-request GenerationConfig, so the override is per-call. Keep 1.0 for artifact and review prose. Go near-greedy when the response is constrained to tool calls. #303 repairs these after the fact. Not emitting them is cheaper.

Same task, second half: files are copied through the token stream. There is no file_copy tool. ethos-icons.svg moved by a file_read of 6124 characters, then a file_write of 6845 characters of tool-call arguments. A READ_BEFORE_WRITE rejection round came first (seq 540), when the model tried to write before reading. About 6 rounds and 13k tokens for two static assets, at temperature 1.0, with silent corruption possible. The project profile says "COPY design-system/ethos.tokens.css and ethos-icons.svg into the project and import them, do NOT read their contents." The toolset offers no way to obey that instruction.

7. The plan contradicts itself, and the kernel sides against the model (#711, open)

The architect emits a stage's prompt and its writes manifest independently. They disagreed twice, and the kernel enforces writes.

  • install_styling_and_ethos prompt step 4: "Copy ... into frontend/src/assets/ethos/." Manifest: [tailwind.config.js, postcss.config.js, src/index.css]. Result at seq 1086: [PATH_OUTSIDE_MANIFEST].
  • configure_tanstack_query prompt step 2: "Create a QueryClient instance and a wrapper component." Manifest: [frontend/src/lib/query-client.ts]. Blocked at seq 1321. QueryClientProvider.tsx was then written 8 times.

Both fields are already parsed by the plan-compile gate, so this is a string check over paths the prompt itself spells out.

8. The stage prompt never updates (#709, open)

Verified against the assembled prompt for round 31 of scaffold_frontend (artifact 47ebf7b0). The system message still read:

Use the shell tool to initialize a new React + TypeScript project using Vite in the frontend/ directory. Run npm create vite@latest frontend -- --template react-ts ...

The project had existed for 30 rounds. The instruction is a fixed string from the execution plan, and nothing recomputes it against what exists. Combined with §4, the model's strongest signal every round is an order to redo step one.

The plan's stage prompts are literal command transcripts written from training memory, and both were wrong:

  • npx tailwindcss init -p. npx is blocked by tool policy, rejected at seq 329 and 988. The policy message helpfully names the alternative. A plan-compile lint rejecting a plan that names a forbidden executable is roughly ten lines.
  • npm install -D tailwindcss postcss autoprefixer plus tailwind.config.js is the Tailwind v3 procedure. npm install tailwindcss now resolves v4, where the PostCSS plugin moved to @tailwindcss/postcss. That is the build failure at seq 609 verbatim, and 22 rounds of thrash followed it.

9. Ninety-four approvals that carried no information (#712, open)

Every ApprovalDecisionResolved in the run is APPROVED with reason: null and userSteering: null. Pause durations: p50 13.6 s, p90 19.0 s, max 20.3 s. Total 19.0 min, 19% of wall clock.

By tool: shell T2 x63, file_write T2 x13, file_edit T3 x10, remainder task and delete calls.

Two problems. The gate is priced per tool call, so it charges 94 interrupts for one operator intention. The attention also lands in the wrong place. All 94 chances to intervene were spent on "may I run npm install". The decision that determined the outcome (seq 236, §1) had no operator checkpoint at all.

Auto-approving reads and writes inside the stage's declared writes manifest costs nothing in safety, because PATH_OUTSIDE_MANIFEST already enforces that boundary. It returns most of the 19 minutes.

Waste breakdown

258 tool calls. 110 wasted, counting each call once (repeat, failed, or policy-rejected).

stage waste / total
discovery 2 / 24 8%
analyst 1 / 8 12%
scaffold_frontend 26 / 60 43%
install_styling_and_ethos 11 / 30 37%
configure_tanstack_query 9 / 31 29%
final_review 55 / 79 70%
recovery 6 / 26 23%

Context composition across all 279 assemblies: toolResult 47.9%, projectProfile 8.4%, assistantToolCall 8.1%, decisionJournal 6.8%, retryFeedback 5.4%. By layer: L2 56.0%, L0 18.0%, L1 13.2%, L3 12.9%.

Ranked remediation

inference, from this run's own timings:

  1. Baseline every gate command at session start (#708). One build at t=0. A command already red before the agent touched anything reports pre-existing and never fails a stage. Recovers final_review plus most of recovery. ~40 min.
  2. Fix the toolchain scope mismatch (#705). One line. Prevents the same 40 minutes by an independent route. Do both.
  3. Deterministic action ledger (#706). ~16 min of duplicate calls, plus the loops they feed. Raising keepLast alone does not fix it.
  4. Auto-approve inside the declared manifest (#712). ~17 min.
  5. Split the sampling config (#710). Near-greedy for tool-call rounds.

Items 1, 3 and 4 account for roughly 70 of the 99 minutes.

The structural point

Every gate in this run was local. Each asked whether a stage did its stage correctly. None asked whether the run was still building the requested thing, until final_review at minute 90 graded against the already-collapsed DoD.

The path from intent to work is five lossy model summarisations: intent, discovery, DoD, plan, stage prompt. Only the last link was ever checked. #699 now checks the DoD-to-discovery link, which is the one that broke here. The general shape remains.

The cheap version needs no model. The intent names eight views. After three of four stages, FileWritten holds nine files and none is a view. That comparison is a set difference over recorded events and costs nothing. It would have fired at minute 25, not failed silently at minute 99.

The deeper item, for #168: this run produced a genuinely valuable artifact and discarded it. A local model correctly diagnosed a Gradle offline-mode misconfiguration in ten minutes from build output alone. That knowledge existed at seq 1994 and was gone by seq 2306. A guard reporting a real file as missing erased it. The event log still has it. Nothing reads it back.

Verification and scope

  • fact: every number above is derived from events rows for session 954da1a9 and from CAS artifacts resolved through ~/.config/correx/artifacts/index.sqlite. No inference re-run.
  • fact: ScopeCoverage (#699) landed at 19:56 on 2026-08-11, after this session ended at 13:31. It was not active during the run.
  • inference: the ./gradlew assemble failure is attributed to ~/.gradle/init.d/offline.gradle on workpc. The build output names offline mode and an uncached io.ktor:ktor-client-websockets. The file itself was not read during this analysis.
  • unknown: whether the 94 approvals were resolved by a human operator or by an auto-approver. The tight 13-20 s clustering suggests a poll interval. No event records the decider.
  • This post-mortem implemented no fixes. It filed Vikunja #705 through #712.