Files
correx/BACKLOG.md
T

218 lines
16 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# CORREX — BACKLOG
> **⚠️ MAINTENANCE RULE — READ FIRST**
> **When an item here is resolved, MOVE it out of this file and into `RETRO.md`.**
> Do not just tick a box and leave it. Cut the entry, paste it into `RETRO.md`
> with the resolving commit hash(es) and the date it landed/was verified. This file
> is the *live* outstanding-work list; `RETRO.md` is the archive of what got done.
> "Resolved" = built **and** verified (live-QA passed where a QA gate is noted).
> Items marked **SHIPPED (pending live-QA)** stay here until their QA gate passes,
> then they move.
>
> **QA RULE — after a feature is done, plan its QA before calling it resolved.**
> Run `/qa-plan <commits|spec>` (or copy `docs/qa/TEMPLATE.md` by hand) to produce a
> `docs/qa/QA-<feature>.md` with observable-evidence checks. A feature is not "resolved"
> until that plan passes live. On PASS → move to `RETRO.md`; on FAIL → refile the misses
> here as numbered findings.
Synthesized 2026-06-14 from the five `docs/plans/2026-06-13-*.md` design specs and
session memory. Pointers are commit hashes / spec sections, not gospel — verify
against current code before acting (memory is point-in-time).
---
## A. Observability spec — `docs/plans/2026-06-13-observability-spec.md`
Metrics + CLI + TUI pane and the health backbone are shipped (fe94140, 4f6bfa8,
f8fd260) — those are RETRO candidates once live-QA'd. Remaining:
- [~] **§4 LLAMA_SERVER health probe** — **SHIPPED (pending live-QA)** `aaf896d`. Liveness via
`GET /health` + tokens/sec degradation watch (windowed avg of recorded `InferenceCompletedEvent`s,
opt-in via `llama_tps_warn_below`). Plugs into the existing `HealthMonitor`. QA gate in §F →
`docs/qa/QA-llama-health-probe.md`.
- [~] **§4 EVENT_STORE health probe** — **SHIPPED (pending live-QA)** `7a1362c`. Times a cheap
`lastGlobalSequence()` read as a responsiveness heartbeat (DEGRADED on throw or over
`event_store_latency_warn_ms`); read-only by design (no probe writes into the log). QA gate in §F →
`docs/qa/QA-event-store-health-probe.md`.
- [~] **§4/§3 live health TUI pane** — **SHIPPED (pending live-QA)** `fd67a6c`. `OverlayHealth` (key `H` /
palette "health checks") mirrors the stats pane: pull-based fetch over WS
(`GetHealthChecks``health.checks`/`HealthReport`), per-subject status with degraded loud, visible
fallback when disabled. QA gate in §F → `docs/qa/QA-health-tui-pane.md`. **With all three probes +
this pane, §4 health-checks is feature-complete in code** — retire as a unit once the QA gates pass.
- [ ] **§3 tier-2 stats** — full session-completion summary pane polish (basic pane shipped).
- 🚫 **§3 tier-3 Prometheus/Grafana export** — spec says "deferred indefinitely". **Do not build.**
## B. Role reliability spec — `docs/plans/2026-06-13-role-reliability-spec.md`
§3 analyst keystone (894969d), entity grounding (b050dc4), §2 write-manifest (4e5e4e5),
and §5 reviewer prompt/needs wiring (3467826) are shipped — RETRO candidates. Remaining:
- [~] **§5 reviewer static-first INFRA** — **SHIPPED (pending live-QA)** `365eb8a`. A stage declares
`static_analysis = [commands]`; after it produces its artifact a deterministic harness gate
(`runStaticAnalysis`) runs each command (compiler/detekt/formatters) in the workspace root,
records a `StaticAnalysisCompletedEvent` (invariant #9; recorded live, folded on replay), and on
any non-clean command fails the stage retryably with the output fed back verbatim (§2). Only
static-clean output reaches the reviewer, so its context excludes static findings *mechanically*
(resolved upstream) — no output-parsing, no shared-findings schema needed. `ProcessStaticAnalysis-
Runner` is the injected adapter (nullable seam; null → logged no-op). Prompt half was `3467826`.
QA gate in §F → `docs/qa/QA-static-first-reviewer.md`.
-**§6 `CritiqueFinding` schema + `CriticCalibrationProjection` (= addenda A3)** — **DEFERRED until
calibration is actually built.** Substrate with no live consumer today (assessed 2026-06-14): zero
`CritiqueFinding` references anywhere (the old "test expectations exist" note was stale — none do),
no `CriticCalibrationProjection`/`CritiqueOutcomeCorrelatedEvent`, and **no plan critic** — so the
spec's "single definition, two consumers" rationale has zero consumers. The one reviewer that exists
emits `review_report` = `verdict` + free-text `notes`; code reads only `verdict`, and the implementer
reads `notes` as prose fine. Structured findings only earn their place when the calibration
projection is built — which needs runtime critique/outcome history first (spec ordering #5) and a
second consumer (a plan critic) to justify sharing. Build the schema *with* calibration, not before.
- [ ] **§4 architect contradiction-check** — L3 retrieval over decision journal + ADRs →
`PossibleContradictionFlag` (display-only in v1). Unbuilt.
- [ ] **§3 symbol grounding** — deferred; needs a real symbol index (file-path grounding done).
- [ ] **§2 dynamic plan-derived diff manifest** — allowed files from the `impl_plan`
artifact. Future refinement; static per-stage `writes = [globs]` form is shipped.
- ⚠️ **§4 architect ADR shape** (`alternativesConsidered` required non-empty) —
**DELIBERATELY NOT TAKEN.** Conflicts with the standing "don't raise alternatives
unless real cost/breakage" preference already encoded in `architect.md`. Revisit only
with an explicit rationale; do not add by default.
## C. Plan pipeline addenda — `docs/plans/2026-06-13-plan-pipeline-addenda.md`
A1 is shipped (bcc59d2) — RETRO candidate once live-QA'd. A2 is **superseded** (see below);
A3 is the same item as §B §6 calibration (tracked there).
- [~] **A1 brief echo-back gate****SHIPPED (pending live-QA)** `bcc59d2`. New `brief_echo` stage
before the planner in `role_pipeline`: the model restates the analyst brief as a structured
artifact; a deterministic gate (`checkBriefEcho`, opt-in `brief_echo=true`) diffs the restatement
vs the brief — dropped requirements (Jaccard coverage < 0.34) or hallucinated files →
`BriefEchoMismatchEvent` + retryable stage failure, so a misread brief never reaches plan
generation. Pure replay-safe diff (`BriefEchoDiff`); symbols recorded but non-blocking in v1.
QA gate in §F → `docs/qa/QA-brief-echo-gate.md`. `freestyle_planning` wiring is a follow-up.
- 🚫 **A2 stage-level plan checkpointing****SUPERSEDED, do not build.** The spec predates the
machinery that now carries this weight. In the current design the "confirmed plan" is the
freestyle `execution_plan``ExecutionPlanCompiler``WorkflowGraph`, and per-stage
reconciliation against the plan's produces-slots *already happens* in phase-2 execution:
`verifyProduces` checks every `stage.produces` slot (sourced from the plan JSON) against
`ArtifactCreatedEvent`s and fails retryably → surfaces on exhaustion; `ManifestContainmentRule`
blocks out-of-plan writes; the validation pipeline checks content vs kind schema. The only net-new
bits A2 adds are a positive `StageCheckpointPassedEvent` + terminal-vs-retryable halt — and the
sole consumer of the former is A3 calibration, which is itself ordering-blocked. Revisit only if a
real consumer for a positive per-stage checkpoint event appears. (Assessed 2026-06-14.)
- [ ] **A3 critique calibration tracking****= §B §6 `CriticCalibrationProjection`** (build once, two
consumers). Tracked under §B §6; needs runtime critique/outcome history first (spec ordering #5).
Not duplicated here.
## D. Research workflow spec — `docs/plans/2026-06-13-research-workflow-spec.md`
Core workflow is **SHIPPED end-to-end, ON by default** (fb1d970..378ee39): extractor +
web_search (T1) + web_fetch (T2) in `:infrastructure:tools`, `research.toml`
(decompose→gather→report), schemas, prompts. Remaining:
- [ ] **Live-QA the full run** — could not validate in-session (no SearXNG/network in sandbox).
Needs SearXNG running; exercise triage→fetch-approval→report→artifact viewer.
- [ ] **§6 web approval client** — the one larger unbuilt piece. Third client on the existing
Ktor REST+WS bus: approval cards + event strip + session list. Behind WireGuard/Tailscale
only, never public. Independent track (spec ordering #5).
- [ ] **§3 batch fetch-approval at source-list level** — currently per-fetch T2; operator
should approve the source list once, not each URL.
- [ ] **Dedicated `SourceFetched` / `LowQualityExtraction` events** — quality +
content_sha256 currently ride inside `ToolExecutionCompletedEvent` metadata only.
- [ ] **Dynamic per-session egress allowlist** — currently static `networkAllowedHosts` + T2 gate.
- [ ] **Fresh-machine runtime install** — copy `research.toml` + `prompts/` + `schemas/` and
the `[[artifacts]]` / `[tools.research]` config entries into `~/.config/correx`
(already done on this box; repo source is `examples/workflows/` + `docs/schemas/`).
- 🚫 No headless browser, no LLM-based extraction, no public approval endpoint (non-responsibilities).
## E. TUI requirements spec — `docs/plans/2026-06-13-tui-requirements-spec.md`
§2 last-event clock + unknown-event fallback (8b896b5) and §5.1 deterministic command-parse
card (65df3f4) are shipped — RETRO candidates. Remaining:
- [ ] **§2 full Go↔Kotlin per-event-type render matrix** — golden-fixture tests asserting a
defined render for every event type (last §2 bullet; partially covered today).
- [ ] **§3/§4 plan comparison view** — top-2 candidates side-by-side, compact stage graphs,
diff-highlighted, lint/critique findings inline. Lands with the planning pipeline;
load-bearing for the future preference ranker.
- [ ] **§3 approval ergonomics audit** — single-key y/n/e for T0T2, mandatory confirm T3+,
queued-approval navigable list (never modal-stacked). Verify against current band impl.
- [ ] **Render markdown in chat turns** (QA feature note b) — TUI currently ignores md.
- [ ] **Session list / resume view** (QA feature note c) — after server restart + TUI reopen
there's no way to see a prior session or its state; `GET /sessions` data exists server-side.
- [ ] **Idea board → profile promotion** — deferred hook; auto-captured ideas promote into the
curated `.correx/project.toml` profile. (Board itself shipped 3ac0a14..02e963d.)
-**§5.2 command-card LLM annotation layer** — deferred by spec ("ship layer 1, live with it,
then decide whether layer 2 earns its place"). Build only after using layer 1.
## F. In-flight QA gates (move to RETRO when each passes)
These are SHIPPED in code but prompt-/server-/network-dependent and not yet live-verified:
- [ ] **Epic 15 live-QA gate** — server up → `correx events` / `correx replay` on a real
session, confirm MDC sessionIds in logs. (Code complete a7d5211..75703ec.)
- [ ] **Idea board live-QA** — needs a real router model to emit the `{"ideas":[…]}` block.
- [ ] **Clarification loop + workflow-propose panel live-QA** — both prompt-dependent (need a
real router/analyst model to emit the structured json blocks).
- [ ] **Research workflow live-QA** — see §D.
- [ ] **LLAMA_SERVER health probe live-QA**`aaf896d`; needs a real llama-server to ping.
Plan: `docs/qa/QA-llama-health-probe.md` (liveness up/down/restore, non-2xx, tps degrade,
restart hysteresis, replay independence).
- [ ] **EVENT_STORE health probe live-QA**`7a1362c`. Plan: `docs/qa/QA-event-store-health-probe.md`
(responsive→HEALTHY, no-log-pollution, slow/throw→DEGRADED, restore hysteresis, replay independence).
- [ ] **Health TUI pane live-QA**`fd67a6c`. Plan: `docs/qa/QA-health-tui-pane.md` (`H` opens/fetches,
matches `correx health`, degraded loud, disabled fallback, resize, live wire-contract values).
- [ ] **Brief echo-back gate live-QA**`bcc59d2`; needs a real planner-capable model + runtime install
(schema/prompt/artifacts block). Plan: `docs/qa/QA-brief-echo-gate.md` (faithful echo→planner runs,
dropped requirement→`BriefEchoMismatchEvent`+retryable fail, hallucinated file flagged, symbols
non-blocking, replay recomputes identically, non-opted pipeline is a no-op).
- [ ] **Static-first reviewer gate live-QA**`365eb8a`; needs a real implementer-capable model + a
workspace whose configured `static_analysis` commands can be made to pass/fail on demand (and the
gate uncommented in the copied `role_pipeline.toml`). Plan: `docs/qa/QA-static-first-reviewer.md`
(clean→reviewer runs, non-clean→`StaticAnalysisCompleted`+retryable fail w/ output fed back,
gate-OFF absence check, no-runner no-op WARN, replay reads recorded findings without re-running).
- [ ] **TUI kernel-steering + AMD gauge fresh live-QA** — approve+note actually revising the
*same* stage's output; AMD VRAM/RAM gauge on the real box.
- [ ] **Unit-tested-only fixes, not live-verified:** `1a7eb05` (verdict-edge rejection),
`e195c79` (stale-parking `OrchestrationPaused ABANDONED_STALE`), `df3ee0b` (bare-id
narration model match).
## G. Low-severity hardening
- [ ] **turnId prefix collision**`startsWith("repomap:$repoRoot")` can collide where one
root path prefixes another (`/repo` vs `/repo2`). Fix = trailing `:` delimiter in 3
places (ProjectMemoryService write+check, L3RepoKnowledgeRetriever filter).
- [ ] **Narration lane lag** — stale pause-trigger narrations blend with current workflow
status; drop queued pause narrations once their approval resolves.
- [ ] **WorkspaceStateProbe fingerprint** — hashes `path|mtime|size` without newline separators
though the docstring says "lines". Internally consistent/deterministic; cosmetic.
- ⚠️ **`PreemptRedirectEvent` has no producer** (QA bug #6) — B3 hard re-route is unreachable
from the TUI. Intended flow: steering → router proposes → operator confirms. Skipped by
design; build only if the hard-reroute path is actually wanted.
## H. Freestyle mode — deferred follow-ups (epic COMPLETE, B4 PASSED)
- [ ] **Graph re-routing by preemptive steering** — only priority-directive surfacing is built
today; mid-execution input cannot re-route the compiled graph.
- [ ] **Child-session `SubagentRunner` impl** — seam exists (`InSessionSubagentRunner`);
child-session is the intended drop-in swap. Sequential only (no-parallel-agents invariant).
- [ ] **Post-lock `ApprovalDecisionResolvedEvent.userSteering` → L0 promotion** — currently
stays L2; out of freestyle's exercised path (gate is pre-lock).
## I. Gamedev / Godot readiness (operator intent)
- [ ] **`.cs` in `RepoMapIndexer` SYMBOL_PATTERNS** — add if the gamedev repo is Godot Mono
(`.gd` and `.md` headings already shipped: 81a21bf, f521ac8).
- [ ] **Structural `.tscn` / `.tres` validator** — parse scene format, node-tree-aware diff
preview. Prerequisite validation layer before any scene editing.
- [ ] **(far future) Scene editing behind approval+diff** + derive a fine-tuning dataset from
the event log. ⚠️ Fine-tuning is on CLAUDE.md's intentionally-NOT-implemented list — needs
an explicit epic + operator de-listing, not an incidental add.
---
## Ops notes (not backlog items, but bite repeatedly)
- Config copies of `examples/*` in `~/.config/correx` go **stale silently** — e.g. a missing
`inject_artifact_kinds` in a copied workflow TOML caused invented artifact kinds. Re-sync
after editing the repo examples.
- **Do not rebuild jars while the server JVM is running** — lazy classloading breaks
(`NoClassDefFoundError`). Stop the server first.