# Maven completion journal This journal tracks the autonomous completion goal started on 2026-08-13. It is an operational index, not a substitute for living subsystem documentation, dated evaluations, Vikunja tasks, or focused caveat entries. ## 2026-08-13 — baseline and backlog reconstruction Goal: make Maven usable end to end with every current and planned feature wired, tested, and polished. Completion requires clean automated gates and successful sessions across local, degraded-ecosystem, and integrated modes. Initial observations: - `HEAD` is `2cf8b7e`, identical to both local and remote `master`, while the checked-out branch is the stale `task/704-...` branch. - The worktree already contained staged documentation/evaluation changes, staged transcript deletions, an unstaged `deploy/mavwaked.service` change, and untracked `deploy/asoundrc`. These are pre-existing work and are being preserved and validated before any commit. - The repository has no prior goal journal. Durable subsystem facts continue to belong under `docs/`; unresolved limits continue to belong under `docs/caveats/` with a task and revisit trigger. Work streams started: - Vikunja project 2: inventory every open task and recover acceptance criteria. - Repository: compare feature plans, caveats, routing/ecosystem contracts, and implementation. - Verification: run build, race tests, simulator, analyzer gates, and inspect skipped hardware/model evaluations separately. - Runtime: exercise the web, IPC, voice, model, and sibling-service paths with real local dependencies where available and explicit degraded-mode probes otherwise. References: `docs/workflow.md`, `docs/qa.md`, `docs/ecosystem.md`, `docs/routing.md`, `docs/caveats/CLAUDE.md`, and Vikunja Maven project 2. ### Backlog correction The first Vikunja page was accidentally read without a `done: false` filter and mixed closed history into the working set. Re-querying all pages strictly open produced 127 records: 89 implementation-open, 20 shipped-but-QA-open, 12 external/owner gates, and 6 duplicate or stale-open records. Closed tasks are used only as commit provenance; they are not work to redo. V-704 was the only open hanging task and was closed after its measured correction landed. ### Model-aware baseline The ordinary `make test` passed but does not set `MAVEN_ONNX_LIB`, so model-aware tests can self-skip. The explicit ONNX boundary gate exposed V-702/V-703: the held-out `я рассказывал тебе про байкал?` was the sole miss at 28/29. A three-neighbour class score fixed the semantic collision without adding a word pattern or copying the held-out sentence. Boundary is now 29/29 and the adjacent topic gate remains 43/43. Measurement: `docs/evals/2026-08-13-personal-boundary-neighbourhood.md`. That narrow result was not accepted as the completion gate. A second agent wrote a balanced 72-case RU/EN matrix across remembered speech, possession, narrative, proper nouns, personal preambles, and advice/current questions. It contains no production seeds and no Baikal paraphrase. The top-three candidate scores only 61/72 (84.7%); top-two reaches 62/72, one-neighbour 56/72, and a whole-class centroid 54/72. V-702 therefore remains open while a principled classifier is developed against the independent matrix. The 29/29 measurement describes the narrow regression set, not general boundary quality. ### Live delivery incident The five-service compose stack was running, but a due reminder was being re-phrased and retried through ntfy every tick. The sink returned HTTP 403 each time. Only secret names were inspected: the configured ntfy and workstation token variables were absent from the deployed environment file; no secret values were read into this journal. The durable outbox records each failed attempt, but the retry path has no backoff or alternate channel and spends the resident model again before every failure. This is active V-651 behavior, with the repeated-phrasing shape related to V-687. V-715 now owns the incident acceptance criteria. In the working tree, phrases, collapsed-group identity, attempts, and next-attempt time are durable; definite failure backs off from one minute to a capped hour; retries and restarts reuse the exact phrase; and away delivery tries ntfy then Telegram, stopping at the first success. The committed deployment explicitly disables the uncredentialed ntfy block. Independent review added a real occurrence key for collapsed bundles, suppresses crash-ambiguous attempts from automatic replay, classifies HTTP 401/403 as permanent, blocks permanently unreachable reminders visibly, and commits the successful outbox result plus every collapsed original in one SQLite transaction. The store, delivery, IPC, loop, config, and mavweb race suites pass. A live rebuild and one-time delivery of the existing backlog are still required before V-715 is closed. ### Explicit integration enablement V-691's deployment boundary was audited against every `${VAR}` reference. The canonical `deploy/telegram.env.example` now names Telegram, ntfy, workstation model, workstation STT, Home Assistant, ambient, CW2, and database-key inputs. Enabled Telegram, ntfy, ambient, non-loopback workstation model, and non-loopback workstation STT paths refuse missing credentials; each arm has an explicit disabled state. The live config disables the currently uncredentialed ntfy and workstation-model arms while retaining the separately credentialed STT arm. CW2 also refuses a non-loopback bind without its token. Focused Go race suites, the Python CW2 startup contract, deploy-config drift test, secret-expanded config validation (values not printed), and `docker compose config --quiet` pass. The workpc is unreachable from this host, so installing the updated CW2 script there remains an external deployment step; the affected model arm is explicitly dark rather than ambiguously half-live. ### Traceable web failures V-689 gives every mavweb response a server-generated request ID and routes every handler failure through one sanitized problem envelope. Stable error codes and the request ID reach the browser; the wrapped internal error reaches only the server log beside the same ID. Degraded inline panels use stable public text rather than backend paths or tokens, and direct ecosystem reads propagate the web request ID as their correlation ID. An AST guard prevents new production handlers from bypassing the contract with `http.Error`. The full mavweb race suite passes, including disclosure, untrusted-ID, log-join, and propagation tests. ### Bounded external responses Three audit defects were repaired and committed directly to `master`: - V-608 (`d7e8804`): llama completion responses are capped at 1 MiB, including the LAN workstation seam. - V-675 (`459fe7a`): remote STT requires nonblank text and an explicit finite confidence in `[0,1]`, caps JSON at 64 KiB, and falls back to mavsttd on a malformed HTTP 200. - V-676 (`7d0250a`): Open-Meteo geocoding and forecasts are bounded, required fields are nullable/validated, and coordinates/weather values are range checked so `{}` cannot become plausible zero-degree weather. Each focused race suite passed and each task was closed only after the commit. ### Transport shutdown V-679 (`de61b75`) adds the listener's `done` channel to TCP `Accept`. A concurrent-close test holds a silent peer in handshake and proves an in-flight accept returns `net.ErrClosed`; the race test passed twenty consecutive runs. V-688 (`80b6068`, caveat retirement `a0e6643`) bounds the browser push-to-talk body at ten minutes of mono PCM and configures header, idle, and read limits on the web server. The unused `/ws` handler was removed instead of retaining a second unauthenticated streaming transport with no browser caller. Focused race tests prove the exact-size request succeeds and an oversized request returns HTTP 413. ### Conversation continuity V-542 (`da9114b`) repairs the five-turn monitor conversation without changing single-turn intent classification. Exact user utterances are now persisted separately from normalized intent slots and retained in chronological order. An anaphoric query with live transcript context reaches the chat path, while non-anaphoric sources are unchanged and acts stay fail-closed. An explicit conversation opener extends the session lifetime through later fact/query/chat routes without suppressing the grounded fact write. The deterministic scenario now names the monitor in all four contextual replies, proves that the original raw turn reached `PhraseChat` four times, stores the fact once, and produces zero unsolicited sends across five turns and one tick. Focused race tests passed for `cmd/mavend`, `internal/dialogue`, `internal/router`, and `internal/lexicon`. Measurement: `docs/evals/2026-08-13-conversation-continuity.md`. ### Personal-data boundary V-702 replaces the narrow nearest-neighbour privacy gate with a frozen, class-balanced logistic head over multilingual-e5-small. It introduces no lexical exception and leaves the decision threshold at 0.5. Historical regressions score 29/29 and the balanced 72-case RU/EN fixture scores 72/72. The first 24-case challenge found one private-configuration miss. That result was treated as model-selection data rather than advertised as independent proof. Shrinkage LDA and an LDA/logistic ensemble repaired it but regressed the 72-case gate, so both were rejected. Increasing the logistic L2 coefficient from 0.0001 to 0.0003 repairs the miss while improving four-fold corpus cross-validation from 97/104 to 99/104 and whole-shape holdout from 91/104 to 92/104. A fresh 24-case challenge written only after that head was frozen scores 24/24 at minimum signed probability margin +0.1718. The original challenge is also 24/24 but its +0.0001 edge remains documented as a regression, not fresh evidence. Full measurement: `docs/evals/2026-08-13-personal-boundary-linear-head.md`. During the audit, running multiple ONNX-backed tests in one `go test` process showed that only the first initializes; later tests self-skip because the runtime is process-global. All V-702 figures were therefore rerun in separate processes. V-716 tracks fixing that harness gap rather than hiding it in this feature. ### Forced dialogue and repair state V-573 closes all four repair seams exposed by the dialogue contract: a correction wins before a parked clarify answer; a repaired decision is checked for required slots before acting; a request completed through clarification is correctable; and declined or stale repairs do not prematurely spend the repair pointer. Same-intent corrections are handled explicitly without redoing the action, so their prose cannot route fresh and overwrite the retained pointer. The independent state audit found two deeper stack losses. A handled repair could leave an older question silently parked with its old TTL, and a repaired request needing clarification could overwrite—or, on completion, delete—the older flow. Repairs now suspend and audibly resume live questions, repaired questions push onto the bounded dialogue stack, and completion/cancellation pops only the active top before resuming the flow underneath. `MAVEN_DIALOGUE_NO_SKIP=1 go test -race ./cmd/mavend -run '^TestDialogueTraces$' -count=1` passes all 22 traces. The complete forced `cmd/mavend` race suite passes in 208.031s. The integrated race command over `cmd/mavend`, `internal/dialogue`, and all `internal/router` packages also passes (162.310s for mavend; every package green). Focused structural possession, repair-pointer, nested-stack, and repaired-clarify tests pass under the race detector. ### ONNX test/runtime lifecycle V-716 found that each embedder constructor tried to initialize ONNX Runtime, while `Close` destroyed only its session. In one package process the first model-aware test ran and later tests converted “already initialized” into a green skip. The router now owns the process-global environment through reference-counted leases held by each embedder and routing-head session; the last owned lease performs cleanup, and close is idempotent. The router and mavend test packages hold a lease across their model gates. `make eval-router` additionally requires proof that both named aggregate gates actually executed. In one process the classifier baseline scored 72/96 and the routing heads 93/96; destination was 11/33 and 25/33 respectively, and ecosystem reach remained 28/30. The lifecycle reacquire test, focused race suite, full aggregate command, and portable no-runtime packages all pass. Measurement: `docs/evals/2026-08-13-onnx-runtime-lifecycle.md`. ### Clarification exhaustion is fail-closed V-717 closes the terminal-policy hole found during the V-573 audit. A request with two required gaps could spend its only question on the first, fill that slot, and then reach `applyAction` with the second still absent. The attempt cap was accidentally acting as permission to execute a partial action. The resolver now rebuilds the pending action and re-runs the canonical `missingFor` schema after every filled gap. One remaining gap produces exactly one next question only while the shared `PendingAction.CanAsk` budget permits it. Exhaustion visibly gives up, removes only the active stack level, and makes no write or action. `finishRebuilt` repeats the same invariant at the execution boundary. Reminder time answers remain separate from the clean payload but are included in the schema decision used for validation. The original `TestClarifySecondGapRespectsTheAttemptCap` now asserts the exact give-up and zero reminders. New tests cover direct boundary refusal and a two-level stack where exhausting the top appends the surviving lower question to the same reply. The focused V-717 race cases pass in 4.529s; every clarify case plus all 22 forced dialogue traces pass under the race detector in 26.202s; `internal/dialogue` passes under race in 2.293s. Routing contract: `docs/routing.md` section “Required slots and attempt exhaustion”. ### A suppressed nudge is identified before it is phrased V-687 closes the phrase-before-dedupe hole in the digestion worker. The dedupe was reported by `EnqueueDigestEntry`, which runs after `PhraseNudge` has already been paid, and the `else if deduped { continue }` meant to skip the cost was the last statement in the loop body. Every tick that kept suppressing the same rule spent the resident model again, against the cache claim in the comment above it. The fix gives a rule a durable semantic identity instead of hashing its prose. A rule eligible for the digest declares `DigestIdentity`, a function of state beside its predicate; `loop.DigestCandidateFingerprint` frames the rule name and severity around it so two rules cannot alias on a shared fact. `BreakRule` anchors on the last completed break rather than on `desk_active`, which is freshness evidence the poller refreshes without the unmet need changing. A rule with no declared identity does not enter the digest, because inventing a generic state hash would either change every tick or ignore an input the rule reads. `tick_digest.go` now looks up `LiveDigestEntry` by rule and fingerprint before phrasing. Migration #26 adds `candidate_fingerprint` with a partial unique index over live pending rows; `EnqueueDigestEntry` expires a matching stale row and inserts inside one transaction, so sweep order is not part of correctness and a second caller cannot race the pre-phrase read into a duplicate. Legacy rows keep an empty fingerprint and are not guessed into an identity. Six tests cover the contract: one phrase call across three suppressed ticks, zero after a daemon restart, and two when the meaning changes, when the entry expires, and when it has been drained. `./cmd/mavend/ -run TestSuppressedCareDigest` passes under race in 4.626s, the digest store and loop cases in 4.123s and 1.046s, and the three full packages in 264.076s, 64.496s and 4.280s. The caveat `docs/caveats/workers.md#nudges` and the `SA4006` baseline entry are deleted. ## 2026-08-15 — authoritative-state continuation The continuation began by treating the checked-out tree and external task state as authoritative. `master` was clean, identical to `origin/master`, and already contained the V-717 and V-687 branch merges. Their interrupted worker messages were therefore not used as evidence of missing work. ### Live reminder catch-up closes V-570 and V-715 The rebuilt stack loaded the resident Qwen model, multilingual ONNX embedder, and routing heads. On its first eligible tick it phrased the three overdue reminders once as one catch-up message. Disabled ntfy left one durable failed attempt for delivery group `reminder #83`; ordered fallback then left exactly one successful Telegram attempt for the same group. All originals became `fired`. Four later ticks produced no second phrase, send, pending attempt, or unknown attempt. `/`, `/reminders`, and `/notifications` each returned HTTP 200 with a server-generated request ID. This supplies the live evidence that was still missing in the 2026-08-13 entry; V-715 and the stale-open V-570 are closed. The first HTTP repro command also exposed a zsh test-harness trap: `path` is a special array tied to `PATH`, so using it as a loop variable removed command lookup inside the loop. The corrected probe used `probe_path` and `wget --no-proxy`; the deployed web process had not crashed. ### Stale-open task reconciliation The open Vikunja list was read with `done: false` and compared with the current tree before choosing work. V-397 already described itself as done and merged; V-557 is present as `ea0eb16` plus the forced missing-slot dialogue contract; V-570 is covered by the live catch-up proof above. Those three records were closed rather than reimplemented. V-651 remains open: its nil-sink and reminder spin halves landed with V-715, but the first failed severity-4 Telegram send is still not represented in the repeat-until-ack stream. ### Stable hash-floor performance evidence [V-718] `TestPersonalBoundaryHashFloorLatency` coupled correctness to ambient machine load while running a numeric training loop under race and coverage instrumentation. It is now a deterministic fit-and-score test that also proves the 1024-dimensional head was built. Elapsed time remains observable through `BenchmarkPersonalBoundaryHashFloorFitAndScore`, where three one-iteration runs on this host measured 75.1 ms, 76.8 ms and 81.4 ms without making those host figures a CI pass condition. The exact race-plus-coverage focused test passed in 9.445s. `make test` then passed on its first run: formatting, vet, CW2 configuration tests, and every internal and command package under race plus coverage; `cmd/mavend` completed in 205.712s. This closes V-718 without raising a brittle timeout. ### Delegation availability All three available subagent slots were filled: clarification exhaustion, durable nudge identity, and a read-only live deployment probe. The first two left complete merges on `master`; all three later reported the same shared Codex usage limit, with capacity unavailable until 2026-08-20. Work continues serially. The temporary constraint and revisit trigger are recorded at `docs/caveats/workers.md#agent-quota` under the V-714 completion umbrella.