From 20aa2d59c9463f6a2dcbf5abfb6f2c4e37eaf79f Mon Sep 17 00:00:00 2001 From: claude Date: Sun, 2 Aug 2026 14:40:22 +0400 Subject: [PATCH] docs: session 3 results and the query-source findings (V-459) --- docs/qa.md | 169 ++++++++++++++++++++++++++++++++++++++++++++++------- 1 file changed, 149 insertions(+), 20 deletions(-) diff --git a/docs/qa.md b/docs/qa.md index 8876454..e5fd079 100644 --- a/docs/qa.md +++ b/docs/qa.md @@ -19,6 +19,31 @@ the redeploy. --- +## What the 02-08-2026 run found + +Sessions 1 and 2 ran, and so did most of session 3. Read these four before +picking anything up. + +- **470: a question writes invented knowledge into memory.** Recall then serves + it back. `что дальше?` lands on `IntentFact` and stores the model's answer as a + `self` fact at confidence 1.00. Two junk rows then claimed seven unrelated + world questions through recall, outranking the search leg. A question about the + capital of Australia was answered `какая последняя версия языка Go?`. Two bad + writes silently disabled world answering, with nothing logged. +- **466: a pending clarify is global.** One unanswerable clarify swallowed the + next three utterances from three separate sessions. With ntfy, telegram and + voice all live, a clarify raised on web chat eats the next telegram message. +- **467: spoken task capture is dead.** The router calls the capture marker an + `act`, and capture is reachable only from the `note` intent. +- **The classifier baseline in this repo was wrong**, and it flattered the + router. See session 2 and **464**. + +Thirteen defects were filed on 02-08-2026: 462 through 474. Six tasks this plan +had written off as blocked turned out to be ready to check. Two of the six passed +the moment they were run. + +--- + ## Before you start Two things bite anyone running these checks on homesrv. @@ -41,7 +66,10 @@ Nothing here has been confirmed since the redeploy, and everything else assumes it works. Do this first. Closes or advances: **44** (conversation), **45** (text chat), **287** (voice -session quality), **321** steps 3-5 (quiet mode), **288** (STT fixtures). +session quality), **321** steps 3-5 (quiet mode), **288** (STT golden audio). + +**288 is not blocked.** The fixtures are committed under `cmd/mavsttd/testdata/` +and `make test-stt-golden` runs today. This plan said otherwise until 02-08-2026. Steps 1 and 3-6 were run on 02-08-2026 and pass. Steps 2 and 7-9 still need a person at the box, because they need a microphone or a nudge to arrive. @@ -213,26 +241,88 @@ exercises it whether you name it or not. **285** is not verification: the bridge framework works and the remaining ask is more adapters. Decide which reach comes next, or park it. +Run 02-08-2026. **280 is blocked.** No morning routine is configured (**472**). +`morning.Item` also has no required-versus-optional field, so behaviour 1 cannot +hold whatever you configure (**473**). **281's digest gap is closed**, and +its presence rule passes on inspection. Three of its five items need traffic the +box has not had. **283 is blocked**: nothing feeds the intake journal. **128 +found the worst defect of the whole session, see below.** + +For **285**, two facts bear on the choice. Synapse is already running on this box +and healthy, so a Matrix reach has a live target and needs no new service. And +mavweb is already a PWA with a service worker, which 285 itself calls the highest +value adapter left. Today's reaches are ntfy, telegram and voice. + **Query sources** (**258**, **286**): ask her something the RSS feeds answer and something only a ZIM answers, with the search block on. Live search leads and the -ZIMs are the fallback since 02-08-2026. A ZIM answer to a current-events question -means the search leg failed silently. **286**'s remaining half is doc and +ZIMs are the fallback since 02-08-2026. **286**'s remaining half is doc and git ingestion, which is build work, not a check. +**Do not read `/trace` for this.** `/trace` is the nudge-rule trace: rule, +severity, predicate, gate, selected. No query-source field exists anywhere in the +codebase. The only evidence of which query source claimed a turn is the +`voice: search:` and `voice: kiwix:` lines in `docker compose logs mavend` +(`actions_query.go:589` and `:660`). + +Run 02-08-2026, 20 turns. **Search leads and the personal boundary holds.** Every +world question that reached the boundary was claimed by search. All three +personal questions produced no search and no kiwix line at all. + +The rest of this sitting went badly. **Kiwix has zero live coverage.** SearXNG +returns four results for everything, including two invented nonsense terms. So +`querySearch` always claims, and Kiwix is unreachable code as deployed. The ZIM +half of the 02-08-2026 decision is unverified. A ZIM answer cannot signal a +silent search failure, because a ZIM answer cannot happen. +**Ordering defects** in feeds and calendar, plus 258 step 1's utterance not +working: **474**. And the sitting independently found stage 2 of **470**. + **Tasks and calendar** (**129**, **130**, **127**, **126**, **246**): capture a task by voice, confirm it lands, check prioritisation ordering is not nonsense. **246** (mail reader) also exercises the `IngestMail` rung that moved to `AuthWrite` this morning. +Run 02-08-2026. **129 passes.** The page and the spoken answer agree on ordering. +The undistinguished task carries no invented reason on either surface, which is +the thing 129 asks for. **130 fails outright** and **127 half fails**: +**467**, **469**. **246 cannot be run**: `mavmaild` is commented out in +`docker-compose.yml` and there is no `email` block, so nothing in steps 4-13 is +reachable. The `IngestMail` rung does sit at `AuthWrite` +(`internal/auth/policy.go:96`, asserted in `auth_test.go:421`), verified by +reading only. + **Routines and patterns** (**43**, **46**, **247**, **254**): these need history to detect against. If the database is thin after the outage, they may have nothing to propose, which is not a failure. Check `/routines` before concluding anything. +Run 02-08-2026. The answer is the middle case: **the detector ran and found +nothing.** The tick loop is live, and `detectPatterns` is called unconditionally +at `cmd/mavend/tick.go:227`. It has run about 25 times since the restart. It +finds nothing because the events table is empty upstream of it. Rows land there +only from `pattern.Extract` at fact-write time, and `Extract` requires the fact +value to match a closed 7-action lexicon. All 200 facts on `/history` are +`page_heartbeat`, `netdata_alarm`, `quiet_hours`, `name`, `service_down` and +`рост`. Not one lexicon hit, so no event can exist, let alone the four one pair +needs. **46 step 5 passes**: `/routines` renders `noticed 0` with the empty state +and the hint string. + +Two things block this sitting, and both are build work. The seeding recipe on +**43** goes through `sqlite3` and cannot work. And `pattern.Detect` has no +minimum-interval floor, so seeding by hand mints a permanent false routine +(**468**). Do not try to seed a pattern with four fast chat turns. + **Ecosystem** (**272**, **273**, **276**): nexus, hexis and praxis are wired and logged clean at boot. **276** is the degraded-mode suite, which means taking siblings down on purpose. Worth doing while you are already in there. +Run 02-08-2026, read-only half. All three answer `/health` 200 and `/ecosystem` +lists 18 Hexis capabilities with correct read-only and mutating badges. **272 and +273 are blocked on empty data**, not on code. Nexus holds no entities, Praxis +holds no attention items, and the Calls panel has never recorded a call. See +**472**, and read its warning first. 273's trace fix has never been validated +here. An empty Calls panel is exactly what the old bug looked like. The page is +`/ecosystem`, not `/siblings`. + **Operations** (**249**, **250**): these bite hardest if they are broken, and nobody has pulled either lever on this box. Roll **249** forward and back once, then swap `phraser.model_path` and confirm **250** reloads without a restart. Do @@ -240,27 +330,66 @@ this sitting last, because both checks can take the box down. --- -## Housekeeping (one sitting, no box needed) +## Housekeeping (done 02-08-2026, and this section was mostly wrong) -Six QA tasks will not close no matter how long they sit, because they are -gated on something that does not exist: +This section claimed eleven tasks were not verification work. **Three were not. +The other eight are.** Every one of the eight has shipped, tested code behind it. +The error ran one way: it wrote off work that is ready to check. Do not trust a +"nothing is built" line in this plan without grepping for the package first. -- **125** zenmoney: needs a token you have not minted. -- **256** Home Assistant: needs HA configured. -- **257** Bluetooth: BLOCKED, no bluez on the box. Says so in the title. -- **288** STT golden audio: needs fixtures generated. -- **14** cold-start unlock: the `-wrapped-key-file` seam exists, the passkey to L3 - half does not. `lockedAPI` was deleted as dead code in PR #50, so there is - nothing to verify. -- **284** replayable full-system simulator: nothing is built. This one is a - design task wearing a `QA:` prefix. +Relabelled to `Blocked:`, claim verified: -Relabel these so they stop reading as backlog. They are not verification work -that is pending, they are work that has not started. +- **125** zenmoney. `internal/zenmoney/` ships and is tested against a fixture. + `deploy/zenmoney.token` does not exist and the compose mount is commented out. + One token unblocks it. +- **256** Home Assistant. `internal/smarthome/` ships, the `smarthome` block sits + in `deploy/mavend.json` at `enabled: false`, and 8123 and 1883 are closed. +- **14** cold-start unlock. The seam is real at `cmd/mavend/main.go:128` and + `internal/webauthn/prf.go` is in place. `lockedAPI` is gone, replaced by + `Server.Check` in `internal/ipc/server.go`. Gated on an authenticator that + implements the WebAuthn PRF extension, which is hardware, not code. -Same treatment for the five plan-only tasks (**251** MCP, **252** vision, -**253** hearing, **255** speaker recognition, **259** crawler). A `QA:` prefix on -a plan is misleading. +Left alone, because the claim here was false: + +- **284** simulator. `cmd/mavend/simulator_test.go`, three scenarios under + `cmd/mavend/testdata/scenarios/`, and a `simulate` target at `Makefile:98`. + **Run 02-08-2026: all three scenarios pass**, plus the determinism and + backwards-step guards. One defect found, see below. +- **288** STT golden audio. Four WAVs and `golden_v1.json` are committed under + `cmd/mavsttd/testdata/`, the make targets exist, and `models/stt/ggml-small.bin` + is on the box. Session 1 lists 288 as blocked on fixtures, which is wrong. + **Run 02-08-2026: all four pass**, WER at or under ceiling with no drift. + +| fixture | transcript | WER | ceiling | +|---|---|---|---| +| ru_reminder | `Напомни мне через час позвонить маме.` | 0.00 | 0.10 | +| ru_fact | `А отметь, что я выпил воды.` | 0.20 | 0.25 | +| ru_query | `Что у меня сегодня по календарю?` | 0.00 | 0.10 | +| en_act | `Restart the web server and check the disk space.` | 0.00 | 0.10 | + +That also settles a session 1 worry indirectly: whisper.cpp works on Vulkan +after the redeploy. Only the mic and the wake path remain unproven. + +**The simulator routes with an empty seed set.** Every `make simulate` run logs +`loaded 0 seed examples from models/seeds`, seven times per scenario. The test +runs from `cmd/mavend`, and the seed path is relative to the repo root. The +scenarios still pass, which means they pass without the classifier having any +seeds to match against. Whatever 284 is proving, it is not proving the routing +the deploy runs. Fix the path before trusting a green simulator. +- **257** Bluetooth. The bluez half is genuinely absent. The LAN-scan half shipped + (`internal/netscan/`), and steps 1-9 run today. Only step 10 is Bluetooth, so + relabelling the whole task would bury real pending work. +- **251** MCP, **253** hearing, **259** crawler. All three ship + (`internal/mcp/`, `internal/capture/`, `internal/crawl/`) with no external gate. + Fully checkable. `259`'s step 1 wants no `crawl` block in `deploy/mavend.json`, + and there is none, so it is already set up correctly. +- **252** vision and **255** speaker recognition. Both ship. Each is blocked only + on a model download: a vision gguf with mmproj, and a speaker embedding model. + Neither is present under `/mnt/hdd1`. Their refusal-path steps run today. + +So the honest split is three blocked on a credential or hardware, two blocked on +a download, and six ready to check. That is roughly a session of real QA this +plan had written off as backlog. ---