docs: session 3 results and the query-source findings (V-459)
This commit is contained in:
+149
-20
@@ -19,6 +19,31 @@ the redeploy.
|
||||
|
||||
---
|
||||
|
||||
## What the 02-08-2026 run found
|
||||
|
||||
Sessions 1 and 2 ran, and so did most of session 3. Read these four before
|
||||
picking anything up.
|
||||
|
||||
- **470: a question writes invented knowledge into memory.** Recall then serves
|
||||
it back. `что дальше?` lands on `IntentFact` and stores the model's answer as a
|
||||
`self` fact at confidence 1.00. Two junk rows then claimed seven unrelated
|
||||
world questions through recall, outranking the search leg. A question about the
|
||||
capital of Australia was answered `какая последняя версия языка Go?`. Two bad
|
||||
writes silently disabled world answering, with nothing logged.
|
||||
- **466: a pending clarify is global.** One unanswerable clarify swallowed the
|
||||
next three utterances from three separate sessions. With ntfy, telegram and
|
||||
voice all live, a clarify raised on web chat eats the next telegram message.
|
||||
- **467: spoken task capture is dead.** The router calls the capture marker an
|
||||
`act`, and capture is reachable only from the `note` intent.
|
||||
- **The classifier baseline in this repo was wrong**, and it flattered the
|
||||
router. See session 2 and **464**.
|
||||
|
||||
Thirteen defects were filed on 02-08-2026: 462 through 474. Six tasks this plan
|
||||
had written off as blocked turned out to be ready to check. Two of the six passed
|
||||
the moment they were run.
|
||||
|
||||
---
|
||||
|
||||
## Before you start
|
||||
|
||||
Two things bite anyone running these checks on homesrv.
|
||||
@@ -41,7 +66,10 @@ Nothing here has been confirmed since the redeploy, and everything else assumes
|
||||
it works. Do this first.
|
||||
|
||||
Closes or advances: **44** (conversation), **45** (text chat), **287** (voice
|
||||
session quality), **321** steps 3-5 (quiet mode), **288** (STT fixtures).
|
||||
session quality), **321** steps 3-5 (quiet mode), **288** (STT golden audio).
|
||||
|
||||
**288 is not blocked.** The fixtures are committed under `cmd/mavsttd/testdata/`
|
||||
and `make test-stt-golden` runs today. This plan said otherwise until 02-08-2026.
|
||||
|
||||
Steps 1 and 3-6 were run on 02-08-2026 and pass. Steps 2 and 7-9 still need a
|
||||
person at the box, because they need a microphone or a nudge to arrive.
|
||||
@@ -213,26 +241,88 @@ exercises it whether you name it or not. **285** is not verification: the bridge
|
||||
framework works and the remaining ask is more adapters. Decide which reach comes
|
||||
next, or park it.
|
||||
|
||||
Run 02-08-2026. **280 is blocked.** No morning routine is configured (**472**).
|
||||
`morning.Item` also has no required-versus-optional field, so behaviour 1 cannot
|
||||
hold whatever you configure (**473**). **281's digest gap is closed**, and
|
||||
its presence rule passes on inspection. Three of its five items need traffic the
|
||||
box has not had. **283 is blocked**: nothing feeds the intake journal. **128
|
||||
found the worst defect of the whole session, see below.**
|
||||
|
||||
For **285**, two facts bear on the choice. Synapse is already running on this box
|
||||
and healthy, so a Matrix reach has a live target and needs no new service. And
|
||||
mavweb is already a PWA with a service worker, which 285 itself calls the highest
|
||||
value adapter left. Today's reaches are ntfy, telegram and voice.
|
||||
|
||||
**Query sources** (**258**, **286**): ask her something the RSS feeds answer and
|
||||
something only a ZIM answers, with the search block on. Live search leads and the
|
||||
ZIMs are the fallback since 02-08-2026. A ZIM answer to a current-events question
|
||||
means the search leg failed silently. **286**'s remaining half is doc and
|
||||
ZIMs are the fallback since 02-08-2026. **286**'s remaining half is doc and
|
||||
git ingestion, which is build work, not a check.
|
||||
|
||||
**Do not read `/trace` for this.** `/trace` is the nudge-rule trace: rule,
|
||||
severity, predicate, gate, selected. No query-source field exists anywhere in the
|
||||
codebase. The only evidence of which query source claimed a turn is the
|
||||
`voice: search:` and `voice: kiwix:` lines in `docker compose logs mavend`
|
||||
(`actions_query.go:589` and `:660`).
|
||||
|
||||
Run 02-08-2026, 20 turns. **Search leads and the personal boundary holds.** Every
|
||||
world question that reached the boundary was claimed by search. All three
|
||||
personal questions produced no search and no kiwix line at all.
|
||||
|
||||
The rest of this sitting went badly. **Kiwix has zero live coverage.** SearXNG
|
||||
returns four results for everything, including two invented nonsense terms. So
|
||||
`querySearch` always claims, and Kiwix is unreachable code as deployed. The ZIM
|
||||
half of the 02-08-2026 decision is unverified. A ZIM answer cannot signal a
|
||||
silent search failure, because a ZIM answer cannot happen.
|
||||
**Ordering defects** in feeds and calendar, plus 258 step 1's utterance not
|
||||
working: **474**. And the sitting independently found stage 2 of **470**.
|
||||
|
||||
**Tasks and calendar** (**129**, **130**, **127**, **126**, **246**): capture a
|
||||
task by voice, confirm it lands, check prioritisation ordering is not nonsense.
|
||||
**246** (mail reader) also exercises the `IngestMail` rung that moved to
|
||||
`AuthWrite` this morning.
|
||||
|
||||
Run 02-08-2026. **129 passes.** The page and the spoken answer agree on ordering.
|
||||
The undistinguished task carries no invented reason on either surface, which is
|
||||
the thing 129 asks for. **130 fails outright** and **127 half fails**:
|
||||
**467**, **469**. **246 cannot be run**: `mavmaild` is commented out in
|
||||
`docker-compose.yml` and there is no `email` block, so nothing in steps 4-13 is
|
||||
reachable. The `IngestMail` rung does sit at `AuthWrite`
|
||||
(`internal/auth/policy.go:96`, asserted in `auth_test.go:421`), verified by
|
||||
reading only.
|
||||
|
||||
**Routines and patterns** (**43**, **46**, **247**, **254**): these need history
|
||||
to detect against. If the database is thin after the outage, they may have
|
||||
nothing to propose, which is not a failure. Check `/routines` before
|
||||
concluding anything.
|
||||
|
||||
Run 02-08-2026. The answer is the middle case: **the detector ran and found
|
||||
nothing.** The tick loop is live, and `detectPatterns` is called unconditionally
|
||||
at `cmd/mavend/tick.go:227`. It has run about 25 times since the restart. It
|
||||
finds nothing because the events table is empty upstream of it. Rows land there
|
||||
only from `pattern.Extract` at fact-write time, and `Extract` requires the fact
|
||||
value to match a closed 7-action lexicon. All 200 facts on `/history` are
|
||||
`page_heartbeat`, `netdata_alarm`, `quiet_hours`, `name`, `service_down` and
|
||||
`рост`. Not one lexicon hit, so no event can exist, let alone the four one pair
|
||||
needs. **46 step 5 passes**: `/routines` renders `noticed 0` with the empty state
|
||||
and the hint string.
|
||||
|
||||
Two things block this sitting, and both are build work. The seeding recipe on
|
||||
**43** goes through `sqlite3` and cannot work. And `pattern.Detect` has no
|
||||
minimum-interval floor, so seeding by hand mints a permanent false routine
|
||||
(**468**). Do not try to seed a pattern with four fast chat turns.
|
||||
|
||||
**Ecosystem** (**272**, **273**, **276**): nexus, hexis and praxis are wired and
|
||||
logged clean at boot. **276** is the degraded-mode suite, which means taking
|
||||
siblings down on purpose. Worth doing while you are already in there.
|
||||
|
||||
Run 02-08-2026, read-only half. All three answer `/health` 200 and `/ecosystem`
|
||||
lists 18 Hexis capabilities with correct read-only and mutating badges. **272 and
|
||||
273 are blocked on empty data**, not on code. Nexus holds no entities, Praxis
|
||||
holds no attention items, and the Calls panel has never recorded a call. See
|
||||
**472**, and read its warning first. 273's trace fix has never been validated
|
||||
here. An empty Calls panel is exactly what the old bug looked like. The page is
|
||||
`/ecosystem`, not `/siblings`.
|
||||
|
||||
**Operations** (**249**, **250**): these bite hardest if they are broken, and
|
||||
nobody has pulled either lever on this box. Roll **249** forward and back once,
|
||||
then swap `phraser.model_path` and confirm **250** reloads without a restart. Do
|
||||
@@ -240,27 +330,66 @@ this sitting last, because both checks can take the box down.
|
||||
|
||||
---
|
||||
|
||||
## Housekeeping (one sitting, no box needed)
|
||||
## Housekeeping (done 02-08-2026, and this section was mostly wrong)
|
||||
|
||||
Six QA tasks will not close no matter how long they sit, because they are
|
||||
gated on something that does not exist:
|
||||
This section claimed eleven tasks were not verification work. **Three were not.
|
||||
The other eight are.** Every one of the eight has shipped, tested code behind it.
|
||||
The error ran one way: it wrote off work that is ready to check. Do not trust a
|
||||
"nothing is built" line in this plan without grepping for the package first.
|
||||
|
||||
- **125** zenmoney: needs a token you have not minted.
|
||||
- **256** Home Assistant: needs HA configured.
|
||||
- **257** Bluetooth: BLOCKED, no bluez on the box. Says so in the title.
|
||||
- **288** STT golden audio: needs fixtures generated.
|
||||
- **14** cold-start unlock: the `-wrapped-key-file` seam exists, the passkey to L3
|
||||
half does not. `lockedAPI` was deleted as dead code in PR #50, so there is
|
||||
nothing to verify.
|
||||
- **284** replayable full-system simulator: nothing is built. This one is a
|
||||
design task wearing a `QA:` prefix.
|
||||
Relabelled to `Blocked:`, claim verified:
|
||||
|
||||
Relabel these so they stop reading as backlog. They are not verification work
|
||||
that is pending, they are work that has not started.
|
||||
- **125** zenmoney. `internal/zenmoney/` ships and is tested against a fixture.
|
||||
`deploy/zenmoney.token` does not exist and the compose mount is commented out.
|
||||
One token unblocks it.
|
||||
- **256** Home Assistant. `internal/smarthome/` ships, the `smarthome` block sits
|
||||
in `deploy/mavend.json` at `enabled: false`, and 8123 and 1883 are closed.
|
||||
- **14** cold-start unlock. The seam is real at `cmd/mavend/main.go:128` and
|
||||
`internal/webauthn/prf.go` is in place. `lockedAPI` is gone, replaced by
|
||||
`Server.Check` in `internal/ipc/server.go`. Gated on an authenticator that
|
||||
implements the WebAuthn PRF extension, which is hardware, not code.
|
||||
|
||||
Same treatment for the five plan-only tasks (**251** MCP, **252** vision,
|
||||
**253** hearing, **255** speaker recognition, **259** crawler). A `QA:` prefix on
|
||||
a plan is misleading.
|
||||
Left alone, because the claim here was false:
|
||||
|
||||
- **284** simulator. `cmd/mavend/simulator_test.go`, three scenarios under
|
||||
`cmd/mavend/testdata/scenarios/`, and a `simulate` target at `Makefile:98`.
|
||||
**Run 02-08-2026: all three scenarios pass**, plus the determinism and
|
||||
backwards-step guards. One defect found, see below.
|
||||
- **288** STT golden audio. Four WAVs and `golden_v1.json` are committed under
|
||||
`cmd/mavsttd/testdata/`, the make targets exist, and `models/stt/ggml-small.bin`
|
||||
is on the box. Session 1 lists 288 as blocked on fixtures, which is wrong.
|
||||
**Run 02-08-2026: all four pass**, WER at or under ceiling with no drift.
|
||||
|
||||
| fixture | transcript | WER | ceiling |
|
||||
|---|---|---|---|
|
||||
| ru_reminder | `Напомни мне через час позвонить маме.` | 0.00 | 0.10 |
|
||||
| ru_fact | `А отметь, что я выпил воды.` | 0.20 | 0.25 |
|
||||
| ru_query | `Что у меня сегодня по календарю?` | 0.00 | 0.10 |
|
||||
| en_act | `Restart the web server and check the disk space.` | 0.00 | 0.10 |
|
||||
|
||||
That also settles a session 1 worry indirectly: whisper.cpp works on Vulkan
|
||||
after the redeploy. Only the mic and the wake path remain unproven.
|
||||
|
||||
**The simulator routes with an empty seed set.** Every `make simulate` run logs
|
||||
`loaded 0 seed examples from models/seeds`, seven times per scenario. The test
|
||||
runs from `cmd/mavend`, and the seed path is relative to the repo root. The
|
||||
scenarios still pass, which means they pass without the classifier having any
|
||||
seeds to match against. Whatever 284 is proving, it is not proving the routing
|
||||
the deploy runs. Fix the path before trusting a green simulator.
|
||||
- **257** Bluetooth. The bluez half is genuinely absent. The LAN-scan half shipped
|
||||
(`internal/netscan/`), and steps 1-9 run today. Only step 10 is Bluetooth, so
|
||||
relabelling the whole task would bury real pending work.
|
||||
- **251** MCP, **253** hearing, **259** crawler. All three ship
|
||||
(`internal/mcp/`, `internal/capture/`, `internal/crawl/`) with no external gate.
|
||||
Fully checkable. `259`'s step 1 wants no `crawl` block in `deploy/mavend.json`,
|
||||
and there is none, so it is already set up correctly.
|
||||
- **252** vision and **255** speaker recognition. Both ship. Each is blocked only
|
||||
on a model download: a vision gguf with mmproj, and a speaker embedding model.
|
||||
Neither is present under `/mnt/hdd1`. Their refusal-path steps run today.
|
||||
|
||||
So the honest split is three blocked on a credential or hardware, two blocked on
|
||||
a download, and six ready to check. That is roughly a session of real QA this
|
||||
plan had written off as backlog.
|
||||
|
||||
---
|
||||
|
||||
|
||||
Reference in New Issue
Block a user