docs: refresh the QA plan against the live task list (V-459)

The plan named 40 task numbers on 2026-08-01. Ten open QA tasks were missing
and two of the named ones had closed, so the 44-of-50 header was wrong twice
over.

- header is 42 of 50, and every open task now appears
- placed the ten unlisted QA tasks: 14, 248, 249, 250, 258, 283, 284, 285,
  286, 323
- new Operations sitting for 249 and 250, and a Query sources sitting for
  258 and 286
- 14 and 284 join housekeeping: both are gated on something unbuilt
- dropped the 317 and 354 rows, closed 01-08-2026, with one line saying what
  landed
- 319's gate recalibration is done; what is left is re-deriving QueryMinMargin
- 323 is down to the 60s startup timeout arm after PR #90
- new "Not this repo" section for 358 (Hexis) and 362 (training workspace)
- router latency is ~27x, not 90x; the 2.7s p50 was contention, not the model

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NJYcaBiuny9UGSpFweQVQ1
This commit is contained in:
2026-08-02 13:56:02 +04:00
parent 5c05163266
commit a1fdfccd61
+76 -24
View File
@@ -1,13 +1,18 @@
# QA plan: checking Maven properly
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
*Last verified: 2026-08-02 @ 5c05163. Living doc: correct it in place, do not append.*
Written 2026-08-01, after the 35-PR stack landed and the box came back up.
Refreshed 2026-08-02 against the live list, after PRs #85-#90.
44 of the 50 open Vikunja tasks are `QA:` tasks. They are verification work, not
42 of the 50 open Vikunja tasks are `QA:` tasks. They are verification work, not
build work. Most sat unverifiable while Maven was down for 11 days. That
blocker is gone.
The plan as written on 2026-08-01 named 40 task numbers. Ten open `QA:` tasks were
missing and two of the named ones had closed. Every open task now appears below,
the eight non-QA ones in the last two sections.
This plan orders them by what unblocks what. Do sessions 1 and 2 first. Almost everything
downstream assumes the voice loop works, and nobody has confirmed that since
the redeploy.
@@ -56,25 +61,25 @@ session quality), **321** steps 3-5 (quiet mode), **288** (STT fixtures).
`snoozed`. Say `потом` again with nothing pending: it must route as an
ordinary utterance, not be swallowed.
8. Wait for the water nudge, then say `выпил воды`. Expect the ordinary fact
reply and nothing extra — she must not congratulate you. Check
reply and nothing extra. She must not congratulate you. Check
`/notifications`: the row reads `acted`. Then trigger another nudge and say
`готово`; expect `отлично, отметила.` and the same outcome.
`готово`. Expect `отлично, отметила.` and the same outcome.
9. Note anything where she is slow, cuts off, or talks over herself. That is
287's whole content and it has no written acceptance criteria yet.
**319 is fixed** (01-08-2026). Single-word Russian utterances no longer come
**319's single-token bug is fixed** (01-08-2026). Single-word Russian utterances no longer come
back as `не совсем поняла — можешь переформулировать?`. `привет` and `поужинал`
both pass now: `thinSingleToken` spares social singles and any token carrying a
verb ending, and only thins a bare nominal like `вода`. If a one-word utterance
still gets clarified during the smoke test, that is a new case for the lexicon,
not the old bug.
verb ending, and only thins a bare nominal like `вода`. A one-word utterance that
still gets clarified in this session is a new case for the lexicon, not the old bug.
---
## Session 2: measurement (half a day, mostly waiting)
Closes or advances: **320** items 2-4, **278** (make the eval lab routine),
**319** (gate recalibration).
Closes or advances: **320** items 2-4, **278** (make the eval lab routine).
Also **248** (memory evaluation), **319** (the margin gate) and **323** (the
startup timeout arm).
The resident llama-server cannot be reached by the eval harness. It binds
`--host 127.0.0.1 --port 0` inside the container, so the port is kernel-assigned
@@ -99,14 +104,25 @@ make eval-recall
A large miss against 72.7% means the deploy differs from the bench harness.
Two things to decide while the numbers are in front of you:
Then three things to decide while the numbers are in front of you:
- **319's gate recalibration.** The single-token rule needs narrowing or
dropping. This needs your judgement, not a threshold sweep. The fixture and the
daemon disagree about what is correct on two of the three false clarifies.
- **319's gate recalibration is done.** 359 gave the LLM path a real confidence
signal. `thinSingleToken` was narrowed on 01-08-2026, and agenda questions moved
to stage 0. Missed clarify sits at 1 of 6 and false clarifies at 2. What is left
on 319 is item 2 point 2: re-derive `QueryMinMargin`, or show the distribution
the hand-picked `0.008` came from.
- **278's real ask** is making the eval lab routine rather than building it. It
is built. Decide whether it runs on a timer, on every merge, or on demand, and
the task can close.
- **248** is the memory evaluation loop. It ships, it writes notes, and it cannot
speak. `make eval-recall` covers the retrieval half. The open question is whether
a written evaluation nobody reads is worth the tick.
**323 is down to one check.** PR #90 covered the spawn path and took phraser
coverage to 76.9%. Only the 60s startup timeout arm is untested, because testing it
needs a `StartupTimeout` field on `Config` rather than a test-only hack. While you
are on the box, time a cold 1.7B load off spinning disk. If it runs near 60s, the
default is too tight and the field earns itself twice.
Item 4 of **320** needs a permission I do not have. Kill the `llama-server`
pid under `maven-mavend-1`, post a turn, and confirm it still completes
@@ -119,10 +135,20 @@ check that the failure floor catches a mid-session model death.
These need real use rather than a command, grouped by what one sitting covers.
**Morning and delivery** (**280**, **281**, **128**, **282**): open `/morning`,
walk the seven required behaviours, then check the four interruption outcomes
and the digest gap. **282** needs the `desk_active` script enabled on the desk
PC first, which is **15** and needs you at that machine.
**Morning and delivery** (**280**, **281**, **128**, **282**, **283**, **285**):
open `/morning`, walk the seven required behaviours, then check the four
interruption outcomes and the digest gap. **282** needs the `desk_active` script
enabled on the desk PC first, which is **15** and needs you at that machine.
**283** is the event intake envelope every reach shares, so a delivery check
exercises it whether you name it or not. **285** is not verification: the bridge
framework works and the remaining ask is more adapters. Decide which reach comes
next, or park it.
**Query sources** (**258**, **286**): ask her something the RSS feeds answer and
something only a ZIM answers, with the search block on. Live search leads and the
ZIMs are the fallback since 02-08-2026. A ZIM answer to a current-events question
means the search leg failed silently. **286**'s remaining half is doc and
git ingestion, which is build work, not a check.
**Tasks and calendar** (**129**, **130**, **127**, **126**, **246**): capture a
task by voice, confirm it lands, check prioritisation ordering is not nonsense.
@@ -138,17 +164,27 @@ concluding anything.
logged clean at boot. **276** is the degraded-mode suite, which means taking
siblings down on purpose. Worth doing while you are already in there.
**Operations** (**249**, **250**): these bite hardest if they are broken, and
nobody has pulled either lever on this box. Roll **249** forward and back once,
then swap `phraser.model_path` and confirm **250** reloads without a restart. Do
this sitting last, because both checks can take the box down.
---
## Housekeeping (one sitting, no box needed)
Four QA tasks will not close no matter how long they sit, because they are
Six QA tasks will not close no matter how long they sit, because they are
gated on something that does not exist:
- **125** zenmoney: needs a token you have not minted.
- **256** Home Assistant: needs HA configured.
- **257** Bluetooth: BLOCKED, no bluez on the box. Says so in the title.
- **288** STT golden audio: needs fixtures generated.
- **14** cold-start unlock: the `-wrapped-key-file` seam exists, the passkey to L3
half does not. `lockedAPI` was deleted as dead code in PR #50, so there is
nothing to verify.
- **284** replayable full-system simulator: nothing is built. This one is a
design task wearing a `QA:` prefix.
Relabel these so they stop reading as backlog. They are not verification work
that is pending, they are work that has not started.
@@ -171,16 +207,32 @@ Not QA. These are blocked on a decision or a credential only you have.
| 355 | Deploy the Hexis auth change. Was blocked on Maven being under construction, which it no longer is. The client half is vendored and wired. |
| 357 | Decide whether entity-existence validation is the permanent target guard or whether blessing lands in Nexus. |
| 275 | Hexis native API and MCP parity. |
| — | Decide on `-require-stepup`. Making it the default needs WebAuthn configured first, or it locks you out of your own admin surfaces. See **317**. |
| — | Three nginx sites bind wildcard `:80` (`acme.conf`, `matrix`, `panel`), so the ecosystem's bind-level protection is not in effect and `allow`/`deny` is carrying it alone. See **354**. |
| — | Decide on `-require-stepup`. Making it the default needs WebAuthn configured first, or it locks you out of your own admin surfaces. |
317 and 354 closed on 01-08-2026. The step-up gate now covers `POST /api/chat` and
`/routines`, and the nginx template is locked down with a `maven.<domain>` block for
mavweb. The `-require-stepup` default is still your call.
---
## Not this repo
Two open tasks sit on the Maven board and are not Maven work. Move them or note
where they land, so the board stops reading as 50 things Maven owes.
- **358** replace the rowid execution cursor with a real seq column. This is Hexis,
and it must land before any execution retention or pruning does.
- **362** mirror the router prompt reorder into the relabelling prompt. This is the
training workspace, enforced by `llm/check_prompt_parity.py` there, not here.
---
## Suggested order
1. Session 1. If the voice loop is broken, nothing else matters.
2. The `-require-stepup` and Kuma decisions. Five minutes, unblocks **317** fully
and **16**.
3. Session 2. The numbers tell you whether the router is worth its 90x latency.
2. The `-require-stepup` and Kuma decisions. Five minutes, and it unblocks **16**.
3. Session 2. The numbers tell you whether the router is worth its ~27× latency.
The 90× figure that stood here came from a 2.7s p50 that was contention, not
the model. Real p50 is 825ms against the classifier's 31ms.
4. Housekeeping. Cheap, and it makes the remaining backlog honest.
5. Session 3, split whichever way suits you.