diff --git a/docs/evals/2026-08-26-capability-baseline.md b/docs/evals/2026-08-26-capability-baseline.md new file mode 100644 index 0000000..1a46203 --- /dev/null +++ b/docs/evals/2026-08-26-capability-baseline.md @@ -0,0 +1,757 @@ +# Capability baseline: what the deployed Maven actually does + +*Measured 2026-08-26 against master `5cae33a` plus the uncommitted +`deploy/mavend.json` model switch. Frozen on the day.* + +The empirical half of `docs/plans/26-capability-ledger-and-baseline.md`. The +ledger at `docs/capabilities/ledger.yaml` says what should happen. This file +says what happened when it was asked. Every `verified` cell in the ledger cites +this file by path. + +Read `docs/capabilities/README.md` for how the artifacts regenerate. + +## In one paragraph + +26 of 146 v1 criteria pass, and **no capability passes all of its own**. Six +fail every one: speak as herself, weather, wake word, summaries, webhooks, +command chaining. What comes closest to working is what never asks the resident +model to write a Russian sentence, which is tasks, feeds, the network scan and +setting a one-shot reminder. The model that does write them, +`maven-instruct-b2-Q4_K_XL`, produces sentences that no longer hold together, +and the phrasing checks that would catch it run in the eval and not on the +outbound path. Four misroutes account for four more failures, including his own +name not being stored as a fact. + +`make test` is green throughout. It is also green with the four `TestONNX*` +measurements silently skipped, which is the second thing this file is about. + +## What was measured, and what that is worth + +Probes ran against the live `maven-mavend-1`, through `POST /api/chat` on +`127.0.0.1:9201`, which drives a real turn: the pre-route ladder, the stage 0 +grammars, the routing heads, the resident model, the query walk, the act path +and the phraser. Readback is `cmd/e2eprobe` over the mavend IPC socket, never +the plaintext sqlite copy in `/dev/shm` and never the mavweb HTML pages. + +**The number this baseline is attributable to is the deployed resident model, +`maven-instruct-b2-Q4_K_XL`, not the Qwen3-1.7B that `CLAUDE.md` names.** The +switch sits uncommitted in `deploy/mavend.json`. A baseline measures one model +on one config, so this file expires the moment either moves. + +## Row counts before anything was written + +Two readings. `mavend -wipe` without `-confirm-wipe` opens the encrypted file at +rest and prints every table. That is the eval's first artifact, and it reads the +23:50 seal, before this session touched anything. + +``` + ack_sends 71 memory_vectors 116 + delivery_attempts 16340 meta 2 + dialogue_sessions 1 notes 180 + digest_entries 0 nudges 5 + ecosystem_traces 86 presence_state 1 + events 17 proposed_routines 0 + facts 12984 reminders 95 + list_items 0 routing_labels 2 + routing_traces 976 + tasks 10 + tools 13 + TOTAL 30899 rows in 19 tables +``` + +Three numbers there are not visible over IPC. **`delivery_attempts` is 16,340**, +not the 200 `e2eprobe` returns. **`routing_traces` is 976**, which matches trace +id 977 exactly: the 14-day retention is working and has swept nothing yet. +**`routing_labels` is 2**, so the correction gesture at `POST /api/correct` has +been used twice in the life of the box, and it is the only supervised signal +this deployment collects. + +`digest_entries` and `proposed_routines` are both 0. + +The second reading is over IPC, before the first probe. + +| store | rows | +| --- | ---: | +| facts | 13018 | +| notes | 180 | +| reminders | 95 | +| tasks, live | 5 | +| nudges | 5 | +| tools | 13 | +| events | 34 | +| ecosystem traces | 86 | +| delivery attempts | >= 200 | +| decision traces | 0 | + +Two of these are findings on their own. + +**13,018 facts.** The architecture pass counted nine writers into `facts`, +including a fetch watermark and a tuning parameter. At this volume the table is +not a model of the owner, it is a log with a model of the owner somewhere in it. + +**Zero decision traces, at trace id 977.** Not a defect and not the 14-day +sweep. `internal/decision/ring.go` sets `ringSize = 25`: the arbitration record +is an in-memory daemon ring, and `storeAPI.TurnDecisions` says so outright, +returning `turn decisions not available via direct store API` (V-564). `mavend` +had restarted 49 minutes earlier, so the ring was empty. + +The consequence is worth stating plainly. `docs/spec.md` requires that every +turn write a decision trace naming the stage that decided, and every turn does. +**That record survives 25 turns and does not survive a restart.** 58 turns ran +during this baseline and 25 rows remain. Anything wanting to explain a turn from +yesterday cannot. + +`e2eprobe pending-reminders` failed with `ipc: unknown method`. The probe binary +is built from master and the deployed image is from 2026-08-19, so this is +build drift in the probe, not a defect in mavend. + +## The deployed stack is not one build + +| service | image | built | +| --- | --- | --- | +| maven-mavend-1 | `sha256:016e80eb` | 2026-08-19 | +| maven-mavweb-1 | `sha256:016e80eb` | 2026-08-19 | +| maven-mavsttd-1 | `sha256:fc338a2f` | 2026-08-08 | +| maven-mavttsd-1 | `sha256:fc338a2f` | 2026-08-08 | +| maven-mavpoll-1 | `sha256:fc338a2f` | 2026-08-08 | + +Three of five run a build eighteen days old. The 2026-08-13 audit found three on +a four-day-old image and the gap has widened, not closed. The outstanding +`docker compose up -d --force-recreate mavsttd mavttsd mavpoll` was deliberately +NOT run before this baseline: running it would have destroyed the observation. + +## The field probes + +25 multi-turn probes across the five vertical slices, drawn from the owner's +real week, the three failing conversations of the 2026-08-13 audit verbatim, the +five existing scenarios, and what `deploy/mavend.json` is configured for. 34 +turns. Full transcript at `docs/capabilities/out/field.transcript.tsv`, raw +output with readback at `docs/capabilities/out/field.raw.jsonl`. + +A field probe never sets a DoD verdict. It is the owner's week, not the spec. + +### The first run measured itself + +`mavweb` hardcodes one conversation id for the whole web reach, so a clarify +parked by one probe is still parked for the next. The first run had no reset +between probes. The park set at turn 8 reached turns 9 through 13, appending its +own question to five consecutive unrelated turns: + +``` +Q меня зовут Ками +A Сейчас 01:07. В какой день? + +Q я работаю в Тинькофф +A не знаю — не нашла у тебя такой записи. На какое время поставить напоминание? +``` + +That run is kept at `out/field.contaminated.jsonl` and scores nothing. The +behaviour it exposed is real and is a finding: **a parked clarify survives and +never resumes.** It outlives every intervening turn, including plain statements, +appending its question to their answers. At turn 14 it was dropped by the +`Прошлую просьбу отпускаю` path, not by being answered and not by `отмена`. It +never asks its original question again. + +`docs/spec.md` asks for the other half of this under "Ask instead of guessing": +a parked clarify should survive an interleaved unrelated turn **and resume**. +Half of that is what is broken. + +The scored run sends `отмена` before every probe. With nothing parked, that +answers `Такую команду я не знаю.` + +### What works + +Four paths answered correctly, fast, and in her own voice. + +| probe | reply | claimed by | +| --- | --- | --- | +| `field:action:capture-a-task` | `внесла в список: заказать корм для собаки`, then the full list ordered by deadline with `просрочено на 17 дней` on the overdue one | `tasks`, 0.01s | +| `field:query:feeds-on-request` | `вот что нового: mklinux-v7.0-mk2 released; [$] Old-school calendaring…` | `feeds`, 1.67s | +| `field:query:network-scan-as-prose` | `нашла 3 устройства, из них 2 с вебом, 2 с ssh. список записала.` | `network`, 27.01s | +| `field:reminder:create-and-persist` | `хорошо, напомню завтра в 09:00.` and reminder id 97 landed `pending` | stage 0, 0.29s | + +Every one of these four is answered by deterministic Go, not by the resident +model. The three claimed by a query source took the model out of the sentence +entirely. This is the shape of the whole run: **the parts of Maven that do not +ask the resident model to write Russian are the parts that work.** + +### The resident model writes broken Russian + +Ten turns reached the resident model for phrasing. What came back: + +``` +Q запиши: ключи от гаража лежат в верхнем ящике стола +A "Записала", "проверила" + +Q где лежат ключи от гаража? +A Ключи к ГАРАЖУ находятся ВЕРХНЕМ ЯСОЧКА СТОЛА. + +Q расскажи про Байкал +A Всего лишь одна фразаЕсли у тебя есть желание узнать больше о Байкале + или его влиянии как гидроэнергетического комплекс, посетите сайт + «Энциклопедии Кирилла и Мефодия» https://www.km.ru/encyclopedia/… + +Q что я говорил тебе про ремонт балкона? +A Уточни пожалуйста имя человека из которого ты будешь учиться или + посмотреть информацию. + +Q меня зовут Ками +A У тебя есть 35 сообщений. +``` + +The first leaks a raw parser fragment to the wire, quotes included. The second +found the right note and could not restate it: `ЯСОЧКА` is not a word and the +case is wrong. The third invents a URL. The fourth is not about anything the +owner said. The fifth is not an answer to anything. + +**Every one of these is attributable to `maven-instruct-b2-Q4_K_XL`.** The +2026-08-13 audit measured Qwen3-1.7B on the same box and its worst reply was +invented physics inside a well-formed Russian sentence. This is a different and +worse failure: the sentences themselves no longer hold together. + +### The persona breaks on the wire, in both directions + +`CheckFeminine`, `CheckAddress` and `CheckCringe` appear only under +`internal/phraser/eval`. They score the phrasing eval and do not run on the +outbound path. **Two of the 34 scored turns** used formal address on the wire. +The masculine reply is from the discarded contaminated run, so it scores nothing +and is shown because it is the same defect the audit found: + +| probe | on the wire | rule broken | +| --- | --- | --- | +| `field:fact:what-do-you-remember` (contaminated run) | `Я ничего о тебе и вы никуда таки не упомянул` | masculine `упомянул` and formal `вы`, both in one reply | +| `field:note:question-is-not-a-statement` | `Давайте проверим.` | formal `давайте` | +| `field:query:arithmetic-must-not-escape` | `Похоже у вас есть интерес к математике! … а вы можете продолжить? Или вам нужны советы по образованию?` | formal `вас`, `вы`, `вам` three times | + +The audit caught `Хорошо, сохранил` the same way. The check exists, passes in the +eval, and is not on the path the owner hears. + +### The act path extracts a preposition as the target + +Three unrelated utterances produced the same reply, in under 60ms, without the +resident model being asked: + +``` +Q перезапусти контейнер с базой A «с» — не знаю такой цели. назови её как в системе. +Q выключи свет в спальне A «с» — не знаю такой цели. назови её как в системе. +Q покажи что запущено в докере A «с» — не знаю такой цели. назови её как в системе. +``` + +The extracted target is `с`, a one-letter Russian preposition, for all three, +including the two that do not contain the word. The eco-trace readback shows no +new Hexis call for any of them. + +The refusal is correct behaviour for the wrong reason. Nexus holds no entities, +so no act naming a target can resolve, and `CLAUDE.md` requires free text to +resolve to a canonical entity id before any mutating call. It does refuse. But +it refuses having decided that the target is `с`, which means the extraction +step is not doing what its message claims, and the safety here rests on Nexus +being empty rather than on the resolution. + +### The Praxis lifecycle path + +``` +Q что требует моего внимания? A Нет ничего нового. +Q отметь первое как выполненное A Такую команду я не знаю. +``` + +Praxis holds nothing, so the mutation had no referent and the second turn never +reached `praxisItemAction.handle`. **The architecture pass's finding that the +path has no tier and no confirm turn is neither confirmed nor refuted by this +run.** It stays a code-reading claim. Seeding Praxis is what would settle it. + +### The reaches + +One reminder was set two minutes out, carrying the marker `MVNPROBE`, and fired. + +| reach | result | evidence | +| --- | --- | --- | +| telegram | **sent** | attempt 16344, `reminder:ea93acef…`, 01:14:05 | +| ntfy | **failed**, then telegram took it in the same second | attempt 16343, same delivery group | +| voice | not exercised by this probe | no listener on the box | + +A failed delivery retrying into another reach works. The audit could not confirm +this; here ntfy failed and telegram carried the same delivery group immediately. + +The 200 most recent attempts, spanning 2026-08-16 to now: + +| channel | status | rows | +| --- | ---: | --- | +| voice | failed | 156 | +| telegram | sent | 28 | +| drop | dropped | 12 | +| ntfy | failed | 4 | + +**149 of those 156 voice failures are one rule, `netdata_critical`.** It retried +once a minute into a reach that has never had a listener, for about three hours: +151 rows from `2026-08-16T22:17:20` to `2026-08-17T01:19:20`, then it stopped. +A bounded burst, not a storm still running. It is the same shape as the ntfy +failure the audit watched run until 03:05. The audit closed V-715 on the +reminder path retrying correctly, and the nudge path is a separate path. + +`ntfy.disabled` is `true` in the config and an ntfy attempt row was still +written and still failed. Whatever `disabled` turns off, it is not the attempt. + +ntfy was deliberately not enabled for this baseline. `docs/spec.md` makes a +write-scoped `NTFY_TOKEN` a precondition of re-enabling, and the owner confirmed +that stands. Its reach verdict is `blocked`, reason `configuration missing`. + +### The fact store is not a model of him + +13,041 facts. By key: + +| key | rows | writer | +| --- | ---: | --- | +| `desk_active` | 12274 | `infer:hyprland` | +| `page_heartbeat` | 491 | `infer:heartbeat` | +| `netdata_alarm` | 111 | `poll:netdata` | +| `rss:latest:lwn` | 69 | `poll:rss` | +| everything else | 96 | mixed | + +By source, **56 facts in 13,041 came from him**: 48 `tap:voice` and 8 +`tap:text`. `desk_active` alone is 94.1%, written every 65 seconds. + +`docs/spec.md` files facts under "memory, model of you". At this ratio the table +is a desk-presence log with a model of him in the noise, and the four-row Memory +cluster having no living doc is the reason nobody decided otherwise. + +And the probe that should have added to those 56 did not: + +``` +Q меня зовут Ками +A У тебя есть 35 сообщений. +``` + +No facts row was written. The 20 most recent facts after the turn are all +`desk_active` and `netdata_alarm`. A `name` fact **does** exist, id 506, written +`2026-08-01` from `tap:voice`, so the path worked once. The audit's finding that +she confirms a write that never happened is unchanged, and this run shows the +same failure without even the false confirmation. + +`field:fact:supersede` never got as far as superseding. All three turns, +including the two plain statements, were claimed by the `personal` query source +and answered `не знаю — не нашла у тебя такой записи`. **Statements are being +routed as questions.** + +## Routing: what the decision traces show + +Claim ladders are in `docs/capabilities/out/field.raw.jsonl` under +`readback.decisions`. 34 turns ran and every one returned a trace id. **22 +ladders were read back**, because the decision ring holds 25 and the probes' +own `отмена` resets consumed the rest. Of the 22, **15 reached the route stage** +and **11 reached the query stage**. The other 7 were decided at pre-route or by +a stage 0 grammar. + +**The classifier answered zero turns.** All 15 traces that carry a classifier +claim read `never_asked`, with the reason `the routing heads answered` or `the +LLM router answered`. `CLAUDE.md` calls it the floor rather than dead code. It is +the floor, and this run never reached it. That is the correct outcome and it +also means this baseline says nothing about whether the floor still works. + +**Stage 0 decided six turns**, and every one of the six is among the correct +answers: `reminder-wakeword` twice, `agenda-query` twice, `task-capture`, +`narrative-query`. + +Four misroutes explain four of the failures outright. + +| utterance | routed to | should have been | decided by | +| --- | --- | --- | --- | +| `меня зовут Ками` | intent `chat` | `remember` | llm-router, score 1.0, after the heads declined at 0.569 | +| `что ты помнишь обо мне?` | intent `chat` | `query` | routing-heads, score 0.866 | +| `что требует моего внимания?` | intent `chat` | `query`, then Praxis | routing-heads, score 0.723 | +| `Самара` | intent `act` | the parked weather clarify | routing-heads thinned at 0.247, then `action:action-handler` | + +The first is why no fact was written for his own name. The second is the audit's +misroute, moved: the audit had it going to `remember`, and now it goes to +`chat`. Two of seven audit probes failed on this shape and the spec makes fixing +it a DoD criterion; it is not fixed, it is different. + +The third is why Praxis was never asked. `PraxisGrammars()` is the only path to +Praxis, the `praxis-attention` grammar declined with `pattern did not match`, +and an intent of `chat` never reaches a query source at all. + +The fourth is the weather follow-up. The audit saw `Самара` die with `Я тебя не +разобрала` and this run reproduces it exactly, with the cause visible: the bare +city name scored `act` and went to the action handler, so the parked weather +clarify was never offered it. + +Two more traces are worth naming. + +**`расскажи про Байкал` was taken by the stage 0 `narrative-query` grammar** and +routed to `query:memory`, which answered from notes and invented a URL. Kiwix +was `never_asked`. The spec's DoD says a Russian question lands on the Russian +book; a stage 0 grammar takes the turn before the question can reach one. + +**`сколько будет два плюс два?` reached external search.** The stage 0 +`arithmetic-query` grammar declined with `pattern did not match`, the LLM router +scored `query` at 1.0, and the answer came back from SearXNG as a chatty +non-answer in formal Russian. `2 + 2` left the box. + +## The wake word has been deaf for seven hours + +Found while checking a verdict, not by a probe. `mavwaked` runs on workpc under +a user unit and was not running during this baseline: + +``` +Active: inactive (dead) since Tue 2026-08-25 18:21:35 +04; 7h ago + Duration: 2h 45min 33.323s + Process: ExecStart=/home/kami/.local/bin/mavwaked -device mavmic ... (code=exited, status=0/SUCCESS) + +Aug 25 18:21:35 bugmachine mavwaked[2535262]: arecord: pcm_read:2285: read error: No such device +``` + +The microphone went away, `arecord` stopped, and `mavwaked` **exited zero**. +systemd read a clean exit and did not restart it. Nothing on either box noticed, +and nothing would have: the always-on listener going silent looks exactly like +the always-on listener having nothing to report. + +`docs/spec.md` says under Hearing that no capture client ships (V-514). **A +capture client does ship and it is `mavwaked`**: it spawns `arecord` for 16kHz +mono PCM, runs silero VAD and the wake head, and sends `PushToTalk` frames to +mavend's voice port. What is absent is `mavenclient`. That line in the spec is +out of date. + +## The speech path was reachable and was not probed + +`POST /api/ptt` is registered on the same `mavweb` mux this baseline drove 58 +turns through (`cmd/mavweb/main.go:241`) and proxies raw PCM onto mavend's voice +port (`cmd/mavweb/voiceproxy.go:47`). It takes audio and runs a real +transcribe, route, reply turn. + +**Nothing in this run posted audio to it.** Every speech and senses criterion +therefore reads `untested`, reason `scenario missing`, and not `deployment +missing`: the deployment is present and the probe was never written. Recording +that as a deployment problem would have sent the next session to the wrong file. + +Two facts about that path were established while correcting it: + +- **Both workstation endpoints refuse.** `192.168.1.105:8080` and `:8081` return + no HTTP status from workpc. `workstation.model_disabled` is also `true`. So + mavend's `stt.Pair` has already fallen to the `mavsttd` floor, and the floor + is the only transcriber in service right now. +- **`internal/ttsnorm` is compiled into `mavend`, not `mavttsd`.** + `ttsnorm.Speakable` runs on the voice reply path and on nudge text. The + deployed mavend carries it on the 2026-08-19 image. + +## What the run does not establish + +- **The voice reach was not exercised.** No listener runs on workpc, so + `voicesink` records `blocked, no listener` and nothing here tests it. 156 of + the last 200 delivery attempts are that reach failing. +- **The Praxis lifecycle gate is untested.** Praxis holds nothing, so the + mutation had no referent. The architecture pass's claim that + `praxisItemAction.handle` calls through with no tier stands as a code reading. +- **The `mavsttd` and `mavttsd` arms were not exercised**, though they were + reachable. See the section above: no audio probe was written. +- **Telegram arrival was not confirmed by the owner.** The attempt row says + `sent`; that the message appeared on his phone is not in this file. +- **The 14-day routing-trace retention was not exercised.** That retention + covers the traces `correct` writes against, which are a table. The decision + ring measured above is a different thing and is capped at 25 turns in memory. + +## Row counts after the field run + +| store | before | after | delta | +| --- | ---: | ---: | ---: | +| facts | 13018 | 13041 | +23 | +| notes | 180 | 185 | +5 | +| reminders | 95 | 98 | +3 | +| tasks, live | 5 | 6 | +1 | +| decision ring | 0 | 25 | ring full | +| events | 34 | 62 | +28 | +| ecosystem traces | 86 | 115 | +29 | +| nudges | 5 | 5 | 0 | +| tools | 13 | 13 | 0 | + +**All 23 new facts are ambient.** 34 probe turns, including one that stated the +owner's name and two that stated where he works, added nothing to the fact +store. The +23 is `desk_active` and `page_heartbeat` continuing at their own +rate through the seven minutes the run took. + +The +5 notes are four RSS items and one captured note. That note is stored as +`запиши: ключи от гаража лежат в верхнем ящике стола`, imperative prefix +included, so the note text is the command rather than the content. + +The +3 reminders are two identical `позвонить в клинику` rows, one from each +run, and the `MVNPROBE` delivery probe. + +## Configuration and deployment, read directly + +These criteria are not observable through a turn. Chasing them through +`POST /api/chat` would measure the router instead, so they were read from the +config, the container and the startup log. + +### The web UI has no notes page and no facts page + +Fourteen pages answer 200, the slowest in 28ms. + +`/`, `/dash`, `/history`, `/trace`, `/notifications`, `/reminders`, `/morning`, +`/events`, `/tasks`, `/chat`, `/ecosystem`, `/tools`, `/routines`, `/models`. + +`docs/spec.md` requires a page for every capability with a surface and names +four: reminders, notes, tasks, facts. **`/notes` and `/facts` do not exist**, and +`cmd/mavweb/main.go` registers no handler for either. The 2026-08-13 audit's +"all ten pages answered 200" was true of the ten that exist. + +This settles the open half of the Notes DoD as well. A note cannot be deleted +from the web UI because there is no page from which to delete one (V-494). + +`/auth/passkey` returns 404, consistent with WebAuthn being unconfigured. + +### Every step-up gate is fail-open + +`mavweb` says so itself, at startup, unprompted: + +``` +SECURITY WARNING: step-up verification is DISABLED (-webauthn-origin/-webauthn-rpid unset). +These surfaces are UNGUARDED: + POST /tools defines arbitrary argv via name+cmd, which internal/tool then EXECUTES + POST /routines accepting schedules recurring firing + POST /models chooses the resident model that routes and words every turn + POST /api/revert voids the latest fact for a key + POST /api/chat reaches the router, the LLM and, through applyAction, the act path + POST /api/ptt the same, from audio +``` + +Six surfaces, and this baseline drove 58 turns through the fifth of them without +authenticating. V-683, unchanged and now measured rather than read. + +### What is configured correctly + +| criterion | evidence | +| --- | --- | +| the voice port stays on homesrv loopback | `docker port` maps `9100/tcp -> 127.0.0.1:9110`. `voice.bind` is `0.0.0.0:9100`, which is the container's own namespace; the publish is what makes it loopback. | +| the database is encrypted at rest, working copy in tmpfs | `/var/lib/maven/maven.db.enc`, 9.2 MB, mode 0600. Plaintext copy in `/dev/shm`, which is tmpfs. | +| the embedder loads at 384 dimensions with the marker check passing | `voice: onnx embedder loaded (384 dim)`, then `voice: embedder marker ok (model_quantized@384/tok2)` | +| the routing heads load from their own file | `router_heads.onnx`, not `model_path`. Refused at config load since V-692. | + +### What is not configured at all + +**There is no `weather` block in `deploy/mavend.json`.** Not a wrong value, not +a disabled flag: the key is absent, so the provider loads as a stub. The audit +found this and it has not changed. `какая сейчас погода?` answers `для какого +города?` and then cannot use the answer, which is the same two-turn failure the +audit recorded, reproduced here with its routing cause visible above. + +### Tests and analyzers + +`make test` is **green** across `./internal/...` and `./cmd/...`, with `-race` +and `-coverprofile`. `cmd/mavend` took 223.8s at 70.0% coverage and +`internal/store` 81.0s at 70.8%. + +**It does not set `MAVEN_ONNX_LIB`.** `Makefile:196` is the whole recipe and the +variable is not in it. Only `make t` and the four `eval-*` targets set it. So +every `TestONNX*` measurement self-skips inside `make test`, and the run prints +`ok` anyway, which is the failure mode `CLAUDE.md` names in as many words. + +Run under the `test` target's own environment, one package: + +``` +--- SKIP: TestONNXPersonalBoundaryStratified (0.00s) +--- SKIP: TestONNXPersonalBoundary (0.00s) +--- SKIP: TestONNXPersonalBoundaryFourFold (0.00s) +--- SKIP: TestONNXPersonalBoundarySemanticGroupHoldout (0.00s) +--- SKIP: TestONNXPersonalBoundaryChallenge (0.00s) +--- SKIP: TestONNXPersonalBoundaryPostRetuneChallenge (0.00s) +--- SKIP: TestONNXPersonalBoundaryLatency (0.00s) +--- SKIP: TestONNXPersonalBoundaryFrozenHeadMatchesCorpusFit (0.00s) +PASS +ok github.com/kami/maven/cmd/mavend 1.045s +``` + +Grepping the `make test` output for `skip` returns nothing, because `go test` +prints no SKIP line without `-v`. A green run and a run where the measurements +never executed are the same eight characters. + +The measurements do pass when they are given the library. Under `make t`, which +sets it and does not use coverage: + +``` +--- PASS: TestONNXPersonalBoundary (5.68s) +--- PASS: TestONNXPersonalBoundaryFourFold (26.30s) +--- PASS: TestONNXPersonalBoundarySemanticGroupHoldout (37.97s) +--- PASS: TestONNXPersonalBoundaryChallenge (5.23s) +--- PASS: TestONNXPersonalBoundaryPostRetuneChallenge (5.35s) +--- PASS: TestONNXPersonalBoundaryLatency (5.17s) +--- PASS: TestONNXPersonalBoundaryFrozenHeadMatchesCorpusFit (13.45s) +--- SKIP: TestONNXPersonalBoundaryStratified (0.00s) +``` + +The latency test V-718 describes as failing only under coverage passed here in +5.17s, in a run that carries no coverage. **This baseline does not reproduce +V-718 and does not refute it.** The one invocation that does use coverage, +`make test`, is the one where this test self-skips. + +**`make analyze` does not pass.** `staticcheck` and `deadcode` produced nothing +over their baselines. `govulncheck` reports seven vulnerabilities, every one in +the standard library, every one fixed in `go1.25.13`, and the vendored toolchain +is `go1.25.12`: + +``` +Your code is affected by 7 vulnerabilities from the Go standard library. +make: *** [Makefile:119: vuln] Error 3 +``` + +Reachable ones include `internal/kiwix/client.go:165` through +`encoding/xml`, `internal/netaddr/netaddr.go:228` through `encoding/asn1`, and +`internal/vision/vision.go:277` through `net/http`. This is a toolchain bump, +not a code fix. + +**This is the point the spec was written to make**, and the suite made it twice +over. It is green while 50 of 146 v1 criteria fail on the running box. It is +also green while the four measurements it is supposed to carry are not running +at all. A green suite has never been evidence that a capability works. + +`mavwaked` was not queried for its build. It runs on workpc under systemd, +outside this stack, and the measuring box could not reach it. + +## Undesigned in v1, by inspection + +Nine capabilities have no design. Their criteria were settled by reading the +deployed artifacts, not by probing, and three readings changed what the spec +says about them. + +**Email and calendar are recorded decisions, not open questions.** +`docker-compose.yml:145` and `:175` each carry a commented service with the +reasoning beside it. The mail block names the blocker (no IMAP account), the +security shape (password from a file, core never sees it, nothing there can +create a reminder) and the steps to enable it. It also records what triage +means: task extraction into candidates he reviews on `/tasks`, explicitly not +the acting variant. The calendar block says outright that mavcaldav "was built, +listed in `make build`, and deployed nowhere, which is the worst of the three +states, this block records the decision instead", then writes out what the +absence costs. Both capabilities' "in compose, or its absence is deliberate and +recorded" criterion **passes on the second branch**. + +**The box has no inbound HTTP intake at all.** +`internal/delivery/telegramsink/intake.go` long-polls `getUpdates` outbound, +precisely because nothing can connect inward and it reaches `api.telegram.org` +through a socks relay. It is not "Telegram's own inbound webhook". + +**Command chaining is half built and live.** `voice.llm_router` is `true` on the +deployed box, the router's grammar contract already returns an array of actions +for a compound utterance, and `parseActions` builds all of them. Nothing +dispatches past the first. The seam is `internal/router/llmrouter.go`, not a +missing package. + +**Learning from mistakes has its stores.** The `nudges` table holds an outcome +per nudge, `store.RecentOutcomes` reads the last N per rule, and +`loop.TuneCooldown` turns a high ignored rate into a longer cooldown. Route +repair is stored twice, as a `routing_labels` row and as a classifier example, +and there are two correction paths, not one: `POST /api/correct` and the spoken +`cmd/mavend/repair.go` (V-455). `routing_labels` held 2 rows before the wipe. +What has no store is a corrected phrasing. + +What genuinely does not exist: summaries as a requestable capability, webhooks in +either direction, any style-learning store, and any document relating the three +schedulers to each other. + +## The tally + +146 v1 criteria across 46 capabilities. The five deferred capabilities keep +their ten criteria at `untested`, reason `deferred past v1`, and were not +probed. + +| verdict | criteria | +| --- | ---: | +| pass | 26 | +| fail | 50 | +| blocked | 15 | +| untested | 51 | +| unknown | 4 | + +`pass` was only ever available to `live` evidence. `untested` is the largest +bucket and most of it is honest scope: a wall-clock trigger a probe cannot +reach, a workpc daemon the measuring box cannot see, an external service that +would have to be taken down on purpose. + +Per capability, in `docs/capabilities/ledger.yaml`. Nothing here is ranked. +Ranking is session 2's job and ranking these numbers without the implementation +mapping behind them would be the reverse of the causal order the plan sets out. + +## What this baseline expires on + +Any of these invalidates every number above. + +- The resident model moving off `maven-instruct-b2-Q4_K_XL`. +- `deploy/mavend.json` changing, including committing the model switch that is + currently uncommitted. +- `docker compose up -d --force-recreate mavsttd mavttsd mavpoll`, which is + outstanding from 2026-08-13 and would close the image drift measured here. +- Seeding Nexus, which would move every `blocked, external dependency + unavailable` verdict on the act path. +- A voice listener appearing on workpc. +- `systemctl --user restart mavwaked` on workpc, once the microphone is back. + The wake word and hearing verdicts are all measured against a dead process. + +## The store was wiped afterwards + +The plan calls for it and the owner confirmed it, knowing the count. Two copies +of the encrypted file were taken first, both outside the repo on homesrv, so the +decision is reversible: + +- `~/maven-preswipe-2026-08-26.db.enc`, the 23:50 seal, before this session. +- `~/maven-preswipe-final-2026-08-26.db.enc`, taken after stopping mavend. + +``` +wiped. 30899 rows gone, the schema is intact, mavend knows nobody. +``` + +Config, models, passkeys and the encryption key are files and were not touched. +Forty seconds after the restart the store held 1 fact, 4 notes, 4 events and 12 +tools: `desk_active` had already fired once, the RSS poller had run, and the tool +registry had re-seeded itself. + +**Every measurement in this file is now unreproducible against the same data.** +That is what a frozen eval is for. + +## Where the raw evidence is + +| file | what it holds | +| --- | --- | +| `docs/capabilities/out/field.transcript.tsv` | 34 turns: probe id, utterance, reply, claiming source, seconds | +| `docs/capabilities/out/field.raw.jsonl` | the same, plus every readback, including the full claim ladder per turn | +| `docs/capabilities/out/field.contaminated.jsonl` | the first run, which measured itself. Kept because the leak is a finding | +| `docs/capabilities/out/counts_after_field.json` | store counts after the probes | +| `docs/capabilities/out/counts_after_wipe.json` | store counts forty seconds after the restart | + +**The transcript carries his real data**: the task list, the note captured +during the run, and the utterances the probes spoke. The store it came from no +longer exists, so this is now the only record of those rows. It sits in the repo +under the same rule as `2026-08-07-week-of-usage-transcript.md`: the transcript +is the evidence and is not summarised anywhere else. + +## How the verdicts were checked + +Every verdict was audited by an independent pass told to refute it, one auditor +per spec section, with `pass` attacked hardest. It found real errors and they +are corrected above rather than argued with. The ones worth naming, because the +same mistakes are easy to repeat: + +- **`make test` was scored `pass` on the `MAVEN_ONNX_LIB` clause.** The recipe + does not set it. Grepping the output for `skip` returned nothing, which is + what a self-skipping test looks like without `-v`. This was the exact trap + `CLAUDE.md` describes, walked into while measuring whether other things had + walked into it. +- **"eleven consecutive turns" was wrong.** The park reached five, turns 9 + through 13, and was released by a path other than `отмена`. The wrong count + had already propagated into two other files. +- **"25 traces" was wrong.** 22 were read back, 15 reached the route stage and + 11 reached the query stage. Claims of the form "in every trace" were false for + the seven turns that never routed. +- **The voice retries were called ongoing.** They are a bounded three-hour burst + on 2026-08-16, 151 rows, then silence. +- **`POST /api/ptt` was called unreachable.** It is on the same mux this + baseline used 58 times. Four speech criteria were filed as `deployment + missing` when the deployment is present and the probe was never written. +- **`internal/worker` was read for Hexis calls.** It is the speech offload wire. + Digestion is `cmd/mavend/tick_digest.go` and `internal/loop`. +- **The masculine reply came from the discarded run**, not from the 34 scored + turns. + +Sixteen verdicts carried an evidence path pointing at a section of this file +that did not exist. `build_ledger.py` now refuses to build when an evidence path +or a `§` heading does not resolve. + +The corrections moved the tally by nine: four passes withdrawn, and five +criteria that had been filed `untested` turned out to be already settled.