What the deployed Maven did when it was asked, measured 2026-08-26 against
master 5cae33a plus the uncommitted deploy/mavend.json model switch. Frozen on
the day and not edited after it.
The number is attributable to the deployed resident model,
maven-instruct-b2-Q4_K_XL, not the Qwen3-1.7B that CLAUDE.md names. What comes
closest to working is what never asks that model to write a Russian sentence.
make test is green throughout. It is also green with the four TestONNX
measurements silently skipped, because the recipe does not set MAVEN_ONNX_LIB.
Every verdict was audited by an independent pass told to refute it, one auditor
per spec section, with pass attacked hardest. The corrections moved the tally by
nine: four passes withdrawn, five criteria filed untested turned out already
settled. The section on how the verdicts were checked names the mistakes, so the
next session does not have to trust that this one got it right.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
37 KiB
Capability baseline: what the deployed Maven actually does
Measured 2026-08-26 against master 5cae33a plus the uncommitted
deploy/mavend.json model switch. Frozen on the day.
The empirical half of docs/plans/26-capability-ledger-and-baseline.md. The
ledger at docs/capabilities/ledger.yaml says what should happen. This file
says what happened when it was asked. Every verified cell in the ledger cites
this file by path.
Read docs/capabilities/README.md for how the artifacts regenerate.
In one paragraph
26 of 146 v1 criteria pass, and no capability passes all of its own. Six
fail every one: speak as herself, weather, wake word, summaries, webhooks,
command chaining. What comes closest to working is what never asks the resident
model to write a Russian sentence, which is tasks, feeds, the network scan and
setting a one-shot reminder. The model that does write them,
maven-instruct-b2-Q4_K_XL, produces sentences that no longer hold together,
and the phrasing checks that would catch it run in the eval and not on the
outbound path. Four misroutes account for four more failures, including his own
name not being stored as a fact.
make test is green throughout. It is also green with the four TestONNX*
measurements silently skipped, which is the second thing this file is about.
What was measured, and what that is worth
Probes ran against the live maven-mavend-1, through POST /api/chat on
127.0.0.1:9201, which drives a real turn: the pre-route ladder, the stage 0
grammars, the routing heads, the resident model, the query walk, the act path
and the phraser. Readback is cmd/e2eprobe over the mavend IPC socket, never
the plaintext sqlite copy in /dev/shm and never the mavweb HTML pages.
The number this baseline is attributable to is the deployed resident model,
maven-instruct-b2-Q4_K_XL, not the Qwen3-1.7B that CLAUDE.md names. The
switch sits uncommitted in deploy/mavend.json. A baseline measures one model
on one config, so this file expires the moment either moves.
Row counts before anything was written
Two readings. mavend -wipe without -confirm-wipe opens the encrypted file at
rest and prints every table. That is the eval's first artifact, and it reads the
23:50 seal, before this session touched anything.
ack_sends 71 memory_vectors 116
delivery_attempts 16340 meta 2
dialogue_sessions 1 notes 180
digest_entries 0 nudges 5
ecosystem_traces 86 presence_state 1
events 17 proposed_routines 0
facts 12984 reminders 95
list_items 0 routing_labels 2
routing_traces 976
tasks 10
tools 13
TOTAL 30899 rows in 19 tables
Three numbers there are not visible over IPC. delivery_attempts is 16,340,
not the 200 e2eprobe returns. routing_traces is 976, which matches trace
id 977 exactly: the 14-day retention is working and has swept nothing yet.
routing_labels is 2, so the correction gesture at POST /api/correct has
been used twice in the life of the box, and it is the only supervised signal
this deployment collects.
digest_entries and proposed_routines are both 0.
The second reading is over IPC, before the first probe.
| store | rows |
|---|---|
| facts | 13018 |
| notes | 180 |
| reminders | 95 |
| tasks, live | 5 |
| nudges | 5 |
| tools | 13 |
| events | 34 |
| ecosystem traces | 86 |
| delivery attempts | >= 200 |
| decision traces | 0 |
Two of these are findings on their own.
13,018 facts. The architecture pass counted nine writers into facts,
including a fetch watermark and a tuning parameter. At this volume the table is
not a model of the owner, it is a log with a model of the owner somewhere in it.
Zero decision traces, at trace id 977. Not a defect and not the 14-day
sweep. internal/decision/ring.go sets ringSize = 25: the arbitration record
is an in-memory daemon ring, and storeAPI.TurnDecisions says so outright,
returning turn decisions not available via direct store API (V-564). mavend
had restarted 49 minutes earlier, so the ring was empty.
The consequence is worth stating plainly. docs/spec.md requires that every
turn write a decision trace naming the stage that decided, and every turn does.
That record survives 25 turns and does not survive a restart. 58 turns ran
during this baseline and 25 rows remain. Anything wanting to explain a turn from
yesterday cannot.
e2eprobe pending-reminders failed with ipc: unknown method. The probe binary
is built from master and the deployed image is from 2026-08-19, so this is
build drift in the probe, not a defect in mavend.
The deployed stack is not one build
| service | image | built |
|---|---|---|
| maven-mavend-1 | sha256:016e80eb |
2026-08-19 |
| maven-mavweb-1 | sha256:016e80eb |
2026-08-19 |
| maven-mavsttd-1 | sha256:fc338a2f |
2026-08-08 |
| maven-mavttsd-1 | sha256:fc338a2f |
2026-08-08 |
| maven-mavpoll-1 | sha256:fc338a2f |
2026-08-08 |
Three of five run a build eighteen days old. The 2026-08-13 audit found three on
a four-day-old image and the gap has widened, not closed. The outstanding
docker compose up -d --force-recreate mavsttd mavttsd mavpoll was deliberately
NOT run before this baseline: running it would have destroyed the observation.
The field probes
25 multi-turn probes across the five vertical slices, drawn from the owner's
real week, the three failing conversations of the 2026-08-13 audit verbatim, the
five existing scenarios, and what deploy/mavend.json is configured for. 34
turns. Full transcript at docs/capabilities/out/field.transcript.tsv, raw
output with readback at docs/capabilities/out/field.raw.jsonl.
A field probe never sets a DoD verdict. It is the owner's week, not the spec.
The first run measured itself
mavweb hardcodes one conversation id for the whole web reach, so a clarify
parked by one probe is still parked for the next. The first run had no reset
between probes. The park set at turn 8 reached turns 9 through 13, appending its
own question to five consecutive unrelated turns:
Q меня зовут Ками
A Сейчас 01:07. В какой день?
Q я работаю в Тинькофф
A не знаю — не нашла у тебя такой записи. На какое время поставить напоминание?
That run is kept at out/field.contaminated.jsonl and scores nothing. The
behaviour it exposed is real and is a finding: a parked clarify survives and
never resumes. It outlives every intervening turn, including plain statements,
appending its question to their answers. At turn 14 it was dropped by the
Прошлую просьбу отпускаю path, not by being answered and not by отмена. It
never asks its original question again.
docs/spec.md asks for the other half of this under "Ask instead of guessing":
a parked clarify should survive an interleaved unrelated turn and resume.
Half of that is what is broken.
The scored run sends отмена before every probe. With nothing parked, that
answers Такую команду я не знаю.
What works
Four paths answered correctly, fast, and in her own voice.
| probe | reply | claimed by |
|---|---|---|
field:action:capture-a-task |
внесла в список: заказать корм для собаки, then the full list ordered by deadline with просрочено на 17 дней on the overdue one |
tasks, 0.01s |
field:query:feeds-on-request |
вот что нового: mklinux-v7.0-mk2 released; [$] Old-school calendaring… |
feeds, 1.67s |
field:query:network-scan-as-prose |
нашла 3 устройства, из них 2 с вебом, 2 с ssh. список записала. |
network, 27.01s |
field:reminder:create-and-persist |
хорошо, напомню завтра в 09:00. and reminder id 97 landed pending |
stage 0, 0.29s |
Every one of these four is answered by deterministic Go, not by the resident model. The three claimed by a query source took the model out of the sentence entirely. This is the shape of the whole run: the parts of Maven that do not ask the resident model to write Russian are the parts that work.
The resident model writes broken Russian
Ten turns reached the resident model for phrasing. What came back:
Q запиши: ключи от гаража лежат в верхнем ящике стола
A "Записала", "проверила"
Q где лежат ключи от гаража?
A Ключи к ГАРАЖУ находятся ВЕРХНЕМ ЯСОЧКА СТОЛА.
Q расскажи про Байкал
A Всего лишь одна фразаЕсли у тебя есть желание узнать больше о Байкале
или его влиянии как гидроэнергетического комплекс, посетите сайт
«Энциклопедии Кирилла и Мефодия» https://www.km.ru/encyclopedia/…
Q что я говорил тебе про ремонт балкона?
A Уточни пожалуйста имя человека из которого ты будешь учиться или
посмотреть информацию.
Q меня зовут Ками
A У тебя есть 35 сообщений.
The first leaks a raw parser fragment to the wire, quotes included. The second
found the right note and could not restate it: ЯСОЧКА is not a word and the
case is wrong. The third invents a URL. The fourth is not about anything the
owner said. The fifth is not an answer to anything.
Every one of these is attributable to maven-instruct-b2-Q4_K_XL. The
2026-08-13 audit measured Qwen3-1.7B on the same box and its worst reply was
invented physics inside a well-formed Russian sentence. This is a different and
worse failure: the sentences themselves no longer hold together.
The persona breaks on the wire, in both directions
CheckFeminine, CheckAddress and CheckCringe appear only under
internal/phraser/eval. They score the phrasing eval and do not run on the
outbound path. Two of the 34 scored turns used formal address on the wire.
The masculine reply is from the discarded contaminated run, so it scores nothing
and is shown because it is the same defect the audit found:
| probe | on the wire | rule broken |
|---|---|---|
field:fact:what-do-you-remember (contaminated run) |
Я ничего о тебе и вы никуда таки не упомянул |
masculine упомянул and formal вы, both in one reply |
field:note:question-is-not-a-statement |
Давайте проверим. |
formal давайте |
field:query:arithmetic-must-not-escape |
Похоже у вас есть интерес к математике! … а вы можете продолжить? Или вам нужны советы по образованию? |
formal вас, вы, вам three times |
The audit caught Хорошо, сохранил the same way. The check exists, passes in the
eval, and is not on the path the owner hears.
The act path extracts a preposition as the target
Three unrelated utterances produced the same reply, in under 60ms, without the resident model being asked:
Q перезапусти контейнер с базой A «с» — не знаю такой цели. назови её как в системе.
Q выключи свет в спальне A «с» — не знаю такой цели. назови её как в системе.
Q покажи что запущено в докере A «с» — не знаю такой цели. назови её как в системе.
The extracted target is с, a one-letter Russian preposition, for all three,
including the two that do not contain the word. The eco-trace readback shows no
new Hexis call for any of them.
The refusal is correct behaviour for the wrong reason. Nexus holds no entities,
so no act naming a target can resolve, and CLAUDE.md requires free text to
resolve to a canonical entity id before any mutating call. It does refuse. But
it refuses having decided that the target is с, which means the extraction
step is not doing what its message claims, and the safety here rests on Nexus
being empty rather than on the resolution.
The Praxis lifecycle path
Q что требует моего внимания? A Нет ничего нового.
Q отметь первое как выполненное A Такую команду я не знаю.
Praxis holds nothing, so the mutation had no referent and the second turn never
reached praxisItemAction.handle. The architecture pass's finding that the
path has no tier and no confirm turn is neither confirmed nor refuted by this
run. It stays a code-reading claim. Seeding Praxis is what would settle it.
The reaches
One reminder was set two minutes out, carrying the marker MVNPROBE, and fired.
| reach | result | evidence |
|---|---|---|
| telegram | sent | attempt 16344, reminder:ea93acef…, 01:14:05 |
| ntfy | failed, then telegram took it in the same second | attempt 16343, same delivery group |
| voice | not exercised by this probe | no listener on the box |
A failed delivery retrying into another reach works. The audit could not confirm this; here ntfy failed and telegram carried the same delivery group immediately.
The 200 most recent attempts, spanning 2026-08-16 to now:
| channel | status | rows |
|---|---|---|
| voice | failed | 156 |
| telegram | sent | 28 |
| drop | dropped | 12 |
| ntfy | failed | 4 |
149 of those 156 voice failures are one rule, netdata_critical. It retried
once a minute into a reach that has never had a listener, for about three hours:
151 rows from 2026-08-16T22:17:20 to 2026-08-17T01:19:20, then it stopped.
A bounded burst, not a storm still running. It is the same shape as the ntfy
failure the audit watched run until 03:05. The audit closed V-715 on the
reminder path retrying correctly, and the nudge path is a separate path.
ntfy.disabled is true in the config and an ntfy attempt row was still
written and still failed. Whatever disabled turns off, it is not the attempt.
ntfy was deliberately not enabled for this baseline. docs/spec.md makes a
write-scoped NTFY_TOKEN a precondition of re-enabling, and the owner confirmed
that stands. Its reach verdict is blocked, reason configuration missing.
The fact store is not a model of him
13,041 facts. By key:
| key | rows | writer |
|---|---|---|
desk_active |
12274 | infer:hyprland |
page_heartbeat |
491 | infer:heartbeat |
netdata_alarm |
111 | poll:netdata |
rss:latest:lwn |
69 | poll:rss |
| everything else | 96 | mixed |
By source, 56 facts in 13,041 came from him: 48 tap:voice and 8
tap:text. desk_active alone is 94.1%, written every 65 seconds.
docs/spec.md files facts under "memory, model of you". At this ratio the table
is a desk-presence log with a model of him in the noise, and the four-row Memory
cluster having no living doc is the reason nobody decided otherwise.
And the probe that should have added to those 56 did not:
Q меня зовут Ками
A У тебя есть 35 сообщений.
No facts row was written. The 20 most recent facts after the turn are all
desk_active and netdata_alarm. A name fact does exist, id 506, written
2026-08-01 from tap:voice, so the path worked once. The audit's finding that
she confirms a write that never happened is unchanged, and this run shows the
same failure without even the false confirmation.
field:fact:supersede never got as far as superseding. All three turns,
including the two plain statements, were claimed by the personal query source
and answered не знаю — не нашла у тебя такой записи. Statements are being
routed as questions.
Routing: what the decision traces show
Claim ladders are in docs/capabilities/out/field.raw.jsonl under
readback.decisions. 34 turns ran and every one returned a trace id. 22
ladders were read back, because the decision ring holds 25 and the probes'
own отмена resets consumed the rest. Of the 22, 15 reached the route stage
and 11 reached the query stage. The other 7 were decided at pre-route or by
a stage 0 grammar.
The classifier answered zero turns. All 15 traces that carry a classifier
claim read never_asked, with the reason the routing heads answered or the LLM router answered. CLAUDE.md calls it the floor rather than dead code. It is
the floor, and this run never reached it. That is the correct outcome and it
also means this baseline says nothing about whether the floor still works.
Stage 0 decided six turns, and every one of the six is among the correct
answers: reminder-wakeword twice, agenda-query twice, task-capture,
narrative-query.
Four misroutes explain four of the failures outright.
| utterance | routed to | should have been | decided by |
|---|---|---|---|
меня зовут Ками |
intent chat |
remember |
llm-router, score 1.0, after the heads declined at 0.569 |
что ты помнишь обо мне? |
intent chat |
query |
routing-heads, score 0.866 |
что требует моего внимания? |
intent chat |
query, then Praxis |
routing-heads, score 0.723 |
Самара |
intent act |
the parked weather clarify | routing-heads thinned at 0.247, then action:action-handler |
The first is why no fact was written for his own name. The second is the audit's
misroute, moved: the audit had it going to remember, and now it goes to
chat. Two of seven audit probes failed on this shape and the spec makes fixing
it a DoD criterion; it is not fixed, it is different.
The third is why Praxis was never asked. PraxisGrammars() is the only path to
Praxis, the praxis-attention grammar declined with pattern did not match,
and an intent of chat never reaches a query source at all.
The fourth is the weather follow-up. The audit saw Самара die with Я тебя не разобрала and this run reproduces it exactly, with the cause visible: the bare
city name scored act and went to the action handler, so the parked weather
clarify was never offered it.
Two more traces are worth naming.
расскажи про Байкал was taken by the stage 0 narrative-query grammar and
routed to query:memory, which answered from notes and invented a URL. Kiwix
was never_asked. The spec's DoD says a Russian question lands on the Russian
book; a stage 0 grammar takes the turn before the question can reach one.
сколько будет два плюс два? reached external search. The stage 0
arithmetic-query grammar declined with pattern did not match, the LLM router
scored query at 1.0, and the answer came back from SearXNG as a chatty
non-answer in formal Russian. 2 + 2 left the box.
The wake word has been deaf for seven hours
Found while checking a verdict, not by a probe. mavwaked runs on workpc under
a user unit and was not running during this baseline:
Active: inactive (dead) since Tue 2026-08-25 18:21:35 +04; 7h ago
Duration: 2h 45min 33.323s
Process: ExecStart=/home/kami/.local/bin/mavwaked -device mavmic ... (code=exited, status=0/SUCCESS)
Aug 25 18:21:35 bugmachine mavwaked[2535262]: arecord: pcm_read:2285: read error: No such device
The microphone went away, arecord stopped, and mavwaked exited zero.
systemd read a clean exit and did not restart it. Nothing on either box noticed,
and nothing would have: the always-on listener going silent looks exactly like
the always-on listener having nothing to report.
docs/spec.md says under Hearing that no capture client ships (V-514). A
capture client does ship and it is mavwaked: it spawns arecord for 16kHz
mono PCM, runs silero VAD and the wake head, and sends PushToTalk frames to
mavend's voice port. What is absent is mavenclient. That line in the spec is
out of date.
The speech path was reachable and was not probed
POST /api/ptt is registered on the same mavweb mux this baseline drove 58
turns through (cmd/mavweb/main.go:241) and proxies raw PCM onto mavend's voice
port (cmd/mavweb/voiceproxy.go:47). It takes audio and runs a real
transcribe, route, reply turn.
Nothing in this run posted audio to it. Every speech and senses criterion
therefore reads untested, reason scenario missing, and not deployment missing: the deployment is present and the probe was never written. Recording
that as a deployment problem would have sent the next session to the wrong file.
Two facts about that path were established while correcting it:
- Both workstation endpoints refuse.
192.168.1.105:8080and:8081return no HTTP status from workpc.workstation.model_disabledis alsotrue. So mavend'sstt.Pairhas already fallen to themavsttdfloor, and the floor is the only transcriber in service right now. internal/ttsnormis compiled intomavend, notmavttsd.ttsnorm.Speakableruns on the voice reply path and on nudge text. The deployed mavend carries it on the 2026-08-19 image.
What the run does not establish
- The voice reach was not exercised. No listener runs on workpc, so
voicesinkrecordsblocked, no listenerand nothing here tests it. 156 of the last 200 delivery attempts are that reach failing. - The Praxis lifecycle gate is untested. Praxis holds nothing, so the
mutation had no referent. The architecture pass's claim that
praxisItemAction.handlecalls through with no tier stands as a code reading. - The
mavsttdandmavttsdarms were not exercised, though they were reachable. See the section above: no audio probe was written. - Telegram arrival was not confirmed by the owner. The attempt row says
sent; that the message appeared on his phone is not in this file. - The 14-day routing-trace retention was not exercised. That retention
covers the traces
correctwrites against, which are a table. The decision ring measured above is a different thing and is capped at 25 turns in memory.
Row counts after the field run
| store | before | after | delta |
|---|---|---|---|
| facts | 13018 | 13041 | +23 |
| notes | 180 | 185 | +5 |
| reminders | 95 | 98 | +3 |
| tasks, live | 5 | 6 | +1 |
| decision ring | 0 | 25 | ring full |
| events | 34 | 62 | +28 |
| ecosystem traces | 86 | 115 | +29 |
| nudges | 5 | 5 | 0 |
| tools | 13 | 13 | 0 |
All 23 new facts are ambient. 34 probe turns, including one that stated the
owner's name and two that stated where he works, added nothing to the fact
store. The +23 is desk_active and page_heartbeat continuing at their own
rate through the seven minutes the run took.
The +5 notes are four RSS items and one captured note. That note is stored as
запиши: ключи от гаража лежат в верхнем ящике стола, imperative prefix
included, so the note text is the command rather than the content.
The +3 reminders are two identical позвонить в клинику rows, one from each
run, and the MVNPROBE delivery probe.
Configuration and deployment, read directly
These criteria are not observable through a turn. Chasing them through
POST /api/chat would measure the router instead, so they were read from the
config, the container and the startup log.
The web UI has no notes page and no facts page
Fourteen pages answer 200, the slowest in 28ms.
/, /dash, /history, /trace, /notifications, /reminders, /morning,
/events, /tasks, /chat, /ecosystem, /tools, /routines, /models.
docs/spec.md requires a page for every capability with a surface and names
four: reminders, notes, tasks, facts. /notes and /facts do not exist, and
cmd/mavweb/main.go registers no handler for either. The 2026-08-13 audit's
"all ten pages answered 200" was true of the ten that exist.
This settles the open half of the Notes DoD as well. A note cannot be deleted from the web UI because there is no page from which to delete one (V-494).
/auth/passkey returns 404, consistent with WebAuthn being unconfigured.
Every step-up gate is fail-open
mavweb says so itself, at startup, unprompted:
SECURITY WARNING: step-up verification is DISABLED (-webauthn-origin/-webauthn-rpid unset).
These surfaces are UNGUARDED:
POST /tools defines arbitrary argv via name+cmd, which internal/tool then EXECUTES
POST /routines accepting schedules recurring firing
POST /models chooses the resident model that routes and words every turn
POST /api/revert voids the latest fact for a key
POST /api/chat reaches the router, the LLM and, through applyAction, the act path
POST /api/ptt the same, from audio
Six surfaces, and this baseline drove 58 turns through the fifth of them without authenticating. V-683, unchanged and now measured rather than read.
What is configured correctly
| criterion | evidence |
|---|---|
| the voice port stays on homesrv loopback | docker port maps 9100/tcp -> 127.0.0.1:9110. voice.bind is 0.0.0.0:9100, which is the container's own namespace; the publish is what makes it loopback. |
| the database is encrypted at rest, working copy in tmpfs | /var/lib/maven/maven.db.enc, 9.2 MB, mode 0600. Plaintext copy in /dev/shm, which is tmpfs. |
| the embedder loads at 384 dimensions with the marker check passing | voice: onnx embedder loaded (384 dim), then voice: embedder marker ok (model_quantized@384/tok2) |
| the routing heads load from their own file | router_heads.onnx, not model_path. Refused at config load since V-692. |
What is not configured at all
There is no weather block in deploy/mavend.json. Not a wrong value, not
a disabled flag: the key is absent, so the provider loads as a stub. The audit
found this and it has not changed. какая сейчас погода? answers для какого города? and then cannot use the answer, which is the same two-turn failure the
audit recorded, reproduced here with its routing cause visible above.
Tests and analyzers
make test is green across ./internal/... and ./cmd/..., with -race
and -coverprofile. cmd/mavend took 223.8s at 70.0% coverage and
internal/store 81.0s at 70.8%.
It does not set MAVEN_ONNX_LIB. Makefile:196 is the whole recipe and the
variable is not in it. Only make t and the four eval-* targets set it. So
every TestONNX* measurement self-skips inside make test, and the run prints
ok anyway, which is the failure mode CLAUDE.md names in as many words.
Run under the test target's own environment, one package:
--- SKIP: TestONNXPersonalBoundaryStratified (0.00s)
--- SKIP: TestONNXPersonalBoundary (0.00s)
--- SKIP: TestONNXPersonalBoundaryFourFold (0.00s)
--- SKIP: TestONNXPersonalBoundarySemanticGroupHoldout (0.00s)
--- SKIP: TestONNXPersonalBoundaryChallenge (0.00s)
--- SKIP: TestONNXPersonalBoundaryPostRetuneChallenge (0.00s)
--- SKIP: TestONNXPersonalBoundaryLatency (0.00s)
--- SKIP: TestONNXPersonalBoundaryFrozenHeadMatchesCorpusFit (0.00s)
PASS
ok github.com/kami/maven/cmd/mavend 1.045s
Grepping the make test output for skip returns nothing, because go test
prints no SKIP line without -v. A green run and a run where the measurements
never executed are the same eight characters.
The measurements do pass when they are given the library. Under make t, which
sets it and does not use coverage:
--- PASS: TestONNXPersonalBoundary (5.68s)
--- PASS: TestONNXPersonalBoundaryFourFold (26.30s)
--- PASS: TestONNXPersonalBoundarySemanticGroupHoldout (37.97s)
--- PASS: TestONNXPersonalBoundaryChallenge (5.23s)
--- PASS: TestONNXPersonalBoundaryPostRetuneChallenge (5.35s)
--- PASS: TestONNXPersonalBoundaryLatency (5.17s)
--- PASS: TestONNXPersonalBoundaryFrozenHeadMatchesCorpusFit (13.45s)
--- SKIP: TestONNXPersonalBoundaryStratified (0.00s)
The latency test V-718 describes as failing only under coverage passed here in
5.17s, in a run that carries no coverage. This baseline does not reproduce
V-718 and does not refute it. The one invocation that does use coverage,
make test, is the one where this test self-skips.
make analyze does not pass. staticcheck and deadcode produced nothing
over their baselines. govulncheck reports seven vulnerabilities, every one in
the standard library, every one fixed in go1.25.13, and the vendored toolchain
is go1.25.12:
Your code is affected by 7 vulnerabilities from the Go standard library.
make: *** [Makefile:119: vuln] Error 3
Reachable ones include internal/kiwix/client.go:165 through
encoding/xml, internal/netaddr/netaddr.go:228 through encoding/asn1, and
internal/vision/vision.go:277 through net/http. This is a toolchain bump,
not a code fix.
This is the point the spec was written to make, and the suite made it twice over. It is green while 50 of 146 v1 criteria fail on the running box. It is also green while the four measurements it is supposed to carry are not running at all. A green suite has never been evidence that a capability works.
mavwaked was not queried for its build. It runs on workpc under systemd,
outside this stack, and the measuring box could not reach it.
Undesigned in v1, by inspection
Nine capabilities have no design. Their criteria were settled by reading the deployed artifacts, not by probing, and three readings changed what the spec says about them.
Email and calendar are recorded decisions, not open questions.
docker-compose.yml:145 and :175 each carry a commented service with the
reasoning beside it. The mail block names the blocker (no IMAP account), the
security shape (password from a file, core never sees it, nothing there can
create a reminder) and the steps to enable it. It also records what triage
means: task extraction into candidates he reviews on /tasks, explicitly not
the acting variant. The calendar block says outright that mavcaldav "was built,
listed in make build, and deployed nowhere, which is the worst of the three
states, this block records the decision instead", then writes out what the
absence costs. Both capabilities' "in compose, or its absence is deliberate and
recorded" criterion passes on the second branch.
The box has no inbound HTTP intake at all.
internal/delivery/telegramsink/intake.go long-polls getUpdates outbound,
precisely because nothing can connect inward and it reaches api.telegram.org
through a socks relay. It is not "Telegram's own inbound webhook".
Command chaining is half built and live. voice.llm_router is true on the
deployed box, the router's grammar contract already returns an array of actions
for a compound utterance, and parseActions builds all of them. Nothing
dispatches past the first. The seam is internal/router/llmrouter.go, not a
missing package.
Learning from mistakes has its stores. The nudges table holds an outcome
per nudge, store.RecentOutcomes reads the last N per rule, and
loop.TuneCooldown turns a high ignored rate into a longer cooldown. Route
repair is stored twice, as a routing_labels row and as a classifier example,
and there are two correction paths, not one: POST /api/correct and the spoken
cmd/mavend/repair.go (V-455). routing_labels held 2 rows before the wipe.
What has no store is a corrected phrasing.
What genuinely does not exist: summaries as a requestable capability, webhooks in either direction, any style-learning store, and any document relating the three schedulers to each other.
The tally
146 v1 criteria across 46 capabilities. The five deferred capabilities keep
their ten criteria at untested, reason deferred past v1, and were not
probed.
| verdict | criteria |
|---|---|
| pass | 26 |
| fail | 50 |
| blocked | 15 |
| untested | 51 |
| unknown | 4 |
pass was only ever available to live evidence. untested is the largest
bucket and most of it is honest scope: a wall-clock trigger a probe cannot
reach, a workpc daemon the measuring box cannot see, an external service that
would have to be taken down on purpose.
Per capability, in docs/capabilities/ledger.yaml. Nothing here is ranked.
Ranking is session 2's job and ranking these numbers without the implementation
mapping behind them would be the reverse of the causal order the plan sets out.
What this baseline expires on
Any of these invalidates every number above.
- The resident model moving off
maven-instruct-b2-Q4_K_XL. deploy/mavend.jsonchanging, including committing the model switch that is currently uncommitted.docker compose up -d --force-recreate mavsttd mavttsd mavpoll, which is outstanding from 2026-08-13 and would close the image drift measured here.- Seeding Nexus, which would move every
blocked, external dependency unavailableverdict on the act path. - A voice listener appearing on workpc.
systemctl --user restart mavwakedon workpc, once the microphone is back. The wake word and hearing verdicts are all measured against a dead process.
The store was wiped afterwards
The plan calls for it and the owner confirmed it, knowing the count. Two copies of the encrypted file were taken first, both outside the repo on homesrv, so the decision is reversible:
~/maven-preswipe-2026-08-26.db.enc, the 23:50 seal, before this session.~/maven-preswipe-final-2026-08-26.db.enc, taken after stopping mavend.
wiped. 30899 rows gone, the schema is intact, mavend knows nobody.
Config, models, passkeys and the encryption key are files and were not touched.
Forty seconds after the restart the store held 1 fact, 4 notes, 4 events and 12
tools: desk_active had already fired once, the RSS poller had run, and the tool
registry had re-seeded itself.
Every measurement in this file is now unreproducible against the same data. That is what a frozen eval is for.
Where the raw evidence is
| file | what it holds |
|---|---|
docs/capabilities/out/field.transcript.tsv |
34 turns: probe id, utterance, reply, claiming source, seconds |
docs/capabilities/out/field.raw.jsonl |
the same, plus every readback, including the full claim ladder per turn |
docs/capabilities/out/field.contaminated.jsonl |
the first run, which measured itself. Kept because the leak is a finding |
docs/capabilities/out/counts_after_field.json |
store counts after the probes |
docs/capabilities/out/counts_after_wipe.json |
store counts forty seconds after the restart |
The transcript carries his real data: the task list, the note captured
during the run, and the utterances the probes spoke. The store it came from no
longer exists, so this is now the only record of those rows. It sits in the repo
under the same rule as 2026-08-07-week-of-usage-transcript.md: the transcript
is the evidence and is not summarised anywhere else.
How the verdicts were checked
Every verdict was audited by an independent pass told to refute it, one auditor
per spec section, with pass attacked hardest. It found real errors and they
are corrected above rather than argued with. The ones worth naming, because the
same mistakes are easy to repeat:
make testwas scoredpasson theMAVEN_ONNX_LIBclause. The recipe does not set it. Grepping the output forskipreturned nothing, which is what a self-skipping test looks like without-v. This was the exact trapCLAUDE.mddescribes, walked into while measuring whether other things had walked into it.- "eleven consecutive turns" was wrong. The park reached five, turns 9
through 13, and was released by a path other than
отмена. The wrong count had already propagated into two other files. - "25 traces" was wrong. 22 were read back, 15 reached the route stage and 11 reached the query stage. Claims of the form "in every trace" were false for the seven turns that never routed.
- The voice retries were called ongoing. They are a bounded three-hour burst on 2026-08-16, 151 rows, then silence.
POST /api/pttwas called unreachable. It is on the same mux this baseline used 58 times. Four speech criteria were filed asdeployment missingwhen the deployment is present and the probe was never written.internal/workerwas read for Hexis calls. It is the speech offload wire. Digestion iscmd/mavend/tick_digest.goandinternal/loop.- The masculine reply came from the discarded run, not from the 34 scored turns.
Sixteen verdicts carried an evidence path pointing at a section of this file
that did not exist. build_ledger.py now refuses to build when an evidence path
or a § heading does not resolve.
The corrections moved the tally by nine: four passes withdrawn, and five
criteria that had been filed untested turned out to be already settled.