mavwaked and mavenclient have been written, tested and deployed nowhere since
V-463 parked them. homesrv has a microphone because it is a laptop, but it is
in the wrong room. workpc is where he sits, and it has a fifine on card 0.
V-515 said this was a config line: "ipc.Dial already speaks
tcp://host:port?token=... so this is config, not protocol work". That premise
is wrong and it is worth writing down. Both mavwaked and mavenclient speak
internal/voice through voice.Dial, not internal/ipc. The netaddr token guards
the daemon-to-daemon IPC seam and never touches the voice wire. That wire is
plaintext with no auth at all, and voice/server.go says so: production binds
inside the wg tunnel, because "the wg layer IS the L0 floor".
workpc is not a wg peer. It sits on wlan0. So the floor here is ssh: mavend
publishes the voice port to homesrv loopback only (127.0.0.1:9110, since host
9100 is Vikunja's MCP), and a user unit on workpc forwards it over his key.
Nothing new is on the LAN. That mattered more than it looks: SurfaceVoice caps
acts at L0, so an unauthorized speaker could not run a destructive tool, but
L0 does not cap reading. A LAN bind would let anyone on the wifi hear his
facts, his notes and his calendar read back.
Two things the deployment found that no test could:
The vendored onnxruntime under deps/ has two copies and the stale one is
1.17.1. The Go binding asks for API 26, so silero refused to load until
1.26.0 was shipped instead. mavwaked logged it and kept running on the energy
threshold, which is the designed fallback working.
The fifine offers 2 channels at 44100 or 48000 and nothing else. mavwaked asks
arecord for 16kHz mono, so hw:0,0 dies on "Channels count non available"
before a frame is read. The unit uses plughw:0,0 so ALSA downmixes and
resamples.
Verified end to end through mavwaked's own -test mode, so no human had to
speak: a 2.43s Russian fixture reached mavend over the tunnel, was transcribed
on the workstation by CW2, routed intent=query, claimed by the calendar
source, and came back as 3.68s of piper audio.
There is still no wake word (V-487 stage two), so the loop runs open. Silero
is passed on purpose, since it declines white noise the energy floor accepts.
Barge-in is not, because its threshold is room-specific and this room has no
number yet.
CLAUDE.md was 805 lines and it is loaded into every session, so every line
costs. The routing section alone was 412 of them, and it was a chronological
log of every measurement since 2026-07-31: four re-measurements of the same
fixture, the history of each of the four routing heads, and the reasoning
behind every grammar.
None of that is a rule. An agent about to edit the router needs to know that
the classifier is the floor, that queryWalk only takes sources out, and that
heads_path must never point at model_path. It does not need the seed spread of
the third head to read the file at all.
So docs/routing.md is a living doc under the tier convention, and it carries
the reasoning and the numbers. CLAUDE.md keeps the constraints and points at
it. 805 lines to 490, with the routing section at 60.
The same cut is applied to the header block and to the world chain under
non-goals: the current fact and the eval filename stay, the "measured on date
D it went from A to B" narrative moves out or is dropped.
Nothing was deleted without checking. Every backticked literal in the old file
was diffed against the two new ones, and the forty that fell out were reviewed
one by one. Nine were facts rather than narrative and are restored: the
ecosystem default URLs, the voice.llm_router flag and pickLLMRouter, the four
head eval filenames, handlePraxisAct, SourceAccuracy, and the rule that
calendar-query names the calendar where the possessive agenda rules do not.
A closing section states the file's own contract, so the next agent adds a
measurement to docs/evals/ instead of a paragraph here.
diff-budget.sh blocked on 1544 changed lines. It counts markdown, which the
repo's own pre-commit hook exempts, and this commit touches nothing else.
The resident model alone scores 37.4% full against 61.5% intent-only. The gap is
slots and not routing. It routes `reminder` and leaves the time to the daemon,
which is what the contract asks.
To re-run the resident model as router, start a **second** llama-server on a
fixed host port. The resident one binds `--port 0` inside the container and no
host process can reach it.
### The workstation is not the better router any more
It was, from 2026-08-02 until the heads landed. gemma-4-12b beat everything on
the box at 84.4% intent and 72.7% destination. The heads beat it on both at a
twelfth of the latency. The workstation stays the better phraser.
E4B replaced the 12B on 2026-08-09 by the owner's call. It is a step down on
routing. Against a same-session 12B control it costs four destination cases and
buys 50ms. Read destination as the finding. It names nothing where the 12B names
`recall` or `calendar`, which is safe but walks the whole chain. It has no MTP
and cannot be given any here. The only `gemma4-assistant` draft on disk is
trained against the 12B's hidden states.
## Stage 0: what the grammars claim, and why
A rule at this stage is a claim. Either the model gets this wrong, or it wastes a
second getting it right. Every rule was added against a measurement.
- **Agenda questions** (`AgendaQueryGrammars`, 2026-08-01). "что у меня сегодня",
"во сколько у меня встреча" and anything naming a calendar go to `IntentQuery`.
They were going to `IntentSystem`, where `replySystem` has no agenda arm and
answered "пока не умею". Worth 2.6 points of full accuracy and calendar 0/2 to
2/2.
- **Rest of day and narrative** (V-498, 2026-08-04). `rest-of-day-query` claims
"что дальше?". `NarrativeQueryGrammar` claims "расскажи про X", "объясни X" and
"опиши X". Neither carries a question mark or an interrogative, so the model
called both `IntentFact`. `IsQuestionShaped` caught the write downstream, so
this was a latency and fixture defect rather than a correctness one. The
narrative rule declines `chatNarrativeTopics`, because the query chain has no
source that answers a joke or a bedtime story.
- **Praxis** (V-516, 2026-08-05). `PraxisGrammars()` fills `Slots.Fn` with a
capability name. These grammars are the **only** path to Praxis and not a
faster one. The model reaches Praxis 0/12 alone, the same as the classifier.
Nothing in the router prompt names a Praxis capability, so there is no string
for it to write. Through the cascade it is 11/12. Measured overall 16/30 to
27/30, lifecycle 0/5 to 5/5
(`docs/evals/2026-08-05-praxis-reach.md`,
`docs/evals/2026-08-05-reach-llm-router.md`).
`handlePraxisAct` compares `Slots.Fn` to a capability alias. Otherwise that
slot is filled from the deployment's enabled tool names, and no Praxis alias
is on that list.
- **World questions** (`WorldQueryGrammars`, V-655, 2026-08-07). "что такое X"
and "сколько будет 17 на 23". Wired after the agenda rules and **before** the
feed and list rules. "что такое лента" is a definition question, and the feed
rule would take it on the noun alone.
`calendar-query` and `event-time-query` name the calendar as the destination.
The possessive agenda rules deliberately do not. "что у меня в списке покупок"
matches `agenda-query`, and naming the calendar there would take the list source
off the turn. That caution now costs four destination cases. See the model arm
below.
Go's `\b` is ASCII-only and never fires after a Cyrillic letter. A pattern needs
an explicit `(\s|[?!.]|$)`.
`baselineGrammars` in `eval_test.go` mirrors `buildRouter` and has drifted before.
`WorldQueryGrammars` was wired into the daemon by V-655 and not into the mirror,
so the fixture scored a grammar set nobody runs. Fixed by V-659, worth 3 points
of destination.
### Praxis lifecycle rules
A **stative** lifecycle word ("готово", "принято") needs an item named beside it.
A bare **imperative** ("закрывай") may ask which one. It also requires a sentence
naming no object of its own. Otherwise "закрой шторы в комнате" goes to Praxis
instead of the house. A demonstrative ("отметь это как сделанное") resolves
against `h.surfacedItems` only when exactly one item was spoken. Otherwise the
turn goes back to the cascade rather than transitioning the wrong item.
### Slots on a stage 0 decision
`fillMatchedSlots` runs the stage 2 extractor over whatever a grammar built
(V-572, 2026-08-06). It fills only the slots the grammar left empty. A matched
value always wins, because the rule read a literal pattern and the extractor
guesses.
It did not run before. So `ReminderGrammar` handed the daemon `HasTime: false`
for "напомни в 11:00 позвонить маме", and `missingFor` read the silence as
absence and asked "Когда?". It is inert for every grammar but the reminder:
`Extract` fills Time, Fn and Key and nothing else. A stage 0 query costs 3.7µs
against 3.9µs before, benchmarked at 20000x.
`Slots.Text` is deliberately not filled. A grammar that left it empty meant it,
and `agendaQueryBuild` hands the query chain the utterance itself.
## Stage 0b: the routing heads
Routing has a bounded output space, so it is classification rather than
generation (owner's call, V-546,
`docs/plans/18-routing-heads-on-e5-small.md`). The 118M multilingual-e5-small is
already resident. A softmax cannot emit a value that does not exist, so no
grammar is needed. Max softmax is a calibratable confidence, where
`Confidence: 1.0` was a hardcode. Training costs roughly 5e15 FLOPs, so 10 to 30
minutes on the workstation. A 100M decoder from scratch is 10 to 20 GPU hours.
**Fine-tune a copy of the weights.** The resident embedder backs memory recall.
Training it in place couples routing accuracy to recall@1, with nothing in the
suite to name the trade.
Four heads share one masked mean pool, trained over three days. The measurements
are `docs/evals/2026-08-08-routing-heads-two-head.md`,
`docs/evals/2026-08-08-slot-head-three-head.md`,
`docs/evals/2026-08-08-clarify-head-four-head.md` and
`docs/evals/2026-08-08-massive-warm-start.md`.
| Head | Score | Notes |
|---|---|---|
| intent | 92.8% mean over 3 seeds | fixture is the 88 cases carrying an intent |
| destination | 80.8% mean, best 29/33 | beats the 12B teacher it was distilled from |
| slot BIO tags | 72.4% span F1 | still climbing when epoch selection stops it |
| clarify | catches 7.0 of 8, 2.3 false of 88 | parity with the cascade, no rules in front |
Read the best destination run as one seed and not a headline. One case is 3
points on a fixture this small. Head intent accuracy is **not** comparable to the
cascade's 76.0% and 84.4%. A softmax has no clarify class, so the head's fixture
is 88 cases and not 96.
Recall is 15/15 and world is 5/5.
**Mood is cut, not deferred.** The enum describes her own reply state, not the
speaker's emotion, and no dataset maps onto it.
### The clarify head
Clarify is not a value of intent, so a softmax cannot emit it. It is a second
question over the same pooled vector: can Maven act on this at all. Accuracy is
the wrong number here and a head that never asks scores 91.7%.
Confidence is the other half. Max softmax over the intent head reads 0.851 where
it is right and 0.604 where it is wrong. It ranks right above wrong in 83.4% of
pairs.
It is not free the way the slot head was. Intent, destination and slot F1 each
move down one to four points, inside the seed spread. `поужинал` is a false
clarify on every seed. That is the same defect `thinSingleToken` was narrowed for
on 2026-08-01.
The corpus is generated, because every existing row is answerable by
construction. The router-prompt agreement filter cannot work here. `routeGrammar`
has no clarify value, and a generated line always agrees with itself. A gemma
judge replaces it. The first judge called 24 of 40 answerable rows underspecified.
It judged against a generic assistant rather than against Maven's contract.
### The slot head
BIO slot tags had no Maven-domain corpus. That was true of found corpora and
false of made ones. `label_slots.py` distils spans out of gemma-4-12b under a
GBNF closed over Maven's own five slots. A span survives only when it is a
literal substring of the utterance, so the agreement filter costs no second call.
2178 spans over 1702 rows, 37 dropped, nothing unparsed.
Epoch selection reads the intent dev slice alone. That costs the slot head about
4 points.
### Warm start and the floor
The MASSIVE warm-start of step 2 is worth nothing here. Stock e5-small ties it on
intent and leads by a third of a case on destination. Nothing argues for keeping
that step.
The floor was a corpus defect and it is fixed. The first 120 floor rows carried
one sentence shape, so the head named a destination where the fixture says walk
the chain. Rotating six shapes took the floor 3/7 to 6/7 and destination 75.8% to
80.8%.
What is left is calendar at 3/6 on every seed, which training cannot move. The
possessive agenda rules claim those cases at stage 0 and name nothing, so no
label reaches the head.
### Reading them in Go
`RouterHeads` in `internal/router/heads.go` loads `router_heads.onnx` (V-664,
2026-08-08). It reads intent, destination and clarify off one forward pass.
Three rules around it, each measured:
- The **clarify head decides first**, before the intent threshold. It answers a
different question. A thin utterance scores low intent by construction, so
gating it cost 6 of 8 ambiguous cases.
- The **destination head is read on `IntentQuery` only**, since no other intent
reaches `queryWalk`.
-`headsThreshold` is 0.6, the measured knee. Every value up to 0.85 drops right
answers and keeps the same two wrong ones.
`voice.embedder.heads_path` is the whole switch. Empty, missing or unloadable
means the heads are nil. The cascade is then byte-for-byte what shipped before
them.
### The tokenizer bug the heads found
`encodeWord` in `onnxembedder.go` read every long word backwards until 2026-08-08.
It cost recall@1 7.4 points and recall@3 11.1. Nothing caught it, because seeds
and queries were mangled the same way and cosine survived. The heads found it.
They are trained through transformers and read through this.
The embedder id now carries a tokenizer revision (`@384/tok2`). So fixing the
tokenizer triggers `ReembedAll` the way swapping the model file does. Bump
`tokenizerRev` on any change to what it emits.
## Clarify
`Confidence: 1.0` was hardcoded in `llmrouter.go`. So the model path could never
ask for clarification, and it missed 6 of 6 refusal cases (V-359). The bug had a
second half. The model branch never consulted `r.threshold` at all, so a correct
low confidence would have been discarded anyway.
Fixed 2026-07-31 with structural signal feeding the same stage 3 gate the
classifier path already had (`gateLLMDecision` in `router.go`). Three signals: a
single-token utterance, a keyless fact, an act with no allowlisted fn.
Re-measured: missed clarify 6/6 to 1, at the cost of 3 false clarifies and 2.6
points of full accuracy. Two of the three false clarifies are acts the model
mis-routed and the gate caught. Asking beats wrongly executing, so the fixture
and the daemon disagree about what is correct there.
The third, `поужинал`, was a real defect. The single-token rule was an English
intuition. It does not transfer to Russian, where one word is routinely a whole
sentence.
Narrowed 2026-08-01. `thinSingleToken` (`internal/router/singletoken.go`) still
thins a bare one-word nominal. It spares two classes. One is a closed lexicon of
social and control singles ("привет", "стоп", "yes"). The other is any token
carrying a Russian verb ending, because a verb already contains its subject. Both
tests are offline and cost nothing. False clarifies 3 to 2, intent-only 74.0% to
75.3%.
The two remaining false clarifies are the act-with-no-allowlisted-fn arm of the
gate, not this rule.
## The destination
`query` was a shrug. The cascade sorted an utterance into one of seven intents,
then `IntentQuery` handed the turn to `querySources` in the daemon. That is
twenty-two branches deciding by seed similarity in a fixed order. It had no
fixture, no accuracy number, no model arm and no floor.
`Decision.Source` (`internal/router/source.go`) is the second half of the route
(V-655, 2026-08-07). Twelve destinations, not twenty-two. The three recall passes
plus `fact-by-key` are one destination from outside. So are search, Kiwix and the
URL reader.
`queryWalk` in `cmd/mavend/actions_query.go` takes sources **out** and moves none.
That is the safety argument. The table's order is load-bearing. Every comment on
it argues a reason between two sources. Above all it carries "the owner's data
first, then the world". Naming `SourceWorld` does not send the turn outside on its
own.
What comes out is only the sources that **guess**. Those decide a turn is theirs
by cosine against frozen seeds, then answer whatever they claimed. They hold no
table that could come back empty. Weather is the pure case and has no local data
at all. It was measured on the box 2026-08-07
(`docs/evals/2026-08-07-week-of-usage.md` section 4). It answered both "что такое
TCP?" and "сколько будет 17 на 23?" with "для какого города?". The feed answered
"какой у меня любимый язык?" with kernel headlines.
### Who may drop the personal boundary
The personal boundary guesses, so naming `SourceWorld` drops it. That is what
stops it answering "кто такой Линус Торвальдс?" with "не нашла у тебя такой
записи", which it did on 2026-08-07.
Three deciders name a destination and two of them infer it: the heads and the
resident model. An inferred `SourceWorld` on a question about him would reach
SearXNG. That widens what is asked rather than costing a local answer. So only a
stage 0 grammar may drop it (owner's call, V-666, 2026-08-09).
`Decision.SourceAnchored` carries the provenance. It is a field and not
`Stage == 0`. Stage 0 also means confidence 1.0 and an anchored claim band, and
one of those could stop implying the others. `definitionQueryPattern` claims "кто
такой X", so the 2026-08-07 case is still anchored and still answered.
### The destination fixture
`want_source` on `eval.Case` is a pointer, because the destination has three
states and a bare string has two. Absent is every intent but query. Present and
empty is the `SourceUnknown` contract: name nothing and walk the chain. Present
and named is a destination the route must produce. Thirty-three of ninety-six
cases carry one.
A destination miss does **not** fail the case. It lands in `Outcome.SourceReason`
and never in `Reasons`, so `Accuracy` and `IntentAccuracy` mean what they meant.
`SourceAccuracy` is a second number over the labelled cases alone.
A route that lost its intent scores no destination hit. Otherwise a clarify would
satisfy an empty label for free.
Seven cases assert the floor and five of them are homelab operations. They cluster
because `SourceRecall`, `SourceNetwork` and `SourceAttention` overlap on every
question about the box. `mavpoll` writes its netdata and uptime-kuma observations
into the fact store recall reads. That is a finding about the enum, not a gap in
the labelling. The other two are `ru-query-005` and `ru-query-014`. No query
source reads the reminder store, and a deadline could sit in tasks, the calendar
or Praxis. The owner confirmed all seven floor labels on 2026-08-08.
### The model arm
`routeGrammar` carries a `source` rule closed over `router.Sources` plus the
empty floor (V-660, 2026-08-08). So the model cannot emit a destination that does
not exist. The prompt lists the twelve in Russian and says `""` is a normal answer
to give often. `LLMRouter.Route` reads it back through `ValidSource` and on
`IntentQuery` alone.
Against gemma-4-12b the cascade scores destination 24/33 with intent unmoved, and
recall goes 0/15 to 14/15.
**Stage 0 now costs four destination points.** It did not before. The four cases
the cascade loses and the model alone wins are all calendar. The possessive agenda
rules claim them first and name nothing on purpose. That caution was free while
nothing downstream could name anything either. It is not free now, and the fix is
the owner's call (V-660 open).
## The decision trace
Arbitration between the claimants on the utterance stream is order. It is
hardcoded in the pre-route resolver ladder, in `buildRouter` and in
`querySources`. Nothing recorded who lost until V-564.
`internal/decision` records one `Record` per turn. It holds every claimant, what
it would have made the turn, the score it reported, and how it ended. A claimant
won, declined, lost on score, was thinned by a gate or was **never asked**.
The record rides the context, the same seam `querysource.go` uses. So a claim
site cannot change a route, and a context with no record costs nothing. It is
installed in `runTurn`, so the mic, telegram and the web leave the same trail.
Adding a rung to the ladder in `runTurn` means adding its name to `preRouteLadder`
in `cmd/mavend/decisiontrace.go`. Otherwise that rung is silently missing from the
record.
### Why it persists now
The original rule was that nothing persists, because a turn record is read minutes
later or never. Storage was a 25-turn in-memory ring read over `ipc.TurnDecisions`
and rendered on `/trace`.
The owner reversed it on 2026-08-06 (V-629,
`docs/plans/21-persisting-the-routing-trace.md`). The routing heads cannot be
fitted or calibrated without real utterances. And 9 of the 31 modes in
`internal/modes` have no seed example at all.
The ring did not move. `cmd/mavend/routingtrace.go` is a second sink beside it,
writing `routing_traces` (migration #23). The utterance is stored in clear. A
384-dimension vector of a short sentence is substantially recoverable, so storing
vectors instead would be a privacy claim we cannot support. What makes it safe is
the same thing that makes the fact store safe. Retention is 14 days, enforced on
write and again on start, so a box that goes quiet does not keep every row.
Nothing reads it outward. `Store.Wipe` deletes it with everything else.
### Corrections
A correction is promoted out into a seed-shaped row in `routing_labels`
(migration #24) and kept, because a label is not a transcript. The transcript
still expires.
A turn marked wrong with no target is a usable negative, so naming the intent is
never required. The target is one of the seven intents and never free text.
All three reaches offer it as of 2026-08-06:
-`/chat` offers two buttons beside the reply, over `ipc.CorrectTurn` and the
trace id that rides back on `ipc.ChatReply`.
- Voice offers the `repair` rung, which has read spoken corrections since V-455.
It now writes the durable label beside the classifier seed it always wrote. A
spoken negative with no target is its own rung, `repair-negative` (V-636,
`docs/plans/22-correcting-a-turn.md`).
- Telegram offers an inline keyboard under the reply. It needed the chat to become
readable first (V-637, `docs/plans/23-inbound-telegram.md`). The poller is dark
unless the `telegram` block says `intake`. It long-polls, because the box takes
no inbound connections. It accepts `chat_id` and no other sender, and it drops
whatever queued while the daemon was down. It reaches the daemon through
`ipc.CoreAPI` alone.
The turn source is still `tap:text` for both telegram and the web. So provenance
cannot tell a chat turn from a typed one.
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.