Seventeen markdown files at the repo root, twelve of them dated one-shot reports sitting next to CLAUDE.md. That is why stale docs read as current: nothing in the path said which was which. Root now keeps CLAUDE.md and AGENTS.md. Living docs move under docs/ and carry a Last verified line. Dated measurements move to docs/evals/ ISO-prefixed, and are never edited after the day, so a newer number is a new file. The senior review moves to docs/archive/. Every reference was rewritten across markdown, Go comments, the Makefile and the recall fixture. The touched Go packages still build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
13 KiB
Deterministic logic around a small model
Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.
Written 2026-08-02. Branch fix/integrated.
The question
Where does deterministic code attach, so that it helps the resident 1.7B now and does not fight a larger model later.
The mistake to avoid
Everything deterministic we have added so far sits in front of the model and preempts it. Stage 0 matches, the model never sees the turn. That shape helps a weak model and blocks a strong one, silently.
The fix is not to remove it. The fix is to know which rules are safe in that position and to have a way to measure the rest.
Four attachment points
Bypass, before the model. The only shape that saves the 2.7s p50. Safe when the rule is a decision procedure over a closed set, not a guess over an open one. Exact match and clock queries qualify.
Evidence, beside the model. Extractors emit candidate slots as a prior, not a verdict. The prompt carries the prior and the validator reuses it. A small model leans on it, a large one overrides it correctly.
Grammar, around the model. GBNF built from live state rather than hardcoded. Costs nothing at runtime and prevents the error instead of catching it.
Repair, after the model. Validation failure re-asks with the specific error rather than overriding. Self-retiring, because a better model trips it less.
The constraint that ranks them
Latency must stay minimal. That demotes repair and promotes grammar.
- Grammar first. Zero runtime cost, immediate gain.
- Bypass keeps its place. It is the only thing that avoids a model call at all.
- Repair only where failure is rare, capped at one retry.
- Evidence is correct but costs a model call where a bypass costs none.
- Ecosystem calls belong in the snapshot path, in parallel, on strict deadlines. Never serial before routing.
The line that never moves
Separate policy from capability compensation. They look alike and age oppositely.
Capability compensation exists because the model is weak. It should be measurable and retirable.
Policy exists because we decided. The personal boundary, the feminine persona, the never-search-his-notes rule, the Hexis allowlist and confirmation binding. None of those yield to a smarter model. A larger model is more dangerous there, not less.
What we keep
Every stage-0 rule stays exactly as it is. Retiring them was the wrong call and it would throw away a day of measured gains.
- exact-match fast path
- clock rules
SystemTimeDateGrammarsAgendaQueryGrammarsthinSingleToken, including the social lexicon and the verb-ending test
Three additions, none of which change behaviour:
- Each rule gets an id and a fixture subset.
- Each rule writes one trace line when it fires.
- Each rule carries a comment saying whether its set is closed or open.
That preserves today's accuracy and buys the option to revisit later with numbers.
What we build
Dynamic grammars from live state. The grammars today are static: routeGrammar
fixes the 7 intents, responseGrammar fixes the mood enum, kiwix queryGrammar
fixes word shape. Everything else is a free string.
Candidates in order of payoff:
- Hexis capability ids. After discovery the exact list is known. As an enum, the model cannot name a capability that does not exist.
- Act fn allowlist. Hits the two remaining false clarifies directly. They are
the act-with-no-allowlisted-fn arm of
gateLLMDecision, firing on invented verbs. - Calendar names. Enumerate the real ones for agenda and query slots.
- Known fact keys. A read-back matches a stored key instead of inventing a synonym. This is the general form of the read side we scrapped on 2026-08-01.
- Nexus display names as act targets. Only while the list stays small.
Two rules or it backfires:
- Always include an escape value. A closed enum with no
otherforces a wrong pick instead of a decline. The escape is what feeds the clarify gate. - Cache the grammar string, keyed on the state that built it. Rebuilding per turn is fine. Recompiling a large grammar per turn is not.
Retrieval over regex
Resolve against stores that already exist rather than adding patterns.
Identity is the worked example. Nexus is authoritative, actionFact already sets
Subject, and cmd/mavend/factenrichment.go resolves it in the background. The
scrapped work invented a parallel key namespace with nothing reconciling the two.
A table that grows with real data beats patterns that grow with our patience.
Note a real gap: the personal boundary in cmd/mavend/actions_query.go guards
Maven's own store only. It does not know Nexus or Praxis exist.
Prerequisites
Both are offline and need no deploy. Neither existed on 2026-08-01, and that is why the day cost what it did.
Failure taxonomy over the 77-case fixture. Classify every miss as model ignorance, contract loss, prompt ambiguity, or our own bug. Only model ignorance deserves deterministic compensation. The other three get fixed once, for every model size. Inference, not verified: much of what we patched was the last two.
Per-assist ablation runner. One command toggles each assist off and reports the accuracy delta on its fixture subset. Then retiring an assist is a config flip and a number, not an argument.
Held
Not now, and nothing gets built for it.
- A 12B or 35B model. It may be cloud or the 16GB workstation, and cloud crosses the current no-third-party line.
- The escalation tier in
gateLLMDecision. - Converting the stage-0 heuristics to evidence.
One thing carries forward for free: deterministic assists emit confidence, never a verdict. A verdict cannot escalate.
Ruling on the idea list
Nineteen ideas, judged on whether they earn a place in Maven. Checked against the tree on 2026-08-02, not from memory.
Build. Absent, and worth it.
SQLite FTS over embeddings. No fts5 anywhere in the tree. Lexical search is
faster than the ONNX embedder, deterministic, and strongest exactly where the
embedder is weakest, which is exact Russian names. Semantic search becomes the
fallback rather than the gate. This is the highest-value absent item.
Cached TTS phrases. No cache in internal/tts. Confirmations, clarifies and
refusals repeat constantly and their text is already fixed. Pre-rendering them is
cheap and pays straight into the minimal-latency constraint.
Synthetic router dataset. Generate Russian tool-calling examples from the same schemas that will build the dynamic grammars. One source of truth for both, so the model is trained on exactly the shapes it will be constrained to at inference.
Command STT separate from dictation STT. One path today. Commands want latency
and a small vocabulary, meeting capture wants accuracy and can take its time.
internal/capture already spools to disk, so the split follows the existing seam.
Medium priority, behind the four above.
Polish. Present, incomplete.
Strict JSON everywhere. Done 02-08-2026. The replier sends
phraser.ResponseGrammar. It is exported once so its two copies cannot drift.
The meeting summariser is wrapped and unwrapped in the daemon's Completer, so
internal/capture stays text-in/text-out. Every model call now carries a grammar.
Evidence-first prompting. Done 02-08-2026, in the evidence branch of
PhraseQuery. Sources arrive numbered, one per line. The system prompt no longer
calls them all "заметки" and no longer lets the model add anything of its own.
Blank sources now take the knowledge branch instead of asking for an answer from
an empty list. Not yet measured against a live model. Run make eval-phrasing,
and watch the case query-notes-do-not-answer.
Progressive inference. What exists is a fallback cascade, not escalation. On
error it drops to something weaker. It never escalates on ambiguity. The upgrade is
held with the bigger-model question, and gateLLMDecision is the hook.
Entity dictionaries. Nexus is the canonical store, behavior_ru.go has
KeyAliases, ecosystem_acts.go has verb aliases. Morphology is the gap, and the
code says so in three places. Russian needs it and there is no stemmer in the repo.
Assistant state machine. confirm.go models pending confirmation and
dialogue.Session models the turn. There is no unified task state. Reuse the
Praxis vocabulary rather than inventing one, because surfaced, acknowledged and
resolved already mean something precise here.
Session memory compiler. The digestion tick, internal/memory and
followUpMerge do parts of this. Not a coherent compile step.
Background memory maintenance. Mostly present, and the absent part is small.
What exists: internal/memeval reads recent memory on a loop and writes
observations, deduped against its own prior output, unable to speak or act. Facts
supersede at write time through voidsID, CorrectValue and VoidLatestFact,
and RecentActiveFactsByKind reads only live rows. Digests, ecosystem traces and
media all prune.
Three gaps remain, all in the fact store:
- Superseding is turn-driven. Nothing reconciles two live facts that contradict unless a turn corrects one of them.
- Duplicates written by different sources or phrasings stay as separate live rows. There is no key-level merge pass.
- Nothing expires. The log is append-only and grows without bound, and no fact ever ages out on its own.
Worth doing, and smaller than it looked. Not urgent.
Qwen CPT then narrow SFT. In flight as #122. One disagreement with the idea as
written: do not drop persona from the SFT. Persona is the stated reason the CPT
exists, because stock writes рад where Maven needs рада. Keep the joint
router-plus-persona tune.
Offline job queue. The tick loop, factEnrichmentWorker with backoff, and the
media prune already defer work. A general queue is tidier, not more capable. Low
priority.
Response templates. Present in clarify.go, in the replySystem arms and in
StubPhraser. Worth extending to high-frequency confirmations, where latency and
persona correctness both matter. Do not extend further. Templating the
conversational reply removes the reason she is worth having.
Reject.
Grammar-first routing as a replacement for the LLM router. Already measured. The classifier scores 36.8% full accuracy against 72.7% through the cascade. Replacing the model with rules halves the accuracy. Grammar-first as an ordering is what stage 0 already is, and that stays.
Hierarchical intent classification. Seven intents is already the coarse layer.
The specialisation stage exists as per-intent slot filling in Extractor.Extract.
Adding a tier buys structure, not accuracy.
Tool-specific micro-models. Contradicts the one-resident-model constraint, needs per-domain training data nobody has, and multiplies model loads on a single Vega iGPU. Dynamic grammars give the same domain narrowing at zero runtime cost.
Local knowledge graph. Nexus owns entities and relationships. Building a second graph in Maven breaks the ecosystem line and creates two answers to one question. If graph traversal is wanted, it is a Nexus feature request.
Preemptible training. Training runs in a separate workspace, not on the serving box. This only becomes real if CPT moves onto homesrv, and that is not the plan.
Known open, carried over from the deleted handoff
Found live on 2026-08-01, not fixed. Everything else in that file was stale.
- Kiwix ranks badly on a correct query. The stop-word pass eats the "and". So
"кто написал войну и мир" reaches Kiwix as "war peace author". The top hit is
"List of peace activists" and she summarises that as the answer. The tag
scrapped/fact-and-kiwix-phrasesfixes the query text. The ranking is the ZIM search and is untouched either way. - Chat drags prior turns into an answer. One live reply mixed the greeting, the height statement and a world question. It named Левитан as the author of Война и мир.
safeKeydrops Cyrillic, so Russian calendar events on one day collide. Vikunja #443 with three fix options. It is a migration, not a patch.
Open after the 02-08-2026 deploy
SearXNG runs on homesrv at http://searxng:9563, on maven_default, and the
rebuilt mavend wires it. "кто написал войну и мир?" now routes to query, takes
4 results off the search, and answers Толстой with the lookup opener. Two things
that turn left unsettled.
- The turn was slow, and nobody knows yet whether that is real. Route 7s,
search 1s, phrasing 15s. The p50 in
docs/evals/2026-07-31-routing.mdis 825ms. It was the first turn after a cold start with the model still warming, so it proves nothing either way. Re-run the same question warm before treating it as a regression. Do not plan latency work off this number. - The personal boundary has never run live.
queryPersonalincmd/mavend/actions_query.gostops a question about him from reaching SearXNG. The question above is not one, so only tests cover it. Ask something about him on the deployed box and confirm from the log that novoice: searchline appears for it.
Sequence
- Failure taxonomy over the fixture.
- Ablation runner, plus ids and trace lines for the existing rules.
- Dynamic grammar for the act fn allowlist.
- Dynamic grammar for Hexis capability ids.
- Dynamic grammar for calendar names and known fact keys.
- Extend the personal boundary to the ecosystem stores.
Steps 1 and 2 come before anything is written in the router.