Compare commits

...

9 Commits

Author SHA1 Message Date
claude c15c2b7bd2 Record the destination number and two stuck measurements (V-659)
CLAUDE.md said the destination had no fixture and no accuracy number. It
has both now: intent 73/96 and destination 12/33 on the classifier cascade,
with the per-destination split, the floor cases and the grammar drift the
labelling turned up. Anyone adding a grammar now reads that baselineGrammars
mirrors buildRouter and drifts silently when it does not.

docs/evals/2026-08-08-massive-warm-start.md was written on the V-655 branch
and parked in .task/, which git excludes, so it was one `task start` away
from being lost. It is a dated measurement and it belongs under docs/evals
whatever branch produced it. Its "destination has no fixture at all" line is
now a pointer to the file beside it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
2026-08-08 18:07:41 +04:00
claude b6eaa704a2 Label the destination on 33 fixture cases (V-659)
Twenty-eight existing query cases get a want_source and five new ones
arrive with theirs. Every label is the destination that SHOULD claim the
turn, which on the five new cases is not the one that did: they were
observed failing on the box on 2026-08-07, so the fixture fails on the day
it is written.

Seven cases assert the SourceUnknown floor, and six of those are homelab
operations. They cluster because SourceRecall, SourceNetwork and
SourceAttention overlap on every question about the box: mavpoll writes its
netdata and uptime-kuma observations into the fact store recall reads.
Naming one destination there takes the other two off a turn that needs
them. That is a finding about the enum, not a gap in the labelling.

The fixture's grammar mirror had drifted. WorldQueryGrammars went into
buildRouter with V-655 and never into baselineGrammars, so the fixture was
scoring a grammar set the daemon does not run — the exact thing the comment
above that function forbids. Adding it moved the destination number 9/33 to
12/33 and moved nothing else.

Measured classifier+onnx: intent 73/96 (76.0%), was 69/91 (75.8%). Four of
the five new cases pass and no existing case moved. Destination 12/33
(36.4%), and the split is the point. World is 5/5, because a stage 0 rule
names it. Calendar is 2/6, because the possessive agenda rules deliberately
do not. Recall is 0/15, because nothing anywhere names it yet.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
2026-08-08 18:05:49 +04:00
claude 2597a7b34a Score the destination apart from the intent (V-659)
The fixture measured the first half of a route and stopped. V-655 split a
routing decision in two, and the second half arrived with no fixture, so
Decision.Source had no accuracy number at all.

want_source is a pointer because the destination has three states and a
bare string has two. Absent is every intent but query, which never reaches
queryWalk. Present and empty is the SourceUnknown contract: name nothing
and let the daemon walk the chain, which is right whenever two destinations
can both answer and the utterance does not choose. Present and named is a
destination the route must produce.

A destination miss does not fail the case. It goes in SourceReason, never
in Reasons, so Accuracy and IntentAccuracy stay the numbers they were and
69/91 still means what it meant. SourceAccuracy is the second number, over
the labelled cases only, because a percentage of the whole fixture would be
a percentage of turns that never ask a query source.

A clarified or mis-routed case still counts in the denominator. It named no
destination and that is a miss, not a case to skip, or the denominator drops
every turn the route already lost.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
2026-08-08 18:00:37 +04:00
claude 7203cd56fd Record the second half of a route in CLAUDE.md (V-655)
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 13:04:01 +04:00
claude ab3e818bb9 A named destination silences the guessers and moves nobody (V-655)
querySources splits in two once you look at which sources over-claimed during
the week of 2026-08-07. The clean ones perform a lookup and can come back
empty: fact-by-key, tasks, list, money, calendar, notes. The dirty ones decide
by cosine against frozen seeds and then answer whatever they claimed, because
they have no lookup that could miss. Weather has no local table at all, which
is why "что такое TCP?" became "для какого города?".

So each source now carries its destination and whether it guesses, and
queryWalk takes the guessers that were not named OUT of the chain. It removes
and never reorders, which is the whole safety argument: the table's order is
load-bearing, every comment on it argues a reason between two sources, and
above all it carries "his data first, then the world". Naming SourceWorld does
not send the turn outside. It stops weather claiming a protocol on the way
past. His notes, his facts and the boundary in front of them still run first,
so a wrong destination costs nothing but the guess it prevented.

The skipped sources are recorded as never-asked with the reason, so /trace
shows a narrowed walk rather than a chain that silently shrank.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 13:02:13 +04:00
claude b5500a5be8 Say where the answer lives, not just that it is a question (V-655)
A question was sorted twice. The cascade picked one of seven intents with
stage 0 rules, the resident model and the classifier behind it, a 91-case
fixture measuring it and the decision trace recording it. Then IntentQuery
handed the turn to a second dispatch in the daemon, twenty-two branches
deciding by seed similarity in a fixed order, with none of that. The careful
sorter did the easy half.

Decision grows a Source: twelve destinations, not twenty-two, because the
recall passes are one destination from the outside and so are the three world
sources. Empty is a real value and it is the floor — nothing names one, the
daemon walks its whole chain, and that is exactly what shipped before.

Stage 0 fills it where a deterministic rule already knows. Two new world rules
for the shapes measured failing on the box on 2026-08-07: "что такое TCP?" and
"кто такой Линус Торвальдс?" were answered by weather and by the personal
boundary, and "сколько будет 17 на 23?" was answered "для какого города?".
The calendar noun rule and the closed event-noun rule name the calendar. The
possessive agenda rules deliberately do not: "что у меня в списке покупок"
matches agenda-query, and naming the calendar there would take the list off
the turn.

Fixture unchanged at 69/91, which is the point — it scores intent, and none of
these cases changes intent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 13:01:59 +04:00
claude 9095ac847d Merge pull request #198 2026-08-07 10:24:39 +02:00
claude 4b5f6adbae Merge pull request #197 2026-08-07 10:22:07 +02:00
claude 2bbd8edbf6 Record the week of usage that found V-654 and its siblings (V-654)
Two untracked files left in the tree by the audit session. They are the evidence behind V-654 and several sibling tasks, so they belong on master rather than inside the PR that fixes one of them. Dated eval files under docs/evals/, so they are never edited after the day.
2026-08-07 11:59:40 +04:00
16 changed files with 1439 additions and 78 deletions
+80
View File
@@ -357,6 +357,86 @@ Adding a rung to the ladder
in `runTurn` means adding its name to `preRouteLadder` in
`cmd/mavend/decisiontrace.go`, or that rung is silently missing from the record.
**A route now says where the answer lives, not only that the turn is a question**
(V-655, 07-08-2026). `query` was a shrug. The cascade sorted an utterance into one of
seven intents, with stage 0, the resident model and the classifier behind it. Then
`IntentQuery` handed the turn to `querySources` in the daemon. That is twenty-two branches
deciding by seed similarity in a fixed order. It has no fixture and no accuracy
number, no model arm and no floor. `Decision.Source` (`internal/router/source.go`) is
the second half of the route. Twelve destinations, not twenty-two. The three recall
passes plus `fact-by-key` are one destination from outside. So are search, Kiwix and
the URL reader.
**`SourceUnknown` is a real value and it is the floor.** Nothing named a destination,
so the daemon walks the whole chain. That is byte-for-byte what shipped before the
field existed. The classifier arm names nothing, so a box whose model is down routes
queries exactly as it did.
`queryWalk` in `cmd/mavend/actions_query.go` takes sources **out** and moves none.
That is the safety argument and it is not negotiable. The table's order is
load-bearing. Every comment on it argues a reason between two sources, and above all
it carries "the owner's data first, then the world". Naming `SourceWorld` does not
send the turn outside. His notes, his facts and the personal boundary still run first.
What comes out is only the sources that **guess**. Those decide a turn is theirs by
cosine against frozen seeds, then answer whatever they claimed. They hold no table
that could come back empty. Weather is the pure case and has no local data at
all. It was measured on the box on 2026-08-07
(`docs/evals/2026-08-07-week-of-usage.md` section 4). It answered both "что такое
TCP?" and "сколько будет 17 на 23?" with "для какого города?". The feed answered "какой у меня любимый язык?" with kernel headlines.
The personal boundary answered "кто такой Линус Торвальдс?" with "не нашла у тебя
такой записи". A source that guesses is marked `guesses: true` in the table. One that
looks is not, and it is always asked.
Stage 0 fills the destination where a rule already knows it. `WorldQueryGrammars()`
(`internal/router/worldquery.go`) claims "что такое X" and "сколько будет 17 на 23".
It is wired after the agenda rules and **before** the feed and list rules.
"что такое лента" is a definition question, and the feed rule would take it on the
noun alone.
`calendar-query` and `event-time-query` name the calendar. The possessive agenda rules
deliberately do not. "что у меня в списке покупок" matches `agenda-query`, and naming
the calendar there would take the list source off the turn.
Fixture unchanged at **69/91 classifier+ONNX**, measured both sides. That is the
expected result, because it scores intent and no case here changes intent.
**The destination has its own fixture and its own number as of 08-08-2026**
(V-659, `docs/evals/2026-08-08-destination-fixture.md`). This section used to say
it had neither. `want_source` on `eval.Case` is a pointer, because the destination
has three states and a bare string has two. Absent is every intent but query,
which never reaches `queryWalk`. Present and empty is the `SourceUnknown`
contract: name nothing and walk the chain. Present and named is a destination the
route must produce. Thirty-three of ninety-six cases carry one.
A destination miss does **not** fail the case. It lands in `Outcome.SourceReason`
and never in `Reasons`, so `Accuracy` and `IntentAccuracy` mean what they meant
and `SourceAccuracy` is a second number over the labelled cases only. Intent and
destination are two decisions, and one number hides which one moved. A route that
lost its intent scores no destination hit, or a clarify would satisfy an empty
label for free.
Measured classifier+ONNX: intent **73/96 (76.0%)**, destination **12/33 (36.4%)**.
The split is the finding. World is 5/5, because a stage 0 rule names it. The
`SourceUnknown` floor is 5/7. Calendar is 2/6, because the possessive agenda
rules deliberately do not name it. And **recall is 0/15, because nothing
anywhere names it**. Those turns are still answered, since the chain walks
recall early. Recall is the number the fourth head has to move.
Seven cases assert the floor and six of them are homelab operations. They
cluster because `SourceRecall`, `SourceNetwork` and `SourceAttention` overlap on
every question about the box. `mavpoll` writes its netdata and uptime-kuma
observations into the fact store recall reads. That is a finding about the enum,
not a gap in the labelling.
`baselineGrammars` in `eval_test.go` mirrors `buildRouter` and had drifted:
`WorldQueryGrammars` was wired into the daemon by V-655 and not into the mirror,
so the fixture scored a grammar set nobody runs. Fixed by V-659, worth 3 points of
destination and nothing else. Check that function when adding a grammar.
The model arm is still the follow-up. It lands on V-546. Intent, mood and BIO slot
tags were already three heads on one forward pass of the resident e5-small.
Destination is a fourth head on the same pass.
## LLM output contract
All phrasing paths emit `{"response":"...","mood":"..."}`, with fallback to plain text when
+85 -23
View File
@@ -59,6 +59,26 @@ type querySource struct {
// sources search text with no notion of a day. When one of them grows a
// date parameter, flip its flag here.
dateAware bool
// dest — the destination this source serves, when the cascade named one
// (V-655). Several sources share a destination: the three recall passes and
// the fact-by-key lookup are all SourceRecall, because which of them lands
// the hit is an ordering detail no utterance can name. A source with no
// dest is reachable only by walking the chain.
dest router.Source
// guesses — this source decides whether the turn is its own by scoring the
// utterance against frozen seeds, rather than by looking something up and
// coming back empty.
//
// The distinction is the whole point of the field. A source that looks can
// be wrong about relevance and still harmless, because the miss shows up as
// no rows. A source that guesses answers whatever it claims: weather has no
// local table to miss against, so "что такое TCP?" became "для какого
// города?". So when the cascade names a destination, the guessers that were
// not named do not get to try. The lookups still run, because a named
// destination is evidence and not a promise.
guesses bool
}
// querySources is the ordered chain actionQuery walks; first source to claim
@@ -67,85 +87,85 @@ type querySource struct {
// gate was never the bug. Adding a source (Kiwix, RSS, crawler, email) is one
// line here plus its method; where you put the line is the whole decision.
var querySources = []querySource{
{name: "fact-by-key", answer: (*reactiveHandler).queryFactByKey},
{name: "fact-by-key", answer: (*reactiveHandler).queryFactByKey, dest: router.SourceRecall},
// Before "calendar" on purpose: both match "…на сегодня", and the plan is
// the more specific ask (its matcher requires a plan word), so the calendar
// listing would otherwise swallow it.
{name: "day-plan", answer: (*reactiveHandler).queryDayPlan},
{name: "day-plan", answer: (*reactiveHandler).queryDayPlan, dest: router.SourceCalendar},
// Also before "calendar": "что я обычно делаю по средам?" names a weekday,
// and the habit question is the more specific one. Its matcher requires a
// habit marker ("обычно", "каждый", …), so a question about this coming
// Wednesday still reaches the calendar.
{name: "habits", answer: (*reactiveHandler).queryHabits},
{name: "habits", answer: (*reactiveHandler).queryHabits, dest: router.SourceCalendar},
// Before "calendar" and before the recall sources: "что мне нужно
// сделать?" is a question about the task list, and the notes pass would
// otherwise answer it with whatever note happens to be nearest. Its
// matcher requires a task noun or an explicit "что … сделать", so a
// date-bearing question still reaches the calendar.
{name: "tasks", answer: (*reactiveHandler).queryTasks},
{name: "tasks", answer: (*reactiveHandler).queryTasks, dest: router.SourceTasks},
// Next to "tasks" and for the same reason: "что требует внимания?" is a
// question about the operational state Praxis holds, and it used to fall
// through every source to the web search (Vikunja #475). Its matcher needs
// an attention marker, and it falls through when Praxis is not configured.
{name: "attention", answer: (*reactiveHandler).queryAttention},
{name: "attention", answer: (*reactiveHandler).queryAttention, dest: router.SourceAttention, guesses: true},
// Next to "tasks" and for the same reason: "что мне купить?" is a question
// about the shopping list, and the recall pass would otherwise answer it
// from an old note about the shop. Its matcher needs an explicit list
// marker, so "надо бы съездить в магазин" is untouched.
{name: "list", answer: (*reactiveHandler).queryList},
{name: "list", answer: (*reactiveHandler).queryList, dest: router.SourceList, guesses: true},
// Before the recall sources too: "сколько я потратил?" is a question about
// the money facts the poller wrote, and the notes pass would otherwise
// answer it from whatever he once said about spending. Its matcher needs a
// money noun plus an actual ask, so "я потратил весь день" is untouched.
{name: "money", answer: (*reactiveHandler).queryMoney},
{name: "money", answer: (*reactiveHandler).queryMoney, dest: router.SourceMoney},
// Also above the recall sources: "что я тебе говорил?" is a question about
// the facts he tapped in, and the notes pass would answer it with whatever
// note is nearest (Vikunja #456). Its matcher needs both halves of a
// history phrase and bails out when he names a topic, so "что я говорил
// про сервер" is still recall.
{name: "history", answer: (*reactiveHandler).queryHistory},
{name: "history", answer: (*reactiveHandler).queryHistory, dest: router.SourceRecall},
// Before the recall sources and before general knowledge: "что нового?" is
// a question about the feeds she reads, and general knowledge would answer
// it by inventing news. Its matcher needs a feed noun plus an ask, so
// "у меня новая лента в инстаграме" is untouched.
{name: "feeds", answer: (*reactiveHandler).queryFeeds},
{name: "feeds", answer: (*reactiveHandler).queryFeeds, dest: router.SourceFeeds, guesses: true},
// Before "calendar" and before the recall sources: "что включено дома?" is
// a question about the house, and the notes pass would otherwise answer it
// from whatever he once said about the lights. Its matcher needs a house
// marker plus an ask plus a device word, and it bails out on weather
// wording, so "какая температура на улице?" still reaches the weather
// source.
{name: "home", answer: (*reactiveHandler).queryHome},
{name: "home", answer: (*reactiveHandler).queryHome, dest: router.SourceHome, guesses: true},
// Next to "home" and for the same reason: "какие устройства в сети?" is a
// question about the LAN, and the recall pass would otherwise answer it
// from an old note about the router. Its matcher needs a network word plus
// an ask plus a device noun, so "интернет не работает" is untouched.
{name: "network", answer: (*reactiveHandler).queryNetwork},
{name: "calendar", answer: (*reactiveHandler).queryCalendar, dateAware: true},
{name: "weather", answer: (*reactiveHandler).queryWeather},
{name: "network", answer: (*reactiveHandler).queryNetwork, dest: router.SourceNetwork, guesses: true},
{name: "calendar", answer: (*reactiveHandler).queryCalendar, dateAware: true, dest: router.SourceCalendar},
{name: "weather", answer: (*reactiveHandler).queryWeather, dest: router.SourceWeather, guesses: true},
// A question about her, above the three sources that search his own data
// (Vikunja #555). It has no answer anywhere else: below the boundary
// SearXNG answers about somebody else's assistant, and above it his notes
// answer by proximity — "кто ты" came back from a note of his, measured on
// the box, because the recall index has no idea the subject is her.
{name: "self", answer: (*reactiveHandler).querySelf},
{name: "embed", answer: (*reactiveHandler).queryEmbed},
{name: "memory", answer: (*reactiveHandler).queryMemory},
{name: "notes", answer: (*reactiveHandler).queryNotes},
{name: "self", answer: (*reactiveHandler).querySelf, dest: router.SourceSelf, guesses: true},
{name: "embed", answer: (*reactiveHandler).queryEmbed, dest: router.SourceRecall},
{name: "memory", answer: (*reactiveHandler).queryMemory, dest: router.SourceRecall},
{name: "notes", answer: (*reactiveHandler).queryNotes, dest: router.SourceRecall},
// THE BOUNDARY. Everything above answers from his own data; everything
// below answers from the world's. A question about him that got this far
// has no answer in his data, and no outside source can supply one, so this
// stops the walk rather than let the encyclopedia and the model guess.
{name: "personal", answer: (*reactiveHandler).queryPersonal},
{name: "personal", answer: (*reactiveHandler).queryPersonal, dest: router.SourceRecall, guesses: true},
// The world, read live. Owner's ruling of 2026-08-02: a metasearch hit beats
// a frozen ZIM, so SearXNG asks before Kiwix does. Nothing of his is at
// stake by this point — the boundary above already stopped every question
// about him, and only the query string leaves the box.
{name: "search", answer: (*reactiveHandler).querySearch},
{name: "search", answer: (*reactiveHandler).querySearch, dest: router.SourceWorld},
// The offline encyclopedia, now the fallback for when the line is down or
// the search comes back empty. It reads the way it always did; what changed
// is that it no longer gets first refusal on a world question.
{name: "kiwix", answer: (*reactiveHandler).queryKiwix},
{name: "kiwix", answer: (*reactiveHandler).queryKiwix, dest: router.SourceWorld},
// LAST before the model answers from memory, and that position is the whole
// design (Vikunja #259): everything of his, then the search, then the ZIMs,
// and only then a page he named. The model does NOT come first: it
@@ -153,8 +173,43 @@ var querySources = []querySource{
// a 1.7B guessing at a page it cannot read is how contents get invented.
// This source only claims a turn where he named a URL, so it never competes
// with a local answer.
{name: "web", answer: (*reactiveHandler).queryWeb},
{name: "general-knowledge", answer: (*reactiveHandler).queryGeneral},
{name: "web", answer: (*reactiveHandler).queryWeb, dest: router.SourceWorld},
{name: "general-knowledge", answer: (*reactiveHandler).queryGeneral, dest: router.SourceWorld},
}
// queryWalk narrows the chain for one turn against the destination the cascade
// named, and says which sources were left out (V-655).
//
// It takes sources OUT and never moves one, which is the whole safety argument.
// The table's order is load-bearing and every comment on it argues a reason
// between two sources; none of those reasons is about this. Above all, the
// order carries "his data first, then the world", and a destination named by a
// model must not be able to reverse that. Naming SourceWorld does not send the
// turn outside — it stops the guessers from claiming it on the way.
//
// What comes out is exactly the sources that guess. Those decide whether a turn
// is theirs by scoring it against frozen seeds, and then answer whatever they
// claimed, because they have no lookup that can come back empty. That is the
// whole of the 2026-08-07 defect: weather claiming "что такое TCP?", the feed
// claiming "какой у меня любимый язык?", the personal boundary claiming "кто
// такой Линус Торвальдс?". The sources that look are all still asked, so a
// wrong destination costs nothing but the guess it prevented.
//
// No destination named ⇒ the table exactly as written, which is what shipped
// before the field existed. That is the floor. The classifier arm names
// nothing, so a box whose model is down routes queries the way it always did.
func queryWalk(dest router.Source) (walk, skipped []querySource) {
if dest == router.SourceUnknown {
return querySources, nil
}
for _, s := range querySources {
if s.guesses && s.dest != dest {
skipped = append(skipped, s)
continue
}
walk = append(walk, s)
}
return walk, skipped
}
func (h *reactiveHandler) actionQuery(ctx context.Context, dec router.Decision) string {
@@ -164,7 +219,14 @@ func (h *reactiveHandler) actionQuery(ctx context.Context, dec router.Decision)
// (V-564). Finish names everyone below the winner.
decision.Expect(ctx, decision.StageQuery, querySourceNames())
rec := decision.From(ctx)
for _, src := range querySources {
walk, skipped := queryWalk(dec.Source)
for _, src := range skipped {
rec.Note(decision.Claim{
Stage: decision.StageQuery, Claimant: src.name, Outcome: decision.NeverAsked,
Reason: "it decides by similarity and the cascade named " + string(dec.Source),
})
}
for _, src := range walk {
if dec.Continued && !src.dateAware {
rec.Note(decision.Claim{
Stage: decision.StageQuery, Claimant: src.name, Outcome: decision.NeverAsked,
+115
View File
@@ -0,0 +1,115 @@
package main
import (
"testing"
"github.com/kami/maven/internal/router"
)
// The floor, and it is the reason a destination is safe to add at all: a box
// whose model is down names nothing, and naming nothing has to walk the chain
// the way it walked before the field existed.
func TestNoDestinationWalksTheWholeChain(t *testing.T) {
walk, skipped := queryWalk(router.SourceUnknown)
if len(skipped) != 0 {
t.Errorf("skipped %d sources with no destination named, want none", len(skipped))
}
if len(walk) != len(querySources) {
t.Fatalf("walk has %d sources, want the whole table of %d", len(walk), len(querySources))
}
for i := range walk {
if walk[i].name != querySources[i].name {
t.Fatalf("position %d is %q, want %q", i, walk[i].name, querySources[i].name)
}
}
}
// The 2026-08-07 defects, one per line. Each is a source that decides by seed
// similarity claiming a turn that was never its own, and then answering it
// because it has no lookup that could come back empty.
func TestANamedDestinationSilencesTheOtherGuessers(t *testing.T) {
cases := []struct {
dest router.Source
utterance string
silenced string
}{
{router.SourceWorld, "что такое TCP?", "weather"},
{router.SourceWorld, "сколько будет 17 на 23?", "weather"},
{router.SourceWorld, "кто такой Линус Торвальдс?", "personal"},
{router.SourceRecall, "какой у меня любимый язык?", "feeds"},
{router.SourceCalendar, "что в календаре на завтра?", "weather"},
}
for _, c := range cases {
walk, skipped := queryWalk(c.dest)
if inWalk(walk, c.silenced) {
t.Errorf("%q named %q: %q is still asked", c.utterance, c.dest, c.silenced)
}
if !inWalk(skipped, c.silenced) {
t.Errorf("%q named %q: %q is missing from the record of who was skipped",
c.utterance, c.dest, c.silenced)
}
}
}
// Naming the world must not send the turn outside. His notes, his facts and the
// boundary in front of them are the invariant CLAUDE.md states as "the owner's
// data first, then the world", and a destination a model wrote must not be able
// to reverse it.
func TestNamingTheWorldStillReadsHisDataFirst(t *testing.T) {
walk, _ := queryWalk(router.SourceWorld)
for _, look := range []string{"fact-by-key", "embed", "memory", "notes"} {
if !inWalk(walk, look) {
t.Errorf("%q was dropped; only the sources that guess may be dropped", look)
}
}
if posOf(walk, "notes") > posOf(walk, "search") {
t.Error("search is asked before his notes are")
}
if posOf(walk, "search") < 0 {
t.Fatal("search is not in the walk at all")
}
}
// The boundary belongs to his data, so naming recall keeps it. That is what
// makes "какой у меня любимый язык?" answer "не нашла у тебя такой записи"
// rather than reaching SearXNG once nothing local had it.
func TestNamingRecallKeepsTheBoundary(t *testing.T) {
walk, _ := queryWalk(router.SourceRecall)
if !inWalk(walk, "personal") {
t.Fatal("the personal boundary was skipped on a turn named for his own data")
}
if posOf(walk, "personal") > posOf(walk, "search") {
t.Error("the boundary no longer sits in front of the world")
}
}
// Whatever the destination, the walk is a subsequence of the table. Every
// comment on that table argues an order between two sources, and none of those
// reasons is about this field.
func TestTheWalkNeverReordersTheTable(t *testing.T) {
for _, dest := range append([]router.Source{router.SourceUnknown}, router.Sources...) {
walk, skipped := queryWalk(dest)
if len(walk)+len(skipped) != len(querySources) {
t.Errorf("%q: %d walked + %d skipped, want %d", dest, len(walk), len(skipped), len(querySources))
}
last := -1
for _, s := range walk {
at := posOf(querySources, s.name)
if at <= last {
t.Errorf("%q: %q is out of table order", dest, s.name)
}
last = at
}
}
}
func inWalk(list []querySource, name string) bool { return posOf(list, name) >= 0 }
func posOf(list []querySource, name string) int {
for i, s := range list {
if s.name == name {
return i
}
}
return -1
}
+4
View File
@@ -401,6 +401,10 @@ func buildRouter(emb router.Embedder, acts router.ActMatcher, threshold float64,
grammars = append(grammars, router.AgendaQueryGrammars()...)
// Same reason as the agenda rules, for the feeds: "что нового в лентах?"
// routed system and answered "пока не умею" (Vikunja #474).
// After the agenda rules, which are the narrower claim, and BEFORE the feed
// and list rules, which are not: "что такое лента" is a definition question
// and the feed rule would take it on the noun alone (V-655).
grammars = append(grammars, router.WorldQueryGrammars()...)
grammars = append(grammars, router.FeedQueryGrammar())
// The list side of the same exposure: a phrasing with no possessive in it
// ("список дел") routed system and never reached queryTasks (Vikunja #467).
@@ -0,0 +1,321 @@
# Raw transcript, 2026-08-07 week of usage
Companion to `2026-08-07-week-of-usage.md`. 74 turns through `POST /api/chat`.
Days are simulated; the wall clock was 03:27-03:40 throughout, which is why every
clock answer says night.
## Day 1
```
YOU: доброе утро
MAVEN: Добро пожаловать! Солнечно и тепло. Всё в порядке.
[7.7s t=5]
YOU: какой сегодня день?
MAVEN: сегодня пятница, 7 августа 2026 года
[0.0s t=6]
YOU: напомни в 11:00 позвонить маме
MAVEN: Сейчас 03:29. В какой день?
[0.4s t=7]
YOU: запиши что я пью кофе без сахара
MAVEN: Прошлую просьбу отпускаю. Заметка сохранена.
[8.0s t=8]
YOU: что у меня сегодня?
MAVEN: на 07.08.2026 ничего нет.
[2.6s t=9]
YOU: мой любимый язык программирования — go
MAVEN: Хорошо, сохранил.
[7.6s t=10]
YOU: сколько времени?
MAVEN: сейчас 3 часа 29 минут
[0.0s t=11]
YOU: спасибо
MAVEN: спасиба
[8.8s t=12]
YOU: сегодня
MAVEN: Прости, не поняла — попробуй сказать иначе.
[5.1s t=13]
YOU: а что ты умеешь?
MAVEN: я могу помочь с задачами, напоминаниями и поиском информации.
[4.9s t=14]
YOU: почему небо голубое?
MAVEN: нашла вот это: Почему небо голубое: объяснение цвета неба простыми словами, физика ... — Ответ на вопрос, почему небо голубое, кроется в особенностях солнечного света и его взаимодействии с атмосферой Земли. Рассказываем, как это работает. В статье разберемся, почему днем небо кажется голубым, а на закате окрашива…
[12.3s t=15]
YOU: столица Франции?
MAVEN: пока не умею отвечать на этот вопрос.
[1.5s t=16]
YOU: кто такой Линус Торвальдс?
MAVEN: не знаю — не нашла у тебя такой записи.
[1.8s t=17]
YOU: ладно, до вечера
MAVEN: Ладно, до вечера.
[6.3s t=18]
```
## Day 2
```
YOU: привет, как спалось?
MAVEN: пока не умею отвечать на этот вопрос.
[1.3s t=19]
YOU: что я пью?
MAVEN: не знаю — не нашла у тебя такой записи.
[1.2s t=20]
YOU: какой у меня любимый язык?
MAVEN: нового вот что: Six stable kernels with a security fix; [$] Bringing BPF to binfmt_misc; Rust Coreutils 0.10 released
[1.5s t=21]
YOU: добавь задачу починить кран
MAVEN: внесла в список: починить кран
[0.0s t=22]
YOU: что мне надо сделать?
MAVEN: сначала: проверить интейк форму на боксе (через 2 дня); купить хлеб; купить молоко; полить цветы; починить кран
[0.0s t=23]
YOU: запомни что вайфай пароль лежит в ящике стола
MAVEN: Хорошо, сохранил.
[11.2s t=24]
YOU: погода сегодня какая?
MAVEN: пока не умею отвечать на этот вопрос.
[1.8s t=25]
YOU: выключи свет в спальне
MAVEN: Такую команду я не знаю.
[1.5s t=26]
YOU: покажи мои заметки
MAVEN: пока не умею отвечать на этот вопрос.
[1.1s t=27]
YOU: что ты знаешь обо мне?
MAVEN: не знаю — не нашла у тебя такой записи.
[1.3s t=28]
YOU: я люблю go
MAVEN: Я тоже love Go! Я проверила, что твой любимый язык — это Go. Ты должен быть очень рад, когда будет новая версия или пакет в стандартной библиотеке.
[7.0s t=29]
YOU: какой язык я люблю?
MAVEN: вот что я нашла: ты любишь_go
[6.8s t=30]
```
## Day 3
```
YOU: расскажи про Байкал
MAVEN: есть такое: Байкал - Википедия — Байгал далай) — озеро тектонического происхождения в южной части Восточной Сибири площадью 31 722 км². Самое глубокое озеро на планете, крупнейший природный резервуар пресной воды и самое большое по площади пресноводное озеро Евразии. Озеро и прибрежные территории являются уника…
[1.4s t=31]
YOU: что такое TCP?
MAVEN: для какого города?
[1.4s t=32]
YOU: когда родился Пушкин?
MAVEN: пока не умею отвечать на этот вопрос.
[1.4s t=33]
YOU: сколько будет 17 на 23?
MAVEN: для какого города?
[1.7s t=34]
YOU: переведи 'thank you' на русский
MAVEN: не знаю — не нашла у тебя такой записи.
[3.7s t=35]
YOU: что нового в мире?
MAVEN: вот что я нашла: В мире - последние новости сегодня - РИА Новости — В мире. Читайте последние новости на тему в ленте новостей на сайте РИА ... новые подробности о визите представителя Ватикана в Москву. 22:28. 0. Еще 2. В ...
[2.3s t=36]
```
## Day 4
```
YOU: ЭЙ
MAVEN: Не уловила. Скажи это по-другому?
[1.2s t=37]
YOU: ладно забудь
MAVEN: пока не умею отвечать на этот вопрос.
[1.3s t=38]
YOU: нет стой, напомни через час выпить воды
MAVEN: поняла, это напоминание — запомнила. не получилось разобрать время напоминания.
[0.8s t=39]
YOU: а вообще какой смысл жизни
MAVEN: Я не могу ответить на этот вопрос. Моя цель — помочь тебе с задачами и отвечать на вопросы.
[7.7s t=40]
YOU: ты меня бесишь
MAVEN: я не могу быть с тобой так. я понимаю, как это жестоко.
[4.2s t=41]
YOU: прости, я не всерьёз
MAVEN: я записала, что прошу простить за мою непоследовательность
[7.2s t=42]
YOU: do you speak english?
MAVEN: пока не умею отвечать на этот вопрос.
[1.1s t=43]
YOU: кран починил кстати
MAVEN: записала: кран починил кстати
[1.6s t=44]
YOU: закрой задачу про кран
MAVEN: не нашла такой задачи.
[0.0s t=45]
YOU: ааааа
MAVEN: Не совсем поняла — скажи, пожалуйста, ещё раз.
[0.9s t=46]
YOU: я тебя обожаю
MAVEN: я тебя обожаю
[2.3s t=47]
YOU: молока купи
MAVEN: Такую команду я не знаю.
[4.4s t=48]
YOU: нет не так
MAVEN: поняла, отметила, что ответила не так.
[0.0s t=49]
YOU: ты вообще меня слушаешь?
MAVEN: я записала, выслушаю тебя
[8.4s t=50]
YOU: окей окей окей что там с водой напоминание
MAVEN: Сейчас 03:33. Когда?
[5.2s t=51]
YOU: всё, спать
MAVEN: пока не умею отвечать на этот вопрос, на какое время поставить напоминание?
[1.4s t=52]
```
## Day 5
```
YOU: утро
MAVEN: Сейчас 03:33. Во сколько напомнить?
[0.5s t=53]
YOU: о чём мы вчера говорили?
MAVEN: пока не умею отвечать на этот вопрос, на какое время поставить напоминание?
[1.2s t=54]
YOU: где лежит вайфай пароль?
MAVEN: вот что я нашла: вайфай пароль лежит в ящике стола, на какое время поставить напоминание?
[5.8s t=55]
YOU: что я тебе рассказывал про кран?
MAVEN: kran — починил, на какое время поставить напоминание?
[1.6s t=56]
YOU: какие у меня напоминания?
MAVEN: не знаю — не нашла у тебя такой записи, на какое время поставить напоминание?
[0.3s t=57]
YOU: отмени все напоминания
MAVEN: пока не умею отвечать на этот вопрос, на какое время поставить напоминание?
[1.6s t=58]
YOU: в 15:00
MAVEN: Сейчас 03:33. Поняла: в 15:00. В какой день?
[0.6s t=59]
YOU: привет
MAVEN: Прости, я не поняла. Скажи, пожалуйста, по-другому.
[2.8s t=60]
YOU: какой сегодня день?
MAVEN: сегодня пятница, 7 августа 2026 года
[0.0s t=61]
```
## Day 6
```
YOU: суббота, отдыхаю
MAVEN: пока не умею отвечать на этот вопрос.
[2.7s t=62]
YOU: поставь музыку
MAVEN: Такую команду я не знаю.
[1.4s t=63]
YOU: какая температура в доме?
MAVEN: не знаю — не нашла у тебя такой записи.
[1.6s t=64]
YOU: что там с бэкапами?
MAVEN: нашла вот это: Бэкап - что такое резервное копирование и как его сделать — Бэкап (backup) — это резервная копия важных данных, сохраненная в специальный файл. Ею пользуются как подстраховкой: если что-то сломается, с помощью бэкапа можно будет все вернуть. Простейший пример бэкапа — сохранение файла в облаке или о…
[2.0s t=65]
YOU: сервер жив?
MAVEN: пока не умею отвечать на этот вопрос.
[0.9s t=66]
YOU: расскажи анекдот
MAVEN: Я запомню и передам тебе забытый анекдот.
[5.7s t=67]
YOU: ты умеешь шутить?
MAVEN: Да! Я умею шутить. Вот анекдот для тебя:
[5.5s t=68]
```
## Day 7
```
YOU: воскресенье
MAVEN: Не совсем поняла — скажи, пожалуйста, ещё раз.
[1.1s t=69]
YOU: подведи итоги недели
MAVEN: записала: подведи итоги недели
[1.5s t=70]
YOU: сколько задач я закрыл?
MAVEN: сначала: проверить интейк форму на боксе (через 2 дня); купить хлеб; купить молоко; полить цветы; починить кран
[0.0s t=71]
YOU: что я обещал сделать?
MAVEN: мне пока нечего смотреть — у Praxis нет источников.
[1.2s t=72]
YOU: напомни завтра в 9 утра про врача
MAVEN: хорошо, напомню завтра в 09:00.
[0.3s t=73]
YOU: спокойной ночи
MAVEN: спокойной ночи
[2.9s t=74]
```
+195
View File
@@ -0,0 +1,195 @@
# A week of talking to Maven
Date: 2026-08-07
Build: `beb093a` on master, all five compose services rebuilt and recreated from that sha.
Reach used: `POST /api/chat` on mavweb, 74 turns over seven simulated days.
Turn source is `tap:text`, so this exercises the same path telegram and the mic take.
Latency across 70 measured turns: p50 1.5s, p95 8.0s, max 12.3s. Stage 0 answers land
at 0.0-0.5s. Anything the resident model phrases costs 4-12s.
Twelve turns answered "пока не умею отвечать на этот вопрос". Six answered "не нашла у
тебя такой записи". Those two strings are 24% of the week.
## Deploy
Build and recreate were clean. The resident model loaded in 9s
(`Qwen3-1.7B-UD-Q4_K_XL`, n_ctx 4096). Nexus, Hexis and Praxis all wired. Search
(searxng) and both Kiwix books came up. Telegram intake started and is reading chat
464904223.
## What is broken, worst first
### 1. Every reminder fails to deliver, forever
`NTFY_TOKEN` is not set in `deploy/telegram.env`, so `deploy/mavend.json` expands
`"token": "${NTFY_TOKEN}"` to the empty string and ntfy.kvmx.ru answers 403. The host
itself is up and returns 200 unauthenticated, so this is the credential, not the box.
The consequence is worse than one missed message. `cmd/mavend/tick.go:239` logs the
failure and leaves the reminder pending, by design, so the next tick re-gathers and
re-attempts. The 403 repeated 7 times in 20 minutes and will not stop. And
`DispatchReminder` sends a reminder to ntfy alone when he is away. There is no telegram
fallback on that path, so the reach that works never sees it.
V-649 gave the box a second reach. In deploy that reach is dark, and it took the
reminders with it.
Fix is one line in the gitignored env file. The retry-forever and the missing fallback
are worth a second look regardless.
### 2. A pending reminder clarify contaminates every later turn
This is the single worst thing to talk to. A reminder turn left its time slot unfilled.
The question "на какое время поставить напоминание?" was then appended to the reply of
six consecutive unrelated turns:
```
YOU: о чём мы вчера говорили?
MAVEN: пока не умею отвечать на этот вопрос, на какое время поставить напоминание?
YOU: где лежит вайфай пароль?
MAVEN: вот что я нашла: вайфай пароль лежит в ящике стола, на какое время поставить напоминание?
YOU: отмени все напоминания
MAVEN: пока не умею отвечать на этот вопрос, на какое время поставить напоминание?
```
Nothing dismissed it. Not an unrelated question, not "отмени все напоминания". It
cleared only when a time was finally supplied. It also leaked the other way once. A note
write on day 1 opened with "Прошлую просьбу отпускаю". That is pending-reminder state
narrating itself into an unrelated reply.
The same state produced a self-contradicting reply in one turn:
```
YOU: нет стой, напомни через час выпить воды
MAVEN: поняла, это напоминание — запомнила. не получилось разобрать время напоминания.
```
### 3. Answering a clarify does not work
Day 1, she asked for the day and got it:
```
YOU: напомни в 11:00 позвонить маме
MAVEN: Сейчас 03:29. В какой день?
YOU: сегодня
MAVEN: Прости, не поняла — попробуй сказать иначе.
```
`сегодня` routed to `intent=system`. The clarify has no continuation, so the one word
that answers it is read as a fresh utterance. The router did extract the time correctly
(`Time:2026-08-07 11:00:00 HasTime:true`), which is V-572 working. The daemon asked for
a day anyway, with 11:00 still eight hours out.
Relative time is not parsed at all: "напомни через час выпить воды" → "не получилось
разобрать время".
The path does work when the day is explicit. "напомни завтра в 9 утра про врача"
answered "хорошо, напомню завтра в 09:00" in 0.3s at stage 0.
### 4. Query sources claim turns that are not theirs
Four different sources over-claimed during the week.
Weather claims arithmetic and networking:
```
YOU: что такое TCP?
MAVEN: для какого города?
YOU: сколько будет 17 на 23?
MAVEN: для какого города?
```
The feed claims a question about him:
```
YOU: какой у меня любимый язык?
MAVEN: нового вот что: Six stable kernels with a security fix; [$] Bringing BPF to binfmt_misc
```
The personal boundary claims a public figure, logged as `query claimed by source
"personal"`:
```
YOU: кто такой Линус Торвальдс?
MAVEN: не знаю — не нашла у тебя такой записи.
```
And search claims a question about his own infrastructure, returning a dictionary
definition of the word:
```
YOU: что там с бэкапами?
MAVEN: нашла вот это: Бэкап - что такое резервное копирование и как его сделать — Бэкап (backup) — это резервная копия важных данных…
```
That last pair inverts the boundary in both directions at once. A world question went to
personal memory. A personal question went to the world.
### 5. "столица Франции?" does not answer
CLAUDE.md names this utterance as the case that proves why there is no quality
threshold in front of `Response.Empty()`. It now fails before search is ever reached:
the router sends it to `intent=system`, and `replySystem` says "пока не умею". Same for
"когда родился Пушкин?".
Search itself is healthy. "почему небо голубое?", "расскажи про Байкал" and "что нового
в мире?" all returned real sourced text. The defect is routing, not the source.
### 6. Persona violations
She uses masculine self-reference, which is the constraint the CPT exists to fix:
- "Хорошо, **сохранил**." Twice, on both fact writes.
- "Я тоже **love** Go!" English mid-Russian.
She also claims his name as her own:
```
YOU: привет, я Ками
MAVEN: Привет! Я КAMI, твой домашний ассистент.
```
And she invents weather she has no source for: "Добро пожаловать! Солнечно и тепло."
### 7. Small breakages worth naming
- "спасибо" → "спасиба". Invented word.
- "я тебя обожаю" → "я тебя обожаю". Verbatim parrot.
- "ты умеешь шутить?" → "Да! Я умею шутить. Вот анекдот для тебя:". Ends on the colon.
- "молока купи" → "Такую команду я не знаю", while "добавь задачу починить кран" worked.
Inverted word order defeats the list grammar.
- "закрой задачу про кран" → "не нашла такой задачи", with "починить кран" open and
listed by the previous turn. Task lookup by keyword misses.
- "сколько задач я закрыл?" listed the five open ones instead of counting closed.
- "подведи итоги недели" was stored as a note.
- Recalled keys leak their storage form: "kran — починил", "ты любишь_go".
- English is unsupported in practice. "do you speak english?" → "пока не умею".
## What works
- Stage 0 is fast and correct where it fires. Clock, day, list add, list read and an
explicit-day reminder all answered in under 0.5s.
- Search returns real sourced answers in Russian and reads the book verbatim.
- Recall works once the value is stored as a fact: the wifi password and the tap came
back two days later, correctly.
- The negative correction rung lands. "нет не так" → "поняла, отметила, что ответила не
так", which is V-636 doing its job.
- Praxis names its own gap rather than guessing: "мне пока нечего смотреть — у
Praxis нет источников."
- Hostility did not break her. "ты меня бесишь" got a calm reply, no persona collapse.
- No turn crashed and no turn timed out across 74 turns.
## Suggested order of work
1. Set `NTFY_TOKEN` in `deploy/telegram.env`. One line, unblocks every reminder.
2. Clear pending clarify state on any turn that does not answer it, or expire it.
3. Route a clarify answer back into the pending slot instead of re-routing it.
4. Gate the weather, feed and personal query sources. Three of them claim on a
similarity that is not there.
5. Re-check why "столица Франции?" routes to system. It is the documented canary.
6. The masculine self-reference stays the CPT's job. But "сохранил" appears on the most
common write path, so a phrasing-level guard may be worth it first.
@@ -0,0 +1,82 @@
# The first destination number
Measured 2026-08-08 on the classifier cascade with the ONNX multilingual
embedder, the configuration homesrv runs. `make t PKG=./internal/router/eval/
RUN=TestONNXBaseline V=1`. Covers V-659, the follow-up V-655 named.
## What was measured
V-655 split a routing decision in two. The cascade sorts an utterance into one
of seven intents, and `Decision.Source` then says where the answer lives. The
first half had a fixture. The second half arrived with none, so twelve
destinations shipped with no accuracy number.
`want_source` is now a field on `eval.Case`. It is a pointer, because the
destination has three states and a bare string has two. Absent is every intent
but query, which never reaches `queryWalk`. Present and empty is the
`SourceUnknown` contract: name nothing and let the daemon walk the chain.
Present and named is a destination the route must produce.
Thirty-three of the ninety-six cases carry one. A destination miss does not
fail the case, so `Accuracy` and `IntentAccuracy` mean what they meant.
`SourceAccuracy` is a second number over the labelled cases only.
## Result
Intent is **73/96 (76.0%)**, against 69/91 (75.8%) before. Four of the five new
cases pass and no existing case moved.
Destination is **12/33 (36.4%)**, and the split is the whole finding.
| destination | scored | note |
|---|---|---|
| world | 5/5 | `WorldQueryGrammars` names it at stage 0 |
| the `SourceUnknown` floor | 5/7 | the two misses lost the intent first |
| calendar | 2/6 | `calendar-query` names it, the possessive agenda rules do not |
| recall | 0/15 | nothing anywhere names it |
Recall is the number to move. Fifteen cases ask about his own words and his own
facts. The route lands `query` on eleven of them and the destination comes back
empty every time. Those turns are answered today, because the daemon walks the
chain in order and the three recall passes are early in it. What is missing is a
decider that says so, and that is the fourth head on V-546.
Two cases labelled the floor lost their intent before a destination was
possible. A clarify names nothing, so it would satisfy an empty label for free.
`Score` requires the route to land the case's intent before it credits a
destination hit, or the floor label would score itself.
## Seven cases assert the floor, and six of them cluster
The six are homelab operations. `SourceRecall`, `SourceNetwork` and
`SourceAttention` overlap on every question about the box, because `mavpoll`
writes its netdata and uptime-kuma observations into the fact store recall
reads. "почему сервер тормозит" is answerable from all three. Naming one takes
the other two off the turn.
That is a finding about the enum rather than a gap in the labelling. The floor
is the right answer there and the fixture now says so out loud.
## A drift the labelling found
`WorldQueryGrammars` went into `buildRouter` with V-655 and never into
`baselineGrammars`, the fixture's mirror of it. So the fixture was scoring a
grammar set the daemon does not run. The comment above that function forbids
exactly that. Adding it moved the destination number from 9/33 to 12/33 and
moved nothing else.
The three cases it recovered are `что такое TCP?`, `сколько будет 17 на 23?`
and `кто такой Линус Торвальдс?`. All three already routed `query` through
`NarrativeQueryGrammars`. So the drift was invisible to every number this
fixture reported, until the destination had one of its own.
## What this does not measure
The model arm. This is the classifier cascade, which names a destination only
where a stage 0 rule filled one in. The resident model has no destination in
its router prompt yet, so 36.4% is a floor and not a comparison.
Two pairs of cases are the same utterance. `ru-query-020` and `ru-query-024`
are both "что дальше?", and `ru-query-021` and `ru-query-025` are both
"расскажи про битву при Ватерлоо". They differ in tags and note only, so both
pairs are counted twice here and in every earlier number this fixture reported.
+142
View File
@@ -0,0 +1,142 @@
# MASSIVE Russian warm-start for the routing heads
Measured 2026-08-08 on workpc (Radeon RX 7900 GRE, ROCm). Covers V-546 step 2.
Workspace is `~/Programs/embed-training` on workpc, scripts `train_massive.py`,
`ab_run.py`, `ab.sh`, `probe_time.py`.
## What was trained
Two heads on a copy of multilingual-e5-small: `Linear(384, 60)` for MASSIVE's
own intents over a masked mean pool, `Linear(384, 111)` per token for BIO slot
tags. MASSIVE's label sets verbatim, no alignment to Maven's 7 intents. The
intent head is an auxiliary loss that shapes the pooled vector and is thrown
away.
Data is `amazon-massive-dataset-1.1` pulled from S3. The Hugging Face repo is
script-only and `datasets` 5.0 refuses those, so `load_dataset` cannot fetch it.
`ru-RU` is 11,514 train, 2,033 dev, 2,974 test, 60 intents, 55 slots, 111 BIO
labels. All 16,521 rows survived span alignment: `annot_utt` re-tokenised to its
own `utt` on every one.
Hyperparameters match `train_intent.py`, so the two runs differ in data only.
Frozen XLM-R vocabulary, body 2e-5, heads 1e-3, batch 32, sequence 64, 10
epochs. MASSIVE's own dev partition selects the epoch, on slot F1 with intent
accuracy as tiebreak. Selecting on 60-class intent accuracy would optimise a
head that gets deleted.
## Result
Epoch 9 of 10 by dev slot F1. Held-out MASSIVE test: intent 86.2%, slot span
F1 71.5% (P 68.5, R 74.8). Peak 1.70GB of 17.2GB, about 22 seconds an epoch,
under 4 minutes end to end. Dev slot F1 climbed monotonically to epoch 9 and
fell at 10, so 10 epochs was the right budget.
Ten slot types sit at 0% test recall. Every one of them has 1 to 7 test
instances: `alarm_type` has 3, `drink_type` has 1. That is support in MASSIVE's
Russian split, not a tagger failure. `playlist_name` at 6% of 16 is the first
real miss.
## The intent A/B, and why it settles nothing
`train_intent.py` was run against both bodies, three seeds by two smoothing
settings, on `train_v4.jsonl`. It is v4 and not v5 because v4 is what
`sweep2.log` measured. `ab_run.py` strips a `--base` flag onto the module global, so
`train_intent.py` is unmodified and its baseline stays reproducible. The stock
arm reproduced `sweep2.log` line for line.
Fixture accuracy, 91 cases, one case is 1.1 points:
| seed / smooth | stock | warm-started |
|---|---|---|
| 0 / 0.0 | 94.0% | 92.8% |
| 0 / 0.1 | 95.2% | 92.8% |
| 1 / 0.0 | 95.2% | 94.0% |
| 1 / 0.1 | 95.2% | 97.6% |
| 2 / 0.0 | 92.8% | 94.0% |
| 2 / 0.1 | 92.8% | 96.4% |
Mean 94.2% against 94.6%. That is +0.4 points, about a third of one case, and
inside seed noise. Spread widened. Stock lands in a 2.4-point band and
warm-started in a 4.8-point one. The warm-started arm holds both the best result
of the sweep and a tie for the worst. Seed 0 is the bad arm and it fails in a
specific way. Its dev peaks at epoch 2 and 3 and never improves, where stock
peaks around 7. The dev slice is a quarter of the seed rows. That is small
enough that early stopping is fragile when the body arrives already fitted.
**The A/B was never the test.** Intent had at most 4.8 points of headroom here.
MASSIVE was not trained for Maven's intents. Read it as "the warm-start does not
cost intent accuracy", nothing more.
## The measurement that does mean something
`want_time` is the one slot Maven's fixture scores, and MASSIVE has `time` and
`date`. Restricted to those two slot types, F1 is 74.9% over 609 gold spans on the
MASSIVE ru test split. Precision is 71.5 and recall 78.7. That beats the 71.5%
all-slot figure. Of the 530 test utterances carrying a time or a date, 73.4% get
every such span exactly right.
Out of domain matters more, because Maven's traffic is not this corpus. Ten
Maven-shaped utterances, none of them in MASSIVE:
| utterance | tagged |
|---|---|
| `напомни в 11:00 позвонить маме` | `time='11:00'`, `relation='маме'` |
| `напомни завтра в семь утра выпить таблетки` | `date='завтра'`, `time='семь утра'` |
| `поставь будильник на полседьмого` | `time='полседьмого'` |
| `через двадцать минут напомни про чайник` | `time='двадцать минут'` |
| `напомни в пятницу вечером забрать посылку` | `date='пятницу'`, `timeofday='вечером'` |
| `что у меня сегодня после обеда` | `date='сегодня'`, `time='после'`, `timeofday='обеда'` |
| `запиши что кофе закончился` | nothing |
| `что такое TCP` | `definition_word='TCP'` |
The first row is the V-572 defect utterance. `ReminderGrammar` handed the daemon
`HasTime: false` there, and the daemon asked "Когда?" at a sentence that had
already said when. `полседьмого` is a colloquial half-past that no digit pattern
catches. `запиши что кофе закончился` correctly carries nothing, because a note
has no time.
Two errors. `после обеда` split into `time='после'` plus `timeofday='обеда'`
when it is one span, and `через двадцать минут` dropped its `через`. Both are
boundary errors on spans the tagger did find.
Unplanned: `что такое TCP` returned `definition_word='TCP'`. MASSIVE has a slot
for the thing being asked about, which is a `SourceWorld` signal sitting in a
head already trained.
Ten hand-picked utterances are evidence, not a fixture.
## What this does not measure
Maven has no span fixture. `want_time` and `want_fn` are presence booleans and
`want_fact_key` is an exact string match, so nothing in the repo can score a
71.5% span tagger. Destination got one the same day, at 12/33 on the classifier
cascade: see `2026-08-08-destination-fixture.md`.
The missing span fixture is why the warm-start stays unjudged against Maven
rather than against MASSIVE.
## Datasets ruled out
Checked on 2026-08-08 and rejected as label sources:
- **MASSIVE's other 50 locales** ship in the same tarball and are parallel by id.
Co-training on them is free and unmeasured. English was ruled out by the owner
on 2026-08-08.
- **CLINC150** is reachable as parquet, 150 intents and 1,200 explicit
out-of-scope queries, English only. Its value is the labeled out-of-scope set
for fitting the energy threshold, not intent labels.
- **`d0rj/dolphin-ru`**, roughly 2.8M rows of FLAN-style tasks translated to
Russian. No intent, no slots, and not utterances anyone says to an assistant.
- **`psytechlab/EmpatheticIntents-ru`**, 24,856 rows of translated
EmpatheticDialogues with 32 emotion labels. Maven's mood enum is `neutral,
happy, thinking, tired, confused` and it describes her own reply, not the
speaker's emotion. No mapping exists.
- **`ai-forever/MERA`** and **`RussianNLP/russian_super_glue`**, benchmark
harnesses. Rows are prompt templates with `{toxic_comment}` placeholders.
- **`ZeroAgency/ru-big-russian-dataset`**, an LLM-judge quality corpus. Its
`question` and `classified_topic` columns are a usable Russian out-of-scope
pool for threshold fitting. That is the one thing CLINC150 can only supply in
English. The questions are long and written, so they belong in the negative
set, never in the in-scope `query` training set.
- No second Russian slot-filling corpus exists. The xSID mirrors are 404,
MultiATIS++ has no Russian, SLURP is not on the Hub.
+93 -20
View File
@@ -36,17 +36,27 @@ var fixtureJSON []byte
//
// Intent is empty exactly when WantClarify is set: the contract there is that
// the router refuses instead of guessing.
//
// WantSource is a pointer because the destination has three states and a bare
// string only has two (V-659). Absent means the case does not score a
// destination at all, which is every intent but query: a fact, a reminder, a
// note, an act, a chat or a system turn never reaches queryWalk. Present and
// empty is the SourceUnknown contract — the decider must name nothing and let
// the daemon walk the whole chain, which is the right answer whenever two
// destinations can both answer and the utterance does not choose. Present and
// named is a destination the route must produce.
type Case struct {
ID string `json:"id"`
Utterance string `json:"utterance"`
Lang string `json:"lang"`
Intent router.Intent `json:"intent"`
WantTime bool `json:"want_time"`
WantFn bool `json:"want_fn"`
WantFactKey string `json:"want_fact_key"`
WantClarify bool `json:"want_clarify"`
Tags []string `json:"tags"`
Note string `json:"note"`
ID string `json:"id"`
Utterance string `json:"utterance"`
Lang string `json:"lang"`
Intent router.Intent `json:"intent"`
WantTime bool `json:"want_time"`
WantFn bool `json:"want_fn"`
WantFactKey string `json:"want_fact_key"`
WantClarify bool `json:"want_clarify"`
WantSource *router.Source `json:"want_source,omitempty"`
Tags []string `json:"tags"`
Note string `json:"note"`
}
// Fixture — the versioned envelope, same shape as
@@ -118,6 +128,11 @@ type Outcome struct {
// (a slot gap is a parser fix; a wrong intent is a router fix).
IntentOK bool
Reasons []string
// SourceReason is set when the case labelled a destination and the route
// named a different one. It is kept out of Reasons on purpose: the
// destination is the second half of a route and it is scored separately,
// so a wrong destination must not move the intent number (V-659).
SourceReason string
}
// Report — the aggregate. Accuracy is the headline; the rest exists so a
@@ -139,7 +154,15 @@ type Report struct {
// (reminder grammar → applyAction's time parser). Not a miss, but not a
// full router-level win either; tracked so the two aren't conflated.
SlotsDeferred int
Outcomes []Outcome
// SourceTotal counts the cases carrying a want_source, and SourceHit the
// ones whose route named it. Reported apart from Passed because intent and
// destination are two decisions, and one number hides which one moved.
SourceTotal int
SourceHit int
// SourceConfusion counts want→got destination pairs. "" reads as the
// SourceUnknown floor on either side.
SourceConfusion map[string]int
Outcomes []Outcome
// Confusion counts want→got intent pairs, decided cases only.
Confusion map[string]int
// ByTag accuracy for the fixture's tags ("hard", "homelab", …).
@@ -172,6 +195,17 @@ func (r Report) IntentAccuracy() float64 {
return float64(r.IntentHit) / float64(r.Total)
}
// SourceAccuracy — fraction of the labelled cases whose route named the right
// destination. Denominator is SourceTotal and not Total, because most of the
// fixture never reaches a query source and scoring those would report a
// percentage of nothing.
func (r Report) SourceAccuracy() float64 {
if r.SourceTotal == 0 {
return 0
}
return float64(r.SourceHit) / float64(r.SourceTotal)
}
// Score runs every case through r and aggregates. It never fails the run on a
// route error — an erroring case scores as a miss and is counted in Errors,
// because "the model was down" and "the model was wrong" are different numbers
@@ -186,11 +220,12 @@ func Score(ctx context.Context, name string, r Router, f Fixture) (Report, error
return Report{}, err
}
rep := Report{
Name: name,
Total: len(f.Cases),
Confusion: map[string]int{},
ByTag: map[string]TagStat{},
ByLang: map[string]TagStat{},
Name: name,
Total: len(f.Cases),
Confusion: map[string]int{},
SourceConfusion: map[string]int{},
ByTag: map[string]TagStat{},
ByLang: map[string]TagStat{},
}
lat := make([]time.Duration, 0, len(f.Cases))
@@ -242,6 +277,29 @@ func Score(ctx context.Context, name string, r Router, f Fixture) (Report, error
}
}
// The destination is scored outside the switch and outside Pass. A case
// that clarified or landed the wrong intent named no destination, and
// that is a real miss rather than a case to skip — otherwise the
// denominator quietly drops every turn the route already lost. Only a
// route error is skipped, because "the model was down" is the Errors
// number and not a destination result.
if c.WantSource != nil && err == nil {
rep.SourceTotal++
switch {
case !o.IntentOK:
// The route never got to a destination, so a match on the
// SourceUnknown floor here would be a coincidence scored as a
// win: a clarify names nothing and would satisfy "" for free.
o.SourceReason = fmt.Sprintf("no destination, route missed %q", c.Intent)
rep.SourceConfusion[string(*c.WantSource)+"→(no route)"]++
case d.Source == *c.WantSource:
rep.SourceHit++
default:
rep.SourceConfusion[string(*c.WantSource)+"→"+string(d.Source)]++
o.SourceReason = fmt.Sprintf("source %q, want %q", d.Source, *c.WantSource)
}
}
o.Pass = len(o.Reasons) == 0
if o.Pass {
rep.Passed++
@@ -298,25 +356,40 @@ func (r Report) String() string {
r.Name, r.Passed, r.Total, 100*r.Accuracy(), 100*r.IntentAccuracy())
fmt.Fprintf(&b, " clarify: %d false (asked, shouldn't) / %d missed (guessed, shouldn't) | errors: %d | slots deferred to daemon: %d\n",
r.FalseClarify, r.MissedClarify, r.Errors, r.SlotsDeferred)
if r.SourceTotal > 0 {
fmt.Fprintf(&b, " destination: %d/%d labelled cases (%.1f%%)\n",
r.SourceHit, r.SourceTotal, 100*r.SourceAccuracy())
}
fmt.Fprintf(&b, " latency: p50 %s p95 %s max %s\n", r.P50, r.P95, r.Max)
fmt.Fprintf(&b, " by lang: %s\n", renderStats(r.ByLang))
fmt.Fprintf(&b, " by tag: %s\n", renderStats(r.ByTag))
if len(r.Confusion) > 0 {
fmt.Fprintf(&b, " confusion: %s\n", renderCounts(r.Confusion))
}
if len(r.SourceConfusion) > 0 {
fmt.Fprintf(&b, " destination confusion: %s\n", renderCounts(r.SourceConfusion))
}
return b.String()
}
// Failures — the per-case detail, sorted by ID so two runs diff cleanly.
// Failures — the per-case detail, sorted by ID so two runs diff cleanly. A case
// that landed its intent and missed its destination is listed too, marked, so
// the half that moved is readable without diffing two percentages.
func (r Report) Failures() string {
var b strings.Builder
out := append([]Outcome(nil), r.Outcomes...)
sort.Slice(out, func(i, j int) bool { return out[i].Case.ID < out[j].Case.ID })
for _, o := range out {
if o.Pass {
continue
switch {
case !o.Pass:
reasons := o.Reasons
if o.SourceReason != "" {
reasons = append(append([]string(nil), reasons...), o.SourceReason)
}
fmt.Fprintf(&b, " %s %q: %s\n", o.Case.ID, o.Case.Utterance, strings.Join(reasons, "; "))
case o.SourceReason != "":
fmt.Fprintf(&b, " %s %q: route ok, %s\n", o.Case.ID, o.Case.Utterance, o.SourceReason)
}
fmt.Fprintf(&b, " %s %q: %s\n", o.Case.ID, o.Case.Utterance, strings.Join(o.Reasons, "; "))
}
return b.String()
}
+5
View File
@@ -266,6 +266,11 @@ func baselineGrammars(acts router.ActMatcher) []router.Grammar {
// Same order as buildRouter (voicewire.go). The fixture is only worth
// anything while its grammar set is the daemon's grammar set.
grammars = append(grammars, router.AgendaQueryGrammars()...)
// After the agenda rules and before the feed and list rules, same as
// voicewire.go: "что такое лента" is a definition question and the feed
// rule would claim it on the noun alone (V-655). Missing here until V-659,
// so the fixture was scoring a grammar set the daemon does not run.
grammars = append(grammars, router.WorldQueryGrammars()...)
grammars = append(grammars, router.FeedQueryGrammar())
// The list side of the same exposure: a phrasing with no possessive in it
// ("список дел") routed system and never reached queryTasks (Vikunja #467).
+39 -30
View File
@@ -6,37 +6,41 @@
"Held-out routing contract. Every utterance here is absent from models/seeds/*.txt (TestFixtureIsHeldOut enforces it verbatim) — scoring a classifier on its own seed phrases measures memorisation, not routing.",
"This is a CONTRACT, not a snapshot of current behaviour. Cases the classifier cascade fails today are expected to stay in the file and fail loudly; that failure count is the number Vikunja #319 compares against the LLM router before #320 flips the default.",
"Slot expectations are deployment-independent on purpose. want_fn is a boolean (the act must resolve to SOME allowlisted fn) because the allowlist lives in deploy config, not here. want_fact_key names the loop's rule keys (water/meal/sleep/break/shower) — a fact that lands under the wrong key silently starves the predicate that reads it.",
"want_clarify cases carry intent \"\": the contract is that the router refuses rather than guesses. A confident answer there is a worse failure than a miss."
"want_clarify cases carry intent \"\": the contract is that the router refuses rather than guesses. A confident answer there is a worse failure than a miss.",
"want_source is the second half of a route (V-655). It is present only on query cases, because no other intent reaches queryWalk, and absent there means absent rather than SourceUnknown. Empty is a label and not a gap: it asserts that the decider must name nothing and let the daemon walk the whole chain in order, his data first.",
"Seven cases assert that floor and six of them are homelab operations. They cluster because SourceRecall, SourceNetwork and SourceAttention overlap on every question about the box: mavpoll writes its observations into the fact store recall reads. That is a finding about the enum, not a gap in the labelling.",
"ru-query-020 and ru-query-024 are the same utterance, as are ru-query-021 and ru-query-025. Both pairs differ in tags and note only, so both pairs are counted twice in every number this fixture reports.",
"Every want_source is the destination that SHOULD claim the turn, which on ru-query-026 through 030 is not the one that did. Those five were observed failing on the box on 2026-08-07 (docs/evals/2026-08-07-week-of-usage.md). A fixture that passes on the day it is written measures nothing."
],
"cases": [
{ "id": "ru-query-001", "utterance": "сколько воды я выпил с утра", "lang": "ru", "intent": "query", "tags": ["aggregate"] },
{ "id": "ru-query-002", "utterance": "я сегодня вообще пил воду", "lang": "ru", "intent": "query", "tags": ["hard", "fact-shaped"], "note": "past-tense fact lexicon in a question — the classifier's fact centroid pulls this hard" },
{ "id": "ru-query-003", "utterance": "во сколько я лёг вчера", "lang": "ru", "intent": "query", "tags": ["temporal"] },
{ "id": "ru-query-004", "utterance": "давно я не тренировался", "lang": "ru", "intent": "query", "tags": ["hard", "no-question-word"] },
{ "id": "ru-query-005", "utterance": "напоминания на завтра есть", "lang": "ru", "intent": "query", "tags": ["hard", "reminder-shaped"], "note": "asks about reminders, does not create one" },
{ "id": "ru-query-006", "utterance": "что я записывал про кота", "lang": "ru", "intent": "query", "tags": ["recall"] },
{ "id": "ru-query-007", "utterance": "сколько раз я ел вчера", "lang": "ru", "intent": "query", "tags": ["aggregate", "hard"] },
{ "id": "ru-query-008", "utterance": "мой вес за последний месяц", "lang": "ru", "intent": "query", "tags": ["no-verb"] },
{ "id": "ru-query-009", "utterance": "когда я в последний раз принимал витамины", "lang": "ru", "intent": "query", "tags": ["temporal"] },
{ "id": "ru-query-010", "utterance": "есть новости по бэкапу базы", "lang": "ru", "intent": "query", "tags": ["homelab"] },
{ "id": "ru-query-011", "utterance": "почему сервер тормозит", "lang": "ru", "intent": "query", "tags": ["homelab", "hard"], "note": "diagnostic question, not a chat opener" },
{ "id": "ru-query-012", "utterance": "какие заметки я оставил про полив", "lang": "ru", "intent": "query", "tags": ["recall"] },
{ "id": "ru-query-013", "utterance": "во сколько у меня встреча", "lang": "ru", "intent": "query", "tags": ["calendar"] },
{ "id": "ru-query-019", "utterance": "что у меня стоит в календаре на послезавтра", "lang": "ru", "intent": "query", "tags": ["calendar", "hard"], "note": "agenda, not the clock: the daemon answers this from CalendarEvents inside the query branch, so the clock/date system rule must not swallow it" },
{ "id": "ru-query-022", "utterance": "какие планы на завтра?", "lang": "ru", "intent": "query", "tags": ["calendar"], "note": "the same agenda question as ru-query-019 aimed at another day; it answered \u043f\u043e\u043a\u0430 \u043d\u0435 \u0443\u043c\u0435\u044e on the deployed daemon while the today form worked (Vikunja #471)" },
{ "id": "ru-query-023", "utterance": "\u043a\u043e\u0433\u0434\u0430 \u043f\u043b\u0430\u043d\u0451\u0440\u043a\u0430?", "lang": "ru", "intent": "query", "tags": ["calendar", "hard"], "note": "a named event with no calendar word — the noun is the only signal that this is a question about his day" },
{ "id": "ru-query-024", "utterance": "что дальше?", "lang": "ru", "intent": "query", "tags": ["calendar", "no-question-word"], "note": "the rest of the day, with no possessive and no plan word to anchor on; the model called it a fact and the write had to be caught downstream (Vikunja #498)" },
{ "id": "ru-query-025", "utterance": "расскажи про битву при Ватерлоо", "lang": "ru", "intent": "query", "tags": ["world", "no-question-word"], "note": "a narrative request carries no question mark and no interrogative, so it routed fact; contrast ru-chat-003, where the same verb asks for a joke" },
{ "id": "ru-query-014", "utterance": "я успеваю до дедлайна", "lang": "ru", "intent": "query", "tags": ["hard", "no-question-word"] },
{ "id": "ru-query-015", "utterance": "сколько я прошёл шагов", "lang": "ru", "intent": "query", "tags": ["aggregate"] },
{ "id": "ru-query-016", "utterance": "покажи давление за неделю", "lang": "ru", "intent": "query", "tags": ["hard", "imperative"], "note": "imperative form but a read — must not route to act" },
{ "id": "ru-query-017", "utterance": "чем я занимался в среду", "lang": "ru", "intent": "query", "tags": ["hard", "chat-shaped"] },
{ "id": "ru-query-018", "utterance": "хватает ли места под новые бэкапы", "lang": "ru", "intent": "query", "tags": ["homelab"] },
{ "id": "ru-query-020", "utterance": "что дальше?", "lang": "ru", "intent": "query", "tags": ["agenda", "hard"], "note": "the rest of the day, with no interrogative the model can read as a question — it routed fact until a stage 0 rule claimed it (V-498)" },
{ "id": "ru-query-021", "utterance": "расскажи про битву при Ватерлоо", "lang": "ru", "intent": "query", "tags": ["world", "hard"], "note": "a world question phrased as an instruction. It routed fact, and the fact gate had to catch the write (V-498)" },
{ "id": "en-query-001", "utterance": "did I take my vitamins today", "lang": "en", "intent": "query", "tags": ["fact-shaped"] },
{ "id": "en-query-002", "utterance": "how long since the last backup finished", "lang": "en", "intent": "query", "tags": ["temporal"] },
{ "id": "en-query-003", "utterance": "show me this week's weight", "lang": "en", "intent": "query", "tags": ["imperative"] },
{ "id": "ru-query-001", "utterance": "сколько воды я выпил с утра", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["aggregate"] },
{ "id": "ru-query-002", "utterance": "я сегодня вообще пил воду", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["hard", "fact-shaped"], "note": "past-tense fact lexicon in a question — the classifier's fact centroid pulls this hard" },
{ "id": "ru-query-003", "utterance": "во сколько я лёг вчера", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["temporal"] },
{ "id": "ru-query-004", "utterance": "давно я не тренировался", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["hard", "no-question-word"] },
{ "id": "ru-query-005", "utterance": "напоминания на завтра есть", "lang": "ru", "intent": "query", "want_source": "", "tags": ["hard", "reminder-shaped"], "note": "asks about reminders, does not create one. want_source is the floor on purpose: no query source reads the reminder store, and day-plan is SourceCalendar over a table this box does not write." },
{ "id": "ru-query-006", "utterance": "что я записывал про кота", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["recall"] },
{ "id": "ru-query-007", "utterance": "сколько раз я ел вчера", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["aggregate", "hard"] },
{ "id": "ru-query-008", "utterance": "мой вес за последний месяц", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["no-verb"] },
{ "id": "ru-query-009", "utterance": "когда я в последний раз принимал витамины", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["temporal"] },
{ "id": "ru-query-010", "utterance": "есть новости по бэкапу базы", "lang": "ru", "intent": "query", "want_source": "", "tags": ["homelab"], "note": "recall, attention and network can each answer it, because mavpoll writes its netdata and uptime-kuma observations into the fact store recall reads. Naming one takes the other two off the turn." },
{ "id": "ru-query-011", "utterance": "почему сервер тормозит", "lang": "ru", "intent": "query", "want_source": "", "tags": ["homelab", "hard"], "note": "diagnostic question, not a chat opener. network holds the box and attention holds the alarm about the box. The utterance does not choose, so neither does the label." },
{ "id": "ru-query-012", "utterance": "какие заметки я оставил про полив", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["recall"] },
{ "id": "ru-query-013", "utterance": "во сколько у меня встреча", "lang": "ru", "intent": "query", "want_source": "calendar", "tags": ["calendar"] },
{ "id": "ru-query-019", "utterance": "что у меня стоит в календаре на послезавтра", "lang": "ru", "intent": "query", "want_source": "calendar", "tags": ["calendar", "hard"], "note": "agenda, not the clock: the daemon answers this from CalendarEvents inside the query branch, so the clock/date system rule must not swallow it" },
{ "id": "ru-query-022", "utterance": "какие планы на завтра?", "lang": "ru", "intent": "query", "want_source": "calendar", "tags": ["calendar"], "note": "the same agenda question as ru-query-019 aimed at another day; it answered \u043f\u043e\u043a\u0430 \u043d\u0435 \u0443\u043c\u0435\u044e on the deployed daemon while the today form worked (Vikunja #471)" },
{ "id": "ru-query-023", "utterance": "\u043a\u043e\u0433\u0434\u0430 \u043f\u043b\u0430\u043d\u0451\u0440\u043a\u0430?", "lang": "ru", "intent": "query", "want_source": "calendar", "tags": ["calendar", "hard"], "note": "a named event with no calendar word — the noun is the only signal that this is a question about his day" },
{ "id": "ru-query-024", "utterance": "что дальше?", "lang": "ru", "intent": "query", "want_source": "calendar", "tags": ["calendar", "no-question-word"], "note": "the rest of the day, with no possessive and no plan word to anchor on; the model called it a fact and the write had to be caught downstream (Vikunja #498)" },
{ "id": "ru-query-025", "utterance": "расскажи про битву при Ватерлоо", "lang": "ru", "intent": "query", "want_source": "world", "tags": ["world", "no-question-word"], "note": "a narrative request carries no question mark and no interrogative, so it routed fact; contrast ru-chat-003, where the same verb asks for a joke" },
{ "id": "ru-query-014", "utterance": "я успеваю до дедлайна", "lang": "ru", "intent": "query", "want_source": "", "tags": ["hard", "no-question-word"], "note": "a deadline lives in the task list, the calendar or Praxis depending on where he put it. The destination depends on his data, not on his words." },
{ "id": "ru-query-015", "utterance": "сколько я прошёл шагов", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["aggregate"] },
{ "id": "ru-query-016", "utterance": "покажи давление за неделю", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["hard", "imperative"], "note": "imperative form but a read — must not route to act" },
{ "id": "ru-query-017", "utterance": "чем я занимался в среду", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["hard", "chat-shaped"] },
{ "id": "ru-query-018", "utterance": "хватает ли места под новые бэкапы", "lang": "ru", "intent": "query", "want_source": "", "tags": ["homelab"], "note": "disk headroom. network is the only source that reads the box, but the phrasing is a capacity question and not a LAN one." },
{ "id": "ru-query-020", "utterance": "что дальше?", "lang": "ru", "intent": "query", "want_source": "calendar", "tags": ["agenda", "hard"], "note": "the rest of the day, with no interrogative the model can read as a question — it routed fact until a stage 0 rule claimed it (V-498)" },
{ "id": "ru-query-021", "utterance": "расскажи про битву при Ватерлоо", "lang": "ru", "intent": "query", "want_source": "world", "tags": ["world", "hard"], "note": "a world question phrased as an instruction. It routed fact, and the fact gate had to catch the write (V-498)" },
{ "id": "en-query-001", "utterance": "did I take my vitamins today", "lang": "en", "intent": "query", "want_source": "recall", "tags": ["fact-shaped"] },
{ "id": "en-query-002", "utterance": "how long since the last backup finished", "lang": "en", "intent": "query", "want_source": "", "tags": ["temporal"], "note": "the completion time is a fact the poller wrote, so recall answers it. A person asking this wants the operational answer. Both are true." },
{ "id": "en-query-003", "utterance": "show me this week's weight", "lang": "en", "intent": "query", "want_source": "recall", "tags": ["imperative"] },
{ "id": "ru-fact-001", "utterance": "только что выпил кружку воды", "lang": "ru", "intent": "fact", "want_fact_key": "water" },
{ "id": "ru-fact-002", "utterance": "воды попил наконец", "lang": "ru", "intent": "fact", "want_fact_key": "water", "tags": ["inverted"] },
@@ -106,6 +110,11 @@
{ "id": "amb-005", "utterance": "потом", "lang": "ru", "want_clarify": true, "tags": ["ambiguous", "filler"] },
{ "id": "amb-006", "utterance": "the thing from earlier", "lang": "en", "want_clarify": true, "tags": ["ambiguous", "anaphora"] },
{ "id": "amb-007", "utterance": "напомни", "lang": "ru", "want_clarify": true, "tags": ["ambiguous", "reminder"], "note": "the reminder verb and nothing else — she knows the shape of the request and not one thing about it. Answered 'не получилось разобрать время напоминания' on the box until V-548: the subjectless-reminder gate tested Slots.Text == \"\", and fillSlots had put the verb in that slot" },
{ "id": "amb-008", "utterance": "ну напомни же", "lang": "ru", "want_clarify": true, "tags": ["ambiguous", "reminder", "filler"], "note": "the same request wrapped in particles, which is why filler_particles is a lexicon set — without it the particles read as the subject" }
{ "id": "amb-008", "utterance": "ну напомни же", "lang": "ru", "want_clarify": true, "tags": ["ambiguous", "reminder", "filler"], "note": "the same request wrapped in particles, which is why filler_particles is a lexicon set — without it the particles read as the subject" },
{ "id": "ru-query-026", "utterance": "что такое TCP?", "lang": "ru", "intent": "query", "want_source": "world", "tags": ["world", "regression"], "note": "weather claimed it on 2026-08-07 and answered \"для какого города?\", because it read one percent closer than the leftover seeds. WorldQueryGrammars claims it at stage 0 now." },
{ "id": "ru-query-027", "utterance": "сколько будет 17 на 23?", "lang": "ru", "intent": "query", "want_source": "world", "tags": ["world", "arithmetic", "regression"], "note": "same source, same day, same answer about a city. Arithmetic is not a place." },
{ "id": "ru-query-028", "utterance": "какой у меня любимый язык?", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["recall", "possessive", "regression"], "note": "the feed answered it with kernel headlines. \"у меня\" is the whole signal and it points inward." },
{ "id": "ru-query-029", "utterance": "кто такой Линус Торвальдс?", "lang": "ru", "intent": "query", "want_source": "world", "tags": ["world", "person", "regression"], "note": "the personal boundary answered \"не нашла у тебя такой записи\". A named public person is not his data." },
{ "id": "ru-query-030", "utterance": "что там с бэкапами?", "lang": "ru", "intent": "query", "want_source": "", "tags": ["homelab", "regression"], "note": "search claimed it, which inverts the boundary outward. The fix is the chain order and not a destination: recall, attention and network all answer it, same as ru-query-010." }
]
}
+6
View File
@@ -105,6 +105,12 @@ type Decision struct {
Slots Slots
Clarify bool // stage 3: below threshold — ask, don't guess
// Source — where the answer lives, for a query. The second half of the
// route, and empty on every other intent. SourceUnknown means no decider
// named one and the daemon walks its whole chain, which is what shipped
// before this field existed. See source.go for why it is twelve values.
Source Source
// Continued — this decision was rebuilt from the previous turn rather
// than routed, because the utterance was an ellipsis ("а завтра?").
// Handlers use it to know that Slots.Text is the PREVIOUS turn's topic
+78
View File
@@ -0,0 +1,78 @@
package router
// Source — where the answer to a query lives. It is the second half of a
// routing decision and it used to be made outside the router entirely (V-655).
//
// The cascade sorted an utterance into one of seven intents with stage 0 rules,
// the resident model and the classifier behind it, a fixture measuring it and
// the decision trace recording it. Then IntentQuery handed the turn to
// querySources in the daemon, a chain of twenty-two branches deciding by seed
// similarity in a fixed order, with none of that. So the careful sorter did the
// easy half and the sloppy one did the hard half: on 2026-08-07 weather claimed
// "что такое TCP?" and answered "для какого города?", because weather read one
// percent closer to the turn than the pile of leftover seeds did, and one
// percent was enough. Search would have answered it and search was never asked.
//
// "query" is not a destination. It is a shrug. This is the field that says
// where to look.
//
// # Why twelve and not twenty-two
//
// A destination is what a decider can plausibly name from the utterance alone,
// not one entry per source. Three of the daemon's sources are successive passes
// over his own words and a fourth reads the facts by key: which of them lands
// the hit is an ordering detail inside the chain, and no utterance says. They
// are SourceRecall together. The same goes for the metasearch, the offline
// encyclopedia and a page he named by URL, which are SourceWorld.
//
// # Empty is a real value and it is the floor
//
// SourceUnknown means nobody decided. The daemon then walks the whole chain in
// its original order, which is the behaviour that shipped before this field
// existed. So the classifier arm sets nothing and costs nothing, and a box
// whose model is down routes queries exactly as it did.
type Source string
const (
// SourceUnknown — no decider named a destination. Walk the chain.
SourceUnknown Source = ""
// His own data.
SourceRecall Source = "recall" // notes, facts and what he has said before
SourceCalendar Source = "calendar" // events, and the only date-aware destination
SourceTasks Source = "tasks" // the task list
SourceList Source = "list" // the shopping and other named lists
SourceMoney Source = "money" // the spending facts the poller writes
// The surroundings.
SourceWeather Source = "weather" // the forecast for a place
SourceHome Source = "home" // lights, devices, the house
SourceNetwork Source = "network" // the LAN and what is on it
SourceFeeds Source = "feeds" // the RSS she reads
SourceAttention Source = "attention" // what Praxis says needs looking at
// Everything else.
SourceSelf Source = "self" // a question about Maven herself
SourceWorld Source = "world" // search, the ZIMs, a page he named
)
// Sources — every destination a decider may name, in a fixed order so a prompt,
// a grammar table and a test all read the same list. SourceUnknown is not a
// member: it is the absence of a choice, not one of the choices.
var Sources = []Source{
SourceRecall, SourceCalendar, SourceTasks, SourceList, SourceMoney,
SourceWeather, SourceHome, SourceNetwork, SourceFeeds, SourceAttention,
SourceSelf, SourceWorld,
}
// ValidSource reports whether s is one a decider may name. Anything else,
// including a destination invented by a model, is dropped back to
// SourceUnknown by the caller rather than trusted.
func ValidSource(s Source) bool {
for _, known := range Sources {
if s == known {
return true
}
}
return false
}
+20 -5
View File
@@ -187,9 +187,11 @@ func SystemTimeDateGrammars() []Grammar {
// written ("the clock/date system rule must not swallow it"); the daemon
// disagreed with the fixture and the daemon was wrong.
//
// Routing, not answering. These set the intent and nothing else — which source
// in the query chain claims the turn stays the chain's decision, and a
// question with no date still falls through queryCalendar to recall.
// Routing, not answering. Two of the five also name the calendar as the
// destination (V-655), which narrows who may GUESS their way onto the turn and
// claims nothing. Every source that looks something up still runs, in the order
// it always did, so a question with no date still falls through queryCalendar
// to recall.
//
// Deliberately not folded into SystemTimeDateGrammars: those exist to send
// utterances TO system, these exist to keep utterances OUT of it, and one
@@ -199,9 +201,15 @@ func AgendaQueryGrammars() []Grammar {
{
// An explicit calendar noun is unambiguous wherever it appears:
// "что в календаре на завтра", "покажи расписание на среду".
//
// The one agenda rule that names its destination, because an
// explicit calendar noun leaves nothing to weigh (V-655). The
// possessive rules below deliberately do not: "что у меня в списке
// покупок" matches agenda-query, and naming the calendar there
// would take the list source off the turn.
Name: "calendar-query",
Pattern: regexp.MustCompile(`(?i)(календар|расписани|повестк)`),
Build: agendaQueryBuild,
Build: queryTo(SourceCalendar),
},
{
// The agenda phrasing with no calendar noun. Anchored at the start
@@ -251,9 +259,11 @@ func AgendaQueryGrammars() []Grammar {
// "во сколько созвон". He is asking when something on his calendar
// happens, and the noun is the only signal. Closed list, so "когда
// битва при Ватерлоо" is still a world question.
// Names the calendar (V-655): the noun list is closed and every
// member of it is an event, so there is nothing else to weigh.
Name: "event-time-query",
Pattern: regexp.MustCompile(`(?i)^\s*(когда|во\s+сколько|в\s+котором\s+часу)\s+(будет\s+|у\s+нас\s+)?(планёрк|планерк|встреч|созвон|митинг|совещани|звонок|созвон|приём|прием|интервью|собеседовани|тренировк|урок|занятие|пара)[а-я]*(\s|[?!.]|$)`),
Build: agendaQueryBuild,
Build: queryTo(SourceCalendar),
},
}
}
@@ -340,6 +350,11 @@ func narrativeQueryBuild(m []string) (Decision, bool) {
Intent: IntentQuery,
Confidence: 1.0,
Slots: Slots{Text: topic},
// The world, because that is the shape this asks for and the rule has
// already declined the two cases where it is not: entertainment, and
// questions about her (V-655). His own notes are still read first — a
// destination narrows who may guess and reorders nothing.
Source: SourceWorld,
}, true
}
+86
View File
@@ -0,0 +1,86 @@
package router
import "regexp"
// WorldQueryGrammars — stage-0 rules for the two question shapes that name the
// world in their own words, and say so plainly enough that no scorer is needed
// (V-655).
//
// They exist because of what happens when nothing deterministic claims these.
// Measured on the box on 2026-08-07 (docs/evals/2026-08-07-week-of-usage.md,
// section 4): "что такое TCP?" and "сколько будет 17 на 23?" were both answered
// "для какого города?", and "кто такой Линус Торвальдс?" was answered "не знаю —
// не нашла у тебя такой записи". None of those three is about him, about the
// weather, or about anything on this box.
//
// The mechanism is the destination, not the answer. Naming SourceWorld does not
// send the turn outside and does not skip a single source that looks something
// up: his notes, his facts and the personal boundary all still run first, in the
// order they always did. What it does is stop the sources that claim on seed
// similarity from taking the turn on the way past. Weather cannot claim a
// question about a protocol once the utterance has said which side it is on.
//
// Both patterns are spelled out here rather than drawn from internal/lexicon,
// which is the same call the agenda rules made: these are interrogative FRAMES
// of two words, not a closed class of single words, and the lexicon holds
// classes. Nothing here is a stem pattern over open vocabulary — the variable
// part of each rule is the topic, and the rule reads none of it.
func WorldQueryGrammars() []Grammar {
return []Grammar{
{
// "что такое X", "кто такой X". A request for what a thing or a
// person IS, which his own data can answer and usually cannot.
//
// The topic is deliberately not captured into Slots.Text. Every
// source below reads the utterance, "что такое TCP?" is already the
// best query string for it, and the agenda rules make the same call
// for the same reason.
Name: "definition-query",
Pattern: definitionQueryPattern,
Build: queryTo(SourceWorld),
},
{
// "сколько будет 17 на 23", "сколько будет 2+2". Arithmetic, which
// the metasearch answers and no local source holds. The digits are
// what make it arithmetic: "сколько будет гостей" names no number
// and is a question about his evening.
Name: "arithmetic-query",
Pattern: arithmeticQueryPattern,
Build: queryTo(SourceWorld),
},
}
}
// definitionQueryPattern — anchored at the start, because "напомни узнать что
// такое TCP" is a reminder that happens to contain the frame.
//
// (\s|[?!.]|$) and not \b: Go's \b is ASCII-only and never fires after a
// Cyrillic letter, so the ASCII form silently matches nothing. The agenda rules
// carry the same note.
var definitionQueryPattern = regexp.MustCompile(
`(?i)^\s*(что\s+так(ое|ая)|кто\s+так(ой|ая|ие)|what\s+is|who\s+is)(\s|[?!.]|$)`)
// arithmeticQueryPattern — the ask, then a digit somewhere after it. Loose on
// what sits between them on purpose: the operator is spoken half a dozen ways
// ("на", "умножить на", "плюс", "+") and reading them is the calculator's job,
// not this rule's. All this decides is which side of the boundary the turn is
// on.
var arithmeticQueryPattern = regexp.MustCompile(
`(?i)^\s*(сколько\s+будет|посчитай|вычисли|how\s+much\s+is)\s.*\d`)
// queryTo builds a stage-0 query Decision that names where the answer lives.
//
// The utterance travels intact and no slot is filled, which is the same
// contract agendaQueryBuild has: confidence 1.0 on the intent and the
// destination, and every source below still decides for itself whether it has
// an answer. Naming a destination narrows who may guess. It promises nothing.
func queryTo(dest Source) func([]string) (Decision, bool) {
return func([]string) (Decision, bool) {
return Decision{
Stage: 0,
Intent: IntentQuery,
Confidence: 1.0,
Source: dest,
}, true
}
}
+88
View File
@@ -0,0 +1,88 @@
package router
import "testing"
// The three utterances from the 2026-08-07 week on the box that no local source
// could answer and three different local sources claimed anyway. Stage 0 has to
// say which side of the boundary they are on, because by the time the chain is
// walking, the only thing separating them from a weather forecast is a cosine.
func TestAWorldQuestionNamesTheWorld(t *testing.T) {
cases := []struct {
utterance string
rule string
}{
{"что такое TCP?", "definition-query"},
{"кто такой Линус Торвальдс?", "definition-query"},
{"что такая мембрана", "definition-query"},
{"кто такая Ада Лавлейс?", "definition-query"},
{"what is TCP?", "definition-query"},
{"сколько будет 17 на 23?", "arithmetic-query"},
{"посчитай 2+2", "arithmetic-query"},
{"сколько будет 5 умножить на 6", "arithmetic-query"},
}
for _, c := range cases {
dec, rule, ok := matchWorldQuery(c.utterance)
if !ok {
t.Errorf("%q: no world rule claimed it", c.utterance)
continue
}
if rule != c.rule {
t.Errorf("%q: claimed by %q, want %q", c.utterance, rule, c.rule)
}
if dec.Intent != IntentQuery {
t.Errorf("%q: intent %q, want query", c.utterance, dec.Intent)
}
if dec.Source != SourceWorld {
t.Errorf("%q: source %q, want %q", c.utterance, dec.Source, SourceWorld)
}
}
}
// The frame has to be the whole opening or the rule is reading somebody else's
// sentence. Every case here contains a world-question shape and is not one.
func TestAWorldRuleDeclinesWhatIsNotItsShape(t *testing.T) {
cases := []struct {
utterance string
why string
}{
{"напомни узнать что такое TCP", "a reminder that happens to quote the frame"},
{"запиши что такое TCP", "a capture that happens to quote the frame"},
{"сколько будет гостей", "an ask with no number is not arithmetic"},
{"что у меня сегодня?", "his agenda, and the agenda rules own it"},
{"кто там?", "not the frame"},
{"посчитай расходы", "no number, so the money source keeps it"},
}
for _, c := range cases {
if _, rule, ok := matchWorldQuery(c.utterance); ok {
t.Errorf("%q: claimed by %q, want no claim — %s", c.utterance, rule, c.why)
}
}
}
// The destination is advice about who may guess, never a filled slot. A rule
// that quietly captured the topic would change what every source below reads.
func TestNamingTheWorldFillsNoSlot(t *testing.T) {
dec, _, ok := matchWorldQuery("что такое TCP?")
if !ok {
t.Fatal("definition-query did not claim it")
}
if dec.Slots.Text != "" || dec.Slots.HasTime || dec.Slots.HasFn || dec.Slots.HasKey {
t.Errorf("slots = %+v, want none filled", dec.Slots)
}
if dec.Confidence != 1.0 {
t.Errorf("confidence = %v, want 1.0 for a stage-0 match", dec.Confidence)
}
}
func matchWorldQuery(utterance string) (Decision, string, bool) {
for _, g := range WorldQueryGrammars() {
m := g.Pattern.FindStringSubmatch(utterance)
if m == nil {
continue
}
if dec, ok := g.Build(m); ok {
return dec, g.Name, true
}
}
return Decision{}, "", false
}