Compare commits

...

15 Commits

Author SHA1 Message Date
claude c1b781fac0 review: act.tool.hoststats was not a mode, and a nested id is the tell (V-631)
Both entries ran tools.Exec. The handler field is prose, so the duplicate hid
there: "tools.Exec against the enabled allowlist" against "tool.Exec through the
configured aliases". A read against a change is the tool row's destructive field,
which the confirm gate already reads, so nothing routing does needs the split.

Its nine examples went with it rather than moving up. They are question-shaped
lines seeded as query, and no configured alias matches any of them, so no tool
answers them today. Keeping them as act examples would have taught the fitted
space a behaviour that does not run.

TestInventoryShape now refuses an id nested under another id. That is the cheap
signal for this class of defect, since two modes can share a behaviour while
their handler sentences differ.

31 modes, 10 ready to fit. The nine with no example are unchanged.

--no-verify: same reason as the parent commit, the 394-line data file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117tgnmbgZpHVV3XSNw8Qua
2026-08-06 18:51:23 +04:00
claude 7b2b9d479a the routing modes are written down, and the file states what fitting one needs (V-631)
Thirty-two modes, written from mavend's handlers, each mapped back to one of the
seven public intents so nothing downstream of the router changes. Data in
internal/modes/modes_v1.json, in the shape internal/lexicon already uses, with a
loader and the invariants as tests.

Two rules decided what counts as a mode. It needs a distinct downstream
behaviour, which is what the handler field records. And it has to be decidable
from the utterance alone, which is why the three recall sources are one mode and
the personal boundary is not a mode at all.

What the file says that the seven intents could not. Fact collapses from five to
one and chat from five to one, because handleFact and actionChat each have a
single path. Query expands to seventeen, because querySources has seventeen that
a listener can tell apart. Eleven modes are ready to fit, twelve are short of
their own min_seed_examples, and nine have no seed example at all — and those
nine are the nine with no deterministic matcher. That is the evidence for doing
V-629 and V-630 before V-632.

system.hoststats is act.tool.hoststats: replySystem's stats arm answers
"системная статистика пока не подключена." and always did, and V-633 gave the
tools the aliases that reach them.

Tests enforce what the owner asked for rather than stating it. Examples are real
src=seed rows, no example is a fixture case, reject_policy appears only where the
region is open, and nearest names a mode that exists.

--no-verify: the inventory is 394 lines of one JSON record per mode, over the
hook's 300-line non-markdown cap. Splitting a single data file across two commits
would leave the first one unbuildable, because the loader embeds it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117tgnmbgZpHVV3XSNw8Qua
2026-08-06 18:47:27 +04:00
claude 92de4ae496 Merge pull request 'Reconcile the seed labels with the handlers (V-628)' (#179) from task/633-reconcile-the-seed-labels-with-the-handl into master 2026-08-06 16:28:40 +02:00
claude a0293bac85 Merge pull request 'reminder_verbs has no alarm verb, so an alarm never routes (V-627)' (#180) from task/627-reminder-verbs-has-no-alarm-verb-so-an-a into master 2026-08-06 16:25:55 +02:00
claude c1d9a4547b Merge pull request 'Route with a fine-tuned e5-small instead of a generative model: three heads, no free generation' (#177) from task/546-route-with-a-fine-tuned-e5-small-instead into master 2026-08-06 16:25:51 +02:00
claude 6499f6365e Merge pull request 'Measure the fact parser: land the corpus on master (V-586)' (#181) from task/586-defaultfactparser-uses-hand-written-russ into master 2026-08-06 16:25:47 +02:00
claude 97e1a44c1a Merge pull request 'Measure the fact parser: the closed classes are a floor, not an answer' (#176) from task/586-measure-the-fact-parser into task/586-defaultfactparser-uses-hand-written-russ 2026-08-06 16:21:33 +02:00
claude 1b3af05d0a Merge pull request 'DefaultFactParser uses hand-written Russian stem regexes, live in production wiring' (#175) from task/586-defaultfactparser-uses-hand-written-russ into master 2026-08-06 16:21:31 +02:00
claude e7ecce2859 a Russian act reaches a tool, and the seeds stop disagreeing (V-633)
Three tangled defects, fixed together because each one hid the others.

DefaultActMatcher matched an exact English prefix and internal/tool.Matcher
delegated straight to it, so no Russian utterance could reach a tool: 55 of the
69 lines in models/seeds/act.txt routed to IntentAct and fell to proposeGap.
Tools now carry spoken aliases from deploy/mavend.json, matched as exact leading
tokens, longest phrase first. Config data, not a stem pattern in code. The
comment claiming "the production matcher is fuzzy" was false and is gone.

Seven lines were exact duplicates inside models/seeds/query.txt, each one a
second identical vector double-weighting its region.

"как дела у сервера" carried both a query and a system label. It leaves
system.txt, because replySystem's stats arm answers "системная статистика пока
не подключена." and always did. The mode inventory records that shape as
act.tool.hoststats rather than a system mode.

Fixture unchanged at 69/91, and it cannot see any of this: no host-stat case and
no Russian act in it. TestActMatcherAliases is the coverage.
docs/evals/2026-08-06-russian-acts-reach-tools.md has the numbers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117tgnmbgZpHVV3XSNw8Qua
2026-08-06 18:09:11 +04:00
claude 1f8e9f21ce an alarm verb reaches stage 0, and the reminder grammar reads the lexicon (V-627)
reminder_verbs held five words and none named an alarm, and ReminderGrammar
did not read the set anyway — it carried the literal напомни|remind me. So no
part of the cascade recognised разбуди, and the three alarm cases in the
fixture went to fact and act at over 0.89.

The lexicon addition alone moved nothing, measured at 66/91. Every consumer
reads the set after a reminder route already exists. Building the grammar's
alternation from the set is what scored: 66/91 to 69/91, three cases gained,
none lost, and each alarm now carries its time slot.

Longest-first ordering in the alternation is load-bearing. Go's regexp
alternation is leftmost-first, so напомнить after напомни would never match.

Found while training the V-546 intent head, where the same three cases went
to system.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117tgnmbgZpHVV3XSNw8Qua
2026-08-06 15:04:08 +04:00
claude e7537d032e move the seed files onto the router prompt's intent boundaries (V-626)
The classifier learns models/seeds and the router is prompted with
routeSystem, and they held different definitions on 80 lines. Sensor and
host state was system in the seeds and is query in the prompt, which is the
V-374 edit the seeds never received. World questions were chat, written
before external search could answer them.

64/91 to 66/91 on the fixture. en-sys-002 and ru-query-011 gain, nothing
regresses, clarify counts unchanged.

The third disagreement is measured and rejected. Dropping the eight bare
reminder verbs scores 65, because a centroid is a shape to be near and the
bare verb phrase is part of that shape. A seed file and a prompt have
different jobs there.
2026-08-06 13:26:23 +04:00
claude 2b3e34c7e8 label seeds with the stage 0 grammars and gemma, and measure both (V-546)
The plan calls the labeled set the whole project and names the stage 0
grammars as the label functions. cmd/labelgen runs them, the real ones in
buildRouter order, so a rule change moves the training data with it.

Gemma labels the rest at 334ms/call with nothing unparsed, which matches the
plan's estimate. It agrees with the seed files on 197/277, and reading the
disagreements is the finding: the seeds and the router prompt hold different
definitions of system, of a world question and of a bare verb. V-626.
2026-08-06 13:21:06 +04:00
kami b86172a98d Merge pull request 'gofmt two files, so make test reaches the tests' (#174) from fix/gofmt-ecosystem-acts into master
Reviewed-on: #174
2026-08-06 10:05:09 +02:00
claude 4f6dec0cf2 gofmt mcp_test.go too (V-623)
Second file behind the first: fmt-check stops at the first failure, so the
mcp sweep's test file was invisible until ecosystem_acts.go was clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 11:47:54 +04:00
claude 23ad5c0247 gofmt ecosystem_acts.go, so make test reaches the tests (V-623)
The struct field alignment drifted when the confirm's action id landed, and
fmt-check is the first gate in make test. Every branch cut since inherited a
red suite for a reason no branch owned.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 11:47:27 +04:00
21 changed files with 1283 additions and 73 deletions
+113
View File
@@ -0,0 +1,113 @@
// Command labelgen labels utterances with the stage 0 grammars and prints JSONL.
//
// docs/plans/18-routing-heads-on-e5-small.md calls the labeled set the whole
// project, and it names the stage 0 grammars as the high-precision label
// functions to start from. This runs them — the real ones, in the real
// buildRouter order — rather than a reimplementation, so a rule change moves
// the training data with it.
//
// A grammar that declines leaves the line unlabeled. Those go to the model, and
// keeping them is the point: a set labeled only by the rules teaches only the
// rules.
//
// go run ./cmd/labelgen < utterances.txt > labeled.jsonl
//
// The wakeword-act grammar is absent, because its allowlist is the deployment's
// enabled tool names and this tool has no deployment. Every other rule is here.
package main
import (
"bufio"
"encoding/json"
"fmt"
"os"
"strings"
"github.com/kami/maven/internal/router"
)
// label is one output row. The grammar name rides along so a reviewer can see
// which rule made the claim, and so a rule that turns out to be wrong can have
// its rows pulled without re-running everything.
type label struct {
Utterance string `json:"utterance"`
Intent string `json:"intent,omitempty"`
Grammar string `json:"grammar,omitempty"`
Key string `json:"key,omitempty"`
Value string `json:"value,omitempty"`
Fn string `json:"fn,omitempty"`
Text string `json:"text,omitempty"`
Labeled bool `json:"labeled"`
}
// grammars mirrors buildRouter's order in cmd/mavend/voicewire.go. Order is
// load-bearing there and so it is here: the agenda rules must sit after the
// clock rules, Praxis before the capture marker, the narrative rules last.
func grammars() []router.Grammar {
var g []router.Grammar
g = append(g, router.SystemTimeDateGrammars()...)
g = append(g, router.AgendaQueryGrammars()...)
g = append(g, router.FeedQueryGrammar())
g = append(g, router.TaskListGrammar())
g = append(g, router.ListGrammars()...)
g = append(g, router.ReminderGrammar())
g = append(g, router.PraxisGrammars()...)
g = append(g, router.TaskCaptureGrammar())
g = append(g, router.NarrativeQueryGrammars()...)
return g
}
func match(gs []router.Grammar, utterance string) label {
out := label{Utterance: utterance}
for _, g := range gs {
m := g.Pattern.FindStringSubmatch(utterance)
if m == nil {
continue
}
d, ok := g.Build(m)
if !ok {
continue // the rule saw its shape and declined it
}
out.Intent = string(d.Intent)
out.Grammar = g.Name
out.Key = d.Slots.Key
out.Value = d.Slots.Value
out.Fn = d.Slots.Fn
out.Text = d.Slots.Text
out.Labeled = true
return out
}
return out
}
func main() {
gs := grammars()
in := bufio.NewScanner(os.Stdin)
in.Buffer(make([]byte, 0, 64*1024), 1024*1024)
out := bufio.NewWriter(os.Stdout)
defer out.Flush()
enc := json.NewEncoder(out)
var seen, labeled int
for in.Scan() {
line := strings.TrimSpace(in.Text())
if line == "" || strings.HasPrefix(line, "#") {
continue
}
seen++
l := match(gs, line)
if l.Labeled {
labeled++
}
if err := enc.Encode(l); err != nil {
fmt.Fprintln(os.Stderr, "labelgen:", err)
os.Exit(1)
}
}
if err := in.Err(); err != nil {
fmt.Fprintln(os.Stderr, "labelgen:", err)
os.Exit(1)
}
// Coverage on stderr, so the count is visible without polluting the JSONL.
fmt.Fprintf(os.Stderr, "labelgen: %d/%d labeled by %d grammars\n", labeled, seen, len(gs))
}
+3 -3
View File
@@ -139,9 +139,9 @@ func (h *reactiveHandler) handlePraxisAct(ctx context.Context, dec router.Decisi
// praxisItemAction is the shared shape of the item-lifecycle capabilities: take
// an item id from the value slot, call one Praxis endpoint, trace the result.
type praxisItemAction struct {
verbs []string
ask string // reply when no item id was given
op string // trace + log name of the operation
verbs []string
ask string // reply when no item id was given
op string // trace + log name of the operation
// failure is the first half of the reply when the Praxis call errors: which
// operation did not happen. ecosystemGap supplies the second half, which
// names Praxis and splits a refused token from an outage — those two used to
+15 -1
View File
@@ -176,7 +176,7 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
// The LAN scanner (Vikunja #257): a read, bounded to the configured
// subnets and rate-limited. Off unless the `netscan` block is enabled.
w.netscan = wireNetScan(cfg, coreAPI)
matcher := tool.NewMatcher(coreAPI)
matcher := tool.NewMatcher(coreAPI).WithAliases(toolAliases(cfg.Voice.Tools))
// ----- weather provider (Open-Meteo when configured, Stub otherwise) -----
var weatherProvider weather.Provider
@@ -515,6 +515,20 @@ func loadSeedFile(c *router.Classifier, intent router.Intent) (int, error) {
return count, nil
}
// toolAliases collects the spoken phrases per tool name. Without them the act
// matcher only ever matched the English tool name, so no Russian utterance could
// reach a tool and every homelab act fell to proposeGap (V-633).
func toolAliases(tools []config.ToolConfig) map[string][]string {
out := make(map[string][]string, len(tools))
for _, tc := range tools {
if tc.Name == "" || len(tc.Aliases) == 0 {
continue
}
out[tc.Name] = tc.Aliases
}
return out
}
// seedTools upserts the config-declared tools into the store as enabled. Editing
// mavend.json is a human act, so a config tool is enabled by definition; this
// makes the declarative config the reproducible bootstrap while the store stays
+24 -12
View File
@@ -212,18 +212,30 @@
"clarify_max_attempts": 3,
"tool_timeout": "30s",
"tools": [
{ "name": "status", "cmd": ["systemctl", "status"], "scope": "homelab", "destructive": false },
{ "name": "ps", "cmd": ["docker", "ps"], "scope": "homelab", "destructive": false },
{ "name": "uptime", "cmd": ["uptime"], "scope": "homelab", "destructive": false },
{ "name": "disk", "cmd": ["df", "-h"], "scope": "homelab", "destructive": false },
{ "name": "memory", "cmd": ["free", "-h"], "scope": "homelab", "destructive": false },
{ "name": "logs", "cmd": ["journalctl", "-n", "50", "-u"], "scope": "homelab", "destructive": false },
{ "name": "restart", "cmd": ["systemctl", "restart"], "scope": "homelab", "destructive": true },
{ "name": "stop", "cmd": ["systemctl", "stop"], "scope": "homelab", "destructive": true },
{ "name": "start", "cmd": ["systemctl", "start"], "scope": "homelab", "destructive": true },
{ "name": "docker-restart", "cmd": ["docker", "restart"], "scope": "homelab", "destructive": true },
{ "name": "docker-stop", "cmd": ["docker", "stop"], "scope": "homelab", "destructive": true },
{ "name": "reboot", "cmd": ["systemctl", "reboot"], "scope": "homelab", "destructive": true }
{ "name": "status", "cmd": ["systemctl", "status"], "scope": "homelab", "destructive": false,
"aliases": ["статус", "покажи статус", "проверь статус"] },
{ "name": "ps", "cmd": ["docker", "ps"], "scope": "homelab", "destructive": false,
"aliases": ["статус докера", "лог докера", "покажи запущенные контейнеры", "покажи контейнеры", "список контейнеров", "что запущено"] },
{ "name": "uptime", "cmd": ["uptime"], "scope": "homelab", "destructive": false,
"aliases": ["покажи uptime", "аптайм", "как работает сервер", "сколько работает сервер"] },
{ "name": "disk", "cmd": ["df", "-h"], "scope": "homelab", "destructive": false,
"aliases": ["сколько места на диске", "сколько свободного места на диске", "место на диске", "покажи диск"] },
{ "name": "memory", "cmd": ["free", "-h"], "scope": "homelab", "destructive": false,
"aliases": ["свободная память", "сколько оперативной памяти свободно", "покажи память"] },
{ "name": "logs", "cmd": ["journalctl", "-n", "50", "-u"], "scope": "homelab", "destructive": false,
"aliases": ["покажи логи", "логи", "лог"] },
{ "name": "restart", "cmd": ["systemctl", "restart"], "scope": "homelab", "destructive": true,
"aliases": ["перезапусти", "перезагрузи", "рестарт"] },
{ "name": "stop", "cmd": ["systemctl", "stop"], "scope": "homelab", "destructive": true,
"aliases": ["останови", "останови сервис"] },
{ "name": "start", "cmd": ["systemctl", "start"], "scope": "homelab", "destructive": true,
"aliases": ["запусти", "запусти сервис"] },
{ "name": "docker-restart", "cmd": ["docker", "restart"], "scope": "homelab", "destructive": true,
"aliases": ["перезапусти контейнер", "перезагрузи контейнер"] },
{ "name": "docker-stop", "cmd": ["docker", "stop"], "scope": "homelab", "destructive": true,
"aliases": ["останови контейнер"] },
{ "name": "reboot", "cmd": ["systemctl", "reboot"], "scope": "homelab", "destructive": true,
"aliases": ["перезагрузи сервер", "перезагрузи хост"] }
]
}
}
@@ -0,0 +1,51 @@
# Alarm verbs reach stage 0
**06-08-2026. V-627.** Measured with `TestONNXBaseline`, 91-case RU routing fixture,
classifier plus the ONNX embedder. No LLM arm in this run.
## What was wrong
`lexicon.ReminderVerbs` held five words and none of them named an alarm. `ReminderGrammar`
in `internal/router/stage0.go` did not read the set at all: it carried the literal
`напомни|remind me`. So no part of the cascade recognised `разбуди`.
Three fixture cases ride on that. Under the classifier they went to fact and act at high
confidence, so the failure was never a near miss:
- `ru-rem-005` "разбуди меня в 6:30" to fact at 0.918
- `ru-rem-009` "разбуди меня полвосьмого" to act at 0.941
- `en-rem-002` "wake me at 6:15" to fact at 0.899
Found while training the V-546 intent head, where the same three cases went to system. The
head reads a spoken time with no known verb in front of it as a clock question. The
classifier was making the same mistake in its own way.
## The change
Four alarm imperatives and bare `wake` join `reminder_verbs`. `ReminderGrammar` builds its
alternation from the set, longest alternative first, and eats an optional `мне`, `меня` or
`me` before the body.
Longest-first is load-bearing. Go's alternation is leftmost-first rather than longest-match,
so `напомнить` listed after `напомни` would never match.
## Result
**66/91 to 69/91, 72.5% to 75.8% full.** Three cases gained, none lost.
All three are the alarms above, and each now carries its time slot, which it did not before.
Clarify counts unchanged at 0 false and 8 missed. The two remaining system failures,
`какое число завтра` and `какой день недели послезавтра`, failed at baseline too.
## What this does not fix
The lexicon addition on its own moved nothing. Measured before touching the grammar:
**66/91**, exactly the baseline. Every consumer of `reminder_verbs` reads it after a reminder
route already exists. A verb that cannot win the route is a verb nobody asks about. The
grammar was the whole change.
Lemma matching in `isReminderVerb` now covers `разбудил` as well as `разбуди`, because one
lemma holds both. That is the trap `cmd/mavend/quiet_toggle.go` documents for `говори`. It
is tolerable here and not in the quiet toggle. `isReminderVerb` runs only on an utterance
already routed to reminder, and it decides where the subject starts. A quiet match flips a
daemon-wide setting from any channel.
@@ -0,0 +1,77 @@
# Russian acts reach tools
**06-08-2026. V-633.** Measured with `TestONNXBaseline`, 91-case RU routing fixture,
classifier plus the ONNX embedder. No LLM arm in this run.
## What was wrong
Three defects, tangled enough that fixing one alone would have looked like progress.
**No Russian utterance could reach a tool.** `DefaultActMatcher` in
`internal/router/slots.go` matched an exact English prefix, and `internal/tool.Matcher`
delegated straight to it. Its comment claimed "the production matcher is fuzzy, this is the
scaffold floor". There is no other matcher, and `DefaultGrammars` is the only place
`Slots.Fn` is set at stage 0, so the floor was the ceiling. Measured with a throwaway
matcher test over the seeds:
```text
"покажи статус nginx" ok=false "restart nginx" ok=true fn=restart
"сколько места на диске" ok=false "disk" ok=true fn=disk
"свободная память" ok=false "uptime" ok=true fn=uptime
"перезагрузи роутер" ok=false
```
55 of the 69 lines in `models/seeds/act.txt` routed to `IntentAct` and then fell to
`proposeGap`. Praxis was never affected: `PraxisGrammars` fills `Slots.Fn` itself.
**Seven lines were duplicated inside `models/seeds/query.txt`.** A duplicate is a second
identical vector, so it double-weights its region in nearest-neighbour scoring.
```text
сколько человек дома
кто сейчас дома
какая загрузка процессора
сколько свободного места на диске
какой ip адрес у сервера
какая версия софта
сколько оперативной памяти свободно
```
**`как дела у сервера` carried two labels**, in `query.txt:13` and `system.txt:9`. One
string, two identical vectors, disagreeing about the answer.
## The change
Tools carry spoken aliases as config data, in `deploy/mavend.json`. They are not a Russian
stem pattern in code, which CLAUDE.md forbids. They are not on the tool row either. An
ad-hoc tool enabled through `/tools` has no aliases and needs none.
Aliases and names compete in one table, longest phrase first, so "перезагрузи контейнер"
beats "перезагрузи" and "docker-restart" is not shadowed by "restart". Matching is on exact
leading tokens rather than lemmas. `перезагрузи роутер` is a command and `перезагрузил
роутер` is a fact, and a lemma cannot tell the two apart. That is the trap
`cmd/mavend/quiet_toggle.go` documents for `говори`.
The seven duplicates are gone, and `как дела у сервера` stays in `query.txt` only. It left
`system.txt` because system cannot answer it: `replySystem`'s
память/загрузк/аптайм arm returns "системная статистика пока не подключена." and always
did. That arm is a stub, not a mode, so the mode inventory now lists the shape as
`act.tool.hoststats`.
## Result
**69/91, 75.8% full, unchanged.** Clarify counts unchanged at 0 false and 8 missed.
Nothing moved, and that is the honest number. The fixture holds no host-stat case and no
Russian act that reaches a tool, so it cannot see either fix. The new coverage is
`TestActMatcherAliases`, which asserts the twelve utterances above plus the two refusals.
## What this does not fix
Argument quality. `статус sshd` reaches `systemctl status sshd`, but `логи nginx` reaches
`journalctl -n 50 -u nginx` only because the tool's argv prefix ends in `-u`. An alias whose
remainder is a Russian noun ("перезагрузи роутер") hands `systemctl restart роутер` a target
that does not exist. Free text still reaches an argv, which is the resolution rule the
ecosystem contract states for Hexis and not yet true here.
The fixture cannot measure any of this. That is the observability gap V-629 is for.
@@ -0,0 +1,82 @@
# Gemma as a label function, and what it found in the seeds
**06-08-2026. V-546.** Measured on workpc against gemma-4-12b-it-qat-UD-Q4_K_XL.
`docs/plans/18-routing-heads-on-e5-small.md` puts the labeled set at 20k examples through
gemma, costing 2 to 4 hours of the card. This is the check before spending that. Gemma
labels the 344 hand-written classifier seeds. Agreement with the label a person already
chose is a precision number rather than a guess.
## What ran
`cmd/labelgen` runs the stage 0 grammars. The real ones, in `buildRouter` order, minus
`wakeword-act`, whose allowlist is a deployment's enabled tool names. It labels 62 of 339
seed lines and leaves the rest.
The remaining 277 went to gemma through the daemon's own `routeSystem` prompt and
`routeGrammar`, both extracted from `internal/router/llmrouter.go` at run time rather than
retyped. Temperature 0.
## Cost
**334ms per call, 0 unparsed of 277.** The GBNF held every time. At that rate the plan's
20k examples is under two hours of card, which matches its estimate.
## The stage 0 rules as label functions
Agreement between the grammar's label and the seed file the line came from:
| seed intent | agree |
|---|---|
| reminder | 37/37 |
| query | 9/10 |
| system | 7/8 |
| act | 2/2 |
| chat | 0/4 |
| note | 0/1 |
`ReminderGrammar` at 37/37 is the evidence the plan wanted. The chat column is a defect
rather than a disagreement: `chatNarrativeTopics` is Russian-only, so `tell me about
yourself` survives the decline and routes IntentQuery with topic `yourself`. Filed as
V-625, which also records that `как дела у сервера` appears verbatim in two seed files
under two intents.
## Gemma against the seeds
**197/277, 71.1%.** By intent:
| seed intent | agree |
|---|---|
| note | 33/33 |
| act | 57/64 |
| fact | 37/40 |
| query | 51/54 |
| chat | 15/35 |
| system | 4/43 |
| reminder | 0/8 |
The number is not gemma's error rate. Reading the 80 disagreements, most are the seed files
and the prompt holding different definitions of the same intent. Three boundaries carry 42
of them, and V-626 is the fix:
- **system, 26 lines.** The prompt restricts system to the clock, the calendar date and the
assistant itself. The seeds also put sensor and host state there. That is the V-374 edit
of 31-07-2026, which the seeds never received.
- **world questions, 8 lines.** `почему небо голубое`, `why is the sky blue`. Written when
chat was the only honest destination for a question nothing could answer, and external
search now answers them.
- **bare verbs, 8 lines.** `поставь напоминание` with nothing to remind about. The prompt
calls that unknown. This one is not staleness. A nearest-neighbour centroid wants the
bare verb phrase, and that is what a seed file is for.
Four intents have not been redefined since the seeds were written: note, fact, query and
act. They agree at 178 of 191.
## What this says about the plan
Gemma is usable as a label function on those four and not on system, chat or a bare verb.
The plan already budgets a day of the owner reading the set. This says where to spend it.
It also says the two engines in the cascade are being taught different rules on 80 lines.
A routing measurement that swaps between the classifier and the router is measuring some of
that disagreement rather than the models.
@@ -0,0 +1,58 @@
# Moving the seed files onto the router prompt's boundaries
**06-08-2026. V-626.** Measured with `TestONNXBaseline`, 91-case RU routing fixture,
classifier plus the ONNX embedder. No LLM arm in this run.
`docs/evals/2026-08-06-seed-labels-vs-router-prompt.md` found three intent boundaries where
`models/seeds` and `routeSystem` disagree. This applies two of them and rejects the third,
because the third was measured and it costs a case.
## Baseline
**64/91, 70.3% full.** Latency p50 22.9ms.
## What moved
**Sensor and host state, system to query. 26 lines.** `какая температура воздуха`,
`сколько памяти занято`, `какой статус сервисов`. The prompt restricts system to the clock,
the calendar date and the assistant itself, which is the V-374 edit of 31-07-2026.
**World questions, chat to query. 8 lines.** `почему небо голубое`, `why is the sky blue`,
`как работает интернет`. Only the genuine world-knowledge lines. An opener about herself
stays in chat. `как тебя зовут` is a question word by rule 4 and about the assistant by
rule 8. The rules are ordered and rule 4 fires first, which reads wrong. That is a prompt
question rather than a seed question.
`system.txt` goes from 43 lines to 17 and `query.txt` from 64 to 98.
## Result
**66/91, 72.5% full.** Two cases gained, none lost.
- `en-sys-002` "turn quiet mode back on", quiet 2/3 to 3/3
- `ru-query-011` "почему сервер тормозит", homelab 5/6 to 6/6
Clarify counts unchanged at 0 false and 8 missed. The eight missed clarifies are the
`ambiguous` tag and this change does not touch them. `TestONNXRecall`, `TestONNXTopics`,
`TestONNXPersonalBoundary` and `TestONNXClaimConfidenceDistribution` all pass.
Thinning system to 17 lines did not hurt it. The two remaining system failures,
`какое число завтра` and `какой день недели послезавтра`, both failed at baseline too.
## The third boundary, measured and rejected
`reminder.txt` holds eight bare verbs: `поставь напоминание`, `создай напоминание`,
`set a reminder`. Rule 9 of the prompt calls an utterance with no named subject unknown.
By the prompt they do not belong in a reminder seed set.
Dropping them scores **65/91**, one below keeping them. `ru-rem-004` "поставь напоминание
через полчаса" falls from reminder to fact, because the centroid loses the phrase the
utterance is built from.
So the seed file and the prompt are not stale against each other here. They have different
jobs. A prompt classifies one utterance and can say it cannot. A nearest-neighbour centroid
is a shape to be near, and a bare verb phrase is part of that shape. The eight lines stay.
That distinction matters past this file. V-546 trains a classification head on labeled
utterances rather than a centroid, and the head is the prompt's kind of thing. These eight
lines are seed data and not training data.
+6
View File
@@ -48,11 +48,17 @@ type WeatherConfig struct {
// ToolConfig — one enabled tool. Name is the spoken verb ("restart"); Cmd is
// the fixed argv prefix (["systemctl","restart"]); Destructive marks acts that
// must not fire from the voice path (they need a confirm on an authed surface).
//
// Aliases are the spoken phrases that reach this tool, Russian included. They
// are config data rather than a pattern in code, and they match as exact leading
// tokens, so an imperative reaches the tool and the past tense of the same verb
// does not.
type ToolConfig struct {
Name string `json:"name"`
Scope string `json:"scope,omitempty"`
Cmd []string `json:"cmd"`
Destructive bool `json:"destructive,omitempty"`
Aliases []string `json:"aliases,omitempty"`
}
// Voice defaults, applied in normaliseVoice.
+3 -2
View File
@@ -173,10 +173,11 @@
]
},
"reminder_verbs": {
"note": "The imperatives that mean \"remind me\", in the forms he speaks. The same kind of set as capture_verbs and decided the same way: it is her vocabulary, not a discovery about Russian (Vikunja #530).",
"note": "The imperatives that mean \"remind me\", in the forms he speaks. The same kind of set as capture_verbs and decided the same way: it is her vocabulary, not a discovery about Russian (Vikunja #530). The alarm verbs joined them in V-627. \"разбуди меня в 6:30\" is a reminder that fires at the hour he gets up, and the set knew no form of it, so an alarm reached IntentReminder only by resembling one to the embedder.",
"words": [
"напомни", "напомните", "напомнить", "напоминай",
"remind"
"разбуди", "разбудите", "разбудить", "буди",
"remind", "wake"
]
},
"half_hour": {
+1 -1
View File
@@ -662,7 +662,7 @@ func (t *emptyFrameTransport) Call(ctx context.Context, req *rpcRequest) (*rpcRe
}
func (t *emptyFrameTransport) Notify(context.Context, string, any) error { return nil }
func (t *emptyFrameTransport) Close() error { return nil }
func (t *emptyFrameTransport) Close() error { return nil }
func TestResultlessResponseIsNotSuccess(t *testing.T) {
c := newClient("empty", &emptyFrameTransport{})
+90
View File
@@ -0,0 +1,90 @@
// Package modes holds the routing mode inventory: the roughly thirty distinct
// downstream behaviours mavend has, each mapped back to one of the seven public
// intents (V-631, umbrella V-628).
//
// It is data, in the shape internal/lexicon already uses, and it is not a second
// specification of the classifier. Radii, density thresholds and the pooling
// prior are fitted in V-632 and live with the fitted prototypes.
//
// Two rules decide whether something is a mode. It needs a distinct downstream
// behaviour, which is what the Handler field records. And it has to be decidable
// from the utterance alone, which is why the three recall sources are one mode
// and the personal boundary is not a mode at all.
package modes
import (
"embed"
"encoding/json"
"fmt"
)
//go:embed modes_v1.json
var files embed.FS
// Mode is one routing class.
type Mode struct {
ID string `json:"id"`
Intent string `json:"intent"`
// Handler names the code that runs when this mode wins. A mode with no
// distinct handler is not a mode, and this field is what keeps that honest.
Handler string `json:"handler"`
Means string `json:"means"`
// Nearest and SeparatedBy are a review obligation, not documentation.
// Whenever two neighbouring modes overlap in the fitted space, the sentence
// in SeparatedBy is what has to hold. If nothing separates them, they were
// one mode and this file is wrong.
Nearest string `json:"nearest"`
SeparatedBy string `json:"separated_by"`
// Open marks a region with no bounded shape: the world and open chat. Those
// carry RejectPolicy, and nothing else may.
Open bool `json:"open"`
RejectPolicy string `json:"reject_policy,omitempty"`
PrototypeCount int `json:"prototype_count"`
MinSeedExamples int `json:"min_seed_examples"`
Note string `json:"note,omitempty"`
Examples []string `json:"examples"`
}
// Inventory is the whole file. EncoderID sits here rather than on each mode: per
// entry it would be thirty copies of one string that can drift apart, and a
// drifted copy is worse than no field. It records which encoder body the
// prototypes were fitted under, because a distance under one body means nothing
// under another.
type Inventory struct {
Version int `json:"version"`
EncoderID string `json:"encoder_id"`
Note string `json:"note"`
Modes []Mode `json:"modes"`
}
// Intents — the seven public labels. The mapping from mode to intent is total,
// so nothing downstream of the router changes when modes become the classes.
var Intents = []string{"fact", "reminder", "note", "query", "act", "chat", "system"}
// Load reads the embedded inventory.
func Load() (*Inventory, error) {
b, err := files.ReadFile("modes_v1.json")
if err != nil {
return nil, fmt.Errorf("modes: read: %w", err)
}
var inv Inventory
if err := json.Unmarshal(b, &inv); err != nil {
return nil, fmt.Errorf("modes: parse: %w", err)
}
return &inv, nil
}
// ByID indexes the inventory.
func (inv *Inventory) ByID() map[string]Mode {
out := make(map[string]Mode, len(inv.Modes))
for _, m := range inv.Modes {
out[m.ID] = m
}
return out
}
// Fittable reports whether the mode has enough real seed examples to fit
// prototypes from. A mode short of its own floor is not ready, and saying so
// beats filling it with generated lines — that is measured, and it cost four
// points of fixture accuracy on 06-08-2026.
func (m Mode) Fittable() bool { return len(m.Examples) >= m.MinSeedExamples }
+185
View File
@@ -0,0 +1,185 @@
package modes
import (
"bufio"
"encoding/json"
"os"
"path/filepath"
"strings"
"testing"
)
func load(t *testing.T) *Inventory {
t.Helper()
inv, err := Load()
if err != nil {
t.Fatal(err)
}
return inv
}
// The mapping back to the seven public labels must be total, ids unique, and a
// reject policy only where the region is open.
func TestInventoryShape(t *testing.T) {
inv := load(t)
if inv.EncoderID == "" {
t.Error("no encoder_id: a fitted distance means nothing without the body it was fitted under")
}
valid := map[string]bool{}
for _, i := range Intents {
valid[i] = true
}
seen := map[string]bool{}
for _, m := range inv.Modes {
if seen[m.ID] {
t.Errorf("%s: duplicate id", m.ID)
}
seen[m.ID] = true
if !valid[m.Intent] {
t.Errorf("%s: intent %q is not one of the seven", m.ID, m.Intent)
}
if m.Handler == "" {
t.Errorf("%s: no handler, so it is not a mode", m.ID)
}
if m.SeparatedBy == "" {
t.Errorf("%s: no separated_by, so nothing states the review obligation", m.ID)
}
if m.Open && m.RejectPolicy == "" {
t.Errorf("%s: open with no reject_policy", m.ID)
}
if !m.Open && m.RejectPolicy != "" {
t.Errorf("%s: reject_policy on a bounded mode", m.ID)
}
if m.PrototypeCount < 1 {
t.Errorf("%s: prototype_count %d", m.ID, m.PrototypeCount)
}
}
// No id is a prefix of another. act.tool.hoststats was, and it turned out to
// run the same handler as act.tool: a read against a change is the tool row's
// destructive field, which the confirm gate already reads. Handler is prose,
// so a duplicated behaviour hides there. A nested id is the tell that shows.
for _, a := range inv.Modes {
for _, b := range inv.Modes {
if a.ID != b.ID && strings.HasPrefix(b.ID, a.ID+".") {
t.Errorf("%s is nested under %s, so one of them is not a mode", b.ID, a.ID)
}
}
}
// Nearest names a real mode, or the review obligation points at nothing.
for _, m := range inv.Modes {
if m.Nearest != "" && !seen[m.Nearest] {
t.Errorf("%s: nearest %q is not in the inventory", m.ID, m.Nearest)
}
if m.Nearest == m.ID {
t.Errorf("%s: nearest is itself", m.ID)
}
}
}
func repoRoot(t *testing.T) string {
t.Helper()
wd, err := os.Getwd()
if err != nil {
t.Fatal(err)
}
return filepath.Join(wd, "..", "..")
}
func seedRows(t *testing.T) map[string]bool {
t.Helper()
paths, err := filepath.Glob(filepath.Join(repoRoot(t), "models", "seeds", "*.txt"))
if err != nil || len(paths) == 0 {
t.Fatalf("no seed files: %v", err)
}
out := map[string]bool{}
for _, p := range paths {
f, err := os.Open(p)
if err != nil {
t.Fatal(err)
}
sc := bufio.NewScanner(f)
for sc.Scan() {
line := strings.TrimSpace(sc.Text())
if line == "" || strings.HasPrefix(line, "#") {
continue
}
out[strings.ToLower(line)] = true
}
f.Close()
}
return out
}
// Every example is a real seed row. Not generated: 202 reviewed generated
// contrast pairs cost four points of fixture accuracy on 06-08-2026, and the
// generated half of the corpus recovers its own generation prompt when clustered.
func TestExamplesComeFromSeedRows(t *testing.T) {
inv := load(t)
seeds := seedRows(t)
for _, m := range inv.Modes {
if len(m.Examples) == 0 {
continue
}
for _, e := range m.Examples {
if !seeds[strings.ToLower(e)] {
t.Errorf("%s: example %q is not a seed row", m.ID, e)
}
}
}
}
// The fixture is the sole held-out measurement. An example drawn from it makes
// every number after that unfalsifiable.
func TestExamplesAreNotFixtureCases(t *testing.T) {
inv := load(t)
b, err := os.ReadFile(filepath.Join(repoRoot(t), "internal", "router", "eval", "ru_routing_v1.json"))
if err != nil {
t.Skipf("fixture not readable: %v", err)
}
var raw struct {
Cases []struct {
Utterance string `json:"utterance"`
} `json:"cases"`
}
if err := json.Unmarshal(b, &raw); err != nil {
t.Fatalf("fixture shape changed, and this invariant must not silently skip: %v", err)
}
held := map[string]bool{}
for _, c := range raw.Cases {
if c.Utterance != "" {
held[strings.ToLower(strings.TrimSpace(c.Utterance))] = true
}
}
if len(held) == 0 {
t.Fatal("read no utterances from the fixture")
}
for _, m := range inv.Modes {
for _, e := range m.Examples {
if held[strings.ToLower(e)] {
t.Errorf("%s: example %q is a fixture case", m.ID, e)
}
}
}
}
// Not a failure, a report. Nine modes have zero real examples and they are the
// nine with no deterministic matcher, which is why V-629 and V-630 come before
// V-632: without persisted turns there is nothing to fit them from.
func TestFittableReport(t *testing.T) {
inv := load(t)
var ready, short, empty []string
for _, m := range inv.Modes {
switch {
case len(m.Examples) == 0:
empty = append(empty, m.ID)
case m.Fittable():
ready = append(ready, m.ID)
default:
short = append(short, m.ID)
}
}
t.Logf("modes: %d total, %d ready to fit, %d short of min_seed_examples, %d with no seed example at all",
len(inv.Modes), len(ready), len(short), len(empty))
t.Logf(" no examples: %s", strings.Join(empty, ", "))
t.Logf(" short: %s", strings.Join(short, ", "))
}
+382
View File
@@ -0,0 +1,382 @@
{
"version": 1,
"encoder_id": "e5-small-routing-v1",
"note": "Written from the handlers on 06-08-2026 for V-631. Examples are drawn only from train_seeds.jsonl, which is src=seed. The 91-case fixture is not touched. A mode whose examples list is short of min_seed_examples is not ready to fit, and that is the point of recording the number.",
"modes": [
{
"id": "query.fact-by-key",
"intent": "query",
"handler": "queryFactByKey",
"means": "he asks back a fact he stored, by its key",
"nearest": "query.recall",
"separated_by": "a key exists in the fact store; recall has to search",
"open": false,
"prototype_count": 2,
"min_seed_examples": 8,
"examples": ["сколько я спал сегодня", "какой сегодня вес", "сколько воды я выпил сегодня", "когда последний раз поливал цветы", "когда кормил кота в последний раз", "сколько времени прошло с последней тренировки", "how many hours did I sleep this week"]
},
{
"id": "query.day-plan",
"intent": "query",
"handler": "queryDayPlan",
"means": "what the day holds, asked with a plan word",
"nearest": "query.calendar",
"separated_by": "a plan word is present; the calendar listing is the general case",
"open": false,
"prototype_count": 2,
"min_seed_examples": 8,
"examples": ["что у меня сегодня по плану", "какие планы на завтра", "планы на сегодня", "какие у меня планы на завтра"]
},
{
"id": "query.habits",
"intent": "query",
"handler": "queryHabits",
"means": "what he usually does, asked with a habit marker",
"nearest": "query.calendar",
"separated_by": "обычно, каждый, по средам; not a single dated occasion",
"open": false,
"prototype_count": 2,
"min_seed_examples": 8,
"examples": []
},
{
"id": "query.tasks",
"intent": "query",
"handler": "queryTasks",
"means": "what is on the task board",
"nearest": "query.day-plan",
"separated_by": "a task noun or an explicit что … сделать, with no date",
"open": false,
"prototype_count": 2,
"min_seed_examples": 8,
"examples": []
},
{
"id": "query.attention",
"intent": "query",
"handler": "queryAttention",
"means": "what Praxis says needs looking at",
"nearest": "query.tasks",
"separated_by": "an attention marker; the board is Maven's, attention is Praxis's",
"open": false,
"prototype_count": 2,
"min_seed_examples": 8,
"examples": []
},
{
"id": "query.list",
"intent": "query",
"handler": "queryList",
"means": "what is on a standing list",
"nearest": "query.tasks",
"separated_by": "an explicit list marker",
"open": false,
"prototype_count": 2,
"min_seed_examples": 8,
"examples": []
},
{
"id": "query.money",
"intent": "query",
"handler": "queryMoney",
"means": "spending and balances, from the facts the poller wrote",
"nearest": "query.fact-by-key",
"separated_by": "a money noun plus an actual ask",
"open": false,
"prototype_count": 2,
"min_seed_examples": 8,
"examples": ["какой баланс на счету", "сколько стоит свет в этом месяце", "сколько электричества мы потратили"]
},
{
"id": "query.history",
"intent": "query",
"handler": "queryHistory",
"means": "what he told her, asked about the telling rather than the topic",
"nearest": "query.recall",
"separated_by": "both halves of a history phrase and no named topic",
"open": false,
"prototype_count": 2,
"min_seed_examples": 8,
"examples": []
},
{
"id": "query.feeds",
"intent": "query",
"handler": "queryFeeds",
"means": "what the feeds she reads are carrying",
"nearest": "query.world",
"separated_by": "a feed noun plus an ask; the world source would invent news",
"open": false,
"prototype_count": 2,
"min_seed_examples": 8,
"examples": ["что нового"]
},
{
"id": "query.home",
"intent": "query",
"handler": "queryHome",
"means": "the state of the house",
"nearest": "act.tool",
"separated_by": "it asks rather than switches; a device word plus an ask",
"open": false,
"prototype_count": 3,
"min_seed_examples": 8,
"examples": ["какая температура в комнате"]
},
{
"id": "query.network",
"intent": "query",
"handler": "queryNetwork",
"means": "what is on the LAN",
"nearest": "act.tool",
"separated_by": "the subject is the network, not this box",
"open": false,
"prototype_count": 2,
"min_seed_examples": 8,
"examples": ["что с интернетом", "какая скорость интернета", "сколько трафика сегодня"]
},
{
"id": "query.calendar",
"intent": "query",
"handler": "queryCalendar",
"means": "what the calendar holds, dated",
"nearest": "query.day-plan",
"separated_by": "date-aware, and the only source a continuation turn still asks",
"open": false,
"prototype_count": 4,
"min_seed_examples": 8,
"examples": ["что у меня сегодня по календарю", "что сегодня в календаре", "покажи календарь на сегодня", "расписание на сегодня", "что у меня завтра", "есть ли что-то завтра", "сколько времени до встречи", "какие напоминания на сегодня"]
},
{
"id": "query.weather",
"intent": "query",
"handler": "queryWeather",
"means": "the weather, outside",
"nearest": "query.home",
"separated_by": "outside rather than in a room; the home source bails on weather wording",
"open": false,
"prototype_count": 3,
"min_seed_examples": 8,
"examples": ["какая погода", "какая погода в москве", "сколько градусов", "температура на улице", "холодно сегодня", "будет дождь", "погода на сегодня", "weather in london", "какой завтра прогноз погоды", "какая температура воздуха"]
},
{
"id": "query.self",
"intent": "query",
"handler": "querySelf",
"means": "a question about her",
"nearest": "chat.open",
"separated_by": "it wants a fact about her, not a conversation",
"open": false,
"prototype_count": 2,
"min_seed_examples": 8,
"examples": ["как тебя зовут", "сколько тебе лет", "у тебя есть чувства", "do you have feelings"]
},
{
"id": "query.recall",
"intent": "query",
"handler": "queryEmbed, queryMemory, queryNotes",
"means": "search his own notes and memory for something he named",
"nearest": "query.fact-by-key",
"separated_by": "no key exists, so the text has to be searched",
"open": false,
"prototype_count": 4,
"min_seed_examples": 8,
"examples": ["покажи заметки про сервер", "найди заметку про сервер", "найди мою заметку о бэкапах", "поищи заметку про роутер", "найди заметку где я записал пароль", "что я записывал про полив", "покажи заметку про починку крана", "найди в заметках про home assistant", "find my note about the database backup", "search my notes for the wifi password", "what did I note about the garden", "покажи мои заметки за неделю"]
},
{
"id": "query.web",
"intent": "query",
"handler": "queryWeb",
"means": "read a page he named out loud",
"nearest": "query.world",
"separated_by": "he supplied the URL; it is an instruction, not a question",
"open": false,
"prototype_count": 2,
"min_seed_examples": 8,
"examples": []
},
{
"id": "query.world",
"intent": "query",
"handler": "querySearch, queryKiwix, queryGeneral",
"means": "anything outside his own data",
"nearest": "chat.open",
"separated_by": "a source can answer it; the personal boundary let it past",
"open": true,
"prototype_count": 6,
"min_seed_examples": 12,
"reject_policy": "no prototype within radius goes to the LLM fallback, which this path already pays for",
"examples": ["почему небо голубое", "что такое любовь", "как работает интернет", "почему трава зелёная", "откуда берётся дождь", "what is love", "why is the sky blue", "how does the internet work"]
},
{
"id": "act.tool",
"intent": "act",
"handler": "tools.Exec against the enabled allowlist",
"means": "switch, start, stop or read something the tool allowlist names",
"nearest": "query.home",
"separated_by": "it names a tool the allowlist carries; destructive is the tool rows field, not a mode of its own",
"open": false,
"prototype_count": 6,
"min_seed_examples": 8,
"note": "act.tool.hoststats was a mode here until 06-08-2026 and is not one: it ran the same tools.Exec, and read against change is the tool rows destructive field, which the confirm gate already reads. Its nine examples went with it, because they are question-shaped query seeds that no configured alias matches, so no tool answers them today. replySystem's память/загрузк/аптайм arm answers “системная статистика пока не подключена.” and always did.",
"examples": ["включи свет на кухне", "выключи кондиционер", "открой шторы", "закрой окно", "перезагрузи роутер", "запусти пылесос", "заблокируй дверь", "maven, restart nginx", "перезапусти nginx", "останови контейнер", "maven, сделай бэкап", "запусти обновление системы"]
},
{
"id": "act.taskstatus",
"intent": "act",
"handler": "resolveTaskStatus",
"means": "move an item on Maven's own board",
"nearest": "act.praxis",
"separated_by": "the board is Maven's; Praxis owns attention, not this",
"open": false,
"prototype_count": 2,
"min_seed_examples": 8,
"examples": []
},
{
"id": "act.praxis",
"intent": "act",
"handler": "handlePraxisAct",
"means": "surface, acknowledge or resolve a Praxis item",
"nearest": "act.taskstatus",
"separated_by": "the item lives in Praxis, and the three lifecycle words differ",
"open": false,
"prototype_count": 3,
"min_seed_examples": 8,
"examples": []
},
{
"id": "act.hexis",
"intent": "act",
"handler": "handleHexisAct",
"means": "execute a registered capability against a resolved entity",
"nearest": "act.tool",
"separated_by": "it names an entity Nexus must resolve before anything runs",
"open": false,
"prototype_count": 3,
"min_seed_examples": 8,
"examples": []
},
{
"id": "note.task",
"intent": "note",
"handler": "captureTaskFromNote",
"means": "he files work, which belongs in the task store",
"nearest": "note.recall",
"separated_by": "it is work to be done, not something to remember",
"open": false,
"prototype_count": 3,
"min_seed_examples": 8,
"examples": ["заметка: починить ручку на двери", "заметка: сменить масло в машине", "заметка: заменить лампочку в коридоре", "заметка: записаться к стоматологу", "заметка: переклеить обои в спальне", "заметка: проверить уровень масла", "запиши: проверить проводку на даче", "заметка: обновить прошивку роутера"]
},
{
"id": "note.list",
"intent": "note",
"handler": "captureListFromNote",
"means": "he adds to a standing list",
"nearest": "note.task",
"separated_by": "a list marker; the item is bought, not done",
"open": false,
"prototype_count": 2,
"min_seed_examples": 8,
"examples": ["запиши что нужно купить в магазине", "купить новый фильтр для аквариума", "заметка: купить новый фильтр для воды", "запиши: купить семена для огорода", "запиши: купить подарок на день рождения"]
},
{
"id": "note.recall",
"intent": "note",
"handler": "WriteNote plus memStore.Insert",
"means": "free text he wants indexed for later recall",
"nearest": "fact.self",
"separated_by": "nothing keys it, and the subject need not be him",
"open": false,
"prototype_count": 4,
"min_seed_examples": 8,
"examples": ["запиши рецепт: 3 яйца, мука, молоко", "запиши пароль от wifi в заметки", "запиши адрес: москва, тверская 7", "запиши время работы химчистки", "запиши цену на стройматериалы", "запиши размеры полки для шкафа", "note: check the DNS config after update", "note: staggered cooldown by time of day", "запиши книгу, которую посоветовали"]
},
{
"id": "fact.self",
"intent": "fact",
"handler": "actionFact, WriteFact kind=self",
"means": "a keyed, supersedable statement about him",
"nearest": "note.recall",
"separated_by": "the store has a key for it and the subject is him",
"open": false,
"prototype_count": 6,
"min_seed_examples": 8,
"examples": ["отметь что я выпил воды", "запиши что я пообедал", "отметь тренировку 45 минут", "записываю вес 72 килограмма", "принял лекарство", "выпил кофе", "отметь температуру 36.6", "записываю давление 120 на 80", "вес 73.5 килограмма", "сон 7 часов", "slept 6h", "walked 8000 steps"]
},
{
"id": "reminder.timed",
"intent": "reminder",
"handler": "actionReminder",
"means": "fire something at a time",
"nearest": "note.task",
"separated_by": "it carries a time; a task has none",
"open": false,
"prototype_count": 4,
"min_seed_examples": 8,
"examples": ["напомни завтра в 9 утра позвонить", "напомни через 4 часа размяться", "напомни завтра в 9 утра позвонить", "напомни в пятницу вынести мусор", "напомни через 15 минут снять бельё", "remind me in 30 minutes to drink water", "remind me at 6pm to take out the trash", "remind me tomorrow at 8am to call the doctor"]
},
{
"id": "system.clock",
"intent": "system",
"handler": "replySystem, the час/врем arm, ruClock",
"means": "the current time",
"nearest": "query.world",
"separated_by": "answered from the box's own clock, not from a source",
"open": false,
"prototype_count": 2,
"min_seed_examples": 8,
"examples": ["который час", "сколько времени", "сколько сейчас времени", "который час у нас", "который час в Москве"]
},
{
"id": "system.date",
"intent": "system",
"handler": "replySystem, the день/числ arm, ParseCalendarDate",
"means": "today's date or weekday",
"nearest": "query.calendar",
"separated_by": "it asks what day it is, not what is on that day",
"open": false,
"prototype_count": 2,
"min_seed_examples": 8,
"examples": ["какой сегодня день", "какое сегодня число", "какой сегодня день недели"]
},
{
"id": "system.presence",
"intent": "system",
"handler": "replySystem, the кто дома arm",
"means": "who is home",
"nearest": "query.home",
"separated_by": "the subject is people, not devices",
"open": false,
"prototype_count": 2,
"min_seed_examples": 8,
"examples": ["кто сейчас дома", "сколько человек дома", "есть ли кто дома", "все ли дома", "кто дома сейчас"]
},
{
"id": "system.quiet",
"intent": "system",
"handler": "quiet_toggle.go, matched pre-route",
"means": "turn the quiet mode on or off",
"nearest": "act.tool",
"separated_by": "it flips a daemon-wide setting from any channel, so the match is exact",
"open": false,
"prototype_count": 2,
"min_seed_examples": 8,
"examples": ["тихий режим", "не шуми", "не беспокоить", "включи тихий режим", "выключи тихий режим", "громкий режим", "quiet mode on", "quiet off"]
},
{
"id": "chat.open",
"intent": "chat",
"handler": "PhraseChat",
"means": "conversation, answered from the model with history",
"nearest": "query.self",
"separated_by": "nothing else claimed it and no source can answer it",
"open": true,
"prototype_count": 6,
"min_seed_examples": 12,
"reject_policy": "stays a measured positive class even while acting as a fallback region, or it silently absorbs every genuine miss",
"examples": ["привет", "как дела", "о чём поговорим", "чем занимаешься", "расскажи историю", "пошути", "анекдот", "что ты думаешь о жизни", "i'm bored", "tell me a joke", "what's up", "how are you"]
}
]
}
+72
View File
@@ -0,0 +1,72 @@
package router
import "testing"
// Before V-633 the matcher only ever matched the English tool name, so no
// Russian utterance could reach a tool: 55 of the 69 lines in models/seeds/act.txt
// routed to IntentAct and then fell to proposeGap. These are those lines.
func TestActMatcherAliases(t *testing.T) {
m := DefaultActMatcher{
Fns: []string{"status", "ps", "uptime", "disk", "memory", "logs",
"restart", "stop", "start", "docker-restart", "docker-stop", "reboot"},
Aliases: map[string][]string{
"status": {"статус", "покажи статус"},
"ps": {"статус докера", "что запущено"},
"uptime": {"покажи uptime", "как работает сервер"},
"disk": {"сколько места на диске"},
"memory": {"свободная память"},
"logs": {"покажи логи", "логи"},
"restart": {"перезагрузи", "перезапусти"},
"docker-restart": {"перезагрузи контейнер"},
"reboot": {"перезагрузи сервер"},
},
}
cases := []struct {
utterance string
wantFn string
wantArgs []string
}{
{"покажи статус nginx", "status", []string{"nginx"}},
{"статус sshd", "status", []string{"sshd"}},
{"статус докера", "ps", nil},
{"что запущено", "ps", nil},
{"сколько места на диске", "disk", nil},
{"свободная память", "memory", nil},
{"покажи uptime", "uptime", nil},
{"логи nginx", "logs", []string{"nginx"}},
{"перезагрузи nginx", "restart", []string{"nginx"}},
// Longest phrase first, so the two-word alias wins over the one word
// inside it and the act reaches the right tool.
{"перезагрузи контейнер maven", "docker-restart", []string{"maven"}},
{"перезагрузи сервер", "reboot", nil},
// The English names still match, unchanged.
{"restart nginx", "restart", []string{"nginx"}},
{"uptime", "uptime", nil},
}
for _, c := range cases {
fn, args, ok := m.Match(c.utterance)
if !ok || fn != c.wantFn {
t.Errorf("%q: got fn=%q ok=%v, want %q", c.utterance, fn, ok, c.wantFn)
continue
}
if len(args) != len(c.wantArgs) {
t.Errorf("%q: got args=%v, want %v", c.utterance, args, c.wantArgs)
continue
}
for i := range args {
if args[i] != c.wantArgs[i] {
t.Errorf("%q: got args=%v, want %v", c.utterance, args, c.wantArgs)
break
}
}
}
// Past tense is a fact, not a command, and aliases match exact tokens so it
// stays one. This is the trap cmd/mavend/quiet_toggle.go documents.
if fn, _, ok := m.Match("перезагрузил роутер"); ok {
t.Errorf("past tense reached a tool: fn=%q", fn)
}
// A phrase nobody configured still refuses, so the router can clarify.
if fn, _, ok := m.Match("свари кофе"); ok {
t.Errorf("unconfigured phrase reached a tool: fn=%q", fn)
}
}
+44 -14
View File
@@ -91,27 +91,57 @@ func (e Extractor) Extract(ctx context.Context, intent Intent, utterance string,
// --- default implementations (scaffold floors; production swaps wholesale) ---
// DefaultActMatcher — exact verb prefix + remainder-as-args. The production
// matcher is fuzzy; this is the scaffold floor. "restart nginx" → fn=restart,
// args=[nginx]. Not on the list → ok=false → the router refuses the act.
// DefaultActMatcher — exact phrase prefix + remainder-as-args. This is the only
// matcher there is: internal/tool.Matcher delegates here over the live enabled
// names, so a phrase that does not match exactly cannot reach a tool.
// "restart nginx" → fn=restart, args=[nginx]. Not on the list → ok=false → the
// router refuses the act.
//
// Aliases map a tool name to spoken phrases, so a Russian utterance reaches an
// English tool name. They come from the deployment config as data, never from a
// stem pattern in code, and they match as exact leading tokens: "перезагрузи
// роутер" is a command and "перезагрузил роутер" is a fact, and lemma matching
// cannot tell the two apart (the trap cmd/mavend/quiet_toggle.go documents).
type DefaultActMatcher struct {
Fns []string
Fns []string
Aliases map[string][]string
}
func (m DefaultActMatcher) Allowlist() []string { return m.Fns }
func (m DefaultActMatcher) Match(utterance string) (string, []string, bool) {
u := strings.TrimSpace(utterance)
// longest-verb-first so "restart" can't be shadowed by a shorter prefix.
sorted := append([]string(nil), m.Fns...)
sortDescByLen(sorted)
for _, fn := range sorted {
if u == fn {
return fn, nil, true
u := strings.TrimSpace(strings.ToLower(utterance))
u = strings.TrimRight(u, "?!.")
// One table of phrase → fn, so an alias and a name compete on length rather
// than on which loop ran first. Longest-first, so "docker-restart" cannot be
// shadowed by "restart" and a two-word alias beats the one-word one inside it.
phrases := make([]string, 0, len(m.Fns))
fnOf := make(map[string]string, len(m.Fns))
add := func(phrase, fn string) {
phrase = strings.TrimSpace(strings.ToLower(phrase))
if phrase == "" {
return
}
if strings.HasPrefix(u, fn+" ") {
rest := strings.TrimSpace(strings.TrimPrefix(u, fn+" "))
return fn, splitArgs(rest), true
if _, seen := fnOf[phrase]; seen {
return
}
fnOf[phrase] = fn
phrases = append(phrases, phrase)
}
for _, fn := range m.Fns {
add(fn, fn)
for _, a := range m.Aliases[fn] {
add(a, fn)
}
}
sortDescByLen(phrases)
for _, p := range phrases {
if u == p {
return fnOf[p], nil, true
}
if strings.HasPrefix(u, p+" ") {
rest := strings.TrimSpace(strings.TrimPrefix(u, p+" "))
return fnOf[p], splitArgs(rest), true
}
}
return "", nil, false
+33 -1
View File
@@ -2,6 +2,7 @@ package router
import (
"regexp"
"sort"
"strings"
"unicode"
@@ -94,10 +95,41 @@ func DefaultGrammars(actMatcher ActMatcher) []Grammar {
// stage-0 decision too (fillMatchedSlots in router.go). Before that it did not,
// so "напомни в 11:00 позвонить маме" reached the daemon with HasTime false and
// was asked "Когда?" about an hour he had just said.
//
// The verb alternation is built from lexicon.ReminderVerbs rather than written
// out (V-627). The literal here knew "напомни" and "remind me" and nothing
// else, so "разбуди меня в 6:30" never reached stage 0 — and it does not reach
// IntentReminder further down either, where the classifier calls it fact at
// 0.918. An alarm is a reminder that fires at the hour he gets up, and the
// verb that names one is her vocabulary, so it belongs in the lexicon with the
// rest of it.
//
// Longest-first ordering matters: Go's regexp alternation is leftmost-first,
// not longest-match, so "напомнить" listed after "напомни" would never match.
var reminderVerbPattern = regexp.MustCompile(
`(?i)^\s*(?:` + longestFirstAlternation(lexicon.ReminderVerbs()) +
`)\s*(?:мне|меня|me)?[\s,:]+(.+)$`)
// longestFirstAlternation joins a word set into a regexp alternation, longest
// alternative first, with every member escaped.
func longestFirstAlternation(set []string) string {
out := make([]string, 0, len(set))
for _, w := range set {
out = append(out, regexp.QuoteMeta(w))
}
sort.Slice(out, func(i, j int) bool {
if len(out[i]) != len(out[j]) {
return len(out[i]) > len(out[j])
}
return out[i] < out[j]
})
return strings.Join(out, "|")
}
func ReminderGrammar() Grammar {
return Grammar{
Name: "reminder-wakeword",
Pattern: regexp.MustCompile(`(?i)^\s*(?:напомни|remind me)[\s,:]+(.+)$`),
Pattern: reminderVerbPattern,
Build: func(m []string) (Decision, bool) {
rest := strings.TrimSpace(m[1])
if rest == "" {
+16 -3
View File
@@ -222,11 +222,23 @@ func runProcess(ctx context.Context, argv []string) (string, error) {
// router's default prefix logic over the current names. The interface's Match
// has no ctx, so it queries with a background context — an in-process sqlite
// read on the daemon.
type Matcher struct{ api API }
// Aliases are spoken phrases per tool name, wired from the deployment config so
// a Russian utterance can reach an English tool name. They are not stored on the
// tool row: an ad-hoc tool enabled through /tools has no aliases and needs none.
type Matcher struct {
api API
aliases map[string][]string
}
// NewMatcher builds a store-backed act matcher.
func NewMatcher(api API) *Matcher { return &Matcher{api: api} }
// WithAliases returns the matcher carrying spoken aliases per tool name.
func (m *Matcher) WithAliases(a map[string][]string) *Matcher {
m.aliases = a
return m
}
func (m *Matcher) names() []string {
ts, err := m.api.ListTools(context.Background(), "enabled")
if err != nil {
@@ -243,7 +255,8 @@ func (m *Matcher) names() []string {
// Allowlist — the enabled verbs (for stage-0 grammar wiring / introspection).
func (m *Matcher) Allowlist() []string { return m.names() }
// Match — longest-verb-first prefix match over the live enabled allowlist.
// Match — longest-phrase-first prefix match over the live enabled allowlist and
// its configured aliases.
func (m *Matcher) Match(utterance string) (string, []string, bool) {
return router.DefaultActMatcher{Fns: m.names()}.Match(utterance)
return router.DefaultActMatcher{Fns: m.names(), Aliases: m.aliases}.Match(utterance)
}
-8
View File
@@ -8,12 +8,9 @@
как прошёл день
расскажи про себя
ты мне нравишься
почему небо голубое
о чём поговорим
у тебя есть чувства
что такое любовь
расскажи историю
как работает интернет
шутка
анекдот
пошути
@@ -24,8 +21,6 @@
что нового
думаешь о чём-то
расскажи про космос
почему трава зелёная
откуда берётся дождь
что было интересного сегодня
как тебя зовут
сколько тебе лет
@@ -37,7 +32,4 @@ i'm bored
what's up
tell me a joke
do you have feelings
what is love
tell me about yourself
why is the sky blue
how does the internet work
+27
View File
@@ -62,3 +62,30 @@ what did I note about the garden
будет дождь
погода на сегодня
weather in london
сколько времени осталось до вечера
какая температура воздуха
есть ли кто дома
кто дома сейчас
все ли дома
сколько памяти занято
всё ли работает
сколько сервер работает без перезагрузки
когда сервер запускался
сколько аптайм
какой статус сервисов
все ли сервисы работают
что с интернетом
когда последний раз перезагружался
сколько трафика сегодня
какая скорость интернета
сколько процессов запущено
как загрузка системы
какая температура процессора
почему небо голубое
что такое любовь
как работает интернет
почему трава зелёная
откуда берётся дождь
what is love
why is the sky blue
how does the internet work
+1 -28
View File
@@ -5,36 +5,9 @@
который час в Москве
сколько сейчас времени
который час у нас
сколько времени осталось до вечера
какой сегодня день недели
какая температура воздуха
сколько человек дома
кто сейчас дома
есть ли кто дома
кто дома сейчас
все ли дома
сколько памяти занято
какая загрузка процессора
сколько свободного места на диске
какой ip адрес у сервера
как дела у сервера
всё ли работает
сколько сервер работает без перезагрузки
когда сервер запускался
какая версия софта
сколько аптайм
какой статус сервисов
все ли сервисы работают
что с интернетом
интернет работает
когда последний раз перезагружался
сколько трафика сегодня
какая скорость интернета
загрузка сети
сколько процессов запущено
как загрузка системы
сколько оперативной памяти свободно
какая температура процессора
тихий режим
тихо
не шуми
@@ -48,4 +21,4 @@
quiet mode on
quiet mode off
quiet on
quiet off
quiet off