Compare commits
5 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 00f9239ef9 | |||
| 3513e508b7 | |||
| c15c2b7bd2 | |||
| b6eaa704a2 | |||
| 2597a7b34a |
@@ -398,11 +398,60 @@ deliberately do not. "что у меня в списке покупок" matches
|
|||||||
the calendar there would take the list source off the turn.
|
the calendar there would take the list source off the turn.
|
||||||
|
|
||||||
Fixture unchanged at **69/91 classifier+ONNX**, measured both sides. That is the
|
Fixture unchanged at **69/91 classifier+ONNX**, measured both sides. That is the
|
||||||
expected result, because it scores intent and no case here changes intent. **The
|
expected result, because it scores intent and no case here changes intent.
|
||||||
destination has no fixture yet, so it has no accuracy number.** That and the model arm
|
|
||||||
are the follow-ups. The field is designed so a decider naming nothing costs nothing.
|
**The destination has its own fixture and its own number as of 08-08-2026**
|
||||||
It lands on V-546. Intent, mood and BIO slot tags were already three heads on one
|
(V-659, `docs/evals/2026-08-08-destination-fixture.md`). This section used to say
|
||||||
forward pass of the resident e5-small. Destination is a fourth head on the same pass.
|
it had neither. `want_source` on `eval.Case` is a pointer, because the destination
|
||||||
|
has three states and a bare string has two. Absent is every intent but query,
|
||||||
|
which never reaches `queryWalk`. Present and empty is the `SourceUnknown`
|
||||||
|
contract: name nothing and walk the chain. Present and named is a destination the
|
||||||
|
route must produce. Thirty-three of ninety-six cases carry one.
|
||||||
|
|
||||||
|
A destination miss does **not** fail the case. It lands in `Outcome.SourceReason`
|
||||||
|
and never in `Reasons`, so `Accuracy` and `IntentAccuracy` mean what they meant
|
||||||
|
and `SourceAccuracy` is a second number over the labelled cases only. Intent and
|
||||||
|
destination are two decisions, and one number hides which one moved. A route that
|
||||||
|
lost its intent scores no destination hit, or a clarify would satisfy an empty
|
||||||
|
label for free.
|
||||||
|
|
||||||
|
Measured classifier+ONNX: intent **73/96 (76.0%)**, destination **12/33 (36.4%)**.
|
||||||
|
The split is the finding. World is 5/5, because a stage 0 rule names it. The
|
||||||
|
`SourceUnknown` floor is 5/7. Calendar is 2/6, because the possessive agenda
|
||||||
|
rules deliberately do not name it. And **recall is 0/15, because nothing
|
||||||
|
anywhere names it**. Those turns are still answered, since the chain walks
|
||||||
|
recall early. Recall is the number the fourth head has to move.
|
||||||
|
|
||||||
|
Seven cases assert the floor and six of them are homelab operations. They
|
||||||
|
cluster because `SourceRecall`, `SourceNetwork` and `SourceAttention` overlap on
|
||||||
|
every question about the box. `mavpoll` writes its netdata and uptime-kuma
|
||||||
|
observations into the fact store recall reads. That is a finding about the enum,
|
||||||
|
not a gap in the labelling.
|
||||||
|
|
||||||
|
`baselineGrammars` in `eval_test.go` mirrors `buildRouter` and had drifted:
|
||||||
|
`WorldQueryGrammars` was wired into the daemon by V-655 and not into the mirror,
|
||||||
|
so the fixture scored a grammar set nobody runs. Fixed by V-659, worth 3 points of
|
||||||
|
destination and nothing else. Check that function when adding a grammar.
|
||||||
|
|
||||||
|
**The model arm landed the same day** (V-660,
|
||||||
|
`docs/evals/2026-08-08-destination-model-arm.md`). `routeGrammar` carries a
|
||||||
|
`source` rule closed over `router.Sources` plus the empty floor, so the model
|
||||||
|
cannot emit a destination that does not exist. The prompt lists the twelve in
|
||||||
|
Russian and says `""` is a normal answer to give often. `LLMRouter.Route` reads it
|
||||||
|
back through `ValidSource` and on `IntentQuery` alone. Against gemma-4-12b on the
|
||||||
|
workstation the cascade scores destination **24/33 (72.7%)** with intent unmoved
|
||||||
|
at 84.4%, and **recall goes 0/15 to 14/15**. The resident Qwen3-1.7B is
|
||||||
|
unmeasured, because it binds `--port 0` inside the container.
|
||||||
|
|
||||||
|
**Stage 0 now costs four destination points.** It did not before. The four cases
|
||||||
|
the cascade loses and the model alone wins are all calendar. The possessive
|
||||||
|
agenda rules claim them first and name nothing on purpose. That caution was free
|
||||||
|
while nothing downstream could name anything either. It is not free now, and the
|
||||||
|
fix is the owner's call rather than a quiet edit.
|
||||||
|
|
||||||
|
The last arm is V-546. Intent, mood and BIO slot tags were already three heads on
|
||||||
|
one forward pass of the resident e5-small. Destination is a fourth head on the
|
||||||
|
same pass, and 72.7% from a 12B teacher is the label source for training it.
|
||||||
|
|
||||||
## LLM output contract
|
## LLM output contract
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,82 @@
|
|||||||
|
# The first destination number
|
||||||
|
|
||||||
|
Measured 2026-08-08 on the classifier cascade with the ONNX multilingual
|
||||||
|
embedder, the configuration homesrv runs. `make t PKG=./internal/router/eval/
|
||||||
|
RUN=TestONNXBaseline V=1`. Covers V-659, the follow-up V-655 named.
|
||||||
|
|
||||||
|
## What was measured
|
||||||
|
|
||||||
|
V-655 split a routing decision in two. The cascade sorts an utterance into one
|
||||||
|
of seven intents, and `Decision.Source` then says where the answer lives. The
|
||||||
|
first half had a fixture. The second half arrived with none, so twelve
|
||||||
|
destinations shipped with no accuracy number.
|
||||||
|
|
||||||
|
`want_source` is now a field on `eval.Case`. It is a pointer, because the
|
||||||
|
destination has three states and a bare string has two. Absent is every intent
|
||||||
|
but query, which never reaches `queryWalk`. Present and empty is the
|
||||||
|
`SourceUnknown` contract: name nothing and let the daemon walk the chain.
|
||||||
|
Present and named is a destination the route must produce.
|
||||||
|
|
||||||
|
Thirty-three of the ninety-six cases carry one. A destination miss does not
|
||||||
|
fail the case, so `Accuracy` and `IntentAccuracy` mean what they meant.
|
||||||
|
`SourceAccuracy` is a second number over the labelled cases only.
|
||||||
|
|
||||||
|
## Result
|
||||||
|
|
||||||
|
Intent is **73/96 (76.0%)**, against 69/91 (75.8%) before. Four of the five new
|
||||||
|
cases pass and no existing case moved.
|
||||||
|
|
||||||
|
Destination is **12/33 (36.4%)**, and the split is the whole finding.
|
||||||
|
|
||||||
|
| destination | scored | note |
|
||||||
|
|---|---|---|
|
||||||
|
| world | 5/5 | `WorldQueryGrammars` names it at stage 0 |
|
||||||
|
| the `SourceUnknown` floor | 5/7 | the two misses lost the intent first |
|
||||||
|
| calendar | 2/6 | `calendar-query` names it, the possessive agenda rules do not |
|
||||||
|
| recall | 0/15 | nothing anywhere names it |
|
||||||
|
|
||||||
|
Recall is the number to move. Fifteen cases ask about his own words and his own
|
||||||
|
facts. The route lands `query` on eleven of them and the destination comes back
|
||||||
|
empty every time. Those turns are answered today, because the daemon walks the
|
||||||
|
chain in order and the three recall passes are early in it. What is missing is a
|
||||||
|
decider that says so, and that is the fourth head on V-546.
|
||||||
|
|
||||||
|
Two cases labelled the floor lost their intent before a destination was
|
||||||
|
possible. A clarify names nothing, so it would satisfy an empty label for free.
|
||||||
|
`Score` requires the route to land the case's intent before it credits a
|
||||||
|
destination hit, or the floor label would score itself.
|
||||||
|
|
||||||
|
## Seven cases assert the floor, and six of them cluster
|
||||||
|
|
||||||
|
The six are homelab operations. `SourceRecall`, `SourceNetwork` and
|
||||||
|
`SourceAttention` overlap on every question about the box, because `mavpoll`
|
||||||
|
writes its netdata and uptime-kuma observations into the fact store recall
|
||||||
|
reads. "почему сервер тормозит" is answerable from all three. Naming one takes
|
||||||
|
the other two off the turn.
|
||||||
|
|
||||||
|
That is a finding about the enum rather than a gap in the labelling. The floor
|
||||||
|
is the right answer there and the fixture now says so out loud.
|
||||||
|
|
||||||
|
## A drift the labelling found
|
||||||
|
|
||||||
|
`WorldQueryGrammars` went into `buildRouter` with V-655 and never into
|
||||||
|
`baselineGrammars`, the fixture's mirror of it. So the fixture was scoring a
|
||||||
|
grammar set the daemon does not run. The comment above that function forbids
|
||||||
|
exactly that. Adding it moved the destination number from 9/33 to 12/33 and
|
||||||
|
moved nothing else.
|
||||||
|
|
||||||
|
The three cases it recovered are `что такое TCP?`, `сколько будет 17 на 23?`
|
||||||
|
and `кто такой Линус Торвальдс?`. All three already routed `query` through
|
||||||
|
`NarrativeQueryGrammars`. So the drift was invisible to every number this
|
||||||
|
fixture reported, until the destination had one of its own.
|
||||||
|
|
||||||
|
## What this does not measure
|
||||||
|
|
||||||
|
The model arm. This is the classifier cascade, which names a destination only
|
||||||
|
where a stage 0 rule filled one in. The resident model has no destination in
|
||||||
|
its router prompt yet, so 36.4% is a floor and not a comparison.
|
||||||
|
|
||||||
|
Two pairs of cases are the same utterance. `ru-query-020` and `ru-query-024`
|
||||||
|
are both "что дальше?", and `ru-query-021` and `ru-query-025` are both
|
||||||
|
"расскажи про битву при Ватерлоо". They differ in tags and note only, so both
|
||||||
|
pairs are counted twice here and in every earlier number this fixture reported.
|
||||||
@@ -0,0 +1,66 @@
|
|||||||
|
# The destination, with a model that can name one
|
||||||
|
|
||||||
|
Measured 2026-08-08 against gemma-4-12b on the workstation, the same 96-case
|
||||||
|
fixture V-659 built. Covers V-660.
|
||||||
|
|
||||||
|
```sh
|
||||||
|
no_proxy='*' MAVEN_LLM_URL=http://192.168.1.105:8080 \
|
||||||
|
make t PKG=./internal/router/eval/ RUN=TestLLMRouterBaseline V=1
|
||||||
|
```
|
||||||
|
|
||||||
|
## The gap was structural
|
||||||
|
|
||||||
|
V-659 measured the destination at 12/33 on the classifier cascade, with recall
|
||||||
|
at 0/15. Nothing in `routeSystem` named a `Source` and `routeGrammar` could not
|
||||||
|
emit one, so the resident model had no string to write. That is the shape V-517
|
||||||
|
measured for Praxis reach at 0/12: not a weak model, an absent contract.
|
||||||
|
|
||||||
|
`routeGrammar` now carries a `source` rule closed over `router.Sources` plus the
|
||||||
|
empty floor. The prompt lists the twelve destinations in Russian and says that
|
||||||
|
`""` is a normal answer to give often.
|
||||||
|
|
||||||
|
## Result
|
||||||
|
|
||||||
|
| run | intent | destination |
|
||||||
|
|---|---|---|
|
||||||
|
| classifier + ONNX (V-659) | 73/96 (76.0%) | 12/33 (36.4%) |
|
||||||
|
| gemma-4-12b alone | 79/96 intent-only (82.3%) | 26/33 (78.8%) |
|
||||||
|
| cascade + gemma-4-12b + hash fallback | 81/96 (84.4%) | 24/33 (72.7%) |
|
||||||
|
|
||||||
|
Recall is the move: 0/15 to 14/15. Intent did not shift, which was the
|
||||||
|
constraint. The prompt is shared, so a destination rule that costs routing
|
||||||
|
points is not a win.
|
||||||
|
|
||||||
|
The eight llm-only errors are the eight `want_clarify` cases. The model returned
|
||||||
|
`unknown` on every one, which is correct, and the llm-only harness surfaces a
|
||||||
|
decline as an error by design.
|
||||||
|
|
||||||
|
## Stage 0 now costs four destination points
|
||||||
|
|
||||||
|
The four cases the cascade loses and the model alone wins are all calendar. The
|
||||||
|
possessive agenda rules claim them at stage 0 and deliberately name nothing.
|
||||||
|
"что у меня в списке покупок" matches the same rule. Naming the calendar there
|
||||||
|
would take the list source off the turn (V-655).
|
||||||
|
|
||||||
|
So a rule written to be careful about the list now blocks a model that would
|
||||||
|
have named the calendar correctly. Before V-660 that caution was free, because
|
||||||
|
nothing downstream of stage 0 could name anything either.
|
||||||
|
|
||||||
|
Three ways out, and each costs something. Split the possessive rule so the
|
||||||
|
calendar-shaped half names its destination. Let a later stage overwrite an empty
|
||||||
|
destination a grammar left behind, which reverses "a matched value always wins".
|
||||||
|
Or leave it, on the argument that four points is cheap next to a wrong
|
||||||
|
destination on a shopping list. This wants the owner's call rather than a quiet
|
||||||
|
edit.
|
||||||
|
|
||||||
|
## What this does not measure
|
||||||
|
|
||||||
|
The resident Qwen3-1.7B, which is what homesrv runs. It binds `--port 0` inside
|
||||||
|
the container and no host process can reach it. Scoring it needs a second
|
||||||
|
llama-server on a fixed port. The workstation is never assumed
|
||||||
|
up, so the homesrv number is the one that decides whether this ships on by
|
||||||
|
default.
|
||||||
|
|
||||||
|
The fixture is 33 labelled destinations over twelve values. Recall carries 15 of
|
||||||
|
them and five destinations carry none at all. A per-destination number below
|
||||||
|
world, recall, calendar and the floor is not supported by this fixture.
|
||||||
@@ -0,0 +1,142 @@
|
|||||||
|
# MASSIVE Russian warm-start for the routing heads
|
||||||
|
|
||||||
|
Measured 2026-08-08 on workpc (Radeon RX 7900 GRE, ROCm). Covers V-546 step 2.
|
||||||
|
Workspace is `~/Programs/embed-training` on workpc, scripts `train_massive.py`,
|
||||||
|
`ab_run.py`, `ab.sh`, `probe_time.py`.
|
||||||
|
|
||||||
|
## What was trained
|
||||||
|
|
||||||
|
Two heads on a copy of multilingual-e5-small: `Linear(384, 60)` for MASSIVE's
|
||||||
|
own intents over a masked mean pool, `Linear(384, 111)` per token for BIO slot
|
||||||
|
tags. MASSIVE's label sets verbatim, no alignment to Maven's 7 intents. The
|
||||||
|
intent head is an auxiliary loss that shapes the pooled vector and is thrown
|
||||||
|
away.
|
||||||
|
|
||||||
|
Data is `amazon-massive-dataset-1.1` pulled from S3. The Hugging Face repo is
|
||||||
|
script-only and `datasets` 5.0 refuses those, so `load_dataset` cannot fetch it.
|
||||||
|
`ru-RU` is 11,514 train, 2,033 dev, 2,974 test, 60 intents, 55 slots, 111 BIO
|
||||||
|
labels. All 16,521 rows survived span alignment: `annot_utt` re-tokenised to its
|
||||||
|
own `utt` on every one.
|
||||||
|
|
||||||
|
Hyperparameters match `train_intent.py`, so the two runs differ in data only.
|
||||||
|
Frozen XLM-R vocabulary, body 2e-5, heads 1e-3, batch 32, sequence 64, 10
|
||||||
|
epochs. MASSIVE's own dev partition selects the epoch, on slot F1 with intent
|
||||||
|
accuracy as tiebreak. Selecting on 60-class intent accuracy would optimise a
|
||||||
|
head that gets deleted.
|
||||||
|
|
||||||
|
## Result
|
||||||
|
|
||||||
|
Epoch 9 of 10 by dev slot F1. Held-out MASSIVE test: intent 86.2%, slot span
|
||||||
|
F1 71.5% (P 68.5, R 74.8). Peak 1.70GB of 17.2GB, about 22 seconds an epoch,
|
||||||
|
under 4 minutes end to end. Dev slot F1 climbed monotonically to epoch 9 and
|
||||||
|
fell at 10, so 10 epochs was the right budget.
|
||||||
|
|
||||||
|
Ten slot types sit at 0% test recall. Every one of them has 1 to 7 test
|
||||||
|
instances: `alarm_type` has 3, `drink_type` has 1. That is support in MASSIVE's
|
||||||
|
Russian split, not a tagger failure. `playlist_name` at 6% of 16 is the first
|
||||||
|
real miss.
|
||||||
|
|
||||||
|
## The intent A/B, and why it settles nothing
|
||||||
|
|
||||||
|
`train_intent.py` was run against both bodies, three seeds by two smoothing
|
||||||
|
settings, on `train_v4.jsonl`. It is v4 and not v5 because v4 is what
|
||||||
|
`sweep2.log` measured. `ab_run.py` strips a `--base` flag onto the module global, so
|
||||||
|
`train_intent.py` is unmodified and its baseline stays reproducible. The stock
|
||||||
|
arm reproduced `sweep2.log` line for line.
|
||||||
|
|
||||||
|
Fixture accuracy, 91 cases, one case is 1.1 points:
|
||||||
|
|
||||||
|
| seed / smooth | stock | warm-started |
|
||||||
|
|---|---|---|
|
||||||
|
| 0 / 0.0 | 94.0% | 92.8% |
|
||||||
|
| 0 / 0.1 | 95.2% | 92.8% |
|
||||||
|
| 1 / 0.0 | 95.2% | 94.0% |
|
||||||
|
| 1 / 0.1 | 95.2% | 97.6% |
|
||||||
|
| 2 / 0.0 | 92.8% | 94.0% |
|
||||||
|
| 2 / 0.1 | 92.8% | 96.4% |
|
||||||
|
|
||||||
|
Mean 94.2% against 94.6%. That is +0.4 points, about a third of one case, and
|
||||||
|
inside seed noise. Spread widened. Stock lands in a 2.4-point band and
|
||||||
|
warm-started in a 4.8-point one. The warm-started arm holds both the best result
|
||||||
|
of the sweep and a tie for the worst. Seed 0 is the bad arm and it fails in a
|
||||||
|
specific way. Its dev peaks at epoch 2 and 3 and never improves, where stock
|
||||||
|
peaks around 7. The dev slice is a quarter of the seed rows. That is small
|
||||||
|
enough that early stopping is fragile when the body arrives already fitted.
|
||||||
|
|
||||||
|
**The A/B was never the test.** Intent had at most 4.8 points of headroom here.
|
||||||
|
MASSIVE was not trained for Maven's intents. Read it as "the warm-start does not
|
||||||
|
cost intent accuracy", nothing more.
|
||||||
|
|
||||||
|
## The measurement that does mean something
|
||||||
|
|
||||||
|
`want_time` is the one slot Maven's fixture scores, and MASSIVE has `time` and
|
||||||
|
`date`. Restricted to those two slot types, F1 is 74.9% over 609 gold spans on the
|
||||||
|
MASSIVE ru test split. Precision is 71.5 and recall 78.7. That beats the 71.5%
|
||||||
|
all-slot figure. Of the 530 test utterances carrying a time or a date, 73.4% get
|
||||||
|
every such span exactly right.
|
||||||
|
|
||||||
|
Out of domain matters more, because Maven's traffic is not this corpus. Ten
|
||||||
|
Maven-shaped utterances, none of them in MASSIVE:
|
||||||
|
|
||||||
|
| utterance | tagged |
|
||||||
|
|---|---|
|
||||||
|
| `напомни в 11:00 позвонить маме` | `time='11:00'`, `relation='маме'` |
|
||||||
|
| `напомни завтра в семь утра выпить таблетки` | `date='завтра'`, `time='семь утра'` |
|
||||||
|
| `поставь будильник на полседьмого` | `time='полседьмого'` |
|
||||||
|
| `через двадцать минут напомни про чайник` | `time='двадцать минут'` |
|
||||||
|
| `напомни в пятницу вечером забрать посылку` | `date='пятницу'`, `timeofday='вечером'` |
|
||||||
|
| `что у меня сегодня после обеда` | `date='сегодня'`, `time='после'`, `timeofday='обеда'` |
|
||||||
|
| `запиши что кофе закончился` | nothing |
|
||||||
|
| `что такое TCP` | `definition_word='TCP'` |
|
||||||
|
|
||||||
|
The first row is the V-572 defect utterance. `ReminderGrammar` handed the daemon
|
||||||
|
`HasTime: false` there, and the daemon asked "Когда?" at a sentence that had
|
||||||
|
already said when. `полседьмого` is a colloquial half-past that no digit pattern
|
||||||
|
catches. `запиши что кофе закончился` correctly carries nothing, because a note
|
||||||
|
has no time.
|
||||||
|
|
||||||
|
Two errors. `после обеда` split into `time='после'` plus `timeofday='обеда'`
|
||||||
|
when it is one span, and `через двадцать минут` dropped its `через`. Both are
|
||||||
|
boundary errors on spans the tagger did find.
|
||||||
|
|
||||||
|
Unplanned: `что такое TCP` returned `definition_word='TCP'`. MASSIVE has a slot
|
||||||
|
for the thing being asked about, which is a `SourceWorld` signal sitting in a
|
||||||
|
head already trained.
|
||||||
|
|
||||||
|
Ten hand-picked utterances are evidence, not a fixture.
|
||||||
|
|
||||||
|
## What this does not measure
|
||||||
|
|
||||||
|
Maven has no span fixture. `want_time` and `want_fn` are presence booleans and
|
||||||
|
`want_fact_key` is an exact string match, so nothing in the repo can score a
|
||||||
|
71.5% span tagger. Destination got one the same day, at 12/33 on the classifier
|
||||||
|
cascade: see `2026-08-08-destination-fixture.md`.
|
||||||
|
|
||||||
|
The missing span fixture is why the warm-start stays unjudged against Maven
|
||||||
|
rather than against MASSIVE.
|
||||||
|
|
||||||
|
## Datasets ruled out
|
||||||
|
|
||||||
|
Checked on 2026-08-08 and rejected as label sources:
|
||||||
|
|
||||||
|
- **MASSIVE's other 50 locales** ship in the same tarball and are parallel by id.
|
||||||
|
Co-training on them is free and unmeasured. English was ruled out by the owner
|
||||||
|
on 2026-08-08.
|
||||||
|
- **CLINC150** is reachable as parquet, 150 intents and 1,200 explicit
|
||||||
|
out-of-scope queries, English only. Its value is the labeled out-of-scope set
|
||||||
|
for fitting the energy threshold, not intent labels.
|
||||||
|
- **`d0rj/dolphin-ru`**, roughly 2.8M rows of FLAN-style tasks translated to
|
||||||
|
Russian. No intent, no slots, and not utterances anyone says to an assistant.
|
||||||
|
- **`psytechlab/EmpatheticIntents-ru`**, 24,856 rows of translated
|
||||||
|
EmpatheticDialogues with 32 emotion labels. Maven's mood enum is `neutral,
|
||||||
|
happy, thinking, tired, confused` and it describes her own reply, not the
|
||||||
|
speaker's emotion. No mapping exists.
|
||||||
|
- **`ai-forever/MERA`** and **`RussianNLP/russian_super_glue`**, benchmark
|
||||||
|
harnesses. Rows are prompt templates with `{toxic_comment}` placeholders.
|
||||||
|
- **`ZeroAgency/ru-big-russian-dataset`**, an LLM-judge quality corpus. Its
|
||||||
|
`question` and `classified_topic` columns are a usable Russian out-of-scope
|
||||||
|
pool for threshold fitting. That is the one thing CLINC150 can only supply in
|
||||||
|
English. The questions are long and written, so they belong in the negative
|
||||||
|
set, never in the in-scope `query` training set.
|
||||||
|
- No second Russian slot-filling corpus exists. The xSID mirrors are 404,
|
||||||
|
MultiATIS++ has no Russian, SLURP is not on the Hub.
|
||||||
@@ -36,17 +36,27 @@ var fixtureJSON []byte
|
|||||||
//
|
//
|
||||||
// Intent is empty exactly when WantClarify is set: the contract there is that
|
// Intent is empty exactly when WantClarify is set: the contract there is that
|
||||||
// the router refuses instead of guessing.
|
// the router refuses instead of guessing.
|
||||||
|
//
|
||||||
|
// WantSource is a pointer because the destination has three states and a bare
|
||||||
|
// string only has two (V-659). Absent means the case does not score a
|
||||||
|
// destination at all, which is every intent but query: a fact, a reminder, a
|
||||||
|
// note, an act, a chat or a system turn never reaches queryWalk. Present and
|
||||||
|
// empty is the SourceUnknown contract — the decider must name nothing and let
|
||||||
|
// the daemon walk the whole chain, which is the right answer whenever two
|
||||||
|
// destinations can both answer and the utterance does not choose. Present and
|
||||||
|
// named is a destination the route must produce.
|
||||||
type Case struct {
|
type Case struct {
|
||||||
ID string `json:"id"`
|
ID string `json:"id"`
|
||||||
Utterance string `json:"utterance"`
|
Utterance string `json:"utterance"`
|
||||||
Lang string `json:"lang"`
|
Lang string `json:"lang"`
|
||||||
Intent router.Intent `json:"intent"`
|
Intent router.Intent `json:"intent"`
|
||||||
WantTime bool `json:"want_time"`
|
WantTime bool `json:"want_time"`
|
||||||
WantFn bool `json:"want_fn"`
|
WantFn bool `json:"want_fn"`
|
||||||
WantFactKey string `json:"want_fact_key"`
|
WantFactKey string `json:"want_fact_key"`
|
||||||
WantClarify bool `json:"want_clarify"`
|
WantClarify bool `json:"want_clarify"`
|
||||||
Tags []string `json:"tags"`
|
WantSource *router.Source `json:"want_source,omitempty"`
|
||||||
Note string `json:"note"`
|
Tags []string `json:"tags"`
|
||||||
|
Note string `json:"note"`
|
||||||
}
|
}
|
||||||
|
|
||||||
// Fixture — the versioned envelope, same shape as
|
// Fixture — the versioned envelope, same shape as
|
||||||
@@ -118,6 +128,11 @@ type Outcome struct {
|
|||||||
// (a slot gap is a parser fix; a wrong intent is a router fix).
|
// (a slot gap is a parser fix; a wrong intent is a router fix).
|
||||||
IntentOK bool
|
IntentOK bool
|
||||||
Reasons []string
|
Reasons []string
|
||||||
|
// SourceReason is set when the case labelled a destination and the route
|
||||||
|
// named a different one. It is kept out of Reasons on purpose: the
|
||||||
|
// destination is the second half of a route and it is scored separately,
|
||||||
|
// so a wrong destination must not move the intent number (V-659).
|
||||||
|
SourceReason string
|
||||||
}
|
}
|
||||||
|
|
||||||
// Report — the aggregate. Accuracy is the headline; the rest exists so a
|
// Report — the aggregate. Accuracy is the headline; the rest exists so a
|
||||||
@@ -139,7 +154,15 @@ type Report struct {
|
|||||||
// (reminder grammar → applyAction's time parser). Not a miss, but not a
|
// (reminder grammar → applyAction's time parser). Not a miss, but not a
|
||||||
// full router-level win either; tracked so the two aren't conflated.
|
// full router-level win either; tracked so the two aren't conflated.
|
||||||
SlotsDeferred int
|
SlotsDeferred int
|
||||||
Outcomes []Outcome
|
// SourceTotal counts the cases carrying a want_source, and SourceHit the
|
||||||
|
// ones whose route named it. Reported apart from Passed because intent and
|
||||||
|
// destination are two decisions, and one number hides which one moved.
|
||||||
|
SourceTotal int
|
||||||
|
SourceHit int
|
||||||
|
// SourceConfusion counts want→got destination pairs. "" reads as the
|
||||||
|
// SourceUnknown floor on either side.
|
||||||
|
SourceConfusion map[string]int
|
||||||
|
Outcomes []Outcome
|
||||||
// Confusion counts want→got intent pairs, decided cases only.
|
// Confusion counts want→got intent pairs, decided cases only.
|
||||||
Confusion map[string]int
|
Confusion map[string]int
|
||||||
// ByTag accuracy for the fixture's tags ("hard", "homelab", …).
|
// ByTag accuracy for the fixture's tags ("hard", "homelab", …).
|
||||||
@@ -172,6 +195,17 @@ func (r Report) IntentAccuracy() float64 {
|
|||||||
return float64(r.IntentHit) / float64(r.Total)
|
return float64(r.IntentHit) / float64(r.Total)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// SourceAccuracy — fraction of the labelled cases whose route named the right
|
||||||
|
// destination. Denominator is SourceTotal and not Total, because most of the
|
||||||
|
// fixture never reaches a query source and scoring those would report a
|
||||||
|
// percentage of nothing.
|
||||||
|
func (r Report) SourceAccuracy() float64 {
|
||||||
|
if r.SourceTotal == 0 {
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
return float64(r.SourceHit) / float64(r.SourceTotal)
|
||||||
|
}
|
||||||
|
|
||||||
// Score runs every case through r and aggregates. It never fails the run on a
|
// Score runs every case through r and aggregates. It never fails the run on a
|
||||||
// route error — an erroring case scores as a miss and is counted in Errors,
|
// route error — an erroring case scores as a miss and is counted in Errors,
|
||||||
// because "the model was down" and "the model was wrong" are different numbers
|
// because "the model was down" and "the model was wrong" are different numbers
|
||||||
@@ -186,11 +220,12 @@ func Score(ctx context.Context, name string, r Router, f Fixture) (Report, error
|
|||||||
return Report{}, err
|
return Report{}, err
|
||||||
}
|
}
|
||||||
rep := Report{
|
rep := Report{
|
||||||
Name: name,
|
Name: name,
|
||||||
Total: len(f.Cases),
|
Total: len(f.Cases),
|
||||||
Confusion: map[string]int{},
|
Confusion: map[string]int{},
|
||||||
ByTag: map[string]TagStat{},
|
SourceConfusion: map[string]int{},
|
||||||
ByLang: map[string]TagStat{},
|
ByTag: map[string]TagStat{},
|
||||||
|
ByLang: map[string]TagStat{},
|
||||||
}
|
}
|
||||||
lat := make([]time.Duration, 0, len(f.Cases))
|
lat := make([]time.Duration, 0, len(f.Cases))
|
||||||
|
|
||||||
@@ -242,6 +277,29 @@ func Score(ctx context.Context, name string, r Router, f Fixture) (Report, error
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// The destination is scored outside the switch and outside Pass. A case
|
||||||
|
// that clarified or landed the wrong intent named no destination, and
|
||||||
|
// that is a real miss rather than a case to skip — otherwise the
|
||||||
|
// denominator quietly drops every turn the route already lost. Only a
|
||||||
|
// route error is skipped, because "the model was down" is the Errors
|
||||||
|
// number and not a destination result.
|
||||||
|
if c.WantSource != nil && err == nil {
|
||||||
|
rep.SourceTotal++
|
||||||
|
switch {
|
||||||
|
case !o.IntentOK:
|
||||||
|
// The route never got to a destination, so a match on the
|
||||||
|
// SourceUnknown floor here would be a coincidence scored as a
|
||||||
|
// win: a clarify names nothing and would satisfy "" for free.
|
||||||
|
o.SourceReason = fmt.Sprintf("no destination, route missed %q", c.Intent)
|
||||||
|
rep.SourceConfusion[string(*c.WantSource)+"→(no route)"]++
|
||||||
|
case d.Source == *c.WantSource:
|
||||||
|
rep.SourceHit++
|
||||||
|
default:
|
||||||
|
rep.SourceConfusion[string(*c.WantSource)+"→"+string(d.Source)]++
|
||||||
|
o.SourceReason = fmt.Sprintf("source %q, want %q", d.Source, *c.WantSource)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
o.Pass = len(o.Reasons) == 0
|
o.Pass = len(o.Reasons) == 0
|
||||||
if o.Pass {
|
if o.Pass {
|
||||||
rep.Passed++
|
rep.Passed++
|
||||||
@@ -298,25 +356,40 @@ func (r Report) String() string {
|
|||||||
r.Name, r.Passed, r.Total, 100*r.Accuracy(), 100*r.IntentAccuracy())
|
r.Name, r.Passed, r.Total, 100*r.Accuracy(), 100*r.IntentAccuracy())
|
||||||
fmt.Fprintf(&b, " clarify: %d false (asked, shouldn't) / %d missed (guessed, shouldn't) | errors: %d | slots deferred to daemon: %d\n",
|
fmt.Fprintf(&b, " clarify: %d false (asked, shouldn't) / %d missed (guessed, shouldn't) | errors: %d | slots deferred to daemon: %d\n",
|
||||||
r.FalseClarify, r.MissedClarify, r.Errors, r.SlotsDeferred)
|
r.FalseClarify, r.MissedClarify, r.Errors, r.SlotsDeferred)
|
||||||
|
if r.SourceTotal > 0 {
|
||||||
|
fmt.Fprintf(&b, " destination: %d/%d labelled cases (%.1f%%)\n",
|
||||||
|
r.SourceHit, r.SourceTotal, 100*r.SourceAccuracy())
|
||||||
|
}
|
||||||
fmt.Fprintf(&b, " latency: p50 %s p95 %s max %s\n", r.P50, r.P95, r.Max)
|
fmt.Fprintf(&b, " latency: p50 %s p95 %s max %s\n", r.P50, r.P95, r.Max)
|
||||||
fmt.Fprintf(&b, " by lang: %s\n", renderStats(r.ByLang))
|
fmt.Fprintf(&b, " by lang: %s\n", renderStats(r.ByLang))
|
||||||
fmt.Fprintf(&b, " by tag: %s\n", renderStats(r.ByTag))
|
fmt.Fprintf(&b, " by tag: %s\n", renderStats(r.ByTag))
|
||||||
if len(r.Confusion) > 0 {
|
if len(r.Confusion) > 0 {
|
||||||
fmt.Fprintf(&b, " confusion: %s\n", renderCounts(r.Confusion))
|
fmt.Fprintf(&b, " confusion: %s\n", renderCounts(r.Confusion))
|
||||||
}
|
}
|
||||||
|
if len(r.SourceConfusion) > 0 {
|
||||||
|
fmt.Fprintf(&b, " destination confusion: %s\n", renderCounts(r.SourceConfusion))
|
||||||
|
}
|
||||||
return b.String()
|
return b.String()
|
||||||
}
|
}
|
||||||
|
|
||||||
// Failures — the per-case detail, sorted by ID so two runs diff cleanly.
|
// Failures — the per-case detail, sorted by ID so two runs diff cleanly. A case
|
||||||
|
// that landed its intent and missed its destination is listed too, marked, so
|
||||||
|
// the half that moved is readable without diffing two percentages.
|
||||||
func (r Report) Failures() string {
|
func (r Report) Failures() string {
|
||||||
var b strings.Builder
|
var b strings.Builder
|
||||||
out := append([]Outcome(nil), r.Outcomes...)
|
out := append([]Outcome(nil), r.Outcomes...)
|
||||||
sort.Slice(out, func(i, j int) bool { return out[i].Case.ID < out[j].Case.ID })
|
sort.Slice(out, func(i, j int) bool { return out[i].Case.ID < out[j].Case.ID })
|
||||||
for _, o := range out {
|
for _, o := range out {
|
||||||
if o.Pass {
|
switch {
|
||||||
continue
|
case !o.Pass:
|
||||||
|
reasons := o.Reasons
|
||||||
|
if o.SourceReason != "" {
|
||||||
|
reasons = append(append([]string(nil), reasons...), o.SourceReason)
|
||||||
|
}
|
||||||
|
fmt.Fprintf(&b, " %s %q: %s\n", o.Case.ID, o.Case.Utterance, strings.Join(reasons, "; "))
|
||||||
|
case o.SourceReason != "":
|
||||||
|
fmt.Fprintf(&b, " %s %q: route ok, %s\n", o.Case.ID, o.Case.Utterance, o.SourceReason)
|
||||||
}
|
}
|
||||||
fmt.Fprintf(&b, " %s %q: %s\n", o.Case.ID, o.Case.Utterance, strings.Join(o.Reasons, "; "))
|
|
||||||
}
|
}
|
||||||
return b.String()
|
return b.String()
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -266,6 +266,11 @@ func baselineGrammars(acts router.ActMatcher) []router.Grammar {
|
|||||||
// Same order as buildRouter (voicewire.go). The fixture is only worth
|
// Same order as buildRouter (voicewire.go). The fixture is only worth
|
||||||
// anything while its grammar set is the daemon's grammar set.
|
// anything while its grammar set is the daemon's grammar set.
|
||||||
grammars = append(grammars, router.AgendaQueryGrammars()...)
|
grammars = append(grammars, router.AgendaQueryGrammars()...)
|
||||||
|
// After the agenda rules and before the feed and list rules, same as
|
||||||
|
// voicewire.go: "что такое лента" is a definition question and the feed
|
||||||
|
// rule would claim it on the noun alone (V-655). Missing here until V-659,
|
||||||
|
// so the fixture was scoring a grammar set the daemon does not run.
|
||||||
|
grammars = append(grammars, router.WorldQueryGrammars()...)
|
||||||
grammars = append(grammars, router.FeedQueryGrammar())
|
grammars = append(grammars, router.FeedQueryGrammar())
|
||||||
// The list side of the same exposure: a phrasing with no possessive in it
|
// The list side of the same exposure: a phrasing with no possessive in it
|
||||||
// ("список дел") routed system and never reached queryTasks (Vikunja #467).
|
// ("список дел") routed system and never reached queryTasks (Vikunja #467).
|
||||||
|
|||||||
@@ -6,37 +6,41 @@
|
|||||||
"Held-out routing contract. Every utterance here is absent from models/seeds/*.txt (TestFixtureIsHeldOut enforces it verbatim) — scoring a classifier on its own seed phrases measures memorisation, not routing.",
|
"Held-out routing contract. Every utterance here is absent from models/seeds/*.txt (TestFixtureIsHeldOut enforces it verbatim) — scoring a classifier on its own seed phrases measures memorisation, not routing.",
|
||||||
"This is a CONTRACT, not a snapshot of current behaviour. Cases the classifier cascade fails today are expected to stay in the file and fail loudly; that failure count is the number Vikunja #319 compares against the LLM router before #320 flips the default.",
|
"This is a CONTRACT, not a snapshot of current behaviour. Cases the classifier cascade fails today are expected to stay in the file and fail loudly; that failure count is the number Vikunja #319 compares against the LLM router before #320 flips the default.",
|
||||||
"Slot expectations are deployment-independent on purpose. want_fn is a boolean (the act must resolve to SOME allowlisted fn) because the allowlist lives in deploy config, not here. want_fact_key names the loop's rule keys (water/meal/sleep/break/shower) — a fact that lands under the wrong key silently starves the predicate that reads it.",
|
"Slot expectations are deployment-independent on purpose. want_fn is a boolean (the act must resolve to SOME allowlisted fn) because the allowlist lives in deploy config, not here. want_fact_key names the loop's rule keys (water/meal/sleep/break/shower) — a fact that lands under the wrong key silently starves the predicate that reads it.",
|
||||||
"want_clarify cases carry intent \"\": the contract is that the router refuses rather than guesses. A confident answer there is a worse failure than a miss."
|
"want_clarify cases carry intent \"\": the contract is that the router refuses rather than guesses. A confident answer there is a worse failure than a miss.",
|
||||||
|
"want_source is the second half of a route (V-655). It is present only on query cases, because no other intent reaches queryWalk, and absent there means absent rather than SourceUnknown. Empty is a label and not a gap: it asserts that the decider must name nothing and let the daemon walk the whole chain in order, his data first.",
|
||||||
|
"Seven cases assert that floor and six of them are homelab operations. They cluster because SourceRecall, SourceNetwork and SourceAttention overlap on every question about the box: mavpoll writes its observations into the fact store recall reads. That is a finding about the enum, not a gap in the labelling.",
|
||||||
|
"ru-query-020 and ru-query-024 are the same utterance, as are ru-query-021 and ru-query-025. Both pairs differ in tags and note only, so both pairs are counted twice in every number this fixture reports.",
|
||||||
|
"Every want_source is the destination that SHOULD claim the turn, which on ru-query-026 through 030 is not the one that did. Those five were observed failing on the box on 2026-08-07 (docs/evals/2026-08-07-week-of-usage.md). A fixture that passes on the day it is written measures nothing."
|
||||||
],
|
],
|
||||||
"cases": [
|
"cases": [
|
||||||
{ "id": "ru-query-001", "utterance": "сколько воды я выпил с утра", "lang": "ru", "intent": "query", "tags": ["aggregate"] },
|
{ "id": "ru-query-001", "utterance": "сколько воды я выпил с утра", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["aggregate"] },
|
||||||
{ "id": "ru-query-002", "utterance": "я сегодня вообще пил воду", "lang": "ru", "intent": "query", "tags": ["hard", "fact-shaped"], "note": "past-tense fact lexicon in a question — the classifier's fact centroid pulls this hard" },
|
{ "id": "ru-query-002", "utterance": "я сегодня вообще пил воду", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["hard", "fact-shaped"], "note": "past-tense fact lexicon in a question — the classifier's fact centroid pulls this hard" },
|
||||||
{ "id": "ru-query-003", "utterance": "во сколько я лёг вчера", "lang": "ru", "intent": "query", "tags": ["temporal"] },
|
{ "id": "ru-query-003", "utterance": "во сколько я лёг вчера", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["temporal"] },
|
||||||
{ "id": "ru-query-004", "utterance": "давно я не тренировался", "lang": "ru", "intent": "query", "tags": ["hard", "no-question-word"] },
|
{ "id": "ru-query-004", "utterance": "давно я не тренировался", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["hard", "no-question-word"] },
|
||||||
{ "id": "ru-query-005", "utterance": "напоминания на завтра есть", "lang": "ru", "intent": "query", "tags": ["hard", "reminder-shaped"], "note": "asks about reminders, does not create one" },
|
{ "id": "ru-query-005", "utterance": "напоминания на завтра есть", "lang": "ru", "intent": "query", "want_source": "", "tags": ["hard", "reminder-shaped"], "note": "asks about reminders, does not create one. want_source is the floor on purpose: no query source reads the reminder store, and day-plan is SourceCalendar over a table this box does not write." },
|
||||||
{ "id": "ru-query-006", "utterance": "что я записывал про кота", "lang": "ru", "intent": "query", "tags": ["recall"] },
|
{ "id": "ru-query-006", "utterance": "что я записывал про кота", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["recall"] },
|
||||||
{ "id": "ru-query-007", "utterance": "сколько раз я ел вчера", "lang": "ru", "intent": "query", "tags": ["aggregate", "hard"] },
|
{ "id": "ru-query-007", "utterance": "сколько раз я ел вчера", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["aggregate", "hard"] },
|
||||||
{ "id": "ru-query-008", "utterance": "мой вес за последний месяц", "lang": "ru", "intent": "query", "tags": ["no-verb"] },
|
{ "id": "ru-query-008", "utterance": "мой вес за последний месяц", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["no-verb"] },
|
||||||
{ "id": "ru-query-009", "utterance": "когда я в последний раз принимал витамины", "lang": "ru", "intent": "query", "tags": ["temporal"] },
|
{ "id": "ru-query-009", "utterance": "когда я в последний раз принимал витамины", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["temporal"] },
|
||||||
{ "id": "ru-query-010", "utterance": "есть новости по бэкапу базы", "lang": "ru", "intent": "query", "tags": ["homelab"] },
|
{ "id": "ru-query-010", "utterance": "есть новости по бэкапу базы", "lang": "ru", "intent": "query", "want_source": "", "tags": ["homelab"], "note": "recall, attention and network can each answer it, because mavpoll writes its netdata and uptime-kuma observations into the fact store recall reads. Naming one takes the other two off the turn." },
|
||||||
{ "id": "ru-query-011", "utterance": "почему сервер тормозит", "lang": "ru", "intent": "query", "tags": ["homelab", "hard"], "note": "diagnostic question, not a chat opener" },
|
{ "id": "ru-query-011", "utterance": "почему сервер тормозит", "lang": "ru", "intent": "query", "want_source": "", "tags": ["homelab", "hard"], "note": "diagnostic question, not a chat opener. network holds the box and attention holds the alarm about the box. The utterance does not choose, so neither does the label." },
|
||||||
{ "id": "ru-query-012", "utterance": "какие заметки я оставил про полив", "lang": "ru", "intent": "query", "tags": ["recall"] },
|
{ "id": "ru-query-012", "utterance": "какие заметки я оставил про полив", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["recall"] },
|
||||||
{ "id": "ru-query-013", "utterance": "во сколько у меня встреча", "lang": "ru", "intent": "query", "tags": ["calendar"] },
|
{ "id": "ru-query-013", "utterance": "во сколько у меня встреча", "lang": "ru", "intent": "query", "want_source": "calendar", "tags": ["calendar"] },
|
||||||
{ "id": "ru-query-019", "utterance": "что у меня стоит в календаре на послезавтра", "lang": "ru", "intent": "query", "tags": ["calendar", "hard"], "note": "agenda, not the clock: the daemon answers this from CalendarEvents inside the query branch, so the clock/date system rule must not swallow it" },
|
{ "id": "ru-query-019", "utterance": "что у меня стоит в календаре на послезавтра", "lang": "ru", "intent": "query", "want_source": "calendar", "tags": ["calendar", "hard"], "note": "agenda, not the clock: the daemon answers this from CalendarEvents inside the query branch, so the clock/date system rule must not swallow it" },
|
||||||
{ "id": "ru-query-022", "utterance": "какие планы на завтра?", "lang": "ru", "intent": "query", "tags": ["calendar"], "note": "the same agenda question as ru-query-019 aimed at another day; it answered \u043f\u043e\u043a\u0430 \u043d\u0435 \u0443\u043c\u0435\u044e on the deployed daemon while the today form worked (Vikunja #471)" },
|
{ "id": "ru-query-022", "utterance": "какие планы на завтра?", "lang": "ru", "intent": "query", "want_source": "calendar", "tags": ["calendar"], "note": "the same agenda question as ru-query-019 aimed at another day; it answered \u043f\u043e\u043a\u0430 \u043d\u0435 \u0443\u043c\u0435\u044e on the deployed daemon while the today form worked (Vikunja #471)" },
|
||||||
{ "id": "ru-query-023", "utterance": "\u043a\u043e\u0433\u0434\u0430 \u043f\u043b\u0430\u043d\u0451\u0440\u043a\u0430?", "lang": "ru", "intent": "query", "tags": ["calendar", "hard"], "note": "a named event with no calendar word — the noun is the only signal that this is a question about his day" },
|
{ "id": "ru-query-023", "utterance": "\u043a\u043e\u0433\u0434\u0430 \u043f\u043b\u0430\u043d\u0451\u0440\u043a\u0430?", "lang": "ru", "intent": "query", "want_source": "calendar", "tags": ["calendar", "hard"], "note": "a named event with no calendar word — the noun is the only signal that this is a question about his day" },
|
||||||
{ "id": "ru-query-024", "utterance": "что дальше?", "lang": "ru", "intent": "query", "tags": ["calendar", "no-question-word"], "note": "the rest of the day, with no possessive and no plan word to anchor on; the model called it a fact and the write had to be caught downstream (Vikunja #498)" },
|
{ "id": "ru-query-024", "utterance": "что дальше?", "lang": "ru", "intent": "query", "want_source": "calendar", "tags": ["calendar", "no-question-word"], "note": "the rest of the day, with no possessive and no plan word to anchor on; the model called it a fact and the write had to be caught downstream (Vikunja #498)" },
|
||||||
{ "id": "ru-query-025", "utterance": "расскажи про битву при Ватерлоо", "lang": "ru", "intent": "query", "tags": ["world", "no-question-word"], "note": "a narrative request carries no question mark and no interrogative, so it routed fact; contrast ru-chat-003, where the same verb asks for a joke" },
|
{ "id": "ru-query-025", "utterance": "расскажи про битву при Ватерлоо", "lang": "ru", "intent": "query", "want_source": "world", "tags": ["world", "no-question-word"], "note": "a narrative request carries no question mark and no interrogative, so it routed fact; contrast ru-chat-003, where the same verb asks for a joke" },
|
||||||
{ "id": "ru-query-014", "utterance": "я успеваю до дедлайна", "lang": "ru", "intent": "query", "tags": ["hard", "no-question-word"] },
|
{ "id": "ru-query-014", "utterance": "я успеваю до дедлайна", "lang": "ru", "intent": "query", "want_source": "", "tags": ["hard", "no-question-word"], "note": "a deadline lives in the task list, the calendar or Praxis depending on where he put it. The destination depends on his data, not on his words." },
|
||||||
{ "id": "ru-query-015", "utterance": "сколько я прошёл шагов", "lang": "ru", "intent": "query", "tags": ["aggregate"] },
|
{ "id": "ru-query-015", "utterance": "сколько я прошёл шагов", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["aggregate"] },
|
||||||
{ "id": "ru-query-016", "utterance": "покажи давление за неделю", "lang": "ru", "intent": "query", "tags": ["hard", "imperative"], "note": "imperative form but a read — must not route to act" },
|
{ "id": "ru-query-016", "utterance": "покажи давление за неделю", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["hard", "imperative"], "note": "imperative form but a read — must not route to act" },
|
||||||
{ "id": "ru-query-017", "utterance": "чем я занимался в среду", "lang": "ru", "intent": "query", "tags": ["hard", "chat-shaped"] },
|
{ "id": "ru-query-017", "utterance": "чем я занимался в среду", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["hard", "chat-shaped"] },
|
||||||
{ "id": "ru-query-018", "utterance": "хватает ли места под новые бэкапы", "lang": "ru", "intent": "query", "tags": ["homelab"] },
|
{ "id": "ru-query-018", "utterance": "хватает ли места под новые бэкапы", "lang": "ru", "intent": "query", "want_source": "", "tags": ["homelab"], "note": "disk headroom. network is the only source that reads the box, but the phrasing is a capacity question and not a LAN one." },
|
||||||
{ "id": "ru-query-020", "utterance": "что дальше?", "lang": "ru", "intent": "query", "tags": ["agenda", "hard"], "note": "the rest of the day, with no interrogative the model can read as a question — it routed fact until a stage 0 rule claimed it (V-498)" },
|
{ "id": "ru-query-020", "utterance": "что дальше?", "lang": "ru", "intent": "query", "want_source": "calendar", "tags": ["agenda", "hard"], "note": "the rest of the day, with no interrogative the model can read as a question — it routed fact until a stage 0 rule claimed it (V-498)" },
|
||||||
{ "id": "ru-query-021", "utterance": "расскажи про битву при Ватерлоо", "lang": "ru", "intent": "query", "tags": ["world", "hard"], "note": "a world question phrased as an instruction. It routed fact, and the fact gate had to catch the write (V-498)" },
|
{ "id": "ru-query-021", "utterance": "расскажи про битву при Ватерлоо", "lang": "ru", "intent": "query", "want_source": "world", "tags": ["world", "hard"], "note": "a world question phrased as an instruction. It routed fact, and the fact gate had to catch the write (V-498)" },
|
||||||
{ "id": "en-query-001", "utterance": "did I take my vitamins today", "lang": "en", "intent": "query", "tags": ["fact-shaped"] },
|
{ "id": "en-query-001", "utterance": "did I take my vitamins today", "lang": "en", "intent": "query", "want_source": "recall", "tags": ["fact-shaped"] },
|
||||||
{ "id": "en-query-002", "utterance": "how long since the last backup finished", "lang": "en", "intent": "query", "tags": ["temporal"] },
|
{ "id": "en-query-002", "utterance": "how long since the last backup finished", "lang": "en", "intent": "query", "want_source": "", "tags": ["temporal"], "note": "the completion time is a fact the poller wrote, so recall answers it. A person asking this wants the operational answer. Both are true." },
|
||||||
{ "id": "en-query-003", "utterance": "show me this week's weight", "lang": "en", "intent": "query", "tags": ["imperative"] },
|
{ "id": "en-query-003", "utterance": "show me this week's weight", "lang": "en", "intent": "query", "want_source": "recall", "tags": ["imperative"] },
|
||||||
|
|
||||||
{ "id": "ru-fact-001", "utterance": "только что выпил кружку воды", "lang": "ru", "intent": "fact", "want_fact_key": "water" },
|
{ "id": "ru-fact-001", "utterance": "только что выпил кружку воды", "lang": "ru", "intent": "fact", "want_fact_key": "water" },
|
||||||
{ "id": "ru-fact-002", "utterance": "воды попил наконец", "lang": "ru", "intent": "fact", "want_fact_key": "water", "tags": ["inverted"] },
|
{ "id": "ru-fact-002", "utterance": "воды попил наконец", "lang": "ru", "intent": "fact", "want_fact_key": "water", "tags": ["inverted"] },
|
||||||
@@ -106,6 +110,11 @@
|
|||||||
{ "id": "amb-005", "utterance": "потом", "lang": "ru", "want_clarify": true, "tags": ["ambiguous", "filler"] },
|
{ "id": "amb-005", "utterance": "потом", "lang": "ru", "want_clarify": true, "tags": ["ambiguous", "filler"] },
|
||||||
{ "id": "amb-006", "utterance": "the thing from earlier", "lang": "en", "want_clarify": true, "tags": ["ambiguous", "anaphora"] },
|
{ "id": "amb-006", "utterance": "the thing from earlier", "lang": "en", "want_clarify": true, "tags": ["ambiguous", "anaphora"] },
|
||||||
{ "id": "amb-007", "utterance": "напомни", "lang": "ru", "want_clarify": true, "tags": ["ambiguous", "reminder"], "note": "the reminder verb and nothing else — she knows the shape of the request and not one thing about it. Answered 'не получилось разобрать время напоминания' on the box until V-548: the subjectless-reminder gate tested Slots.Text == \"\", and fillSlots had put the verb in that slot" },
|
{ "id": "amb-007", "utterance": "напомни", "lang": "ru", "want_clarify": true, "tags": ["ambiguous", "reminder"], "note": "the reminder verb and nothing else — she knows the shape of the request and not one thing about it. Answered 'не получилось разобрать время напоминания' on the box until V-548: the subjectless-reminder gate tested Slots.Text == \"\", and fillSlots had put the verb in that slot" },
|
||||||
{ "id": "amb-008", "utterance": "ну напомни же", "lang": "ru", "want_clarify": true, "tags": ["ambiguous", "reminder", "filler"], "note": "the same request wrapped in particles, which is why filler_particles is a lexicon set — without it the particles read as the subject" }
|
{ "id": "amb-008", "utterance": "ну напомни же", "lang": "ru", "want_clarify": true, "tags": ["ambiguous", "reminder", "filler"], "note": "the same request wrapped in particles, which is why filler_particles is a lexicon set — without it the particles read as the subject" },
|
||||||
|
{ "id": "ru-query-026", "utterance": "что такое TCP?", "lang": "ru", "intent": "query", "want_source": "world", "tags": ["world", "regression"], "note": "weather claimed it on 2026-08-07 and answered \"для какого города?\", because it read one percent closer than the leftover seeds. WorldQueryGrammars claims it at stage 0 now." },
|
||||||
|
{ "id": "ru-query-027", "utterance": "сколько будет 17 на 23?", "lang": "ru", "intent": "query", "want_source": "world", "tags": ["world", "arithmetic", "regression"], "note": "same source, same day, same answer about a city. Arithmetic is not a place." },
|
||||||
|
{ "id": "ru-query-028", "utterance": "какой у меня любимый язык?", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["recall", "possessive", "regression"], "note": "the feed answered it with kernel headlines. \"у меня\" is the whole signal and it points inward." },
|
||||||
|
{ "id": "ru-query-029", "utterance": "кто такой Линус Торвальдс?", "lang": "ru", "intent": "query", "want_source": "world", "tags": ["world", "person", "regression"], "note": "the personal boundary answered \"не нашла у тебя такой записи\". A named public person is not his data." },
|
||||||
|
{ "id": "ru-query-030", "utterance": "что там с бэкапами?", "lang": "ru", "intent": "query", "want_source": "", "tags": ["homelab", "regression"], "note": "search claimed it, which inverts the boundary outward. The fix is the chain order and not a destination: recall, attention and network all answer it, same as ru-query-010." }
|
||||||
]
|
]
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -45,12 +45,19 @@ const routeGrammar = `
|
|||||||
root ::= "[" ws action ("," ws action)* ws "]"
|
root ::= "[" ws action ("," ws action)* ws "]"
|
||||||
action ::= "{" ws "\"intent\"" ws ":" ws intent ("," ws field)* ws "}"
|
action ::= "{" ws "\"intent\"" ws ":" ws intent ("," ws field)* ws "}"
|
||||||
intent ::= "\"fact\"" | "\"reminder\"" | "\"note\"" | "\"query\"" | "\"act\"" | "\"chat\"" | "\"system\"" | "\"unknown\""
|
intent ::= "\"fact\"" | "\"reminder\"" | "\"note\"" | "\"query\"" | "\"act\"" | "\"chat\"" | "\"system\"" | "\"unknown\""
|
||||||
field ::= key ws ":" ws string
|
field ::= (key ws ":" ws string) | ("\"source\"" ws ":" ws source)
|
||||||
key ::= "\"key\"" | "\"value\"" | "\"text\"" | "\"verb\""
|
key ::= "\"key\"" | "\"value\"" | "\"text\"" | "\"verb\""
|
||||||
|
source ::= "\"recall\"" | "\"calendar\"" | "\"tasks\"" | "\"list\"" | "\"money\"" | "\"weather\"" | "\"home\"" | "\"network\"" | "\"feeds\"" | "\"attention\"" | "\"self\"" | "\"world\"" | "\"\""
|
||||||
string ::= "\"" ([^"\\\x00-\x1F] | "\\" ["\\/bfnrt] | "\\u" [0-9a-fA-F]{4}){0,120} "\""
|
string ::= "\"" ([^"\\\x00-\x1F] | "\\" ["\\/bfnrt] | "\\u" [0-9a-fA-F]{4}){0,120} "\""
|
||||||
ws ::= [ \t\n]{0,4}
|
ws ::= [ \t\n]{0,4}
|
||||||
`
|
`
|
||||||
|
|
||||||
|
// TestRouteGrammarCoversSources holds the source rule above to router.Sources.
|
||||||
|
// The enum is the point: a grammar cannot emit a destination that does not
|
||||||
|
// exist, which is the guarantee V-546 wants from a softmax and gets here for
|
||||||
|
// free. Empty is the thirteenth alternative and it is not an oversight — it is
|
||||||
|
// the SourceUnknown floor, and the model must be able to decline.
|
||||||
|
|
||||||
// routeSystem — the router prompt. Changed 31-07-2026: the query test now sits
|
// routeSystem — the router prompt. Changed 31-07-2026: the query test now sits
|
||||||
// above the fact test and there is an explicit question test. Before that, a
|
// above the fact test and there is an explicit question test. Before that, a
|
||||||
// question naming a fact key ("сколько воды я выпил с утра") matched the fact
|
// question naming a fact key ("сколько воды я выпил с утра") matched the fact
|
||||||
@@ -121,6 +128,31 @@ const routeSystem = `Классифицируй ровно одно сообще
|
|||||||
"что такое кватернион?" → {"intent":"query","text":"что такое кватернион"}
|
"что такое кватернион?" → {"intent":"query","text":"что такое кватернион"}
|
||||||
"ага" → {"intent":"chat","text":"ага"}
|
"ага" → {"intent":"chat","text":"ага"}
|
||||||
|
|
||||||
|
Только для query добавь поле source — где лежит ответ:
|
||||||
|
- recall — его заметки, факты и то, что он раньше говорил
|
||||||
|
- calendar — встречи и события
|
||||||
|
- tasks — список задач
|
||||||
|
- list — списки покупок и другие именованные списки
|
||||||
|
- money — траты
|
||||||
|
- weather — погода
|
||||||
|
- home — свет, устройства, дом
|
||||||
|
- network — локальная сеть, сервер, диски
|
||||||
|
- feeds — новостные ленты
|
||||||
|
- attention — что требует внимания сейчас
|
||||||
|
- self — вопрос про самого ассистента
|
||||||
|
- world — всё остальное: определения, счёт, люди, факты о мире
|
||||||
|
|
||||||
|
Пустое значение "" — нормальный ответ и его надо ставить часто. Ставь "", если ответ могут дать сразу два источника или если не уверен: тогда проверяются все по порядку, и это правильно. Никогда не угадывай.
|
||||||
|
|
||||||
|
"сколько воды я выпил с утра" → {"intent":"query","text":"сколько воды я выпил с утра","source":"recall"}
|
||||||
|
"что я записывал про кота" → {"intent":"query","text":"что я записывал про кота","source":"recall"}
|
||||||
|
"во сколько у меня встреча" → {"intent":"query","text":"во сколько у меня встреча","source":"calendar"}
|
||||||
|
"что такое docker?" → {"intent":"query","text":"что такое docker","source":"world"}
|
||||||
|
"кто такой Линус Торвальдс?" → {"intent":"query","text":"кто такой Линус Торвальдс","source":"world"}
|
||||||
|
"сколько будет 17 на 23?" → {"intent":"query","text":"сколько будет 17 на 23","source":"world"}
|
||||||
|
"почему сервер тормозит" → {"intent":"query","text":"почему сервер тормозит","source":""}
|
||||||
|
"есть новости по бэкапу базы" → {"intent":"query","text":"есть новости по бэкапу базы","source":""}
|
||||||
|
|
||||||
Ответ — JSON-массив: по одному объекту на каждую просьбу. Обычно один. Если в реплике несколько просьб — по объекту на каждую. "напомни купить молоко, и запиши что кофе кончился" → [{"intent":"reminder","text":"купить молоко"},{"intent":"note","text":"кофе кончился"}]. Только JSON, без пояснений.`
|
Ответ — JSON-массив: по одному объекту на каждую просьбу. Обычно один. Если в реплике несколько просьб — по объекту на каждую. "напомни купить молоко, и запиши что кофе кончился" → [{"intent":"reminder","text":"купить молоко"},{"intent":"note","text":"кофе кончился"}]. Только JSON, без пояснений.`
|
||||||
|
|
||||||
// routeRepeatPenalty — the sub-1B model loops one sentence inside the text field
|
// routeRepeatPenalty — the sub-1B model loops one sentence inside the text field
|
||||||
@@ -172,6 +204,7 @@ type routeAction struct {
|
|||||||
Value string `json:"value"`
|
Value string `json:"value"`
|
||||||
Text string `json:"text"`
|
Text string `json:"text"`
|
||||||
Verb string `json:"verb"`
|
Verb string `json:"verb"`
|
||||||
|
Source string `json:"source"`
|
||||||
}
|
}
|
||||||
|
|
||||||
// Route asks the model for one decision. The bool is false when there is no
|
// Route asks the model for one decision. The bool is false when there is no
|
||||||
@@ -240,6 +273,14 @@ func (lr *LLMRouter) Route(ctx context.Context, utterance string, now time.Time)
|
|||||||
case IntentQuery:
|
case IntentQuery:
|
||||||
d.Intent = IntentQuery
|
d.Intent = IntentQuery
|
||||||
d.Slots.Text = firstNonEmpty(a.Text, utterance)
|
d.Slots.Text = firstNonEmpty(a.Text, utterance)
|
||||||
|
// Through ValidSource, and on query alone. The grammar already bounds
|
||||||
|
// the enum, but the grammar is a request to a server that may be
|
||||||
|
// running a different build, and a destination this binary does not
|
||||||
|
// know would take real query sources off the turn. Anything unknown
|
||||||
|
// drops to SourceUnknown, which is the floor and costs nothing.
|
||||||
|
if ValidSource(Source(a.Source)) {
|
||||||
|
d.Source = Source(a.Source)
|
||||||
|
}
|
||||||
case IntentAct:
|
case IntentAct:
|
||||||
d.Intent = IntentAct
|
d.Intent = IntentAct
|
||||||
d.Slots.Text = firstNonEmpty(a.Verb, utterance)
|
d.Slots.Text = firstNonEmpty(a.Verb, utterance)
|
||||||
|
|||||||
@@ -398,3 +398,70 @@ func TestLLMReminderWithSubjectIsNotGated(t *testing.T) {
|
|||||||
t.Fatalf("a complete reminder was sent back as a question: %+v", d.Slots)
|
t.Fatalf("a complete reminder was sent back as a question: %+v", d.Slots)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// TestRouteGrammarCoversSources — the grammar enum and router.Sources are two
|
||||||
|
// hand-written lists of the same twelve destinations, and nothing else notices
|
||||||
|
// when one grows. A destination missing from the grammar is a destination the
|
||||||
|
// model is structurally unable to name, which is the exact defect V-517
|
||||||
|
// measured for Praxis: not a weak model, an absent string.
|
||||||
|
func TestRouteGrammarCoversSources(t *testing.T) {
|
||||||
|
for _, s := range Sources {
|
||||||
|
if !strings.Contains(routeGrammar, `"\"`+string(s)+`\""`) {
|
||||||
|
t.Errorf("routeGrammar cannot emit %q — the model can never name it", s)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
// The floor has to be reachable too, or the model is forced to pick one.
|
||||||
|
if !strings.Contains(routeGrammar, `"\"\""`) {
|
||||||
|
t.Error(`routeGrammar cannot emit "" — the model cannot decline a destination`)
|
||||||
|
}
|
||||||
|
// Count the alternatives on the source rule: an extra one is a destination
|
||||||
|
// the daemon would drop to SourceUnknown after the model spent tokens on it.
|
||||||
|
for _, line := range strings.Split(routeGrammar, "\n") {
|
||||||
|
if !strings.HasPrefix(line, "source ") {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if got, want := strings.Count(line, "|")+1, len(Sources)+1; got != want {
|
||||||
|
t.Errorf("source rule has %d alternatives, want %d (Sources plus the floor)", got, want)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The destination is read back only through ValidSource. A model on an older or
|
||||||
|
// newer build can write a string this binary does not know, and trusting it
|
||||||
|
// would take real query sources off the turn for a name nothing answers.
|
||||||
|
func TestLLMUnknownSourceFallsToTheFloor(t *testing.T) {
|
||||||
|
r := newLLMTestRouter(t, `{"intent":"query","text":"что там с бэкапами","source":"praxis"}`)
|
||||||
|
d, err := r.Route(context.Background(), "что там с бэкапами", refNow())
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("route: %v", err)
|
||||||
|
}
|
||||||
|
if d.Source != SourceUnknown {
|
||||||
|
t.Fatalf("invented destination %q was trusted, want the floor", d.Source)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// And a known one survives, or the read-back is just a filter.
|
||||||
|
func TestLLMNamedSourceSurvives(t *testing.T) {
|
||||||
|
r := newLLMTestRouter(t, `{"intent":"query","text":"кто такой Линус Торвальдс","source":"world"}`)
|
||||||
|
d, err := r.Route(context.Background(), "кто такой Линус Торвальдс?", refNow())
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("route: %v", err)
|
||||||
|
}
|
||||||
|
if d.Source != SourceWorld {
|
||||||
|
t.Fatalf("source %q, want %q", d.Source, SourceWorld)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// A destination on anything but a query is dropped. Only IntentQuery reaches
|
||||||
|
// queryWalk, so a source elsewhere is a field nobody reads and a claim nobody
|
||||||
|
// checks.
|
||||||
|
func TestLLMSourceIsQueryOnly(t *testing.T) {
|
||||||
|
r := newLLMTestRouter(t, `{"intent":"note","text":"кофе кончился","source":"recall"}`)
|
||||||
|
d, err := r.Route(context.Background(), "запиши что кофе кончился", refNow())
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("route: %v", err)
|
||||||
|
}
|
||||||
|
if d.Source != SourceUnknown {
|
||||||
|
t.Fatalf("a note carried destination %q", d.Source)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|||||||
Reference in New Issue
Block a user