Record the destination number and two stuck measurements (V-659)
CLAUDE.md said the destination had no fixture and no accuracy number. It has both now: intent 73/96 and destination 12/33 on the classifier cascade, with the per-destination split, the floor cases and the grammar drift the labelling turned up. Anyone adding a grammar now reads that baselineGrammars mirrors buildRouter and drifts silently when it does not. docs/evals/2026-08-08-massive-warm-start.md was written on the V-655 branch and parked in .task/, which git excludes, so it was one `task start` away from being lost. It is a dated measurement and it belongs under docs/evals whatever branch produced it. Its "destination has no fixture at all" line is now a pointer to the file beside it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
This commit is contained in:
@@ -398,11 +398,44 @@ deliberately do not. "что у меня в списке покупок" matches
|
||||
the calendar there would take the list source off the turn.
|
||||
|
||||
Fixture unchanged at **69/91 classifier+ONNX**, measured both sides. That is the
|
||||
expected result, because it scores intent and no case here changes intent. **The
|
||||
destination has no fixture yet, so it has no accuracy number.** That and the model arm
|
||||
are the follow-ups. The field is designed so a decider naming nothing costs nothing.
|
||||
It lands on V-546. Intent, mood and BIO slot tags were already three heads on one
|
||||
forward pass of the resident e5-small. Destination is a fourth head on the same pass.
|
||||
expected result, because it scores intent and no case here changes intent.
|
||||
|
||||
**The destination has its own fixture and its own number as of 08-08-2026**
|
||||
(V-659, `docs/evals/2026-08-08-destination-fixture.md`). This section used to say
|
||||
it had neither. `want_source` on `eval.Case` is a pointer, because the destination
|
||||
has three states and a bare string has two. Absent is every intent but query,
|
||||
which never reaches `queryWalk`. Present and empty is the `SourceUnknown`
|
||||
contract: name nothing and walk the chain. Present and named is a destination the
|
||||
route must produce. Thirty-three of ninety-six cases carry one.
|
||||
|
||||
A destination miss does **not** fail the case. It lands in `Outcome.SourceReason`
|
||||
and never in `Reasons`, so `Accuracy` and `IntentAccuracy` mean what they meant
|
||||
and `SourceAccuracy` is a second number over the labelled cases only. Intent and
|
||||
destination are two decisions, and one number hides which one moved. A route that
|
||||
lost its intent scores no destination hit, or a clarify would satisfy an empty
|
||||
label for free.
|
||||
|
||||
Measured classifier+ONNX: intent **73/96 (76.0%)**, destination **12/33 (36.4%)**.
|
||||
The split is the finding. World is 5/5, because a stage 0 rule names it. The
|
||||
`SourceUnknown` floor is 5/7. Calendar is 2/6, because the possessive agenda
|
||||
rules deliberately do not name it. And **recall is 0/15, because nothing
|
||||
anywhere names it**. Those turns are still answered, since the chain walks
|
||||
recall early. Recall is the number the fourth head has to move.
|
||||
|
||||
Seven cases assert the floor and six of them are homelab operations. They
|
||||
cluster because `SourceRecall`, `SourceNetwork` and `SourceAttention` overlap on
|
||||
every question about the box. `mavpoll` writes its netdata and uptime-kuma
|
||||
observations into the fact store recall reads. That is a finding about the enum,
|
||||
not a gap in the labelling.
|
||||
|
||||
`baselineGrammars` in `eval_test.go` mirrors `buildRouter` and had drifted:
|
||||
`WorldQueryGrammars` was wired into the daemon by V-655 and not into the mirror,
|
||||
so the fixture scored a grammar set nobody runs. Fixed by V-659, worth 3 points of
|
||||
destination and nothing else. Check that function when adding a grammar.
|
||||
|
||||
The model arm is still the follow-up. It lands on V-546. Intent, mood and BIO slot
|
||||
tags were already three heads on one forward pass of the resident e5-small.
|
||||
Destination is a fourth head on the same pass.
|
||||
|
||||
## LLM output contract
|
||||
|
||||
|
||||
@@ -0,0 +1,82 @@
|
||||
# The first destination number
|
||||
|
||||
Measured 2026-08-08 on the classifier cascade with the ONNX multilingual
|
||||
embedder, the configuration homesrv runs. `make t PKG=./internal/router/eval/
|
||||
RUN=TestONNXBaseline V=1`. Covers V-659, the follow-up V-655 named.
|
||||
|
||||
## What was measured
|
||||
|
||||
V-655 split a routing decision in two. The cascade sorts an utterance into one
|
||||
of seven intents, and `Decision.Source` then says where the answer lives. The
|
||||
first half had a fixture. The second half arrived with none, so twelve
|
||||
destinations shipped with no accuracy number.
|
||||
|
||||
`want_source` is now a field on `eval.Case`. It is a pointer, because the
|
||||
destination has three states and a bare string has two. Absent is every intent
|
||||
but query, which never reaches `queryWalk`. Present and empty is the
|
||||
`SourceUnknown` contract: name nothing and let the daemon walk the chain.
|
||||
Present and named is a destination the route must produce.
|
||||
|
||||
Thirty-three of the ninety-six cases carry one. A destination miss does not
|
||||
fail the case, so `Accuracy` and `IntentAccuracy` mean what they meant.
|
||||
`SourceAccuracy` is a second number over the labelled cases only.
|
||||
|
||||
## Result
|
||||
|
||||
Intent is **73/96 (76.0%)**, against 69/91 (75.8%) before. Four of the five new
|
||||
cases pass and no existing case moved.
|
||||
|
||||
Destination is **12/33 (36.4%)**, and the split is the whole finding.
|
||||
|
||||
| destination | scored | note |
|
||||
|---|---|---|
|
||||
| world | 5/5 | `WorldQueryGrammars` names it at stage 0 |
|
||||
| the `SourceUnknown` floor | 5/7 | the two misses lost the intent first |
|
||||
| calendar | 2/6 | `calendar-query` names it, the possessive agenda rules do not |
|
||||
| recall | 0/15 | nothing anywhere names it |
|
||||
|
||||
Recall is the number to move. Fifteen cases ask about his own words and his own
|
||||
facts. The route lands `query` on eleven of them and the destination comes back
|
||||
empty every time. Those turns are answered today, because the daemon walks the
|
||||
chain in order and the three recall passes are early in it. What is missing is a
|
||||
decider that says so, and that is the fourth head on V-546.
|
||||
|
||||
Two cases labelled the floor lost their intent before a destination was
|
||||
possible. A clarify names nothing, so it would satisfy an empty label for free.
|
||||
`Score` requires the route to land the case's intent before it credits a
|
||||
destination hit, or the floor label would score itself.
|
||||
|
||||
## Seven cases assert the floor, and six of them cluster
|
||||
|
||||
The six are homelab operations. `SourceRecall`, `SourceNetwork` and
|
||||
`SourceAttention` overlap on every question about the box, because `mavpoll`
|
||||
writes its netdata and uptime-kuma observations into the fact store recall
|
||||
reads. "почему сервер тормозит" is answerable from all three. Naming one takes
|
||||
the other two off the turn.
|
||||
|
||||
That is a finding about the enum rather than a gap in the labelling. The floor
|
||||
is the right answer there and the fixture now says so out loud.
|
||||
|
||||
## A drift the labelling found
|
||||
|
||||
`WorldQueryGrammars` went into `buildRouter` with V-655 and never into
|
||||
`baselineGrammars`, the fixture's mirror of it. So the fixture was scoring a
|
||||
grammar set the daemon does not run. The comment above that function forbids
|
||||
exactly that. Adding it moved the destination number from 9/33 to 12/33 and
|
||||
moved nothing else.
|
||||
|
||||
The three cases it recovered are `что такое TCP?`, `сколько будет 17 на 23?`
|
||||
and `кто такой Линус Торвальдс?`. All three already routed `query` through
|
||||
`NarrativeQueryGrammars`. So the drift was invisible to every number this
|
||||
fixture reported, until the destination had one of its own.
|
||||
|
||||
## What this does not measure
|
||||
|
||||
The model arm. This is the classifier cascade, which names a destination only
|
||||
where a stage 0 rule filled one in. The resident model has no destination in
|
||||
its router prompt yet, so 36.4% is a floor and not a comparison.
|
||||
|
||||
Two pairs of cases are the same utterance. `ru-query-020` and `ru-query-024`
|
||||
are both "что дальше?", and `ru-query-021` and `ru-query-025` are both
|
||||
"расскажи про битву при Ватерлоо". They differ in tags and note only, so both
|
||||
pairs are counted twice here and in every earlier number this fixture reported.
|
||||
@@ -0,0 +1,142 @@
|
||||
# MASSIVE Russian warm-start for the routing heads
|
||||
|
||||
Measured 2026-08-08 on workpc (Radeon RX 7900 GRE, ROCm). Covers V-546 step 2.
|
||||
Workspace is `~/Programs/embed-training` on workpc, scripts `train_massive.py`,
|
||||
`ab_run.py`, `ab.sh`, `probe_time.py`.
|
||||
|
||||
## What was trained
|
||||
|
||||
Two heads on a copy of multilingual-e5-small: `Linear(384, 60)` for MASSIVE's
|
||||
own intents over a masked mean pool, `Linear(384, 111)` per token for BIO slot
|
||||
tags. MASSIVE's label sets verbatim, no alignment to Maven's 7 intents. The
|
||||
intent head is an auxiliary loss that shapes the pooled vector and is thrown
|
||||
away.
|
||||
|
||||
Data is `amazon-massive-dataset-1.1` pulled from S3. The Hugging Face repo is
|
||||
script-only and `datasets` 5.0 refuses those, so `load_dataset` cannot fetch it.
|
||||
`ru-RU` is 11,514 train, 2,033 dev, 2,974 test, 60 intents, 55 slots, 111 BIO
|
||||
labels. All 16,521 rows survived span alignment: `annot_utt` re-tokenised to its
|
||||
own `utt` on every one.
|
||||
|
||||
Hyperparameters match `train_intent.py`, so the two runs differ in data only.
|
||||
Frozen XLM-R vocabulary, body 2e-5, heads 1e-3, batch 32, sequence 64, 10
|
||||
epochs. MASSIVE's own dev partition selects the epoch, on slot F1 with intent
|
||||
accuracy as tiebreak. Selecting on 60-class intent accuracy would optimise a
|
||||
head that gets deleted.
|
||||
|
||||
## Result
|
||||
|
||||
Epoch 9 of 10 by dev slot F1. Held-out MASSIVE test: intent 86.2%, slot span
|
||||
F1 71.5% (P 68.5, R 74.8). Peak 1.70GB of 17.2GB, about 22 seconds an epoch,
|
||||
under 4 minutes end to end. Dev slot F1 climbed monotonically to epoch 9 and
|
||||
fell at 10, so 10 epochs was the right budget.
|
||||
|
||||
Ten slot types sit at 0% test recall. Every one of them has 1 to 7 test
|
||||
instances: `alarm_type` has 3, `drink_type` has 1. That is support in MASSIVE's
|
||||
Russian split, not a tagger failure. `playlist_name` at 6% of 16 is the first
|
||||
real miss.
|
||||
|
||||
## The intent A/B, and why it settles nothing
|
||||
|
||||
`train_intent.py` was run against both bodies, three seeds by two smoothing
|
||||
settings, on `train_v4.jsonl`. It is v4 and not v5 because v4 is what
|
||||
`sweep2.log` measured. `ab_run.py` strips a `--base` flag onto the module global, so
|
||||
`train_intent.py` is unmodified and its baseline stays reproducible. The stock
|
||||
arm reproduced `sweep2.log` line for line.
|
||||
|
||||
Fixture accuracy, 91 cases, one case is 1.1 points:
|
||||
|
||||
| seed / smooth | stock | warm-started |
|
||||
|---|---|---|
|
||||
| 0 / 0.0 | 94.0% | 92.8% |
|
||||
| 0 / 0.1 | 95.2% | 92.8% |
|
||||
| 1 / 0.0 | 95.2% | 94.0% |
|
||||
| 1 / 0.1 | 95.2% | 97.6% |
|
||||
| 2 / 0.0 | 92.8% | 94.0% |
|
||||
| 2 / 0.1 | 92.8% | 96.4% |
|
||||
|
||||
Mean 94.2% against 94.6%. That is +0.4 points, about a third of one case, and
|
||||
inside seed noise. Spread widened. Stock lands in a 2.4-point band and
|
||||
warm-started in a 4.8-point one. The warm-started arm holds both the best result
|
||||
of the sweep and a tie for the worst. Seed 0 is the bad arm and it fails in a
|
||||
specific way. Its dev peaks at epoch 2 and 3 and never improves, where stock
|
||||
peaks around 7. The dev slice is a quarter of the seed rows. That is small
|
||||
enough that early stopping is fragile when the body arrives already fitted.
|
||||
|
||||
**The A/B was never the test.** Intent had at most 4.8 points of headroom here.
|
||||
MASSIVE was not trained for Maven's intents. Read it as "the warm-start does not
|
||||
cost intent accuracy", nothing more.
|
||||
|
||||
## The measurement that does mean something
|
||||
|
||||
`want_time` is the one slot Maven's fixture scores, and MASSIVE has `time` and
|
||||
`date`. Restricted to those two slot types, F1 is 74.9% over 609 gold spans on the
|
||||
MASSIVE ru test split. Precision is 71.5 and recall 78.7. That beats the 71.5%
|
||||
all-slot figure. Of the 530 test utterances carrying a time or a date, 73.4% get
|
||||
every such span exactly right.
|
||||
|
||||
Out of domain matters more, because Maven's traffic is not this corpus. Ten
|
||||
Maven-shaped utterances, none of them in MASSIVE:
|
||||
|
||||
| utterance | tagged |
|
||||
|---|---|
|
||||
| `напомни в 11:00 позвонить маме` | `time='11:00'`, `relation='маме'` |
|
||||
| `напомни завтра в семь утра выпить таблетки` | `date='завтра'`, `time='семь утра'` |
|
||||
| `поставь будильник на полседьмого` | `time='полседьмого'` |
|
||||
| `через двадцать минут напомни про чайник` | `time='двадцать минут'` |
|
||||
| `напомни в пятницу вечером забрать посылку` | `date='пятницу'`, `timeofday='вечером'` |
|
||||
| `что у меня сегодня после обеда` | `date='сегодня'`, `time='после'`, `timeofday='обеда'` |
|
||||
| `запиши что кофе закончился` | nothing |
|
||||
| `что такое TCP` | `definition_word='TCP'` |
|
||||
|
||||
The first row is the V-572 defect utterance. `ReminderGrammar` handed the daemon
|
||||
`HasTime: false` there, and the daemon asked "Когда?" at a sentence that had
|
||||
already said when. `полседьмого` is a colloquial half-past that no digit pattern
|
||||
catches. `запиши что кофе закончился` correctly carries nothing, because a note
|
||||
has no time.
|
||||
|
||||
Two errors. `после обеда` split into `time='после'` plus `timeofday='обеда'`
|
||||
when it is one span, and `через двадцать минут` dropped its `через`. Both are
|
||||
boundary errors on spans the tagger did find.
|
||||
|
||||
Unplanned: `что такое TCP` returned `definition_word='TCP'`. MASSIVE has a slot
|
||||
for the thing being asked about, which is a `SourceWorld` signal sitting in a
|
||||
head already trained.
|
||||
|
||||
Ten hand-picked utterances are evidence, not a fixture.
|
||||
|
||||
## What this does not measure
|
||||
|
||||
Maven has no span fixture. `want_time` and `want_fn` are presence booleans and
|
||||
`want_fact_key` is an exact string match, so nothing in the repo can score a
|
||||
71.5% span tagger. Destination got one the same day, at 12/33 on the classifier
|
||||
cascade: see `2026-08-08-destination-fixture.md`.
|
||||
|
||||
The missing span fixture is why the warm-start stays unjudged against Maven
|
||||
rather than against MASSIVE.
|
||||
|
||||
## Datasets ruled out
|
||||
|
||||
Checked on 2026-08-08 and rejected as label sources:
|
||||
|
||||
- **MASSIVE's other 50 locales** ship in the same tarball and are parallel by id.
|
||||
Co-training on them is free and unmeasured. English was ruled out by the owner
|
||||
on 2026-08-08.
|
||||
- **CLINC150** is reachable as parquet, 150 intents and 1,200 explicit
|
||||
out-of-scope queries, English only. Its value is the labeled out-of-scope set
|
||||
for fitting the energy threshold, not intent labels.
|
||||
- **`d0rj/dolphin-ru`**, roughly 2.8M rows of FLAN-style tasks translated to
|
||||
Russian. No intent, no slots, and not utterances anyone says to an assistant.
|
||||
- **`psytechlab/EmpatheticIntents-ru`**, 24,856 rows of translated
|
||||
EmpatheticDialogues with 32 emotion labels. Maven's mood enum is `neutral,
|
||||
happy, thinking, tired, confused` and it describes her own reply, not the
|
||||
speaker's emotion. No mapping exists.
|
||||
- **`ai-forever/MERA`** and **`RussianNLP/russian_super_glue`**, benchmark
|
||||
harnesses. Rows are prompt templates with `{toxic_comment}` placeholders.
|
||||
- **`ZeroAgency/ru-big-russian-dataset`**, an LLM-judge quality corpus. Its
|
||||
`question` and `classified_topic` columns are a usable Russian out-of-scope
|
||||
pool for threshold fitting. That is the one thing CLINC150 can only supply in
|
||||
English. The questions are long and written, so they belong in the negative
|
||||
set, never in the in-scope `query` training set.
|
||||
- No second Russian slot-filling corpus exists. The xSID mirrors are 404,
|
||||
MultiATIS++ has no Russian, SLURP is not on the Hub.
|
||||
Reference in New Issue
Block a user