Merge remote-tracking branch 'origin/master' into task/486-deploy-the-workstation-transcriber
This commit is contained in:
@@ -0,0 +1,146 @@
|
||||
# The routing heads, running in Go
|
||||
|
||||
Date: 2026-08-08. Vikunja V-664.
|
||||
Weights: `router_heads.onnx`, fp32, exported from `heads.pt` on workpc.
|
||||
Fixture: `internal/router/eval/ru_routing_v1.json`, 96 cases, 33 carrying a destination.
|
||||
Runner: `make t PKG=./internal/router/eval/ RUN=TestONNXRoutingHeads`.
|
||||
|
||||
The four heads of V-661 ran nowhere. This is the number they score through the
|
||||
Go cascade. Same fixture and same grammars as `TestONNXBaseline`, and only the
|
||||
middle stage varies.
|
||||
|
||||
## Headline
|
||||
|
||||
| | classifier + ONNX | heads + classifier | gemma-4-12b cascade |
|
||||
|---|---|---|---|
|
||||
| intent | 75.0% (72/96) | **96.9% (93/96)** | 84.4% |
|
||||
| destination | 33.3% (11/33) | **75.8% (25/33)** | 72.7% |
|
||||
| false clarify | 0 | 1 | 2 |
|
||||
| missed clarify | 8 | 1 | 1 |
|
||||
| p50 | 24.5ms | 27.9ms | 329ms |
|
||||
|
||||
A 118M encoder beats the 12B teacher it was distilled from. It wins on both
|
||||
halves of the route, at a twelfth of the latency. The workstation stays the
|
||||
better phraser and is no longer the better router.
|
||||
|
||||
The p50 is not the heads. Most of it is the classifier's own embedder pass on
|
||||
the turns the heads decline, plus process warm-up on the first case. The heads'
|
||||
own forward pass measures 7.3ms on workpc.
|
||||
|
||||
## Two defects were in the way, and the first was not in the heads
|
||||
|
||||
**The tokenizer read every long word backwards.** `encodeWord` backtracks the
|
||||
Viterbi path from the end of a word and prepends each piece. That puts them back
|
||||
in reading order, and a second reverse after the loop undid it. So
|
||||
`query: вода` tokenized to `[0 12 1294 41 12489 2]` where the reference
|
||||
tokenizer gives `[0 41 1294 12 12489 2]`.
|
||||
|
||||
It was found here and only here. The heads were trained through transformers and
|
||||
are read through the hand-written tokenizer. So a mismatch shows up as a score
|
||||
far below what Python measured on the same weights. Nothing else in the suite
|
||||
compares the two.
|
||||
|
||||
Measured on the recall fixture, same 27 cases either way:
|
||||
|
||||
| | reversed | fixed |
|
||||
|---|---|---|
|
||||
| recall@1 | 70.4% (19/27) | **77.8% (21/27)** |
|
||||
| recall@3 | 85.2% (23/27) | **96.3% (26/27)** |
|
||||
| answered after gate | 63.0% | 66.7% |
|
||||
| wrong note on top | 8 | 6 |
|
||||
| false recall | 0/5 | 1/5 |
|
||||
|
||||
The classifier barely moved, 76.0% to 75.0%, and destination 36.4% to 33.3%.
|
||||
Both are one case on 96 and neither is a finding. Seeds and queries were mangled
|
||||
the same way, so cosine survived it. Recall is where it cost, because a stored
|
||||
passage and a live query are different lengths and break differently.
|
||||
|
||||
The one new false recall is the honest cost and it is not being hidden. A
|
||||
sharper embedder scores every candidate higher, including the ones that should
|
||||
have stayed under the gate. That is the same trade `2026-08-04-recall-e5-small.md`
|
||||
recorded when e5-small replaced MiniLM.
|
||||
|
||||
The embedder id now carries a tokenizer revision, `model_quantized@384/tok2`.
|
||||
Stored vectors were written under rev 1 and no longer sit in the same space as a
|
||||
query embedded now. The model file's name never moved, so nothing would have
|
||||
triggered `ReembedAll`. On the box the marker fired on start, and the re-embed
|
||||
rewrote 65 notes and 19 facts in 5 seconds.
|
||||
|
||||
**The clarify head was being thrown away.** It was read only when the intent head
|
||||
cleared its own threshold. That cost 6 of the 8 ambiguous cases. `вода` reads as intent
|
||||
`act` at 0.233 and clarify at 0.983. Burying that handed the turn to the
|
||||
classifier, which routed it confidently and never asked. The clarify head answers
|
||||
a different question, which is whether there is enough here to act on at all. So
|
||||
it decides on its own and decides first.
|
||||
|
||||
| | intent-gated | clarify decides first |
|
||||
|---|---|---|
|
||||
| intent | 90.6% | 96.9% |
|
||||
| missed clarify | 7 | 1 |
|
||||
| false clarify | 0 | 1 |
|
||||
|
||||
## The threshold is measured, not chosen
|
||||
|
||||
Max softmax over the intent head, on the 88 cases carrying an intent:
|
||||
|
||||
| threshold | kept | accuracy kept | wrong kept | right dropped |
|
||||
|---|---|---|---|---|
|
||||
| 0.5 | 84 | 96.4% | 3 | 2 |
|
||||
| **0.6** | **81** | **97.5%** | **2** | **4** |
|
||||
| 0.7 | 75 | 97.3% | 2 | 10 |
|
||||
| 0.8 | 64 | 96.9% | 2 | 21 |
|
||||
| 0.9 | 46 | 100.0% | 0 | 37 |
|
||||
|
||||
0.6 is the knee. Every value from 0.7 to 0.85 drops right answers and keeps the
|
||||
same two wrong ones. 0.9 is the only value that clears them, and it costs 37
|
||||
correct routes to do it.
|
||||
|
||||
## Quantization was measured and rejected
|
||||
|
||||
| build | size | intent | destination | p50 |
|
||||
|---|---|---|---|---|
|
||||
| fp32 | 470MB | 83/88 (94.3%) | 28/33 (84.8%) | 7.3ms |
|
||||
| int8 | 118MB | 79/88 (89.8%) | 26/33 (78.8%) | 4.0ms |
|
||||
| fp16 | 235MB | will not load | — | — |
|
||||
|
||||
Python numbers, on the heads alone rather than through the cascade. int8 costs
|
||||
4.5 points of intent and 6 of destination to save 3ms. The cascade around it has
|
||||
a p50 over a second when the resident model answers. The fp16 graph is broken:
|
||||
`convert_float_to_float16` leaves a Cast node emitting float16 where the graph
|
||||
expects float, and onnxruntime refuses the session. It was not worth fixing.
|
||||
|
||||
The exporter also had to be told to write one file. It splits weights into a
|
||||
`.onnx.data` sidecar by default. This onnxruntime resolves that path against the
|
||||
process working directory rather than the model. A split graph loads from one
|
||||
directory only.
|
||||
|
||||
## What is still wrong
|
||||
|
||||
**Four of the eight destination misses are calendar.** Training cannot move them.
|
||||
The possessive agenda rules claim those cases at stage 0 and name nothing on
|
||||
purpose. That caution was free while nothing downstream could name anything
|
||||
either. It has now cost four points in three separate measurements. The call is
|
||||
the owner's and it is still open.
|
||||
|
||||
**The slot head is exported and not read.** Slots come from the stage-2
|
||||
extractor. Mapping BIO tags back to text needs character offsets the unigram
|
||||
tokenizer does not keep, which is its own piece of work.
|
||||
|
||||
**`поужинал` is a false clarify**, which is the same defect `thinSingleToken`
|
||||
was narrowed for on 2026-08-01, arriving now from a different direction.
|
||||
|
||||
## On the box
|
||||
|
||||
Deployed to homesrv the same day. `voice: routing heads loaded` on start, and
|
||||
`/trace` shows `routing-heads` winning or thinning every turn. The resident model
|
||||
and the classifier are both marked never asked. Live probes:
|
||||
|
||||
```text
|
||||
что такое TCP? -> kiwix a real definition
|
||||
кто такой Линус Торвальдс? -> kiwix a real answer
|
||||
во сколько я лёг вчера -> personal не нашла у тебя такой записи
|
||||
вода -> thinned to clarify at 0.233 / 0.983
|
||||
```
|
||||
|
||||
A missing or broken weights file logs and leaves the heads nil, which is
|
||||
byte-for-byte the cascade that shipped before this.
|
||||
@@ -0,0 +1,110 @@
|
||||
# CrisperWhisper 2.0 in Russian, measured
|
||||
|
||||
Date: 2026-08-09. Vikunja V-665.
|
||||
Corpus: `bond005/sberdevices_golos_10h_crowd`, test split, first 200 clips.
|
||||
Harness: `~/Programs/cw2-eval` on workpc, not in this repo.
|
||||
Runner: `./.venv/bin/python run_asr.py <arm>...` then `score.py`.
|
||||
|
||||
The model card benchmarks disfluency F1 in German and English. It never names
|
||||
Russian and publishes no per-language WER. So the measurement came before the
|
||||
wiring.
|
||||
|
||||
## The corpus
|
||||
|
||||
200 clips, 13.7 minutes, 1001 reference words. Median clip 3.91s, range 1.04s
|
||||
to 13.5s. Golos crowd is short crowd-sourced Russian spoken close to the
|
||||
microphone, which is the nearest public thing to someone talking to Maven. The
|
||||
alternatives are read speech, which flatters every model equally.
|
||||
|
||||
Two rows carry a null transcription and are skipped.
|
||||
|
||||
Scoring normalizes both sides: lowercase, `ё` to `е`, punctuation stripped, and
|
||||
digits expanded to Russian words through num2words. Without that last step a
|
||||
model is penalized for writing `60000` where the reference says
|
||||
`шестьдесят тысяч`. Thousands separators are joined before expansion, or
|
||||
`60 000` expands to `шестьдесят ноль`.
|
||||
|
||||
## Headline
|
||||
|
||||
| arm | WER | CER | exact | empty | RTF |
|
||||
|---|---|---|---|---|---|
|
||||
| cw2-turbo-intended | **10.4%** | 3.4% | 65.5% | 0 | 0.065 |
|
||||
| cw2-turbo-verbatim | 10.8% | **3.1%** | **66.5%** | 0 | 0.065 |
|
||||
| whisper-turbo | 11.8% | 4.1% | 64.0% | 0 | 0.031 |
|
||||
| cw2-large-intended | 12.3% | 3.8% | 63.5% | 0 | 0.107 |
|
||||
| whisper-small | 27.5% | 9.8% | 35.0% | 0 | 0.026 |
|
||||
|
||||
`whisper-small` is the floor, because `ggml-small.bin` is what mavsttd loads on
|
||||
homesrv today. CW2 turbo beats it by 17 points of WER and takes exact matches
|
||||
from 35.0% to 65.5%.
|
||||
|
||||
Two results are worth naming beyond the winner. CW2 turbo beats its own base
|
||||
model, whisper-large-v3-turbo, by 1.4 points. And it beats CW2 large by 1.9
|
||||
points, which inverts what the card implies by calling turbo a degraded draft.
|
||||
No arm returned an empty transcript.
|
||||
|
||||
## Intended and verbatim are closer than the mode names suggest
|
||||
|
||||
The two modes disagree on 70 of the 200 clips before normalization and on 29
|
||||
after it. So the raw difference is mostly casing and punctuation, which
|
||||
normalization removes and which Maven does not read either.
|
||||
|
||||
Verbatim scores worse on WER and better on CER and exact matches. The reason is
|
||||
script, not disfluency:
|
||||
|
||||
```text
|
||||
ref: футбольный матч челси брайтон
|
||||
int: Футбольный матч Chelsea-Брайтон.
|
||||
ver: Футбольный матч Челси Брайтон.
|
||||
```
|
||||
|
||||
Intended writes foreign entity names in Latin script and verbatim
|
||||
transliterates them. Golos references are Cyrillic throughout, so verbatim
|
||||
collects the exact matches. That is a property of this corpus rather than a
|
||||
quality difference.
|
||||
|
||||
**This corpus cannot settle the mode choice.** Golos crowd is clean short
|
||||
commands with almost no disfluency. The two modes have nothing to disagree
|
||||
about here. They separate on spontaneous speech with fillers, restarts and
|
||||
repairs, which is what the owner speaks. Intended stays the choice for the
|
||||
reason it was always the choice. Maven wants what was meant, not every stumble
|
||||
on the way there.
|
||||
|
||||
The Latin-script habit is the one finding here that touches routing. The
|
||||
routing heads were trained on Cyrillic utterances, so an entity name arriving
|
||||
in Latin script is out of distribution for them. Nothing measures that yet.
|
||||
|
||||
## The runtime is workpc, because whisper.cpp cannot load CW2
|
||||
|
||||
`num_languages()` in `deps/whisper.cpp/src/whisper.cpp` derives the language
|
||||
count from the vocabulary size:
|
||||
|
||||
```cpp
|
||||
return n_vocab - 51765 - (is_multilingual() ? 1 : 0);
|
||||
```
|
||||
|
||||
CW2 carries 31 extra tokens, so `n_vocab` is 51897 and this yields 131
|
||||
languages. The derived `dt` offset becomes 33 and shifts seven special token
|
||||
ids, including `token_beg` and `token_transcribe`. The architecture is
|
||||
otherwise byte-identical to whisper-large-v3-turbo, and the new tokens sit
|
||||
above every whisper special id.
|
||||
|
||||
So loading CW2 in whisper.cpp is a patch to a vendored dependency, not a port.
|
||||
It was not taken, because STT is moving to workpc anyway under V-486. CW2 turbo
|
||||
becomes the preferred remote and `ggml-small.bin` on homesrv stays the floor,
|
||||
which is the shape `modelSeam` already uses for routing and replies. The 27.5%
|
||||
floor is what a turn falls back to when the workstation is down, and this table
|
||||
is what that costs.
|
||||
|
||||
## License
|
||||
|
||||
Standard CW2 weights carry `nyra-health-non-commercial-research`. The Pro
|
||||
variants are commercial-license only. Maven is personal and self-hosted, so the
|
||||
standard weights are usable and the Pro ones are not free to take.
|
||||
|
||||
## What is not measured
|
||||
|
||||
Disfluent spontaneous speech, which is the whole reason to prefer Intended.
|
||||
Long-form audio beyond 13.5s. Far-field or noisy microphones. English, which
|
||||
Maven also speaks. The ONNX turbo export, which was never run, since the
|
||||
transformers path already meets the latency budget at RTF 0.065.
|
||||
Reference in New Issue
Block a user