Merge pull request 'Measure CrisperWhisper 2.0 turbo in Russian before wiring a runtime for it' (#207) from task/665-crisperwhisper-2-russian into master

This commit was merged in pull request #207.
This commit is contained in:
2026-08-08 23:21:13 +02:00
@@ -0,0 +1,110 @@
# CrisperWhisper 2.0 in Russian, measured
Date: 2026-08-09. Vikunja V-665.
Corpus: `bond005/sberdevices_golos_10h_crowd`, test split, first 200 clips.
Harness: `~/Programs/cw2-eval` on workpc, not in this repo.
Runner: `./.venv/bin/python run_asr.py <arm>...` then `score.py`.
The model card benchmarks disfluency F1 in German and English. It never names
Russian and publishes no per-language WER. So the measurement came before the
wiring.
## The corpus
200 clips, 13.7 minutes, 1001 reference words. Median clip 3.91s, range 1.04s
to 13.5s. Golos crowd is short crowd-sourced Russian spoken close to the
microphone, which is the nearest public thing to someone talking to Maven. The
alternatives are read speech, which flatters every model equally.
Two rows carry a null transcription and are skipped.
Scoring normalizes both sides: lowercase, `ё` to `е`, punctuation stripped, and
digits expanded to Russian words through num2words. Without that last step a
model is penalized for writing `60000` where the reference says
`шестьдесят тысяч`. Thousands separators are joined before expansion, or
`60 000` expands to `шестьдесят ноль`.
## Headline
| arm | WER | CER | exact | empty | RTF |
|---|---|---|---|---|---|
| cw2-turbo-intended | **10.4%** | 3.4% | 65.5% | 0 | 0.065 |
| cw2-turbo-verbatim | 10.8% | **3.1%** | **66.5%** | 0 | 0.065 |
| whisper-turbo | 11.8% | 4.1% | 64.0% | 0 | 0.031 |
| cw2-large-intended | 12.3% | 3.8% | 63.5% | 0 | 0.107 |
| whisper-small | 27.5% | 9.8% | 35.0% | 0 | 0.026 |
`whisper-small` is the floor, because `ggml-small.bin` is what mavsttd loads on
homesrv today. CW2 turbo beats it by 17 points of WER and takes exact matches
from 35.0% to 65.5%.
Two results are worth naming beyond the winner. CW2 turbo beats its own base
model, whisper-large-v3-turbo, by 1.4 points. And it beats CW2 large by 1.9
points, which inverts what the card implies by calling turbo a degraded draft.
No arm returned an empty transcript.
## Intended and verbatim are closer than the mode names suggest
The two modes disagree on 70 of the 200 clips before normalization and on 29
after it. So the raw difference is mostly casing and punctuation, which
normalization removes and which Maven does not read either.
Verbatim scores worse on WER and better on CER and exact matches. The reason is
script, not disfluency:
```text
ref: футбольный матч челси брайтон
int: Футбольный матч Chelsea-Брайтон.
ver: Футбольный матч Челси Брайтон.
```
Intended writes foreign entity names in Latin script and verbatim
transliterates them. Golos references are Cyrillic throughout, so verbatim
collects the exact matches. That is a property of this corpus rather than a
quality difference.
**This corpus cannot settle the mode choice.** Golos crowd is clean short
commands with almost no disfluency. The two modes have nothing to disagree
about here. They separate on spontaneous speech with fillers, restarts and
repairs, which is what the owner speaks. Intended stays the choice for the
reason it was always the choice. Maven wants what was meant, not every stumble
on the way there.
The Latin-script habit is the one finding here that touches routing. The
routing heads were trained on Cyrillic utterances, so an entity name arriving
in Latin script is out of distribution for them. Nothing measures that yet.
## The runtime is workpc, because whisper.cpp cannot load CW2
`num_languages()` in `deps/whisper.cpp/src/whisper.cpp` derives the language
count from the vocabulary size:
```cpp
return n_vocab - 51765 - (is_multilingual() ? 1 : 0);
```
CW2 carries 31 extra tokens, so `n_vocab` is 51897 and this yields 131
languages. The derived `dt` offset becomes 33 and shifts seven special token
ids, including `token_beg` and `token_transcribe`. The architecture is
otherwise byte-identical to whisper-large-v3-turbo, and the new tokens sit
above every whisper special id.
So loading CW2 in whisper.cpp is a patch to a vendored dependency, not a port.
It was not taken, because STT is moving to workpc anyway under V-486. CW2 turbo
becomes the preferred remote and `ggml-small.bin` on homesrv stays the floor,
which is the shape `modelSeam` already uses for routing and replies. The 27.5%
floor is what a turn falls back to when the workstation is down, and this table
is what that costs.
## License
Standard CW2 weights carry `nyra-health-non-commercial-research`. The Pro
variants are commercial-license only. Maven is personal and self-hosted, so the
standard weights are usable and the Pro ones are not free to take.
## What is not measured
Disfluent spontaneous speech, which is the whole reason to prefer Intended.
Long-form audio beyond 13.5s. Far-field or noisy microphones. English, which
Maven also speaks. The ONNX turbo export, which was never run, since the
transformers path already meets the latency budget at RTF 0.065.