Four paths on the same 72 RU cases with the daemon's own router prompt. Text in scores 90.3% intent-only. Whisper then route scores 84.7% at p50 1372ms. The workstation transcribing then routing scores 83.3% at p50 997ms. One call from audio straight to a route scores 54.2%. The one-call number is not a transcription failure. Four clips it transcribes word for word it then routes wrong or refuses, and the emitted slot holds the tail of the sentence with the interrogative head gone. A 3.5k-character classification prompt and an audio part compete for attention, so transcription needs its own call with a short instruction. The two speech-to-text paths differ by one case, which is noise on 72, so the choice is latency and transcript quality. The workstation wins both. mavgpud.json on the workstation is restored to its text-only args.
5.1 KiB
Routing from audio: four paths, one fixture
05-08-2026. Vikunja #486. Workstation gemma-4-12B-it-qat-UD-Q4_K_XL with
mmproj-F16.gguf, homesrv whisper ggml-small, piper ru_RU-irina-medium.
Verdict: transcribe, then route. One call from audio straight to a route loses 36 points, so it is not a candidate. Moving speech-to-text to the workstation buys 375ms and better transcripts at no measurable accuracy cost. So #486 proceeds on the two-call shape.
The numbers
72 Russian cases from internal/router/eval/ru_routing_v1.json, rendered by piper at
16kHz mono, 153.9s of audio, mean 2.14s per clip. Every path used the daemon's own
routeSystem prompt and routeGrammar, read out of internal/router/llmrouter.go at run
time, at temperature 0 and enable_thinking:false.
| Path | Intent-only | Verbatim transcripts | p50 | p95 |
|---|---|---|---|---|
| text in, the ceiling | 90.3% (65/72) | — | 361ms | 495ms |
| whisper on homesrv, then route | 84.7% (61/72) | 29/72 | 1372ms | 1546ms |
| workstation transcribes, then routes | 83.3% (60/72) | 48/72 | 997ms | 1177ms |
| workstation, one call from audio | 54.2% (39/72) | — | 425ms | 756ms |
The 90.3% ceiling is the same model on the same 72 cases with the utterance as text. It is
not the 93.5% in docs/evals/2026-08-02-workstation-gemma4-12b.md, which scored all 87
cases including the English ones.
The two speech-to-text paths differ by one case, which is noise on 72. So the choice between them is latency and transcript quality, and the workstation wins both.
One call from audio is not a transcription failure
The obvious reading of 54.2% is that the audio encoder cannot hear Russian. It can. Eight of the failing clips were sent back with a transcribe instruction instead of the router prompt:
| Clip | Said | Heard, transcribing | Routing from audio |
|---|---|---|---|
| ru-sys-002 | какое число завтра | Какое число завтра? | unknown |
| ru-sys-003 | переходи в тихий режим | Переходи в тихий режим. | unknown |
| ru-query-001 | сколько воды я выпил с утра | Сколько воды я выпил с утра? | fact, value "выпил с утра" |
| ru-act-002 | выключи свет в спальне | Выключи свет в спальне. | fact, value "включен" |
Four clips it transcribes word for word, and routes wrong or refuses. The ru-query-001
row shows the mechanism: the emitted slot holds the tail of the sentence and the
interrogative head is gone. The model is not deaf, it stops attending to the audio once it
is also holding a 3.5k-character classification prompt.
That pattern decides the whole task. A long system prompt and an audio part compete, so the
transcription has to be its own call with a short instruction. It also means the number
would not be rescued by a better prompt, a longer clip, or a bigger mmproj.
The failures cluster where the head of the sentence carries the intent: ru-act 1/6,
ru-sys 2/5, ru-query 12/25. Reminders scored 10/10, because "напомни" is the first word
and nothing after it changes the answer.
Transcript quality and routing accuracy come apart
The workstation transcribes 48 of 72 verbatim against whisper's 29, and routes one case worse. Both directions of that appear in the same run:
ru-chat-002: whisper heard "Кто думаешь про переезд", the workstation heard "Что ты думаешь про переезд". The correct transcript routed tochat, the broken one toquery.ru-query-020: whisper heard "Кто дальше?", the workstation heard the correct "Что дальше?". The broken transcript routed correctly and the correct one missed.
A word error rate is not a proxy for routing accuracy here. Judge a speech-to-text change on the routing fixture, not on transcripts.
Three cases only the text path gets right. No speech-to-text path recovers them, so they are lost in the rendering rather than in the model.
Latency
Whisper ggml-small on homesrv CPU costs p50 998ms for a 2.14s clip, which is nearly
all of that path's 1372ms. The workstation does the same job inside its 997ms end-to-end
total for two calls. So the transfer is worth about 375ms per turn at p50, and more at p95.
Both are above the one-call 425ms, and that is the trade the table settles: 29 points of accuracy for 572ms.
Notes for the next run
--mmproj /mnt/D/AI/gemma4/mmproj-F16.ggufhas to be inllama_argsin~/.config/mavgpud.json, or/propsreportsmodalities.audio: falseand every audio part is dropped silently. It was added for this measurement and removed afterwards, so the box is back to the text-only config.enable_thinking:falseis mandatory. It was set for all 224 calls here.- The degenerate
<|channel>thoughtoutput recorded against #486 did not reproduce, in 80 transcribe calls or in 144 routing calls. - Piper renders at 22050Hz mono. Every clip was resampled with
ffmpeg -ar 16000 -ac 1 -c:a pcm_s16le, because 16kHz is whataudio.PCM16kMonodeclares and what the earlier measurement used.