# Routing from audio: four paths, one fixture **05-08-2026. Vikunja #486.** Workstation `gemma-4-12B-it-qat-UD-Q4_K_XL` with `mmproj-F16.gguf`, homesrv whisper `ggml-small`, piper `ru_RU-irina-medium`. **Verdict: transcribe, then route.** One call from audio straight to a route loses 36 points, so it is not a candidate. Moving speech-to-text to the workstation buys 375ms and better transcripts at no measurable accuracy cost. So #486 proceeds on the two-call shape. ## The numbers 72 Russian cases from `internal/router/eval/ru_routing_v1.json`, rendered by piper at 16kHz mono, 153.9s of audio, mean 2.14s per clip. Every path used the daemon's own `routeSystem` prompt and `routeGrammar`, read out of `internal/router/llmrouter.go` at run time, at `temperature 0` and `enable_thinking:false`. | Path | Intent-only | Verbatim transcripts | p50 | p95 | |---|---|---|---|---| | text in, the ceiling | **90.3%** (65/72) | — | 361ms | 495ms | | whisper on homesrv, then route | **84.7%** (61/72) | 29/72 | 1372ms | 1546ms | | workstation transcribes, then routes | **83.3%** (60/72) | 48/72 | 997ms | 1177ms | | workstation, one call from audio | **54.2%** (39/72) | — | 425ms | 756ms | The 90.3% ceiling is the same model on the same 72 cases with the utterance as text. It is not the 93.5% in `docs/evals/2026-08-02-workstation-gemma4-12b.md`, which scored all 87 cases including the English ones. The two speech-to-text paths differ by one case, which is noise on 72. So the choice between them is latency and transcript quality, and the workstation wins both. ## One call from audio is not a transcription failure The obvious reading of 54.2% is that the audio encoder cannot hear Russian. It can. Eight of the failing clips were sent back with a transcribe instruction instead of the router prompt: | Clip | Said | Heard, transcribing | Routing from audio | |---|---|---|---| | ru-sys-002 | какое число завтра | Какое число завтра? | `unknown` | | ru-sys-003 | переходи в тихий режим | Переходи в тихий режим. | `unknown` | | ru-query-001 | сколько воды я выпил с утра | Сколько воды я выпил с утра? | `fact`, value "выпил с утра" | | ru-act-002 | выключи свет в спальне | Выключи свет в спальне. | `fact`, value "включен" | Four clips it transcribes word for word, and routes wrong or refuses. The `ru-query-001` row shows the mechanism: the emitted slot holds the tail of the sentence and the interrogative head is gone. The model is not deaf, it stops attending to the audio once it is also holding a 3.5k-character classification prompt. That pattern decides the whole task. A long system prompt and an audio part compete, so the transcription has to be its own call with a short instruction. It also means the number would not be rescued by a better prompt, a longer clip, or a bigger `mmproj`. The failures cluster where the head of the sentence carries the intent: `ru-act` 1/6, `ru-sys` 2/5, `ru-query` 12/25. Reminders scored 10/10, because "напомни" is the first word and nothing after it changes the answer. ## Transcript quality and routing accuracy come apart The workstation transcribes 48 of 72 verbatim against whisper's 29, and routes one case worse. Both directions of that appear in the same run: - `ru-chat-002`: whisper heard "Кто думаешь про переезд", the workstation heard "Что ты думаешь про переезд". The correct transcript routed to `chat`, the broken one to `query`. - `ru-query-020`: whisper heard "Кто дальше?", the workstation heard the correct "Что дальше?". The **broken** transcript routed correctly and the correct one missed. A word error rate is not a proxy for routing accuracy here. Judge a speech-to-text change on the routing fixture, not on transcripts. Three cases only the text path gets right. No speech-to-text path recovers them, so they are lost in the rendering rather than in the model. ## Latency Whisper `ggml-small` on homesrv CPU costs p50 998ms for a 2.14s clip, which is nearly all of that path's 1372ms. The workstation does the same job inside its 997ms end-to-end total for two calls. So the transfer is worth about 375ms per turn at p50, and more at p95. Both are above the one-call 425ms, and that is the trade the table settles: 29 points of accuracy for 572ms. ## Notes for the next run - `--mmproj /mnt/D/AI/gemma4/mmproj-F16.gguf` has to be in `llama_args` in `~/.config/mavgpud.json`, or `/props` reports `modalities.audio: false` and every audio part is dropped silently. It was added for this measurement and removed afterwards, so the box is back to the text-only config. - `enable_thinking:false` is mandatory. It was set for all 224 calls here. - The degenerate `<|channel>thought` output recorded against #486 did not reproduce, in 80 transcribe calls or in 144 routing calls. - Piper renders at 22050Hz mono. Every clip was resampled with `ffmpeg -ar 16000 -ac 1 -c:a pcm_s16le`, because 16kHz is what `audio.PCM16kMono` declares and what the earlier measurement used.