Files
Maven/docs/evals/2026-08-05-audio-in-routing.md
claude b789676244 audio-in routing measured: transcribe then route (V-486)
Four paths on the same 72 RU cases with the daemon's own router prompt.
Text in scores 90.3% intent-only. Whisper then route scores 84.7% at p50
1372ms. The workstation transcribing then routing scores 83.3% at p50 997ms.
One call from audio straight to a route scores 54.2%.

The one-call number is not a transcription failure. Four clips it
transcribes word for word it then routes wrong or refuses, and the emitted
slot holds the tail of the sentence with the interrogative head gone. A
3.5k-character classification prompt and an audio part compete for
attention, so transcription needs its own call with a short instruction.

The two speech-to-text paths differ by one case, which is noise on 72, so
the choice is latency and transcript quality. The workstation wins both.
mavgpud.json on the workstation is restored to its text-only args.
2026-08-05 17:04:45 +04:00

5.1 KiB
Raw Permalink Blame History

Routing from audio: four paths, one fixture

05-08-2026. Vikunja #486. Workstation gemma-4-12B-it-qat-UD-Q4_K_XL with mmproj-F16.gguf, homesrv whisper ggml-small, piper ru_RU-irina-medium.

Verdict: transcribe, then route. One call from audio straight to a route loses 36 points, so it is not a candidate. Moving speech-to-text to the workstation buys 375ms and better transcripts at no measurable accuracy cost. So #486 proceeds on the two-call shape.

The numbers

72 Russian cases from internal/router/eval/ru_routing_v1.json, rendered by piper at 16kHz mono, 153.9s of audio, mean 2.14s per clip. Every path used the daemon's own routeSystem prompt and routeGrammar, read out of internal/router/llmrouter.go at run time, at temperature 0 and enable_thinking:false.

Path Intent-only Verbatim transcripts p50 p95
text in, the ceiling 90.3% (65/72) 361ms 495ms
whisper on homesrv, then route 84.7% (61/72) 29/72 1372ms 1546ms
workstation transcribes, then routes 83.3% (60/72) 48/72 997ms 1177ms
workstation, one call from audio 54.2% (39/72) 425ms 756ms

The 90.3% ceiling is the same model on the same 72 cases with the utterance as text. It is not the 93.5% in docs/evals/2026-08-02-workstation-gemma4-12b.md, which scored all 87 cases including the English ones.

The two speech-to-text paths differ by one case, which is noise on 72. So the choice between them is latency and transcript quality, and the workstation wins both.

One call from audio is not a transcription failure

The obvious reading of 54.2% is that the audio encoder cannot hear Russian. It can. Eight of the failing clips were sent back with a transcribe instruction instead of the router prompt:

Clip Said Heard, transcribing Routing from audio
ru-sys-002 какое число завтра Какое число завтра? unknown
ru-sys-003 переходи в тихий режим Переходи в тихий режим. unknown
ru-query-001 сколько воды я выпил с утра Сколько воды я выпил с утра? fact, value "выпил с утра"
ru-act-002 выключи свет в спальне Выключи свет в спальне. fact, value "включен"

Four clips it transcribes word for word, and routes wrong or refuses. The ru-query-001 row shows the mechanism: the emitted slot holds the tail of the sentence and the interrogative head is gone. The model is not deaf, it stops attending to the audio once it is also holding a 3.5k-character classification prompt.

That pattern decides the whole task. A long system prompt and an audio part compete, so the transcription has to be its own call with a short instruction. It also means the number would not be rescued by a better prompt, a longer clip, or a bigger mmproj.

The failures cluster where the head of the sentence carries the intent: ru-act 1/6, ru-sys 2/5, ru-query 12/25. Reminders scored 10/10, because "напомни" is the first word and nothing after it changes the answer.

Transcript quality and routing accuracy come apart

The workstation transcribes 48 of 72 verbatim against whisper's 29, and routes one case worse. Both directions of that appear in the same run:

  • ru-chat-002: whisper heard "Кто думаешь про переезд", the workstation heard "Что ты думаешь про переезд". The correct transcript routed to chat, the broken one to query.
  • ru-query-020: whisper heard "Кто дальше?", the workstation heard the correct "Что дальше?". The broken transcript routed correctly and the correct one missed.

A word error rate is not a proxy for routing accuracy here. Judge a speech-to-text change on the routing fixture, not on transcripts.

Three cases only the text path gets right. No speech-to-text path recovers them, so they are lost in the rendering rather than in the model.

Latency

Whisper ggml-small on homesrv CPU costs p50 998ms for a 2.14s clip, which is nearly all of that path's 1372ms. The workstation does the same job inside its 997ms end-to-end total for two calls. So the transfer is worth about 375ms per turn at p50, and more at p95.

Both are above the one-call 425ms, and that is the trade the table settles: 29 points of accuracy for 572ms.

Notes for the next run

  • --mmproj /mnt/D/AI/gemma4/mmproj-F16.gguf has to be in llama_args in ~/.config/mavgpud.json, or /props reports modalities.audio: false and every audio part is dropped silently. It was added for this measurement and removed afterwards, so the box is back to the text-only config.
  • enable_thinking:false is mandatory. It was set for all 224 calls here.
  • The degenerate <|channel>thought output recorded against #486 did not reproduce, in 80 transcribe calls or in 144 routing calls.
  • Piper renders at 22050Hz mono. Every clip was resampled with ffmpeg -ar 16000 -ac 1 -c:a pcm_s16le, because 16kHz is what audio.PCM16kMono declares and what the earlier measurement used.