Golden-audio STT tests against real whisper.cpp (#288) #75
Closed
claude
wants to merge 1 commits from
overnight/stt-golden-audio into overnight/senses-speaker
pull from: overnight/stt-golden-audio
merge into: kami:overnight/senses-speaker
kami:master
kami:task/725-capability-ledger-and-empirical-baseline
kami:task/692-heads-path-may-equal-model-path-and-noth
kami:task/694-staticcheck-and-deadcode-are-still-not-i
kami:task/682-go-1-25-5-and-x-text-0-14-0-carry-20-rea
kami:task/674-caveats
kami:task/673-mavgpud-serves-the-model-to-the-whole-la
kami:task/487-capture-device-doc
kami:task/487-capture-device
kami:task/487-wake-word-deploy
kami:task/487-wake-word-threshold
kami:task/487-wake-word-stage-two
kami:task/671-mavwaked-registers-as-a-voice-consumer-i
kami:task/670-cut-claude-md-to-200-lines
kami:task/515-deploy-mavwaked-workpc
kami:task/669-prune-claude-md
kami:task/668-e4b-phrasing
kami:task/668-title-capital
kami:task/668-kiwix-answers-a-question-it-cannot-answe
kami:task/666-only-a-stage-0-grammar-may-take-the-pers
kami:task/487-mavwaked-has-no-wake-word-only-an-energy
kami:task/486-deploy-the-workstation-transcriber
kami:task/486-move-stt-and-tts-to-the-workstation-wher
kami:task/665-crisperwhisper-2-russian
kami:task/664-routing-heads-in-go
kami:task/662-usage-harness-source-badge
kami:task/661-post-merge-usage-rerun
kami:task/661-routing-heads-step-3-train-the-multi-hea
kami:task/660-router-prompt-destination
kami:task/659-destination-fixture
kami:task/655-query-source-is-a-routing-decision-made
kami:task/654-a-pending-clarify-has-no-way-out-neither
kami:task/654-week-of-usage-eval-docs
kami:task/649-needs-kami-telegram-is-the-only-reach-an
kami:task/643-memorystore-search-decodes-and-unmarshal
kami:task/641-two-maps-grow-for-the-process-lifetime-w
kami:task/644-mavcaldav-is-built-documented-as-running
kami:task/642-the-store-caps-sqlite-at-one-connection
kami:task/647-factenrichmentworker-walks-the-pending-q
kami:task/646-v-637-follow-up-telegram-intake-has-no-d
kami:task/638-no-deadline-survives-the-turn-path-from
kami:task/637-inbound-telegram-turns-and-corrections-f
kami:task/636-correcting-a-turn-from-telegram-and-from
kami:task/634-an-act-alias-resolves-the-verb-but-not-t
kami:task/630-one-gesture-correction-on-chat-v-628
kami:task/629-persist-the-routing-trace-and-record-it
kami:task/631-mode-inventory-written-from-the-handlers
kami:task/586-defaultfactparser-uses-hand-written-russ
kami:task/633-reconcile-the-seed-labels-with-the-handl
kami:task/627-reminder-verbs-has-no-alarm-verb-so-an-a
kami:task/626-the-classifier-seeds-teach-an-older-inte
kami:task/546-route-with-a-fine-tuned-e5-small-instead
kami:task/586-measure-the-fact-parser
kami:fix/gofmt-ecosystem-acts
kami:task/584-media-store-a-failed-write-leaks-its-bud
kami:task/518-no-write-path-for-a-backdated-event-so-t
kami:task/287-qa-voice-session-quality-polish
kami:task/492-qa-plan-reconcile
kami:task/530-sweep-tail-four-files-the-russian-sweep
kami:task/405-score-how-often-a-real-utterance-reaches
kami:task/529-money-and-list-pick-a-mechanism
kami:task/528-sweep-tail-the-three-files-on-467
kami:task/527-embedder-open-set-phrasings-stop-being-r
kami:task/526-morphology-a-dictionary-answers-the-gram
kami:task/525-lexicons-the-finite-russian-sets-move-to
kami:task/524-entity-reference-ask-nexus-about-every-l
kami:task/523-risk-tiers-take-hexis-s-tier-for-a-hexis
kami:task/521-review-pr-111-query-strings-declension-h
kami:task/491-llama-server-core-dumps-on-every-sigterm
kami:task/479-bug-an-unconfigured-capability-does-not
kami:task/467-bug-spoken-task-capture-is-dead-the-rout
kami:task/463-deploy-mavwaked-and-mavenclient-run-nowh
kami:task/480-hearing-no-shipped-client-can-start-a-re
kami:task/432-ambient-calendar-intake-is-fragile-and-p
kami:task/431-board-surface-maven-holds-the-work-board
kami:task/433-reactivehandler-has-30-fields-and-is-pas
kami:task/371-swap-the-embedder-for-an-asymmetric-retr
kami:task/408-review-31-07-split-the-30-method-coreapi
kami:task/410-review-31-07-hand-rolled-string-enums-st
kami:task/423-review-pr50-split-internal-ipc-server-go
kami:task/422-review-pr50-split-cmd-mavend-tick-go-860
kami:task/409-review-31-07-finish-moving-mavweb-markup
kami:task/482-ambient-ingest-reads-a-notification-s-ti
kami:task/444-kuma-a-fact-per-monitor-so-she-can-name
kami:task/452-capability-model-homelab-docker-restart
kami:task/449-destructive-confirm-policy-risk-tiers-no
kami:task/453-grocery-list-items-table-fourth-append-o
kami:task/399-run-the-persona-checks-inside-the-daemon
kami:task/448-bounded-follow-up-state-pending-candidat
kami:task/455-conversation-repair-name-the-misroute-co
kami:task/454-go-mod-tidy
kami:task/458-pronunciation-dictionary-for-piper
kami:task/456-command-history-read-only-query-over-exi
kami:task/457-clarification-templates-for-the-router-s
kami:task/474-query-source-ordering-feeds-and-calendar
kami:task/469-reminders-spelled-out-times-fail-the-bod
kami:task/475-bug-the-praxis-attention-capability-is-u
kami:task/481-bug-a-transient-complaint-is-stored-as-a
kami:task/476-bug-the-router-transliterates-latin-enti
kami:task/385-decide-whether-a-parked-clarify-question
kami:task/377-backfill-routines
kami:task/421-weather-geocoder
kami:task/390-no-read-path-for-delivery-attempts
kami:task/386-recall-fixture-filler-note-ids
kami:task/473-bug-morning-item-has-no-required-flag
kami:task/465-bug-make-simulate-routes-with-an-empty
kami:task/467-bug-spoken-task-capture-is-dead
kami:task/466-bug-a-pending-clarify-is-global-so-one-u
kami:task/468-bug-pattern-detect-has-no-minimum-interv
kami:task/462-bug-checkfeminine-flags-second-person-ma
kami:task/443-safekey-drops-cyrillic-so-russian-calend
kami:task/471-bug-agendaquerygrammars-covers-today-but
kami:task/383-slottext-in-clarify-answer-would-clobber
kami:task/323-qa-phraser-coverage-is-65-3-but-the-llam
kami:task/498-bug-and-x-reach-the-model-with-no-determ
kami:task/506-strings-family-6-summaries-and-reports-i
kami:task/504-strings-family-4-act-and-smart-home-repl
kami:task/503-strings-family-3-query-answers-and-gaps
kami:task/502-strings-family-2-capture-acknowledgement
kami:task/501-strings-family-1-phrasing-fallbacks-into
kami:task/397-phrasechat-and-phrasequery-hide-model-fa
kami:task/396-the-reply-path-can-t-be-tested-llmreplie
kami:task/496-recall-a-cross-language-question-loses-i
kami:task/495-bug-x-escapes-the-personal-boundary-and
kami:task/499-llama-server-holds-7-9gb-rss-for-a-1-1gb
kami:task/470-bug-a-question-writes-invented-knowledge
kami:task/493-bug-the-memory-index-stores-the-raw-utte
kami:task/490-name-the-gap-world-questions-through-the
kami:task/485-run-the-big-model-on-the-workstation-wit
kami:task/489-workstation-deploy-mavgpud-on-workpc-and
kami:task/488-workstation-a-supervisor-that-keeps-llam
kami:task/483-docs-offload-design
kami:task/483-design-offload-ml-to-the-workstation-kee
kami:task/459-docs-refresh-the-qa-plan-against-the-liv
kami:task/446-doc-reorg-tier-the-tree-retire-the-three
kami:fix/367-voice-parks-routine-accept
kami:task/365-dialogue-slots-and-router-slots-are-hand
kami:task/364-snooze-does-nothing-at-runtime-the-gate
kami:task/447-retire-progress-md-the-backlog-and-the-f
kami:task/445-session-workflow
kami:overnight/eco-versioned-traces
kami:overnight/eco-entity-refs
kami:overnight/eco-degraded-suite
kami:overnight/netscan
kami:overnight/smarthome
kami:overnight/replay-simulator
kami:overnight/event-envelope
kami:overnight/coldstart-unlock
kami:overnight/voice-barge-in
kami:overnight/senses-speaker
kami:overnight/senses-hearing
kami:overnight/senses-media-vision
kami:overnight/mcp-tools
kami:overnight/mcp-client
kami:overnight/self-update
kami:overnight/model-swap
kami:overnight/web-crawler
kami:overnight/rss-feeds
kami:overnight/email-poller
kami:overnight/email-extract
kami:overnight/email-imap
kami:overnight/money-zenmoney
kami:overnight/task-priority
kami:overnight/task-capture
kami:overnight/behavior-profile
kami:overnight/day-plan
kami:overnight/ambient-calendar
kami:overnight/local-calendar
kami:overnight/memory-eval
kami:overnight/proactive-proposals
kami:overnight/split-voice-quiet
kami:overnight/nginx-maven-block
kami:overnight/stepup-chat-surface
kami:integration/small-batch
kami:docs/fix-drift
kami:fix/ru-wording
kami:integration/jul31
kami:overnight/resident-1.7b
kami:overnight/nudge-templates
kami:overnight/kiwix-rewrite
kami:overnight/eval-writeup
kami:overnight/fix-truncation
kami:overnight/kiwix-client
kami:overnight/ru-prompts
kami:overnight/external-data
kami:overnight/phrasing-grammar
kami:overnight/talk-eval
kami:overnight/prompt-context
kami:overnight/prompt-address
kami:overnight/eval-label-kill
kami:overnight/delivery-boundary
kami:overnight/address-check
kami:overnight/system-replies-pr
kami:overnight/clock-intent-pr
kami:overnight/embedder-backfill-pr
kami:overnight/embedder-marker-pr
kami:overnight/note-recall-pr
kami:overnight/thinking-off-pr
kami:overnight/dialogue-persist-pr
kami:overnight/persona-2p-pr
kami:overnight/clarify-expiry-pr
kami:overnight/clarify-rework
kami:overnight/phrasing
kami:overnight/bakeoff
kami:overnight/recall-margin
kami:overnight/router-on
kami:overnight/slot-extract
kami:overnight/embedder-e5
kami:overnight/router-refusal
kami:overnight/eval-rerun
kami:overnight/eval-harnesses
kami:overnight/eval-rerun-base
kami:overnight/fmt-gate
kami:overnight/routines-fire
kami:overnight/router-prompt
kami:overnight/away-leak
kami:overnight/recall-eval
kami:overnight/snooze-works
kami:overnight/clarify-wiring
kami:overnight/delivery-tests
kami:overnight/routine-accept
kami:overnight/llm-router-flag
kami:overnight/clarify-data-layer
kami:overnight/loop-rule-tests
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "overnight/stt-golden-audio"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
What changed
cmd/mavsttd/golden_test.gopushes four committed WAV fixtures through the realwhisperHandler— the same CGO whisper.cpp binding mavsttd runs in production — and scores the transcripts.cmd/mavsttd/testdata/golden_v1.json— the manifest: wav, lang, reference transcript, intent keywords, WER ceiling.cmd/mavsttd/testdata/*.wav— three Russian clips, one English. 360K total.scripts/gen-stt-fixtures.sh— regenerates them from piper.make stt-fixtures,make test-stt-golden.Why the fixtures are committed, and why they are small
whisper.cpp needs real audio; there is no way to fake it and still test it. But the fixtures are synthesised, not recorded: the script drives the vendored
deps/piper/piperwithmodels/tts/ru_RU-irina-medium.onnx— the voice Maven already speaks with — so nothing of the owner's voice is in the repo and any fixture rebuilds from the script plus the voice model. Each clip is ~2s of 16 kHz mono s16le, 60-100K.The English voice (
en_US-lessac-medium) is not vendored; the script finds it under~/esp-server/voicesand skips the English fixture when it is absent.Why matching is tolerant
Golden transcripts are model-dependent. An exact-string assertion would turn every whisper model swap into a fixture rewrite, the same way the phraser swap moved every phrasing baseline. So each case asserts two things: the intent-carrying keywords are present (prefix match, so
водыmatchesводуbutчасdoes not matchчасть), and the word error rate against the reference stays under a per-case ceiling. Both are pure functions, unit-tested in the same file without any model.How it was verified
make build— exit 0.make test— exit 0,cmd/mavsttdcoverage 7.8% to 53.9%.make test-stt-goldenwithmodels/stt/ggml-small.binpresent, all four cases pass:ru_reminder.wavtoНапомни мне через час позвонить маме.(conf 0.71)ru_fact.wavtoА отметь, что я выпил воды.(conf 0.77)ru_query.wavtoЧто у меня сегодня по календарю?(conf 0.76)en_act.wavtoRestart the web server and check the disk space.(conf 0.78)MAVEN_WHISPER_MODELat a nonexistent file:TestGoldenAudioTranscriptionskips,TestGoldenFixturesAreCanonicalstill runs and passes.make testnever breaks on a box without models.Not in this PR
Tier 2 of the task — audio to real STT to router to phraser through the fake-ecosystem harness — is not here. It needs a live llama-server for the router leg, so it would be an env-gated eval rather than a
make testtest, and it belongs with #284's replayable simulator. The fixture format (manifest plus canonical WAV) is deliberately the one #284 can reuse.Vikunja #288
Synthesising the fixtures instead of recording them is the right trade. Nothing of his voice is committed. The WAVs are regenerable from the script plus a voice model. The
nginxnote ingen-stt-fixtures.shshows the artefact was hit and worked around, not guessed at.TestGoldenFixturesAreCanonicalrunning the fixtures throughgateReasonbefore the model test uses them is the check that stops a silent-fixture false pass.1. Three of the four regressions the header claims to catch are not caught
The file comment names four regressions caught by
make test: a bad model path, a wrong language hint, a broken resample, a regressed silence gate. Walk each one.Bad model path.
TestGoldenAudioTranscriptionopens withos.Stat(model)andt.Skipfon failure. A wrong path is the exact condition that makes the test disappear.make teststays green and prints a skip nobody reads. The same is true for a fixture the generator failed to write:t.Skipf("fixture %s absent"), though the canonical test does catch that one.Broken resample. Nothing in this path resamples.
PCMFromWAVrefuses anything that is not 16 kHz mono s16. Its own comment says so: "Refused at the seam rather than resampled". The fixtures arrive at 16 kHz becausegen-stt-fixtures.shruns ffmpeg. There is no resample step between the WAV andwhisper_full.Wrong language hint.
c.Langcomes out of the manifest and goes straight intoTranscribeReq. The test hardcodes the correct hint for each case, so nothing about how mavsttd chooses a language is exercised.What the test really covers is the model plus the silence gate. That is worth having. Say that in the comment instead.
2.
looseWordMatchlets a different word satisfy a keywordThe prefix rule is
n = len(want)-1for words of 6 runes or fewer. The comment argues the short-word case is safe because "час" cannot pass for "часть", which is true. The 4-rune case is not.воды→n = 3→ prefixвод.водкаmatches. So theru_factkeyword assertion passes if whisper hears "выпил водки".disk→n = 3→ prefixdis.distance,displayanddiscussall match theen_actkeyword.TestMissingKeywordsonly tests the 3-rune boundary and the inflection case it was designed for. Add{"воды", "водка"}to it and it fails. Raise the floor to four retained runes. Or cap how much longer the hypothesis word may be than the keyword.3. The spoken text lives in two files with nothing tying them together
gen-stt-fixtures.shhardcodes"Отметь, что я выпил воды."andgolden_v1.jsonseparately carries"отметь что я выпил воды". Change the script line, runmake stt-fixtures, and the manifest is now a reference for audio that no longer exists. WER 0.34 on a five-word reference tolerates one wrong word. A small edit drifts silently. A large one fails with a confusing diff.Have the script read
cases[].textout ofgolden_v1.jsonwithjqand synthesise from that. One source of truth, and the punctuation the script adds stops mattering becausenormalizeTranscriptstrips it anyway.4.
max_weris loose enough to pass a real regressionEvery case is 0.34 against references of five to eight words. That is one or two wrong words. Clean piper speech through ggml-small should score at or near 0. The ceiling leaves most of the range unguarded. The per-case field holds the same number four times.
Record the WER each case measures today in the manifest. Set the ceiling just above it. A model swap then shows up as a diff to a number rather than as silence. That is what the per-case ceiling was for.
Smaller notes
pcmToF32in the test duplicates the identical conversion inwhisper_handler.go. The canonical test says the fixture "must clear mavsttd's own silence gate", but it feedsgateReasonits own copy of the conversion. A regression in the production loop, say a/32767divisor, leaves the assertion green. Export the daemon's conversion and call it.make test-stt-goldenuses-run TestGolden, which also picks upTestGoldenFixturesAreCanonical. That is probably what you want, but the Makefile comment names onlyTestGoldenAudioTranscription.max_werbut not fortextorlang. A case with an emptytextmakeswordErrorRatetake itslen(ref) == 0branch and return 1 for every hypothesis. The WER assertion then fires with no useful message.Landed on master. The stack was one linear chain, so #84 carried every commit from #50 up, and master now contains this branch in full. Merging this PR on its own is an empty diff, so it is closed rather than merged. The review findings for it were fixed in the 2026-08-01 pass and are on master as commits on the stack tip, not on this branch.
Pull request closed