Hear speech instead of loudness in mavwaked (V-487)
silero-vad replaces the energy threshold when -vad-model points at it. Everything after the speech decision is the same state machine: the speech hold, the silence hold, the length cap and the utterance buffer. The model window is 512 samples and the capture frame is 480, so silero.go re-chunks across frames. main.go claimed the two matched, which was true of silero v4. Stage two, the wake word, is not here. It needs a Russian keyword model that does not exist yet.
This commit is contained in:
@@ -0,0 +1,63 @@
|
||||
# silero-vad against the energy threshold in mavwaked
|
||||
|
||||
*Measured 2026-08-09 on homesrv. V-487, stage one of two.*
|
||||
|
||||
mavwaked decided an utterance had started by comparing frame energy to an
|
||||
adaptive floor. That answers "is this frame loud". A fan, a door and a
|
||||
television are all loud, and every utterance mavwaked accepts becomes a turn.
|
||||
|
||||
silero-vad answers "is this frame speech". It is 2.3MB of ONNX and it replaces
|
||||
the comparison and nothing else. The speech hold, the silence hold, the length
|
||||
cap and the utterance buffer are the same state machine either way.
|
||||
|
||||
## What it declines
|
||||
|
||||
Speech is the four piper fixtures `mavsttd` already scores against, so nothing
|
||||
of the owner's voice is committed. Non-speech is white noise at the same RMS as the clip beside it. That is the
|
||||
cheapest thing that fools an energy floor.
|
||||
|
||||
| clip | silero, speech frames on speech | silero on noise | energy on noise |
|
||||
|---|---|---|---|
|
||||
| ru_fact.wav | 59 | 0 | 68 |
|
||||
| ru_query.wav | 69 | 0 | 79 |
|
||||
| ru_reminder.wav | 80 | 0 | 89 |
|
||||
| en_act.wav | 90 | 0 | 99 |
|
||||
|
||||
The energy threshold accepts every noise clip as a complete utterance. Silero
|
||||
calls not one frame of any of them speech, and still hears all four spoken
|
||||
clips. `TestSileroHearsSpeechAndDeclinesNoise` is that table.
|
||||
|
||||
White noise is a floor, not a proof. It says nothing about a television, which
|
||||
is speech, or about a fan, which is narrowband. Those need room recordings and
|
||||
this box has none.
|
||||
|
||||
## What it costs
|
||||
|
||||
`BenchmarkSileroFrame` on the homesrv laptop (Ryzen 5 5600U), one 30ms frame
|
||||
through the model including the re-chunking:
|
||||
|
||||
509µs per frame
|
||||
|
||||
That is 1.7% of one core, on the slower of the two machines. The detector runs
|
||||
on the workstation beside the microphone, never on the GPU. This number is what
|
||||
says it does not need one.
|
||||
|
||||
## The window is 512 samples, not 480
|
||||
|
||||
`cmd/mavwaked/main.go` claimed the frame contract matched silero's input
|
||||
exactly. That was true of silero v4. Version 5 takes exactly 512 samples at 16kHz, plus 64 samples of context from
|
||||
the previous window. So `sileroVAD` buffers across capture frames, and a frame
|
||||
completing no window inherits the previous probability. `TestSileroRechunksAcrossFrames` pins it.
|
||||
|
||||
## Still an energy gate by default
|
||||
|
||||
`-vad-model` is empty in the code default, so a deployment that does not pass
|
||||
it runs exactly what shipped before. Barge-in is untouched and deliberately so. It reads frame energy while she is
|
||||
speaking, which is a different question from whether the frame is speech.
|
||||
|
||||
## Not done here
|
||||
|
||||
The wake word. This is stage one of the two V-487 asks for. The second needs a
|
||||
keyword model that does not exist yet. The pretrained openWakeWord keywords are
|
||||
English, and a Russian one has to be trained. Until then anything spoken near
|
||||
the microphone still becomes a turn. It is now merely required to be speech.
|
||||
Reference in New Issue
Block a user