mavwaked has no wake word, only an energy VAD — add silero-vad and a keyword gate #220
@@ -21,11 +21,15 @@ import (
|
||||
"sync"
|
||||
)
|
||||
|
||||
// defaultWakeThreshold — score above which the keyword was said. Picked from
|
||||
// the false-accept rate on held-out Russian speech, not from accuracy: a miss
|
||||
// costs him a repeat, a false accept costs a turn nobody asked for. See
|
||||
// docs/evals for the wakes-per-hour this buys.
|
||||
const defaultWakeThreshold = 0.99
|
||||
// defaultWakeThreshold — score above which the keyword was said.
|
||||
//
|
||||
// Picked from the false-accept rate on held-out Russian speech, not from
|
||||
// accuracy: a miss costs him a repeat, a false accept costs a turn nobody
|
||||
// asked for. Over 65 minutes of Common Voice, 0.99 woke her three times and
|
||||
// 0.999 once, and the difference in recall was one render out of 126. So the
|
||||
// default is the strict one. `docs/evals/2026-08-09-wake-word.md` has both
|
||||
// tables.
|
||||
const defaultWakeThreshold = 0.999
|
||||
|
||||
// wakeWord is the streaming state around wakeModels. It is fed the same
|
||||
// capture frames the VAD sees and answers whether the keyword has just been
|
||||
|
||||
@@ -0,0 +1,104 @@
|
||||
# The "Мэйвен" wake word: what it hears and what it invents
|
||||
|
||||
*Measured 2026-08-09 on workpc and homesrv. V-487, stage two of two.*
|
||||
|
||||
Stage one gave mavwaked silero-vad, which answers "is this frame speech".
|
||||
Nothing answered "was this said to her", so every utterance near the
|
||||
microphone became a turn. SurfaceVoice caps acts at L0, which made that safe
|
||||
rather than expensive. L0 does not cap reading, so the room could still hear
|
||||
his facts read back.
|
||||
|
||||
The keyword is "Мэйвен". openWakeWord's two frozen feature models do the
|
||||
hearing and a 100KB head trained here draws the boundary. It runs on CPU
|
||||
beside silero and never touches the GPU.
|
||||
|
||||
## Why a per-window accuracy is not a number anyone can act on
|
||||
|
||||
The gate scores every 80ms. A 1.7% false-accept rate per window sounds small
|
||||
and means a wake every few seconds. The useful question is how many times an hour
|
||||
it wakes on speech that was not the keyword. So every table below counts
|
||||
threshold crossings over whole clips and divides by the audio duration.
|
||||
|
||||
A crossing, not a window above the threshold. A keyword held high for half a
|
||||
second is one wake, not six.
|
||||
|
||||
## The data
|
||||
|
||||
Positives are 600 silero TTS renders of three stressings of the keyword, six
|
||||
speakers, ten trailing phrases, augmented eight ways each. Hard negatives are
|
||||
560 renders of confusable Russian words. Real speech is Common Voice ru and Golos.
|
||||
The 74257 Common Voice clips were already on workpc from the CrisperWhisper
|
||||
work. The 200 Golos clips came from the CW2 WER eval.
|
||||
|
||||
Splits are by source file. Augmented copies of one render on both sides of a
|
||||
split would measure memorisation.
|
||||
|
||||
Golos was never trained on at any stage, so it answers the harder question:
|
||||
does this survive a change of speakers and rooms.
|
||||
|
||||
## Three heads
|
||||
|
||||
Each row is a full retrain. The false-accept column is 8.89 hours of Common
|
||||
Voice that no stage of training had seen.
|
||||
|
||||
| trained on | recall (window) | false wakes/hour @0.99 |
|
||||
|---|---|---|
|
||||
| TTS + 13.7 min of Golos | 0.869 | not measurable |
|
||||
| + 4000 Common Voice clips | 0.836 | 21.9 |
|
||||
| + 3837 mined hard negatives | 0.784 | 4.2 |
|
||||
| + 753 more mined | 0.810 | 3.4 |
|
||||
|
||||
The first row is why the second exists. Thirteen minutes of held-out speech
|
||||
cannot measure a rate for a gate that scores twelve times a second. A head
|
||||
trained only against TTS learns to tell TTS from not-TTS.
|
||||
|
||||
Mining is the whole story after that. Random negatives teach the head what
|
||||
most speech sounds like. They do not teach it the few syllable sequences that
|
||||
score high, because 4000 clips barely contain them. So the current head was
|
||||
run over 20000 fresh clips, keeping every window it scored above 0.05. That
|
||||
found 3837 windows in 855512. Repeating those ten times in the next training
|
||||
run cut the rate five-fold.
|
||||
|
||||
The second round found 753 in 852240, a fifth of the yield, and bought a
|
||||
further 20%. It also recovered recall, which the first round had cost. Whether
|
||||
a third round is worth 25 minutes of workpc is untested.
|
||||
|
||||
## Where the threshold came from
|
||||
|
||||
Both columns are held out. Positives are the 126 renders in the test split.
|
||||
Speech is 65.1 minutes of Common Voice, disjoint from every training and
|
||||
mining pool. Both were run through the built `mavwaked` binary reading PCM from a
|
||||
file, not through the python that trained the head.
|
||||
|
||||
| threshold | renders shipped | false wakes/hour |
|
||||
|---|---|---|
|
||||
| 0.99 | 116 / 126 | 2.8 |
|
||||
| 0.999 | 115 / 126 | 0.9 |
|
||||
|
||||
One render against a third of the false wakes. `defaultWakeThreshold` is
|
||||
0.999.
|
||||
|
||||
Golos disagrees. It gave 2 wakes in 14 minutes at every threshold, which is
|
||||
8.7 per hour. Two events is not a rate. What it does say is that a handful of real utterances score above 0.999
|
||||
and no threshold will move them.
|
||||
|
||||
## What it costs him
|
||||
|
||||
Ten of the 126 held-out renders were heard and still dropped, and every one
|
||||
was an utterance shorter than 1.32s. The head scores 16 embeddings, or 1.28s of
|
||||
audio. The score therefore peaks up to a second after a short keyword ends.
|
||||
By then the VAD has closed the utterance and dispatch has already asked.
|
||||
|
||||
Real commands are "Мэйвен, <request>" and run past two seconds, which gives
|
||||
the head the whole request to peak during. A bare "Мэйвен" with nothing after
|
||||
it is the case that fails. One fix would hold an ignored utterance for a grace
|
||||
period and ship it if the keyword lands late. It is not built.
|
||||
|
||||
## What was not measured
|
||||
|
||||
No room recordings. Every negative above is a clean corpus clip. This gate
|
||||
will live among a television, a fan and the far side of a kitchen. None of
|
||||
those are in these numbers.
|
||||
|
||||
No measurement of him. Training on his voice means copying his transcripts off
|
||||
homesrv, which is his call and has not been asked.
|
||||
Reference in New Issue
Block a user