silero-vad replaces the energy threshold when -vad-model points at it.
Everything after the speech decision is the same state machine: the speech
hold, the silence hold, the length cap and the utterance buffer.
The model window is 512 samples and the capture frame is 480, so silero.go
re-chunks across frames. main.go claimed the two matched, which was true of
silero v4.
Stage two, the wake word, is not here. It needs a Russian keyword model that
does not exist yet.
Nothing reads the microphone while Send is in flight, so the audio piles
up in arecord's pipe and arrives in a burst the moment dispatch returns.
A round-trip is p50 2.7s through the LLM router, which is about 90
frames of room, of him finishing his sentence, of the television.
The old code reset the VAD on the reply path only, and for a reason that
was not true: the comment said the VAD had been accumulating during the
round-trip, when its state is exactly what Feed left it as. The two
paths with no reset are the ones that mattered, because neither starts
playback and so neither is covered by the half-duplex gate. A text-only
turn fed the whole backlog into the VAD, and a Send error did the same
on every failed turn, so a dead socket drove a retry loop off backlog
alone.
The backlog was scored for barge-in too. Five frames delivered in
microseconds cut her off with audio recorded before she started
speaking, which is the opposite of what the five-frame guard is for.
Both are fixed by the same mechanism: measure the wall time the
round-trip took, convert it to frames, and discard that many before
anything looks at them.
Barge-in also threw away the 150ms that proved he was talking. The VAD
started from the next frame, so the first word of a short interruption
was clipped before whisper saw it. Those frames are kept in a small ring
and replayed after the reset.
A stuck aplay was worse than before this feature existed. Playing()
gates all capture, so a wedged child made her deaf rather than silent,
for the full 30s ceiling inherited from the fire-and-forget version. The
mute window is bounded by the reply's own duration plus a margin now.
Three smaller ones. "-barge-in -barge-in-rms 0" logged "barge-in on" and
then did nothing. The sent counter incremented before the error check,
so failed round-trips counted as shipped. And the threshold the operator
has to guess is now reported: mavwaked logs the mean energy of the
frames it suppressed while speaking, so he can set it from data.
Found in review of #76.
Playback was `go playAudio(reply)` — fire and forget, nobody holding the
process handle. Two audible consequences fell out of that.
She answered herself. The capture loop kept feeding the VAD while the
speaker was running, so her own reply came back in through the mic,
tripped the VAD, and was shipped to the daemon as a fresh command. There
is no acoustic echo canceller in this pipeline, so the fix is
half-duplex: while she is speaking, the capture side is muted. That part
is unconditional — it repairs a defect, it is not a new capability.
And talking over her did nothing, because there was no handle to cancel.
-barge-in now cuts playback when sustained energy clears a room-tuned
threshold (-barge-in-rms, default 0.12 normalised, over -barge-in-frames
consecutive frames, default 5). It is off by default: without an echo
canceller the only way to tell "he is talking over her" from "the mic is
hearing her" is that he is much louder, and how much louder depends on
where the mic sits.
The frame decision moved out of main.go into session.feed, behind a
player and an utteranceSender interface, so all of it is testable with
no mic, no speaker and no daemon. Nine tests cover the self-hearing
case, the off-by-default case, the consecutive-frame requirement,
speaker-leak-level audio not triggering, capturing the interrupting
utterance after a cut, and failed round-trips not starting playback.
The other seven items on #287 (partial STT, per-segment retry, mic
profiles, noise-floor calibration, short-response-while-speaking) are
untouched and stay on the task.
New cmd/mavwaked — always-on voice listening client that:
- Captures PCM from arecord subprocess (16kHz mono int16)
- Runs energy-based VAD in 30ms windows (RMS threshold, adaptive floor)
- Buffers utterances (300ms min speech, 800ms silence end, 10s max)
- Sends complete utterances as PushToTalk with Surface=SurfaceVoice (L0)
- Plays reply audio through aplay subprocess
- No new CGo/onnxruntime deps — pure Go
- 10 VAD tests with -race (speech detect, silence, max duration, reset, adaptive floor)
- Makefile build-waked target + Dockerfile integration + alsa-utils runtime dep