Stop mavwaked from hearing itself, and add barge-in (#287)

Playback was `go playAudio(reply)` — fire and forget, nobody holding the
process handle. Two audible consequences fell out of that.

She answered herself. The capture loop kept feeding the VAD while the
speaker was running, so her own reply came back in through the mic,
tripped the VAD, and was shipped to the daemon as a fresh command. There
is no acoustic echo canceller in this pipeline, so the fix is
half-duplex: while she is speaking, the capture side is muted. That part
is unconditional — it repairs a defect, it is not a new capability.

And talking over her did nothing, because there was no handle to cancel.
-barge-in now cuts playback when sustained energy clears a room-tuned
threshold (-barge-in-rms, default 0.12 normalised, over -barge-in-frames
consecutive frames, default 5). It is off by default: without an echo
canceller the only way to tell "he is talking over her" from "the mic is
hearing her" is that he is much louder, and how much louder depends on
where the mic sits.

The frame decision moved out of main.go into session.feed, behind a
player and an utteranceSender interface, so all of it is testable with
no mic, no speaker and no daemon. Nine tests cover the self-hearing
case, the off-by-default case, the consecutive-frame requirement,
speaker-leak-level audio not triggering, capturing the interrupting
utterance after a cut, and failed round-trips not starting playback.

The other seven items on #287 (partial STT, per-segment retry, mic
profiles, noise-floor calibration, short-response-while-speaking) are
untouched and stay on the task.
This commit is contained in:
kami
2026-08-01 05:36:13 +04:00
parent 62cc072f8c
commit fed33a4e16
6 changed files with 619 additions and 92 deletions
+9
View File
@@ -23,6 +23,15 @@ const (
defaultSilenceMs = 800 // silence hold before declaring end-of-utterance
defaultMaxMs = 10000 // cap single utterance at 10s
defaultMinRMS = 0.01 // RMS floor (same as mavsttd)
// Barge-in thresholds. Only used when -barge-in is passed. The RMS is
// x10000 like -min-rms, and sits an order of magnitude above the VAD's
// own floor on purpose: with no acoustic echo canceller, a frame only
// counts as "he is talking over her" if it is far louder than what the
// speaker leaks back into the mic. 5 frames is 150ms — long enough that
// a door or a cough does not cut her off mid-sentence.
defaultBargeRMS = 1200 // 0.12 normalised RMS
defaultBargeFrames = 5
)
// frameSamples — samples per 30ms frame at 16kHz.