Stop mavwaked from hearing itself, and add barge-in (#287)
Playback was `go playAudio(reply)` — fire and forget, nobody holding the process handle. Two audible consequences fell out of that. She answered herself. The capture loop kept feeding the VAD while the speaker was running, so her own reply came back in through the mic, tripped the VAD, and was shipped to the daemon as a fresh command. There is no acoustic echo canceller in this pipeline, so the fix is half-duplex: while she is speaking, the capture side is muted. That part is unconditional — it repairs a defect, it is not a new capability. And talking over her did nothing, because there was no handle to cancel. -barge-in now cuts playback when sustained energy clears a room-tuned threshold (-barge-in-rms, default 0.12 normalised, over -barge-in-frames consecutive frames, default 5). It is off by default: without an echo canceller the only way to tell "he is talking over her" from "the mic is hearing her" is that he is much louder, and how much louder depends on where the mic sits. The frame decision moved out of main.go into session.feed, behind a player and an utteranceSender interface, so all of it is testable with no mic, no speaker and no daemon. Nine tests cover the self-hearing case, the off-by-default case, the consecutive-frame requirement, speaker-leak-level audio not triggering, capturing the interrupting utterance after a cut, and failed round-trips not starting playback. The other seven items on #287 (partial STT, per-segment retry, mic profiles, noise-floor calibration, short-response-while-speaking) are untouched and stay on the task.
This commit is contained in:
@@ -23,6 +23,15 @@ const (
|
||||
defaultSilenceMs = 800 // silence hold before declaring end-of-utterance
|
||||
defaultMaxMs = 10000 // cap single utterance at 10s
|
||||
defaultMinRMS = 0.01 // RMS floor (same as mavsttd)
|
||||
|
||||
// Barge-in thresholds. Only used when -barge-in is passed. The RMS is
|
||||
// x10000 like -min-rms, and sits an order of magnitude above the VAD's
|
||||
// own floor on purpose: with no acoustic echo canceller, a frame only
|
||||
// counts as "he is talking over her" if it is far louder than what the
|
||||
// speaker leaks back into the mic. 5 frames is 150ms — long enough that
|
||||
// a door or a cough does not cut her off mid-sentence.
|
||||
defaultBargeRMS = 1200 // 0.12 normalised RMS
|
||||
defaultBargeFrames = 5
|
||||
)
|
||||
|
||||
// frameSamples — samples per 30ms frame at 16kHz.
|
||||
|
||||
Reference in New Issue
Block a user