Compare commits

...

8 Commits

Author SHA1 Message Date
claude 999a5ad562 Record what Kiwix returns and why the gate is not one (V-668)
gofmt on cmd/mavwaked/silero.go came in with a99932b and blocked make test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
2026-08-09 10:46:15 +04:00
claude 2ea39a3d41 Search Kiwix for the topic, not the whole sentence (V-668)
Kiwix ranks by keyword overlap, which the package doc has said since it was
written: "why is the sky blue" finds a TV episode. queryKiwix sent the whole
Russian sentence, because the verbatim path added by V-508 skips the rewriter
that would have reduced it.

Measured against the Russian ZIM on 2026-08-09, over eight questions. Four
reach the right article where they did not: TCP was "Перехват TCP-соединения"
and is TCP, фотосинтез was "C4-фотосинтез" and is Фотосинтез, Линус Торвальдс
was "Tux", and "кто написал Войну и мир" was "Радуйся, мир (Доктор Кто)".
Two were already right and stay right. Two are still wrong and were wrong
before. Nothing regressed.

kiwix.Topic drops the narrative request, the interrogative and a verb behind
one, and keeps everything else. A word it cannot classify is more likely the
topic than noise. TitlePath tries the exact article first, since a ZIM is
addressable by title and a wrong title is a 404.

The gate this task set out to build does not exist. Query-to-passage cosine
scored 0.79-0.91 on answerable questions and 0.75-0.84 on unanswerable ones,
and the sets overlap. The wrong TCP article scored 0.8653, above five of six
unanswerable rows. e5 measures topic, not whether the passage answers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
2026-08-09 10:42:45 +04:00
claude 8fb6f2154d Only a literal pattern may take the personal boundary off a turn (#211) 2026-08-09 00:01:51 +02:00
claude c938148619 Only a literal pattern may take the personal boundary off a turn (V-666)
Naming a destination takes the guessing query sources off a turn, and the
personal boundary is one of them. Every other guesser costs an answer when it
is wrongly dropped. This one costs the rule that a question about him never
reaches an upstream engine.

Three deciders name a destination now and two of them infer it: the routing
heads and the resident model. Decision.SourceAnchored says a stage 0 grammar
read the words instead. queryWalk honours it for the source marked
boundary: true and for no other, so the rest of the table is unchanged.

Owner's call of 2026-08-09.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013ptwopxyo3Z2kwFckHkLvN
2026-08-09 01:56:46 +04:00
claude 2512d686a1 Give mavwaked a speech model instead of an energy threshold (#210) 2026-08-08 23:49:09 +02:00
claude f44abcc526 Measure what silero declines that the threshold accepts (V-487)
Speech is the four piper fixtures mavsttd already scores against, so nothing
of the owner's voice is committed. Non-speech is white noise at the same RMS
as the clip beside it.

Silero calls 0 noise frames speech where the energy threshold calls 68 to 99,
and hears all four spoken clips. 509us per 30ms frame, 1.7% of one core on
the slower machine.

White noise is a floor and not a proof. It says nothing about a television,
which is speech, or a fan, which is narrowband.
2026-08-09 01:43:41 +04:00
claude a99932b427 Hear speech instead of loudness in mavwaked (V-487)
silero-vad replaces the energy threshold when -vad-model points at it.
Everything after the speech decision is the same state machine: the speech
hold, the silence hold, the length cap and the utterance buffer.

The model window is 512 samples and the capture frame is 480, so silero.go
re-chunks across frames. main.go claimed the two matched, which was true of
silero v4.

Stage two, the wake word, is not here. It needs a Russian keyword model that
does not exist yet.
2026-08-09 01:43:32 +04:00
claude 6d5801bb1f Merge pull request 'Move STT and TTS to the workstation, where the microphone already is' (#209) from task/486-deploy-the-workstation-transcriber into master 2026-08-08 23:26:31 +02:00
17 changed files with 924 additions and 34 deletions
+2
View File
@@ -70,3 +70,5 @@ coverage.out
# root .env — MAVEN_AMBIENT_TOKEN and friends, same class as deploy/telegram.env
.env
# silero-vad, downloaded (see AGENTS.md)
/models/vad/
+16
View File
@@ -95,6 +95,22 @@ model: the code puts `query: ` in front of a question and `passage: ` in front
of a stored note, which is how e5 was trained. The quantized file is the one
that is downloaded, deployed and measured.
## Voice activity model for mavwaked
`mavwaked` decides an utterance has started with silero-vad when `-vad-model`
points at it, and with an energy threshold when it does not. The model is 2.3MB
and is not committed:
```sh
mkdir -p models/vad
curl -sL -o models/vad/silero_vad.onnx \
https://github.com/snakers4/silero-vad/raw/master/src/silero_vad/data/silero_vad.onnx
```
It needs the same `libonnxruntime.so` the embedder needs, passed as `-onnx-lib`
or read from `MAVEN_ONNX_LIB`. The measurement is
`docs/evals/2026-08-09-silero-vad.md`, and the tests skip without the file.
**Also need ONNX Runtime** (`libonnxruntime.so`):
```sh
+30 -8
View File
@@ -510,13 +510,21 @@ they look rather than guess.
**The personal boundary is the one exception and it is deliberate.** It guesses,
so naming `SourceWorld` drops it. That is what stops it answering "кто такой
Линус Торвальдс?" with "не нашла у тебя такой записи", which it did on
2026-08-07. The cost is that a destination a model wrote can now take the
boundary off a turn. A question about him that the model calls `world` reaches
SearXNG, where today the boundary stops it. Only the utterance leaves the box,
never his notes or history, so this widens what is asked and not what is sent.
`TestNamingRecallKeepsTheBoundary` pins the other half: naming `SourceRecall`
keeps the boundary in front of the world. Whether a model may drop it at all is
the owner's call and has not been made.
2026-08-07. `TestNamingRecallKeepsTheBoundary` pins the other half: naming
`SourceRecall` keeps the boundary in front of the world.
**Only a stage 0 grammar may drop it** (owner's call, 09-08-2026, V-666). The
question of who is allowed to was open until then. Three deciders name a
destination and two of them infer it: the routing heads and the resident model.
An inferred `SourceWorld` on a question about him would reach SearXNG, and that
widens what is asked rather than costing a local answer. So `Decision.SourceAnchored`
carries the provenance. It is a field and not `Stage == 0`. Stage 0 also means
confidence 1.0 and an anchored claim band, and one of those could stop implying
the others. `queryWalk` reads it for the source marked `boundary: true` and for
no other. So every other guesser still comes off the turn, whoever named the
destination. `TestOnlyAGrammarMayDropTheBoundary` pins both directions.
`definitionQueryPattern` claims "кто такой X", so the 2026-08-07 case is still
anchored and still answered.
What comes out is only the sources that **guess**. Those decide a turn is theirs by
cosine against frozen seeds, then answer whatever they claimed. They hold no table
@@ -667,11 +675,25 @@ world questions, so she needs to read external sources. What replaces it:
`wikipedia_ru_all_maxi_2026-02` verbatim** through `kiwix.book_ru`. The RU→EN rewriter
is the workaround for an English book and is skipped there. Kiwix catalog names come
from the filename, not the `<name>` field.
**That verbatim path sent the whole sentence to a keyword engine until 09-08-2026**
(V-668, `docs/evals/2026-08-09-kiwix-topic-retrieval.md`). Kiwix ranks by keyword
overlap, so the question words outrank the one word naming the article. "что такое TCP"
returned "Перехват TCP-соединения". "кто написал Войну и мир" returned an episode of
Doctor Who. `kiwix.Topic` drops the narrative request, the interrogative and a verb
behind one. `kiwix.TitlePath` tries the exact article first, since a ZIM is addressable
by title and a wrong title is a 404. Four of eight questions reach the right article
where they did not, two were already right, and nothing regressed. Both apply on the
verbatim path alone. The rewriter already reduces a question, and reducing twice takes
the topic off its input.
`Response.Empty()` is the whole gate and there is no quality threshold in front of it:
the three signals one could read were measured on 2026-08-05 and none of them separate a
real question from an invented one. Token overlap would cost "столица Франции" its
answer, because the answer is Париж and that word is not in the question. See
`docs/evals/2026-08-05-search-quality-signals.md` (V-539). **Which query source claimed
`docs/evals/2026-08-05-search-quality-signals.md` (V-539). **The embedder is not a
fourth signal**, measured 2026-08-09 (V-668). Query-to-passage cosine scores 0.79 to
0.91 on answerable questions and 0.75 to 0.84 on unanswerable ones, and the sets
overlap. The wrong TCP article scored 0.8653, above five of six unanswerable rows. It
measures topic and not whether the passage answers, so no threshold splits them. **Which query source claimed
a turn is readable on `/chat`** as a badge beside the reply, carried on
`ipc.ChatReply.Source` and noted by `noteQuerySource` in `cmd/mavend/querysource.go`. It
rides the context, so `handleText` keeps the one string signature the mic, telegram and
+47 -8
View File
@@ -13,6 +13,7 @@ import (
"github.com/kami/maven/internal/crawl"
"github.com/kami/maven/internal/decision"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/kiwix"
"github.com/kami/maven/internal/memory"
"github.com/kami/maven/internal/morning"
"github.com/kami/maven/internal/phraser"
@@ -79,6 +80,15 @@ type querySource struct {
// not named do not get to try. The lookups still run, because a named
// destination is evidence and not a promise.
guesses bool
// boundary — dropping this source widens what leaves the box, so only a
// literal pattern may do it (V-666, owner's call of 2026-08-09).
//
// Every other guesser costs an answer when it is wrongly taken off a turn.
// This one costs the rule that a question about him never reaches an
// upstream engine. A grammar read the words to name a destination. A model
// and a softmax both inferred one, and neither may spend that.
boundary bool
}
// querySources is the ordered chain actionQuery walks; first source to claim
@@ -156,7 +166,7 @@ var querySources = []querySource{
// below answers from the world's. A question about him that got this far
// has no answer in his data, and no outside source can supply one, so this
// stops the walk rather than let the encyclopedia and the model guess.
{name: "personal", answer: (*reactiveHandler).queryPersonal, dest: router.SourceRecall, guesses: true},
{name: "personal", answer: (*reactiveHandler).queryPersonal, dest: router.SourceRecall, guesses: true, boundary: true},
// The world, read live. Owner's ruling of 2026-08-02: a metasearch hit beats
// a frozen ZIM, so SearXNG asks before Kiwix does. Nothing of his is at
// stake by this point — the boundary above already stopped every question
@@ -198,12 +208,17 @@ var querySources = []querySource{
// No destination named ⇒ the table exactly as written, which is what shipped
// before the field existed. That is the floor. The classifier arm names
// nothing, so a box whose model is down routes queries the way it always did.
func queryWalk(dest router.Source) (walk, skipped []querySource) {
// The personal boundary is the one exception, and anchored is what buys it
// (V-666). A grammar matched a literal pattern to name the destination. The
// routing heads and the resident model inferred one, and an inferred SourceWorld
// takes the boundary off a question about him. That widens what is asked
// upstream rather than costing a local answer, so those two keep it.
func queryWalk(dest router.Source, anchored bool) (walk, skipped []querySource) {
if dest == router.SourceUnknown {
return querySources, nil
}
for _, s := range querySources {
if s.guesses && s.dest != dest {
if s.guesses && s.dest != dest && (anchored || !s.boundary) {
skipped = append(skipped, s)
continue
}
@@ -219,7 +234,7 @@ func (h *reactiveHandler) actionQuery(ctx context.Context, dec router.Decision)
// (V-564). Finish names everyone below the winner.
decision.Expect(ctx, decision.StageQuery, querySourceNames())
rec := decision.From(ctx)
walk, skipped := queryWalk(dec.Source)
walk, skipped := queryWalk(dec.Source, dec.SourceAnchored)
for _, src := range skipped {
rec.Note(decision.Claim{
Stage: decision.StageQuery, Claimant: src.name, Outcome: decision.NeverAsked,
@@ -912,6 +927,26 @@ func (h *reactiveHandler) queryKiwix(ctx context.Context, t *queryTurn) (string,
}
}
// The topic, not the sentence (V-668). Kiwix ranks by keyword overlap, so
// the question words outrank the one word that names the article: measured
// on 2026-08-09, "что такое TCP" returns "Перехват TCP-соединения" and
// "TCP" returns TCP. Only the verbatim path needs this. The rewriter
// already reduces a question to English keywords, and reducing twice would
// take the topic off the input it reads.
if verbatim {
if topic := kiwix.Topic(pattern); topic != "" {
// The article named exactly, before any ranking runs. A ZIM is
// addressable by title and a wrong title is a 404, so this either
// answers or costs one request that says nothing.
page, err := h.kiwix.client.Article(ctxK, kiwix.TitlePath(book, topic), h.kiwix.runes)
if err == nil && page.Text != "" {
log.Printf("voice: kiwix: %q in %q → title hit %q", topic, book, page.Title)
return h.kiwixReply(ctx, t, page.Title, page.Text)
}
pattern = topic
}
}
hits, err := h.kiwix.client.Search(ctxK, pattern, book, h.kiwix.max)
if err != nil {
log.Printf("voice: kiwix: search %q: %v", pattern, err)
@@ -944,14 +979,18 @@ func (h *reactiveHandler) queryKiwix(ctx context.Context, t *queryTurn) (string,
}
page = crawl.Page{Title: top.Title, Text: top.Snippet}
}
// Handed over the same way a note or a page is: context for the question he
// asked, not something to recite.
snippet := top.Title + "\n" + crawl.TrimRunes(page.Text, h.kiwix.runes)
return h.kiwixReply(ctx, t, top.Title, page.Text)
}
// kiwixReply hands one article over the same way a note or a page is handed
// over: context for the question he asked, not something to recite.
func (h *reactiveHandler) kiwixReply(ctx context.Context, t *queryTurn, title, text string) (string, bool) {
snippet := title + "\n" + crawl.TrimRunes(text, h.kiwix.runes)
reply := h.phraseSource(ctx, "kiwix", t.dec.Utterance, []string{snippet})
if reply == "" {
// No phraser, or it failed. Read back the best hit rather than pretend
// the search did not happen.
return readBack(top.Title + " — " + page.Text), true
return readBack(title + " — " + text), true
}
return reply, true
}
+32 -10
View File
@@ -10,7 +10,7 @@ import (
// whose model is down names nothing, and naming nothing has to walk the chain
// the way it walked before the field existed.
func TestNoDestinationWalksTheWholeChain(t *testing.T) {
walk, skipped := queryWalk(router.SourceUnknown)
walk, skipped := queryWalk(router.SourceUnknown, false)
if len(skipped) != 0 {
t.Errorf("skipped %d sources with no destination named, want none", len(skipped))
}
@@ -32,15 +32,16 @@ func TestANamedDestinationSilencesTheOtherGuessers(t *testing.T) {
dest router.Source
utterance string
silenced string
anchored bool // a stage 0 grammar named the destination
}{
{router.SourceWorld, "что такое TCP?", "weather"},
{router.SourceWorld, "сколько будет 17 на 23?", "weather"},
{router.SourceWorld, "кто такой Линус Торвальдс?", "personal"},
{router.SourceRecall, "какой у меня любимый язык?", "feeds"},
{router.SourceCalendar, "что в календаре на завтра?", "weather"},
{router.SourceWorld, "что такое TCP?", "weather", true},
{router.SourceWorld, "сколько будет 17 на 23?", "weather", true},
{router.SourceWorld, "кто такой Линус Торвальдс?", "personal", true},
{router.SourceRecall, "какой у меня любимый язык?", "feeds", false},
{router.SourceCalendar, "что в календаре на завтра?", "weather", true},
}
for _, c := range cases {
walk, skipped := queryWalk(c.dest)
walk, skipped := queryWalk(c.dest, c.anchored)
if inWalk(walk, c.silenced) {
t.Errorf("%q named %q: %q is still asked", c.utterance, c.dest, c.silenced)
}
@@ -56,7 +57,7 @@ func TestANamedDestinationSilencesTheOtherGuessers(t *testing.T) {
// data first, then the world", and a destination a model wrote must not be able
// to reverse it.
func TestNamingTheWorldStillReadsHisDataFirst(t *testing.T) {
walk, _ := queryWalk(router.SourceWorld)
walk, _ := queryWalk(router.SourceWorld, true)
for _, look := range []string{"fact-by-key", "embed", "memory", "notes"} {
if !inWalk(walk, look) {
t.Errorf("%q was dropped; only the sources that guess may be dropped", look)
@@ -74,7 +75,7 @@ func TestNamingTheWorldStillReadsHisDataFirst(t *testing.T) {
// makes "какой у меня любимый язык?" answer "не нашла у тебя такой записи"
// rather than reaching SearXNG once nothing local had it.
func TestNamingRecallKeepsTheBoundary(t *testing.T) {
walk, _ := queryWalk(router.SourceRecall)
walk, _ := queryWalk(router.SourceRecall, true)
if !inWalk(walk, "personal") {
t.Fatal("the personal boundary was skipped on a turn named for his own data")
}
@@ -83,12 +84,33 @@ func TestNamingRecallKeepsTheBoundary(t *testing.T) {
}
}
// The owner's call of 2026-08-09 (V-666): only a stage 0 grammar may take the
// personal boundary off a turn. The routing heads and the resident model both
// name a destination by inference, and an inferred SourceWorld would send a
// question about him upstream. Every other guesser still goes.
func TestOnlyAGrammarMayDropTheBoundary(t *testing.T) {
walk, skipped := queryWalk(router.SourceWorld, false)
if !inWalk(walk, "personal") {
t.Error("an inferred destination took the boundary off the turn")
}
if !inWalk(skipped, "weather") {
t.Error("weather is still asked; the rule covers the boundary alone")
}
if posOf(walk, "personal") > posOf(walk, "search") {
t.Error("the boundary no longer sits in front of the world")
}
if anchored, _ := queryWalk(router.SourceWorld, true); inWalk(anchored, "personal") {
t.Error(`a grammar named the world and the boundary stayed: ` +
`"кто такой Линус Торвальдс?" is answered "не нашла у тебя такой записи" again`)
}
}
// Whatever the destination, the walk is a subsequence of the table. Every
// comment on that table argues an order between two sources, and none of those
// reasons is about this field.
func TestTheWalkNeverReordersTheTable(t *testing.T) {
for _, dest := range append([]router.Source{router.SourceUnknown}, router.Sources...) {
walk, skipped := queryWalk(dest)
walk, skipped := queryWalk(dest, true)
if len(walk)+len(skipped) != len(querySources) {
t.Errorf("%q: %d walked + %d skipped, want %d", dest, len(walk), len(skipped), len(querySources))
}
+27 -7
View File
@@ -5,12 +5,17 @@
// is detected sends it as a PushToTalk frame to the voice server. The reply
// audio is played back through aplay(1).
//
// No wake-word model yet (MVP uses voice-activity-only trigger). The
// SurfaceVoice auth layer caps all commands at L0 (no destructive acts),
// making accidental triggers safe by design. A proper wake-word engine
// (openWakeWord / Silero VAD ONNX) is the planned upgrade — the VAD shape
// (30ms frames, 16kHz PCM) matches silero-vad's input interface exactly, so
// swapping energy-threshold for ONNX-inference is a local change in vad.go.
// Voice activity is silero-vad when -vad-model points at the graph, and an
// energy threshold when it does not. Silero declines noise the threshold
// accepts: 0 frames against 68 to 99 on the four fixtures, measured in
// docs/evals/2026-08-09-silero-vad.md. Note that the model window is 512
// samples and the capture frame is 480, so silero.go re-chunks. This comment
// used to say the two matched, which was true of silero v4.
//
// There is still no wake-word model, so anything spoken near the microphone
// becomes a turn (V-487 stage two). The SurfaceVoice auth layer caps all
// commands at L0 (no destructive acts), which is what makes an accidental
// trigger safe rather than expensive.
//
// While a reply is playing the capture side is muted (half-duplex): without
// it, Maven's own voice comes back in through the mic and she answers
@@ -73,6 +78,9 @@ func run(args []string) error {
bargeIn := flag.Bool("barge-in", false, "cut Maven off when he talks over her (needs a room-tuned -barge-in-rms)")
bargeRMS := flag.Int("barge-in-rms", defaultBargeRMS, "RMS x10000 a frame must clear to count as barge-in")
bargeFrames := flag.Int("barge-in-frames", defaultBargeFrames, "consecutive frames over -barge-in-rms before playback is cut")
vadModel := flag.String("vad-model", "", "silero-vad onnx file; empty runs the energy threshold instead")
vadThreshold := flag.Float64("vad-threshold", defaultSileroThreshold, "speech probability a frame must clear")
onnxLib := flag.String("onnx-lib", os.Getenv("MAVEN_ONNX_LIB"), "libonnxruntime.so, needed with -vad-model")
flag.CommandLine.Parse(args)
ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGINT, syscall.SIGTERM, syscall.SIGHUP)
@@ -82,8 +90,20 @@ func run(args []string) error {
vc := voice.Dial(*addr)
defer vc.Close()
// VAD engine.
// VAD engine. A model that will not load is logged and not fatal: the
// energy threshold is worse, and it is a great deal better than a
// listening client that refuses to start.
vad := NewVAD(*minRMS, *speechMs, *silenceMs, *maxMs)
if *vadModel != "" {
s, err := newSileroVAD(*vadModel, *onnxLib)
if err != nil {
log.Printf("mavwaked: silero unavailable, energy threshold unchanged: %v", err)
} else {
defer s.Close()
vad.UseSilero(s, *vadThreshold)
log.Printf("mavwaked: silero-vad from %s, threshold %.2f", *vadModel, *vadThreshold)
}
}
// Audio source.
var src io.ReadCloser
+171
View File
@@ -0,0 +1,171 @@
package main
// silero-vad, the speech detector that replaces the energy threshold (V-487).
//
// Why an energy threshold is not a voice activity detector. It answers "is
// this frame loud", and a fan, a door and a television are all loud. mavwaked
// sends every utterance it accepts to speech-to-text and then to the daemon,
// so a false trigger is a turn Maven takes on something nobody said to her.
// Silero answers "is this frame speech", which is the question.
//
// It is 2.3MB of ONNX and runs on one CPU core in real time. That is not an
// aside: this is the one model in the system that may never be offloaded or
// gated on GPU admission, because a wake path that waits on a card is not a
// wake path.
//
// Nil is a working value. Without -vad-model the daemon runs the energy VAD
// exactly as it did before this file existed.
import (
"fmt"
"sync"
ort "github.com/yalue/onnxruntime_go"
)
const (
// sileroWindow — samples per inference at 16kHz. The model is fixed at
// 512 and does not accept another size, which is why this file
// re-chunks rather than reusing the 480-sample capture frame. main.go
// used to claim the two matched; that was true of silero v4.
sileroWindow = 512
// sileroContext — samples of the previous window prepended to each
// inference, as the reference implementation does. Without it the first
// milliseconds of every window are judged with no history and speech
// onsets score low.
sileroContext = 64
// sileroState — the LSTM state carried between windows, [2][1][128].
sileroStateDim = 128
// defaultSileroThreshold — probability above which a window is speech.
// 0.5 is the reference default. Raising it costs speech onsets, which
// are the quietest part of an utterance.
defaultSileroThreshold = 0.5
)
// sileroVAD holds one ONNX session and the streaming state around it. It is
// fed 30ms capture frames and answers per frame, buffering across calls
// because 480 samples never line up with a 512-sample window.
type sileroVAD struct {
mu sync.Mutex
session *ort.DynamicAdvancedSession
pending []float32 // samples not yet part of a full window
context [sileroContext]float32 // tail of the previous window
state []float32 // [2][1][128], carried between windows
last float64 // most recent probability, held between windows
sr []int64
}
// newSileroVAD loads the graph. The ONNX environment is initialised here when
// nothing else has done it, because mavwaked has no embedder to do it first.
func newSileroVAD(modelPath, libPath string) (*sileroVAD, error) {
if !ort.IsInitialized() {
if libPath != "" {
ort.SetSharedLibraryPath(libPath)
}
if err := ort.InitializeEnvironment(); err != nil {
return nil, fmt.Errorf("silero: onnx runtime: %w", err)
}
}
s, err := ort.NewDynamicAdvancedSession(modelPath,
[]string{"input", "state", "sr"}, []string{"output", "stateN"}, nil)
if err != nil {
return nil, fmt.Errorf("silero: load %s: %w", modelPath, err)
}
return &sileroVAD{
session: s,
state: make([]float32, 2*sileroStateDim),
sr: []int64{16000},
}, nil
}
// Speech reports whether the frame carries speech, and the probability behind
// that answer. A frame that completes no window inherits the previous
// probability, so the caller sees one answer per frame either way.
func (s *sileroVAD) Speech(frame []int16, threshold float64) (bool, float64) {
s.mu.Lock()
defer s.mu.Unlock()
for _, v := range frame {
s.pending = append(s.pending, float32(v)/32768.0)
}
for len(s.pending) >= sileroWindow {
p, err := s.infer(s.pending[:sileroWindow])
if err != nil {
// A failed inference must not silence the microphone. Hold the
// last answer and let the next window try again.
break
}
s.last = p
s.pending = s.pending[sileroWindow:]
}
return s.last >= threshold, s.last
}
// infer runs one window and rolls the state and the context forward.
func (s *sileroVAD) infer(window []float32) (float64, error) {
in := make([]float32, sileroContext+sileroWindow)
copy(in, s.context[:])
copy(in[sileroContext:], window)
inT, err := ort.NewTensor(ort.NewShape(1, int64(len(in))), in)
if err != nil {
return 0, err
}
defer inT.Destroy()
stT, err := ort.NewTensor(ort.NewShape(2, 1, sileroStateDim), s.state)
if err != nil {
return 0, err
}
defer stT.Destroy()
srT, err := ort.NewTensor(ort.NewShape(1), s.sr)
if err != nil {
return 0, err
}
defer srT.Destroy()
out, err := ort.NewEmptyTensor[float32](ort.NewShape(1, 1))
if err != nil {
return 0, err
}
defer out.Destroy()
next, err := ort.NewEmptyTensor[float32](ort.NewShape(2, 1, sileroStateDim))
if err != nil {
return 0, err
}
defer next.Destroy()
if err := s.session.Run(
[]ort.Value{inT, stT, srT},
[]ort.Value{out, next},
); err != nil {
return 0, err
}
copy(s.state, next.GetData())
copy(s.context[:], in[len(in)-sileroContext:])
return float64(out.GetData()[0]), nil
}
// Reset drops the streaming state. Called at every utterance boundary and
// after barge-in, so echo-era history never scores the next sentence.
func (s *sileroVAD) Reset() {
s.mu.Lock()
defer s.mu.Unlock()
s.pending = s.pending[:0]
s.context = [sileroContext]float32{}
for i := range s.state {
s.state[i] = 0
}
s.last = 0
}
// Close releases the session.
func (s *sileroVAD) Close() error {
if s == nil || s.session == nil {
return nil
}
return s.session.Destroy()
}
+149
View File
@@ -0,0 +1,149 @@
package main
// What this measures. The energy threshold cannot tell a voice from a
// television, and every utterance it accepts becomes a turn. So the test that
// matters is not "does silero find speech" — it is "does it decline what the
// energy threshold accepts".
//
// Speech is the four piper fixtures mavsttd already scores against. They are
// synthesised, so nothing of the owner's voice is committed. Non-speech is
// white noise at the same loudness, which is the cheapest thing that fools an
// energy floor and the honest floor for this claim.
//
// Both halves skip without models/vad/silero_vad.onnx and MAVEN_ONNX_LIB,
// like the TestONNX measurements in internal/router/eval.
import (
"math"
"math/rand"
"os"
"path/filepath"
"testing"
)
const wavHeader = 44 // 16kHz mono s16le, written by piper
func loadSilero(t *testing.T) *sileroVAD {
t.Helper()
model := filepath.Join("..", "..", "models", "vad", "silero_vad.onnx")
lib := os.Getenv("MAVEN_ONNX_LIB")
if _, err := os.Stat(model); err != nil {
t.Skipf("missing %s: %v", model, err)
}
if lib == "" {
t.Skip("MAVEN_ONNX_LIB unset")
}
s, err := newSileroVAD(model, lib)
if err != nil {
t.Skipf("silero unavailable: %v", err)
}
return s
}
// feedAll runs a whole clip through a VAD and reports how many utterances it
// produced and how many frames it called speech.
func feedAll(v *VAD, pcm []int16) (utterances, speechFrames int) {
for i := 0; i+frameSamples <= len(pcm); i += frameSamples {
frame := pcm[i : i+frameSamples]
utt, state := v.Feed(frame)
if state == StateSpeech {
speechFrames++
}
if utt.Bytes != nil {
utterances++
}
}
return utterances, speechFrames
}
func readFixture(t *testing.T, name string) []int16 {
t.Helper()
raw, err := os.ReadFile(filepath.Join("..", "mavsttd", "testdata", name))
if err != nil {
t.Skipf("missing fixture %s: %v", name, err)
}
if len(raw) <= wavHeader {
t.Fatalf("%s: %d bytes, no audio", name, len(raw))
}
return PCMToI16(raw[wavHeader:])
}
// noise returns white noise scaled to the same RMS as ref. Same loudness,
// nothing said.
func noise(ref []int16, seed int64) []int16 {
target := frameRMS(ref)
r := rand.New(rand.NewSource(seed))
out := make([]int16, len(ref))
for i := range out {
out[i] = int16(r.NormFloat64() * target * 32768.0)
}
return out
}
func TestSileroHearsSpeechAndDeclinesNoise(t *testing.T) {
s := loadSilero(t)
defer s.Close()
for _, name := range []string{"ru_fact.wav", "ru_query.wav", "ru_reminder.wav", "en_act.wav"} {
pcm := readFixture(t, name)
v := NewVAD(0, 0, 0, 0)
v.UseSilero(s, defaultSileroThreshold)
_, spoke := feedAll(v, pcm)
if spoke == 0 {
t.Errorf("%s: silero heard no speech in a spoken clip", name)
}
s.Reset()
v2 := NewVAD(0, 0, 0, 0)
v2.UseSilero(s, defaultSileroThreshold)
_, heard := feedAll(v2, noise(pcm, 7))
energy := NewVAD(0, 0, 0, 0)
_, energyHeard := feedAll(energy, noise(pcm, 7))
t.Logf("%s: speech frames — silero on speech %d, silero on noise %d, energy on noise %d",
name, spoke, heard, energyHeard)
if heard >= energyHeard {
t.Errorf("%s: silero called %d noise frames speech, energy called %d — no improvement",
name, heard, energyHeard)
}
s.Reset()
}
}
// BenchmarkSileroFrame answers the only performance question that matters
// here: one 30ms frame must cost far less than 30ms on one core, or the
// detector cannot run always-on beside everything else on that machine.
func BenchmarkSileroFrame(b *testing.B) {
s := loadSilero(&testing.T{})
if s == nil {
b.Skip("silero unavailable")
}
defer s.Close()
frame := make([]int16, frameSamples)
for i := range frame {
frame[i] = int16(i%400 - 200)
}
for i := 0; i < b.N; i++ {
s.Speech(frame, defaultSileroThreshold)
}
}
// TestSileroRechunksAcrossFrames pins the reason this file exists. The capture
// frame is 480 samples and the model window is 512, so a detector that ran one
// inference per frame would be feeding the model a shape it does not accept.
func TestSileroRechunksAcrossFrames(t *testing.T) {
s := loadSilero(t)
defer s.Close()
silence := make([]int16, frameSamples)
for i := 0; i < 20; i++ {
if _, p := s.Speech(silence, defaultSileroThreshold); math.IsNaN(p) {
t.Fatalf("frame %d: probability is NaN", i)
}
}
if len(s.pending) >= sileroWindow {
t.Errorf("pending grew to %d samples, so windows are not being consumed", len(s.pending))
}
}
+35 -1
View File
@@ -68,6 +68,37 @@ type VAD struct {
// follows the room's ambient level. Initialised to minRMS; updated
// on each silence frame.
floorRMS float64
// speech is silero-vad, or nil. When it is set the energy floor decides
// nothing: the question becomes "is this speech" rather than "is this
// loud", and the noise floor is not even tracked. Everything after that
// answer — the speech hold, the silence hold, the length cap, the
// buffer — is the same state machine either way, which is why the
// detector goes here and not around this type.
speech *sileroVAD
speechMin float64
}
// UseSilero swaps the energy threshold for the model. Passing nil is a
// no-op, so a caller that could not load the graph keeps a working VAD.
func (v *VAD) UseSilero(s *sileroVAD, threshold float64) {
if s == nil {
return
}
if threshold <= 0 {
threshold = defaultSileroThreshold
}
v.speech = s
v.speechMin = threshold
}
// isSpeech answers the one question the state machine asks of a frame.
func (v *VAD) isSpeech(frame []int16, rms float64) bool {
if v.speech != nil {
ok, _ := v.speech.Speech(frame, v.speechMin)
return ok
}
return rms >= v.floorRMS
}
// NewVAD creates a VAD with the given thresholds. Zero values use defaults.
@@ -110,7 +141,7 @@ func (v *VAD) State() SpeechState { return v.state }
// should send the audio to the voice server before feeding more frames.
func (v *VAD) Feed(frame []int16) (_ audio.Audio, state SpeechState) {
rms := frameRMS(frame)
isSpeech := rms >= v.floorRMS
isSpeech := v.isSpeech(frame, rms)
switch v.state {
case StateSilence:
@@ -175,6 +206,9 @@ func (v *VAD) Feed(frame []int16) (_ audio.Audio, state SpeechState) {
func (v *VAD) Reset() { v.reset() }
func (v *VAD) reset() {
if v.speech != nil {
v.speech.Reset()
}
v.state = StateSilence
v.speechFrames = 0
v.silenceFrames = 0
@@ -0,0 +1,103 @@
# Kiwix answered the wrong question, and the fix was not a relevance gate
Date: 2026-08-09. Task: V-668. Box: homesrv, workstation off.
Book: `wikipedia_ru_all_maxi_2026-02` on `127.0.0.1:8034`.
## What started it
Two turns on 2026-08-09 came back wrong from the offline encyclopedia.
"почему небо голубое" was answered off the song "Город золотой". "что такое
TCP?" was answered off "Перехват TCP-соединения". Both were phrased
confidently, because `queryKiwix` claims a turn whenever the search returns
anything and `len(hits) == 0` is its only gate.
The plan was a relevance gate. multilingual-e5-small is asymmetric and trained
for exactly this, `query:` against `passage:`, and the query vector is already
held on the turn. The 2026-08-05 measurement that killed a search-quality gate
killed three lexical signals. It says in its own words that it never probed
Kiwix.
## The gate does not exist
Fourteen Russian questions, eight the encyclopedia can answer and six it
cannot. Each question was searched, the top article read, and the cosine of
`EmbedQuery(question)` against `EmbedPassage(article)` recorded.
| set | n | min | mean | max |
|---|---|---|---|---|
| answerable | 8 | 0.7934 | 0.8400 | 0.9087 |
| not answerable | 6 | 0.7480 | 0.7852 | 0.8367 |
Two of the six unanswerable score above the weakest answerable one. That alone
would be a poor threshold. The log killed it outright: seven of the eight
answerable questions got a **wrong** article back, and those wrong articles
scored high. The TCP hijacking article scored 0.8653, above five of the six
unanswerable rows.
The finding is that this cosine measures topic and not answerhood. A page about
hijacking TCP sessions is about TCP. No threshold separates it from a page that
defines TCP, and one that tried would take the definition with it.
## The defect is retrieval
`internal/kiwix/client.go` has said it since it was written: ranking is keyword
based, "why is the sky blue" finds a TV episode. `queryKiwix` sends the whole
sentence. The English path has a rewriter that reduces a question to keywords
with a model call. The Russian path reads the book verbatim (V-508) and had
nothing. So the question words compete with the one word that names the article.
Dropping the question words changes the answer:
| sent | first hit |
|---|---|
| `кто написал Войну и мир` | Радуйся, мир (Доктор Кто) |
| `Война и мир` | Война и мир |
| `что такое TCP` | Перехват TCP-соединения |
| `TCP` | TCP |
A ZIM is also addressable by title, which nothing here used. `/A/Франция`,
`/A/TCP` and `/A/Небо` are 200. `/A/Трюмбальная_нидроскопия` is 404. So an
exact title is safe to try first: it either answers or costs one request that
says nothing.
## What shipped, measured
`kiwix.Topic` drops the narrative request, the interrogative and a verb sitting
behind one. It keeps everything else, because a word it cannot classify is more
likely the topic than noise. `kiwix.TitlePath` tries the exact article before
any ranking runs. Both apply on the verbatim path only, since reducing twice
would take the topic off the rewriter's input.
| question | before | after |
|---|---|---|
| что такое TCP? | Перехват TCP-соединения | **TCP** (by title) |
| что такое фотосинтез | C4-фотосинтез | **Фотосинтез** |
| кто такой Линус Торвальдс? | Tux | **Торвальдс, Линус** (by title) |
| кто написал Войну и мир | Радуйся, мир (Доктор Кто) | **Война и мир** |
| что такое чёрная дыра | Чёрная дыра | Чёрная дыра |
| почему небо голубое | Город золотой | Под небом голубым… (фильм) |
| почему трава зелёная | Сено | Зелень |
| столица Франции | Список столиц Олимпийских игр | Список столиц Олимпийских игр |
Four questions reach the right article where they did not. Two were already
right and stay right. Nothing regressed.
## What is still wrong
Two of the eight are still not answered, and both are the same shape. The
question names no article. "почему небо голубое" is answered by Rayleigh
scattering and "столица Франции" by the lead of Франция. Neither title is in
the question. Keyword retrieval cannot bridge that and neither can a
threshold. The candidates are a semantic index over titles, or asking the
resident model for the article title rather than for keywords.
`Response.Empty()` is still the whole gate. A wrong article that the search
does return is still spoken. What this change buys is that the article is
usually right, not that a wrong one is caught.
## Not measured here
The English path, which still goes through the rewriter and was not touched.
SearXNG, where the same question about answerhood is open and the 2026-08-05
result stands. The cascade end to end, since the workstation is off and the
phrasing arm is the resident model.
+63
View File
@@ -0,0 +1,63 @@
# silero-vad against the energy threshold in mavwaked
*Measured 2026-08-09 on homesrv. V-487, stage one of two.*
mavwaked decided an utterance had started by comparing frame energy to an
adaptive floor. That answers "is this frame loud". A fan, a door and a
television are all loud, and every utterance mavwaked accepts becomes a turn.
silero-vad answers "is this frame speech". It is 2.3MB of ONNX and it replaces
the comparison and nothing else. The speech hold, the silence hold, the length
cap and the utterance buffer are the same state machine either way.
## What it declines
Speech is the four piper fixtures `mavsttd` already scores against, so nothing
of the owner's voice is committed. Non-speech is white noise at the same RMS as the clip beside it. That is the
cheapest thing that fools an energy floor.
| clip | silero, speech frames on speech | silero on noise | energy on noise |
|---|---|---|---|
| ru_fact.wav | 59 | 0 | 68 |
| ru_query.wav | 69 | 0 | 79 |
| ru_reminder.wav | 80 | 0 | 89 |
| en_act.wav | 90 | 0 | 99 |
The energy threshold accepts every noise clip as a complete utterance. Silero
calls not one frame of any of them speech, and still hears all four spoken
clips. `TestSileroHearsSpeechAndDeclinesNoise` is that table.
White noise is a floor, not a proof. It says nothing about a television, which
is speech, or about a fan, which is narrowband. Those need room recordings and
this box has none.
## What it costs
`BenchmarkSileroFrame` on the homesrv laptop (Ryzen 5 5600U), one 30ms frame
through the model including the re-chunking:
509µs per frame
That is 1.7% of one core, on the slower of the two machines. The detector runs
on the workstation beside the microphone, never on the GPU. This number is what
says it does not need one.
## The window is 512 samples, not 480
`cmd/mavwaked/main.go` claimed the frame contract matched silero's input
exactly. That was true of silero v4. Version 5 takes exactly 512 samples at 16kHz, plus 64 samples of context from
the previous window. So `sileroVAD` buffers across capture frames, and a frame
completing no window inherits the previous probability. `TestSileroRechunksAcrossFrames` pins it.
## Still an energy gate by default
`-vad-model` is empty in the code default, so a deployment that does not pass
it runs exactly what shipped before. Barge-in is untouched and deliberately so. It reads frame energy while she is
speaking, which is a different question from whether the frame is speech.
## Not done here
The wake word. This is stage one of the two V-487 asks for. The second needs a
keyword model that does not exist yet. The pretrained openWakeWord keywords are
English, and a Russian one has to be trained. Until then anything spoken near
the microphone still becomes a turn. It is now merely required to be speech.
+16
View File
@@ -127,6 +127,22 @@ func (c *Client) Article(ctx context.Context, path string, maxRunes int) (crawl.
return crawl.Extract(u, body, maxRunes), nil
}
// TitlePath is the article path for an exact title, for Article to fetch.
//
// It exists because a ZIM is addressable by title and the full-text index is
// not the only way in. "Франция", "TCP" and "Небо" resolve; "Трюмбальная
// нидроскопия" is a 404, which is the honest answer and the reason this is
// safe to try first. Measured on 2026-08-09, keyword search on the same terms
// returns "Список пэров Франции" and "Список портов TCP и UDP" instead.
//
// A miss is normal rather than a failure. An article whose title inverts a name
// ("Торвальдс, Линус") is a 404 here and the first hit in search, so the caller
// falls through and loses nothing.
func TitlePath(book, title string) string {
t := strings.ReplaceAll(strings.TrimSpace(title), " ", "_")
return "/content/" + url.PathEscape(book) + "/A/" + url.PathEscape(t)
}
// rss mirrors just the bits of the RSS 2.0 reply we use.
type rss struct {
Items []struct {
+65
View File
@@ -0,0 +1,65 @@
package kiwix
import (
"context"
"os"
"testing"
"time"
)
// TestLiveTopicBeatsTheSentence — the measurement V-668 turned on, kept as a
// test so the claim can be re-run rather than believed.
//
// It prints the article the old path returned and the article the new one
// returns, for the same question. It asserts nothing about which is better,
// because "is this the right article" is a human's call. It fails only if the
// two paths agree on every case, which would mean the change does nothing.
//
// MAVEN_KIWIX_URL=http://127.0.0.1:8034 make t PKG=./internal/kiwix/ RUN=TestLive V=1
func TestLiveTopicBeatsTheSentence(t *testing.T) {
base := os.Getenv("MAVEN_KIWIX_URL")
if base == "" {
t.Skip("MAVEN_KIWIX_URL unset — point it at the kiwix-server host port")
}
const book = "wikipedia_ru_all_maxi_2026-02"
c := New(base)
questions := []string{
"что такое TCP?",
"что такое фотосинтез",
"кто такой Линус Торвальдс?",
"кто написал Войну и мир",
"что такое чёрная дыра",
"почему небо голубое",
"почему трава зелёная",
"столица Франции",
}
moved := 0
for _, q := range questions {
ctx, cancel := context.WithTimeout(context.Background(), 20*time.Second)
before := firstTitle(ctx, c, q, book)
topic := Topic(q)
after := ""
if page, err := c.Article(ctx, TitlePath(book, topic), 400); err == nil && page.Text != "" {
after = page.Title + " (by title)"
} else {
after = firstTitle(ctx, c, topic, book)
}
cancel()
if before != after {
moved++
}
t.Logf("%-30s before=%-34q after=%q", q, before, after)
}
t.Logf("%d of %d questions reach a different article", moved, len(questions))
if moved == 0 {
t.Error("the topic path returns exactly what the sentence path returned")
}
}
func firstTitle(ctx context.Context, c *Client, pattern, book string) string {
hits, err := c.Search(ctx, pattern, book, 3)
if err != nil || len(hits) == 0 {
return "(nothing)"
}
return hits[0].Title
}
+98
View File
@@ -0,0 +1,98 @@
package kiwix
import (
"strings"
"unicode"
"github.com/kami/maven/internal/lexicon"
"github.com/kami/maven/internal/morph"
)
// Topic reduces a question to the thing it is about, because Kiwix ranks by
// keyword overlap and a whole sentence buries the keyword that matters.
//
// This package's own doc says it: "why is the sky blue" finds a TV episode.
// Measured against the Russian ZIM on 2026-08-09, the sentence and the topic
// return different articles for the same question. "кто написал Войну и мир"
// returns "Радуйся, мир (Доктор Кто)"; "Войну и мир" returns the novel first.
// "что такое TCP" returns "Перехват TCP-соединения"; "TCP" returns TCP. The
// English path had a rewriter doing this with a model call. The Russian path
// reads the book verbatim (V-508) and had nothing.
//
// It drops three things off the front and stops: the narrative request, the
// interrogative, and a verb sitting between them and the noun. Everything else
// is kept, because a word this cannot classify is more likely the topic than
// noise. An empty return means the utterance was question words alone, and the
// caller searches the sentence as before.
func Topic(utterance string) string {
words := strings.Fields(strings.TrimSpace(utterance))
cut := 0
for cut < len(words) {
w := strings.Trim(strings.ToLower(words[cut]), ".,!?…:;\"'«»")
if w == "" {
cut++
continue
}
switch {
case inList(lexicon.NarrativeRequests(), w),
inList(lexicon.Interrogatives(), w),
inList(lexicon.FirstPerson(), w),
// "что ТАКОЕ x", "кто ТАКОЙ x" — the copula that only ever follows
// an interrogative, and never a topic on its own.
cut > 0 && isCopula(w),
// "расскажи ПРО x", "о x". One-letter and two-letter prepositions
// are not a closed class worth a lexicon set of their own.
cut > 0 && isLeadingPreposition(w),
// "кто НАПИСАЛ Войну и мир". A verb here is the question's own
// verb, not part of the title. Only after something was already
// dropped, so "написал отчёт" as a topic survives intact.
cut > 0 && morph.IsVerbForm(w):
cut++
default:
// The question mark is the sentence's, not the title's, and Kiwix
// carries it into the keyword match.
topic := strings.TrimRight(strings.Join(words[cut:], " "), " .,!?…:;\"'«»")
if !hasLetter(topic) {
return ""
}
return topic
}
}
return ""
}
func isCopula(w string) bool {
switch w {
case "такое", "такой", "такая", "такие", "is", "are", "was", "were":
return true
}
return false
}
func isLeadingPreposition(w string) bool {
switch w {
case "про", "о", "об", "обо", "по", "about", "of", "on":
return true
}
return false
}
func inList(list []string, w string) bool {
for _, x := range list {
if x == w {
return true
}
}
return false
}
// hasLetter is the guard against a topic that reduced to punctuation or digits
// alone, which no ZIM title matches.
func hasLetter(s string) bool {
for _, r := range s {
if unicode.IsLetter(r) {
return true
}
}
return false
}
+51
View File
@@ -0,0 +1,51 @@
package kiwix
import "testing"
// The cases the 2026-08-09 measurement turned on, plus the ones a topic must
// not damage. Each left column returned a wrong article when it was sent whole.
func TestTopicKeepsTheThingTheQuestionIsAbout(t *testing.T) {
cases := []struct{ utterance, want string }{
{"что такое TCP?", "TCP"},
{"что такое фотосинтез", "фотосинтез"},
{"кто такой Линус Торвальдс?", "Линус Торвальдс"},
{"кто написал Войну и мир", "Войну и мир"},
{"расскажи про битву при Ватерлоо", "битву при Ватерлоо"},
{"what is photosynthesis", "photosynthesis"},
// No question word, so there is nothing to drop. The topic is the
// whole utterance and the search is what it was before.
{"столица Франции", "столица Франции"},
{"почему небо голубое", "небо голубое"},
}
for _, c := range cases {
if got := Topic(c.utterance); got != c.want {
t.Errorf("Topic(%q) = %q, want %q", c.utterance, got, c.want)
}
}
}
// A verb only goes when a question word already went. Otherwise "написал
// отчёт" loses the verb that names what he means.
func TestTopicDropsAVerbOnlyBehindAQuestionWord(t *testing.T) {
if got := Topic("написал отчёт"); got != "написал отчёт" {
t.Errorf("Topic dropped a leading verb with no question word: %q", got)
}
}
// Question words alone reduce to nothing, and the caller reads that as "no
// topic" and searches the sentence rather than searching the empty string.
func TestTopicIsEmptyWhenNothingIsLeft(t *testing.T) {
for _, q := range []string{"что такое?", "кто?", "почему", "???"} {
if got := Topic(q); got != "" {
t.Errorf("Topic(%q) = %q, want empty", q, got)
}
}
}
func TestTitlePathEscapesAndUnderscores(t *testing.T) {
got := TitlePath("wikipedia_ru_all_maxi_2026-02", "Чёрная дыра")
want := "/content/wikipedia_ru_all_maxi_2026-02/A/%D0%A7%D1%91%D1%80%D0%BD%D0%B0%D1%8F_%D0%B4%D1%8B%D1%80%D0%B0"
if got != want {
t.Errorf("TitlePath = %q, want %q", got, want)
}
}
+15
View File
@@ -111,6 +111,21 @@ type Decision struct {
// before this field existed. See source.go for why it is twelve values.
Source Source
// SourceAnchored — a stage 0 grammar named that destination, matching a
// literal pattern to do it. Only the router sets this, and only there.
//
// It exists because one thing downstream is not reversible by evidence
// (V-666). Naming a destination normally takes guessing sources off a turn,
// and one of those is the personal boundary, which is what stops a question
// about him from reaching the world. A grammar that read "что такое X" may
// take it off. A model or a softmax may not, because a wrong destination
// there widens what leaves the box rather than costing an answer.
//
// Read Stage instead and the two decisions get coupled: stage 0 also means
// confidence 1.0 and an anchored claim band, and a later cascade change
// could make one true where the other is not.
SourceAnchored bool
// Continued — this decision was rebuilt from the previous turn rather
// than routed, because the utterance was an ellipsis ("а завтра?").
// Handlers use it to know that Slots.Text is the PREVIOUS turn's topic
+4
View File
@@ -98,6 +98,10 @@ func (r *Router) Route(ctx context.Context, utterance string, now time.Time) (De
continue // grammar matched shape but not content → fall through
}
d.Utterance = utterance
// A literal pattern named that destination, which is the one provenance
// allowed to take the personal boundary off a turn (V-666). Set here and
// nowhere else, so no other arm of the cascade can claim it.
d.SourceAnchored = d.Source != SourceUnknown
// The grammar decided the intent; the extractor fills the slots it did
// not match (V-572). See fillMatchedSlots for why every grammar gets it.
r.fillMatchedSlots(ctx, &d, now)