Files
Maven/docs/plans/10-speaker-recognition.md
kami 7c7bd8ceeb Ship voice enrolment, and report recognition as blocked (#255)
Maven can now be told who someone is. She cannot yet tell who is speaking,
and this commit is careful to say so rather than pretend otherwise.

What works: profiles are enrolled from several deliberately recorded samples,
listed, and deleted. They live in the existing memory_vectors table under a
"speaker:" id prefix, so there is no migration; what that needed was a wider
interface than memory.Store, hence memory.Catalog with ByPrefix and Delete.
Delete is the load-bearing half — a voiceprint someone asked to be rid of has
to actually go, and a search-only store cannot do that. InMemoryStore.Insert
became an upsert by id to match what the persistent store already did.

What does not work, and why it is not faked: there is no speaker-embedding
model on this box. Sixteen ggufs in /mnt/hdd1/llms, all text; no ECAPA, no
x-vector, no titanet, no wespeaker, no .onnx anywhere under /mnt/hdd1. So
newSpeakerEmbedder returns nil, internal/speaker falls back to
speaker.Disabled, Identify answers ErrDisabled, and the daemon logs which
half is off at startup. The plan's "simple MFCC + GMM" floor is refused in
the package comment: MFCC cosine distance detects channel and loudness as
much as voice, and a biometric that is confidently wrong writes false claims
about named people into his memory. A bad floor is worse than none here.

Refused as well, and the reason is in enroll.go's doc comment: the plan asked
for unknown speakers to be enrolled on first interaction with a TTS "кто
это?". There is no request shape in the protocol that could express that.
Taking a biometric of whoever walks past the microphone does it to guests who
are not party to the exchange, and a synthesised question into a room is not
consent from whoever answers.

Authority: enrolment is AuthStepUp, because it is a deliberate sit-down act
that writes a biometric of a named person and never something done by voice
mid-conversation. Deletion is one rung lower at AuthWrite, deliberately
inverting the usual pattern — getting rid of a biometric must never be the
harder half. Listing is AuthRead and never returns the vectors themselves.

Off unless configured: no speaker block means the three methods answer
ErrUnknownMethod, so a default box has no wire path that takes a voiceprint.

make build and make test pass.

Vikunja #255

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 05:23:03 +04:00

7.2 KiB

Plan: Speaker Recognition

Goal: Maven can tell who is speaking on the voice channel, and tag what she writes with who said it.

Status (2026-08-01, Vikunja #255): the enrolment half is shipped. The recognising half is BLOCKED on a model download — there is no speaker-embedding model on this box, and one was not invented to fill the gap. See "Blocked, and on what" below.

What shipped

Piece Where State
Recognizer — identify, list, get, forget internal/speaker/recognizer.go done; Identify answers ErrDisabled until a model exists
Enrolment — several samples, averaged, re-normalised internal/speaker/enroll.go done
Profile shape, id validation, cosine similarity internal/speaker/speaker.go done
Profile storage as speaker:<id> vectors internal/memory Catalog + internal/store/memory.go done, no schema migration
Config block, off by default internal/config SpeakerConfig done
enroll_speaker / list_speakers / forget_speaker internal/ipc done, absent unless configured
Authority rows internal/auth/policy.go done — enrol step-up, forget write, list read
Daemon wiring + honest startup log cmd/mavend/speaker.go done
Embedding backend newSpeakerEmbedder BLOCKED — returns nil, seam only
Tagging voice writes with the speaker cmd/mavend/voice.go not wired; nothing to tag with yet

Blocked, and on what

A voiceprint needs a speaker-embedding model. The box was searched: /mnt/hdd1/llms holds sixteen ggufs across seven families and every one of them is a text model. There is no ECAPA, no x-vector, no titanet, no wespeaker, and no .onnx under /mnt/hdd1 at all. There are also no enrolment samples, because nothing has ever recorded any.

To unblock, two things are needed and neither can be done from inside the repo:

  1. A model. SpeechBrain ECAPA-TDNN exported to ONNX (speechbrain/spkrec-ecapa-voxceleb, 192-dim) is the usual choice and runs on CPU in well under a second for a few seconds of audio. Download it per the recipe in AGENTS.md, put it beside the other models so the bind mount picks it up, and point speaker.model_path at it.
  2. An implementation of one function. newSpeakerEmbedder in cmd/mavend/speaker.go is the entire seam: give it an ONNX session that turns audio.Audio into a []float32 and Identify starts working. Nothing else changes — not the store, not the protocol, not the authority table, not the handlers. internal/onnx already loads the e5 embedder, so the runtime wiring exists to copy.
  3. Enrolment samples, three or more per person, recorded deliberately.

Why there is no fallback

The original plan offered "a simple MFCC + GMM" as the floor. That is refused. MFCC cosine distance is a channel and loudness detector as much as a voice detector: it will happily match two different people who sit at the same distance from the same microphone, and it drifts when the room changes. A general classifier that is sometimes wrong is a nuisance; a biometric that is confidently wrong writes false claims about named people into his memory, and then those claims get recalled as fact. For this capability a bad floor is worse than none, so the shipped state is honest absence: speaker.Disabled, ErrDisabled, and a startup line saying so.

The refusals, and why

  • Unknown speakers are NOT enrolled on first interaction. The plan's fourth "done when" bullet asked for exactly that, with a TTS "кто это?" prompt. It is refused in enroll.go's doc comment and there is no request shape in the protocol that could express it. Enrolling a voice is taking a biometric of a person; doing it automatically to whoever walks past the microphone does it to guests who are not party to the exchange, and a synthesised question into a room is not consent from whoever happens to answer. Enrolment is an explicit act: an id, a name, and samples recorded for the purpose.
  • One sample is not enough. Three separate utterances and nine seconds minimum. A profile built from one sentence encodes that sentence as much as the person, and the threshold then behaves unpredictably against everything else.
  • An unknown voice stays unknown. Below threshold, Identify returns ErrUnknown naming the closest profile in the error text for diagnosis, never as an answer. Guessing who is in the room is how false memories about people get written.
  • Deletion is one authority rung below enrolment. Everywhere else in policy.go the destructive direction is gated at least as hard as the constructive one. Here that would be backwards: getting rid of a biometric must never be the harder half.
  • The voiceprint never crosses the socket. ListSpeakersResp carries ids, names, dates and sample counts. The vector stays in core.
  • Off unless configured. No speaker block ⇒ the three methods answer ErrUnknownMethod. There is no wire path on a default box that takes a voiceprint.

Storage

Profiles live in the existing memory_vectors table under the speaker: id prefix, as the plan intended, so there is no migration. What that needed was a wider interface than memory.Store: memory.Catalog adds ByPrefix and Delete. Delete is the load-bearing one — a voiceprint someone asked to be rid of has to actually go, and a search-only store cannot do that. InMemoryStore.Insert also became an upsert by id, matching what the persistent store already did, so re-enrolling replaces a profile instead of stacking a second one behind the first.

Profiles do not collide with note or fact vectors: they are only ever read through ByPrefix("speaker:"), and a note search never returns one because the prefix is not in its query path.

Config

"speaker": {
  "enabled": true,
  "model_path": "/opt/maven/models/spk/ecapa-voxceleb.onnx",
  "lib_path": "/opt/maven/lib",
  "threshold": 0.7,
  "min_seconds": 2.0
}

Recognizes() requires both enabled and a model_path, so a half-filled block reads as off rather than as a capability that fails every turn. With enabled and no model the daemon still attaches the three methods — profiles can be created, listed and deleted — and logs that recognition is blocked.

Still open

  • The embedding backend (above). Everything below waits on it.
  • Tagging voice writes. Profile.Source("tap:voice") already produces tap:voice:speaker:kami, which is the shape step 6 asked for, but nothing calls it yet: with no recogniser there is no id to tag with. When the model lands, the hook is in the voice path before STT.
  • Speaker as router/phraser context. Same dependency. Note the persona constraint when it arrives: Maven addresses the owner informally and speaks to him, so "ok, " needs care for anyone who is not him.
  • An enrolment surface. The three IPC methods exist; no page drives them. Enrolment is step-up, so it belongs on /dash behind a passkey, with a per-profile forget button next to each row — that button is the reason list_speakers exists.
  • A speaker column on the meeting recorder (#253). Attributing lines in a transcript is the obvious pairing, and it is the place where getting attribution wrong is most damaging, so it waits for a real model too.