Maven can record a meeting when she is told to, transcribe it through the STT she already has, and write a summary note. The audio lives in the blob store #252 introduced, under the same retention loop. Nothing here listens. Recorder.Append is the only way audio enters and it refuses every frame unless someone explicitly started a session, so audio arriving at an idle core is dropped rather than buffered. The plan document asked for a keyword trigger ("maven record" heard in the room) and that is refused: noticing a keyword means listening to the room, which is the one behaviour this capability must not have. Off unless configured twice over. No media block means nowhere to keep audio, no capture block means no recorder, and in either case the four IPC methods answer ErrUnknownMethod. On an unconfigured box there is no wire path that begins a recording at all. A forgotten session ends itself at max_minutes, checked on every append, and the audio collected before the cap is kept. Stop with discard set is what "забудь, не записывай" maps to and it leaves nothing behind. The verbatim transcript is not saved unless save_transcript says so; the summary is. Long audio against n_ctx 4096 is handled by map-reduce over 3000-rune windows rather than by truncation, because a truncated meeting summary reads as complete and is not. Transcription is windowed at five minutes so the whisper worker stays responsive to the voice path. No second STT: internal/capture takes the stt.Transcriber the voice path already holds. Capture with voice off is refused rather than degraded, since hours of unreadable audio of other people is worse than no recording. The three write methods are AuthWrite, not AuthStepUp: step-up needs a passkey gesture the voice path cannot make, which would leave "запиши встречу" impossible by voice. capture_status is AuthRead. make build and make test both pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
6.7 KiB
Plan: Hearing — Meeting Capture & Summarisation
Goal: "Maven, запиши встречу" starts a recording, "хватит" stops it, and she writes a summary note. The audio stays on the box, is pruned by retention, and nothing is recorded that nobody asked for.
Status (2026-08-01): the recorder, the storage, the chunked transcription, the map-reduce summariser, the config seam and the four IPC methods are shipped and tested. What is not shipped is the workpc-side microphone agent and the router intent — see "Still open".
What shipped
| Piece | Where |
|---|---|
| Session state machine: start / append / stop / abort / status | internal/capture/capture.go |
Map-reduce summarisation against n_ctx 4096 |
internal/capture/summarize.go |
Audio blobs in the shared store, pruned by media.retention |
internal/media (from #252) |
Config block capture, off by default |
internal/config/config.go |
IPC capture_start / capture_append / capture_stop / capture_status |
internal/ipc/{wire,api,client,server}.go |
Authority: the three write methods AuthWrite, status AuthRead |
internal/auth/policy.go |
| Daemon wiring, note write, STT reuse | cmd/mavend/capture.go |
The audio lands in the same content-addressed blob store as images, under the same retention loop, because #252 and #253 have the same intake problem and solving it twice would mean two directories to remember to prune.
The refusals, and why
Nothing listens. The original step 8 called for capture "triggered by voice command
(IntentCapture) or configurable keyword ('maven record')". The keyword half is refused.
Noticing a keyword requires listening to the room continuously, which is precisely the
behaviour this capability must not have, and the refusal is in the code rather than in a
comment: Recorder.Append is the only way audio enters, and it returns ErrNoSession unless
someone explicitly started a session. Audio arriving at an idle core is dropped, not buffered
"just in case".
Off unless configured, twice over. No media block ⇒ nowhere to keep audio ⇒ the four
methods do not exist. No capture block with enabled: true ⇒ they still do not exist. On an
unconfigured box there is no wire path at all that begins a recording. That is the only
guarantee worth making here, and it is the reason the hooks use the nil-hook ⇒
ErrUnknownMethod pattern rather than an in-handler check.
A forgotten session ends itself. max_minutes defaults to 120 and is checked on every
append, not on a timer that could be missed. Past the cap Append returns ErrExpired
permanently, so a client that ignores the error cannot grow the recording; the audio collected
before the cap is kept and Stop still works.
"Забудь, не записывай" leaves nothing behind. capture_stop with discard: true throws
the session away without storing, transcribing or summarising anything — not a blob with a note
saying it was abandoned. Nothing.
The transcript is not saved by default. The summary is written where he will read it; the
verbatim record of what other people said in a room is a heavier thing to keep and takes a
deliberate save_transcript: true. The audio blob is pruned by media.retention either way.
No second STT. Step 3 of the original plan extended the Transcriber interface with
streaming. Not needed and not done: whisper.cpp already runs as mavsttd, and internal/capture
takes the ordinary stt.Transcriber the voice path already holds (exposed as
voiceWiring.transcriber). Long recordings are handed over in five-minute windows —
chunkAudio, cut on sample boundaries — for the same reason whisper itself works in 30-second
windows: an hour of PCM in one call either times out or blocks the voice path for minutes.
Capture with voice off is refused rather than degraded, because storing hours of unreadable
audio of other people is worse than not recording.
Not AuthStepUp. Recording people is invasive enough to argue for the top rung, and it is
still wrong: step-up needs a passkey gesture, which the voice path cannot make, so
"запиши встречу" could never work by voice — the only way he will actually use this. AuthWrite
plus the off-unless-configured gate is the honest combination.
Long audio against a 4096-token context
The resident model is a Thinking variant at n_ctx 4096, so an hour of transcript does not fit
in one prompt and never will. summarize.go does map-reduce and nothing cleverer: split the
transcript on sentence boundaries into 3000-rune windows (about 1100 Qwen tokens of Russian,
leaving room for the persona block, the reasoning and the answer), summarise each, then
summarise the summaries. A transcript that fits in one window skips the reduce step.
Truncation was the alternative and is rejected: a truncated meeting summary reads as complete
and is not, and he would act on it. Past max_chunks (40, roughly the two-hour cap) the
transcript is cut, and the summary says so in the note.
Two degradations are deliberate and both are reported rather than hidden:
- No llama-server ⇒ transcript, no summary. The words exist.
- The reduce call fails ⇒ the per-chunk summaries are returned joined. Real work, not thrown away over the last call.
The map and reduce prompts contain no first person at all, so the persona's feminine-form rules have nothing to get wrong in them; the reply she actually gives him is phrased by the ordinary replier, which does carry the persona.
Config
"media": { "dir": "media", "retention": "168h" },
"capture": {
"enabled": true,
"max_minutes": 120,
"stt_window": "5m",
"chunk_runes": 3000,
"max_chunks": 40,
"save_transcript": false
}
Both absent by default. capture alone does nothing without media.
Still open
cmd/mavheard— the workpc-side microphone agent. Deferred, not refused: the core half is the part with the invariants in it, and a mic client is straightforward once there is a stable wire to stream at. It should be an explicit-start process, not a resident one, for the same reason the recorder has no keyword trigger. The four IPC methods are the wire it will use;mavenclientalready has the mic plumbing to borrow.- Router intent. "запиши встречу" / "хватит" does not route anywhere yet. It needs the
systemintent plus slots, and it needs care: "хватит" is also how someone tells her to stop talking, so the recorder's stop and the speech barge-in must not collide. - A
/dashpanel showing a running session, so a recording is visible on a surface and not only in a log line. - Speaker attribution — who said what — is #255 and is blocked on a model; see
docs/plans/10-speaker-recognition.md.