Files
Maven/docs/plans/08-hearing.md
T
kami aa1a26532c Add meeting capture with explicit start and stop (#253)
Maven can record a meeting when she is told to, transcribe it through the
STT she already has, and write a summary note. The audio lives in the blob
store #252 introduced, under the same retention loop.

Nothing here listens. Recorder.Append is the only way audio enters and it
refuses every frame unless someone explicitly started a session, so audio
arriving at an idle core is dropped rather than buffered. The plan document
asked for a keyword trigger ("maven record" heard in the room) and that is
refused: noticing a keyword means listening to the room, which is the one
behaviour this capability must not have.

Off unless configured twice over. No media block means nowhere to keep
audio, no capture block means no recorder, and in either case the four IPC
methods answer ErrUnknownMethod. On an unconfigured box there is no wire
path that begins a recording at all.

A forgotten session ends itself at max_minutes, checked on every append,
and the audio collected before the cap is kept. Stop with discard set is
what "забудь, не записывай" maps to and it leaves nothing behind. The
verbatim transcript is not saved unless save_transcript says so; the
summary is.

Long audio against n_ctx 4096 is handled by map-reduce over 3000-rune
windows rather than by truncation, because a truncated meeting summary
reads as complete and is not. Transcription is windowed at five minutes so
the whisper worker stays responsive to the voice path.

No second STT: internal/capture takes the stt.Transcriber the voice path
already holds. Capture with voice off is refused rather than degraded,
since hours of unreadable audio of other people is worse than no recording.

The three write methods are AuthWrite, not AuthStepUp: step-up needs a
passkey gesture the voice path cannot make, which would leave "запиши
встречу" impossible by voice. capture_status is AuthRead.

make build and make test both pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
2026-08-01 05:08:08 +04:00

122 lines
6.7 KiB
Markdown

# Plan: Hearing — Meeting Capture & Summarisation
**Goal:** "Maven, запиши встречу" starts a recording, "хватит" stops it, and she writes a
summary note. The audio stays on the box, is pruned by retention, and nothing is recorded that
nobody asked for.
**Status (2026-08-01):** the recorder, the storage, the chunked transcription, the map-reduce
summariser, the config seam and the four IPC methods are shipped and tested. What is not
shipped is the workpc-side microphone agent and the router intent — see "Still open".
## What shipped
| Piece | Where |
|---|---|
| Session state machine: start / append / stop / abort / status | `internal/capture/capture.go` |
| Map-reduce summarisation against `n_ctx` 4096 | `internal/capture/summarize.go` |
| Audio blobs in the shared store, pruned by `media.retention` | `internal/media` (from #252) |
| Config block `capture`, off by default | `internal/config/config.go` |
| IPC `capture_start` / `capture_append` / `capture_stop` / `capture_status` | `internal/ipc/{wire,api,client,server}.go` |
| Authority: the three write methods `AuthWrite`, status `AuthRead` | `internal/auth/policy.go` |
| Daemon wiring, note write, STT reuse | `cmd/mavend/capture.go` |
The audio lands in the same content-addressed blob store as images, under the same retention
loop, because #252 and #253 have the same intake problem and solving it twice would mean two
directories to remember to prune.
## The refusals, and why
**Nothing listens.** The original step 8 called for capture "triggered by voice command
(IntentCapture) **or configurable keyword ('maven record')**". The keyword half is refused.
Noticing a keyword requires listening to the room continuously, which is precisely the
behaviour this capability must not have, and the refusal is in the code rather than in a
comment: `Recorder.Append` is the only way audio enters, and it returns `ErrNoSession` unless
someone explicitly started a session. Audio arriving at an idle core is dropped, not buffered
"just in case".
**Off unless configured, twice over.** No `media` block ⇒ nowhere to keep audio ⇒ the four
methods do not exist. No `capture` block with `enabled: true` ⇒ they still do not exist. On an
unconfigured box there is no wire path at all that begins a recording. That is the only
guarantee worth making here, and it is the reason the hooks use the nil-hook ⇒
`ErrUnknownMethod` pattern rather than an in-handler check.
**A forgotten session ends itself.** `max_minutes` defaults to 120 and is checked on every
append, not on a timer that could be missed. Past the cap `Append` returns `ErrExpired`
permanently, so a client that ignores the error cannot grow the recording; the audio collected
before the cap is kept and `Stop` still works.
**"Забудь, не записывай" leaves nothing behind.** `capture_stop` with `discard: true` throws
the session away without storing, transcribing or summarising anything — not a blob with a note
saying it was abandoned. Nothing.
**The transcript is not saved by default.** The summary is written where he will read it; the
verbatim record of what other people said in a room is a heavier thing to keep and takes a
deliberate `save_transcript: true`. The audio blob is pruned by `media.retention` either way.
**No second STT.** Step 3 of the original plan extended the `Transcriber` interface with
streaming. Not needed and not done: whisper.cpp already runs as `mavsttd`, and `internal/capture`
takes the ordinary `stt.Transcriber` the voice path already holds (exposed as
`voiceWiring.transcriber`). Long recordings are handed over in five-minute windows —
`chunkAudio`, cut on sample boundaries — for the same reason whisper itself works in 30-second
windows: an hour of PCM in one call either times out or blocks the voice path for minutes.
Capture with voice off is refused rather than degraded, because storing hours of unreadable
audio of other people is worse than not recording.
**Not `AuthStepUp`.** Recording people is invasive enough to argue for the top rung, and it is
still wrong: step-up needs a passkey gesture, which the voice path cannot make, so
"запиши встречу" could never work by voice — the only way he will actually use this. `AuthWrite`
plus the off-unless-configured gate is the honest combination.
## Long audio against a 4096-token context
The resident model is a Thinking variant at `n_ctx` 4096, so an hour of transcript does not fit
in one prompt and never will. `summarize.go` does map-reduce and nothing cleverer: split the
transcript on sentence boundaries into 3000-rune windows (about 1100 Qwen tokens of Russian,
leaving room for the persona block, the reasoning and the answer), summarise each, then
summarise the summaries. A transcript that fits in one window skips the reduce step.
Truncation was the alternative and is rejected: a truncated meeting summary reads as complete
and is not, and he would act on it. Past `max_chunks` (40, roughly the two-hour cap) the
transcript *is* cut, and the summary says so in the note.
Two degradations are deliberate and both are reported rather than hidden:
- No llama-server ⇒ transcript, no summary. The words exist.
- The reduce call fails ⇒ the per-chunk summaries are returned joined. Real work, not thrown
away over the last call.
The map and reduce prompts contain no first person at all, so the persona's feminine-form rules
have nothing to get wrong in them; the reply she actually gives him is phrased by the ordinary
replier, which does carry the persona.
## Config
```json
"media": { "dir": "media", "retention": "168h" },
"capture": {
"enabled": true,
"max_minutes": 120,
"stt_window": "5m",
"chunk_runes": 3000,
"max_chunks": 40,
"save_transcript": false
}
```
Both absent by default. `capture` alone does nothing without `media`.
## Still open
- **`cmd/mavheard`** — the workpc-side microphone agent. Deferred, not refused: the core half
is the part with the invariants in it, and a mic client is straightforward once there is a
stable wire to stream at. It should be an explicit-start process, not a resident one, for the
same reason the recorder has no keyword trigger. The four IPC methods are the wire it will
use; `mavenclient` already has the mic plumbing to borrow.
- **Router intent.** "запиши встречу" / "хватит" does not route anywhere yet. It needs the
`system` intent plus slots, and it needs care: "хватит" is also how someone tells her to stop
talking, so the recorder's stop and the speech barge-in must not collide.
- **A `/dash` panel** showing a running session, so a recording is visible on a surface and not
only in a log line.
- **Speaker attribution** — who said what — is #255 and is blocked on a model; see
`docs/plans/10-speaker-recognition.md`.