diff --git a/deploy/mavend.json b/deploy/mavend.json index a7371e3..2881cb6 100644 --- a/deploy/mavend.json +++ b/deploy/mavend.json @@ -78,10 +78,30 @@ "Addressed by LAN address, not container name: mavgpud runs on another", "machine and there is no shared docker network to name it on." ], + "//workstation.stt": [ + "CrisperWhisper 2.0 turbo on the same machine, a second service on port", + "8081 and not a second endpoint on mavgpud. whisper.cpp cannot load CW2 at", + "all: it derives its language count from the vocabulary size, and CW2's", + "51897 tokens shift seven special token ids. So it runs under transformers", + "there and mavsttd stays whisper.cpp here.", + "Worth the second service: CW2 turbo scores 10.4% WER in Russian against", + "27.5% for the ggml-small.bin mavsttd loads, measured on 200 Golos clips", + "in docs/evals/2026-08-09-crisperwhisper2-russian-wer.md.", + "Deleting this block sends every utterance to mavsttd, which is what the", + "box did before it existed. A worse transcript is still a turn, so the", + "fallback is silent and Kami is never told which machine heard him.", + "The token is what stops anything on the LAN posting audio to that port." + ], "workstation": { "url": "http://192.168.1.105:8080", "probe": "15s", - "timeout": "90s" + "timeout": "90s", + "stt": { + "url": "http://192.168.1.105:8081/transcribe", + "token": "${MAVEN_STT_TOKEN}", + "probe": "15s", + "timeout": "10s" + } }, "//search": [ diff --git a/docs/offload.md b/docs/offload.md index 0d4868d..66b9bb1 100644 --- a/docs/offload.md +++ b/docs/offload.md @@ -1,6 +1,6 @@ # Offloading model work to the workstation -*Last verified: 2026-08-05 @ b789676. Living doc: correct it in place, do not append.* +*Last verified: 2026-08-09 @ 50c6637. Living doc: correct it in place, do not append.* Owner's call, 2026-08-02. Vikunja #483 is the umbrella. Tasks #484 to #487 are the work, and this file holds the shape and the rules all four must obey. @@ -149,6 +149,25 @@ flips. It is wired anyway: `PhraseReminder` is on the same transport and is on. Then the embedder above, **whisper.cpp** in `mavsttd`, and **piper** in `mavttsd`. `mavwaked` uses no model at all: an energy-threshold VAD over 30ms frames. +Speech-to-text is wired as of 09-08-2026, and it takes only the silent half of the +rule. A worse transcript is still a turn, so there is nothing to name a gap about +and `stt.Pair` has no `TranscribeRemote`. `sttSeam` in `cmd/mavend/voicewire.go` +builds it, beside `modelSeam` and at the same place in `wireVoice`, so the voice +path and the meeting recorder still share one transcriber. + +The remote is not a second endpoint on mavgpud. whisper.cpp cannot load +CrisperWhisper 2.0 at all. It derives its language count from the vocabulary +size, and CW2's 51897 tokens shift seven special token ids. So CW2 runs under +transformers as its own service on port 8081, and `stt.HTTPTranscriber` is the +second transport for the same seam. It posts raw PCM with the format in headers. +It carries a bearer token, because audio is the most sensitive thing that +crosses here. + +The margin is the reason: CW2 turbo scores 10.4% WER in Russian against 27.5% for +the `ggml-small.bin` mavsttd loads, over 200 Golos clips +(`docs/evals/2026-08-09-crisperwhisper2-russian-wer.md`). Text-to-speech has not +moved and piper on homesrv is still the only synthesizer. + Speech-to-text stays two stages when it moves. One call carrying both a clip and the router prompt was measured on 05-08-2026. It scores 54.2% intent-only against 84.7% for whisper on homesrv, on the same 72 cases. The model transcribes clips it then routes wrong, so a long @@ -173,8 +192,9 @@ cleaner transcripts, not accuracy. See `docs/evals/2026-08-05-audio-in-routing.m which fixes what the 1.7B gets wrong: world knowledge, and the persona the CPT targets. The degradation path is already written and measured, since the classifier scores 68.8% full accuracy at p50 16.6µs on its own. -3. **Speech-to-text and text-to-speech** (#486). They gain a real margin, but on - quality alone, and both already work. +3. **Speech-to-text and text-to-speech** (#486). Speech-to-text is wired, see + above. Text-to-speech is not, and piper is good enough that nothing argues + for moving it yet. 4. **The wake word** (#487). Independent of all of the above. ## Assumptions