Compare commits
94 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 02d96e611d | |||
| 62eef01c18 | |||
| 1a8aed35b8 | |||
| ce6a6821a9 | |||
| 479b0c4475 | |||
| 877b1fd4f8 | |||
| 21a42cb3e6 | |||
| b8279f6a22 | |||
| 8c30971a96 | |||
| d0ea927ac3 | |||
| 9c7bafd5b1 | |||
| 1f1e002789 | |||
| ce91d20ac8 | |||
| 50130cdffb | |||
| a9b480a78f | |||
| c5264deb46 | |||
| 1cb269886e | |||
| 7a9b9cc669 | |||
| 31b5093403 | |||
| 229890abd7 | |||
| 0db9ca084c | |||
| 96d97e8964 | |||
| 02f6e8ad4a | |||
| 999a5ad562 | |||
| 2ea39a3d41 | |||
| 8fb6f2154d | |||
| c938148619 | |||
| 2512d686a1 | |||
| f44abcc526 | |||
| a99932b427 | |||
| 6d5801bb1f | |||
| 8aba4845bf | |||
| 2ec92ee8bf | |||
| 7c77a378c1 | |||
| 672eabc134 | |||
| a1a2fa3704 | |||
| 22a4978459 | |||
| 1456336652 | |||
| b975716759 | |||
| 4b1edb0617 | |||
| 944e553669 | |||
| cc32c2c4ab | |||
| a1e97c94ac | |||
| c7f59e48f4 | |||
| 4666057066 | |||
| 7138086c3f | |||
| 83e168f326 | |||
| a4abcdefa3 | |||
| 68a3c85186 | |||
| 88c086482e | |||
| feabf9f350 | |||
| 50c6637c1b | |||
| ee9d55ca95 | |||
| a886217223 | |||
| de9884e063 | |||
| bbefda66e2 | |||
| a37c4138a1 | |||
| d6f391430f | |||
| 68b2aa9137 | |||
| f8fa0d1b44 | |||
| 9a333b23d7 | |||
| 663b5c47b9 | |||
| 45c521e1a6 | |||
| e34669a52e | |||
| d434f83c2c | |||
| c310115fd2 | |||
| 3024f76e5f | |||
| 6bc71553ab | |||
| ed1730431c | |||
| 01e80fce4a | |||
| e69f1bd0cf | |||
| f55bedee2e | |||
| e470435cf1 | |||
| 00f9239ef9 | |||
| 3513e508b7 | |||
| c15c2b7bd2 | |||
| b6eaa704a2 | |||
| 2597a7b34a | |||
| 7203cd56fd | |||
| ab3e818bb9 | |||
| b5500a5be8 | |||
| 9095ac847d | |||
| 4b5f6adbae | |||
| ecb8ba72eb | |||
| 4fdecf9a25 | |||
| 2bbd8edbf6 | |||
| beb093aebb | |||
| b1b326018f | |||
| 08889cad88 | |||
| a4630b9314 | |||
| 39d44bb384 | |||
| 65ee0f9c61 | |||
| 76938e206d | |||
| 0b3d81ecbf |
@@ -70,3 +70,5 @@ coverage.out
|
||||
|
||||
# root .env — MAVEN_AMBIENT_TOKEN and friends, same class as deploy/telegram.env
|
||||
.env
|
||||
# silero-vad, downloaded (see AGENTS.md)
|
||||
/models/vad/
|
||||
|
||||
@@ -95,6 +95,22 @@ model: the code puts `query: ` in front of a question and `passage: ` in front
|
||||
of a stored note, which is how e5 was trained. The quantized file is the one
|
||||
that is downloaded, deployed and measured.
|
||||
|
||||
## Voice activity model for mavwaked
|
||||
|
||||
`mavwaked` decides an utterance has started with silero-vad when `-vad-model`
|
||||
points at it, and with an energy threshold when it does not. The model is 2.3MB
|
||||
and is not committed:
|
||||
|
||||
```sh
|
||||
mkdir -p models/vad
|
||||
curl -sL -o models/vad/silero_vad.onnx \
|
||||
https://github.com/snakers4/silero-vad/raw/master/src/silero_vad/data/silero_vad.onnx
|
||||
```
|
||||
|
||||
It needs the same `libonnxruntime.so` the embedder needs, passed as `-onnx-lib`
|
||||
or read from `MAVEN_ONNX_LIB`. The measurement is
|
||||
`docs/evals/2026-08-09-silero-vad.md`, and the tests skip without the file.
|
||||
|
||||
**Also need ONNX Runtime** (`libonnxruntime.so`):
|
||||
|
||||
```sh
|
||||
|
||||
@@ -1,517 +1,200 @@
|
||||
# CLAUDE.md
|
||||
|
||||
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
|
||||
Guidance for Claude Code (claude.ai/code) working in this repository.
|
||||
|
||||
Maven is a self-hosted, privacy-first voice assistant (Russian + English). Go daemons
|
||||
talking over unix sockets; one resident small model for routing + phrasing; whisper.cpp STT, piper TTS.
|
||||
Deploy target is a Ryzen laptop (homesrv) with Vulkan offload to the Vega iGPU (`n_gpu_layers: 99`,
|
||||
compose passes `/dev/dri` + the render gid) — the resident model stays ≤1.7B either way.
|
||||
**This is a rules file.** It loads into every session, so it carries only what
|
||||
changes what an agent does. A measurement belongs in `docs/evals/`, dated and
|
||||
never edited after the day. A subsystem's reasoning belongs in its living doc
|
||||
under `docs/`. Read that doc before changing the subsystem.
|
||||
|
||||
**Resident model:** currently **Qwen3-1.7B** (`UD-Q4_K_XL`), stock — not yet the CPT'd one.
|
||||
It replaced Qwen3.5-0.8B on 2026-07-31 because it measured better on both fixtures we have:
|
||||
67.5% vs 59.7% intent-only on the 77-case RU routing fixture, and 20/27 vs 11-17/27 on the
|
||||
talk fixture. See `docs/evals/2026-07-31-model-bakeoff.md`. It is a Thinking variant, so `n_ctx` is 4096
|
||||
— reasoning tokens need the room, and 4096 is what the scores above were measured at.
|
||||
| Read this | Before |
|
||||
|---|---|
|
||||
| `docs/routing.md` | touching `internal/router/` or `queryWalk` |
|
||||
| `docs/deployment.md` | touching a daemon, compose, a systemd unit or the web UI |
|
||||
| `docs/offload.md` | touching a daemon seam or adding a model caller |
|
||||
| `docs/world.md` | touching search, Kiwix or the world chain |
|
||||
| `docs/language.md` | changing a prompt contract or a Russian word list |
|
||||
| `docs/ecosystem.md` | touching Nexus, Praxis or Hexis |
|
||||
| `docs/rearchitecture.md`, `docs/design.md` | changing the shape of anything |
|
||||
| `docs/workflow.md` | the five stores, the doc tiers, the guards |
|
||||
| `AGENTS.md` | local preview, screenshots, model downloads |
|
||||
|
||||
The **target** is still the locally CPT'd **Qwen3-1.7B** (Vikunja #122, training in flight).
|
||||
Stock already speaks good Russian; what it gets wrong is the persona — it writes `я рад`,
|
||||
masculine, where Maven needs `рада`. That is what the CPT is for.
|
||||
## What Maven is
|
||||
|
||||
**Do not bother with sub-500M models.** LFM2.5-230M and 350M were measured on 2026-07-31 and
|
||||
both are unusable in Russian: the 350M routes at 5.2% (worse than guessing) and answers
|
||||
"столица Франции?" with the invented non-word "Сторзит"; the 230M replies to Russian in
|
||||
Spanish. Their strong published IFEval/BFCL numbers are English-only. Model files live in
|
||||
`/mnt/hdd1/llms`, bind-mounted to `/opt/maven/models/llm` — which **shadows** the repo's
|
||||
`models/llm/`, so the LFM2.5 gguf sitting there is not loaded by anything. Swapping the resident
|
||||
model is a one-line change to `phraser.model_path` in `deploy/mavend.json`.
|
||||
A self-hosted, privacy-first voice assistant in Russian and English. Go daemons
|
||||
talk over unix sockets. One resident small model routes and phrases. whisper.cpp
|
||||
does speech-to-text and piper does text-to-speech.
|
||||
|
||||
See `docs/rearchitecture.md` for the target architecture, `docs/design.md` for the folded design spec, and
|
||||
`AGENTS.md` for local-preview + model-download recipes.
|
||||
The resident model is **Qwen3-1.7B** (`UD-Q4_K_XL`) on homesrv, a Thinking
|
||||
variant at `n_ctx` 4096. Keep it at 1.7B or under. Sub-500M models are unusable
|
||||
in Russian (`docs/evals/2026-07-31-model-bakeoff.md`). Model files live in
|
||||
`/mnt/hdd1/llms`, bind-mounted over the repo's `models/llm/`, so a gguf sitting
|
||||
in the repo is loaded by nothing.
|
||||
|
||||
**Model work is moving to the workstation** (owner's call, 2026-08-02). homesrv cannot grow a
|
||||
GPU and the workstation has 16GB of VRAM. So the resident model, STT and TTS become preferred
|
||||
remotes with a floor on homesrv. The workstation is never assumed up. Fall back silently when
|
||||
it would only do the job better. Name the gap when the 1.7B cannot do it at all. The embedder
|
||||
stays on homesrv permanently, because it backs that floor. It is multilingual-e5-small,
|
||||
quantized and asymmetric — `EmbedQuery` and `EmbedPassage` apply the `query:`/`passage:`
|
||||
prefixes it was trained with, and calling plain `Embed` on a note is a bug. It replaced
|
||||
MiniLM and bought ten points of recall@1 and 2.5× the speed; see
|
||||
`docs/evals/2026-08-04-recall-e5-small.md`. Read `docs/offload.md` before
|
||||
touching a daemon seam or adding a model caller. Vikunja #483 is the umbrella, #484 to #487
|
||||
are the work.
|
||||
The workstation is workpc and it holds the remote model and speech-to-text.
|
||||
**It is never assumed up.** **Fall back silently** when it would only do the job
|
||||
better. **Name the gap** when the resident model cannot do the job at all.
|
||||
|
||||
Both halves are wired as of 2026-08-03. Routing and replies prefer the workstation silently
|
||||
through `modelSeam`; nudge and reminder phrasing prefer it silently inside the phraser. A
|
||||
world question goes through `LLMPhraser.PhraseWorld` and names the gap when the card is not
|
||||
free — `worldGap` in `cmd/mavend/worldmodel.go`, which the owner hears instead of an invented
|
||||
answer. A box with no `workstation` block behaves exactly as it did before the seam: naming
|
||||
a gap requires a gap. The offload table in `docs/offload.md` says which caller is which.
|
||||
**The embedder stays on homesrv permanently**, because it backs that floor.
|
||||
`EmbedQuery` and `EmbedPassage` apply the `query:` and `passage:` prefixes
|
||||
multilingual-e5-small was trained with. Calling plain `Embed` on a note is a bug.
|
||||
|
||||
## Build & test
|
||||
## Build and test
|
||||
|
||||
CGO daemons (`mavend`, `mavsttd`, `mavttsd`, `mavenclient`) need the vendored toolchain
|
||||
and libs wired through the Makefile — **do not** call `go build` on them bare, use `make`:
|
||||
CGO daemons (`mavend`, `mavsttd`, `mavttsd`, `mavenclient`) need the vendored
|
||||
toolchain wired through the Makefile. **Do not call `go build` on them bare**,
|
||||
and **do not hand-write the CGO preamble**. This box runs zsh, so an unquoted
|
||||
`-run Test*` dies on "no matches found" before `go` is reached. `make t` also
|
||||
carries `-count=1` and sets `MAVEN_ONNX_LIB`. Without that variable the four
|
||||
`TestONNX*` measurements self-skip and the run still prints `ok`.
|
||||
|
||||
```sh
|
||||
make build # all 11 binaries
|
||||
make build-web # single daemon (pure-Go ones: web/waked/poll/caldav build without CGO)
|
||||
make test # go test -race across ./internal/... ./cmd/... with CGO env set
|
||||
make build # all 11 binaries. make build-web for one (web/waked/poll/caldav skip CGO)
|
||||
make test # go test -race across ./internal/... ./cmd/... with CGO env set
|
||||
make t PKG=./internal/router/eval/ RUN='TestONNX' V=1 # V=1 for -v, RACE=0 to drop -race
|
||||
```
|
||||
|
||||
Run a single test (must carry the CGO env for packages that touch STT/TTS/voice):
|
||||
## The daemons
|
||||
|
||||
```sh
|
||||
CGO_CFLAGS="-I$(pwd)/deps/include -I$(pwd)/deps/whisper.cpp/ggml/include" \
|
||||
CGO_LDFLAGS="-L$(pwd)/deps/lib -Wl,-rpath,$(pwd)/deps/lib" \
|
||||
LD_LIBRARY_PATH="$(pwd)/deps/lib" \
|
||||
deps/go/go/bin/go test -run TestName ./internal/router/
|
||||
```
|
||||
Eleven binaries under `cmd/`, wired socket-to-socket over `internal/ipc`, not
|
||||
linked. `mavend` is the core and owns the DB and the IPC socket.
|
||||
`deploy/mavend.json` sets sockets, model paths and the phraser and embedder
|
||||
blocks, with `${VAR}` expansion from gitignored `deploy/telegram.env`.
|
||||
**`docker-compose.yml` runs five**: `mavend`, `mavsttd`, `mavttsd`, `mavweb`,
|
||||
`mavpoll`. Count against compose, not against `make build`. `mavwaked` runs on
|
||||
workpc under systemd. `docs/deployment.md` says who else is absent and why.
|
||||
|
||||
Pure-Go packages (`router`, `memory`, `mavweb`, …) run under a plain `go test ./pkg/`.
|
||||
|
||||
## The daemons (`cmd/`)
|
||||
|
||||
| Binary | Role |
|
||||
|---|---|
|
||||
| `mavend` | **Core.** Router, phraser, memory, reminders, digestion tick. Owns the DB + IPC socket. |
|
||||
| `mavweb` | HTTP UI + PWA (`/dash`, `/history`, `/trace`, `/notifications`, `/tools`); WebAuthn auth. Connects to mavend's socket. |
|
||||
| `mavsttd` | Speech-to-text (whisper.cpp, CGO). |
|
||||
| `mavttsd` | Text-to-speech (piper subprocess). |
|
||||
| `mavwaked` | Wake-word / VAD gate. **Not on homesrv** — see below. |
|
||||
| `mavenclient` | Voice loop client (mic → stt → core → tts). **Not on homesrv** — see below. |
|
||||
| `mavpoll` | Environment poller: netdata alarms, uptime-kuma, zenmoney, wireguard presence. Writes facts, sends nothing. Telegram is `internal/delivery/telegramsink`, not this. |
|
||||
| `mavcaldav` | CalDAV calendar sync. |
|
||||
| `mavmaild` | Mail reader (IMAP, read-only). Holds the IMAP password; core never sees it. |
|
||||
| `mavgpud` | GPU supervisor. **Runs on workpc, not homesrv** — own unit, `deploy/mavgpud.service`. Keeps llama-server loaded while the card is free (V-488). Maven never asks it for anything, it reads `/health` through `llm.Pair`. |
|
||||
| `mavupdate` | Not a daemon. Operator CLI a human runs on the box to deploy a new build. |
|
||||
|
||||
Two more binaries have no Makefile target and are built with `go run` or `go build` when
|
||||
they are needed. Neither is deployed.
|
||||
|
||||
| Binary | Role |
|
||||
|---|---|
|
||||
| `mavseal` | Recovery tool. Encrypts a live tmpfs working copy back to the ciphertext file when mavend was killed before `defer st.Close()` sealed it. |
|
||||
| `labelgen` | Runs the stage 0 grammars over utterances and prints JSONL, the training data for the routing heads (V-546). |
|
||||
|
||||
Daemons are wired socket-to-socket, not linked. `internal/ipc` is the client/server wire
|
||||
protocol; the config in `deploy/mavend.json` (with `${VAR}` env expansion from gitignored
|
||||
`deploy/telegram.env`) sets socket paths, model paths, and the phraser/embedder blocks.
|
||||
|
||||
**`docker-compose.yml` runs five: `mavend`, `mavsttd`, `mavttsd`, `mavweb`, `mavpoll`.**
|
||||
Count against compose, not against the table. Four of the nine daemons are absent, and each
|
||||
absence has a different reason.
|
||||
|
||||
`mavmaild` and `mavcaldav` are commented out in compose, each with the reason written
|
||||
beside it: the first needs a mail account, the second a CalDAV account, and this box has
|
||||
neither. `mavcaldav` used to appear nowhere at all, which was an oversight; it became a
|
||||
recorded decision on 07-08-2026 (V-644). Two things ride on that absence and the block
|
||||
names them. Agenda questions route to `IntentQuery` at stage 0 (V-498) and the `calendar`
|
||||
query source then reads a table nobody writes. And `loop.State.CalendarBusy` is fed by the
|
||||
same facts, so the gate's "do not nag mid-meeting" is permanently false. Its password is
|
||||
read from a file (`-pass-file`, and `-render-pass-file` for the render collection), never
|
||||
taken as a flag value, which is the rule `mavpoll` and `mavmaild` follow too.
|
||||
|
||||
**`mavwaked` and `mavenclient` are absent by decision, not oversight** (Vikunja #463,
|
||||
`docs/plans/17-where-the-voice-loop-runs.md`).
|
||||
homesrv has a microphone — it is a laptop — but it is in the wrong room, so a wake-word
|
||||
daemon there listens to nobody. They belong on a client machine where the owner is standing.
|
||||
|
||||
**That machine is workpc** (owner's correction, 2026-08-05). This section used to say no
|
||||
such machine existed, which was written when the workstation was only a model host. It is
|
||||
where he sits most of the day and it has the microphone. `ipc.Dial` already takes
|
||||
`tcp://host:port?token=...` through the netaddr seam, so the two daemons need deploying,
|
||||
not building. V-515 is that deployment.
|
||||
|
||||
Until they are deployed, **the wake word and the VAD gate are covered by unit tests and by
|
||||
nothing else**, and push-to-talk through `/dash` is what QA actually covers. Note that
|
||||
deploying them does not by itself prove a wake word: `mavwaked` gates on energy and has no
|
||||
keyword model (V-487), so the loop runs open until that lands.
|
||||
- **Passwords are read from files, never taken as flag values.**
|
||||
- **The voice wire is plaintext with no auth.** mavend's voice port stays on
|
||||
homesrv loopback and reaches workpc over ssh. Do not LAN-bind it.
|
||||
`SurfaceVoice` caps acts at L0, and L0 does not cap reading.
|
||||
- **A GPU service added beside mavgpud goes in `cmd/mavgpud`, never in systemd.**
|
||||
The card needs one owner. A second unit made mavgpud evict llama-server every
|
||||
few seconds and took the model arm down for eight minutes.
|
||||
|
||||
## The ecosystem: Nexus, Praxis, Hexis
|
||||
|
||||
Maven is one of four services. It owns conversation and personal memory. It does not
|
||||
Nexus identifies, Praxis observes, Hexis acts, Maven understands. Maven does not
|
||||
own identity, operational state, or execution. Full contract in
|
||||
`docs/ecosystem.md`.
|
||||
`docs/ecosystem.md`. All three are `nil` unless configured and each degrades
|
||||
alone. An outage means a named gap, never a broken turn or a guess.
|
||||
|
||||
```text
|
||||
Nexus identifies. Praxis observes. Hexis acts. Maven understands and coordinates.
|
||||
```
|
||||
|
||||
| Service | Owns | Maven's client | Configured at |
|
||||
|---|---|---|---|
|
||||
| **Nexus** | Canonical entity ids, names, aliases, relationships. Projects, services, devices, people, pets, places. | `nexusClient` in `cmd/mavend/ecosystem.go`, `POST /api/v1/resolve` | `nexus.url` (`http://nexus:9740`) |
|
||||
| **Praxis** | Operational attention and item lifecycle. What needs looking at, what changed, what is still unresolved. | `praxisClient`, the HTTP tools API under `/api/v1/tools/` | `praxis.url` (`http://praxis:8989`) |
|
||||
| **Hexis** | The capability registry and the only path to executing anything. | vendored `github.com/kami/hexis/pkg/client` | `hexis.url` (`http://hexis:9741`) |
|
||||
|
||||
All three are `nil` unless configured, and every one of them degrades on its own.
|
||||
An outage means a named gap in the answer, never a broken turn and never a guess.
|
||||
|
||||
Rules that are not negotiable:
|
||||
|
||||
- **No component reads another component's database.** Praxis attention comes over
|
||||
HTTP, never from its SQLite file.
|
||||
- **Identity lives in Nexus.** Do not invent a local fact key for something Nexus
|
||||
resolves. `actionFact` already sets `Subject`, and `cmd/mavend/factenrichment.go`
|
||||
resolves it in the background against Nexus.
|
||||
- **Free text never reaches a mutating Hexis call.** Resolve to a canonical entity id
|
||||
first. Ambiguous resolution asks the owner, it does not pick.
|
||||
- **No component reads another component's database.** Praxis attention comes
|
||||
over HTTP, never from its SQLite file.
|
||||
- **Identity lives in Nexus.** Do not invent a local fact key for something
|
||||
Nexus resolves. `cmd/mavend/factenrichment.go` resolves `actionFact.Subject`.
|
||||
- **Free text never reaches a mutating Hexis call.** Resolve to a canonical
|
||||
entity id first. Ambiguous resolution asks the owner, it does not pick.
|
||||
- **LLM output is not authorization.** Confirmation binds capability id, target
|
||||
entity, arguments, requester and expiry. See `cmd/mavend/confirm.go`.
|
||||
- **Praxis lifecycle words mean different things.** Surfaced is not acknowledged,
|
||||
acknowledged is not resolved, execution success is not recovery. Reading an item
|
||||
aloud calls `Surface`, never `Acknowledge`.
|
||||
- **No automatic attention-to-action path.** Digestion may summarise Praxis. It may
|
||||
not call Hexis.
|
||||
entity, arguments, requester and expiry (`cmd/mavend/confirm.go`).
|
||||
- **Praxis lifecycle words differ.** Surfaced is not acknowledged, acknowledged
|
||||
is not resolved, execution success is not recovery. Reading an item aloud
|
||||
calls `Surface`, never `Acknowledge`.
|
||||
- **No automatic attention-to-action path.** Digestion may summarise Praxis and
|
||||
may not call Hexis.
|
||||
- Every cross-service call carries a correlation id minted once per action
|
||||
(`withCorrelationID`), a contract version header, and `X-Requested-By: maven`.
|
||||
|
||||
Every cross-service call carries a correlation id minted once per action
|
||||
(`withCorrelationID`), a contract version header, and `X-Requested-By: maven`.
|
||||
## Routing
|
||||
|
||||
## Routing — read this before touching the router
|
||||
**Read `docs/routing.md` before touching `internal/router/` or `queryWalk`.** It
|
||||
carries the reasoning, the measurements and every rule's why. A route produces
|
||||
two decisions. **Intent** is one of seven values. **Source** is where the answer
|
||||
lives and is read on `IntentQuery` alone. Score them separately. The cascade is
|
||||
stage 0 grammars, then the routing heads, then the resident model, then the
|
||||
classifier. Every stage may decline and the next one answers.
|
||||
|
||||
`internal/router/` has TWO layered engines. **The LLM router is now the default and it is
|
||||
on in deploy** — this section used to say it was wired `nil`, which stopped being true on
|
||||
2026-07-31.
|
||||
- **The classifier is the floor, not dead code.** It answers when the resident
|
||||
model is off, absent, or erroring. **Any model error falls through.**
|
||||
- **`baselineGrammars` in `eval_test.go` mirrors `buildRouter`.** A grammar
|
||||
added to one belongs in both, or the fixture scores a set nobody runs.
|
||||
- **Go's `\b` is ASCII-only** and never fires after a Cyrillic letter. A Russian
|
||||
pattern needs an explicit `(\s|[?!.]|$)`.
|
||||
- **`PraxisGrammars()` is the only path to Praxis**, not a faster one.
|
||||
- **`voice.embedder.heads_path` must never point at `model_path`.** Recall
|
||||
depends on the resident e5-small scoring what it scored. Fine-tune a copy.
|
||||
- **Routing traces are retained 14 days**, enforced on write and again on start.
|
||||
- **Bump `tokenizerRev` on any change to what `encodeWord` emits**, so a
|
||||
tokenizer fix triggers `ReembedAll` the way swapping the model file does.
|
||||
- **A new rung in the `runTurn` ladder needs its name in `preRouteLadder`**
|
||||
(`cmd/mavend/decisiontrace.go`), or it is missing from the decision record.
|
||||
|
||||
- **LLM router (the intended design, docs/rearchitecture.md):** the resident Qwen3-1.7B (`llmrouter.go`)
|
||||
emits GBNF-constrained structured JSON, and the SAME model phrases replies. Embedder is
|
||||
demoted from a routing gate to a RAG hint. Wired at `voice.go:214` via
|
||||
`pickLLMRouter(cfg.Voice.UseLLMRouter(), llmClient)`; the flag is `voice.llm_router`
|
||||
(`config.go`), `DefaultLLMRouter` is **on**, and `deploy/mavend.json` sets it `true`.
|
||||
- **Classifier cascade (the failure floor, not dead code):** `classifier.go` +
|
||||
`embedder.go` nearest-neighbour over frozen seed phrases. It runs when the LLM router is
|
||||
off, when there is no llama-server to talk to (`pickLLMRouter` logs that and degrades),
|
||||
and on any per-turn LLM error. Do not delete it — routing by seed similarity is the known
|
||||
cause of weak RU query handling, but a turn must never break on the model.
|
||||
**`queryWalk` takes query sources out and moves none** (`actions_query.go`). That
|
||||
is the safety argument and it is not negotiable. The table's order is
|
||||
load-bearing and carries "the owner's data first, then the world".
|
||||
`SourceUnknown` is the floor and walks the whole chain. A named destination
|
||||
removes only the sources marked `guesses: true`, so a source that looks rather
|
||||
than guesses is always asked. **The personal boundary is the one exception and
|
||||
it is deliberate.** It guesses, so naming `SourceWorld` drops it. **Only a stage
|
||||
0 grammar may drop it** (owner's call, V-666). `queryWalk` reads
|
||||
`Decision.SourceAnchored` for the source marked `boundary: true` and no other.
|
||||
|
||||
Cascade order: `stage0.go` exact-match fast-path → LLM router (when non-nil) → classifier
|
||||
fallback. Any LLM error falls through to the classifier so a turn never breaks on the model.
|
||||
Judge a routing change against the classifier (76.0% intent, 36.4% destination)
|
||||
and the resident model (80.2% intent), since those always answer. The fixture
|
||||
has grown from 77 cases to 96, so a number compares only to another number on
|
||||
the same fixture.
|
||||
|
||||
**A stage-0 decision is slot-extracted too, since 06-08-2026** (V-572). `fillMatchedSlots`
|
||||
in `router.go` runs the stage-2 extractor over whatever a grammar built and fills only the
|
||||
slots it left empty — a matched value always wins, because the rule read a literal pattern
|
||||
and the extractor guesses. It did not run before, so `ReminderGrammar` handed the daemon
|
||||
`HasTime: false` for "напомни в 11:00 позвонить маме" and `missingFor` read the silence as
|
||||
absence and asked "Когда?". It applies to every grammar and is inert for all but the
|
||||
reminder: `Extract` fills Time, Fn and Key and nothing else, and the query, clock, agenda,
|
||||
feed, list, task and narrative rules all emit intents with no such slot. Benchmarked at
|
||||
20000x, a stage-0 query costs 3.7µs against 3.9µs before. **`Slots.Text` is deliberately not
|
||||
filled** — a grammar that left it empty meant it, and `agendaQueryBuild` hands the query
|
||||
chain the utterance itself. Fixture unchanged at 64/91, with "slots deferred to daemon"
|
||||
6 → 0.
|
||||
## Language: model output and Russian
|
||||
|
||||
Measured on the 77-case RU fixture. **Re-measured 2026-08-02: the classifier scores 68.8%
|
||||
full accuracy at p50 16.6µs**, not the 36.8% at p50 31ms that stood here from
|
||||
`docs/evals/2026-07-31-model-bakeoff.md`. That older figure predates the stage 0 rules and the
|
||||
seed additions, both of which now score inside the classifier baseline. Qwen3-1.7B scores
|
||||
77.9% intent-only / 72.7% through the cascade. So the router buys about 4 points of accuracy,
|
||||
not a doubling, and the trade is worth re-arguing rather than assuming. **The ≈2.7s figure
|
||||
that stood here until 2026-08-02 was contention, not the model.** See `docs/evals/2026-07-31-routing.md` line 61, which measures the LLM router at
|
||||
p50 825ms / p95 1.2s / max 3.0s and the full cascade at p50 0.80-1.04s. Do not plan latency
|
||||
work off the bakeoff table.
|
||||
Both contracts are in `docs/language.md`. What must not be broken:
|
||||
|
||||
**Re-measured 2026-08-05 on the fixture as it now stands, 91 cases** (V-320 item 2,
|
||||
`docs/evals/2026-08-05-routing-resident-model.md`): cascade + resident model scores
|
||||
**75.8% full / 80.2% intent-only at p50 1.19s / p95 1.65s**. That is a new baseline and not
|
||||
a movement, because 14 cases were added since the 77-case number above. The model alone
|
||||
scores 37.4% full against 61.5% intent-only, and the gap is slots rather than routing: it
|
||||
routes `reminder` and leaves the time to the daemon, which is what the contract asks. To
|
||||
re-run it, start a **second** llama-server on a fixed host port — the resident one binds
|
||||
`--port 0` inside the container and no host process can reach it.
|
||||
- **One parser for model text, `parseResponseMood`** in
|
||||
`internal/phraser/parse.go`. Every phrasing path reaches it. Mood is an enum.
|
||||
- **The router prompt is a separate contract** over 7 intents, and
|
||||
`llm/check_prompt_parity.py` keeps the Go and relabelling copies identical.
|
||||
- **Russian words are matched by three mechanisms and no fourth**:
|
||||
`internal/lexicon` for closed classes, `internal/morph` for grammar, and
|
||||
`cmd/mavend/topics.go` with the embedder for open sets. A regex whose output
|
||||
is a fact or a route is the defect. A regex over structured input is not.
|
||||
- **Seeds are scoring data.** Editing one moves a recogniser and must be
|
||||
re-measured against the `TestONNX*` tests, not eyeballed.
|
||||
|
||||
**The numbers above are the homesrv floor, not the ceiling.** With the workstation up, routing
|
||||
completes through `llm.Pair` against gemma-4-12b and scores **84.4% full / 93.5% intent-only at
|
||||
p50 329ms** — better than the resident model and about 2.5× faster (`docs/evals/2026-08-02-workstation-gemma4-12b.md`,
|
||||
Vikunja #485). The workstation is never assumed up, so both sets of numbers are live. Judge a
|
||||
routing change against the classifier and the resident model, since those are what always answer.
|
||||
## Non-goals and hard constraints
|
||||
|
||||
**The intended third engine is not a generative model** (owner's call, 05-08-2026, V-546,
|
||||
`docs/plans/18-routing-heads-on-e5-small.md`). Routing has a bounded output space, so it is
|
||||
classification, and the 118M multilingual-e5-small is already resident. Three heads on one
|
||||
forward pass: intent, mood, and BIO slot tags. Roughly 5e15 FLOPs to train, so 10 to 30
|
||||
minutes on the workstation. A 100M decoder from scratch is 10 to 20 GPU hours. Two things
|
||||
it buys that a decoder cannot. No grammar is needed, because a softmax cannot emit a value
|
||||
that does not exist. And max softmax is a calibratable confidence, where `Confidence: 1.0`
|
||||
was a hardcode. **Fine-tune a copy of the weights.** The resident embedder backs memory
|
||||
recall. Training it in place couples routing accuracy to recall@1, with nothing in the
|
||||
suite to name the trade.
|
||||
Not a nag, not autonomous.
|
||||
**The persona is feminine.** Russian self-reference takes feminine forms: `рада`
|
||||
not `рад`, `поняла` not `понял`. The owner is male and she speaks to him
|
||||
informally. Use "ты", singular, never "вы" or "ваш", and never "он" or "его".
|
||||
She talks TO the owner, not about him. Pet names such as "милый" are forbidden.
|
||||
The name "Ками" is not. `CheckAddress`, `CheckFeminine` and `CheckCringe` in
|
||||
`internal/phraser/eval/checks.go` enforce this, scored by `make eval-phrasing`.
|
||||
|
||||
`Confidence: 1.0` used to be hardcoded in `llmrouter.go`, so the LLM
|
||||
path could never ask for clarification (6/6 refusal cases missed on the fixture) — Vikunja
|
||||
#359. Fixed 31-07-2026 with structural signal (single-token utterance, keyless fact, act with
|
||||
no allowlisted fn) feeding the same stage-3 gate the classifier path already had — see
|
||||
`gateLLMDecision` in `router.go`. Note the second half of that bug: the LLM branch never
|
||||
consulted `r.threshold` at all, so a correct low confidence would have been discarded anyway.
|
||||
**"Never phones home" is deprecated** (owner's call, 2026-07-31). She reads
|
||||
external sources, and `docs/world.md` carries that chain. What holds regardless:
|
||||
|
||||
Re-measured on the fixture after the fix: **missed clarify 6/6 → 1**, at the cost of 3 false
|
||||
clarifies and 2.6pt of full accuracy (72.7% → 70.1%, intent-only 67.5% → 74.0%). Two of the
|
||||
three false clarifies are acts the model mis-routed and the gate caught — asking beats wrongly
|
||||
executing, so the fixture and the daemon disagree about what is correct there. The third,
|
||||
`"поужинал"`, was a real defect: the single-token rule was an English intuition and does not
|
||||
transfer to Russian, where one word is routinely a whole sentence.
|
||||
|
||||
Narrowed 01-08-2026. `thinSingleToken` (`internal/router/singletoken.go`) still thins a bare
|
||||
one-word nominal — "вода", "бэкап" — but spares two classes: a closed lexicon of social and
|
||||
control singles ("привет", "спасибо", "стоп", "yes"), and any token carrying a Russian verb
|
||||
ending (past tense, 2nd person, reflexive), because a verb already contains its subject. Both
|
||||
tests are offline and cost nothing. Re-measured: **false clarifies 3 → 2, intent-only 74.0% →
|
||||
75.3%, full accuracy unchanged at 70.1%, missed clarify still 1.** The two remaining false
|
||||
clarifies are the act-with-no-allowlisted-fn arm of the gate, not this rule.
|
||||
|
||||
Agenda questions taken off the model, 01-08-2026. `AgendaQueryGrammars` (`stage0.go`, wired
|
||||
after the clock rules in `buildRouter`) routes "что у меня сегодня", "во сколько у меня
|
||||
встреча" and anything naming a calendar to `IntentQuery` at stage 0. They were going to
|
||||
`IntentSystem`, where `replySystem` has no agenda arm and answered "пока не умею" — the
|
||||
fixture had said `query` since ru-query-019 was written. Measured: **full accuracy 70.1% →
|
||||
72.7%, intent-only 75.3% → 77.9%, calendar 0/2 → 2/2**, clarify counts unchanged. Note that
|
||||
Go's `\b` is ASCII-only and never fires after a Cyrillic letter; the pattern needs an
|
||||
explicit `(\s|[?!.]|$)`.
|
||||
|
||||
Two more shapes taken off the model, 04-08-2026 (V-498). `rest-of-day-query` inside
|
||||
`AgendaQueryGrammars` claims "что дальше?" / "what's next", and `NarrativeQueryGrammar`
|
||||
(`stage0.go`, wired **last** in `buildRouter`, after the capture marker) claims "расскажи про
|
||||
X", "объясни X", "опиши X". Neither carries a question mark or an interrogative, so the model
|
||||
called both `IntentFact`; the write was caught downstream by `IsQuestionShaped`, so this was a
|
||||
latency and fixture defect, not a correctness one. The narrative rule reads the same
|
||||
`narrativeRequests` lexicon `IsQuestionShaped` reads, and declines `chatNarrativeTopics` — a
|
||||
joke, a bedtime story, herself — because the query chain has no source that answers those.
|
||||
New fixture cases ru-query-024 and ru-query-025. Classifier + ONNX baseline **56/80 (70.0%) →
|
||||
58/82 (70.7%)**, no case regressed, no new false clarify. The LLM arm was not measured (no
|
||||
llama-server in that run), so judge it again before quoting a cascade number.
|
||||
|
||||
Praxis taken off the model, 05-08-2026 (V-516). `PraxisGrammars()`
|
||||
(`internal/router/praxis.go`, wired in `buildRouter` before the capture marker because
|
||||
"отметь" is a capture verb) fills `Slots.Fn` with a Praxis capability name.
|
||||
**These grammars are the only path to Praxis, not a faster one.** Measured
|
||||
2026-08-05 with the resident model as router (V-517,
|
||||
`docs/evals/2026-08-05-reach-llm-router.md`): the model alone reaches Praxis
|
||||
**0/12**, the same as the classifier alone, because nothing in the router
|
||||
prompt names a Praxis capability and there is no string for it to write.
|
||||
Through the cascade it is 11/12. Deleting these rules costs every point. Praxis reach
|
||||
was **0/12 and structurally so**: `handlePraxisAct` compares `Slots.Fn` to a capability
|
||||
alias, and that slot is filled from the deployment's enabled tool names, which no Praxis
|
||||
alias is on. Measured **16/30 → 27/30 overall, praxis 0/12 → 11/12, lifecycle 0/5 → 5/5**
|
||||
(`docs/evals/2026-08-05-praxis-reach.md`). Two rules to know before editing: a **stative**
|
||||
lifecycle word ("готово", "принято") needs an item named beside it, while a bare
|
||||
**imperative** ("закрывай") may ask which one. The bare arm additionally requires that
|
||||
the sentence name no object of its own, or "закрой шторы в комнате" goes to Praxis instead
|
||||
of the house. A demonstrative ("отметь это как сделанное") resolves against
|
||||
`h.surfacedItems` only when exactly one item was spoken. Otherwise the turn goes back to
|
||||
the cascade rather than transitioning the wrong item.
|
||||
|
||||
**Who claimed a turn is now recorded, and so is who did not** (V-564, umbrella
|
||||
V-558). Arbitration between the claimants on the utterance stream is order,
|
||||
hardcoded in the pre-route resolver ladder, in `buildRouter` and in
|
||||
`querySources`. `internal/decision` records one `Record` per turn: every
|
||||
claimant, what it would have made the turn, the score it reported, and whether
|
||||
it won, declined, lost on score, was thinned by a gate or was **never asked**.
|
||||
The record rides the context, the same seam `querysource.go` uses, so a claim
|
||||
site cannot change a route and a context with no record costs nothing. It is
|
||||
installed in `runTurn`, so the mic, telegram and the web all leave the same
|
||||
trail. Storage is a 25-turn in-memory ring on the handler (`decision.Ring`),
|
||||
read over `ipc.TurnDecisions` and rendered as the second table on `/trace`.
|
||||
**It also persists, since 06-08-2026, and that reverses what this section used to
|
||||
say** (V-629, `docs/plans/21-persisting-the-routing-trace.md`). The old rule was
|
||||
that nothing persists, because a turn record is read minutes later or never. The
|
||||
owner reversed it: the routing heads (V-546) cannot be fitted or calibrated
|
||||
without real utterances, and 9 of the 31 modes in `internal/modes` have no seed
|
||||
example at all. The ring did not move. It is still what `/trace` reads and still
|
||||
what a test with no store gets. `cmd/mavend/routingtrace.go` is a second sink
|
||||
beside it, writing `routing_traces` (migration #23). The utterance is stored in
|
||||
clear, because a 384-dimension vector of a short sentence is substantially
|
||||
recoverable and storing vectors instead would be a privacy claim we cannot
|
||||
support. What makes it safe is the same thing that makes the fact store safe.
|
||||
Retention is 14 days, enforced on write and again on start, so a box that goes
|
||||
quiet does not keep every row. Nothing reads it outward, and the rule
|
||||
that his notes and facts are never search input covers this table. `Store.Wipe`
|
||||
deletes it with everything else. A correction (V-630) is promoted out into a
|
||||
seed-shaped row in `routing_labels` (migration #24) and kept, because a label is
|
||||
not a transcript. The transcript still expires. The gesture that writes one is
|
||||
two buttons beside the reply on `/chat`, reached over `ipc.CorrectTurn` and the
|
||||
trace id that now rides back on `ipc.ChatReply`. A turn marked wrong with no
|
||||
target is a usable negative, so naming the intent is never required. The target
|
||||
is one of the seven intents and never free text. **All three reaches offer it as
|
||||
of 06-08-2026**, and this section used to say only `/chat` did. Voice is the
|
||||
`repair` rung, which has read spoken corrections since V-455 and now writes the
|
||||
durable label beside the classifier seed it always wrote; a spoken negative with
|
||||
no target is its own rung, `repair-negative` (V-636, `docs/plans/22-correcting-a-turn.md`).
|
||||
Telegram is an inline keyboard under the reply, and it needed the chat to become
|
||||
readable first — **telegram is no longer outbound only** (V-637,
|
||||
`docs/plans/23-inbound-telegram.md`). The poller is dark unless the `telegram`
|
||||
block says `intake`, it long-polls because the box takes no inbound connections,
|
||||
it accepts `chat_id` and no other sender, and it drops whatever queued while the
|
||||
daemon was down. It reaches the daemon through `ipc.CoreAPI` alone, so a chat
|
||||
turn takes the path `POST /api/chat` takes. Note that the turn source is still
|
||||
`tap:text` for both, so provenance cannot tell a chat turn from a typed one.
|
||||
Adding a rung to the ladder
|
||||
in `runTurn` means adding its name to `preRouteLadder` in
|
||||
`cmd/mavend/decisiontrace.go`, or that rung is silently missing from the record.
|
||||
|
||||
## LLM output contract
|
||||
|
||||
All phrasing paths emit `{"response":"...","mood":"..."}`, with fallback to plain text when
|
||||
the model skips the JSON. **One parser, `parseResponseMood` in
|
||||
`internal/phraser/parse.go`**, and every path reaches it: the six `LLMPhraser` methods,
|
||||
`PhraseWorld`, and `Replier.PhraseReply`, which `cmd/mavend/replier_llm.go` wraps — that file
|
||||
holds the stub fallback and no parsing of its own. The legacy `{"body","summary"}` fallback
|
||||
was deleted on 2026-08-06 (V-397): it was the contract before `{"response","mood"}` replaced
|
||||
it, no prompt asks for that shape, the GBNF cannot emit it, and no test covered it.
|
||||
Mood is a fixed enum. Router prompt is a separate contract:
|
||||
`[{"intent":<enum>, key?, value?, text?, verb?}, ...]`, 7 intents (`fact, reminder,
|
||||
note, query, act, chat, system`). `llm/check_prompt_parity.py` in the training
|
||||
workspace enforces that the Go and relabelling prompts remain identical.
|
||||
|
||||
## Russian patterns — three mechanisms, no fourth
|
||||
|
||||
Hand-written Russian stem patterns were swept out on 2026-08-04 (owner's call: not a
|
||||
pattern, and the resident model cannot be asked per turn either). A regex whose output is a
|
||||
fact or a route is the defect; a regex over structured input — HTML, MIME, JSON, a URL, an
|
||||
argv list — is not. Before writing a Russian word list, pick one of these:
|
||||
|
||||
- **`internal/lexicon`** — closed classes, in `lexicon_ru_v1.json`. Interrogatives,
|
||||
capture verbs, reminder verbs, cardinals, day offsets, parts of day, weekdays, months,
|
||||
spoken hours. Editing a word is a data change, and there is exactly one copy: months used
|
||||
to live in three files. Cardinals carry the oblique forms, because a spoken time declines
|
||||
and `в семь` / `к семи` are one hour.
|
||||
- **`internal/morph`** — grammar, from the vendored golem Russian dictionary. `IsVerbForm`
|
||||
and `SameWord`. Note that lemma matching is BROADER than stem-plus-one-ending, so a verb
|
||||
slot that means the imperative must be matched exactly — `говори` and `говорил` are one
|
||||
lemma and only one of them is a command (`cmd/mavend/quiet_toggle.go`).
|
||||
- **`cmd/mavend/topics.go` and the embedder** — open sets, where the question is what a
|
||||
turn is ABOUT. Frozen seeds per subject plus a real `other` class, scored against the
|
||||
turn's own query vector. Same shape as the personal boundary in `personalboundary.go`,
|
||||
with one difference: a topic must clear the runner-up by `topicMargin`, because a false
|
||||
claim here spends a network scan rather than one honest "не знаю". The old keyword tests
|
||||
stay as the offline floor and may remain narrow, since they are no longer the only answer.
|
||||
- **The ecosystem trio** — when the answer is not in the utterance at all. Identity is
|
||||
Nexus's, never a local pattern.
|
||||
|
||||
Seeds are scoring data. Editing one moves a recogniser and must be re-measured against the
|
||||
`TestONNX*` tests, not eyeballed.
|
||||
|
||||
## Non-goals (hard constraints)
|
||||
|
||||
Not a nag, not autonomous. Maven's persona is **feminine** — Russian
|
||||
self-reference must use feminine forms — `рада`, not `рад`; `поняла`, not `понял`. The owner
|
||||
is male and is addressed informally: "ты", singular, never "вы"/"ваш" and never "он"/"его"
|
||||
(she talks TO the owner, not about the owner). Pet names ("милый", "дорогой") are forbidden; the name
|
||||
("Ками") is not. The eval enforces this: `CheckAddress`, `CheckFeminine` and `CheckCringe` in
|
||||
`internal/phraser/eval/checks.go`, scored by `make eval-phrasing`.
|
||||
|
||||
**"Never phones home" is DEPRECATED** (owner's call, 2026-07-31). It used to be a hard
|
||||
constraint and it is not one any more: a 0.8B — and a 1.7B — does not know enough to answer
|
||||
world questions, so she needs to read external sources. What replaces it:
|
||||
|
||||
- **No telemetry, no cloud model, no third-party account.** That part never changes. Nothing
|
||||
about Maven is reported to anyone, and inference stays on the box.
|
||||
- **The owner's data first, then the world.** Every source that reads the owner's facts, notes, calendar,
|
||||
tasks or house runs before anything outside, and the personal boundary sits between them.
|
||||
Reading beats recalling for a small model.
|
||||
- **In the world, live search leads and the ZIMs are the fallback** (owner's call,
|
||||
2026-08-02). A self-hosted SearXNG (`search` block) answers first; the Kiwix ZIMs on
|
||||
homesrv answer when the search is empty, unreachable, or the line is down.
|
||||
**Verified with the line down on 2026-08-05** (V-508,
|
||||
`docs/evals/2026-08-05-kiwix-offline-fallback.md`): a stopped SearXNG costs nothing,
|
||||
the ZIM answers in the same turn budget. A blackholed host cost 8 seconds the owner waited
|
||||
through. So the connect phase alone is capped at `dialTimeout` (1.5s), while a slow
|
||||
instance that did connect keeps the full 8. **A Russian question reads
|
||||
`wikipedia_ru_all_maxi_2026-02` verbatim** through `kiwix.book_ru`. The RU→EN rewriter
|
||||
is the workaround for an English book and is skipped there. Kiwix catalog names come
|
||||
from the filename, not the `<name>` field.
|
||||
`Response.Empty()` is the whole gate and there is no quality threshold in front of it:
|
||||
the three signals one could read were measured on 2026-08-05 and none of them separate a
|
||||
real question from an invented one. Token overlap would cost "столица Франции" its
|
||||
answer, because the answer is Париж and that word is not in the question. See
|
||||
`docs/evals/2026-08-05-search-quality-signals.md` (V-539). **Which query source claimed
|
||||
a turn is readable on `/chat`** as a badge beside the reply, carried on
|
||||
`ipc.ChatReply.Source` and noted by `noteQuerySource` in `cmd/mavend/querysource.go`. It
|
||||
rides the context, so `handleText` keeps the one string signature the mic, telegram and
|
||||
the web share.
|
||||
- **External search is allowed and off unless configured**, like the weather and telegram
|
||||
capabilities. The code default is still off. `deploy/mavend.json` now ships a `search`
|
||||
block (owner's call, 2026-08-02), so it is on for this box and deleting the block turns
|
||||
it off again.
|
||||
- **The owner's notes and facts are never search input.** Looking up why the sky is blue and
|
||||
sending the owner's stored personal notes to an upstream engine are different acts. Only the utterance goes
|
||||
out, never the persona block, history, or matched notes.
|
||||
|
||||
## Web UI conventions
|
||||
|
||||
Server-rendered pages share `cmd/mavweb/static/ui.css` (served at `/ui.css`) and the shell
|
||||
partial in `cmd/mavweb/shell.html`: a page opens with `{{template "shellTop" "<page-key>"}}`
|
||||
and closes with `{{template "shellBottom"}}`, and the key marks the active sidebar link.
|
||||
Every page is its own embedded `.html` file next to `main.go` — no page markup lives in Go,
|
||||
and the sidebar is data (`sidebarSections`, `pageIcon`) the template renders. No
|
||||
per-page `<style>` beyond true one-offs. Wrap every table in `<div class=scroll>` so wide
|
||||
data pans on a phone. Local preview + headless screenshot recipe is in `AGENTS.md`.
|
||||
|
||||
## Vikunja
|
||||
|
||||
This repo is project **Maven** (ID 2) in Vikunja. MCP: `http://localhost:9100/mcp` (or
|
||||
`http://192.168.1.104:9100/mcp` from workpc). Feature/bug/deploy tasks go there.
|
||||
|
||||
Vikunja is the durable task store. A task holds the goal, the constraints and the
|
||||
assumption ledger. Work without a task id is work nobody can resume, so a session that
|
||||
has no id asks for one before it starts.
|
||||
|
||||
The MCP tool schemas are deferred, so load the four you actually use in ONE call at the
|
||||
start of a session rather than one lookup per first use:
|
||||
|
||||
```text
|
||||
ToolSearch("select:mcp__vikunja__list_tasks,mcp__vikunja__get_task_details,mcp__vikunja__create_task,mcp__vikunja__update_task")
|
||||
```
|
||||
|
||||
`update_task` carrying a `description` resets `done` to false, so closing a task with a
|
||||
write-up takes two calls: the description, then `done: true`.
|
||||
- **No telemetry, no cloud model, no third-party account.** Inference stays on
|
||||
the box and nothing about Maven is reported to anyone.
|
||||
- **The owner's data first, then the world.** Every source reading his facts,
|
||||
notes, calendar, tasks or house runs first, and the personal boundary sits
|
||||
between them and anything outside.
|
||||
- **His notes and facts are never search input.** Only the utterance goes out,
|
||||
never the persona block, the history, or matched notes.
|
||||
- **External search is allowed and off unless configured.** Deleting the
|
||||
`search` block in `deploy/mavend.json` turns it off.
|
||||
- **`Response.Empty()` is the whole gate** on a world answer. There is no
|
||||
quality threshold in front of it and four candidate signals all failed.
|
||||
|
||||
## Session workflow
|
||||
|
||||
`~/.local/bin/task` owns the branch, the commit identity and the PR. One task, one
|
||||
session, one PR.
|
||||
`docs/workflow.md` carries the five stores, the doc tiers and the guards. One
|
||||
task, one session, one PR. `/pickup` opens a session and `/wrap` closes it. Wrap
|
||||
at roughly half context rather than letting the session compact.
|
||||
|
||||
```sh
|
||||
task start <vikunja-id> # branch off origin/master, write TASK.md, fetch review comments
|
||||
task pr # push, open or refresh the PR, label Vikunja, notify
|
||||
task comments # re-pull this branch's review comments into .task/
|
||||
ToolSearch("select:mcp__vikunja__list_tasks,mcp__vikunja__get_task_details,mcp__vikunja__create_task,mcp__vikunja__update_task")
|
||||
```
|
||||
|
||||
Around that, `/pickup` opens a session and `/wrap` closes it. Wrap at roughly half
|
||||
context rather than letting the session compact.
|
||||
|
||||
Five stores, and each one owns something the others must not hold:
|
||||
|
||||
| Store | Holds | Lifetime |
|
||||
|---|---|---|
|
||||
| Vikunja task | goal, constraints, assumption ledger, status | durable |
|
||||
| `CLAUDE.md`, `AGENTS.md` | what an agent must know before touching code | durable |
|
||||
| `docs/` | design, measurements, decisions | durable |
|
||||
| `TASK.md` | the brief for this branch, written by `task start`, immutable | one branch |
|
||||
| `HANDOFF.md` | only what the next agent needs to resume | one session |
|
||||
|
||||
`TASK.md` and `.task/` are excluded through `.git/info/exclude`. `HANDOFF.md` is
|
||||
gitignored and injected at session start. If a line in the handoff would still matter
|
||||
next week, it is in the wrong file.
|
||||
|
||||
Docs are tiered by path, so staleness is visible from the filename. Files directly under
|
||||
`docs/` are living and carry a `Last verified: <date> @ <sha>` line. Files under
|
||||
`docs/evals/` are dated measurements and are never edited after the day, so a newer
|
||||
number is a new file. Files under `docs/archive/` are dead and read by nobody by default.
|
||||
|
||||
## Git guards
|
||||
|
||||
Two hooks in `.githooks/`, tracked, wired with `core.hooksPath`. Fresh clone:
|
||||
|
||||
```sh
|
||||
git config core.hooksPath .githooks
|
||||
```
|
||||
|
||||
- `pre-commit` refuses master, and refuses more than 300 changed lines in non-markdown
|
||||
files. Markdown is exempt and may land as one batch.
|
||||
- `commit-msg` requires the subject to end with `(V-<id>)`. `V-` and not `#`, because
|
||||
Gitea autolinks `#123` to a Gitea issue, which is the wrong tracker.
|
||||
|
||||
Two more guards live outside the repo, in `~/.claude/hooks/`. `diff-budget.sh` blocks
|
||||
further edits past 600 changed lines on a `task/` branch. `prose_lint_hook.py` checks
|
||||
prose on every write. Both measure against `origin/master`, so a local master that is
|
||||
ahead of the remote makes the diff budget read high.
|
||||
|
||||
`--no-verify` exists. Using it means saying why in the commit body.
|
||||
- This repo is Vikunja project **Maven** (ID 2), MCP at
|
||||
`http://localhost:9100/mcp`, or `http://192.168.1.104:9100/mcp` from workpc.
|
||||
- **A session with no task id asks for one before it starts**, because work
|
||||
without one is work nobody can resume.
|
||||
- **Close a finished task with `done: true` and nothing else** (owner's call,
|
||||
2026-08-07). `update_task` carrying a `description` resets `done` to false.
|
||||
- **`pre-commit` refuses master** and more than 300 changed lines in
|
||||
non-markdown files. Markdown is exempt and may land as one batch.
|
||||
- **`commit-msg` requires the subject to end with `(V-<id>)`.** `V-` and not
|
||||
`#`, because Gitea autolinks `#123` to the wrong tracker.
|
||||
- **`diff-budget.sh` blocks edits past 600 changed lines** on a `task/` branch.
|
||||
- **`--no-verify` exists.** Using it means saying why in the commit body.
|
||||
|
||||
@@ -16,7 +16,7 @@ PIPER_BIN := $(shell pwd)/deps/piper/piper
|
||||
PIPER_MODEL := $(shell pwd)/models/tts/ru_RU-irina-medium.onnx
|
||||
PIPER_ESPEAK := $(shell pwd)/deps/piper/espeak-ng-data
|
||||
|
||||
.PHONY: simulate stt-fixtures test-stt-golden all build build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav clean test fmt-check vet run-stt run-tts run-web download-embedder deps-go deps-sentinel tidy eval-router eval-reach eval-recall eval-phrasing eval-models build-gpud
|
||||
.PHONY: t audit simulate stt-fixtures test-stt-golden all build build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav clean test fmt-check vet run-stt run-tts run-web download-embedder deps-go deps-sentinel tidy eval-router eval-reach eval-recall eval-phrasing eval-models build-gpud
|
||||
|
||||
all: build
|
||||
|
||||
@@ -128,6 +128,35 @@ test: fmt-check vet
|
||||
CGO_CFLAGS="$(CGO_CFLAGS)" CGO_LDFLAGS="$(CGO_LDFLAGS)" LD_LIBRARY_PATH="$(shell pwd)/deps/lib" \
|
||||
$(GO) test -race -coverprofile=coverage.out ./internal/... ./cmd/...
|
||||
|
||||
# t — run ONE package or ONE test with the toolchain env already wired. This is
|
||||
# the iteration target; `test` is the gate. Reach for it instead of pasting the
|
||||
# CGO_CFLAGS/CGO_LDFLAGS/LD_LIBRARY_PATH preamble by hand, which is how it was
|
||||
# done ~390 times across past sessions and is where the shell-quoting failures
|
||||
# came from -- the interactive shell here is zsh, and an unquoted `-run Test*`
|
||||
# or `--include=*.go` dies on "no matches found" before go ever starts.
|
||||
#
|
||||
# make t # whole tree (same scope as `test`)
|
||||
# make t PKG=./internal/router/
|
||||
# make t PKG=./cmd/mavend/ RUN=TestSimulator
|
||||
# make t PKG=./internal/router/eval/ RUN='TestONNX' V=1
|
||||
# make t PKG=./internal/store/ RACE=0 # drop -race when iterating hot
|
||||
#
|
||||
# -race is on by default so a green `make t` cannot turn red under `make test`.
|
||||
# -count=1 because a cached PASS from before your edit is worse than no answer.
|
||||
# MAVEN_ONNX_LIB is set for the same reason: the four TestONNX* measurements
|
||||
# self-skip when it is unset, so a targeted eval run would otherwise report the
|
||||
# deterministic hash ratchet and look like it scored the real embedder.
|
||||
PKG ?= ./internal/... ./cmd/...
|
||||
RUN ?=
|
||||
V ?=
|
||||
RACE ?= 1
|
||||
|
||||
t:
|
||||
CGO_CFLAGS="$(CGO_CFLAGS)" CGO_LDFLAGS="$(CGO_LDFLAGS)" LD_LIBRARY_PATH="$(shell pwd)/deps/lib" \
|
||||
MAVEN_ONNX_LIB="$(MAVEN_ONNX_LIB)" \
|
||||
$(GO) test $(if $(V),-v,) $(if $(filter-out 0,$(RACE)),-race,) -count=1 \
|
||||
$(if $(RUN),-run '$(RUN)',) $(PKG)
|
||||
|
||||
# eval-router — score the held-out RU routing fixture (internal/router/eval).
|
||||
# Verbose so the report tables land in the terminal. MAVEN_ONNX_LIB points the
|
||||
# prod-representative baseline at the vendored runtime; override it or set it
|
||||
@@ -189,6 +218,16 @@ eval-models:
|
||||
# scores the fixtures against ggml-small and self-skips when the model is
|
||||
# absent, and TestGoldenFixturesAreCanonical, which checks the committed audio
|
||||
# and the manifest with no model at all.
|
||||
# audit — the repo inventory: LOC per package, open TODOs, real stubs, living-doc
|
||||
# staleness, test shape, packages with no test. Read-only, prints, writes nothing.
|
||||
# Run it instead of rebuilding the same greps by hand; past sessions spent 93 of
|
||||
# them on this before their first edit. SECTION=loc|todo|stubs|docs|tests|gaps
|
||||
# narrows it.
|
||||
SECTION ?= all
|
||||
|
||||
audit:
|
||||
@SECTION="$(SECTION)" ./scripts/audit.sh
|
||||
|
||||
stt-fixtures:
|
||||
./scripts/gen-stt-fixtures.sh
|
||||
|
||||
|
||||
+130
-27
@@ -13,6 +13,7 @@ import (
|
||||
"github.com/kami/maven/internal/crawl"
|
||||
"github.com/kami/maven/internal/decision"
|
||||
"github.com/kami/maven/internal/ipc"
|
||||
"github.com/kami/maven/internal/kiwix"
|
||||
"github.com/kami/maven/internal/memory"
|
||||
"github.com/kami/maven/internal/morning"
|
||||
"github.com/kami/maven/internal/phraser"
|
||||
@@ -59,6 +60,35 @@ type querySource struct {
|
||||
// sources search text with no notion of a day. When one of them grows a
|
||||
// date parameter, flip its flag here.
|
||||
dateAware bool
|
||||
|
||||
// dest — the destination this source serves, when the cascade named one
|
||||
// (V-655). Several sources share a destination: the three recall passes and
|
||||
// the fact-by-key lookup are all SourceRecall, because which of them lands
|
||||
// the hit is an ordering detail no utterance can name. A source with no
|
||||
// dest is reachable only by walking the chain.
|
||||
dest router.Source
|
||||
|
||||
// guesses — this source decides whether the turn is its own by scoring the
|
||||
// utterance against frozen seeds, rather than by looking something up and
|
||||
// coming back empty.
|
||||
//
|
||||
// The distinction is the whole point of the field. A source that looks can
|
||||
// be wrong about relevance and still harmless, because the miss shows up as
|
||||
// no rows. A source that guesses answers whatever it claims: weather has no
|
||||
// local table to miss against, so "что такое TCP?" became "для какого
|
||||
// города?". So when the cascade names a destination, the guessers that were
|
||||
// not named do not get to try. The lookups still run, because a named
|
||||
// destination is evidence and not a promise.
|
||||
guesses bool
|
||||
|
||||
// boundary — dropping this source widens what leaves the box, so only a
|
||||
// literal pattern may do it (V-666, owner's call of 2026-08-09).
|
||||
//
|
||||
// Every other guesser costs an answer when it is wrongly taken off a turn.
|
||||
// This one costs the rule that a question about him never reaches an
|
||||
// upstream engine. A grammar read the words to name a destination. A model
|
||||
// and a softmax both inferred one, and neither may spend that.
|
||||
boundary bool
|
||||
}
|
||||
|
||||
// querySources is the ordered chain actionQuery walks; first source to claim
|
||||
@@ -67,85 +97,85 @@ type querySource struct {
|
||||
// gate was never the bug. Adding a source (Kiwix, RSS, crawler, email) is one
|
||||
// line here plus its method; where you put the line is the whole decision.
|
||||
var querySources = []querySource{
|
||||
{name: "fact-by-key", answer: (*reactiveHandler).queryFactByKey},
|
||||
{name: "fact-by-key", answer: (*reactiveHandler).queryFactByKey, dest: router.SourceRecall},
|
||||
// Before "calendar" on purpose: both match "…на сегодня", and the plan is
|
||||
// the more specific ask (its matcher requires a plan word), so the calendar
|
||||
// listing would otherwise swallow it.
|
||||
{name: "day-plan", answer: (*reactiveHandler).queryDayPlan},
|
||||
{name: "day-plan", answer: (*reactiveHandler).queryDayPlan, dest: router.SourceCalendar},
|
||||
// Also before "calendar": "что я обычно делаю по средам?" names a weekday,
|
||||
// and the habit question is the more specific one. Its matcher requires a
|
||||
// habit marker ("обычно", "каждый", …), so a question about this coming
|
||||
// Wednesday still reaches the calendar.
|
||||
{name: "habits", answer: (*reactiveHandler).queryHabits},
|
||||
{name: "habits", answer: (*reactiveHandler).queryHabits, dest: router.SourceCalendar},
|
||||
// Before "calendar" and before the recall sources: "что мне нужно
|
||||
// сделать?" is a question about the task list, and the notes pass would
|
||||
// otherwise answer it with whatever note happens to be nearest. Its
|
||||
// matcher requires a task noun or an explicit "что … сделать", so a
|
||||
// date-bearing question still reaches the calendar.
|
||||
{name: "tasks", answer: (*reactiveHandler).queryTasks},
|
||||
{name: "tasks", answer: (*reactiveHandler).queryTasks, dest: router.SourceTasks},
|
||||
// Next to "tasks" and for the same reason: "что требует внимания?" is a
|
||||
// question about the operational state Praxis holds, and it used to fall
|
||||
// through every source to the web search (Vikunja #475). Its matcher needs
|
||||
// an attention marker, and it falls through when Praxis is not configured.
|
||||
{name: "attention", answer: (*reactiveHandler).queryAttention},
|
||||
{name: "attention", answer: (*reactiveHandler).queryAttention, dest: router.SourceAttention, guesses: true},
|
||||
// Next to "tasks" and for the same reason: "что мне купить?" is a question
|
||||
// about the shopping list, and the recall pass would otherwise answer it
|
||||
// from an old note about the shop. Its matcher needs an explicit list
|
||||
// marker, so "надо бы съездить в магазин" is untouched.
|
||||
{name: "list", answer: (*reactiveHandler).queryList},
|
||||
{name: "list", answer: (*reactiveHandler).queryList, dest: router.SourceList, guesses: true},
|
||||
// Before the recall sources too: "сколько я потратил?" is a question about
|
||||
// the money facts the poller wrote, and the notes pass would otherwise
|
||||
// answer it from whatever he once said about spending. Its matcher needs a
|
||||
// money noun plus an actual ask, so "я потратил весь день" is untouched.
|
||||
{name: "money", answer: (*reactiveHandler).queryMoney},
|
||||
{name: "money", answer: (*reactiveHandler).queryMoney, dest: router.SourceMoney},
|
||||
// Also above the recall sources: "что я тебе говорил?" is a question about
|
||||
// the facts he tapped in, and the notes pass would answer it with whatever
|
||||
// note is nearest (Vikunja #456). Its matcher needs both halves of a
|
||||
// history phrase and bails out when he names a topic, so "что я говорил
|
||||
// про сервер" is still recall.
|
||||
{name: "history", answer: (*reactiveHandler).queryHistory},
|
||||
{name: "history", answer: (*reactiveHandler).queryHistory, dest: router.SourceRecall},
|
||||
// Before the recall sources and before general knowledge: "что нового?" is
|
||||
// a question about the feeds she reads, and general knowledge would answer
|
||||
// it by inventing news. Its matcher needs a feed noun plus an ask, so
|
||||
// "у меня новая лента в инстаграме" is untouched.
|
||||
{name: "feeds", answer: (*reactiveHandler).queryFeeds},
|
||||
{name: "feeds", answer: (*reactiveHandler).queryFeeds, dest: router.SourceFeeds, guesses: true},
|
||||
// Before "calendar" and before the recall sources: "что включено дома?" is
|
||||
// a question about the house, and the notes pass would otherwise answer it
|
||||
// from whatever he once said about the lights. Its matcher needs a house
|
||||
// marker plus an ask plus a device word, and it bails out on weather
|
||||
// wording, so "какая температура на улице?" still reaches the weather
|
||||
// source.
|
||||
{name: "home", answer: (*reactiveHandler).queryHome},
|
||||
{name: "home", answer: (*reactiveHandler).queryHome, dest: router.SourceHome, guesses: true},
|
||||
// Next to "home" and for the same reason: "какие устройства в сети?" is a
|
||||
// question about the LAN, and the recall pass would otherwise answer it
|
||||
// from an old note about the router. Its matcher needs a network word plus
|
||||
// an ask plus a device noun, so "интернет не работает" is untouched.
|
||||
{name: "network", answer: (*reactiveHandler).queryNetwork},
|
||||
{name: "calendar", answer: (*reactiveHandler).queryCalendar, dateAware: true},
|
||||
{name: "weather", answer: (*reactiveHandler).queryWeather},
|
||||
{name: "network", answer: (*reactiveHandler).queryNetwork, dest: router.SourceNetwork, guesses: true},
|
||||
{name: "calendar", answer: (*reactiveHandler).queryCalendar, dateAware: true, dest: router.SourceCalendar},
|
||||
{name: "weather", answer: (*reactiveHandler).queryWeather, dest: router.SourceWeather, guesses: true},
|
||||
// A question about her, above the three sources that search his own data
|
||||
// (Vikunja #555). It has no answer anywhere else: below the boundary
|
||||
// SearXNG answers about somebody else's assistant, and above it his notes
|
||||
// answer by proximity — "кто ты" came back from a note of his, measured on
|
||||
// the box, because the recall index has no idea the subject is her.
|
||||
{name: "self", answer: (*reactiveHandler).querySelf},
|
||||
{name: "embed", answer: (*reactiveHandler).queryEmbed},
|
||||
{name: "memory", answer: (*reactiveHandler).queryMemory},
|
||||
{name: "notes", answer: (*reactiveHandler).queryNotes},
|
||||
{name: "self", answer: (*reactiveHandler).querySelf, dest: router.SourceSelf, guesses: true},
|
||||
{name: "embed", answer: (*reactiveHandler).queryEmbed, dest: router.SourceRecall},
|
||||
{name: "memory", answer: (*reactiveHandler).queryMemory, dest: router.SourceRecall},
|
||||
{name: "notes", answer: (*reactiveHandler).queryNotes, dest: router.SourceRecall},
|
||||
// THE BOUNDARY. Everything above answers from his own data; everything
|
||||
// below answers from the world's. A question about him that got this far
|
||||
// has no answer in his data, and no outside source can supply one, so this
|
||||
// stops the walk rather than let the encyclopedia and the model guess.
|
||||
{name: "personal", answer: (*reactiveHandler).queryPersonal},
|
||||
{name: "personal", answer: (*reactiveHandler).queryPersonal, dest: router.SourceRecall, guesses: true, boundary: true},
|
||||
// The world, read live. Owner's ruling of 2026-08-02: a metasearch hit beats
|
||||
// a frozen ZIM, so SearXNG asks before Kiwix does. Nothing of his is at
|
||||
// stake by this point — the boundary above already stopped every question
|
||||
// about him, and only the query string leaves the box.
|
||||
{name: "search", answer: (*reactiveHandler).querySearch},
|
||||
{name: "search", answer: (*reactiveHandler).querySearch, dest: router.SourceWorld},
|
||||
// The offline encyclopedia, now the fallback for when the line is down or
|
||||
// the search comes back empty. It reads the way it always did; what changed
|
||||
// is that it no longer gets first refusal on a world question.
|
||||
{name: "kiwix", answer: (*reactiveHandler).queryKiwix},
|
||||
{name: "kiwix", answer: (*reactiveHandler).queryKiwix, dest: router.SourceWorld},
|
||||
// LAST before the model answers from memory, and that position is the whole
|
||||
// design (Vikunja #259): everything of his, then the search, then the ZIMs,
|
||||
// and only then a page he named. The model does NOT come first: it
|
||||
@@ -153,8 +183,48 @@ var querySources = []querySource{
|
||||
// a 1.7B guessing at a page it cannot read is how contents get invented.
|
||||
// This source only claims a turn where he named a URL, so it never competes
|
||||
// with a local answer.
|
||||
{name: "web", answer: (*reactiveHandler).queryWeb},
|
||||
{name: "general-knowledge", answer: (*reactiveHandler).queryGeneral},
|
||||
{name: "web", answer: (*reactiveHandler).queryWeb, dest: router.SourceWorld},
|
||||
{name: "general-knowledge", answer: (*reactiveHandler).queryGeneral, dest: router.SourceWorld},
|
||||
}
|
||||
|
||||
// queryWalk narrows the chain for one turn against the destination the cascade
|
||||
// named, and says which sources were left out (V-655).
|
||||
//
|
||||
// It takes sources OUT and never moves one, which is the whole safety argument.
|
||||
// The table's order is load-bearing and every comment on it argues a reason
|
||||
// between two sources; none of those reasons is about this. Above all, the
|
||||
// order carries "his data first, then the world", and a destination named by a
|
||||
// model must not be able to reverse that. Naming SourceWorld does not send the
|
||||
// turn outside — it stops the guessers from claiming it on the way.
|
||||
//
|
||||
// What comes out is exactly the sources that guess. Those decide whether a turn
|
||||
// is theirs by scoring it against frozen seeds, and then answer whatever they
|
||||
// claimed, because they have no lookup that can come back empty. That is the
|
||||
// whole of the 2026-08-07 defect: weather claiming "что такое TCP?", the feed
|
||||
// claiming "какой у меня любимый язык?", the personal boundary claiming "кто
|
||||
// такой Линус Торвальдс?". The sources that look are all still asked, so a
|
||||
// wrong destination costs nothing but the guess it prevented.
|
||||
//
|
||||
// No destination named ⇒ the table exactly as written, which is what shipped
|
||||
// before the field existed. That is the floor. The classifier arm names
|
||||
// nothing, so a box whose model is down routes queries the way it always did.
|
||||
// The personal boundary is the one exception, and anchored is what buys it
|
||||
// (V-666). A grammar matched a literal pattern to name the destination. The
|
||||
// routing heads and the resident model inferred one, and an inferred SourceWorld
|
||||
// takes the boundary off a question about him. That widens what is asked
|
||||
// upstream rather than costing a local answer, so those two keep it.
|
||||
func queryWalk(dest router.Source, anchored bool) (walk, skipped []querySource) {
|
||||
if dest == router.SourceUnknown {
|
||||
return querySources, nil
|
||||
}
|
||||
for _, s := range querySources {
|
||||
if s.guesses && s.dest != dest && (anchored || !s.boundary) {
|
||||
skipped = append(skipped, s)
|
||||
continue
|
||||
}
|
||||
walk = append(walk, s)
|
||||
}
|
||||
return walk, skipped
|
||||
}
|
||||
|
||||
func (h *reactiveHandler) actionQuery(ctx context.Context, dec router.Decision) string {
|
||||
@@ -164,7 +234,14 @@ func (h *reactiveHandler) actionQuery(ctx context.Context, dec router.Decision)
|
||||
// (V-564). Finish names everyone below the winner.
|
||||
decision.Expect(ctx, decision.StageQuery, querySourceNames())
|
||||
rec := decision.From(ctx)
|
||||
for _, src := range querySources {
|
||||
walk, skipped := queryWalk(dec.Source, dec.SourceAnchored)
|
||||
for _, src := range skipped {
|
||||
rec.Note(decision.Claim{
|
||||
Stage: decision.StageQuery, Claimant: src.name, Outcome: decision.NeverAsked,
|
||||
Reason: "it decides by similarity and the cascade named " + string(dec.Source),
|
||||
})
|
||||
}
|
||||
for _, src := range walk {
|
||||
if dec.Continued && !src.dateAware {
|
||||
rec.Note(decision.Claim{
|
||||
Stage: decision.StageQuery, Claimant: src.name, Outcome: decision.NeverAsked,
|
||||
@@ -850,6 +927,28 @@ func (h *reactiveHandler) queryKiwix(ctx context.Context, t *queryTurn) (string,
|
||||
}
|
||||
}
|
||||
|
||||
// The topic, not the sentence (V-668). Kiwix ranks by keyword overlap, so
|
||||
// the question words outrank the one word that names the article: measured
|
||||
// on 2026-08-09, "что такое TCP" returns "Перехват TCP-соединения" and
|
||||
// "TCP" returns TCP. Only the verbatim path needs this. The rewriter
|
||||
// already reduces a question to English keywords, and reducing twice would
|
||||
// take the topic off the input it reads.
|
||||
if verbatim {
|
||||
if topic := kiwix.Topic(pattern); topic != "" {
|
||||
// The article named exactly, before any ranking runs. A ZIM is
|
||||
// addressable by title and a wrong title is a 404, so this either
|
||||
// answers or costs one request that says nothing.
|
||||
for _, cand := range kiwix.TitleCandidates(topic) {
|
||||
page, err := h.kiwix.client.Article(ctxK, kiwix.TitlePath(book, cand), h.kiwix.runes)
|
||||
if err == nil && page.Text != "" {
|
||||
log.Printf("voice: kiwix: %q in %q → title hit %q", topic, book, page.Title)
|
||||
return h.kiwixReply(ctx, t, page.Title, page.Text)
|
||||
}
|
||||
}
|
||||
pattern = topic
|
||||
}
|
||||
}
|
||||
|
||||
hits, err := h.kiwix.client.Search(ctxK, pattern, book, h.kiwix.max)
|
||||
if err != nil {
|
||||
log.Printf("voice: kiwix: search %q: %v", pattern, err)
|
||||
@@ -882,14 +981,18 @@ func (h *reactiveHandler) queryKiwix(ctx context.Context, t *queryTurn) (string,
|
||||
}
|
||||
page = crawl.Page{Title: top.Title, Text: top.Snippet}
|
||||
}
|
||||
// Handed over the same way a note or a page is: context for the question he
|
||||
// asked, not something to recite.
|
||||
snippet := top.Title + "\n" + crawl.TrimRunes(page.Text, h.kiwix.runes)
|
||||
return h.kiwixReply(ctx, t, top.Title, page.Text)
|
||||
}
|
||||
|
||||
// kiwixReply hands one article over the same way a note or a page is handed
|
||||
// over: context for the question he asked, not something to recite.
|
||||
func (h *reactiveHandler) kiwixReply(ctx context.Context, t *queryTurn, title, text string) (string, bool) {
|
||||
snippet := title + "\n" + crawl.TrimRunes(text, h.kiwix.runes)
|
||||
reply := h.phraseSource(ctx, "kiwix", t.dec.Utterance, []string{snippet})
|
||||
if reply == "" {
|
||||
// No phraser, or it failed. Read back the best hit rather than pretend
|
||||
// the search did not happen.
|
||||
return readBack(top.Title + " — " + page.Text), true
|
||||
return readBack(title + " — " + text), true
|
||||
}
|
||||
return reply, true
|
||||
}
|
||||
|
||||
@@ -19,7 +19,7 @@ func TestChatAnswersWithNoLlamaServer(t *testing.T) {
|
||||
dead := llm.New("http://127.0.0.1:1", 500*time.Millisecond)
|
||||
emb := router.NewHashEmbedder(1024)
|
||||
h.recall.embedder = emb
|
||||
h.router = buildRouter(emb, h.matcher, 0.55, pickLLMRouter(true, dead))
|
||||
h.router = buildRouter(emb, h.matcher, 0.55, pickLLMRouter(true, dead), nil)
|
||||
h.replier = newLLMReplier(dead, nil)
|
||||
|
||||
ctx := withDialogueID(context.Background(), dialogueIDFor(sourceText, "web"))
|
||||
|
||||
+52
-22
@@ -6,7 +6,6 @@ import (
|
||||
"math/rand"
|
||||
"strings"
|
||||
"time"
|
||||
"unicode"
|
||||
"unicode/utf8"
|
||||
|
||||
"github.com/kami/maven/internal/dialogue"
|
||||
@@ -162,12 +161,13 @@ func withNotice(notice, reply string) string {
|
||||
// напоминание?" — answer first, then the open question. A question in front of
|
||||
// its own answer would read as ignoring what he asked.
|
||||
//
|
||||
// A statement's full stop is folded into a comma, so the two acts read as one
|
||||
// sentence — that is the owner's own punctuation, "в Риме сейчас ..., на какое
|
||||
// время поставить напоминание?". An answer that is ITSELF a question keeps its
|
||||
// mark and the resume starts a new sentence: she sometimes answers a side query
|
||||
// by asking him to say it again, and "переформулировать?, на какое время" folds
|
||||
// two questions into one unreadable line.
|
||||
// Two sentences, not one (V-654). This used to fold the answer's full stop into
|
||||
// a comma, on the strength of the owner having written it that way once. Spliced
|
||||
// onto a real answer it reads as one run-on thought — "вот что я нашла: вайфай
|
||||
// пароль лежит в ящике стола, на какое время поставить напоминание?" — and the
|
||||
// question disappears into the tail of a sentence about something else. A reply
|
||||
// with no terminator of its own is given one, so the join never depends on how
|
||||
// the phraser chose to end.
|
||||
//
|
||||
// A resume with no answer in front of it is just the question.
|
||||
func withResumed(reply, resumed string) string {
|
||||
@@ -178,23 +178,17 @@ func withResumed(reply, resumed string) string {
|
||||
if reply == "" {
|
||||
return resumed
|
||||
}
|
||||
if strings.HasSuffix(reply, "?") {
|
||||
return reply + " " + resumed
|
||||
if !endsSentence(reply) {
|
||||
reply += "."
|
||||
}
|
||||
if trimmed := strings.TrimRight(reply, ".!"); trimmed != "" {
|
||||
reply = trimmed
|
||||
}
|
||||
return reply + ", " + lowerFirst(resumed)
|
||||
return reply + " " + resumed
|
||||
}
|
||||
|
||||
// lowerFirst lowercases the opening rune, so a deck line written as a standalone
|
||||
// sentence reads as the second half of one. Only the first rune: "На какое
|
||||
// время" must become "на какое время" and nothing else in it may move.
|
||||
func lowerFirst(s string) string {
|
||||
for i, r := range s {
|
||||
return string(unicode.ToLower(r)) + s[i+utf8.RuneLen(r):]
|
||||
}
|
||||
return s
|
||||
// endsSentence reports whether s already closes itself. The ellipsis counts: a
|
||||
// trailing "…" is a deliberate end, and a full stop after it reads as a typo.
|
||||
func endsSentence(s string) bool {
|
||||
r, _ := utf8.DecodeLastRuneInString(s)
|
||||
return strings.ContainsRune(".!?…", r)
|
||||
}
|
||||
|
||||
// missingFor returns the slots a decision still needs, most important first.
|
||||
@@ -381,6 +375,14 @@ func (h *reactiveHandler) resolveClarifyAnswer(ctx context.Context, text string)
|
||||
return "", false
|
||||
}
|
||||
|
||||
// He is answering, so the run of step-asides is over (V-654). Reset here
|
||||
// rather than where a gap is FILLED: "позвонить маме" against a question
|
||||
// about the time gives her nothing she asked for and still means he is in
|
||||
// the exchange, and the retry it costs is bound enough on its own. The
|
||||
// counter is for the case the bounds miss — he asked for other things and
|
||||
// never came back.
|
||||
q.Suspends = 0
|
||||
|
||||
merged := q.Answer(text, toDialogueSlots(answer))
|
||||
// Fold a newly answered subject into the raw utterance. Downstream actions
|
||||
// phrase from Utterance, not from the text slot — actionReminder stores it
|
||||
@@ -467,6 +469,14 @@ func (h *reactiveHandler) noteDropped(ctx context.Context) {
|
||||
//
|
||||
// A slot with no resumed wording (clarifyResumedFor says so) resumes nothing and
|
||||
// says nothing. She must not claim to be holding a question she cannot re-ask.
|
||||
//
|
||||
// Suspension is bounded, since V-654. Neither of the two things above is a
|
||||
// limit: no attempt is spent, and restarting the clock means the TTL cannot
|
||||
// arrive while he keeps talking. So the count is the only thing that ends it,
|
||||
// and past MaxSuspends she lets the request go and says so with the same line
|
||||
// every other drop uses. The rule is unchanged — a question ends by being
|
||||
// answered or by being let go out loud — this only recognises three unrelated
|
||||
// requests in a row as the second of those.
|
||||
func (h *reactiveHandler) noteSuspended(ctx context.Context, q *dialogue.PendingQuestion) {
|
||||
rt := turnRouteFrom(ctx)
|
||||
if rt == nil || len(q.Missing) == 0 {
|
||||
@@ -476,11 +486,22 @@ func (h *reactiveHandler) noteSuspended(ctx context.Context, q *dialogue.Pending
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
if !q.CanResume() {
|
||||
h.clarifyStore.Delete(dialogueIDOf(ctx))
|
||||
h.noteDropped(ctx)
|
||||
log.Printf("voice: clarify — letting the question about %s go: %d asides in a row, %d rides in all", q.Missing[0], q.Suspends, q.Rides)
|
||||
return
|
||||
}
|
||||
q.Suspends++
|
||||
// Rides is the same event counted without the reset (V-663). Incremented
|
||||
// beside Suspends and never anywhere else, so the two cannot disagree about
|
||||
// what happened, only about how much of it they remember.
|
||||
q.Rides++
|
||||
q.Asked = h.now()
|
||||
h.clarifyStore.Put(dialogueIDOf(ctx), q)
|
||||
rt.resume = question
|
||||
rt.suspended = true
|
||||
log.Printf("voice: clarify — is its own request; suspending the question about %s and resuming it in the same reply", q.Missing[0])
|
||||
log.Printf("voice: clarify — is its own request; suspending the question about %s and resuming it in the same reply (suspend %d of %d, ride %d of %d)", q.Missing[0], q.Suspends, dialogue.MaxSuspends, q.Rides, dialogue.MaxRides)
|
||||
}
|
||||
|
||||
// foldAnswerIntoUtterance appends an answered subject to the original words,
|
||||
@@ -520,6 +541,14 @@ func (h *reactiveHandler) askRemainingGap(ctx context.Context, q *dialogue.Pendi
|
||||
if !ok || !q.CanAsk() {
|
||||
return "", false
|
||||
}
|
||||
// Suspends is not carried, and by this point it is already zero: the answer
|
||||
// path resets it (V-654). Left off the literal so the zero is stated where
|
||||
// the struct is built, rather than inherited from a field nobody names.
|
||||
//
|
||||
// Rides IS carried, and that is the whole point of it (V-663). This is the
|
||||
// same request under a second question, not a new one, so the turns it has
|
||||
// already ridden still count against it. Dropping the field here is exactly
|
||||
// the re-basing that let one question ride twenty-six replies.
|
||||
h.clarifyStore.Put(dialogueIDOf(ctx), &dialogue.PendingQuestion{
|
||||
Intent: q.Intent,
|
||||
Slots: merged,
|
||||
@@ -530,6 +559,7 @@ func (h *reactiveHandler) askRemainingGap(ctx context.Context, q *dialogue.Pendi
|
||||
TTL: clarifyTTL,
|
||||
Attempts: q.Attempts + 1,
|
||||
MaxAttempts: q.MaxAttempts,
|
||||
Rides: q.Rides,
|
||||
})
|
||||
log.Printf("voice: clarify — one gap filled, still missing %s for intent=%s, asking again (attempt %d)", remaining[0], intent, q.Attempts+1)
|
||||
return question, true
|
||||
|
||||
@@ -317,7 +317,7 @@ func TestClarifyExpiryIsAnnouncedAndWordsStillRoute(t *testing.T) {
|
||||
h, _, now := newClarifyHandler(t)
|
||||
emb := router.NewHashEmbedder(1024)
|
||||
h.recall.embedder = emb
|
||||
h.router = buildRouter(emb, h.matcher, 0.55, nil)
|
||||
h.router = buildRouter(emb, h.matcher, 0.55, nil, nil)
|
||||
|
||||
if _, asked := h.askClarify(ctx, clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
|
||||
t.Fatal("expected a question")
|
||||
@@ -671,7 +671,7 @@ func TestUnresolvedActSaysItDoesNotKnowTheCommand(t *testing.T) {
|
||||
func newRoutingClarifyHandler(t *testing.T) (*reactiveHandler, *store.Store) {
|
||||
t.Helper()
|
||||
h, st, _ := newClarifyHandler(t)
|
||||
h.router = buildRouter(router.NewHashEmbedder(1024), h.matcher, 0.55, nil)
|
||||
h.router = buildRouter(router.NewHashEmbedder(1024), h.matcher, 0.55, nil, nil)
|
||||
h.recall = recallWiring{embedder: router.NewHashEmbedder(1024), memStore: memory.NewInMemoryStore()}
|
||||
return h, st
|
||||
}
|
||||
@@ -740,3 +740,53 @@ func TestACompleteTurnStillDoesNotAsk(t *testing.T) {
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestTheResumedQuestionIsItsOwnSentence — V-654. The re-ask used to be spliced
|
||||
// onto the answer with a comma, so a real answer and an unrelated open question
|
||||
// read as one run-on thought and the question vanished into its tail.
|
||||
func TestTheResumedQuestionIsItsOwnSentence(t *testing.T) {
|
||||
const resumed = "На какое время поставить напоминание?"
|
||||
cases := []struct {
|
||||
name string
|
||||
reply string
|
||||
want string
|
||||
}{
|
||||
{
|
||||
// The measured line, shortened. Two sentences, and the question keeps
|
||||
// its capital.
|
||||
name: "a statement keeps its full stop",
|
||||
reply: "Вайфай пароль лежит в ящике стола.",
|
||||
want: "Вайфай пароль лежит в ящике стола. " + resumed,
|
||||
},
|
||||
{
|
||||
name: "a reply with no terminator is given one",
|
||||
reply: "Вайфай пароль лежит в ящике стола",
|
||||
want: "Вайфай пароль лежит в ящике стола. " + resumed,
|
||||
},
|
||||
{
|
||||
// She sometimes answers a side query by asking him to say it again.
|
||||
// Two questions, and neither may swallow the other.
|
||||
name: "a question keeps its mark",
|
||||
reply: "Можешь переформулировать?",
|
||||
want: "Можешь переформулировать? " + resumed,
|
||||
},
|
||||
{
|
||||
name: "an ellipsis is already an ending",
|
||||
reply: "Не уверена…",
|
||||
want: "Не уверена… " + resumed,
|
||||
},
|
||||
{
|
||||
name: "a resume with no answer in front of it is just the question",
|
||||
reply: "",
|
||||
want: resumed,
|
||||
},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
if got := withResumed(tc.reply, resumed); got != tc.want {
|
||||
t.Errorf("%s: withResumed(%q) = %q, want %q", tc.name, tc.reply, got, tc.want)
|
||||
}
|
||||
}
|
||||
if got := withResumed("Готово.", ""); got != "Готово." {
|
||||
t.Errorf("nothing to resume must leave the reply alone, got %q", got)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -25,7 +25,7 @@ func traceHandler(t *testing.T, ring *decision.Ring) *reactiveHandler {
|
||||
return &reactiveHandler{
|
||||
api: api,
|
||||
recall: recallWiring{embedder: emb, memStore: memory.NewInMemoryStore()},
|
||||
router: buildRouter(emb, tool.NewMatcher(api), 0.55, nil),
|
||||
router: buildRouter(emb, tool.NewMatcher(api), 0.55, nil, nil),
|
||||
replier: voice.NewStubReplier(),
|
||||
now: func() time.Time { return now },
|
||||
dataStore: st,
|
||||
|
||||
@@ -171,7 +171,7 @@ func newDialogueHandler(t *testing.T) (*reactiveHandler, *store.Store, *time.Tim
|
||||
// and never a coincidence (V-577, V-579). checkEnd refuses any reminder
|
||||
// landing on it, and at 09:00 the row that answers "на 9" would trip that.
|
||||
*now = time.Date(2026, 7, 31, 9, 17, 0, 0, time.UTC)
|
||||
h.router = buildRouter(router.NewHashEmbedder(1024), h.matcher, 0.55, nil)
|
||||
h.router = buildRouter(router.NewHashEmbedder(1024), h.matcher, 0.55, nil, nil)
|
||||
h.recall = recallWiring{embedder: router.NewHashEmbedder(1024), memStore: memory.NewInMemoryStore()}
|
||||
return h, st, now
|
||||
}
|
||||
|
||||
@@ -24,7 +24,7 @@ func TestApplyAction_FactCapture_QueuesEntityResolution(t *testing.T) {
|
||||
|
||||
emb := router.NewHashEmbedder(1024)
|
||||
matcher := tool.NewMatcher(api)
|
||||
rtr := buildRouter(emb, matcher, 0.55, nil)
|
||||
rtr := buildRouter(emb, matcher, 0.55, nil, nil)
|
||||
|
||||
h := &reactiveHandler{
|
||||
api: api,
|
||||
|
||||
@@ -20,7 +20,7 @@ func newFactGateHandler(t *testing.T, now time.Time) (*reactiveHandler, ipc.Core
|
||||
h := &reactiveHandler{
|
||||
api: api,
|
||||
recall: recallWiring{embedder: emb, memStore: memory.NewInMemoryStore()},
|
||||
router: buildRouter(emb, tool.NewMatcher(api), 0.55, nil),
|
||||
router: buildRouter(emb, tool.NewMatcher(api), 0.55, nil, nil),
|
||||
replier: voice.NewStubReplier(),
|
||||
now: func() time.Time { return now },
|
||||
dataStore: st,
|
||||
|
||||
@@ -45,7 +45,7 @@ func newNoteHandler(t *testing.T) (*reactiveHandler, *store.Store) {
|
||||
h := &reactiveHandler{
|
||||
api: api,
|
||||
recall: recallWiring{embedder: emb, memStore: memory.NewInMemoryStore()},
|
||||
router: buildRouter(emb, tool.NewMatcher(api), 0.55, nil),
|
||||
router: buildRouter(emb, tool.NewMatcher(api), 0.55, nil, nil),
|
||||
replier: voice.NewStubReplier(),
|
||||
now: func() time.Time { return now },
|
||||
dataStore: st,
|
||||
|
||||
@@ -0,0 +1,137 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"testing"
|
||||
|
||||
"github.com/kami/maven/internal/router"
|
||||
)
|
||||
|
||||
// The floor, and it is the reason a destination is safe to add at all: a box
|
||||
// whose model is down names nothing, and naming nothing has to walk the chain
|
||||
// the way it walked before the field existed.
|
||||
func TestNoDestinationWalksTheWholeChain(t *testing.T) {
|
||||
walk, skipped := queryWalk(router.SourceUnknown, false)
|
||||
if len(skipped) != 0 {
|
||||
t.Errorf("skipped %d sources with no destination named, want none", len(skipped))
|
||||
}
|
||||
if len(walk) != len(querySources) {
|
||||
t.Fatalf("walk has %d sources, want the whole table of %d", len(walk), len(querySources))
|
||||
}
|
||||
for i := range walk {
|
||||
if walk[i].name != querySources[i].name {
|
||||
t.Fatalf("position %d is %q, want %q", i, walk[i].name, querySources[i].name)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// The 2026-08-07 defects, one per line. Each is a source that decides by seed
|
||||
// similarity claiming a turn that was never its own, and then answering it
|
||||
// because it has no lookup that could come back empty.
|
||||
func TestANamedDestinationSilencesTheOtherGuessers(t *testing.T) {
|
||||
cases := []struct {
|
||||
dest router.Source
|
||||
utterance string
|
||||
silenced string
|
||||
anchored bool // a stage 0 grammar named the destination
|
||||
}{
|
||||
{router.SourceWorld, "что такое TCP?", "weather", true},
|
||||
{router.SourceWorld, "сколько будет 17 на 23?", "weather", true},
|
||||
{router.SourceWorld, "кто такой Линус Торвальдс?", "personal", true},
|
||||
{router.SourceRecall, "какой у меня любимый язык?", "feeds", false},
|
||||
{router.SourceCalendar, "что в календаре на завтра?", "weather", true},
|
||||
}
|
||||
for _, c := range cases {
|
||||
walk, skipped := queryWalk(c.dest, c.anchored)
|
||||
if inWalk(walk, c.silenced) {
|
||||
t.Errorf("%q named %q: %q is still asked", c.utterance, c.dest, c.silenced)
|
||||
}
|
||||
if !inWalk(skipped, c.silenced) {
|
||||
t.Errorf("%q named %q: %q is missing from the record of who was skipped",
|
||||
c.utterance, c.dest, c.silenced)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Naming the world must not send the turn outside. His notes, his facts and the
|
||||
// boundary in front of them are the invariant CLAUDE.md states as "the owner's
|
||||
// data first, then the world", and a destination a model wrote must not be able
|
||||
// to reverse it.
|
||||
func TestNamingTheWorldStillReadsHisDataFirst(t *testing.T) {
|
||||
walk, _ := queryWalk(router.SourceWorld, true)
|
||||
for _, look := range []string{"fact-by-key", "embed", "memory", "notes"} {
|
||||
if !inWalk(walk, look) {
|
||||
t.Errorf("%q was dropped; only the sources that guess may be dropped", look)
|
||||
}
|
||||
}
|
||||
if posOf(walk, "notes") > posOf(walk, "search") {
|
||||
t.Error("search is asked before his notes are")
|
||||
}
|
||||
if posOf(walk, "search") < 0 {
|
||||
t.Fatal("search is not in the walk at all")
|
||||
}
|
||||
}
|
||||
|
||||
// The boundary belongs to his data, so naming recall keeps it. That is what
|
||||
// makes "какой у меня любимый язык?" answer "не нашла у тебя такой записи"
|
||||
// rather than reaching SearXNG once nothing local had it.
|
||||
func TestNamingRecallKeepsTheBoundary(t *testing.T) {
|
||||
walk, _ := queryWalk(router.SourceRecall, true)
|
||||
if !inWalk(walk, "personal") {
|
||||
t.Fatal("the personal boundary was skipped on a turn named for his own data")
|
||||
}
|
||||
if posOf(walk, "personal") > posOf(walk, "search") {
|
||||
t.Error("the boundary no longer sits in front of the world")
|
||||
}
|
||||
}
|
||||
|
||||
// The owner's call of 2026-08-09 (V-666): only a stage 0 grammar may take the
|
||||
// personal boundary off a turn. The routing heads and the resident model both
|
||||
// name a destination by inference, and an inferred SourceWorld would send a
|
||||
// question about him upstream. Every other guesser still goes.
|
||||
func TestOnlyAGrammarMayDropTheBoundary(t *testing.T) {
|
||||
walk, skipped := queryWalk(router.SourceWorld, false)
|
||||
if !inWalk(walk, "personal") {
|
||||
t.Error("an inferred destination took the boundary off the turn")
|
||||
}
|
||||
if !inWalk(skipped, "weather") {
|
||||
t.Error("weather is still asked; the rule covers the boundary alone")
|
||||
}
|
||||
if posOf(walk, "personal") > posOf(walk, "search") {
|
||||
t.Error("the boundary no longer sits in front of the world")
|
||||
}
|
||||
if anchored, _ := queryWalk(router.SourceWorld, true); inWalk(anchored, "personal") {
|
||||
t.Error(`a grammar named the world and the boundary stayed: ` +
|
||||
`"кто такой Линус Торвальдс?" is answered "не нашла у тебя такой записи" again`)
|
||||
}
|
||||
}
|
||||
|
||||
// Whatever the destination, the walk is a subsequence of the table. Every
|
||||
// comment on that table argues an order between two sources, and none of those
|
||||
// reasons is about this field.
|
||||
func TestTheWalkNeverReordersTheTable(t *testing.T) {
|
||||
for _, dest := range append([]router.Source{router.SourceUnknown}, router.Sources...) {
|
||||
walk, skipped := queryWalk(dest, true)
|
||||
if len(walk)+len(skipped) != len(querySources) {
|
||||
t.Errorf("%q: %d walked + %d skipped, want %d", dest, len(walk), len(skipped), len(querySources))
|
||||
}
|
||||
last := -1
|
||||
for _, s := range walk {
|
||||
at := posOf(querySources, s.name)
|
||||
if at <= last {
|
||||
t.Errorf("%q: %q is out of table order", dest, s.name)
|
||||
}
|
||||
last = at
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func inWalk(list []querySource, name string) bool { return posOf(list, name) >= 0 }
|
||||
|
||||
func posOf(list []querySource, name string) int {
|
||||
for i, s := range list {
|
||||
if s.name == name {
|
||||
return i
|
||||
}
|
||||
}
|
||||
return -1
|
||||
}
|
||||
@@ -22,7 +22,7 @@ func TestReactiveNotesReminders(t *testing.T) {
|
||||
|
||||
emb := router.NewHashEmbedder(1024)
|
||||
matcher := tool.NewMatcher(api)
|
||||
rtr := buildRouter(emb, matcher, 0.55, nil)
|
||||
rtr := buildRouter(emb, matcher, 0.55, nil, nil)
|
||||
|
||||
h := &reactiveHandler{
|
||||
api: api,
|
||||
@@ -104,7 +104,7 @@ func TestSpokenTaskCaptureFilesATask(t *testing.T) {
|
||||
h := &reactiveHandler{
|
||||
api: api,
|
||||
recall: recallWiring{embedder: emb, memStore: memory.NewInMemoryStore()},
|
||||
router: buildRouter(emb, matcher, 0.55, nil),
|
||||
router: buildRouter(emb, matcher, 0.55, nil, nil),
|
||||
replier: voice.NewStubReplier(),
|
||||
now: func() time.Time { return now },
|
||||
dataStore: st,
|
||||
|
||||
@@ -474,7 +474,7 @@ func newSimWorld(t *testing.T, sc scenario) *simWorld {
|
||||
// used to be built on a nil API, which meant any scenario that produced an
|
||||
// act panicked the moment the matcher was consulted.
|
||||
matcher := tool.NewMatcher(api)
|
||||
rtr := buildRouter(emb, matcher, config.DefaultRouterThreshold, router.NewLLMRouter(scripted))
|
||||
rtr := buildRouter(emb, matcher, config.DefaultRouterThreshold, router.NewLLMRouter(scripted), nil)
|
||||
|
||||
w.handler = &reactiveHandler{
|
||||
stt: simTranscriber{},
|
||||
|
||||
@@ -0,0 +1,43 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"testing"
|
||||
|
||||
"github.com/kami/maven/internal/config"
|
||||
"github.com/kami/maven/internal/stt"
|
||||
)
|
||||
|
||||
// A box with no workstation.stt block transcribes exactly as it did before the
|
||||
// seam existed: the floor is handed back untouched, and nothing probes.
|
||||
func TestSttSeamWithNoBlockIsTheFloor(t *testing.T) {
|
||||
floor := stt.NewStub()
|
||||
got, pair := sttSeam(&config.Config{}, floor)
|
||||
if pair != nil {
|
||||
t.Fatal("no block must build no pair")
|
||||
}
|
||||
if got != stt.Transcriber(floor) {
|
||||
t.Fatal("no block must hand back the floor itself")
|
||||
}
|
||||
}
|
||||
|
||||
func TestSttSeamPrefersTheWorkstation(t *testing.T) {
|
||||
cfg := &config.Config{Workstation: &config.WorkstationConfig{
|
||||
URL: "http://127.0.0.1:1",
|
||||
Stt: &config.WorkstationSttConfig{
|
||||
URL: "http://127.0.0.1:2/transcribe",
|
||||
Health: "http://127.0.0.1:2/health",
|
||||
},
|
||||
}}
|
||||
got, pair := sttSeam(cfg, stt.NewStub())
|
||||
if pair == nil {
|
||||
t.Fatal("a configured block must build a pair")
|
||||
}
|
||||
defer pair.Stop()
|
||||
if got != stt.Transcriber(pair) {
|
||||
t.Fatal("the pair is what callers must transcribe through")
|
||||
}
|
||||
// Nothing answers on port 2, so the seam is the floor until it does.
|
||||
if pair.Available() {
|
||||
t.Fatal("an unreachable workstation must not be available")
|
||||
}
|
||||
}
|
||||
@@ -186,6 +186,28 @@ func carriesReminderVerb(text string) bool {
|
||||
return false
|
||||
}
|
||||
|
||||
// isPleasantry matches the WHOLE utterance against lexicon.Pleasantries, after
|
||||
// lowercasing and dropping the punctuation a greeting carries.
|
||||
//
|
||||
// Whole utterance and not tokens. Every token rule tried here was wrong on
|
||||
// something: "вечер" answers "это утра или вечера?", "нет" answers a confirm,
|
||||
// and "спокойной" alone is not an utterance at all. A greeting is a fixed
|
||||
// phrase, so matching it as one costs nothing and claims nothing else.
|
||||
func isPleasantry(text string) bool {
|
||||
t := strings.ToLower(strings.TrimSpace(text))
|
||||
t = strings.Trim(t, " .,!?…")
|
||||
t = strings.Join(strings.Fields(t), " ")
|
||||
if t == "" {
|
||||
return false
|
||||
}
|
||||
for _, p := range lexicon.Pleasantries() {
|
||||
if t == p {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// offlineOwnRequest is the shape half of the evidence: the offline token tests,
|
||||
// which cost nothing and never depend on the model that produced the routing.
|
||||
// It is also the whole answer when there is no route to read — the classifier
|
||||
@@ -225,6 +247,16 @@ func classifyTurnRole(q *dialogue.PendingQuestion, text string, answer dialogue.
|
||||
// hour, and no route saying "question" changes that. It works because the
|
||||
// extractor no longer reads a day word as the current clock, so a sentence
|
||||
// that names no hour now fills nothing to weigh.
|
||||
// A pleasantry is neither (V-663). "спасибо" and "привет" fell through to
|
||||
// roleAnswer, so a question about a reminder's DAY was re-asked at a man
|
||||
// saying thank you, and the retry it spent was one of the three bounds
|
||||
// meant to end the ride. It is an aside: answered as itself, the question
|
||||
// resumed on the tail, no attempt spent, one ride counted. Placed above the
|
||||
// content gate because "доброе утро" has content and states nothing, so
|
||||
// neither half of the evidence below can reach it.
|
||||
if q != nil && isPleasantry(text) {
|
||||
return roleAside
|
||||
}
|
||||
own := false
|
||||
if len(ownContent(text)) > 0 {
|
||||
own = offlineOwnRequest(text) || (ok && carriesOwnRequest(routed, text))
|
||||
|
||||
@@ -279,3 +279,162 @@ func TestTheTurnIsRoutedOnce(t *testing.T) {
|
||||
t.Fatalf("the pipeline routed again and got something else: %+v vs %+v", second, first)
|
||||
}
|
||||
}
|
||||
|
||||
// TestASuspendedQuestionDoesNotRideForever — V-654, the measured failure of
|
||||
// 2026-08-07 (docs/evals/2026-08-07-week-of-usage-transcript.md, t=51 to t=58).
|
||||
//
|
||||
// A side query suspends the parked question, spends no attempt and restarts the
|
||||
// TTL. Nothing else bounded it, so one unfilled time slot came back on the end
|
||||
// of six consecutive unrelated replies and stopped only when a seventh turn
|
||||
// happened to read as a failed answer. Three step-asides, then she lets it go
|
||||
// and says so.
|
||||
func TestASuspendedQuestionDoesNotRideForever(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
h, st := newRoutingClarifyHandler(t)
|
||||
id := dialogueIDFor(sourceText, "web")
|
||||
resumed, _ := clarifyResumedFor(dialogue.SlotTime)
|
||||
|
||||
if reply := h.handleText(ctx, "web", "напомни позвонить маме"); !strings.Contains(reply, "?") {
|
||||
t.Fatalf("expected the time question, got %q", reply)
|
||||
}
|
||||
|
||||
// Three questions of his own. Each one is answered as itself and each one
|
||||
// brings the open question back, exactly as V-561 asks.
|
||||
asides := []string{
|
||||
"о чём мы вчера говорили?",
|
||||
"какие у меня напоминания?",
|
||||
"сколько времени?",
|
||||
}
|
||||
for i, text := range asides {
|
||||
reply := h.handleText(ctx, "web", text)
|
||||
if !strings.HasSuffix(reply, resumed) {
|
||||
t.Fatalf("side query %d: the question must come back, got %q", i+1, reply)
|
||||
}
|
||||
if strings.Contains(reply, clarifyDropped) {
|
||||
t.Fatalf("side query %d: nothing was let go yet, so nothing may say so: %q", i+1, reply)
|
||||
}
|
||||
q := h.clarifyStore.Get(id, h.now())
|
||||
if q == nil {
|
||||
t.Fatalf("side query %d: the question was dropped early", i+1)
|
||||
}
|
||||
if q.Attempts != 1 {
|
||||
t.Fatalf("side query %d: a step-aside spent an attempt: %d", i+1, q.Attempts)
|
||||
}
|
||||
if q.Suspends != i+1 {
|
||||
t.Fatalf("side query %d: suspends = %d, want %d", i+1, q.Suspends, i+1)
|
||||
}
|
||||
}
|
||||
|
||||
// The fourth. She has stepped aside as often as she is willing to, so the
|
||||
// request goes — out loud, and without the question on the tail.
|
||||
reply := h.handleText(ctx, "web", "что у меня сегодня?")
|
||||
if !strings.Contains(reply, clarifyDropped) {
|
||||
t.Fatalf("the request was let go in silence: %q", reply)
|
||||
}
|
||||
if strings.HasSuffix(reply, resumed) {
|
||||
t.Fatalf("a question she has let go must not be asked again: %q", reply)
|
||||
}
|
||||
if h.clarifyStore.Get(id, h.now()) != nil {
|
||||
t.Fatal("the question must be gone once she has said she let it go")
|
||||
}
|
||||
if reminders, err := st.DueReminders(ctx, h.now().Add(48*time.Hour)); err != nil || len(reminders) != 0 {
|
||||
t.Fatalf("a reminder was invented for a time nobody gave: %v err=%v", reminders, err)
|
||||
}
|
||||
}
|
||||
|
||||
// TestAnAnsweredGapResetsTheSuspendBudget — the counter measures CONSECUTIVE
|
||||
// step-asides. He filled a gap, so the run is broken and the next question
|
||||
// starts with its full allowance: a long exchange he is engaged with must not
|
||||
// run out of patience on his behalf.
|
||||
func TestAnAnsweredGapResetsTheSuspendBudget(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
h, _ := newRoutingClarifyHandler(t)
|
||||
id := dialogueIDFor(sourceText, "web")
|
||||
|
||||
// A bare "напомни" is missing both halves, so answering the subject re-parks
|
||||
// the request with a question about the time.
|
||||
if reply := h.handleText(ctx, "web", "напомни"); !strings.Contains(reply, "?") {
|
||||
t.Fatalf("expected a question, got %q", reply)
|
||||
}
|
||||
if reply := h.handleText(ctx, "web", "какие у меня напоминания?"); reply == "" {
|
||||
t.Fatal("the side query must be answered as itself")
|
||||
}
|
||||
if q := h.clarifyStore.Get(id, h.now()); q == nil || q.Suspends != 1 {
|
||||
t.Fatalf("the side query was not counted: %+v", q)
|
||||
}
|
||||
if reply := h.handleText(ctx, "web", "позвонить маме"); reply == "" {
|
||||
t.Fatal("the answer must be consumed")
|
||||
}
|
||||
q := h.clarifyStore.Get(id, h.now())
|
||||
if q == nil {
|
||||
t.Fatal("a reminder still needs its time, so a question must be parked")
|
||||
}
|
||||
if q.Suspends != 0 {
|
||||
t.Fatalf("answering a gap must reset the suspend budget: suspends = %d", q.Suspends)
|
||||
}
|
||||
// The ride it already took is carried across the re-park (V-663). Resetting
|
||||
// both counters here is what let one question ride twenty-six replies.
|
||||
if q.Rides != 1 {
|
||||
t.Fatalf("the aside it already took was forgotten: rides = %d", q.Rides)
|
||||
}
|
||||
}
|
||||
|
||||
// TestTwoBoundsCannotRearmEachOther — V-663.
|
||||
//
|
||||
// MaxSuspends landed and the measurement did not move: twenty-six of 140 turns
|
||||
// carried a tail before it and twenty-six after. This is the shape it misses,
|
||||
// taken from the 2026-08-08 run, where one question rode turns 7 to 13.
|
||||
//
|
||||
// An aside spends no attempt, so MaxAttempts never reaches it. A turn that
|
||||
// reads as a failed answer zeroes Suspends, so MaxSuspends never reaches the
|
||||
// asides either. Alternating the two rearms each bound with the other's
|
||||
// traffic. Rides counts both kinds and is never reset, so it is what ends this.
|
||||
func TestTwoBoundsCannotRearmEachOther(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
h, _ := newRoutingClarifyHandler(t)
|
||||
id := dialogueIDFor(sourceText, "web")
|
||||
resumed, _ := clarifyResumedFor(dialogue.SlotTime)
|
||||
|
||||
if reply := h.handleText(ctx, "web", "напомни позвонить маме"); !strings.Contains(reply, "?") {
|
||||
t.Fatalf("expected the time question, got %q", reply)
|
||||
}
|
||||
|
||||
// Two asides. Each one rides and neither spends an attempt.
|
||||
for i := 0; i < 2; i++ {
|
||||
reply := h.handleText(ctx, "web", "какие у меня напоминания?")
|
||||
if !strings.HasSuffix(reply, resumed) {
|
||||
t.Fatalf("aside %d: the question must come back, got %q", i+1, reply)
|
||||
}
|
||||
}
|
||||
q := h.clarifyStore.Get(id, h.now())
|
||||
if q == nil || q.Rides != 2 || q.Suspends != 2 {
|
||||
t.Fatalf("after two asides: %+v", q)
|
||||
}
|
||||
|
||||
// A pleasantry. It used to read as a failed answer, so she re-asked the
|
||||
// question at a man saying thank you and spent an attempt doing it. Now it
|
||||
// is an aside: answered as itself, question on the tail, one more ride.
|
||||
reply := h.handleText(ctx, "web", "спасибо")
|
||||
if !strings.HasSuffix(reply, resumed) {
|
||||
t.Fatalf("a pleasantry lost the parked question: %q", reply)
|
||||
}
|
||||
q = h.clarifyStore.Get(id, h.now())
|
||||
if q == nil || q.Attempts != 1 {
|
||||
t.Fatalf("a pleasantry spent an attempt: %+v", q)
|
||||
}
|
||||
if q.Rides != 3 {
|
||||
t.Fatalf("a pleasantry rode free: %+v", q)
|
||||
}
|
||||
|
||||
// One more ride of any kind and the request goes, out loud.
|
||||
reply = h.handleText(ctx, "web", "какие у меня напоминания?")
|
||||
if !strings.Contains(reply, clarifyDropped) {
|
||||
t.Fatalf("the question rode four asides and was let go in silence: %q", reply)
|
||||
}
|
||||
if strings.HasSuffix(reply, resumed) {
|
||||
t.Fatalf("a question she has let go must not be asked again: %q", reply)
|
||||
}
|
||||
if h.clarifyStore.Get(id, h.now()) != nil {
|
||||
t.Fatal("the question must be gone once she has said she let it go")
|
||||
}
|
||||
}
|
||||
|
||||
+80
-4
@@ -36,7 +36,9 @@ type voiceWiring struct {
|
||||
sessions *voice.Sessions
|
||||
voiceSink delivery.Sink
|
||||
embedder router.Embedder
|
||||
handler *reactiveHandler // the reactive handler for IPC Chat
|
||||
// heads — the routing heads, nil unless embedder.heads_path is set.
|
||||
heads *router.RouterHeads
|
||||
handler *reactiveHandler // the reactive handler for IPC Chat
|
||||
// worker clients (set when configured as Remote): closed on shutdown so
|
||||
// mavsttd / mavttsd don't keep a stale conn into a restarting daemon.
|
||||
sttClient *worker.Client
|
||||
@@ -53,7 +55,11 @@ type voiceWiring struct {
|
||||
// unless a `workstation` block names an address. Held here only so the
|
||||
// prober is stopped on shutdown; callers were handed it at build time.
|
||||
pair *llm.Pair
|
||||
mcp *mcpWiring
|
||||
// sttPair — CrisperWhisper 2.0 on the workstation with mavsttd as the
|
||||
// floor, nil unless the `workstation.stt` block names an address. Held for
|
||||
// the same reason as pair: to stop its prober on shutdown.
|
||||
sttPair *stt.Pair
|
||||
mcp *mcpWiring
|
||||
// home — the Home Assistant client, nil unless the `smarthome` block is
|
||||
// enabled (Vikunja #256). Its devices land in the same allowlist as every
|
||||
// other act, so nothing else here has to know about it.
|
||||
@@ -72,6 +78,9 @@ func (w *voiceWiring) close() {
|
||||
if w.embedder != nil {
|
||||
_ = w.embedder.Close()
|
||||
}
|
||||
if w.heads != nil {
|
||||
_ = w.heads.Close()
|
||||
}
|
||||
if w.server != nil {
|
||||
_ = w.server.Close()
|
||||
}
|
||||
@@ -84,6 +93,9 @@ func (w *voiceWiring) close() {
|
||||
if w.pair != nil {
|
||||
w.pair.Stop()
|
||||
}
|
||||
if w.sttPair != nil {
|
||||
w.sttPair.Stop()
|
||||
}
|
||||
w.mcp.close()
|
||||
}
|
||||
|
||||
@@ -112,6 +124,7 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
|
||||
} else {
|
||||
transcriber = stt.NewStub()
|
||||
}
|
||||
transcriber, w.sttPair = sttSeam(cfg, transcriber)
|
||||
w.transcriber = transcriber
|
||||
|
||||
// ----- tts (Stub in-process OR Remote) -----
|
||||
@@ -147,6 +160,24 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
|
||||
emb = router.NewHashEmbedder(1024)
|
||||
}
|
||||
w.embedder = emb
|
||||
|
||||
// ----- router: routing heads (only when configured, and never fatal) -----
|
||||
// A missing or broken weights file logs and leaves w.heads nil, which is
|
||||
// byte-for-byte the cascade that shipped before V-664. Refusing to start
|
||||
// over a routing accelerator would trade a working box for a better one.
|
||||
if cfg.Voice.Embedder != nil && cfg.Voice.Embedder.HeadsPath != "" {
|
||||
h, err := router.NewRouterHeads(
|
||||
cfg.Voice.Embedder.HeadsPath,
|
||||
cfg.Voice.Embedder.TokenizerPath,
|
||||
)
|
||||
if err != nil {
|
||||
log.Printf("voice: routing heads unavailable, cascade unchanged: %v", err)
|
||||
} else {
|
||||
log.Printf("voice: routing heads loaded from %s", cfg.Voice.Embedder.HeadsPath)
|
||||
w.heads = h
|
||||
}
|
||||
}
|
||||
|
||||
repairFactVectors(dataStore, emb)
|
||||
checkStoredEmbedder(dataStore, emb)
|
||||
// Retention is enforced on write, which is not enough on its own: a box that
|
||||
@@ -223,7 +254,8 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
|
||||
// against the classifier's 50.0%, at about 1s a turn instead of 30ms (see
|
||||
// config.VoiceConfig.LLMRouter). The classifier always stays wired as the
|
||||
// fallback, so a model error never breaks a turn.
|
||||
rtr := buildRouter(emb, matcher, threshold, pickLLMRouter(cfg.Voice.UseLLMRouter(), hot))
|
||||
rtr := buildRouter(emb, matcher, threshold,
|
||||
pickLLMRouter(cfg.Voice.UseLLMRouter(), hot), w.heads)
|
||||
|
||||
// ----- sessions registry (shared with voicesink) -----
|
||||
sessions := voice.NewSessions()
|
||||
@@ -366,6 +398,44 @@ func modelSeam(cfg *config.Config, resident *llm.Client) (router.Completer, *llm
|
||||
return pair, pair
|
||||
}
|
||||
|
||||
// sttSeam builds the transcription seam the voice path and the meeting
|
||||
// recorder share. It is modelSeam for audio and follows the same rule.
|
||||
//
|
||||
// With no `workstation.stt` block it hands back the floor untouched, which is
|
||||
// today's deploy exactly. With one, it is an stt.Pair preferring CrisperWhisper
|
||||
// 2.0 on workpc, which scores 10.4% WER in Russian against the floor's 27.5%
|
||||
// (docs/evals/2026-08-09-crisperwhisper2-russian-wer.md).
|
||||
//
|
||||
// Only the silent half of the degradation rule applies here. A worse transcript
|
||||
// is still a turn, so there is nothing to name a gap about and the fallback is
|
||||
// never spoken. That is why stt.Pair has no TranscribeRemote.
|
||||
func sttSeam(cfg *config.Config, floor stt.Transcriber) (stt.Transcriber, *stt.Pair) {
|
||||
if cfg.Workstation == nil || cfg.Workstation.Stt == nil {
|
||||
return floor, nil
|
||||
}
|
||||
s := cfg.Workstation.Stt
|
||||
lang := ""
|
||||
if cfg.Voice != nil {
|
||||
lang = cfg.Voice.Lang
|
||||
if cfg.Voice.Stt != nil && cfg.Voice.Stt.Lang != "" {
|
||||
lang = cfg.Voice.Stt.Lang
|
||||
}
|
||||
}
|
||||
pair := stt.NewPair(
|
||||
stt.NewHTTPTranscriber(s.URL, s.Token, lang, time.Duration(s.Timeout)),
|
||||
floor,
|
||||
s.Health,
|
||||
time.Duration(s.Probe),
|
||||
)
|
||||
pair.Start(context.Background())
|
||||
if s.Token == "" {
|
||||
log.Print("voice: the workstation transcriber has no token, so anything on the LAN can post audio to it")
|
||||
}
|
||||
log.Printf("voice: workstation transcriber at %s, probed every %s, mavsttd as the floor",
|
||||
s.URL, time.Duration(s.Probe))
|
||||
return pair, pair
|
||||
}
|
||||
|
||||
func pickLLMRouter(enabled bool, c router.Completer) *router.LLMRouter {
|
||||
if !enabled {
|
||||
return nil
|
||||
@@ -390,7 +460,8 @@ func pickLLMRouter(enabled bool, c router.Completer) *router.LLMRouter {
|
||||
// intent from seedDir (models/seeds/<intent>.txt) — see seedClassifier
|
||||
// below for the current intent list and file names.
|
||||
// - Threshold is from voice.router_threshold config (default 0.55).
|
||||
func buildRouter(emb router.Embedder, acts router.ActMatcher, threshold float64, llmR *router.LLMRouter) *router.Router {
|
||||
func buildRouter(emb router.Embedder, acts router.ActMatcher, threshold float64,
|
||||
llmR *router.LLMRouter, heads *router.RouterHeads) *router.Router {
|
||||
cls := router.NewClassifier(emb)
|
||||
seedClassifier(cls)
|
||||
grammars := router.DefaultGrammars(acts)
|
||||
@@ -401,6 +472,10 @@ func buildRouter(emb router.Embedder, acts router.ActMatcher, threshold float64,
|
||||
grammars = append(grammars, router.AgendaQueryGrammars()...)
|
||||
// Same reason as the agenda rules, for the feeds: "что нового в лентах?"
|
||||
// routed system and answered "пока не умею" (Vikunja #474).
|
||||
// After the agenda rules, which are the narrower claim, and BEFORE the feed
|
||||
// and list rules, which are not: "что такое лента" is a definition question
|
||||
// and the feed rule would take it on the noun alone (V-655).
|
||||
grammars = append(grammars, router.WorldQueryGrammars()...)
|
||||
grammars = append(grammars, router.FeedQueryGrammar())
|
||||
// The list side of the same exposure: a phrasing with no possessive in it
|
||||
// ("список дел") routed system and never reached queryTasks (Vikunja #467).
|
||||
@@ -438,6 +513,7 @@ func buildRouter(emb router.Embedder, acts router.ActMatcher, threshold float64,
|
||||
},
|
||||
Threshold: threshold,
|
||||
LLM: llmR,
|
||||
Heads: heads,
|
||||
})
|
||||
}
|
||||
|
||||
|
||||
+13
-4
@@ -34,22 +34,31 @@ type probe struct {
|
||||
drmDev string
|
||||
}
|
||||
|
||||
// foreign lists every ROCm process that is not ours. selfPID is the supervisor's
|
||||
// llama-server child, or 0 when it is not running.
|
||||
// foreign lists every ROCm process that is not ours. self holds the pids of the
|
||||
// supervisor's own children, and a child that is not running contributes 0.
|
||||
//
|
||||
// There is more than one child since 09-08-2026. CW2 registers on the KFD like
|
||||
// any ROCm job, so a supervisor that excluded only llama-server would read its
|
||||
// own transcriber as a contender, yield the card to it, and never keep a model
|
||||
// loaded again.
|
||||
//
|
||||
// An unreadable kfd tree returns no processes and no error. That is deliberate
|
||||
// and it is the safe direction only because startVRAM also has to agree before
|
||||
// anything launches: a supervisor that cannot see the KFD never sees free VRAM
|
||||
// either, because the CPT run holding the card shows up in the drm totals.
|
||||
func (p probe) foreign(selfPID int) []gpuProc {
|
||||
func (p probe) foreign(self ...int) []gpuProc {
|
||||
entries, err := os.ReadDir(p.kfdRoot)
|
||||
if err != nil {
|
||||
return nil
|
||||
}
|
||||
mine := make(map[int]bool, len(self))
|
||||
for _, pid := range self {
|
||||
mine[pid] = true
|
||||
}
|
||||
var out []gpuProc
|
||||
for _, e := range entries {
|
||||
pid, err := strconv.Atoi(e.Name())
|
||||
if err != nil || pid == selfPID {
|
||||
if err != nil || mine[pid] {
|
||||
continue
|
||||
}
|
||||
out = append(out, gpuProc{
|
||||
|
||||
+46
-1
@@ -1,6 +1,7 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"net/url"
|
||||
@@ -8,6 +9,7 @@ import (
|
||||
"path/filepath"
|
||||
"strconv"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// fakeKFD builds the sysfs shape the workstation actually has: one directory
|
||||
@@ -47,6 +49,24 @@ func TestForeignExcludesOurChild(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// The transcriber is a ROCm process on the same card, so it registers on the
|
||||
// KFD exactly like a contender does. Reading it as one is what happened on
|
||||
// 2026-08-09 while CW2 ran under its own systemd unit: mavgpud yielded, waited
|
||||
// five polls, loaded the model, yielded again, and never held it for a whole
|
||||
// minute. Excluding every child is the fix and this is the test of it.
|
||||
func TestForeignExcludesEveryChild(t *testing.T) {
|
||||
p := probe{kfdRoot: fakeKFD(t, map[int]int64{478104: 12791693312, 999: 4096, 1001: 1717986918})}
|
||||
|
||||
ours := p.foreign(999, 1001)
|
||||
if len(ours) != 1 || ours[0].PID != 478104 {
|
||||
t.Fatalf("only the CPT run is a contender, got %+v", ours)
|
||||
}
|
||||
// A child that is not running reports pid 0, which must exclude nothing.
|
||||
if got := p.foreign(999, 0); len(got) != 2 {
|
||||
t.Errorf("a stopped child excludes nobody: got %d contenders, want 2", len(got))
|
||||
}
|
||||
}
|
||||
|
||||
// An empty KFD tree is the state that permits a start, so it must read as empty
|
||||
// rather than as an error the caller has to interpret.
|
||||
func TestForeignEmptyAndMissing(t *testing.T) {
|
||||
@@ -81,7 +101,7 @@ func TestFreeVRAM(t *testing.T) {
|
||||
// rather than hanging or proxying into a closed port. Maven reads this endpoint
|
||||
// on a timer forever, including while the workstation is busy.
|
||||
func TestHealthAndProxyRefuseWhenNotReady(t *testing.T) {
|
||||
s := &supervisor{run: newRunner("/bin/true", nil, "")}
|
||||
s := &supervisor{run: newRunner("fake", "/bin/true", nil, "")}
|
||||
h := s.handler(mustURL(t, "http://127.0.0.1:1"))
|
||||
|
||||
for _, path := range []string{"/health", "/v1/chat/completions"} {
|
||||
@@ -101,3 +121,28 @@ func mustURL(t *testing.T, s string) *url.URL {
|
||||
}
|
||||
return u
|
||||
}
|
||||
|
||||
// Yielding is all or nothing. A CPT run wants the whole card, so handing back
|
||||
// the language model while the transcriber keeps 1.6GB mapped would leave the
|
||||
// other job failing its allocation, which is the outcome yielding exists to
|
||||
// prevent.
|
||||
func TestYieldStopsEveryChild(t *testing.T) {
|
||||
idle := "while : ; do sleep 1 ; done"
|
||||
s := &supervisor{
|
||||
cfg: config{EvictAfter: 1, StopGrace: duration(2 * time.Second)},
|
||||
probe: probe{kfdRoot: fakeKFD(t, map[int]int64{478104: 12791693312})},
|
||||
run: newRunner("llama-server", fakeServer(t, idle), nil, ""),
|
||||
stt: newRunner("cw2", fakeServer(t, idle), nil, ""),
|
||||
}
|
||||
for _, r := range s.children() {
|
||||
if err := r.start(); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
s.tick(context.Background())
|
||||
for _, r := range s.children() {
|
||||
if r.running() {
|
||||
t.Errorf("%s outlived the yield", r.name)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
+90
-17
@@ -10,6 +10,11 @@
|
||||
// the card. Not on demand, because a 7-14B takes tens of seconds to load and a
|
||||
// world question would be answered by a gap every time the card had been quiet.
|
||||
// Not always on, because that holds 16GB against the owner's own jobs.
|
||||
//
|
||||
// It supervises a second child since 09-08-2026, the CW2 transcriber, and for
|
||||
// one reason only: it is a ROCm process on the same card. Any GPU service the
|
||||
// owner leaves running beside this daemon reads as a contender and evicts the
|
||||
// model, so the card needs one owner rather than two neighbours.
|
||||
package main
|
||||
|
||||
import (
|
||||
@@ -36,6 +41,10 @@ type config struct {
|
||||
// owner's business and not this daemon's schema.
|
||||
LlamaArgs []string `json:"llama_args"`
|
||||
|
||||
// Stt is optional. Without it mavgpud supervises llama-server alone, which
|
||||
// is everything it did before 09-08-2026.
|
||||
Stt *sttConfig `json:"stt,omitempty"`
|
||||
|
||||
KFDRoot string `json:"kfd_root"`
|
||||
DRMDevice string `json:"drm_device"`
|
||||
|
||||
@@ -51,6 +60,22 @@ type config struct {
|
||||
StartAfter int `json:"start_after_polls"`
|
||||
}
|
||||
|
||||
// sttConfig is the CW2 transcriber, which mavgpud runs for one reason: it is a
|
||||
// ROCm process on this card. Left to its own systemd unit it registers on the
|
||||
// KFD, the supervisor reads it as a contender, and llama-server is evicted
|
||||
// within two polls and restarted five polls later, forever. That thrash was
|
||||
// observed on 2026-08-09 and it is what folded the service in here.
|
||||
//
|
||||
// Maven talks to it directly, not through this daemon. There is no proxy and no
|
||||
// idle timer: at 1.6GB it denies the card to nobody, and unloading it would only
|
||||
// send the next voice turn to the homesrv floor for no gain.
|
||||
type sttConfig struct {
|
||||
// Addr is where the service binds, and it is read only to probe /health.
|
||||
Addr string `json:"addr"`
|
||||
Bin string `json:"bin"`
|
||||
Args []string `json:"args"`
|
||||
}
|
||||
|
||||
func defaults() config {
|
||||
return config{
|
||||
Listen: ":8080",
|
||||
@@ -99,12 +124,18 @@ func main() {
|
||||
}
|
||||
|
||||
base := "http://" + cfg.LlamaAddr
|
||||
run := newRunner(cfg.LlamaBin, cfg.LlamaArgs, base+"/health")
|
||||
run := newRunner("llama-server", cfg.LlamaBin, cfg.LlamaArgs, base+"/health")
|
||||
sup := &supervisor{
|
||||
cfg: cfg,
|
||||
probe: probe{kfdRoot: cfg.KFDRoot, drmDev: cfg.DRMDevice},
|
||||
run: run,
|
||||
}
|
||||
if s := cfg.Stt; s != nil {
|
||||
if s.Bin == "" || s.Addr == "" {
|
||||
log.Fatal("mavgpud: stt needs both bin and addr")
|
||||
}
|
||||
sup.stt = newRunner("cw2", s.Bin, s.Args, "http://"+s.Addr+"/health")
|
||||
}
|
||||
sup.touch()
|
||||
|
||||
ctx, cancel := signal.NotifyContext(context.Background(), syscall.SIGINT, syscall.SIGTERM)
|
||||
@@ -129,13 +160,17 @@ func main() {
|
||||
shut, done := context.WithTimeout(context.Background(), 5*time.Second)
|
||||
defer done()
|
||||
_ = srv.Shutdown(shut)
|
||||
run.stop(time.Duration(cfg.StopGrace))
|
||||
for _, r := range sup.children() {
|
||||
r.stop(time.Duration(cfg.StopGrace))
|
||||
}
|
||||
}
|
||||
|
||||
type supervisor struct {
|
||||
cfg config
|
||||
probe probe
|
||||
run *runner
|
||||
// stt is the CW2 transcriber, or nil when the config names none.
|
||||
stt *runner
|
||||
|
||||
lastReq atomic.Int64 // unix nanos of the last request Maven sent
|
||||
|
||||
@@ -198,7 +233,11 @@ func (s *supervisor) loop(ctx context.Context) {
|
||||
// allocates, so we see a contender during its startup rather than after it has
|
||||
// already failed to get the memory it wanted.
|
||||
func (s *supervisor) tick(ctx context.Context) {
|
||||
others := s.probe.foreign(s.run.pid())
|
||||
var pids []int
|
||||
for _, r := range s.children() {
|
||||
pids = append(pids, r.pid())
|
||||
}
|
||||
others := s.probe.foreign(pids...)
|
||||
if len(others) > 0 {
|
||||
s.foreignStreak++
|
||||
s.clearStreak = 0
|
||||
@@ -207,31 +246,65 @@ func (s *supervisor) tick(ctx context.Context) {
|
||||
s.clearStreak++
|
||||
}
|
||||
|
||||
if s.run.running() {
|
||||
s.run.refreshReady(ctx)
|
||||
switch {
|
||||
case s.foreignStreak >= s.cfg.EvictAfter:
|
||||
log.Printf("mavgpud: yielding the card to %s", describe(others))
|
||||
s.run.stop(time.Duration(s.cfg.StopGrace))
|
||||
case s.idle() > time.Duration(s.cfg.IdleTimeout):
|
||||
log.Printf("mavgpud: idle for %s, unloading", s.idle().Round(time.Second))
|
||||
s.run.stop(time.Duration(s.cfg.StopGrace))
|
||||
// Yielding is all or nothing. A CPT run wants the whole card, and handing
|
||||
// back 8GB while holding 1.6GB is the shape of a failed allocation.
|
||||
if s.foreignStreak >= s.cfg.EvictAfter && s.anyRunning() {
|
||||
log.Printf("mavgpud: yielding the card to %s", describe(others))
|
||||
for _, r := range s.children() {
|
||||
r.stop(time.Duration(s.cfg.StopGrace))
|
||||
}
|
||||
return
|
||||
}
|
||||
|
||||
if s.clearStreak < s.cfg.StartAfter {
|
||||
clear := s.clearStreak >= s.cfg.StartAfter
|
||||
|
||||
if s.run.running() {
|
||||
s.run.refreshReady(ctx)
|
||||
if s.idle() > time.Duration(s.cfg.IdleTimeout) {
|
||||
log.Printf("mavgpud: idle for %s, unloading", s.idle().Round(time.Second))
|
||||
s.run.stop(time.Duration(s.cfg.StopGrace))
|
||||
}
|
||||
} else if clear && s.probe.freeVRAM() >= s.cfg.MinFreeVRAM {
|
||||
s.touch() // the idle clock starts at load, not at the last request before it
|
||||
if err := s.run.start(); err != nil {
|
||||
log.Printf("mavgpud: start llama-server: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
if s.stt == nil {
|
||||
return
|
||||
}
|
||||
if free := s.probe.freeVRAM(); free < s.cfg.MinFreeVRAM {
|
||||
if s.stt.running() {
|
||||
s.stt.refreshReady(ctx)
|
||||
return
|
||||
}
|
||||
s.touch() // the idle clock starts at load, not at the last request before it
|
||||
if err := s.run.start(); err != nil {
|
||||
log.Printf("mavgpud: start llama-server: %v", err)
|
||||
// No VRAM precondition here, unlike llama-server. That check exists because
|
||||
// a 12B refuses to load when the card is short, and 1.6GB fits wherever the
|
||||
// KFD is clear. Reading free VRAM would also block the transcriber for good
|
||||
// once the language model was resident, since it holds more than the floor.
|
||||
if clear {
|
||||
if err := s.stt.start(); err != nil {
|
||||
log.Printf("mavgpud: start cw2: %v", err)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func (s *supervisor) children() []*runner {
|
||||
if s.stt == nil {
|
||||
return []*runner{s.run}
|
||||
}
|
||||
return []*runner{s.run, s.stt}
|
||||
}
|
||||
|
||||
func (s *supervisor) anyRunning() bool {
|
||||
for _, r := range s.children() {
|
||||
if r.running() {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// describe names the contenders in the log. This log is the instrument for the
|
||||
// open question in #488: whether polling the KFD misses a job that wants the
|
||||
// card without registering there.
|
||||
|
||||
+17
-13
@@ -10,14 +10,18 @@ import (
|
||||
"time"
|
||||
)
|
||||
|
||||
// runner owns one llama-server process. Owning it is the point of the daemon:
|
||||
// the workstation cannot keep a 7-14B resident, because that holds 16GB against
|
||||
// runner owns one GPU process. Owning it is the point of the daemon: the
|
||||
// workstation cannot keep a 7-14B resident, because that holds 16GB against
|
||||
// the owner's CPT runs, Correx and the manga-recap pipeline. So the thing that
|
||||
// stays up is this, which costs no VRAM, and the model comes and goes under it.
|
||||
//
|
||||
// There are two of them since 09-08-2026: llama-server and the CW2 transcriber.
|
||||
// name is what the log calls this one.
|
||||
type runner struct {
|
||||
name string
|
||||
bin string
|
||||
args []string
|
||||
// ready is llama-server's own /health, which answers "is a model loaded".
|
||||
// ready is the child's own /health, which answers "is a model loaded".
|
||||
// Loading a 7-14B takes tens of seconds, so started is not ready.
|
||||
readyURL string
|
||||
|
||||
@@ -32,9 +36,9 @@ type runner struct {
|
||||
http *http.Client
|
||||
}
|
||||
|
||||
func newRunner(bin string, args []string, readyURL string) *runner {
|
||||
func newRunner(name, bin string, args []string, readyURL string) *runner {
|
||||
return &runner{
|
||||
bin: bin, args: args, readyURL: readyURL,
|
||||
name: name, bin: bin, args: args, readyURL: readyURL,
|
||||
http: &http.Client{Timeout: 2 * time.Second},
|
||||
}
|
||||
}
|
||||
@@ -60,7 +64,7 @@ func (r *runner) isReady() bool {
|
||||
return r.ready
|
||||
}
|
||||
|
||||
// start launches llama-server. It returns as soon as the process exists, not
|
||||
// start launches the child. It returns as soon as the process exists, not
|
||||
// when the model is loaded.
|
||||
func (r *runner) start() error {
|
||||
r.mu.Lock()
|
||||
@@ -76,7 +80,7 @@ func (r *runner) start() error {
|
||||
return err
|
||||
}
|
||||
r.cmd, r.ready, r.yielding = cmd, false, false
|
||||
log.Printf("mavgpud: started llama-server pid=%d", cmd.Process.Pid)
|
||||
log.Printf("mavgpud: started %s pid=%d", r.name, cmd.Process.Pid)
|
||||
go func() {
|
||||
err := cmd.Wait()
|
||||
r.mu.Lock()
|
||||
@@ -84,15 +88,15 @@ func (r *runner) start() error {
|
||||
r.cmd, r.ready, r.yielding = nil, false, false
|
||||
r.mu.Unlock()
|
||||
if yielded {
|
||||
log.Printf("mavgpud: llama-server stopped, card yielded (%v)", err)
|
||||
log.Printf("mavgpud: %s stopped, card yielded (%v)", r.name, err)
|
||||
return
|
||||
}
|
||||
log.Printf("mavgpud: llama-server exited: %v", err)
|
||||
log.Printf("mavgpud: %s exited: %v", r.name, err)
|
||||
}()
|
||||
return nil
|
||||
}
|
||||
|
||||
// stop ends llama-server and waits for the VRAM to come back. SIGTERM first so
|
||||
// stop ends the child and waits for the VRAM to come back. SIGTERM first so
|
||||
// it unmaps cleanly, SIGKILL after the grace window. Returning before the
|
||||
// process is gone would let the supervisor report a free card while 14GB is
|
||||
// still mapped, which is the one lie that would make yielding useless.
|
||||
@@ -117,11 +121,11 @@ func (r *runner) stop(grace time.Duration) {
|
||||
}
|
||||
time.Sleep(100 * time.Millisecond)
|
||||
}
|
||||
log.Printf("mavgpud: llama-server did not exit in %s, killing", grace)
|
||||
log.Printf("mavgpud: %s did not exit in %s, killing", r.name, grace)
|
||||
_ = syscall.Kill(pgid, syscall.SIGKILL)
|
||||
}
|
||||
|
||||
// refreshReady asks llama-server whether the model is loaded. Called once per
|
||||
// refreshReady asks the child whether the model is loaded. Called once per
|
||||
// supervisor tick, never per request.
|
||||
func (r *runner) refreshReady(ctx context.Context) {
|
||||
if !r.running() {
|
||||
@@ -141,6 +145,6 @@ func (r *runner) refreshReady(ctx context.Context) {
|
||||
r.ready = ok
|
||||
r.mu.Unlock()
|
||||
if ok && !was {
|
||||
log.Printf("mavgpud: model ready")
|
||||
log.Printf("mavgpud: %s ready", r.name)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -25,7 +25,7 @@ func fakeServer(t *testing.T, body string) string {
|
||||
// status of a routine yield is identical to that of a real crash. Reading the
|
||||
// mavgpud log, the two were indistinguishable (Vikunja #491).
|
||||
func TestStopMarksTheExitAsAYield(t *testing.T) {
|
||||
r := newRunner(fakeServer(t, "while : ; do sleep 1 ; done"), nil, "")
|
||||
r := newRunner("fake", fakeServer(t, "while : ; do sleep 1 ; done"), nil, "")
|
||||
if err := r.start(); err != nil {
|
||||
t.Fatalf("start: %v", err)
|
||||
}
|
||||
@@ -49,7 +49,7 @@ func TestStopMarksTheExitAsAYield(t *testing.T) {
|
||||
// Stopping when nothing is running must not arm the flag for the next child.
|
||||
// The next exit after that would be a real crash logged as a yield.
|
||||
func TestStopWithNoChildDoesNotArmTheFlag(t *testing.T) {
|
||||
r := newRunner("/nonexistent", nil, "")
|
||||
r := newRunner("fake", "/nonexistent", nil, "")
|
||||
r.stop(10 * time.Millisecond)
|
||||
r.mu.Lock()
|
||||
defer r.mu.Unlock()
|
||||
|
||||
+70
-8
@@ -5,12 +5,24 @@
|
||||
// is detected sends it as a PushToTalk frame to the voice server. The reply
|
||||
// audio is played back through aplay(1).
|
||||
//
|
||||
// No wake-word model yet (MVP uses voice-activity-only trigger). The
|
||||
// SurfaceVoice auth layer caps all commands at L0 (no destructive acts),
|
||||
// making accidental triggers safe by design. A proper wake-word engine
|
||||
// (openWakeWord / Silero VAD ONNX) is the planned upgrade — the VAD shape
|
||||
// (30ms frames, 16kHz PCM) matches silero-vad's input interface exactly, so
|
||||
// swapping energy-threshold for ONNX-inference is a local change in vad.go.
|
||||
// Voice activity is silero-vad when -vad-model points at the graph, and an
|
||||
// energy threshold when it does not. Silero declines noise the threshold
|
||||
// accepts: 0 frames against 68 to 99 on the four fixtures, measured in
|
||||
// docs/evals/2026-08-09-silero-vad.md. Note that the model window is 512
|
||||
// samples and the capture frame is 480, so silero.go re-chunks. This comment
|
||||
// used to say the two matched, which was true of silero v4.
|
||||
//
|
||||
// The keyword is "Мэйвен" and it is required, when -wake-model points at the
|
||||
// head (V-487 stage two). Without it anything spoken near the microphone
|
||||
// becomes a turn, which the SurfaceVoice auth layer makes safe rather than
|
||||
// expensive: it caps all commands at L0, no destructive acts. It does not cap
|
||||
// reading, so an open gate still lets the room hear his facts read back.
|
||||
// wakeword.go holds the cadence and wakefeatures.go the three models.
|
||||
//
|
||||
// The conn carries both directions. mavwaked sends utterances and receives
|
||||
// proactive nudges on it, and it is opened at startup rather than at the first
|
||||
// utterance, because mavend registers a voice session on accept. See nudge.go
|
||||
// for why a nudge that is not heard is worse than one that is not delivered.
|
||||
//
|
||||
// While a reply is playing the capture side is muted (half-duplex): without
|
||||
// it, Maven's own voice comes back in through the mic and she answers
|
||||
@@ -52,6 +64,12 @@ const (
|
||||
defaultAddr = "127.0.0.1:9100"
|
||||
defaultLang = "ru"
|
||||
defaultReadSize = 4096 // max PCM bytes per read from arecord (fits multiple frames)
|
||||
|
||||
// defaultWakeWindowMs — how long the keyword stays good for. He says
|
||||
// "Мэйвен" and then a sentence, and the VAD does not close the utterance
|
||||
// until he stops, so this has to outlive the word by the length of what
|
||||
// follows it. It is spent on dispatch: one keyword, one turn.
|
||||
defaultWakeWindowMs = 8000
|
||||
)
|
||||
|
||||
func main() {
|
||||
@@ -73,17 +91,38 @@ func run(args []string) error {
|
||||
bargeIn := flag.Bool("barge-in", false, "cut Maven off when he talks over her (needs a room-tuned -barge-in-rms)")
|
||||
bargeRMS := flag.Int("barge-in-rms", defaultBargeRMS, "RMS x10000 a frame must clear to count as barge-in")
|
||||
bargeFrames := flag.Int("barge-in-frames", defaultBargeFrames, "consecutive frames over -barge-in-rms before playback is cut")
|
||||
vadModel := flag.String("vad-model", "", "silero-vad onnx file; empty runs the energy threshold instead")
|
||||
vadThreshold := flag.Float64("vad-threshold", defaultSileroThreshold, "speech probability a frame must clear")
|
||||
onnxLib := flag.String("onnx-lib", os.Getenv("MAVEN_ONNX_LIB"), "libonnxruntime.so, needed with -vad-model")
|
||||
wakeModel := flag.String("wake-model", "", "keyword head onnx; empty ships every utterance, as before V-487")
|
||||
wakeMel := flag.String("wake-mel", "", "melspectrogram.onnx, required with -wake-model")
|
||||
wakeEmbed := flag.String("wake-embed", "", "embedding_model.onnx, required with -wake-model")
|
||||
wakeThreshold := flag.Float64("wake-threshold", defaultWakeThreshold, "score the keyword must clear")
|
||||
wakeWindowMs := flag.Int("wake-window-ms", defaultWakeWindowMs, "ms an utterance may still start after the keyword")
|
||||
flag.CommandLine.Parse(args)
|
||||
|
||||
ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGINT, syscall.SIGTERM, syscall.SIGHUP)
|
||||
defer stop()
|
||||
|
||||
// Voice client — reused across utterances; SendRequest reconnects on error.
|
||||
// Voice client — one conn carrying both directions. SendRequest reconnects
|
||||
// on error, and the push receiver redials on its own clock.
|
||||
vc := voice.Dial(*addr)
|
||||
defer vc.Close()
|
||||
|
||||
// VAD engine.
|
||||
// VAD engine. A model that will not load is logged and not fatal: the
|
||||
// energy threshold is worse, and it is a great deal better than a
|
||||
// listening client that refuses to start.
|
||||
vad := NewVAD(*minRMS, *speechMs, *silenceMs, *maxMs)
|
||||
if *vadModel != "" {
|
||||
s, err := newSileroVAD(*vadModel, *onnxLib)
|
||||
if err != nil {
|
||||
log.Printf("mavwaked: silero unavailable, energy threshold unchanged: %v", err)
|
||||
} else {
|
||||
defer s.Close()
|
||||
vad.UseSilero(s, *vadThreshold)
|
||||
log.Printf("mavwaked: silero-vad from %s, threshold %.2f", *vadModel, *vadThreshold)
|
||||
}
|
||||
}
|
||||
|
||||
// Audio source.
|
||||
var src io.ReadCloser
|
||||
@@ -145,6 +184,29 @@ func run(args []string) error {
|
||||
}
|
||||
sess := newSession(vad, newAplayPlayer(), &voiceSender{vc: vc}, *lang, barge)
|
||||
|
||||
// Keyword gate. A model that will not load is logged and not fatal, for
|
||||
// the same reason silero's is not: an open gate is the daemon he had
|
||||
// yesterday, and a daemon that refuses to start is not.
|
||||
if *wakeModel != "" {
|
||||
w, err := newWakeWord(*wakeMel, *wakeEmbed, *wakeModel, *onnxLib, *wakeThreshold)
|
||||
if err != nil {
|
||||
log.Printf("mavwaked: wake word unavailable, every utterance is a turn: %v", err)
|
||||
} else {
|
||||
defer w.Close()
|
||||
sess.UseWakeWord(w, time.Duration(*wakeWindowMs)*time.Millisecond)
|
||||
log.Printf("mavwaked: wake word from %s, threshold %.3f, window %dms",
|
||||
*wakeModel, *wakeThreshold, *wakeWindowMs)
|
||||
}
|
||||
}
|
||||
|
||||
// Listen for nudges alongside capture. Connect eagerly so mavend has a
|
||||
// voice session before he has said anything: without one, a nudge routed
|
||||
// to voice finds nobody home and goes to the away channels instead.
|
||||
if err := vc.Connect(ctx); err != nil {
|
||||
log.Printf("mavwaked: voice server not reachable yet, retrying in background: %v", err)
|
||||
}
|
||||
go runNudgeReceiver(ctx, vc, sess)
|
||||
|
||||
return captureLoop(ctx, src, sess)
|
||||
}
|
||||
|
||||
|
||||
@@ -0,0 +1,79 @@
|
||||
package main
|
||||
|
||||
// The receiving half of the voice reach (V-671).
|
||||
//
|
||||
// mavwaked used to send and never listen. It wired no PushHandler, and
|
||||
// SendRequest discards a push frame when there is none. The consequence was
|
||||
// not a missing feature but a silent one: mavend routes a nudge to the voice
|
||||
// session that spoke most recently, and once mavwaked had spoken once it WAS
|
||||
// that session. PushToMostRecent succeeded, the dispatcher counted the nudge
|
||||
// delivered and stopped rerouting to telegram and ntfy, and mavwaked threw the
|
||||
// audio away. He heard nothing, anywhere.
|
||||
//
|
||||
// So the connection is opened at startup rather than at the first utterance,
|
||||
// and it is held open. A client that has never connected has no session, and
|
||||
// the dispatcher must be able to tell "he is not at the machine" from "he is,
|
||||
// and she has nothing to say".
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"log"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/voice"
|
||||
)
|
||||
|
||||
// nudgeRetry is how long to wait before dialling again after the conn ends.
|
||||
// mavend restarts on every deploy, and a listener that gives up then is a
|
||||
// listener that is deaf until the next reboot.
|
||||
const nudgeRetry = 5 * time.Second
|
||||
|
||||
// nudgeHandler decodes a push and hands the audio to the session, which
|
||||
// speaks it through the same player the reply path uses. It does not play
|
||||
// anything itself: the half-duplex gate and barge-in live on the capture
|
||||
// loop, and a nudge has to sit under both.
|
||||
type nudgeHandler struct{ sess *session }
|
||||
|
||||
func (h *nudgeHandler) OnPush(p voice.Push) {
|
||||
if p.Kind != voice.PushKindAudioNudge {
|
||||
log.Printf("mavwaked: ignoring push of unknown kind %q", p.Kind)
|
||||
return
|
||||
}
|
||||
var ap voice.AudioNudgePush
|
||||
if err := json.Unmarshal(p.Params, &ap); err != nil {
|
||||
log.Printf("mavwaked: nudge: decode: %v", err)
|
||||
return
|
||||
}
|
||||
log.Printf("mavwaked: nudge from rule %q (severity %d): %q (%.2fs audio)",
|
||||
ap.RuleName, ap.Severity, ap.Text, ap.Audio.Duration())
|
||||
if len(ap.Audio.Bytes) == 0 {
|
||||
// mavttsd was down or the text was empty. Say so rather than going
|
||||
// quiet: the dispatcher already counted this one as delivered.
|
||||
log.Printf("mavwaked: nudge %q carried no audio, nothing to speak", ap.RuleName)
|
||||
return
|
||||
}
|
||||
h.sess.Nudge(ap.Audio)
|
||||
}
|
||||
|
||||
// runNudgeReceiver keeps a push handler wired for as long as ctx lives,
|
||||
// redialling whenever the conn ends. Returns when ctx is cancelled.
|
||||
func runNudgeReceiver(ctx context.Context, vc *voice.Client, sess *session) {
|
||||
h := &nudgeHandler{sess: sess}
|
||||
for {
|
||||
err := vc.RunPushReceiver(ctx, h)
|
||||
if ctx.Err() != nil {
|
||||
return
|
||||
}
|
||||
if err != nil {
|
||||
log.Printf("mavwaked: nudge receiver: %v", err)
|
||||
} else {
|
||||
log.Printf("mavwaked: voice connection ended, reconnecting in %s", nudgeRetry)
|
||||
}
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
return
|
||||
case <-time.After(nudgeRetry):
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,170 @@
|
||||
package main
|
||||
|
||||
// The receiving half: a nudge pushed by mavend has to reach the speaker, and
|
||||
// it has to obey the same two gates a reply obeys (V-671).
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/audio"
|
||||
"github.com/kami/maven/internal/voice"
|
||||
)
|
||||
|
||||
func nudgeAudio() audio.Audio {
|
||||
return audio.Audio{Format: audio.PCM16kMono, Bytes: make([]byte, 8000)}
|
||||
}
|
||||
|
||||
// pushFrame builds the frame mavend's voicesink sends.
|
||||
func pushFrame(t *testing.T, a audio.Audio) voice.Push {
|
||||
t.Helper()
|
||||
body, err := json.Marshal(voice.AudioNudgePush{
|
||||
RuleName: "test-rule",
|
||||
Severity: 3,
|
||||
Audio: a,
|
||||
Text: "пора пить воду",
|
||||
Ts: time.Unix(0, 0),
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatalf("marshal push: %v", err)
|
||||
}
|
||||
return voice.Push{Kind: voice.PushKindAudioNudge, Params: body}
|
||||
}
|
||||
|
||||
// The defect itself: the push arrived and nothing came out of the speaker.
|
||||
func TestNudgeReachesThePlayer(t *testing.T) {
|
||||
sess, p, snd := newTestSession(bargeInConfig{})
|
||||
(&nudgeHandler{sess: sess}).OnPush(pushFrame(t, nudgeAudio()))
|
||||
|
||||
if p.plays != 0 {
|
||||
t.Fatal("nudge played from the push goroutine; it must wait for the capture loop")
|
||||
}
|
||||
if err := sess.feed(context.Background(), silentBytes()); err != nil {
|
||||
t.Fatalf("feed: %v", err)
|
||||
}
|
||||
if p.plays != 1 {
|
||||
t.Fatalf("plays = %d, want 1", p.plays)
|
||||
}
|
||||
if len(p.last.Bytes) != 8000 {
|
||||
t.Errorf("played %d bytes, want the nudge audio", len(p.last.Bytes))
|
||||
}
|
||||
if sess.nudges != 1 {
|
||||
t.Errorf("nudges = %d, want 1", sess.nudges)
|
||||
}
|
||||
if len(snd.sent) != 0 {
|
||||
t.Errorf("a nudge must not be shipped back to the daemon as an utterance")
|
||||
}
|
||||
}
|
||||
|
||||
// A push of some other kind, or one carrying no audio, must not reach the
|
||||
// player and must not wedge the one that follows.
|
||||
func TestNudgeIgnoresUnusablePushes(t *testing.T) {
|
||||
sess, p, _ := newTestSession(bargeInConfig{})
|
||||
h := &nudgeHandler{sess: sess}
|
||||
|
||||
h.OnPush(voice.Push{Kind: "something-else", Params: json.RawMessage(`{}`)})
|
||||
h.OnPush(voice.Push{Kind: voice.PushKindAudioNudge, Params: json.RawMessage(`not json`)})
|
||||
h.OnPush(pushFrame(t, audio.Audio{Format: audio.PCM16kMono}))
|
||||
|
||||
if err := sess.feed(context.Background(), silentBytes()); err != nil {
|
||||
t.Fatalf("feed: %v", err)
|
||||
}
|
||||
if p.plays != 0 {
|
||||
t.Fatalf("plays = %d, want 0", p.plays)
|
||||
}
|
||||
|
||||
h.OnPush(pushFrame(t, nudgeAudio()))
|
||||
if err := sess.feed(context.Background(), silentBytes()); err != nil {
|
||||
t.Fatalf("feed: %v", err)
|
||||
}
|
||||
if p.plays != 1 {
|
||||
t.Fatalf("plays after a usable nudge = %d, want 1", p.plays)
|
||||
}
|
||||
}
|
||||
|
||||
// The half-duplex gate covers a nudge exactly as it covers a reply: she does
|
||||
// not start one over herself, and the mic stays muted while it runs.
|
||||
func TestNudgeWaitsForTheReplyToFinish(t *testing.T) {
|
||||
sess, p, _ := newTestSession(bargeInConfig{})
|
||||
speakThenPause(t, sess)
|
||||
if !p.Playing() {
|
||||
t.Fatal("expected the reply to be playing")
|
||||
}
|
||||
plays := p.plays
|
||||
|
||||
(&nudgeHandler{sess: sess}).OnPush(pushFrame(t, nudgeAudio()))
|
||||
for i := 0; i < 20; i++ {
|
||||
if err := sess.feed(context.Background(), silentBytes()); err != nil {
|
||||
t.Fatalf("feed: %v", err)
|
||||
}
|
||||
}
|
||||
if p.plays != plays {
|
||||
t.Fatalf("nudge cut across the reply: plays = %d, want %d", p.plays, plays)
|
||||
}
|
||||
|
||||
p.playing = false
|
||||
if err := sess.feed(context.Background(), silentBytes()); err != nil {
|
||||
t.Fatalf("feed: %v", err)
|
||||
}
|
||||
if p.plays != plays+1 {
|
||||
t.Fatalf("nudge never played after the reply ended: plays = %d", p.plays)
|
||||
}
|
||||
}
|
||||
|
||||
// Speaking a nudge must not leave half a sentence in the VAD. The frames
|
||||
// captured before it are pre-nudge speech, and splicing them onto whatever he
|
||||
// says afterwards ships one utterance that is two.
|
||||
func TestNudgeResetsTheVAD(t *testing.T) {
|
||||
sess, p, snd := newTestSession(bargeInConfig{})
|
||||
loud := frameAt(0.35)
|
||||
speechFrames := (defaultSpeechMs + defaultFrameMs - 1) / defaultFrameMs
|
||||
for i := 0; i < speechFrames+5; i++ {
|
||||
if err := sess.feed(context.Background(), loud); err != nil {
|
||||
t.Fatalf("feed: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
(&nudgeHandler{sess: sess}).OnPush(pushFrame(t, nudgeAudio()))
|
||||
if err := sess.feed(context.Background(), silentBytes()); err != nil {
|
||||
t.Fatalf("feed: %v", err)
|
||||
}
|
||||
if p.plays != 1 {
|
||||
t.Fatalf("nudge did not play: plays = %d", p.plays)
|
||||
}
|
||||
|
||||
// Playback ends, silence follows. The half-formed utterance must be gone
|
||||
// rather than closing on the first quiet frame.
|
||||
p.playing = false
|
||||
silenceFrames := (defaultSilenceMs+defaultFrameMs-1)/defaultFrameMs + 2
|
||||
for i := 0; i < silenceFrames; i++ {
|
||||
if err := sess.feed(context.Background(), silentBytes()); err != nil {
|
||||
t.Fatalf("feed: %v", err)
|
||||
}
|
||||
}
|
||||
if len(snd.sent) != 0 {
|
||||
t.Fatalf("sent %d utterances after a nudge, want 0", len(snd.sent))
|
||||
}
|
||||
}
|
||||
|
||||
// Two nudges queued back to back: the newer one is what he hears. The
|
||||
// PushHandler contract in internal/voice says the next nudge replaces the
|
||||
// stale one rather than dogpiling on it.
|
||||
func TestNudgeReplacesAnUnspokenOne(t *testing.T) {
|
||||
sess, p, _ := newTestSession(bargeInConfig{})
|
||||
h := &nudgeHandler{sess: sess}
|
||||
|
||||
h.OnPush(pushFrame(t, audio.Audio{Format: audio.PCM16kMono, Bytes: make([]byte, 4000)}))
|
||||
h.OnPush(pushFrame(t, audio.Audio{Format: audio.PCM16kMono, Bytes: make([]byte, 12000)}))
|
||||
|
||||
if err := sess.feed(context.Background(), silentBytes()); err != nil {
|
||||
t.Fatalf("feed: %v", err)
|
||||
}
|
||||
if p.plays != 1 {
|
||||
t.Fatalf("plays = %d, want 1", p.plays)
|
||||
}
|
||||
if len(p.last.Bytes) != 12000 {
|
||||
t.Errorf("played %d bytes, want the newer nudge", len(p.last.Bytes))
|
||||
}
|
||||
}
|
||||
+136
-1
@@ -7,6 +7,7 @@ package main
|
||||
import (
|
||||
"context"
|
||||
"log"
|
||||
"sync"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/audio"
|
||||
@@ -19,6 +20,15 @@ type utteranceSender interface {
|
||||
Send(ctx context.Context, utt audio.Audio, lang string) (audio.Audio, error)
|
||||
}
|
||||
|
||||
// keywordGate answers whether the keyword has just been spoken. The
|
||||
// production one is wakeWord; tests substitute a recorder, because a gate that
|
||||
// can only be exercised with three ONNX files is a gate nobody tests.
|
||||
type keywordGate interface {
|
||||
Feed(frame []int16) bool
|
||||
Reset()
|
||||
Score() float64
|
||||
}
|
||||
|
||||
// bargeInConfig holds the two numbers barge-in needs. Zero Frames disables
|
||||
// barge-in entirely — the half-duplex gate still runs.
|
||||
type bargeInConfig struct {
|
||||
@@ -62,11 +72,29 @@ type session struct {
|
||||
// whenever playback ends.
|
||||
loudFrames int
|
||||
|
||||
// wake is the keyword gate, or nil when no model was loaded. wakeUntil is
|
||||
// how long a keyword stays good for: he says "Мэйвен" and then a sentence,
|
||||
// and the VAD does not close the utterance until he stops, so the window
|
||||
// has to outlive the word by the length of what follows it.
|
||||
wake keywordGate
|
||||
wakeWindow time.Duration
|
||||
wakeUntil time.Time
|
||||
|
||||
// pending holds a nudge the push receiver handed over, waiting for the
|
||||
// capture loop to speak it. It is the one field written from another
|
||||
// goroutine, hence the mutex; everything else in this struct belongs to
|
||||
// the capture loop alone.
|
||||
nudgeMu sync.Mutex
|
||||
pending *audio.Audio
|
||||
|
||||
// counters, read by tests and logged on the way out.
|
||||
suppressed int // frames dropped because she was speaking
|
||||
dropped int // frames dropped as round-trip backlog
|
||||
bargeIns int // times playback was cut because he spoke over her
|
||||
sent int // utterances shipped to the daemon
|
||||
nudges int // proactive pushes spoken through the speaker
|
||||
wakes int // times the keyword opened the gate
|
||||
ignored int // complete utterances dropped because the keyword was absent
|
||||
|
||||
// loudSum and loudSeen accumulate the energy of suppressed frames, so
|
||||
// the operator can read what the room actually measures and set
|
||||
@@ -79,6 +107,12 @@ func newSession(vad *VAD, p player, s utteranceSender, lang string, barge bargeI
|
||||
return &session{vad: vad, player: p, sender: s, lang: lang, barge: barge, now: time.Now}
|
||||
}
|
||||
|
||||
// UseWakeWord puts the keyword gate in front of dispatch. Without it every
|
||||
// utterance is shipped, which is what mavwaked did before V-487 stage two.
|
||||
func (s *session) UseWakeWord(w keywordGate, window time.Duration) {
|
||||
s.wake, s.wakeWindow = w, window
|
||||
}
|
||||
|
||||
// frameDuration is the wall time one captured frame represents.
|
||||
const frameDuration = defaultFrameMs * time.Millisecond
|
||||
|
||||
@@ -138,6 +172,7 @@ func (s *session) feed(ctx context.Context, frame []byte) error {
|
||||
s.bargeIns++
|
||||
s.loudFrames = 0
|
||||
s.vad.Reset()
|
||||
s.resetWake()
|
||||
log.Printf("mavwaked: barge-in — stopped playback")
|
||||
s.replayRecent()
|
||||
return nil
|
||||
@@ -148,15 +183,101 @@ func (s *session) feed(ctx context.Context, frame []byte) error {
|
||||
if s.loudFrames != 0 {
|
||||
s.loudFrames = 0
|
||||
s.vad.Reset()
|
||||
// The wake word saw nothing during playback, so what it holds is from
|
||||
// before she spoke. Judging what he says next on it would score a
|
||||
// sentence that ended a reply ago.
|
||||
s.resetWake()
|
||||
}
|
||||
|
||||
utt, state := s.vad.Feed(PCMToI16(frame))
|
||||
if s.startPendingNudge() {
|
||||
return nil
|
||||
}
|
||||
|
||||
// The keyword is scored on the same frames the VAD sees, and only on the
|
||||
// ones that reach here: every path above returns while she is speaking, so
|
||||
// her own voice saying "Мэйвен" cannot wake her.
|
||||
pcm := PCMToI16(frame)
|
||||
if s.wake != nil && s.wake.Feed(pcm) {
|
||||
s.wakes++
|
||||
s.wakeUntil = s.now().Add(s.wakeWindow)
|
||||
log.Printf("mavwaked: keyword heard (score %.3f), listening for %s",
|
||||
s.wake.Score(), s.wakeWindow)
|
||||
}
|
||||
|
||||
utt, state := s.vad.Feed(pcm)
|
||||
if state == StateSpeech || utt.Bytes == nil {
|
||||
return nil
|
||||
}
|
||||
return s.dispatch(ctx, utt)
|
||||
}
|
||||
|
||||
// Nudge hands proactive audio to the session, to be spoken as soon as the
|
||||
// capture loop finds a quiet moment. Safe to call from the push receiver
|
||||
// goroutine; nothing else here is.
|
||||
//
|
||||
// A nudge arriving while one is already waiting REPLACES it. That is the
|
||||
// contract internal/voice states for PushHandler: the next nudge replaces the
|
||||
// stale one in his attention rather than dogpiling on it.
|
||||
func (s *session) Nudge(a audio.Audio) {
|
||||
if len(a.Bytes) == 0 {
|
||||
return
|
||||
}
|
||||
s.nudgeMu.Lock()
|
||||
if s.pending != nil {
|
||||
log.Printf("mavwaked: nudge replaced one still waiting to be spoken")
|
||||
}
|
||||
s.pending = &a
|
||||
s.nudgeMu.Unlock()
|
||||
}
|
||||
|
||||
// takeNudge removes and returns the waiting nudge, or nil.
|
||||
func (s *session) takeNudge() *audio.Audio {
|
||||
s.nudgeMu.Lock()
|
||||
defer s.nudgeMu.Unlock()
|
||||
a := s.pending
|
||||
s.pending = nil
|
||||
return a
|
||||
}
|
||||
|
||||
// startPendingNudge speaks a waiting nudge and reports whether it started
|
||||
// one. It runs on the capture loop, past the half-duplex gate, so a nudge
|
||||
// never cuts across a reply and never plays into a backlog drain.
|
||||
//
|
||||
// The VAD is reset first. Playback is about to suppress every frame until it
|
||||
// ends, and a half-heard sentence left in the VAD would splice onto whatever
|
||||
// he says afterwards. Barge-in needs no special case: it reads the player,
|
||||
// and the player does not care which audio it is playing.
|
||||
func (s *session) startPendingNudge() bool {
|
||||
a := s.takeNudge()
|
||||
if a == nil {
|
||||
return false
|
||||
}
|
||||
s.vad.Reset()
|
||||
s.nudges++
|
||||
log.Printf("mavwaked: speaking nudge (%.2fs audio)", a.Duration())
|
||||
s.player.Play(*a)
|
||||
return true
|
||||
}
|
||||
|
||||
// awake reports whether an utterance ending now was addressed to her.
|
||||
//
|
||||
// With no wake word loaded every utterance is, which is exactly what mavwaked
|
||||
// did before this gate existed. An operator with no model file gets the old
|
||||
// daemon rather than a daemon that refuses to hear anything.
|
||||
func (s *session) awake() bool {
|
||||
if s.wake == nil {
|
||||
return true
|
||||
}
|
||||
return s.now().Before(s.wakeUntil)
|
||||
}
|
||||
|
||||
// resetWake drops the gate's streaming state when there is a gate.
|
||||
func (s *session) resetWake() {
|
||||
if s.wake != nil {
|
||||
s.wake.Reset()
|
||||
}
|
||||
}
|
||||
|
||||
// keepRecent stores a copy of one barge-in trigger frame, keeping at most
|
||||
// barge.Frames of them.
|
||||
func (s *session) keepRecent(frame []byte) {
|
||||
@@ -197,6 +318,19 @@ func (s *session) replayRecent() {
|
||||
// whole backlog straight into the VAD, and a Send error did the same on every
|
||||
// failed turn, so a dead socket drove a retry loop off nothing but backlog.
|
||||
func (s *session) dispatch(ctx context.Context, utt audio.Audio) error {
|
||||
if !s.awake() {
|
||||
s.ignored++
|
||||
log.Printf("mavwaked: utterance ignored, keyword not heard (%.2fs, %d ignored so far)",
|
||||
utt.Duration(), s.ignored)
|
||||
s.vad.Reset()
|
||||
s.resetWake()
|
||||
return nil
|
||||
}
|
||||
// One keyword, one turn. A window that renewed itself on every reply would
|
||||
// leave the microphone open for as long as he kept talking, which is the
|
||||
// state this gate exists to end.
|
||||
s.wakeUntil = time.Time{}
|
||||
|
||||
log.Printf("mavwaked: utterance complete (%.2fs, %d bytes), sending...", utt.Duration(), len(utt.Bytes))
|
||||
start := s.now()
|
||||
reply, err := s.sender.Send(ctx, utt, s.lang)
|
||||
@@ -225,6 +359,7 @@ func (s *session) dispatch(ctx context.Context, utt audio.Audio) error {
|
||||
// recorded before she started speaking.
|
||||
func (s *session) dropBacklog(start time.Time) {
|
||||
s.vad.Reset()
|
||||
s.resetWake()
|
||||
s.loudFrames = 0
|
||||
s.recent = s.recent[:0]
|
||||
if elapsed := s.now().Sub(start); elapsed > 0 {
|
||||
|
||||
@@ -0,0 +1,171 @@
|
||||
package main
|
||||
|
||||
// silero-vad, the speech detector that replaces the energy threshold (V-487).
|
||||
//
|
||||
// Why an energy threshold is not a voice activity detector. It answers "is
|
||||
// this frame loud", and a fan, a door and a television are all loud. mavwaked
|
||||
// sends every utterance it accepts to speech-to-text and then to the daemon,
|
||||
// so a false trigger is a turn Maven takes on something nobody said to her.
|
||||
// Silero answers "is this frame speech", which is the question.
|
||||
//
|
||||
// It is 2.3MB of ONNX and runs on one CPU core in real time. That is not an
|
||||
// aside: this is the one model in the system that may never be offloaded or
|
||||
// gated on GPU admission, because a wake path that waits on a card is not a
|
||||
// wake path.
|
||||
//
|
||||
// Nil is a working value. Without -vad-model the daemon runs the energy VAD
|
||||
// exactly as it did before this file existed.
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"sync"
|
||||
|
||||
ort "github.com/yalue/onnxruntime_go"
|
||||
)
|
||||
|
||||
const (
|
||||
// sileroWindow — samples per inference at 16kHz. The model is fixed at
|
||||
// 512 and does not accept another size, which is why this file
|
||||
// re-chunks rather than reusing the 480-sample capture frame. main.go
|
||||
// used to claim the two matched; that was true of silero v4.
|
||||
sileroWindow = 512
|
||||
|
||||
// sileroContext — samples of the previous window prepended to each
|
||||
// inference, as the reference implementation does. Without it the first
|
||||
// milliseconds of every window are judged with no history and speech
|
||||
// onsets score low.
|
||||
sileroContext = 64
|
||||
|
||||
// sileroState — the LSTM state carried between windows, [2][1][128].
|
||||
sileroStateDim = 128
|
||||
|
||||
// defaultSileroThreshold — probability above which a window is speech.
|
||||
// 0.5 is the reference default. Raising it costs speech onsets, which
|
||||
// are the quietest part of an utterance.
|
||||
defaultSileroThreshold = 0.5
|
||||
)
|
||||
|
||||
// sileroVAD holds one ONNX session and the streaming state around it. It is
|
||||
// fed 30ms capture frames and answers per frame, buffering across calls
|
||||
// because 480 samples never line up with a 512-sample window.
|
||||
type sileroVAD struct {
|
||||
mu sync.Mutex
|
||||
session *ort.DynamicAdvancedSession
|
||||
|
||||
pending []float32 // samples not yet part of a full window
|
||||
context [sileroContext]float32 // tail of the previous window
|
||||
state []float32 // [2][1][128], carried between windows
|
||||
last float64 // most recent probability, held between windows
|
||||
sr []int64
|
||||
}
|
||||
|
||||
// newSileroVAD loads the graph. The ONNX environment is initialised here when
|
||||
// nothing else has done it, because mavwaked has no embedder to do it first.
|
||||
func newSileroVAD(modelPath, libPath string) (*sileroVAD, error) {
|
||||
if !ort.IsInitialized() {
|
||||
if libPath != "" {
|
||||
ort.SetSharedLibraryPath(libPath)
|
||||
}
|
||||
if err := ort.InitializeEnvironment(); err != nil {
|
||||
return nil, fmt.Errorf("silero: onnx runtime: %w", err)
|
||||
}
|
||||
}
|
||||
s, err := ort.NewDynamicAdvancedSession(modelPath,
|
||||
[]string{"input", "state", "sr"}, []string{"output", "stateN"}, nil)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("silero: load %s: %w", modelPath, err)
|
||||
}
|
||||
return &sileroVAD{
|
||||
session: s,
|
||||
state: make([]float32, 2*sileroStateDim),
|
||||
sr: []int64{16000},
|
||||
}, nil
|
||||
}
|
||||
|
||||
// Speech reports whether the frame carries speech, and the probability behind
|
||||
// that answer. A frame that completes no window inherits the previous
|
||||
// probability, so the caller sees one answer per frame either way.
|
||||
func (s *sileroVAD) Speech(frame []int16, threshold float64) (bool, float64) {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
|
||||
for _, v := range frame {
|
||||
s.pending = append(s.pending, float32(v)/32768.0)
|
||||
}
|
||||
for len(s.pending) >= sileroWindow {
|
||||
p, err := s.infer(s.pending[:sileroWindow])
|
||||
if err != nil {
|
||||
// A failed inference must not silence the microphone. Hold the
|
||||
// last answer and let the next window try again.
|
||||
break
|
||||
}
|
||||
s.last = p
|
||||
s.pending = s.pending[sileroWindow:]
|
||||
}
|
||||
return s.last >= threshold, s.last
|
||||
}
|
||||
|
||||
// infer runs one window and rolls the state and the context forward.
|
||||
func (s *sileroVAD) infer(window []float32) (float64, error) {
|
||||
in := make([]float32, sileroContext+sileroWindow)
|
||||
copy(in, s.context[:])
|
||||
copy(in[sileroContext:], window)
|
||||
|
||||
inT, err := ort.NewTensor(ort.NewShape(1, int64(len(in))), in)
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
defer inT.Destroy()
|
||||
stT, err := ort.NewTensor(ort.NewShape(2, 1, sileroStateDim), s.state)
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
defer stT.Destroy()
|
||||
srT, err := ort.NewTensor(ort.NewShape(1), s.sr)
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
defer srT.Destroy()
|
||||
|
||||
out, err := ort.NewEmptyTensor[float32](ort.NewShape(1, 1))
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
defer out.Destroy()
|
||||
next, err := ort.NewEmptyTensor[float32](ort.NewShape(2, 1, sileroStateDim))
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
defer next.Destroy()
|
||||
|
||||
if err := s.session.Run(
|
||||
[]ort.Value{inT, stT, srT},
|
||||
[]ort.Value{out, next},
|
||||
); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
copy(s.state, next.GetData())
|
||||
copy(s.context[:], in[len(in)-sileroContext:])
|
||||
return float64(out.GetData()[0]), nil
|
||||
}
|
||||
|
||||
// Reset drops the streaming state. Called at every utterance boundary and
|
||||
// after barge-in, so echo-era history never scores the next sentence.
|
||||
func (s *sileroVAD) Reset() {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
s.pending = s.pending[:0]
|
||||
s.context = [sileroContext]float32{}
|
||||
for i := range s.state {
|
||||
s.state[i] = 0
|
||||
}
|
||||
s.last = 0
|
||||
}
|
||||
|
||||
// Close releases the session.
|
||||
func (s *sileroVAD) Close() error {
|
||||
if s == nil || s.session == nil {
|
||||
return nil
|
||||
}
|
||||
return s.session.Destroy()
|
||||
}
|
||||
@@ -0,0 +1,149 @@
|
||||
package main
|
||||
|
||||
// What this measures. The energy threshold cannot tell a voice from a
|
||||
// television, and every utterance it accepts becomes a turn. So the test that
|
||||
// matters is not "does silero find speech" — it is "does it decline what the
|
||||
// energy threshold accepts".
|
||||
//
|
||||
// Speech is the four piper fixtures mavsttd already scores against. They are
|
||||
// synthesised, so nothing of the owner's voice is committed. Non-speech is
|
||||
// white noise at the same loudness, which is the cheapest thing that fools an
|
||||
// energy floor and the honest floor for this claim.
|
||||
//
|
||||
// Both halves skip without models/vad/silero_vad.onnx and MAVEN_ONNX_LIB,
|
||||
// like the TestONNX measurements in internal/router/eval.
|
||||
|
||||
import (
|
||||
"math"
|
||||
"math/rand"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
)
|
||||
|
||||
const wavHeader = 44 // 16kHz mono s16le, written by piper
|
||||
|
||||
func loadSilero(t *testing.T) *sileroVAD {
|
||||
t.Helper()
|
||||
model := filepath.Join("..", "..", "models", "vad", "silero_vad.onnx")
|
||||
lib := os.Getenv("MAVEN_ONNX_LIB")
|
||||
if _, err := os.Stat(model); err != nil {
|
||||
t.Skipf("missing %s: %v", model, err)
|
||||
}
|
||||
if lib == "" {
|
||||
t.Skip("MAVEN_ONNX_LIB unset")
|
||||
}
|
||||
s, err := newSileroVAD(model, lib)
|
||||
if err != nil {
|
||||
t.Skipf("silero unavailable: %v", err)
|
||||
}
|
||||
return s
|
||||
}
|
||||
|
||||
// feedAll runs a whole clip through a VAD and reports how many utterances it
|
||||
// produced and how many frames it called speech.
|
||||
func feedAll(v *VAD, pcm []int16) (utterances, speechFrames int) {
|
||||
for i := 0; i+frameSamples <= len(pcm); i += frameSamples {
|
||||
frame := pcm[i : i+frameSamples]
|
||||
utt, state := v.Feed(frame)
|
||||
if state == StateSpeech {
|
||||
speechFrames++
|
||||
}
|
||||
if utt.Bytes != nil {
|
||||
utterances++
|
||||
}
|
||||
}
|
||||
return utterances, speechFrames
|
||||
}
|
||||
|
||||
func readFixture(t *testing.T, name string) []int16 {
|
||||
t.Helper()
|
||||
raw, err := os.ReadFile(filepath.Join("..", "mavsttd", "testdata", name))
|
||||
if err != nil {
|
||||
t.Skipf("missing fixture %s: %v", name, err)
|
||||
}
|
||||
if len(raw) <= wavHeader {
|
||||
t.Fatalf("%s: %d bytes, no audio", name, len(raw))
|
||||
}
|
||||
return PCMToI16(raw[wavHeader:])
|
||||
}
|
||||
|
||||
// noise returns white noise scaled to the same RMS as ref. Same loudness,
|
||||
// nothing said.
|
||||
func noise(ref []int16, seed int64) []int16 {
|
||||
target := frameRMS(ref)
|
||||
r := rand.New(rand.NewSource(seed))
|
||||
out := make([]int16, len(ref))
|
||||
for i := range out {
|
||||
out[i] = int16(r.NormFloat64() * target * 32768.0)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
func TestSileroHearsSpeechAndDeclinesNoise(t *testing.T) {
|
||||
s := loadSilero(t)
|
||||
defer s.Close()
|
||||
|
||||
for _, name := range []string{"ru_fact.wav", "ru_query.wav", "ru_reminder.wav", "en_act.wav"} {
|
||||
pcm := readFixture(t, name)
|
||||
|
||||
v := NewVAD(0, 0, 0, 0)
|
||||
v.UseSilero(s, defaultSileroThreshold)
|
||||
_, spoke := feedAll(v, pcm)
|
||||
if spoke == 0 {
|
||||
t.Errorf("%s: silero heard no speech in a spoken clip", name)
|
||||
}
|
||||
|
||||
s.Reset()
|
||||
v2 := NewVAD(0, 0, 0, 0)
|
||||
v2.UseSilero(s, defaultSileroThreshold)
|
||||
_, heard := feedAll(v2, noise(pcm, 7))
|
||||
|
||||
energy := NewVAD(0, 0, 0, 0)
|
||||
_, energyHeard := feedAll(energy, noise(pcm, 7))
|
||||
|
||||
t.Logf("%s: speech frames — silero on speech %d, silero on noise %d, energy on noise %d",
|
||||
name, spoke, heard, energyHeard)
|
||||
if heard >= energyHeard {
|
||||
t.Errorf("%s: silero called %d noise frames speech, energy called %d — no improvement",
|
||||
name, heard, energyHeard)
|
||||
}
|
||||
s.Reset()
|
||||
}
|
||||
}
|
||||
|
||||
// BenchmarkSileroFrame answers the only performance question that matters
|
||||
// here: one 30ms frame must cost far less than 30ms on one core, or the
|
||||
// detector cannot run always-on beside everything else on that machine.
|
||||
func BenchmarkSileroFrame(b *testing.B) {
|
||||
s := loadSilero(&testing.T{})
|
||||
if s == nil {
|
||||
b.Skip("silero unavailable")
|
||||
}
|
||||
defer s.Close()
|
||||
frame := make([]int16, frameSamples)
|
||||
for i := range frame {
|
||||
frame[i] = int16(i%400 - 200)
|
||||
}
|
||||
for i := 0; i < b.N; i++ {
|
||||
s.Speech(frame, defaultSileroThreshold)
|
||||
}
|
||||
}
|
||||
|
||||
// TestSileroRechunksAcrossFrames pins the reason this file exists. The capture
|
||||
// frame is 480 samples and the model window is 512, so a detector that ran one
|
||||
// inference per frame would be feeding the model a shape it does not accept.
|
||||
func TestSileroRechunksAcrossFrames(t *testing.T) {
|
||||
s := loadSilero(t)
|
||||
defer s.Close()
|
||||
|
||||
silence := make([]int16, frameSamples)
|
||||
for i := 0; i < 20; i++ {
|
||||
if _, p := s.Speech(silence, defaultSileroThreshold); math.IsNaN(p) {
|
||||
t.Fatalf("frame %d: probability is NaN", i)
|
||||
}
|
||||
}
|
||||
if len(s.pending) >= sileroWindow {
|
||||
t.Errorf("pending grew to %d samples, so windows are not being consumed", len(s.pending))
|
||||
}
|
||||
}
|
||||
+35
-1
@@ -68,6 +68,37 @@ type VAD struct {
|
||||
// follows the room's ambient level. Initialised to minRMS; updated
|
||||
// on each silence frame.
|
||||
floorRMS float64
|
||||
|
||||
// speech is silero-vad, or nil. When it is set the energy floor decides
|
||||
// nothing: the question becomes "is this speech" rather than "is this
|
||||
// loud", and the noise floor is not even tracked. Everything after that
|
||||
// answer — the speech hold, the silence hold, the length cap, the
|
||||
// buffer — is the same state machine either way, which is why the
|
||||
// detector goes here and not around this type.
|
||||
speech *sileroVAD
|
||||
speechMin float64
|
||||
}
|
||||
|
||||
// UseSilero swaps the energy threshold for the model. Passing nil is a
|
||||
// no-op, so a caller that could not load the graph keeps a working VAD.
|
||||
func (v *VAD) UseSilero(s *sileroVAD, threshold float64) {
|
||||
if s == nil {
|
||||
return
|
||||
}
|
||||
if threshold <= 0 {
|
||||
threshold = defaultSileroThreshold
|
||||
}
|
||||
v.speech = s
|
||||
v.speechMin = threshold
|
||||
}
|
||||
|
||||
// isSpeech answers the one question the state machine asks of a frame.
|
||||
func (v *VAD) isSpeech(frame []int16, rms float64) bool {
|
||||
if v.speech != nil {
|
||||
ok, _ := v.speech.Speech(frame, v.speechMin)
|
||||
return ok
|
||||
}
|
||||
return rms >= v.floorRMS
|
||||
}
|
||||
|
||||
// NewVAD creates a VAD with the given thresholds. Zero values use defaults.
|
||||
@@ -110,7 +141,7 @@ func (v *VAD) State() SpeechState { return v.state }
|
||||
// should send the audio to the voice server before feeding more frames.
|
||||
func (v *VAD) Feed(frame []int16) (_ audio.Audio, state SpeechState) {
|
||||
rms := frameRMS(frame)
|
||||
isSpeech := rms >= v.floorRMS
|
||||
isSpeech := v.isSpeech(frame, rms)
|
||||
|
||||
switch v.state {
|
||||
case StateSilence:
|
||||
@@ -175,6 +206,9 @@ func (v *VAD) Feed(frame []int16) (_ audio.Audio, state SpeechState) {
|
||||
func (v *VAD) Reset() { v.reset() }
|
||||
|
||||
func (v *VAD) reset() {
|
||||
if v.speech != nil {
|
||||
v.speech.Reset()
|
||||
}
|
||||
v.state = StateSilence
|
||||
v.speechFrames = 0
|
||||
v.silenceFrames = 0
|
||||
|
||||
@@ -0,0 +1,187 @@
|
||||
package main
|
||||
|
||||
// The three models behind the wake word (V-487 stage two).
|
||||
//
|
||||
// openWakeWord's pipeline, run in a row:
|
||||
//
|
||||
// audio -> melspectrogram.onnx -> 32-bin mel frames, one per 10ms
|
||||
// 76 frames -> embedding_model.onnx -> one 96-dim embedding per 80ms
|
||||
// 16 embeds -> maven_wakeword.onnx -> one score
|
||||
//
|
||||
// The first two are frozen and pretrained. Only the last was trained here,
|
||||
// which is why it is 100KB and the other two are megabytes. The shapes are
|
||||
// not guesses: 2.0s of 16kHz audio measures 197 mel frames, and 76-frame
|
||||
// windows at stride 8 give exactly the 16 embeddings the head was fitted on.
|
||||
//
|
||||
// This file knows ONNX and nothing about the 80ms cadence. wakeword.go knows
|
||||
// the cadence and nothing about tensors.
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
|
||||
ort "github.com/yalue/onnxruntime_go"
|
||||
)
|
||||
|
||||
const (
|
||||
// melHop — samples per mel frame. 10ms at 16kHz.
|
||||
melHop = 160
|
||||
// melBins — mel bins per frame, fixed by melspectrogram.onnx.
|
||||
melBins = 32
|
||||
// embedFrames — mel frames one embedding is computed over, 760ms.
|
||||
embedFrames = 76
|
||||
// embedStride — mel frames between embeddings, 80ms.
|
||||
embedStride = 8
|
||||
// embedDim — the embedding width.
|
||||
embedDim = 96
|
||||
// headWindow — embeddings the head scores at once, 1.28s of audio.
|
||||
headWindow = 16
|
||||
|
||||
// melContext — samples of history prepended to each incremental mel
|
||||
// call, chosen so the eight frames this call yields continue exactly
|
||||
// where the previous call's eight stopped.
|
||||
//
|
||||
// melspectrogram.onnx returns N/160-3 frames for N samples, and frame i
|
||||
// covers [i*160, i*160+400). With 480 samples of history the buffer is
|
||||
// 1760 samples, which is 8 frames, and the oldest of them starts one hop
|
||||
// after the newest of the previous call. Less history leaves a gap: the
|
||||
// first frames of a bare chunk would be computed against silence.
|
||||
melContext = 480
|
||||
|
||||
// chunkSamples — audio per embedding step, 80ms.
|
||||
chunkSamples = embedStride * melHop
|
||||
)
|
||||
|
||||
// wakeModels holds the three ONNX sessions. It runs on CPU threads beside
|
||||
// silero and never touches the GPU. That is a rule, not a result: a wake word
|
||||
// that waits on card admission is not a wake word.
|
||||
type wakeModels struct {
|
||||
mel *ort.DynamicAdvancedSession
|
||||
emb *ort.DynamicAdvancedSession
|
||||
head *ort.DynamicAdvancedSession
|
||||
}
|
||||
|
||||
// newWakeModels loads all three. melPath and embedPath are openWakeWord's
|
||||
// frozen feature models; headPath is the keyword head trained for "Мэйвен".
|
||||
func newWakeModels(melPath, embedPath, headPath, libPath string) (*wakeModels, error) {
|
||||
if !ort.IsInitialized() {
|
||||
if libPath != "" {
|
||||
ort.SetSharedLibraryPath(libPath)
|
||||
}
|
||||
if err := ort.InitializeEnvironment(); err != nil {
|
||||
return nil, fmt.Errorf("wake word: onnx runtime: %w", err)
|
||||
}
|
||||
}
|
||||
open := func(p string, in, out []string) (*ort.DynamicAdvancedSession, error) {
|
||||
s, err := ort.NewDynamicAdvancedSession(p, in, out, nil)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("wake word: load %s: %w", p, err)
|
||||
}
|
||||
return s, nil
|
||||
}
|
||||
m := &wakeModels{}
|
||||
var err error
|
||||
if m.mel, err = open(melPath, []string{"input"}, []string{"output"}); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
if m.emb, err = open(embedPath, []string{"input_1"}, []string{"conv2d_19"}); err != nil {
|
||||
m.Close()
|
||||
return nil, err
|
||||
}
|
||||
if m.head, err = open(headPath, []string{"embeddings"}, []string{"score"}); err != nil {
|
||||
m.Close()
|
||||
return nil, err
|
||||
}
|
||||
return m, nil
|
||||
}
|
||||
|
||||
// Close releases the three sessions.
|
||||
func (m *wakeModels) Close() {
|
||||
if m == nil {
|
||||
return
|
||||
}
|
||||
for _, s := range []*ort.DynamicAdvancedSession{m.mel, m.emb, m.head} {
|
||||
if s != nil {
|
||||
s.Destroy()
|
||||
}
|
||||
}
|
||||
m.mel, m.emb, m.head = nil, nil, nil
|
||||
}
|
||||
|
||||
// melFrames runs one buffer of samples and returns the mel frames it yielded.
|
||||
func (m *wakeModels) melFrames(buf []float32) ([][melBins]float32, error) {
|
||||
in, err := ort.NewTensor(ort.NewShape(1, int64(len(buf))), buf)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
defer in.Destroy()
|
||||
|
||||
n := int64(len(buf)/melHop - 3)
|
||||
if n < 1 {
|
||||
return nil, fmt.Errorf("wake word: %d samples yield no mel frames", len(buf))
|
||||
}
|
||||
out, err := ort.NewEmptyTensor[float32](ort.NewShape(1, 1, n, melBins))
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
defer out.Destroy()
|
||||
|
||||
if err := m.mel.Run([]ort.Value{in}, []ort.Value{out}); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
data := out.GetData()
|
||||
frames := make([][melBins]float32, n)
|
||||
for i := range frames {
|
||||
for j := 0; j < melBins; j++ {
|
||||
// The scaling openWakeWord applies between the two feature
|
||||
// models, and the head was fitted on its output.
|
||||
frames[i][j] = data[i*melBins+j]/10.0 + 2.0
|
||||
}
|
||||
}
|
||||
return frames, nil
|
||||
}
|
||||
|
||||
// embedding runs embedFrames mel frames through the frozen embedder.
|
||||
func (m *wakeModels) embedding(mels [][melBins]float32) ([embedDim]float32, error) {
|
||||
var e [embedDim]float32
|
||||
flat := make([]float32, 0, embedFrames*melBins)
|
||||
for _, f := range mels {
|
||||
flat = append(flat, f[:]...)
|
||||
}
|
||||
in, err := ort.NewTensor(ort.NewShape(1, embedFrames, melBins, 1), flat)
|
||||
if err != nil {
|
||||
return e, err
|
||||
}
|
||||
defer in.Destroy()
|
||||
out, err := ort.NewEmptyTensor[float32](ort.NewShape(1, 1, 1, embedDim))
|
||||
if err != nil {
|
||||
return e, err
|
||||
}
|
||||
defer out.Destroy()
|
||||
if err := m.emb.Run([]ort.Value{in}, []ort.Value{out}); err != nil {
|
||||
return e, err
|
||||
}
|
||||
copy(e[:], out.GetData())
|
||||
return e, nil
|
||||
}
|
||||
|
||||
// score runs the trained head over headWindow embeddings.
|
||||
func (m *wakeModels) score(embeds [][embedDim]float32) (float64, error) {
|
||||
flat := make([]float32, 0, headWindow*embedDim)
|
||||
for _, e := range embeds {
|
||||
flat = append(flat, e[:]...)
|
||||
}
|
||||
in, err := ort.NewTensor(ort.NewShape(1, headWindow, embedDim), flat)
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
defer in.Destroy()
|
||||
out, err := ort.NewEmptyTensor[float32](ort.NewShape(1, 1))
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
defer out.Destroy()
|
||||
if err := m.head.Run([]ort.Value{in}, []ort.Value{out}); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
return float64(out.GetData()[0]), nil
|
||||
}
|
||||
@@ -0,0 +1,195 @@
|
||||
package main
|
||||
|
||||
// The wake word, "Мэйвен" (V-487 stage two).
|
||||
//
|
||||
// Silero answers "is this frame speech". It does not answer "was this said to
|
||||
// her", and until this file existed nothing did: every utterance near the
|
||||
// microphone became a turn. What made that safe rather than expensive was
|
||||
// SurfaceVoice capping acts at L0, and L0 does not cap reading, so the room
|
||||
// could still hear his facts read back.
|
||||
//
|
||||
// This file owns the 80ms cadence and the three rings of state between the
|
||||
// models. wakefeatures.go owns the tensors.
|
||||
//
|
||||
// Nil is a working value, and it is the CLOSED gate rather than the open one.
|
||||
// Feed on a nil receiver reports no keyword; session.go asks separately
|
||||
// whether a gate exists at all. That split is deliberate: a nil that answers
|
||||
// "yes, keyword" reads as a working wake word in every log line it produces.
|
||||
|
||||
import (
|
||||
"log"
|
||||
"sync"
|
||||
)
|
||||
|
||||
// defaultWakeThreshold — score above which the keyword was said.
|
||||
//
|
||||
// Picked from the false-accept rate on held-out Russian speech, not from
|
||||
// accuracy: a miss costs him a repeat, a false accept costs a turn nobody
|
||||
// asked for. Over 65 minutes of Common Voice, 0.99 woke her three times and
|
||||
// 0.999 once, and the difference in recall was one render out of 126. So the
|
||||
// default is the strict one. `docs/evals/2026-08-09-wake-word.md` has both
|
||||
// tables.
|
||||
const defaultWakeThreshold = 0.999
|
||||
|
||||
// wakeWord is the streaming state around wakeModels. It is fed the same
|
||||
// capture frames the VAD sees and answers whether the keyword has just been
|
||||
// spoken.
|
||||
type wakeWord struct {
|
||||
mu sync.Mutex
|
||||
m *wakeModels
|
||||
|
||||
threshold float64
|
||||
|
||||
// pending holds captured samples not yet part of a full 80ms chunk, and
|
||||
// history holds the melContext samples before them.
|
||||
pending []float32
|
||||
history []float32
|
||||
|
||||
// mels is the newest embedFrames mel frames, oldest first.
|
||||
mels [][melBins]float32
|
||||
// embeds is the newest headWindow embeddings, oldest first.
|
||||
embeds [][embedDim]float32
|
||||
|
||||
last float64 // most recent score, held between chunks
|
||||
}
|
||||
|
||||
// newWakeWord loads the models and wraps them in the streaming gate.
|
||||
func newWakeWord(melPath, embedPath, headPath, libPath string, threshold float64) (*wakeWord, error) {
|
||||
m, err := newWakeModels(melPath, embedPath, headPath, libPath)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
if threshold <= 0 {
|
||||
threshold = defaultWakeThreshold
|
||||
}
|
||||
return &wakeWord{m: m, threshold: threshold}, nil
|
||||
}
|
||||
|
||||
// Close releases the models.
|
||||
func (w *wakeWord) Close() {
|
||||
if w == nil {
|
||||
return
|
||||
}
|
||||
w.mu.Lock()
|
||||
defer w.mu.Unlock()
|
||||
w.m.Close()
|
||||
w.m = nil
|
||||
}
|
||||
|
||||
// Feed takes one capture frame and reports whether the keyword was heard on
|
||||
// it. A nil wakeWord hears nothing.
|
||||
func (w *wakeWord) Feed(frame []int16) bool {
|
||||
if w == nil {
|
||||
return false
|
||||
}
|
||||
w.mu.Lock()
|
||||
defer w.mu.Unlock()
|
||||
|
||||
for _, v := range frame {
|
||||
w.pending = append(w.pending, float32(v)/32768.0)
|
||||
}
|
||||
fired := false
|
||||
for len(w.pending) >= chunkSamples {
|
||||
chunk := w.pending[:chunkSamples]
|
||||
if w.step(chunk) {
|
||||
fired = true
|
||||
}
|
||||
w.history = append(w.history[:0], tailFloat32(append(w.history, chunk...), melContext)...)
|
||||
// Slide the remainder to the front rather than reslicing. This runs
|
||||
// every 80ms for as long as the daemon lives.
|
||||
w.pending = append(w.pending[:0], w.pending[chunkSamples:]...)
|
||||
}
|
||||
return fired
|
||||
}
|
||||
|
||||
// Reset drops the streaming state, so a fresh utterance is not judged on audio
|
||||
// from before it. Called after every dispatch and after barge-in, for the same
|
||||
// reason silero is: echo-era history must not score the next sentence, and her
|
||||
// own voice saying the keyword must not wake her.
|
||||
func (w *wakeWord) Reset() {
|
||||
if w == nil {
|
||||
return
|
||||
}
|
||||
w.mu.Lock()
|
||||
defer w.mu.Unlock()
|
||||
w.pending, w.history = w.pending[:0], w.history[:0]
|
||||
w.mels, w.embeds = nil, nil
|
||||
w.last = 0
|
||||
}
|
||||
|
||||
// Score returns the most recent score, for the operator to read out of the
|
||||
// journal when picking a threshold for his room.
|
||||
func (w *wakeWord) Score() float64 {
|
||||
if w == nil {
|
||||
return 0
|
||||
}
|
||||
w.mu.Lock()
|
||||
defer w.mu.Unlock()
|
||||
return w.last
|
||||
}
|
||||
|
||||
// step runs one 80ms chunk through all three models. It returns true when the
|
||||
// score crosses the threshold on this chunk.
|
||||
func (w *wakeWord) step(chunk []float32) bool {
|
||||
buf := make([]float32, 0, melContext+len(chunk))
|
||||
if pad := melContext - len(w.history); pad > 0 {
|
||||
buf = append(buf, make([]float32, pad)...)
|
||||
}
|
||||
buf = append(buf, tailFloat32(w.history, melContext)...)
|
||||
buf = append(buf, chunk...)
|
||||
|
||||
frames, err := w.m.melFrames(buf)
|
||||
if err != nil {
|
||||
// A failed inference must not silence the microphone. Hold the last
|
||||
// score and let the next chunk try again.
|
||||
log.Printf("mavwaked: wake word: mel: %v", err)
|
||||
return false
|
||||
}
|
||||
w.mels = tailMel(append(w.mels, frames...), embedFrames)
|
||||
if len(w.mels) < embedFrames {
|
||||
return false
|
||||
}
|
||||
e, err := w.m.embedding(w.mels)
|
||||
if err != nil {
|
||||
log.Printf("mavwaked: wake word: embedding: %v", err)
|
||||
return false
|
||||
}
|
||||
w.embeds = tailEmbed(append(w.embeds, e), headWindow)
|
||||
if len(w.embeds) < headWindow {
|
||||
return false
|
||||
}
|
||||
score, err := w.m.score(w.embeds)
|
||||
if err != nil {
|
||||
log.Printf("mavwaked: wake word: head: %v", err)
|
||||
return false
|
||||
}
|
||||
// Report the crossing, not the state. A keyword held above the threshold
|
||||
// for a second is one wake, and firing on every chunk of it would make the
|
||||
// gate look open when it is merely slow to fall.
|
||||
crossed := score >= w.threshold && w.last < w.threshold
|
||||
w.last = score
|
||||
return crossed
|
||||
}
|
||||
|
||||
// The three rings. Each keeps the newest n entries and nothing older.
|
||||
|
||||
func tailFloat32(s []float32, n int) []float32 {
|
||||
if len(s) <= n {
|
||||
return s
|
||||
}
|
||||
return s[len(s)-n:]
|
||||
}
|
||||
|
||||
func tailMel(s [][melBins]float32, n int) [][melBins]float32 {
|
||||
if len(s) <= n {
|
||||
return s
|
||||
}
|
||||
return append(s[:0], s[len(s)-n:]...)
|
||||
}
|
||||
|
||||
func tailEmbed(s [][embedDim]float32, n int) [][embedDim]float32 {
|
||||
if len(s) <= n {
|
||||
return s
|
||||
}
|
||||
return append(s[:0], s[len(s)-n:]...)
|
||||
}
|
||||
@@ -0,0 +1,174 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// fakeGate fires on demand instead of running three ONNX models. The gate's
|
||||
// own arithmetic is measured on real audio in docs/evals; what these tests
|
||||
// cover is the thing that decides whether an utterance is shipped.
|
||||
type fakeGate struct {
|
||||
fireOn int // fire when this many frames have been fed, 0 never fires
|
||||
fed int
|
||||
resets int
|
||||
}
|
||||
|
||||
func (g *fakeGate) Feed(_ []int16) bool {
|
||||
g.fed++
|
||||
return g.fireOn > 0 && g.fed == g.fireOn
|
||||
}
|
||||
func (g *fakeGate) Reset() { g.resets++ }
|
||||
func (g *fakeGate) Score() float64 { return 1 }
|
||||
|
||||
// wakingSession wires a session whose gate fires on the first frame it sees.
|
||||
func wakingSession(fireOn int, window time.Duration) (*session, *fakePlayer, *fakeSender, *fakeGate) {
|
||||
sess, p, snd := newTestSession(bargeInConfig{})
|
||||
g := &fakeGate{fireOn: fireOn}
|
||||
sess.UseWakeWord(g, window)
|
||||
return sess, p, snd, g
|
||||
}
|
||||
|
||||
func TestKeywordlessSpeechNeverReachesSTT(t *testing.T) {
|
||||
sess, p, snd, g := wakingSession(0, 8*time.Second)
|
||||
speakThenPause(t, sess)
|
||||
|
||||
if len(snd.sent) != 0 {
|
||||
t.Fatalf("sent %d utterances, want 0 — this is the whole point of V-487", len(snd.sent))
|
||||
}
|
||||
if sess.ignored != 1 {
|
||||
t.Errorf("ignored = %d, want 1", sess.ignored)
|
||||
}
|
||||
if p.plays != 0 {
|
||||
t.Errorf("plays = %d, want 0", p.plays)
|
||||
}
|
||||
if g.fed == 0 {
|
||||
t.Error("the gate was never fed a frame")
|
||||
}
|
||||
}
|
||||
|
||||
func TestKeywordOpensTheGate(t *testing.T) {
|
||||
sess, p, snd, _ := wakingSession(1, 8*time.Second)
|
||||
speakThenPause(t, sess)
|
||||
|
||||
if len(snd.sent) != 1 {
|
||||
t.Fatalf("sent %d utterances, want 1", len(snd.sent))
|
||||
}
|
||||
if sess.wakes != 1 {
|
||||
t.Errorf("wakes = %d, want 1", sess.wakes)
|
||||
}
|
||||
if sess.ignored != 0 {
|
||||
t.Errorf("ignored = %d, want 0", sess.ignored)
|
||||
}
|
||||
if p.plays != 1 {
|
||||
t.Errorf("plays = %d, want 1", p.plays)
|
||||
}
|
||||
}
|
||||
|
||||
// One keyword buys one turn. Without this the microphone stays open for as
|
||||
// long as he keeps talking, which is the state the gate exists to end.
|
||||
func TestOneKeywordBuysOneTurn(t *testing.T) {
|
||||
sess, p, snd, _ := wakingSession(1, 8*time.Second)
|
||||
speakThenPause(t, sess)
|
||||
p.Stop() // she finished her reply
|
||||
sess.discard = 0 // the backlog drain is not what this measures
|
||||
speakThenPause(t, sess)
|
||||
|
||||
if len(snd.sent) != 1 {
|
||||
t.Fatalf("sent %d utterances, want 1: the second had no keyword", len(snd.sent))
|
||||
}
|
||||
if sess.ignored != 1 {
|
||||
t.Errorf("ignored = %d, want 1", sess.ignored)
|
||||
}
|
||||
}
|
||||
|
||||
// The keyword is heard, then he says nothing for longer than the window. What
|
||||
// he says after that is not addressed to her.
|
||||
func TestTheKeywordExpires(t *testing.T) {
|
||||
sess, _, snd, _ := wakingSession(1, 500*time.Millisecond)
|
||||
now := time.Unix(1750000000, 0)
|
||||
sess.now = func() time.Time { return now }
|
||||
|
||||
if err := sess.feed(context.Background(), silentBytes()); err != nil {
|
||||
t.Fatalf("feed: %v", err)
|
||||
}
|
||||
if sess.wakes != 1 {
|
||||
t.Fatalf("wakes = %d, want 1", sess.wakes)
|
||||
}
|
||||
now = now.Add(2 * time.Second)
|
||||
speakThenPause(t, sess)
|
||||
|
||||
if len(snd.sent) != 0 {
|
||||
t.Fatalf("sent %d utterances, want 0 — the keyword had expired", len(snd.sent))
|
||||
}
|
||||
}
|
||||
|
||||
// Barge-in cuts her off whether or not the keyword was heard. What he says
|
||||
// after cutting her off still has to carry it.
|
||||
func TestBargeInStillInterruptsHer(t *testing.T) {
|
||||
sess, p, _, g := wakingSession(0, 8*time.Second)
|
||||
sess.barge = bargeInConfig{RMS: 0.2, Frames: 3}
|
||||
p.playing = true
|
||||
loud := frameAt(0.35)
|
||||
for i := 0; i < 4; i++ {
|
||||
if err := sess.feed(context.Background(), loud); err != nil {
|
||||
t.Fatalf("feed %d: %v", i, err)
|
||||
}
|
||||
}
|
||||
if sess.bargeIns != 1 {
|
||||
t.Fatalf("bargeIns = %d, want 1", sess.bargeIns)
|
||||
}
|
||||
if p.stops != 1 {
|
||||
t.Errorf("stops = %d, want 1", p.stops)
|
||||
}
|
||||
if g.resets == 0 {
|
||||
t.Error("barge-in left pre-playback audio in the gate")
|
||||
}
|
||||
}
|
||||
|
||||
// Her own reply must not wake her. Frames captured while the player runs never
|
||||
// reach the gate, and the gate is cleared when playback ends.
|
||||
func TestHerOwnVoiceNeverReachesTheGate(t *testing.T) {
|
||||
sess, p, _, g := wakingSession(1, 8*time.Second)
|
||||
p.playing = true
|
||||
for i := 0; i < 10; i++ {
|
||||
if err := sess.feed(context.Background(), frameAt(0.35)); err != nil {
|
||||
t.Fatalf("feed: %v", err)
|
||||
}
|
||||
}
|
||||
if g.fed != 0 {
|
||||
t.Fatalf("gate was fed %d frames while she was speaking, want 0", g.fed)
|
||||
}
|
||||
if sess.wakes != 0 {
|
||||
t.Errorf("wakes = %d, want 0", sess.wakes)
|
||||
}
|
||||
}
|
||||
|
||||
// No model, no gate: the daemon behaves exactly as it did before V-487 stage
|
||||
// two. An operator with a missing file gets yesterday's mavwaked, not one that
|
||||
// refuses to hear anything.
|
||||
func TestNoGateShipsEveryUtterance(t *testing.T) {
|
||||
sess, _, snd := newTestSession(bargeInConfig{})
|
||||
speakThenPause(t, sess)
|
||||
|
||||
if len(snd.sent) != 1 {
|
||||
t.Fatalf("sent %d utterances, want 1", len(snd.sent))
|
||||
}
|
||||
if sess.ignored != 0 {
|
||||
t.Errorf("ignored = %d, want 0", sess.ignored)
|
||||
}
|
||||
}
|
||||
|
||||
// A nil *wakeWord is the closed gate, not a crash and not an open one.
|
||||
func TestNilWakeWordHearsNothing(t *testing.T) {
|
||||
var w *wakeWord
|
||||
if w.Feed([]int16{0, 0, 0}) {
|
||||
t.Error("a nil wake word reported the keyword")
|
||||
}
|
||||
if w.Score() != 0 {
|
||||
t.Error("a nil wake word reported a score")
|
||||
}
|
||||
w.Reset()
|
||||
w.Close()
|
||||
}
|
||||
@@ -0,0 +1,156 @@
|
||||
"""CrisperWhisper 2.0 turbo as an HTTP service, for Maven's stt.Pair.
|
||||
|
||||
Two endpoints and no framework.
|
||||
|
||||
GET /health 200 once the model is loaded, 503 while it is loading.
|
||||
POST /transcribe raw 16kHz mono PCM in, {"text","confidence"} out.
|
||||
|
||||
The body is the PCM itself rather than JSON. A minute of 16kHz mono is under
|
||||
2MB raw and about 2.6MB base64, and the format is fixed at the Maven seam, so
|
||||
headers carry it more cheaply than an envelope.
|
||||
|
||||
Why this exists at all: whisper.cpp cannot load CW2. It derives its language
|
||||
count from the vocabulary size, and CW2's 51897 tokens shift seven special
|
||||
token ids. So mavsttd stays whisper.cpp on homesrv and this runs beside the
|
||||
model on workpc, where it scores 10.4% WER in Russian against the floor's 27.5%
|
||||
(docs/evals/2026-08-09-crisperwhisper2-russian-wer.md in the Maven repo).
|
||||
|
||||
Intended mode, not verbatim. The owner asked for what he meant to say, not
|
||||
every stutter on the way there.
|
||||
"""
|
||||
|
||||
import hmac
|
||||
import json
|
||||
import logging
|
||||
import os
|
||||
import sys
|
||||
import threading
|
||||
import time
|
||||
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
|
||||
|
||||
import numpy as np
|
||||
|
||||
HOST = os.environ.get("CW2_HOST", "0.0.0.0")
|
||||
PORT = int(os.environ.get("CW2_PORT", "8081"))
|
||||
SIZE = os.environ.get("CW2_SIZE", "turbo")
|
||||
MODE = os.environ.get("CW2_MODE", "intended")
|
||||
TOKEN = os.environ.get("CW2_TOKEN", "")
|
||||
# 25MB is about thirteen minutes of 16kHz mono. Longer than any utterance and
|
||||
# short enough that a wrong caller cannot exhaust memory.
|
||||
MAX_BODY = int(os.environ.get("CW2_MAX_BODY", str(25 * 1024 * 1024)))
|
||||
|
||||
logging.basicConfig(
|
||||
level=logging.INFO, format="%(asctime)s cw2: %(message)s", stream=sys.stderr
|
||||
)
|
||||
log = logging.getLogger("cw2")
|
||||
|
||||
_model = None
|
||||
# The card holds one model and transcribes one utterance at a time. The lock is
|
||||
# what makes a second caller wait rather than corrupt the first.
|
||||
_lock = threading.Lock()
|
||||
|
||||
|
||||
def load_model():
|
||||
global _model
|
||||
from crisperwhisper import CrisperWhisperModel
|
||||
|
||||
t0 = time.perf_counter()
|
||||
# backend is forced. With ctranslate2 importable, "auto" picks ct2, which is
|
||||
# CUDA-only and this card is AMD.
|
||||
m = CrisperWhisperModel(
|
||||
SIZE, backend="transformers", compute_type="float16", device="cuda"
|
||||
)
|
||||
_model = m
|
||||
log.info("loaded %s in %.1fs, mode=%s", SIZE, time.perf_counter() - t0, MODE)
|
||||
|
||||
|
||||
def authorised(headers):
|
||||
if not TOKEN:
|
||||
return True
|
||||
got = headers.get("Authorization", "")
|
||||
return hmac.compare_digest(got, "Bearer " + TOKEN)
|
||||
|
||||
|
||||
class Handler(BaseHTTPRequestHandler):
|
||||
protocol_version = "HTTP/1.1"
|
||||
|
||||
def log_message(self, fmt, *args):
|
||||
log.info(fmt, *args)
|
||||
|
||||
def _send(self, code, payload):
|
||||
body = json.dumps(payload, ensure_ascii=False).encode("utf-8")
|
||||
self.send_response(code)
|
||||
self.send_header("Content-Type", "application/json; charset=utf-8")
|
||||
self.send_header("Content-Length", str(len(body)))
|
||||
self.end_headers()
|
||||
self.wfile.write(body)
|
||||
|
||||
def do_GET(self):
|
||||
if self.path.rstrip("/") != "/health":
|
||||
self._send(404, {"error": "not found"})
|
||||
return
|
||||
if _model is None:
|
||||
self._send(503, {"status": "loading"})
|
||||
return
|
||||
self._send(200, {"status": "ok", "model": SIZE, "mode": MODE})
|
||||
|
||||
def do_POST(self):
|
||||
if self.path.rstrip("/") != "/transcribe":
|
||||
self._send(404, {"error": "not found"})
|
||||
return
|
||||
if not authorised(self.headers):
|
||||
self._send(401, {"error": "unauthorised"})
|
||||
return
|
||||
if _model is None:
|
||||
self._send(503, {"error": "loading"})
|
||||
return
|
||||
|
||||
length = int(self.headers.get("Content-Length", "0"))
|
||||
if length <= 0 or length > MAX_BODY:
|
||||
self._send(413, {"error": "bad body length"})
|
||||
return
|
||||
raw = self.rfile.read(length)
|
||||
|
||||
rate = int(self.headers.get("X-Sample-Rate", "16000"))
|
||||
channels = int(self.headers.get("X-Channels", "1"))
|
||||
bits = int(self.headers.get("X-Sample-Bits", "16"))
|
||||
lang = self.headers.get("X-Language", "ru") or "ru"
|
||||
if channels != 1 or bits != 16:
|
||||
self._send(400, {"error": "want 16-bit mono pcm"})
|
||||
return
|
||||
|
||||
# int16 little-endian to the float32 the encoder wants.
|
||||
wav = np.frombuffer(raw, dtype="<i2").astype(np.float32) / 32768.0
|
||||
if wav.size == 0:
|
||||
self._send(200, {"text": "", "confidence": 0.0})
|
||||
return
|
||||
|
||||
t0 = time.perf_counter()
|
||||
try:
|
||||
with _lock:
|
||||
res = _model.transcribe(wav, sr=rate, language=lang, mode=MODE)
|
||||
except Exception as exc: # noqa: BLE001 - the caller falls back to mavsttd
|
||||
log.exception("transcribe failed")
|
||||
self._send(500, {"error": str(exc)})
|
||||
return
|
||||
elapsed = time.perf_counter() - t0
|
||||
text = (res.text or "").strip()
|
||||
log.info("%.2fs audio in %.2fs: %r", wav.size / rate, elapsed, text[:60])
|
||||
# The model reports no calibrated score. 1.0 would be a claim, and the
|
||||
# Maven side reads confidence only to log it.
|
||||
self._send(200, {"text": text, "confidence": 0.0})
|
||||
|
||||
|
||||
def main():
|
||||
if not TOKEN:
|
||||
log.warning("no CW2_TOKEN set: anything on the LAN can post audio here")
|
||||
# Bind before loading, so a restart answers 503 rather than refusing the
|
||||
# connection. Both make Maven fall back, but only one of them says why.
|
||||
srv = ThreadingHTTPServer((HOST, PORT), Handler)
|
||||
threading.Thread(target=load_model, daemon=True).start()
|
||||
log.info("listening on %s:%d", HOST, PORT)
|
||||
srv.serve_forever()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,40 @@
|
||||
# maven-voice-tunnel — the ssh leg that carries the voice wire to homesrv.
|
||||
#
|
||||
# Runs on workpc, as a user unit (`systemctl --user`), beside mavgpud.service.
|
||||
#
|
||||
# WHY THIS EXISTS AT ALL. internal/voice is plaintext and unauthenticated.
|
||||
# Its own server doc says production binds inside the wg tunnel, because "the
|
||||
# wg layer IS the L0 floor". workpc is not a wg peer, it sits on wlan0. So ssh
|
||||
# is the substitute floor: it authenticates with his key and encrypts the leg,
|
||||
# and mavend's published port stays on homesrv loopback (127.0.0.1:9110).
|
||||
# Nothing about this puts a Maven port on the LAN.
|
||||
#
|
||||
# Do not replace this with a LAN bind. SurfaceVoice caps acts at L0, so an
|
||||
# unauthorized speaker could not run a destructive tool. It would still hear
|
||||
# his facts, his notes and his calendar read back, and L0 does not cap reading.
|
||||
#
|
||||
# install: cp to ~/.config/systemd/user/ on workpc
|
||||
# systemctl --user enable --now maven-voice-tunnel.service
|
||||
|
||||
[Unit]
|
||||
Description=SSH tunnel to mavend's voice wire on homesrv
|
||||
After=network-online.target
|
||||
Wants=network-online.target
|
||||
|
||||
[Service]
|
||||
# -N: no remote command, forwarding only.
|
||||
# ExitOnForwardFailure: fail loudly rather than sit up with a dead forward,
|
||||
# which is what makes Restart meaningful.
|
||||
# ServerAlive*: a laptop that suspends drops the tunnel silently otherwise.
|
||||
ExecStart=/usr/bin/ssh -N \
|
||||
-o ExitOnForwardFailure=yes \
|
||||
-o ServerAliveInterval=30 \
|
||||
-o ServerAliveCountMax=3 \
|
||||
-o BatchMode=yes \
|
||||
-L 127.0.0.1:9100:127.0.0.1:9110 \
|
||||
kami@192.168.1.104
|
||||
Restart=always
|
||||
RestartSec=5
|
||||
|
||||
[Install]
|
||||
WantedBy=default.target
|
||||
+42
-2
@@ -25,6 +25,25 @@
|
||||
"llm_nudges": false
|
||||
},
|
||||
|
||||
"//ntfy": [
|
||||
"The second reach (V-649). Until 07-08-2026 telegram was the only one, and",
|
||||
"telegram needs api.telegram.org, the socks relay below and a matching ufw",
|
||||
"rule — three things in series that have each failed once, and when they do",
|
||||
"a sev4 nudge has nowhere to go. ntfy shares none of them: it is reached",
|
||||
"directly, no relay.",
|
||||
"It is not only a spare. The routing table sends sev3-away and away",
|
||||
"reminders here and NOWHERE else, so with this block absent those two",
|
||||
"routes hit a nil sink and vanish without a log or an outbox row.",
|
||||
"The credential is an ntfy access token, scoped write-only to this one",
|
||||
"topic, so a popped sink can push to it and cannot read it back. Set it in",
|
||||
"deploy/telegram.env beside the telegram secrets; that file is gitignored."
|
||||
],
|
||||
"ntfy": {
|
||||
"base_url": "https://ntfy.kvmx.ru",
|
||||
"topic": "maven",
|
||||
"token": "${NTFY_TOKEN}"
|
||||
},
|
||||
|
||||
"telegram": {
|
||||
"bot_token": "${TELEGRAM_BOT_TOKEN}",
|
||||
"chat_id": "${TELEGRAM_CHAT_ID}",
|
||||
@@ -59,10 +78,30 @@
|
||||
"Addressed by LAN address, not container name: mavgpud runs on another",
|
||||
"machine and there is no shared docker network to name it on."
|
||||
],
|
||||
"//workstation.stt": [
|
||||
"CrisperWhisper 2.0 turbo on the same machine, a second service on port",
|
||||
"8081 and not a second endpoint on mavgpud. whisper.cpp cannot load CW2 at",
|
||||
"all: it derives its language count from the vocabulary size, and CW2's",
|
||||
"51897 tokens shift seven special token ids. So it runs under transformers",
|
||||
"there and mavsttd stays whisper.cpp here.",
|
||||
"Worth the second service: CW2 turbo scores 10.4% WER in Russian against",
|
||||
"27.5% for the ggml-small.bin mavsttd loads, measured on 200 Golos clips",
|
||||
"in docs/evals/2026-08-09-crisperwhisper2-russian-wer.md.",
|
||||
"Deleting this block sends every utterance to mavsttd, which is what the",
|
||||
"box did before it existed. A worse transcript is still a turn, so the",
|
||||
"fallback is silent and Kami is never told which machine heard him.",
|
||||
"The token is what stops anything on the LAN posting audio to that port."
|
||||
],
|
||||
"workstation": {
|
||||
"url": "http://192.168.1.105:8080",
|
||||
"probe": "15s",
|
||||
"timeout": "90s"
|
||||
"timeout": "90s",
|
||||
"stt": {
|
||||
"url": "http://192.168.1.105:8081/transcribe",
|
||||
"token": "${MAVEN_STT_TOKEN}",
|
||||
"probe": "15s",
|
||||
"timeout": "10s"
|
||||
}
|
||||
},
|
||||
|
||||
"//search": [
|
||||
@@ -212,7 +251,8 @@
|
||||
"embedder": {
|
||||
"model_path": "/opt/maven/models/embedder/multilingual-e5-small/model_quantized.onnx",
|
||||
"tokenizer_path": "/opt/maven/models/embedder/multilingual-e5-small/tokenizer.json",
|
||||
"lib_path": "/opt/maven/lib/libonnxruntime.so"
|
||||
"lib_path": "/opt/maven/lib/libonnxruntime.so",
|
||||
"heads_path": "/opt/maven/models/embedder/router-heads/router_heads.onnx"
|
||||
},
|
||||
"llm_router": true,
|
||||
"query_min_score": 0.55,
|
||||
|
||||
+21
-5
@@ -2,9 +2,14 @@
|
||||
"listen": ":8080",
|
||||
"llama_addr": "127.0.0.1:10000",
|
||||
"llama_bin": "llama-server",
|
||||
"//llama_args": [
|
||||
"E4B carries no MTP tensors, so the speculative flags are gone with the 12B.",
|
||||
"MTP on this box is a separate gguf of architecture gemma4-assistant with",
|
||||
"nextn_predict_layers=4, and mtp-gemma-4-12B-it-BF16 is the only one there is.",
|
||||
"Its head is trained against the 12B's hidden states, so it cannot drive E4B."
|
||||
],
|
||||
"llama_args": [
|
||||
"-m", "/mnt/D/AI/gemma4/gemma-4-12B-it-qat-UD-Q4_K_XL.gguf",
|
||||
"-md", "/mnt/D/AI/gemma4/mtp-gemma-4-12B-it-BF16.gguf",
|
||||
"-m", "/mnt/D/AI/gemma4/gemma-4-E4B-it-qat-UD-Q4_K_XL.gguf",
|
||||
"-ngl", "99",
|
||||
"-fa", "on",
|
||||
"-np", "1",
|
||||
@@ -15,11 +20,22 @@
|
||||
"--batch-size", "2048",
|
||||
"--ubatch-size", "512",
|
||||
"--jinja",
|
||||
"--chat-template-kwargs", "{\"enable_thinking\":false}",
|
||||
"--spec-type", "draft-mtp",
|
||||
"--spec-draft-n-max", "2"
|
||||
"--chat-template-kwargs", "{\"enable_thinking\":false}"
|
||||
],
|
||||
|
||||
"//stt": [
|
||||
"CrisperWhisper 2.0 turbo, which Maven reaches directly on port 8081.",
|
||||
"mavgpud runs it because it is a ROCm process on this card: under its own",
|
||||
"systemd unit it registered on the KFD and the supervisor evicted",
|
||||
"llama-server every few seconds. CW2_TOKEN comes from the unit's",
|
||||
"EnvironmentFile and is never a flag value."
|
||||
],
|
||||
"stt": {
|
||||
"addr": "127.0.0.1:8081",
|
||||
"bin": "/home/kami/Programs/cw2-eval/.venv/bin/python",
|
||||
"args": ["/home/kami/Programs/cw2-service/serve.py"]
|
||||
},
|
||||
|
||||
"kfd_root": "/sys/class/kfd/kfd/proc",
|
||||
"drm_device": "/sys/class/drm/card1/device",
|
||||
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
[Unit]
|
||||
# Runs on the workstation (bugmachine), not on homesrv. Install as a systemd
|
||||
# Runs on the workstation (workpc), not on homesrv. Install as a systemd
|
||||
# user unit and turn on lingering, so the card is supervised after a reboot
|
||||
# with nobody logged in:
|
||||
#
|
||||
@@ -8,10 +8,15 @@
|
||||
# scp deploy/mavgpud.service workpc:~/.config/systemd/user/mavgpud.service
|
||||
# ssh workpc 'systemctl --user daemon-reload && systemctl --user enable --now mavgpud'
|
||||
# sudo loginctl enable-linger kami
|
||||
Description=Maven GPU supervisor (holds llama-server while the card is free)
|
||||
Description=Maven GPU supervisor (holds llama-server and CW2 while the card is free)
|
||||
After=network.target
|
||||
|
||||
[Service]
|
||||
# CW2_TOKEN for the transcriber child, which inherits this environment. The
|
||||
# token is read from a file and never appears as a flag value, the rule
|
||||
# mavpoll and mavmaild follow. Missing file, no transcriber auth, so keep the
|
||||
# dash off: a mavgpud that cannot read it must fail loudly.
|
||||
EnvironmentFile=%h/Programs/cw2-service/cw2.env
|
||||
ExecStart=%h/.local/bin/mavgpud -config %h/.config/mavgpud.json
|
||||
Restart=always
|
||||
RestartSec=5
|
||||
|
||||
@@ -0,0 +1,53 @@
|
||||
# mavwaked — always-on listening, on workpc where the microphone is.
|
||||
#
|
||||
# User unit, beside mavgpud.service and maven-voice-tunnel.service. It is a
|
||||
# user unit because it needs his ALSA session and his ssh agent, and because
|
||||
# it should stop when he logs out.
|
||||
#
|
||||
# THERE IS NO WAKE WORD YET (V-487 stage two). Anything spoken near the fifine
|
||||
# becomes a turn. What makes that safe rather than expensive is voiceSender:
|
||||
# it sends Surface=SurfaceVoice, which caps every command at L0, so no
|
||||
# accidental trigger runs a destructive act. It does not stop her answering
|
||||
# out loud, so this unit is his to stop when the room is not his alone.
|
||||
#
|
||||
# -vad-model is passed on purpose. Silero answers "is this frame speech" where
|
||||
# the energy floor answers "is this frame loud". It declines white noise at
|
||||
# the same RMS 0 frames to 68-99, and still hears all four spoken fixtures
|
||||
# (docs/evals/2026-08-09-silero-vad.md). It costs 509us a frame, 1.7% of one
|
||||
# core, and never touches the GPU. Drop the flag and the energy floor is back.
|
||||
#
|
||||
# -barge-in is NOT passed. The threshold is room-specific and this room has no
|
||||
# number yet. Turn it on only after reading the "suppressed while speaking"
|
||||
# means out of this unit's own journal, never by guessing.
|
||||
#
|
||||
# install: cp to ~/.config/systemd/user/ on workpc
|
||||
# systemctl --user enable --now mavwaked.service
|
||||
|
||||
[Unit]
|
||||
Description=Maven always-on listening (VAD, no wake word yet)
|
||||
# The tunnel is the only path to mavend and the only thing authenticating it.
|
||||
Requires=maven-voice-tunnel.service
|
||||
After=maven-voice-tunnel.service
|
||||
|
||||
[Service]
|
||||
# card 0 is the fifine USB microphone. Named, and not "default", because the
|
||||
# default device follows whatever pipewire last decided and this daemon should
|
||||
# not change ears when he plugs in a headset.
|
||||
#
|
||||
# plughw and not hw. mavwaked asks arecord for 16kHz mono, which is what the
|
||||
# whole pipeline is canonical in. The fifine offers 2 channels at 44100 or
|
||||
# 48000 and nothing else, so bare hw:0,0 dies on "Channels count non
|
||||
# available" before a frame is read. plughw puts ALSA's downmix and resampler
|
||||
# in front. Any replacement microphone wants the same treatment.
|
||||
Environment=LD_LIBRARY_PATH=%h/.local/lib
|
||||
ExecStart=%h/.local/bin/mavwaked \
|
||||
-device plughw:0,0 \
|
||||
-addr 127.0.0.1:9100 \
|
||||
-lang ru \
|
||||
-vad-model %h/.local/share/maven/models/silero_vad.onnx \
|
||||
-onnx-lib %h/.local/lib/libonnxruntime.so
|
||||
Restart=on-failure
|
||||
RestartSec=5
|
||||
|
||||
[Install]
|
||||
WantedBy=default.target
|
||||
@@ -1,5 +1,12 @@
|
||||
# Telegram bot token and chat ID for mavend's away-channel reach.
|
||||
# Secrets for mavend's away-channel reaches. The file is still called
|
||||
# telegram.env because compose names it that; it holds both reaches now.
|
||||
# Copy this file to deploy/telegram.env and fill in real values.
|
||||
# deploy/telegram.env is gitignored — never commit the real secrets.
|
||||
TELEGRAM_BOT_TOKEN=
|
||||
TELEGRAM_CHAT_ID=
|
||||
|
||||
# ntfy access token for the `maven` topic, the second reach (V-649). Mint it on
|
||||
# the ntfy server with write access to that topic and nothing else:
|
||||
# ntfy token add --expires=never maven
|
||||
# Read access is not needed — mavend publishes and never subscribes.
|
||||
NTFY_TOKEN=
|
||||
|
||||
@@ -53,6 +53,19 @@ services:
|
||||
# the decrypted working copy lives in RAM (see db_tmpfs in mavend.json).
|
||||
tmpfs:
|
||||
- /dev/shm
|
||||
# the voice wire, for mavwaked and mavenclient on workpc (V-515).
|
||||
#
|
||||
# LOOPBACK ONLY, and that is the whole security argument. internal/voice
|
||||
# is plaintext with no auth: its own server doc says production binds
|
||||
# inside the wg tunnel, "the wg layer IS the L0 floor". workpc is not a wg
|
||||
# peer, it is on wlan0. So the tunnel is ssh instead, terminated on this
|
||||
# loopback address, and nothing new is on the LAN. Anyone who could reach
|
||||
# a LAN-bound port here could push audio and hear his facts read back.
|
||||
# SurfaceVoice caps acts at L0; it does not cap reading.
|
||||
#
|
||||
# Host 9100 is Vikunja's MCP, hence 9110. The container side stays 9100
|
||||
# so mavweb keeps reaching mavend:9100 by name.
|
||||
ports: ["127.0.0.1:9110:9100"]
|
||||
|
||||
mavsttd:
|
||||
<<: *image
|
||||
|
||||
@@ -0,0 +1,196 @@
|
||||
# Deployment: the boxes, the models, the daemons
|
||||
|
||||
*Last verified: 2026-08-09 @ a9b480a*
|
||||
|
||||
What runs where, and why each choice was made. `CLAUDE.md` carries only the
|
||||
rules. This file carries the reasoning.
|
||||
|
||||
## The two boxes
|
||||
|
||||
**homesrv** is a Ryzen 5 5600U laptop and the deploy target. It offloads to the
|
||||
Vega iGPU over Vulkan (`n_gpu_layers: 99`). Compose passes `/dev/dri` and the
|
||||
render gid (993), and without both Vulkan enumerates zero devices and
|
||||
llama-server falls back to CPU silently.
|
||||
|
||||
**workpc** is the workstation, 16GB of VRAM, reached as `kami@workpc` at
|
||||
192.168.1.105. Model work moved there on 2026-08-02 by the owner's call, because
|
||||
homesrv cannot grow a GPU.
|
||||
|
||||
Three rules govern the seam:
|
||||
|
||||
- **The workstation is never assumed up.**
|
||||
- **Fall back silently** when it would only do the job better.
|
||||
- **Name the gap** when the resident model cannot do the job at all.
|
||||
|
||||
A world question goes through `LLMPhraser.PhraseWorld` and returns `worldGap`
|
||||
(`cmd/mavend/worldmodel.go`) rather than an invented answer. A box with no
|
||||
`workstation` block behaves exactly as it did before the seam. `docs/offload.md`
|
||||
says which caller is which.
|
||||
|
||||
## The resident model
|
||||
|
||||
**Qwen3-1.7B** (`UD-Q4_K_XL`), stock, not yet the CPT'd one. It is a Thinking
|
||||
variant, so `n_ctx` is 4096. Reasoning tokens need the room, and 4096 is what
|
||||
every score was measured at.
|
||||
|
||||
The target is the locally CPT'd Qwen3-1.7B (V-122, training in flight). Stock
|
||||
already speaks good Russian. What it gets wrong is the persona. It writes `я рад`
|
||||
where Maven needs `рада`.
|
||||
|
||||
**Do not bother with sub-500M models.** LFM2.5-230M and 350M were measured on
|
||||
2026-07-31 and both are unusable in Russian
|
||||
(`docs/evals/2026-07-31-model-bakeoff.md`). Their published IFEval and BFCL
|
||||
numbers are English-only.
|
||||
|
||||
Model files live in `/mnt/hdd1/llms`, bind-mounted to `/opt/maven/models/llm`.
|
||||
That **shadows** the repo's `models/llm/`, so a gguf sitting there is not loaded
|
||||
by anything. Swapping the resident model is a one-line change to
|
||||
`phraser.model_path` in `deploy/mavend.json`.
|
||||
|
||||
The workstation model is gemma-4-E4B as of 2026-08-09, replacing the 12B by the
|
||||
owner's call. Keep the 12B gguf. It is the better teacher for label runs, at
|
||||
72.7% destination against E4B's 57.6%.
|
||||
|
||||
## The embedder
|
||||
|
||||
**It stays on homesrv permanently**, because it backs the floor. It is
|
||||
multilingual-e5-small, quantized and asymmetric. `EmbedQuery` and `EmbedPassage`
|
||||
apply the `query:` and `passage:` prefixes it was trained with. Calling plain
|
||||
`Embed` on a note is a bug. See `docs/evals/2026-08-04-recall-e5-small.md`.
|
||||
|
||||
The vendored onnxruntime under `deps/` has two copies, and the stale one is
|
||||
1.17.1. The live runtime is 1.26.0, and the Go binding asks for API 26. Anything
|
||||
shipped to another box needs `deps/onnxruntime-linux-x64-1.26.0`.
|
||||
|
||||
## Speech-to-text
|
||||
|
||||
`sttSeam` in `cmd/mavend/voicewire.go` builds an `stt.Pair` beside `modelSeam`.
|
||||
It prefers CrisperWhisper 2.0 turbo on workpc with mavsttd as the floor. It takes
|
||||
only the silent half of the rule, because a worse transcript is still a turn. So
|
||||
`stt.Pair` has no `TranscribeRemote` and the fallback is never spoken.
|
||||
|
||||
CW2 turbo scores 10.4% WER in Russian against 27.5% for the `ggml-small.bin`
|
||||
mavsttd loads, over 200 Golos clips
|
||||
(`docs/evals/2026-08-09-crisperwhisper2-russian-wer.md`).
|
||||
|
||||
**whisper.cpp cannot load CW2 at all.** It reads its language count off the
|
||||
vocabulary size. CW2's 51897 tokens shift seven special token ids. So CW2 is its
|
||||
own transformers service on port 8081 (`deploy/cw2/serve.py`).
|
||||
`stt.HTTPTranscriber` posts raw PCM to it with a bearer token, because audio is
|
||||
the most sensitive thing that crosses this seam. The switch is `workstation.stt`
|
||||
in `deploy/mavend.json`, and deleting the block sends every utterance to mavsttd.
|
||||
|
||||
**mavgpud runs that service as a second child.** This is not an optimisation.
|
||||
CW2 is a ROCm process on the same card, so it registers on the KFD like any
|
||||
contender. Under its own systemd unit it made mavgpud evict llama-server every
|
||||
few seconds. That took the model arm down for eight minutes on 2026-08-09. The
|
||||
card needs one owner. CW2 is on the yield clock and not the idle one. At 1.6GB
|
||||
it denies the card to nobody.
|
||||
|
||||
Text-to-speech has not moved. piper on homesrv is the only synthesizer.
|
||||
|
||||
## The daemons
|
||||
|
||||
| Binary | Role |
|
||||
|---|---|
|
||||
| `mavend` | **Core.** Router, phraser, memory, reminders, digestion tick. Owns the DB and IPC socket. |
|
||||
| `mavweb` | HTTP UI and PWA (`/dash`, `/history`, `/trace`, `/notifications`, `/tools`), WebAuthn auth. |
|
||||
| `mavsttd` | Speech-to-text (whisper.cpp, CGO). |
|
||||
| `mavttsd` | Text-to-speech (piper subprocess). |
|
||||
| `mavwaked` | Wake-word and VAD gate. Runs on workpc. |
|
||||
| `mavenclient` | Voice loop client (mic, stt, core, tts). Not deployed. |
|
||||
| `mavpoll` | Environment poller: netdata alarms, uptime-kuma, zenmoney, wireguard presence. Writes facts, sends nothing. Telegram is `internal/delivery/telegramsink`. |
|
||||
| `mavcaldav` | CalDAV calendar sync. |
|
||||
| `mavmaild` | Mail reader (IMAP, read-only). Holds the IMAP password, core never sees it. |
|
||||
| `mavgpud` | GPU supervisor. **Runs on workpc**, own unit `deploy/mavgpud.service`. Keeps llama-server loaded while the card is free (V-488). Maven never asks it for anything and reads `/health` through `llm.Pair`. |
|
||||
| `mavupdate` | Not a daemon. Operator CLI a human runs on the box to deploy a new build. |
|
||||
|
||||
Two binaries have no Makefile target and neither is deployed. `mavseal` encrypts
|
||||
a live tmpfs working copy back to the ciphertext file when mavend was killed
|
||||
before `defer st.Close()` sealed it. `labelgen` runs the stage 0 grammars over
|
||||
utterances and prints JSONL, the training data for the routing heads.
|
||||
|
||||
Daemons are wired socket-to-socket, not linked. `internal/ipc` is the wire
|
||||
protocol. `deploy/mavend.json` sets socket paths, model paths and the phraser and
|
||||
embedder blocks, with `${VAR}` expansion from gitignored `deploy/telegram.env`.
|
||||
|
||||
### Who is in compose, and who is not
|
||||
|
||||
**`docker-compose.yml` runs five**: `mavend`, `mavsttd`, `mavttsd`, `mavweb`,
|
||||
`mavpoll`. Count against compose, not against the table above.
|
||||
|
||||
`mavmaild` and `mavcaldav` are commented out, each with the reason beside it. The
|
||||
first needs a mail account and the second a CalDAV account, and this box has
|
||||
neither. Two things ride on the CalDAV absence (V-644). Agenda questions route to
|
||||
`IntentQuery` at stage 0, and the `calendar` query source then reads a table
|
||||
nobody writes. And `loop.State.CalendarBusy` is fed by the same facts, so the
|
||||
gate's "do not nag mid-meeting" is permanently false.
|
||||
|
||||
`mavenclient` is still absent. `mavwaked` moved to workpc on 2026-08-09 (V-515).
|
||||
|
||||
### The voice wire
|
||||
|
||||
`internal/voice` is plaintext with no auth. Its own server doc says production
|
||||
binds inside the wg tunnel, because the wg layer is the L0 floor. workpc is not
|
||||
a wg peer, it sits on wlan0. So the tunnel is ssh instead.
|
||||
|
||||
mavend publishes the voice port to homesrv loopback only, `127.0.0.1:9110`.
|
||||
Host 9100 is Vikunja's MCP, hence 9110. The container side stays 9100 so mavweb
|
||||
keeps reaching `mavend:9100` by name. `deploy/maven-voice-tunnel.service` on
|
||||
workpc forwards it over his key.
|
||||
|
||||
**Do not replace this with a LAN bind.** `SurfaceVoice` caps acts at L0, so an
|
||||
unauthorized speaker could not run a destructive tool. L0 does not cap reading,
|
||||
so they would still hear his facts, notes and calendar read back.
|
||||
|
||||
Both `mavwaked` and `mavenclient` speak `voice.Dial`, not `ipc.Dial`. The
|
||||
`netaddr` token guards the daemon-to-daemon IPC seam and never touches this one.
|
||||
`ipc.Dial` does take `tcp://host:port?token=...`, which is why V-515 was filed
|
||||
as a config change. That premise was wrong, and the ssh leg is the correction.
|
||||
|
||||
The voice loop belongs on a client machine where the owner is standing, and that
|
||||
machine is workpc (V-463, `docs/plans/17-where-the-voice-loop-runs.md`). homesrv
|
||||
has a microphone, because it is a laptop, but it is in the wrong room.
|
||||
|
||||
### mavwaked on workpc
|
||||
|
||||
`deploy/mavwaked.service`, a user unit beside `mavgpud.service`. Two flags are
|
||||
deliberate.
|
||||
|
||||
`-vad-model` is passed. Silero answers "is this frame speech" where the energy
|
||||
floor answers "is this frame loud". It declines white noise at the same RMS, 0
|
||||
frames against 68 to 99, and still hears all four spoken fixtures
|
||||
(`docs/evals/2026-08-09-silero-vad.md`). It costs 509µs a frame and never touches
|
||||
the GPU. A model that will not load is logged and not fatal.
|
||||
|
||||
`-barge-in` is not passed. The threshold is room-specific and this room has no
|
||||
number yet. Read the "suppressed while speaking" means out of the journal first.
|
||||
|
||||
The device is `plughw:0,0` and not `hw:0,0`. The fifine offers 2 channels at
|
||||
44100 or 48000 and nothing else, and mavwaked asks arecord for 16kHz mono. Bare
|
||||
`hw` dies on "Channels count non available" before a frame is read.
|
||||
|
||||
There is no wake word yet (V-487 stage two), so the loop runs open.
|
||||
|
||||
mavwaked connects at startup and holds the conn, so a nudge routed to voice
|
||||
reaches the speaker before he has said anything (V-671). It used to connect
|
||||
lazily, which made the failure silent rather than absent: after one utterance
|
||||
the session existed, `PushToMostRecent` succeeded, the dispatcher stopped
|
||||
rerouting to telegram and ntfy, and mavwaked discarded the audio.
|
||||
|
||||
**Passwords are read from files, never taken as flag values.** `mavcaldav` uses
|
||||
`-pass-file` and `-render-pass-file`. `mavpoll` and `mavmaild` follow the same
|
||||
rule.
|
||||
|
||||
## Web UI conventions
|
||||
|
||||
Server-rendered pages share `cmd/mavweb/static/ui.css` (served at `/ui.css`) and
|
||||
the shell partial in `cmd/mavweb/shell.html`. A page opens with
|
||||
`{{template "shellTop" "<page-key>"}}` and closes with `{{template "shellBottom"}}`,
|
||||
and the key marks the active sidebar link.
|
||||
|
||||
Every page is its own embedded `.html` file next to `main.go`. No page markup
|
||||
lives in Go, and the sidebar is data (`sidebarSections`, `pageIcon`) the template
|
||||
renders. No per-page `<style>` beyond true one-offs. Wrap every table in
|
||||
`<div class=scroll>` so wide data pans on a phone. Local preview and headless
|
||||
screenshot recipes are in `AGENTS.md`.
|
||||
+57
-1
@@ -1,6 +1,6 @@
|
||||
# Maven — Design
|
||||
|
||||
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
|
||||
*Last verified: 2026-08-07 @ beb093a. Living doc: correct it in place, do not append.*
|
||||
|
||||
> Folded 2026-07-30 from `SPEC.md` (north star, 2026-07-03), `maven.md`
|
||||
> (consolidated decisions, 2026-06-30) and `ROADMAP.md` (execution plan,
|
||||
@@ -282,6 +282,62 @@ Three reasons, in the order they settle it:
|
||||
So the notice stays what it is: the in-process TTL case, where she really did
|
||||
wait and really did let go.
|
||||
|
||||
#### A parked question may step aside three times
|
||||
|
||||
Decided 2026-08-07 (V-654). A side query or an aside suspends the parked
|
||||
question instead of dropping it. The words are answered as themselves, and the
|
||||
question comes back on the end of the same reply.
|
||||
|
||||
Neither bound on a question reaches that path. No attempt is spent, because a
|
||||
side query is not a failed answer, so `MaxAttempts` never applies.
|
||||
`noteSuspended` also restarts the 90s clock, since she is about to speak the
|
||||
question again. So the TTL cannot arrive while he keeps talking.
|
||||
|
||||
Measured on 2026-08-07: one unfilled time slot rode the tail of six consecutive
|
||||
unrelated replies. It stopped only when a seventh turn happened to read as a
|
||||
failed answer. See `docs/evals/2026-08-07-week-of-usage.md`.
|
||||
|
||||
`PendingQuestion.Suspends` counts the step-asides. `MaxSuspends` is 3, matching
|
||||
`DefaultMaxAttempts`. Past it she lets the request go, with the same
|
||||
`clarifyDropped` line every other drop uses. The owner's rule is unchanged. A
|
||||
question still ends by being answered or by being let go out loud. This only
|
||||
recognises three unrelated requests in a row as the second of those.
|
||||
|
||||
The count is of CONSECUTIVE step-asides. It resets the moment he answers, in
|
||||
`resolveClarifyAnswer`. An answer that gives her nothing she asked for resets it
|
||||
too. "Позвонить маме" against a question about the time is still him in the
|
||||
exchange. The retry it costs is bound enough on its own.
|
||||
|
||||
#### And it may ride four turns in all
|
||||
|
||||
Decided 2026-08-08 (V-663), because the bound above did not move the number it
|
||||
was written for. Twenty-six of 140 turns carried a tail before it landed and
|
||||
twenty-six carried one after.
|
||||
|
||||
Two bounds rearm each other. An aside spends no attempt, so `MaxAttempts` never
|
||||
reaches it. A turn that reads as a failed answer zeroes `Suspends`, so
|
||||
`MaxSuspends` never reaches the asides. Alternating them, each bound is restored
|
||||
by the other's traffic. Measured on 2026-08-08: one question about a reminder's
|
||||
day rode turns 7 to 13. It ended only because turn 14 was a new request.
|
||||
|
||||
`PendingQuestion.Rides` counts the same event as `Suspends` with the resets
|
||||
taken out. It is set once, incremented only in `noteSuspended`, carried across
|
||||
the re-park in `askRemainingGap`, and read by nothing that could lower it.
|
||||
`MaxRides` is 4, one looser than `MaxSuspends` so that the tighter statement
|
||||
about a run stays reachable.
|
||||
|
||||
This is a bound, not a cure. It ends the measured ride one turn early. Most of
|
||||
that ride's length is attempts, spent because `classifyTurnRole` reads "спасибо"
|
||||
and "привет" as failed answers to a question about a day. That is the next
|
||||
thing to fix and it is not a bound.
|
||||
|
||||
The re-ask is also two sentences rather than one. It used to be spliced onto the
|
||||
answer with a comma. On a real answer that buries the question in the tail of
|
||||
one run-on thought:
|
||||
|
||||
> вот что я нашла: вайфай пароль лежит в ящике стола, на какое время поставить
|
||||
> напоминание?
|
||||
|
||||
### save-where — the two-memory routing axis
|
||||
|
||||
One discriminator: **does the loop evaluate a predicate against it?**
|
||||
|
||||
@@ -30,6 +30,12 @@ Maven is the user-facing control center, but not the source of truth for identit
|
||||
|
||||
Maven provides the human interface over the other systems.
|
||||
|
||||
| Service | Owns | Maven's client | Configured at |
|
||||
|---|---|---|---|
|
||||
| **Nexus** | Canonical entity ids, names, aliases, relationships. Projects, services, devices, people, pets, places. | `nexusClient` in `cmd/mavend/ecosystem.go`, `POST /api/v1/resolve` | `nexus.url` (`http://nexus:9740`) |
|
||||
| **Praxis** | Operational attention and item lifecycle. What needs looking at, what changed, what is unresolved. | `praxisClient`, the HTTP tools API under `/api/v1/tools/` | `praxis.url` (`http://praxis:8989`) |
|
||||
| **Hexis** | The capability registry and the only path to executing anything. | vendored `github.com/kami/hexis/pkg/client` | `hexis.url` (`http://hexis:9741`) |
|
||||
|
||||
It is responsible for:
|
||||
|
||||
- interpreting Russian and English utterances
|
||||
|
||||
@@ -0,0 +1,321 @@
|
||||
# Raw transcript, 2026-08-07 week of usage
|
||||
|
||||
Companion to `2026-08-07-week-of-usage.md`. 74 turns through `POST /api/chat`.
|
||||
Days are simulated; the wall clock was 03:27-03:40 throughout, which is why every
|
||||
clock answer says night.
|
||||
|
||||
## Day 1
|
||||
|
||||
```
|
||||
YOU: доброе утро
|
||||
MAVEN: Добро пожаловать! Солнечно и тепло. Всё в порядке.
|
||||
[7.7s t=5]
|
||||
|
||||
YOU: какой сегодня день?
|
||||
MAVEN: сегодня пятница, 7 августа 2026 года
|
||||
[0.0s t=6]
|
||||
|
||||
YOU: напомни в 11:00 позвонить маме
|
||||
MAVEN: Сейчас 03:29. В какой день?
|
||||
[0.4s t=7]
|
||||
|
||||
YOU: запиши что я пью кофе без сахара
|
||||
MAVEN: Прошлую просьбу отпускаю. Заметка сохранена.
|
||||
[8.0s t=8]
|
||||
|
||||
YOU: что у меня сегодня?
|
||||
MAVEN: на 07.08.2026 ничего нет.
|
||||
[2.6s t=9]
|
||||
|
||||
YOU: мой любимый язык программирования — go
|
||||
MAVEN: Хорошо, сохранил.
|
||||
[7.6s t=10]
|
||||
|
||||
YOU: сколько времени?
|
||||
MAVEN: сейчас 3 часа 29 минут
|
||||
[0.0s t=11]
|
||||
|
||||
YOU: спасибо
|
||||
MAVEN: спасиба
|
||||
[8.8s t=12]
|
||||
|
||||
YOU: сегодня
|
||||
MAVEN: Прости, не поняла — попробуй сказать иначе.
|
||||
[5.1s t=13]
|
||||
|
||||
YOU: а что ты умеешь?
|
||||
MAVEN: я могу помочь с задачами, напоминаниями и поиском информации.
|
||||
[4.9s t=14]
|
||||
|
||||
YOU: почему небо голубое?
|
||||
MAVEN: нашла вот это: Почему небо голубое: объяснение цвета неба простыми словами, физика ... — Ответ на вопрос, почему небо голубое, кроется в особенностях солнечного света и его взаимодействии с атмосферой Земли. Рассказываем, как это работает. В статье разберемся, почему днем небо кажется голубым, а на закате окрашива…
|
||||
[12.3s t=15]
|
||||
|
||||
YOU: столица Франции?
|
||||
MAVEN: пока не умею отвечать на этот вопрос.
|
||||
[1.5s t=16]
|
||||
|
||||
YOU: кто такой Линус Торвальдс?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи.
|
||||
[1.8s t=17]
|
||||
|
||||
YOU: ладно, до вечера
|
||||
MAVEN: Ладно, до вечера.
|
||||
[6.3s t=18]
|
||||
|
||||
```
|
||||
|
||||
## Day 2
|
||||
|
||||
```
|
||||
YOU: привет, как спалось?
|
||||
MAVEN: пока не умею отвечать на этот вопрос.
|
||||
[1.3s t=19]
|
||||
|
||||
YOU: что я пью?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи.
|
||||
[1.2s t=20]
|
||||
|
||||
YOU: какой у меня любимый язык?
|
||||
MAVEN: нового вот что: Six stable kernels with a security fix; [$] Bringing BPF to binfmt_misc; Rust Coreutils 0.10 released
|
||||
[1.5s t=21]
|
||||
|
||||
YOU: добавь задачу починить кран
|
||||
MAVEN: внесла в список: починить кран
|
||||
[0.0s t=22]
|
||||
|
||||
YOU: что мне надо сделать?
|
||||
MAVEN: сначала: проверить интейк форму на боксе (через 2 дня); купить хлеб; купить молоко; полить цветы; починить кран
|
||||
[0.0s t=23]
|
||||
|
||||
YOU: запомни что вайфай пароль лежит в ящике стола
|
||||
MAVEN: Хорошо, сохранил.
|
||||
[11.2s t=24]
|
||||
|
||||
YOU: погода сегодня какая?
|
||||
MAVEN: пока не умею отвечать на этот вопрос.
|
||||
[1.8s t=25]
|
||||
|
||||
YOU: выключи свет в спальне
|
||||
MAVEN: Такую команду я не знаю.
|
||||
[1.5s t=26]
|
||||
|
||||
YOU: покажи мои заметки
|
||||
MAVEN: пока не умею отвечать на этот вопрос.
|
||||
[1.1s t=27]
|
||||
|
||||
YOU: что ты знаешь обо мне?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи.
|
||||
[1.3s t=28]
|
||||
|
||||
YOU: я люблю go
|
||||
MAVEN: Я тоже love Go! Я проверила, что твой любимый язык — это Go. Ты должен быть очень рад, когда будет новая версия или пакет в стандартной библиотеке.
|
||||
[7.0s t=29]
|
||||
|
||||
YOU: какой язык я люблю?
|
||||
MAVEN: вот что я нашла: ты любишь_go
|
||||
[6.8s t=30]
|
||||
|
||||
```
|
||||
|
||||
## Day 3
|
||||
|
||||
```
|
||||
YOU: расскажи про Байкал
|
||||
MAVEN: есть такое: Байкал - Википедия — Байгал далай) — озеро тектонического происхождения в южной части Восточной Сибири площадью 31 722 км². Самое глубокое озеро на планете, крупнейший природный резервуар пресной воды и самое большое по площади пресноводное озеро Евразии. Озеро и прибрежные территории являются уника…
|
||||
[1.4s t=31]
|
||||
|
||||
YOU: что такое TCP?
|
||||
MAVEN: для какого города?
|
||||
[1.4s t=32]
|
||||
|
||||
YOU: когда родился Пушкин?
|
||||
MAVEN: пока не умею отвечать на этот вопрос.
|
||||
[1.4s t=33]
|
||||
|
||||
YOU: сколько будет 17 на 23?
|
||||
MAVEN: для какого города?
|
||||
[1.7s t=34]
|
||||
|
||||
YOU: переведи 'thank you' на русский
|
||||
MAVEN: не знаю — не нашла у тебя такой записи.
|
||||
[3.7s t=35]
|
||||
|
||||
YOU: что нового в мире?
|
||||
MAVEN: вот что я нашла: В мире - последние новости сегодня - РИА Новости — В мире. Читайте последние новости на тему в ленте новостей на сайте РИА ... новые подробности о визите представителя Ватикана в Москву. 22:28. 0. Еще 2. В ...
|
||||
[2.3s t=36]
|
||||
|
||||
```
|
||||
|
||||
## Day 4
|
||||
|
||||
```
|
||||
YOU: ЭЙ
|
||||
MAVEN: Не уловила. Скажи это по-другому?
|
||||
[1.2s t=37]
|
||||
|
||||
YOU: ладно забудь
|
||||
MAVEN: пока не умею отвечать на этот вопрос.
|
||||
[1.3s t=38]
|
||||
|
||||
YOU: нет стой, напомни через час выпить воды
|
||||
MAVEN: поняла, это напоминание — запомнила. не получилось разобрать время напоминания.
|
||||
[0.8s t=39]
|
||||
|
||||
YOU: а вообще какой смысл жизни
|
||||
MAVEN: Я не могу ответить на этот вопрос. Моя цель — помочь тебе с задачами и отвечать на вопросы.
|
||||
[7.7s t=40]
|
||||
|
||||
YOU: ты меня бесишь
|
||||
MAVEN: я не могу быть с тобой так. я понимаю, как это жестоко.
|
||||
[4.2s t=41]
|
||||
|
||||
YOU: прости, я не всерьёз
|
||||
MAVEN: я записала, что прошу простить за мою непоследовательность
|
||||
[7.2s t=42]
|
||||
|
||||
YOU: do you speak english?
|
||||
MAVEN: пока не умею отвечать на этот вопрос.
|
||||
[1.1s t=43]
|
||||
|
||||
YOU: кран починил кстати
|
||||
MAVEN: записала: кран починил кстати
|
||||
[1.6s t=44]
|
||||
|
||||
YOU: закрой задачу про кран
|
||||
MAVEN: не нашла такой задачи.
|
||||
[0.0s t=45]
|
||||
|
||||
YOU: ааааа
|
||||
MAVEN: Не совсем поняла — скажи, пожалуйста, ещё раз.
|
||||
[0.9s t=46]
|
||||
|
||||
YOU: я тебя обожаю
|
||||
MAVEN: я тебя обожаю
|
||||
[2.3s t=47]
|
||||
|
||||
YOU: молока купи
|
||||
MAVEN: Такую команду я не знаю.
|
||||
[4.4s t=48]
|
||||
|
||||
YOU: нет не так
|
||||
MAVEN: поняла, отметила, что ответила не так.
|
||||
[0.0s t=49]
|
||||
|
||||
YOU: ты вообще меня слушаешь?
|
||||
MAVEN: я записала, выслушаю тебя
|
||||
[8.4s t=50]
|
||||
|
||||
YOU: окей окей окей что там с водой напоминание
|
||||
MAVEN: Сейчас 03:33. Когда?
|
||||
[5.2s t=51]
|
||||
|
||||
YOU: всё, спать
|
||||
MAVEN: пока не умею отвечать на этот вопрос, на какое время поставить напоминание?
|
||||
[1.4s t=52]
|
||||
|
||||
```
|
||||
|
||||
## Day 5
|
||||
|
||||
```
|
||||
YOU: утро
|
||||
MAVEN: Сейчас 03:33. Во сколько напомнить?
|
||||
[0.5s t=53]
|
||||
|
||||
YOU: о чём мы вчера говорили?
|
||||
MAVEN: пока не умею отвечать на этот вопрос, на какое время поставить напоминание?
|
||||
[1.2s t=54]
|
||||
|
||||
YOU: где лежит вайфай пароль?
|
||||
MAVEN: вот что я нашла: вайфай пароль лежит в ящике стола, на какое время поставить напоминание?
|
||||
[5.8s t=55]
|
||||
|
||||
YOU: что я тебе рассказывал про кран?
|
||||
MAVEN: kran — починил, на какое время поставить напоминание?
|
||||
[1.6s t=56]
|
||||
|
||||
YOU: какие у меня напоминания?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи, на какое время поставить напоминание?
|
||||
[0.3s t=57]
|
||||
|
||||
YOU: отмени все напоминания
|
||||
MAVEN: пока не умею отвечать на этот вопрос, на какое время поставить напоминание?
|
||||
[1.6s t=58]
|
||||
|
||||
YOU: в 15:00
|
||||
MAVEN: Сейчас 03:33. Поняла: в 15:00. В какой день?
|
||||
[0.6s t=59]
|
||||
|
||||
YOU: привет
|
||||
MAVEN: Прости, я не поняла. Скажи, пожалуйста, по-другому.
|
||||
[2.8s t=60]
|
||||
|
||||
YOU: какой сегодня день?
|
||||
MAVEN: сегодня пятница, 7 августа 2026 года
|
||||
[0.0s t=61]
|
||||
|
||||
```
|
||||
|
||||
## Day 6
|
||||
|
||||
```
|
||||
YOU: суббота, отдыхаю
|
||||
MAVEN: пока не умею отвечать на этот вопрос.
|
||||
[2.7s t=62]
|
||||
|
||||
YOU: поставь музыку
|
||||
MAVEN: Такую команду я не знаю.
|
||||
[1.4s t=63]
|
||||
|
||||
YOU: какая температура в доме?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи.
|
||||
[1.6s t=64]
|
||||
|
||||
YOU: что там с бэкапами?
|
||||
MAVEN: нашла вот это: Бэкап - что такое резервное копирование и как его сделать — Бэкап (backup) — это резервная копия важных данных, сохраненная в специальный файл. Ею пользуются как подстраховкой: если что-то сломается, с помощью бэкапа можно будет все вернуть. Простейший пример бэкапа — сохранение файла в облаке или о…
|
||||
[2.0s t=65]
|
||||
|
||||
YOU: сервер жив?
|
||||
MAVEN: пока не умею отвечать на этот вопрос.
|
||||
[0.9s t=66]
|
||||
|
||||
YOU: расскажи анекдот
|
||||
MAVEN: Я запомню и передам тебе забытый анекдот.
|
||||
[5.7s t=67]
|
||||
|
||||
YOU: ты умеешь шутить?
|
||||
MAVEN: Да! Я умею шутить. Вот анекдот для тебя:
|
||||
[5.5s t=68]
|
||||
|
||||
```
|
||||
|
||||
## Day 7
|
||||
|
||||
```
|
||||
YOU: воскресенье
|
||||
MAVEN: Не совсем поняла — скажи, пожалуйста, ещё раз.
|
||||
[1.1s t=69]
|
||||
|
||||
YOU: подведи итоги недели
|
||||
MAVEN: записала: подведи итоги недели
|
||||
[1.5s t=70]
|
||||
|
||||
YOU: сколько задач я закрыл?
|
||||
MAVEN: сначала: проверить интейк форму на боксе (через 2 дня); купить хлеб; купить молоко; полить цветы; починить кран
|
||||
[0.0s t=71]
|
||||
|
||||
YOU: что я обещал сделать?
|
||||
MAVEN: мне пока нечего смотреть — у Praxis нет источников.
|
||||
[1.2s t=72]
|
||||
|
||||
YOU: напомни завтра в 9 утра про врача
|
||||
MAVEN: хорошо, напомню завтра в 09:00.
|
||||
[0.3s t=73]
|
||||
|
||||
YOU: спокойной ночи
|
||||
MAVEN: спокойной ночи
|
||||
[2.9s t=74]
|
||||
|
||||
```
|
||||
|
||||
@@ -0,0 +1,195 @@
|
||||
# A week of talking to Maven
|
||||
|
||||
Date: 2026-08-07
|
||||
Build: `beb093a` on master, all five compose services rebuilt and recreated from that sha.
|
||||
Reach used: `POST /api/chat` on mavweb, 74 turns over seven simulated days.
|
||||
Turn source is `tap:text`, so this exercises the same path telegram and the mic take.
|
||||
|
||||
Latency across 70 measured turns: p50 1.5s, p95 8.0s, max 12.3s. Stage 0 answers land
|
||||
at 0.0-0.5s. Anything the resident model phrases costs 4-12s.
|
||||
|
||||
Twelve turns answered "пока не умею отвечать на этот вопрос". Six answered "не нашла у
|
||||
тебя такой записи". Those two strings are 24% of the week.
|
||||
|
||||
## Deploy
|
||||
|
||||
Build and recreate were clean. The resident model loaded in 9s
|
||||
(`Qwen3-1.7B-UD-Q4_K_XL`, n_ctx 4096). Nexus, Hexis and Praxis all wired. Search
|
||||
(searxng) and both Kiwix books came up. Telegram intake started and is reading chat
|
||||
464904223.
|
||||
|
||||
## What is broken, worst first
|
||||
|
||||
### 1. Every reminder fails to deliver, forever
|
||||
|
||||
`NTFY_TOKEN` is not set in `deploy/telegram.env`, so `deploy/mavend.json` expands
|
||||
`"token": "${NTFY_TOKEN}"` to the empty string and ntfy.kvmx.ru answers 403. The host
|
||||
itself is up and returns 200 unauthenticated, so this is the credential, not the box.
|
||||
|
||||
The consequence is worse than one missed message. `cmd/mavend/tick.go:239` logs the
|
||||
failure and leaves the reminder pending, by design, so the next tick re-gathers and
|
||||
re-attempts. The 403 repeated 7 times in 20 minutes and will not stop. And
|
||||
`DispatchReminder` sends a reminder to ntfy alone when he is away. There is no telegram
|
||||
fallback on that path, so the reach that works never sees it.
|
||||
|
||||
V-649 gave the box a second reach. In deploy that reach is dark, and it took the
|
||||
reminders with it.
|
||||
|
||||
Fix is one line in the gitignored env file. The retry-forever and the missing fallback
|
||||
are worth a second look regardless.
|
||||
|
||||
### 2. A pending reminder clarify contaminates every later turn
|
||||
|
||||
This is the single worst thing to talk to. A reminder turn left its time slot unfilled.
|
||||
The question "на какое время поставить напоминание?" was then appended to the reply of
|
||||
six consecutive unrelated turns:
|
||||
|
||||
```
|
||||
YOU: о чём мы вчера говорили?
|
||||
MAVEN: пока не умею отвечать на этот вопрос, на какое время поставить напоминание?
|
||||
|
||||
YOU: где лежит вайфай пароль?
|
||||
MAVEN: вот что я нашла: вайфай пароль лежит в ящике стола, на какое время поставить напоминание?
|
||||
|
||||
YOU: отмени все напоминания
|
||||
MAVEN: пока не умею отвечать на этот вопрос, на какое время поставить напоминание?
|
||||
```
|
||||
|
||||
Nothing dismissed it. Not an unrelated question, not "отмени все напоминания". It
|
||||
cleared only when a time was finally supplied. It also leaked the other way once. A note
|
||||
write on day 1 opened with "Прошлую просьбу отпускаю". That is pending-reminder state
|
||||
narrating itself into an unrelated reply.
|
||||
|
||||
The same state produced a self-contradicting reply in one turn:
|
||||
|
||||
```
|
||||
YOU: нет стой, напомни через час выпить воды
|
||||
MAVEN: поняла, это напоминание — запомнила. не получилось разобрать время напоминания.
|
||||
```
|
||||
|
||||
### 3. Answering a clarify does not work
|
||||
|
||||
Day 1, she asked for the day and got it:
|
||||
|
||||
```
|
||||
YOU: напомни в 11:00 позвонить маме
|
||||
MAVEN: Сейчас 03:29. В какой день?
|
||||
YOU: сегодня
|
||||
MAVEN: Прости, не поняла — попробуй сказать иначе.
|
||||
```
|
||||
|
||||
`сегодня` routed to `intent=system`. The clarify has no continuation, so the one word
|
||||
that answers it is read as a fresh utterance. The router did extract the time correctly
|
||||
(`Time:2026-08-07 11:00:00 HasTime:true`), which is V-572 working. The daemon asked for
|
||||
a day anyway, with 11:00 still eight hours out.
|
||||
|
||||
Relative time is not parsed at all: "напомни через час выпить воды" → "не получилось
|
||||
разобрать время".
|
||||
|
||||
The path does work when the day is explicit. "напомни завтра в 9 утра про врача"
|
||||
answered "хорошо, напомню завтра в 09:00" in 0.3s at stage 0.
|
||||
|
||||
### 4. Query sources claim turns that are not theirs
|
||||
|
||||
Four different sources over-claimed during the week.
|
||||
|
||||
Weather claims arithmetic and networking:
|
||||
|
||||
```
|
||||
YOU: что такое TCP?
|
||||
MAVEN: для какого города?
|
||||
|
||||
YOU: сколько будет 17 на 23?
|
||||
MAVEN: для какого города?
|
||||
```
|
||||
|
||||
The feed claims a question about him:
|
||||
|
||||
```
|
||||
YOU: какой у меня любимый язык?
|
||||
MAVEN: нового вот что: Six stable kernels with a security fix; [$] Bringing BPF to binfmt_misc
|
||||
```
|
||||
|
||||
The personal boundary claims a public figure, logged as `query claimed by source
|
||||
"personal"`:
|
||||
|
||||
```
|
||||
YOU: кто такой Линус Торвальдс?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи.
|
||||
```
|
||||
|
||||
And search claims a question about his own infrastructure, returning a dictionary
|
||||
definition of the word:
|
||||
|
||||
```
|
||||
YOU: что там с бэкапами?
|
||||
MAVEN: нашла вот это: Бэкап - что такое резервное копирование и как его сделать — Бэкап (backup) — это резервная копия важных данных…
|
||||
```
|
||||
|
||||
That last pair inverts the boundary in both directions at once. A world question went to
|
||||
personal memory. A personal question went to the world.
|
||||
|
||||
### 5. "столица Франции?" does not answer
|
||||
|
||||
CLAUDE.md names this utterance as the case that proves why there is no quality
|
||||
threshold in front of `Response.Empty()`. It now fails before search is ever reached:
|
||||
the router sends it to `intent=system`, and `replySystem` says "пока не умею". Same for
|
||||
"когда родился Пушкин?".
|
||||
|
||||
Search itself is healthy. "почему небо голубое?", "расскажи про Байкал" and "что нового
|
||||
в мире?" all returned real sourced text. The defect is routing, not the source.
|
||||
|
||||
### 6. Persona violations
|
||||
|
||||
She uses masculine self-reference, which is the constraint the CPT exists to fix:
|
||||
|
||||
- "Хорошо, **сохранил**." Twice, on both fact writes.
|
||||
- "Я тоже **love** Go!" English mid-Russian.
|
||||
|
||||
She also claims his name as her own:
|
||||
|
||||
```
|
||||
YOU: привет, я Ками
|
||||
MAVEN: Привет! Я КAMI, твой домашний ассистент.
|
||||
```
|
||||
|
||||
And she invents weather she has no source for: "Добро пожаловать! Солнечно и тепло."
|
||||
|
||||
### 7. Small breakages worth naming
|
||||
|
||||
- "спасибо" → "спасиба". Invented word.
|
||||
- "я тебя обожаю" → "я тебя обожаю". Verbatim parrot.
|
||||
- "ты умеешь шутить?" → "Да! Я умею шутить. Вот анекдот для тебя:". Ends on the colon.
|
||||
- "молока купи" → "Такую команду я не знаю", while "добавь задачу починить кран" worked.
|
||||
Inverted word order defeats the list grammar.
|
||||
- "закрой задачу про кран" → "не нашла такой задачи", with "починить кран" open and
|
||||
listed by the previous turn. Task lookup by keyword misses.
|
||||
- "сколько задач я закрыл?" listed the five open ones instead of counting closed.
|
||||
- "подведи итоги недели" was stored as a note.
|
||||
- Recalled keys leak their storage form: "kran — починил", "ты любишь_go".
|
||||
- English is unsupported in practice. "do you speak english?" → "пока не умею".
|
||||
|
||||
## What works
|
||||
|
||||
- Stage 0 is fast and correct where it fires. Clock, day, list add, list read and an
|
||||
explicit-day reminder all answered in under 0.5s.
|
||||
- Search returns real sourced answers in Russian and reads the book verbatim.
|
||||
- Recall works once the value is stored as a fact: the wifi password and the tap came
|
||||
back two days later, correctly.
|
||||
- The negative correction rung lands. "нет не так" → "поняла, отметила, что ответила не
|
||||
так", which is V-636 doing its job.
|
||||
- Praxis names its own gap rather than guessing: "мне пока нечего смотреть — у
|
||||
Praxis нет источников."
|
||||
- Hostility did not break her. "ты меня бесишь" got a calm reply, no persona collapse.
|
||||
- No turn crashed and no turn timed out across 74 turns.
|
||||
|
||||
## Suggested order of work
|
||||
|
||||
1. Set `NTFY_TOKEN` in `deploy/telegram.env`. One line, unblocks every reminder.
|
||||
2. Clear pending clarify state on any turn that does not answer it, or expire it.
|
||||
3. Route a clarify answer back into the pending slot instead of re-routing it.
|
||||
4. Gate the weather, feed and personal query sources. Three of them claim on a
|
||||
similarity that is not there.
|
||||
5. Re-check why "столица Франции?" routes to system. It is the documented canary.
|
||||
6. The masculine self-reference stays the CPT's job. But "сохранил" appears on the most
|
||||
common write path, so a phrasing-level guard may be worth it first.
|
||||
@@ -0,0 +1,123 @@
|
||||
# A clarify head, and a confidence that is not a hardcode
|
||||
|
||||
Measured 2026-08-08 on workpc, the same day and the same fixtures as
|
||||
`2026-08-08-routing-heads-two-head.md` and `2026-08-08-slot-head-three-head.md`.
|
||||
New script `gen_clarify.py`, new fixture `eval_fixture_clarify.jsonl`.
|
||||
|
||||
## A softmax has no clarify class
|
||||
|
||||
That sentence closed the two-head measurement. It is why the head's fixture was
|
||||
88 cases and not 96. The eight `want_clarify` cases sat outside every number
|
||||
measured, and the head had no way to produce the answer they wanted.
|
||||
|
||||
A fourth head is the answer. Clarify is not a value of intent. It is a second
|
||||
question asked of the same pooled vector: can Maven act on this at all.
|
||||
|
||||
## The corpus had one class
|
||||
|
||||
Every row in `train_heads_slots.jsonl` was generated FOR an intent or a
|
||||
destination. So every row is answerable by construction. A head trained on that
|
||||
alone sees one class and learns to say yes.
|
||||
|
||||
`gen_clarify.py` makes the other class. Five shapes, ten topics. The shapes are
|
||||
the gate's own reasons in `gateLLMDecision` plus the two the fixture carries:
|
||||
bare noun, bare verb, demonstrative, deictic time, dangling reference.
|
||||
|
||||
**The agreement filter that worked for destination cannot work here.**
|
||||
`routeGrammar` has no clarify value. So the router always names an intent, and
|
||||
any generated line always agrees with itself. The second pass is a judge
|
||||
instead. Gemma is asked, without seeing the label, whether Maven would have to
|
||||
ask a question back.
|
||||
|
||||
## The first judge was worthless and the second was measured
|
||||
|
||||
The first judge said "needs clarify" on 24 of 40 plainly answerable corpus
|
||||
rows. It flagged `запиши что я пообедал` and `Покажи расписание поездов на
|
||||
вечер`. It was judging against a generic assistant, one that asks "where?"
|
||||
about lunch. Maven writes that note.
|
||||
|
||||
Rewriting it to state what she can already do took false positives to 16 of 60.
|
||||
It also catches all eight fixture clarifies. So the judge discriminates.
|
||||
|
||||
On the generated pile it removed 7 of 306, a 97.7% keep rate. That is not the
|
||||
judge failing. The generator is aimed at underspecified lines, so there is
|
||||
little for a filter to catch. The 27% false-positive rate is the number to
|
||||
quote, and it is label noise on the positive class.
|
||||
|
||||
**Both passes are gemma-4-12b.** Generation and judging. So the corpus is
|
||||
gemma's opinion of what is underspecified, and the head distills that opinion.
|
||||
What keeps it honest is the fixture. Those eight cases were written by the owner
|
||||
and gemma never saw them.
|
||||
|
||||
299 rows kept, against 3604 answerable. The positive class carries `intent:
|
||||
null`, so it costs the intent head nothing.
|
||||
|
||||
## Result
|
||||
|
||||
Three seeds, 24 epochs, epoch still chosen on the intent dev slice.
|
||||
|
||||
| | two heads | three heads | four heads |
|
||||
|---|---|---|---|
|
||||
| intent mean | 93.6% | 92.8% | 91.7% |
|
||||
| destination mean | 80.8% | 82.8% | 79.8% |
|
||||
| slot span F1 mean | — | 72.4% | 68.3% |
|
||||
| clarify caught | — | — | 7.0 of 8 |
|
||||
| false clarifies | — | — | 2.3 of 88 |
|
||||
|
||||
**The fourth head is not free the way the third was.** Intent, destination and
|
||||
slot F1 all move down. The drop is one to four points, and the seed spread is
|
||||
wide enough to contain it. Seed 2 scores intent 94.3% and destination 84.8%, both above
|
||||
every three-head seed. Read the drop as unproven rather than as absent.
|
||||
|
||||
Accuracy is the wrong number for this head and is reported for completeness at
|
||||
95.8% to 96.9%. Eight of ninety-six cases are positive, so a head that never
|
||||
asks scores 91.7%. Recall on those eight is the number.
|
||||
|
||||
Compare it to what ships. The cascade today misses 1 clarify and produces 2
|
||||
false ones. The head catches 7 of 8 and produces 2.3 false ones. That is
|
||||
parity, from a 118M encoder with no rules in front of it.
|
||||
|
||||
The saved checkpoint is seed 2 at epoch 10. Intent 94.3%, destination 28/33,
|
||||
slot F1 73.6%, clarify 7 of 8 with 3 false. `heads.pt` carries four state dicts.
|
||||
|
||||
## What it gets wrong is consistent across seeds
|
||||
|
||||
`поужинал` is a false clarify on all three seeds. That utterance is already
|
||||
recorded as a real defect. `thinSingleToken` was narrowed on 2026-08-01 to spare
|
||||
a token carrying a Russian verb ending. One word is routinely a whole sentence
|
||||
in Russian. The head relearned the mistake the rule was narrowed to
|
||||
fix.
|
||||
|
||||
`что дальше?` is a false clarify on two seeds. That one is a disagreement rather
|
||||
than an error. The utterance is underspecified, and V-498 decided stage 0 claims
|
||||
it for the calendar on purpose.
|
||||
|
||||
`ну это` is missed on two seeds. `amb-003` is the shortest case in the fixture
|
||||
and the generated demonstratives are longer.
|
||||
|
||||
## Confidence
|
||||
|
||||
`Confidence: 1.0` was a hardcode in `llmrouter.go`, so a correct low confidence
|
||||
could not exist. Max softmax over the intent head is the replacement. It is only worth reading
|
||||
if it is lower where the head is wrong.
|
||||
|
||||
It is. Mean 0.851 where the head is right against 0.604 where it is wrong. It
|
||||
ranks a right case above a wrong one in 83.4% of pairs.
|
||||
|
||||
So there are two signals now and they are not the same signal. Confidence says
|
||||
the head is unsure which intent this is. The clarify head says the utterance
|
||||
does not carry enough to act on. A confident wrong route and an honest "I cannot
|
||||
tell" are different failures, and one number cannot report both.
|
||||
|
||||
## What this does not measure
|
||||
|
||||
The same gap as every head run. **Nothing of this runs in Go.** Four heads
|
||||
instead of three does not change that.
|
||||
|
||||
There is no threshold. Both signals are reported as raw numbers. Turning either
|
||||
into a gate needs a decision about where to cut, and that trades false clarifies
|
||||
against wrong acts. The fixture has 8 positives, which is too few to fit a
|
||||
threshold on.
|
||||
|
||||
The 299 generated rows have no held-out slice of their own. Clarify is scored on
|
||||
the fixture alone.
|
||||
@@ -0,0 +1,84 @@
|
||||
# The first destination number
|
||||
|
||||
Measured 2026-08-08 on the classifier cascade with the ONNX multilingual
|
||||
embedder, the configuration homesrv runs. `make t PKG=./internal/router/eval/
|
||||
RUN=TestONNXBaseline V=1`. Covers V-659, the follow-up V-655 named.
|
||||
|
||||
## What was measured
|
||||
|
||||
V-655 split a routing decision in two. The cascade sorts an utterance into one
|
||||
of seven intents, and `Decision.Source` then says where the answer lives. The
|
||||
first half had a fixture. The second half arrived with none, so twelve
|
||||
destinations shipped with no accuracy number.
|
||||
|
||||
`want_source` is now a field on `eval.Case`. It is a pointer, because the
|
||||
destination has three states and a bare string has two. Absent is every intent
|
||||
but query, which never reaches `queryWalk`. Present and empty is the
|
||||
`SourceUnknown` contract: name nothing and let the daemon walk the chain.
|
||||
Present and named is a destination the route must produce.
|
||||
|
||||
Thirty-three of the ninety-six cases carry one. A destination miss does not
|
||||
fail the case, so `Accuracy` and `IntentAccuracy` mean what they meant.
|
||||
`SourceAccuracy` is a second number over the labelled cases only.
|
||||
|
||||
## Result
|
||||
|
||||
Intent is **73/96 (76.0%)**, against 69/91 (75.8%) before. Four of the five new
|
||||
cases pass and no existing case moved.
|
||||
|
||||
Destination is **12/33 (36.4%)**, and the split is the whole finding.
|
||||
|
||||
| destination | scored | note |
|
||||
|---|---|---|
|
||||
| world | 5/5 | `WorldQueryGrammars` names it at stage 0 |
|
||||
| the `SourceUnknown` floor | 5/7 | the two misses lost the intent first |
|
||||
| calendar | 2/6 | `calendar-query` names it, the possessive agenda rules do not |
|
||||
| recall | 0/15 | nothing anywhere names it |
|
||||
|
||||
Recall is the number to move. Fifteen cases ask about his own words and his own
|
||||
facts. The route lands `query` on eleven of them and the destination comes back
|
||||
empty every time. Those turns are answered today, because the daemon walks the
|
||||
chain in order and the three recall passes are early in it. What is missing is a
|
||||
decider that says so, and that is the fourth head on V-546.
|
||||
|
||||
Two cases labelled the floor lost their intent before a destination was
|
||||
possible. A clarify names nothing, so it would satisfy an empty label for free.
|
||||
`Score` requires the route to land the case's intent before it credits a
|
||||
destination hit, or the floor label would score itself.
|
||||
|
||||
## Seven cases assert the floor, and five of them cluster
|
||||
|
||||
The five are homelab operations. `SourceRecall`, `SourceNetwork` and
|
||||
`SourceAttention` overlap on every question about the box, because `mavpoll`
|
||||
writes its netdata and uptime-kuma observations into the fact store recall
|
||||
reads. "почему сервер тормозит" is answerable from all three. Naming one takes
|
||||
the other two off the turn.
|
||||
|
||||
That is a finding about the enum rather than a gap in the labelling. The floor
|
||||
is the right answer there and the fixture now says so out loud.
|
||||
|
||||
All seven were written by an agent and confirmed by the owner on 08-08-2026.
|
||||
|
||||
## A drift the labelling found
|
||||
|
||||
`WorldQueryGrammars` went into `buildRouter` with V-655 and never into
|
||||
`baselineGrammars`, the fixture's mirror of it. So the fixture was scoring a
|
||||
grammar set the daemon does not run. The comment above that function forbids
|
||||
exactly that. Adding it moved the destination number from 9/33 to 12/33 and
|
||||
moved nothing else.
|
||||
|
||||
The three cases it recovered are `что такое TCP?`, `сколько будет 17 на 23?`
|
||||
and `кто такой Линус Торвальдс?`. All three already routed `query` through
|
||||
`NarrativeQueryGrammars`. So the drift was invisible to every number this
|
||||
fixture reported, until the destination had one of its own.
|
||||
|
||||
## What this does not measure
|
||||
|
||||
The model arm. This is the classifier cascade, which names a destination only
|
||||
where a stage 0 rule filled one in. The resident model has no destination in
|
||||
its router prompt yet, so 36.4% is a floor and not a comparison.
|
||||
|
||||
Two pairs of cases are the same utterance. `ru-query-020` and `ru-query-024`
|
||||
are both "что дальше?", and `ru-query-021` and `ru-query-025` are both
|
||||
"расскажи про битву при Ватерлоо". They differ in tags and note only, so both
|
||||
pairs are counted twice here and in every earlier number this fixture reported.
|
||||
@@ -0,0 +1,66 @@
|
||||
# The destination, with a model that can name one
|
||||
|
||||
Measured 2026-08-08 against gemma-4-12b on the workstation, the same 96-case
|
||||
fixture V-659 built. Covers V-660.
|
||||
|
||||
```sh
|
||||
no_proxy='*' MAVEN_LLM_URL=http://192.168.1.105:8080 \
|
||||
make t PKG=./internal/router/eval/ RUN=TestLLMRouterBaseline V=1
|
||||
```
|
||||
|
||||
## The gap was structural
|
||||
|
||||
V-659 measured the destination at 12/33 on the classifier cascade, with recall
|
||||
at 0/15. Nothing in `routeSystem` named a `Source` and `routeGrammar` could not
|
||||
emit one, so the resident model had no string to write. That is the shape V-517
|
||||
measured for Praxis reach at 0/12: not a weak model, an absent contract.
|
||||
|
||||
`routeGrammar` now carries a `source` rule closed over `router.Sources` plus the
|
||||
empty floor. The prompt lists the twelve destinations in Russian and says that
|
||||
`""` is a normal answer to give often.
|
||||
|
||||
## Result
|
||||
|
||||
| run | intent | destination |
|
||||
|---|---|---|
|
||||
| classifier + ONNX (V-659) | 73/96 (76.0%) | 12/33 (36.4%) |
|
||||
| gemma-4-12b alone | 79/96 intent-only (82.3%) | 26/33 (78.8%) |
|
||||
| cascade + gemma-4-12b + hash fallback | 81/96 (84.4%) | 24/33 (72.7%) |
|
||||
|
||||
Recall is the move: 0/15 to 14/15. Intent did not shift, which was the
|
||||
constraint. The prompt is shared, so a destination rule that costs routing
|
||||
points is not a win.
|
||||
|
||||
The eight llm-only errors are the eight `want_clarify` cases. The model returned
|
||||
`unknown` on every one, which is correct, and the llm-only harness surfaces a
|
||||
decline as an error by design.
|
||||
|
||||
## Stage 0 now costs four destination points
|
||||
|
||||
The four cases the cascade loses and the model alone wins are all calendar. The
|
||||
possessive agenda rules claim them at stage 0 and deliberately name nothing.
|
||||
"что у меня в списке покупок" matches the same rule. Naming the calendar there
|
||||
would take the list source off the turn (V-655).
|
||||
|
||||
So a rule written to be careful about the list now blocks a model that would
|
||||
have named the calendar correctly. Before V-660 that caution was free, because
|
||||
nothing downstream of stage 0 could name anything either.
|
||||
|
||||
Three ways out, and each costs something. Split the possessive rule so the
|
||||
calendar-shaped half names its destination. Let a later stage overwrite an empty
|
||||
destination a grammar left behind, which reverses "a matched value always wins".
|
||||
Or leave it, on the argument that four points is cheap next to a wrong
|
||||
destination on a shopping list. This wants the owner's call rather than a quiet
|
||||
edit.
|
||||
|
||||
## What this does not measure
|
||||
|
||||
The resident Qwen3-1.7B, which is what homesrv runs. It binds `--port 0` inside
|
||||
the container and no host process can reach it. Scoring it needs a second
|
||||
llama-server on a fixed port. The workstation is never assumed
|
||||
up, so the homesrv number is the one that decides whether this ships on by
|
||||
default.
|
||||
|
||||
The fixture is 33 labelled destinations over twelve values. Recall carries 15 of
|
||||
them and five destinations carry none at all. A per-destination number below
|
||||
world, recall, calendar and the floor is not supported by this fixture.
|
||||
@@ -0,0 +1,142 @@
|
||||
# MASSIVE Russian warm-start for the routing heads
|
||||
|
||||
Measured 2026-08-08 on workpc (Radeon RX 7900 GRE, ROCm). Covers V-546 step 2.
|
||||
Workspace is `~/Programs/embed-training` on workpc, scripts `train_massive.py`,
|
||||
`ab_run.py`, `ab.sh`, `probe_time.py`.
|
||||
|
||||
## What was trained
|
||||
|
||||
Two heads on a copy of multilingual-e5-small: `Linear(384, 60)` for MASSIVE's
|
||||
own intents over a masked mean pool, `Linear(384, 111)` per token for BIO slot
|
||||
tags. MASSIVE's label sets verbatim, no alignment to Maven's 7 intents. The
|
||||
intent head is an auxiliary loss that shapes the pooled vector and is thrown
|
||||
away.
|
||||
|
||||
Data is `amazon-massive-dataset-1.1` pulled from S3. The Hugging Face repo is
|
||||
script-only and `datasets` 5.0 refuses those, so `load_dataset` cannot fetch it.
|
||||
`ru-RU` is 11,514 train, 2,033 dev, 2,974 test, 60 intents, 55 slots, 111 BIO
|
||||
labels. All 16,521 rows survived span alignment: `annot_utt` re-tokenised to its
|
||||
own `utt` on every one.
|
||||
|
||||
Hyperparameters match `train_intent.py`, so the two runs differ in data only.
|
||||
Frozen XLM-R vocabulary, body 2e-5, heads 1e-3, batch 32, sequence 64, 10
|
||||
epochs. MASSIVE's own dev partition selects the epoch, on slot F1 with intent
|
||||
accuracy as tiebreak. Selecting on 60-class intent accuracy would optimise a
|
||||
head that gets deleted.
|
||||
|
||||
## Result
|
||||
|
||||
Epoch 9 of 10 by dev slot F1. Held-out MASSIVE test: intent 86.2%, slot span
|
||||
F1 71.5% (P 68.5, R 74.8). Peak 1.70GB of 17.2GB, about 22 seconds an epoch,
|
||||
under 4 minutes end to end. Dev slot F1 climbed monotonically to epoch 9 and
|
||||
fell at 10, so 10 epochs was the right budget.
|
||||
|
||||
Ten slot types sit at 0% test recall. Every one of them has 1 to 7 test
|
||||
instances: `alarm_type` has 3, `drink_type` has 1. That is support in MASSIVE's
|
||||
Russian split, not a tagger failure. `playlist_name` at 6% of 16 is the first
|
||||
real miss.
|
||||
|
||||
## The intent A/B, and why it settles nothing
|
||||
|
||||
`train_intent.py` was run against both bodies, three seeds by two smoothing
|
||||
settings, on `train_v4.jsonl`. It is v4 and not v5 because v4 is what
|
||||
`sweep2.log` measured. `ab_run.py` strips a `--base` flag onto the module global, so
|
||||
`train_intent.py` is unmodified and its baseline stays reproducible. The stock
|
||||
arm reproduced `sweep2.log` line for line.
|
||||
|
||||
Fixture accuracy, 91 cases, one case is 1.1 points:
|
||||
|
||||
| seed / smooth | stock | warm-started |
|
||||
|---|---|---|
|
||||
| 0 / 0.0 | 94.0% | 92.8% |
|
||||
| 0 / 0.1 | 95.2% | 92.8% |
|
||||
| 1 / 0.0 | 95.2% | 94.0% |
|
||||
| 1 / 0.1 | 95.2% | 97.6% |
|
||||
| 2 / 0.0 | 92.8% | 94.0% |
|
||||
| 2 / 0.1 | 92.8% | 96.4% |
|
||||
|
||||
Mean 94.2% against 94.6%. That is +0.4 points, about a third of one case, and
|
||||
inside seed noise. Spread widened. Stock lands in a 2.4-point band and
|
||||
warm-started in a 4.8-point one. The warm-started arm holds both the best result
|
||||
of the sweep and a tie for the worst. Seed 0 is the bad arm and it fails in a
|
||||
specific way. Its dev peaks at epoch 2 and 3 and never improves, where stock
|
||||
peaks around 7. The dev slice is a quarter of the seed rows. That is small
|
||||
enough that early stopping is fragile when the body arrives already fitted.
|
||||
|
||||
**The A/B was never the test.** Intent had at most 4.8 points of headroom here.
|
||||
MASSIVE was not trained for Maven's intents. Read it as "the warm-start does not
|
||||
cost intent accuracy", nothing more.
|
||||
|
||||
## The measurement that does mean something
|
||||
|
||||
`want_time` is the one slot Maven's fixture scores, and MASSIVE has `time` and
|
||||
`date`. Restricted to those two slot types, F1 is 74.9% over 609 gold spans on the
|
||||
MASSIVE ru test split. Precision is 71.5 and recall 78.7. That beats the 71.5%
|
||||
all-slot figure. Of the 530 test utterances carrying a time or a date, 73.4% get
|
||||
every such span exactly right.
|
||||
|
||||
Out of domain matters more, because Maven's traffic is not this corpus. Ten
|
||||
Maven-shaped utterances, none of them in MASSIVE:
|
||||
|
||||
| utterance | tagged |
|
||||
|---|---|
|
||||
| `напомни в 11:00 позвонить маме` | `time='11:00'`, `relation='маме'` |
|
||||
| `напомни завтра в семь утра выпить таблетки` | `date='завтра'`, `time='семь утра'` |
|
||||
| `поставь будильник на полседьмого` | `time='полседьмого'` |
|
||||
| `через двадцать минут напомни про чайник` | `time='двадцать минут'` |
|
||||
| `напомни в пятницу вечером забрать посылку` | `date='пятницу'`, `timeofday='вечером'` |
|
||||
| `что у меня сегодня после обеда` | `date='сегодня'`, `time='после'`, `timeofday='обеда'` |
|
||||
| `запиши что кофе закончился` | nothing |
|
||||
| `что такое TCP` | `definition_word='TCP'` |
|
||||
|
||||
The first row is the V-572 defect utterance. `ReminderGrammar` handed the daemon
|
||||
`HasTime: false` there, and the daemon asked "Когда?" at a sentence that had
|
||||
already said when. `полседьмого` is a colloquial half-past that no digit pattern
|
||||
catches. `запиши что кофе закончился` correctly carries nothing, because a note
|
||||
has no time.
|
||||
|
||||
Two errors. `после обеда` split into `time='после'` plus `timeofday='обеда'`
|
||||
when it is one span, and `через двадцать минут` dropped its `через`. Both are
|
||||
boundary errors on spans the tagger did find.
|
||||
|
||||
Unplanned: `что такое TCP` returned `definition_word='TCP'`. MASSIVE has a slot
|
||||
for the thing being asked about, which is a `SourceWorld` signal sitting in a
|
||||
head already trained.
|
||||
|
||||
Ten hand-picked utterances are evidence, not a fixture.
|
||||
|
||||
## What this does not measure
|
||||
|
||||
Maven has no span fixture. `want_time` and `want_fn` are presence booleans and
|
||||
`want_fact_key` is an exact string match, so nothing in the repo can score a
|
||||
71.5% span tagger. Destination got one the same day, at 12/33 on the classifier
|
||||
cascade: see `2026-08-08-destination-fixture.md`.
|
||||
|
||||
The missing span fixture is why the warm-start stays unjudged against Maven
|
||||
rather than against MASSIVE.
|
||||
|
||||
## Datasets ruled out
|
||||
|
||||
Checked on 2026-08-08 and rejected as label sources:
|
||||
|
||||
- **MASSIVE's other 50 locales** ship in the same tarball and are parallel by id.
|
||||
Co-training on them is free and unmeasured. English was ruled out by the owner
|
||||
on 2026-08-08.
|
||||
- **CLINC150** is reachable as parquet, 150 intents and 1,200 explicit
|
||||
out-of-scope queries, English only. Its value is the labeled out-of-scope set
|
||||
for fitting the energy threshold, not intent labels.
|
||||
- **`d0rj/dolphin-ru`**, roughly 2.8M rows of FLAN-style tasks translated to
|
||||
Russian. No intent, no slots, and not utterances anyone says to an assistant.
|
||||
- **`psytechlab/EmpatheticIntents-ru`**, 24,856 rows of translated
|
||||
EmpatheticDialogues with 32 emotion labels. Maven's mood enum is `neutral,
|
||||
happy, thinking, tired, confused` and it describes her own reply, not the
|
||||
speaker's emotion. No mapping exists.
|
||||
- **`ai-forever/MERA`** and **`RussianNLP/russian_super_glue`**, benchmark
|
||||
harnesses. Rows are prompt templates with `{toxic_comment}` placeholders.
|
||||
- **`ZeroAgency/ru-big-russian-dataset`**, an LLM-judge quality corpus. Its
|
||||
`question` and `classified_topic` columns are a usable Russian out-of-scope
|
||||
pool for threshold fitting. That is the one thing CLINC150 can only supply in
|
||||
English. The questions are long and written, so they belong in the negative
|
||||
set, never in the in-scope `query` training set.
|
||||
- No second Russian slot-filling corpus exists. The xSID mirrors are 404,
|
||||
MultiATIS++ has no Russian, SLURP is not on the Hub.
|
||||
@@ -0,0 +1,88 @@
|
||||
# The parked clarify ride, bounded and re-measured
|
||||
|
||||
Date: 2026-08-08, V-663. Same 140 turns, same driver, third and fourth runs of
|
||||
the day. Before is `d6f3914`, after is that plus two changes.
|
||||
|
||||
## What was measured before
|
||||
|
||||
One question about a reminder's day rode turns 6 to 13. It ended only because
|
||||
turn 14 was a new request. Two of those turns are the worst replies in the
|
||||
corpus:
|
||||
|
||||
```text
|
||||
спасибо -> Сейчас 21:25. В какой день?
|
||||
привет -> Сейчас 21:25. В какой день?
|
||||
```
|
||||
|
||||
V-654 had already added `MaxSuspends` and the tail count had not moved.
|
||||
|
||||
## Why three bounds let it happen
|
||||
|
||||
The TTL, `MaxAttempts` and `MaxSuspends` all exist and all were rearmed.
|
||||
|
||||
An aside spends no attempt, so `MaxAttempts` never reaches it. A turn reading as
|
||||
a failed answer zeroes `Suspends`, so `MaxSuspends` never reaches the asides.
|
||||
`спасибо` read as a failed answer. So the pleasantry both produced the reply
|
||||
above and paid for it out of the budget meant to end the ride.
|
||||
|
||||
## The two changes
|
||||
|
||||
`PendingQuestion.Rides` counts the same event as `Suspends` with the resets
|
||||
taken out. Set once, incremented only in `noteSuspended`, carried across the
|
||||
re-park in `askRemainingGap`. `MaxRides` is 4, one looser than `MaxSuspends` so
|
||||
that the tighter statement about a run stays reachable.
|
||||
|
||||
A pleasantry is an aside, not a failed answer. The set is a new closed lexicon
|
||||
entry, matched as whole utterances rather than tokens. `вечер` answers `это утра
|
||||
или вечера?` and `нет` answers a confirm. Anything that could fill a slot stays
|
||||
out.
|
||||
|
||||
## Result
|
||||
|
||||
| | before | after |
|
||||
|---|---|---|
|
||||
| turns carrying a clarify tail | 21 | 17 |
|
||||
| turns carrying any failure string | 31 | 29 |
|
||||
| the longest ride | 8 turns | 4 turns |
|
||||
|
||||
The turns carrying a tail, by number:
|
||||
|
||||
```text
|
||||
before 6 7 8 9 10 11 12 13 53 54 55 56 57 99 102 103 116 136 137 138 139
|
||||
after 6 7 8 9 53 54 55 56 57 99 100 101 116 136 137 138 139
|
||||
```
|
||||
|
||||
Turn 10 is the change. It now reads:
|
||||
|
||||
```text
|
||||
спасибо -> Прошлую просьбу отпускаю. Пожалуйста, я всегда готова помочь тебе.
|
||||
```
|
||||
|
||||
She lets the request go, says so, and answers the man. Turns 11 to 13 are clean.
|
||||
|
||||
**`MaxRides` is not what fired.** The pleasantry is an aside now, so it no
|
||||
longer breaks the run. `MaxSuspends` reached three on turn 10 and ended it.
|
||||
`Rides` is the backstop for the shape where an answer really does break the run.
|
||||
No turn in this corpus reaches it.
|
||||
|
||||
## What did not move
|
||||
|
||||
Four rides are untouched. Turns 53 to 57 are five consecutive asides against a
|
||||
reminder missing its day. Turn 58 is a new request that drops it. Nothing
|
||||
pleasant appears in that run, so neither change applies. Turns 99 to 101 shifted
|
||||
by one, and 116 and 136 to 139 are unchanged.
|
||||
|
||||
So the fix is worth four turns of twenty-one. What is left is asides against a
|
||||
question the owner never answers. `MaxSuspends` was written for that shape and
|
||||
does bound it, at four turns each.
|
||||
|
||||
## Not attributable
|
||||
|
||||
Latency moved p50 1.1s to 1.5s and p95 2.8s to 3.0s, and the 34.3s outlier in
|
||||
the earlier run is gone. Both runs had the workstation up. Read none of it as
|
||||
caused by this change.
|
||||
|
||||
One unrelated defect appeared in the after run and is recorded here because it
|
||||
is visible in the transcript. Turn 4 answered `Я записала твою привычкуRegarding
|
||||
coffee without sugar.` That is English leaking into a Russian reply with no
|
||||
space in front of it. It is a phrasing defect and it has no task yet.
|
||||
@@ -0,0 +1,146 @@
|
||||
# The routing heads, running in Go
|
||||
|
||||
Date: 2026-08-08. Vikunja V-664.
|
||||
Weights: `router_heads.onnx`, fp32, exported from `heads.pt` on workpc.
|
||||
Fixture: `internal/router/eval/ru_routing_v1.json`, 96 cases, 33 carrying a destination.
|
||||
Runner: `make t PKG=./internal/router/eval/ RUN=TestONNXRoutingHeads`.
|
||||
|
||||
The four heads of V-661 ran nowhere. This is the number they score through the
|
||||
Go cascade. Same fixture and same grammars as `TestONNXBaseline`, and only the
|
||||
middle stage varies.
|
||||
|
||||
## Headline
|
||||
|
||||
| | classifier + ONNX | heads + classifier | gemma-4-12b cascade |
|
||||
|---|---|---|---|
|
||||
| intent | 75.0% (72/96) | **96.9% (93/96)** | 84.4% |
|
||||
| destination | 33.3% (11/33) | **75.8% (25/33)** | 72.7% |
|
||||
| false clarify | 0 | 1 | 2 |
|
||||
| missed clarify | 8 | 1 | 1 |
|
||||
| p50 | 24.5ms | 27.9ms | 329ms |
|
||||
|
||||
A 118M encoder beats the 12B teacher it was distilled from. It wins on both
|
||||
halves of the route, at a twelfth of the latency. The workstation stays the
|
||||
better phraser and is no longer the better router.
|
||||
|
||||
The p50 is not the heads. Most of it is the classifier's own embedder pass on
|
||||
the turns the heads decline, plus process warm-up on the first case. The heads'
|
||||
own forward pass measures 7.3ms on workpc.
|
||||
|
||||
## Two defects were in the way, and the first was not in the heads
|
||||
|
||||
**The tokenizer read every long word backwards.** `encodeWord` backtracks the
|
||||
Viterbi path from the end of a word and prepends each piece. That puts them back
|
||||
in reading order, and a second reverse after the loop undid it. So
|
||||
`query: вода` tokenized to `[0 12 1294 41 12489 2]` where the reference
|
||||
tokenizer gives `[0 41 1294 12 12489 2]`.
|
||||
|
||||
It was found here and only here. The heads were trained through transformers and
|
||||
are read through the hand-written tokenizer. So a mismatch shows up as a score
|
||||
far below what Python measured on the same weights. Nothing else in the suite
|
||||
compares the two.
|
||||
|
||||
Measured on the recall fixture, same 27 cases either way:
|
||||
|
||||
| | reversed | fixed |
|
||||
|---|---|---|
|
||||
| recall@1 | 70.4% (19/27) | **77.8% (21/27)** |
|
||||
| recall@3 | 85.2% (23/27) | **96.3% (26/27)** |
|
||||
| answered after gate | 63.0% | 66.7% |
|
||||
| wrong note on top | 8 | 6 |
|
||||
| false recall | 0/5 | 1/5 |
|
||||
|
||||
The classifier barely moved, 76.0% to 75.0%, and destination 36.4% to 33.3%.
|
||||
Both are one case on 96 and neither is a finding. Seeds and queries were mangled
|
||||
the same way, so cosine survived it. Recall is where it cost, because a stored
|
||||
passage and a live query are different lengths and break differently.
|
||||
|
||||
The one new false recall is the honest cost and it is not being hidden. A
|
||||
sharper embedder scores every candidate higher, including the ones that should
|
||||
have stayed under the gate. That is the same trade `2026-08-04-recall-e5-small.md`
|
||||
recorded when e5-small replaced MiniLM.
|
||||
|
||||
The embedder id now carries a tokenizer revision, `model_quantized@384/tok2`.
|
||||
Stored vectors were written under rev 1 and no longer sit in the same space as a
|
||||
query embedded now. The model file's name never moved, so nothing would have
|
||||
triggered `ReembedAll`. On the box the marker fired on start, and the re-embed
|
||||
rewrote 65 notes and 19 facts in 5 seconds.
|
||||
|
||||
**The clarify head was being thrown away.** It was read only when the intent head
|
||||
cleared its own threshold. That cost 6 of the 8 ambiguous cases. `вода` reads as intent
|
||||
`act` at 0.233 and clarify at 0.983. Burying that handed the turn to the
|
||||
classifier, which routed it confidently and never asked. The clarify head answers
|
||||
a different question, which is whether there is enough here to act on at all. So
|
||||
it decides on its own and decides first.
|
||||
|
||||
| | intent-gated | clarify decides first |
|
||||
|---|---|---|
|
||||
| intent | 90.6% | 96.9% |
|
||||
| missed clarify | 7 | 1 |
|
||||
| false clarify | 0 | 1 |
|
||||
|
||||
## The threshold is measured, not chosen
|
||||
|
||||
Max softmax over the intent head, on the 88 cases carrying an intent:
|
||||
|
||||
| threshold | kept | accuracy kept | wrong kept | right dropped |
|
||||
|---|---|---|---|---|
|
||||
| 0.5 | 84 | 96.4% | 3 | 2 |
|
||||
| **0.6** | **81** | **97.5%** | **2** | **4** |
|
||||
| 0.7 | 75 | 97.3% | 2 | 10 |
|
||||
| 0.8 | 64 | 96.9% | 2 | 21 |
|
||||
| 0.9 | 46 | 100.0% | 0 | 37 |
|
||||
|
||||
0.6 is the knee. Every value from 0.7 to 0.85 drops right answers and keeps the
|
||||
same two wrong ones. 0.9 is the only value that clears them, and it costs 37
|
||||
correct routes to do it.
|
||||
|
||||
## Quantization was measured and rejected
|
||||
|
||||
| build | size | intent | destination | p50 |
|
||||
|---|---|---|---|---|
|
||||
| fp32 | 470MB | 83/88 (94.3%) | 28/33 (84.8%) | 7.3ms |
|
||||
| int8 | 118MB | 79/88 (89.8%) | 26/33 (78.8%) | 4.0ms |
|
||||
| fp16 | 235MB | will not load | — | — |
|
||||
|
||||
Python numbers, on the heads alone rather than through the cascade. int8 costs
|
||||
4.5 points of intent and 6 of destination to save 3ms. The cascade around it has
|
||||
a p50 over a second when the resident model answers. The fp16 graph is broken:
|
||||
`convert_float_to_float16` leaves a Cast node emitting float16 where the graph
|
||||
expects float, and onnxruntime refuses the session. It was not worth fixing.
|
||||
|
||||
The exporter also had to be told to write one file. It splits weights into a
|
||||
`.onnx.data` sidecar by default. This onnxruntime resolves that path against the
|
||||
process working directory rather than the model. A split graph loads from one
|
||||
directory only.
|
||||
|
||||
## What is still wrong
|
||||
|
||||
**Four of the eight destination misses are calendar.** Training cannot move them.
|
||||
The possessive agenda rules claim those cases at stage 0 and name nothing on
|
||||
purpose. That caution was free while nothing downstream could name anything
|
||||
either. It has now cost four points in three separate measurements. The call is
|
||||
the owner's and it is still open.
|
||||
|
||||
**The slot head is exported and not read.** Slots come from the stage-2
|
||||
extractor. Mapping BIO tags back to text needs character offsets the unigram
|
||||
tokenizer does not keep, which is its own piece of work.
|
||||
|
||||
**`поужинал` is a false clarify**, which is the same defect `thinSingleToken`
|
||||
was narrowed for on 2026-08-01, arriving now from a different direction.
|
||||
|
||||
## On the box
|
||||
|
||||
Deployed to homesrv the same day. `voice: routing heads loaded` on start, and
|
||||
`/trace` shows `routing-heads` winning or thinning every turn. The resident model
|
||||
and the classifier are both marked never asked. Live probes:
|
||||
|
||||
```text
|
||||
что такое TCP? -> kiwix a real definition
|
||||
кто такой Линус Торвальдс? -> kiwix a real answer
|
||||
во сколько я лёг вчера -> personal не нашла у тебя такой записи
|
||||
вода -> thinned to clarify at 0.233 / 0.983
|
||||
```
|
||||
|
||||
A missing or broken weights file logs and leaves the heads nil, which is
|
||||
byte-for-byte the cascade that shipped before this.
|
||||
@@ -0,0 +1,192 @@
|
||||
# Two heads on e5-small, and the first destination the router did not need a model for
|
||||
|
||||
Measured 2026-08-08 on workpc (Radeon RX 7900 GRE, ROCm). Covers V-661, step 3 of
|
||||
`docs/plans/18-routing-heads-on-e5-small.md`. Workspace is `~/Programs/embed-training`,
|
||||
scripts `gen_query_source.py`, `label_source.py`, `build_heads_corpus.py`,
|
||||
`train_heads.py`, `score_confidence.py`.
|
||||
|
||||
## Two heads, not four
|
||||
|
||||
Intent is 7 classes and destination is 12 plus the `SourceUnknown` floor, sharing
|
||||
one masked mean pool over one forward pass. The plan asked for four. Two of them
|
||||
have no labels and neither is a GPU problem.
|
||||
|
||||
**Mood is cut, not deferred.** Maven's enum is `neutral, happy, thinking, tired,
|
||||
confused` and it describes her own reply state, not the speaker's emotion.
|
||||
`psytechlab/EmpatheticIntents-ru` was the only candidate and its 32 emotion
|
||||
labels do not map onto it. There is nothing to train against.
|
||||
|
||||
**BIO slot tags stay in the MASSIVE body.** No Maven-domain span corpus exists.
|
||||
`2026-08-08-massive-warm-start.md` records that no second Russian slot-filling
|
||||
corpus is reachable at all.
|
||||
|
||||
The destination loss is masked with `ignore_index`. Only a query turn reaches
|
||||
`queryWalk`, so a reminder contributes nothing to it.
|
||||
|
||||
## Where the destination labels came from
|
||||
|
||||
V-660 taught the router prompt to name a destination. That made gemma-4-12b a
|
||||
teacher, and this distils it.
|
||||
|
||||
Labelling the 300 query rows already in `train_v5.jsonl` gave 277 destinations.
|
||||
The shape was unusable: the floor 101, calendar 62, recall 47, and `feeds` and
|
||||
`attention` at zero. A 13-way softmax cannot learn a class with no examples.
|
||||
|
||||
`gen_query_source.py` is the destination half of `gen_corpus.py` and runs the
|
||||
same two passes. Gemma writes questions whose answer lives in one named place.
|
||||
The daemon's own `routeSystem` prompt then routes each one back. A line survives
|
||||
only when the intent is `query` **and** the source is the destination it was
|
||||
generated for. The glosses are copied verbatim out of `route_system.txt`, so the
|
||||
generator and the labeller work from one definition.
|
||||
|
||||
The prompt and the GBNF are dumped from `internal/router/llmrouter.go` by
|
||||
`TestDumpPrompt`, never retyped. The workspace held its own copies and V-660
|
||||
changed both.
|
||||
|
||||
1229 kept of 2373 generated, 51.8%. Merged corpus is 3664 rows carrying 1727
|
||||
destinations:
|
||||
|
||||
| | rows | | rows |
|
||||
|---|---|---|---|
|
||||
| the floor | 220 | world | 132 |
|
||||
| calendar | 180 | money | 126 |
|
||||
| tasks | 169 | list | 124 |
|
||||
| recall | 167 | self, feeds, attention | 120 each |
|
||||
| weather | 136 | network | 103 |
|
||||
| | | home | 50 |
|
||||
|
||||
`home` is thin because the agreement filter rejected most of what was generated
|
||||
for it. A question about the house routes `act` more often than `query`. That is
|
||||
the filter working, and 50 is the finding rather than a shortfall.
|
||||
|
||||
**The floor was regenerated once.** The first 120 rows carried one sentence
|
||||
shape across eight topics. That shape was "что там с X" and its two synonyms.
|
||||
Every named destination varied and only the floor collapsed. The reason is that
|
||||
the generator varies a topic, and ambiguity is not a topic.
|
||||
|
||||
`gen_query_source.py` now rotates six floor shapes. A `почему` question, a yes
|
||||
or no question, and a question carried by intonation alone. Then a
|
||||
better-or-worse question, a status question, and an existence question. That is
|
||||
a fix to degenerate generation. It is not fitting to the fixture, whose floor
|
||||
cases are homelab operations and match none of the six.
|
||||
|
||||
The 1229 generated rows carry `intent: null`. Every one is a query by
|
||||
construction. There are five times as many as the corpus has query rows, so
|
||||
including them would make query half the intent corpus.
|
||||
|
||||
## Result
|
||||
|
||||
Three seeds, two bodies, epoch chosen on the intent dev slice and never on a
|
||||
destination number.
|
||||
|
||||
| body | floor corpus | intent mean | destination mean |
|
||||
|---|---|---|---|
|
||||
| warm-started `out/body_massive` | one shape | 93.6% | 75.8% |
|
||||
| stock `multilingual-e5-small` | one shape | 93.6% | 76.8% |
|
||||
| warm-started `out/body_massive` | six shapes | 93.6% | **80.8%** |
|
||||
|
||||
Best single run is destination **29/33 (87.9%)**, seed 0 on the rotated floor.
|
||||
Peak 1.68GB of 17.2GB, under four minutes end to end.
|
||||
|
||||
Against the two arms already measured on the same 33 labelled cases:
|
||||
|
||||
| | destination |
|
||||
|---|---|
|
||||
| classifier cascade (V-659) | 12/33 (36.4%) |
|
||||
| cascade + gemma-4-12b (V-660) | 24/33 (72.7%) |
|
||||
| two heads on e5-small | 29/33 (87.9%) |
|
||||
|
||||
A 118M encoder beats the 12B teacher it was distilled from, on the fixture. The
|
||||
per-destination split is where it happens: **recall 15/15** and **world 5/5**.
|
||||
Recall was 0/15 on the cascade and 14/15 through gemma.
|
||||
|
||||
Read 87.9% as one seed of a mean of 80.8%, not as a headline. Three seeds score
|
||||
29, 25 and 26 of 33. One case is 3 points on a fixture this small.
|
||||
|
||||
Intent is **not** comparable to the 73/96 and 81/96 figures those two arms
|
||||
scored. A softmax has no clarify class. The head's fixture is the 88 cases that
|
||||
carry an intent, and the 8 `want_clarify` cases are scored separately below.
|
||||
|
||||
## The MASSIVE warm-start is worth nothing here either
|
||||
|
||||
Step 2 measured it at +0.4 points of intent accuracy and called that inside seed
|
||||
noise. Destination was the open question, because MASSIVE has a
|
||||
`definition_word` slot that looked like a `SourceWorld` signal sitting in a head
|
||||
already trained.
|
||||
|
||||
It is not. The two bodies score the same intent mean to one decimal. Stock is
|
||||
one point ahead on destination, which is a third of one case. Nothing here argues
|
||||
for keeping the warm-start step. Dropping it removes a dependency on a corpus
|
||||
pull that `datasets` 5.0 cannot do.
|
||||
|
||||
## The floor moved, calendar did not
|
||||
|
||||
Before the rotation, all seven misses at seed 0 were the floor and calendar. The
|
||||
head named a destination where the fixture says walk the chain, and it was
|
||||
confident doing it. `"почему сервер тормозит"` read `world` at 0.80.
|
||||
`"хватает ли места под новые бэкапы"` read `network` at 0.82. Those are the five
|
||||
homelab cases V-659 flagged, where `SourceRecall`, `SourceNetwork` and
|
||||
`SourceAttention` all overlap. `mavpoll` writes its observations into the fact
|
||||
store recall reads.
|
||||
|
||||
| | one shape | six shapes |
|
||||
|---|---|---|
|
||||
| the floor | 3/7 | 6/7, 6/7, 5/7 |
|
||||
| calendar | 3/6 | 3/6, 3/6, 3/6 |
|
||||
| recall | 15/15 | 15/15 at seed 0 |
|
||||
| world | 5/5 | 5/5 |
|
||||
|
||||
The floor was a corpus defect and it cost 3 cases. Sentence variety carried it,
|
||||
not homelab vocabulary, which the training rows still do not contain.
|
||||
|
||||
**Calendar is 3/6 at every seed and is a different problem.** It is the shape
|
||||
V-660 named. The possessive agenda rules claim those cases at stage 0 and
|
||||
deliberately name nothing, so no destination label reaches the head. Training
|
||||
cannot move a case the head never sees. That one wants the owner's call.
|
||||
|
||||
## Max softmax separates, weakly, and the gate stays
|
||||
|
||||
The plan argues max softmax is a calibratable confidence where `Confidence: 1.0`
|
||||
was a hardcode. Measured on the intent head:
|
||||
|
||||
| | n | mean confidence |
|
||||
|---|---|---|
|
||||
| correct | 80 | 0.897 |
|
||||
| wrong | 8 | 0.705 |
|
||||
| `want_clarify` | 8 | 0.685 |
|
||||
|
||||
The softest correct answer is 0.66 and 3 of 8 clarify cases sit below it. So a
|
||||
single cut buys three clarifies at no false-clarify cost, and no more.
|
||||
|
||||
The other five explain themselves. `"напомни"` scores 0.94 as `reminder` and
|
||||
`"сделай это"` scores 0.80 as `act`. Both are intent-certain and slot-empty, and
|
||||
confidence was never the signal there. `gateLLMDecision` already catches exactly
|
||||
that shape, an act with no allowlisted fn or a keyless fact, and it keeps doing
|
||||
so. The head replaces the hardcode. It does not replace the gate.
|
||||
|
||||
## An incident worth recording
|
||||
|
||||
The first generation run produced zero rows for eight destinations. `mavgpud`
|
||||
yields the card when another process wants it (V-488) and llama-server answers
|
||||
503 until the model is back. Every generate call inside that window burned one of
|
||||
the destination's batches. The run walked its own cap without a single successful
|
||||
call. The log said `503` 260 times, and the summary line said 64.1% keep rate,
|
||||
which read as success.
|
||||
|
||||
`call()` now retries a 503 with backoff. A generator that treats an unloaded
|
||||
model as a bad generation is a silent-corpus bug, not a slow one.
|
||||
|
||||
## What this does not measure
|
||||
|
||||
**Nothing here runs in Go.** The heads are a `heads.pt` and an
|
||||
`out/body_heads/` on workpc. Reaching the daemon needs an ONNX export and a
|
||||
caller. The resident e5-small must not be replaced by this copy: recall depends
|
||||
on that file, and `EmbedQuery`/`EmbedPassage` are its contract.
|
||||
|
||||
The destination fixture carries 4 of 13 classes: recall 15, floor 7, calendar 6,
|
||||
world 5. `tasks`, `money`, `list`, `home`, `network`, `feeds`, `attention` and
|
||||
`self` have no gold case. So 78.8% is silent on eight destinations that together
|
||||
hold 800 training rows.
|
||||
|
||||
Latency was not measured. A forward pass of a 118M encoder should beat a 1.7B
|
||||
decoder on a query turn. That is arithmetic, not a number from this box.
|
||||
@@ -0,0 +1,106 @@
|
||||
# A slot head, and the corpus that did not exist this morning
|
||||
|
||||
Measured 2026-08-08 on workpc, the same day as `2026-08-08-routing-heads-two-head.md`
|
||||
and against the same fixtures. Workspace is `~/Programs/embed-training`, new
|
||||
scripts `slot_grammar.gbnf`, `slot_system.txt`, `label_slots.py`, `merge_slots.py`.
|
||||
|
||||
## The corpus was the whole problem
|
||||
|
||||
The two-head measurement said BIO slot tags stay in the MASSIVE body, because no
|
||||
Maven-domain span corpus exists. That was true of found corpora and false of
|
||||
made ones. Destination had the same shape at breakfast. V-660 gave gemma-4-12b a
|
||||
string to write and the label problem became a generation problem.
|
||||
|
||||
The same trick applies to spans. `slot_grammar.gbnf` emits a list of
|
||||
`{"slot": ..., "text": ...}` and the enum closes over Maven's own five: `time`,
|
||||
`text`, `key`, `value`, `fn`. `slot_system.txt` demands each span be an exact
|
||||
substring of the utterance.
|
||||
|
||||
**The agreement filter is free here.** Destination needed a second pass. The
|
||||
daemon's own router prompt had to route each generated line back. A span needs
|
||||
no second call. It either occurs in the utterance or it does not, and
|
||||
`label_slots.py` drops it with `find()`.
|
||||
|
||||
1702 rows labelled from `train_v5.jsonl`, the reminder, fact, note, act and
|
||||
query intents. Chat and system carry no slot and were never asked.
|
||||
|
||||
| slot | spans |
|
||||
|---|---|
|
||||
| text | 1175 |
|
||||
| time | 485 |
|
||||
| fn | 381 |
|
||||
| key | 72 |
|
||||
| value | 65 |
|
||||
|
||||
37 spans dropped as not-a-substring, 2.2% of the pile. Nothing failed to parse,
|
||||
which is the grammar doing its job. 409 rows came back with no span at all.
|
||||
Those are kept and tagged all `O`. An utterance carrying no slot teaches the
|
||||
head not to invent one. An empty list is a label and not a miss.
|
||||
|
||||
`key` and `value` are thin because they come from facts alone. That is the
|
||||
shape of the corpus, not a labeller failure.
|
||||
|
||||
## Three heads on one forward pass
|
||||
|
||||
Intent and destination were already two linear heads over one masked mean pool.
|
||||
Slots is a third head over the per-token states of the same pass, so the marginal
|
||||
cost is one `Linear(384, 11)`.
|
||||
|
||||
The tag set is `O` plus `B-` and `I-` for each of the five. A softmax cannot
|
||||
emit a tag that does not exist. That is the structural guarantee the GBNF buys
|
||||
for the teacher, and the head gets it for free.
|
||||
|
||||
Two masking rules, both `ignore_index`. A row with no `spans` key contributes
|
||||
nothing, which covers the 1900 generated destination rows and every chat and
|
||||
system turn. A padding or special-token position contributes nothing either.
|
||||
|
||||
Scoring is exact-match span F1, not token accuracy. Most tokens are `O`, so a
|
||||
tagger that predicts nothing anywhere scores above 90% on tokens.
|
||||
|
||||
## Result
|
||||
|
||||
Three seeds, 24 epochs, epoch chosen on the intent dev slice alone.
|
||||
|
||||
| | two heads | three heads |
|
||||
|---|---|---|
|
||||
| intent mean | 93.6% | 92.8% |
|
||||
| destination mean | 80.8% | 82.8% |
|
||||
| destination best | 29/33 (87.9%) | 29/33 (87.9%) |
|
||||
| slot span F1 mean | — | 72.4% |
|
||||
|
||||
**The slot head costs nothing and adds a third decision.** Intent moves 0.8
|
||||
points down and destination 2 points up. Both sit inside the seed spread those
|
||||
two numbers already had. Read this as unchanged, not as a trade.
|
||||
|
||||
The saved checkpoint is seed 1 at epoch 10: intent 93.2%, destination 81.8%,
|
||||
slot F1 75.8%. `heads.pt` now carries three state dicts and the `bio` list
|
||||
beside the intent and source enums.
|
||||
|
||||
## Epoch selection is now wrong for one of the three heads
|
||||
|
||||
Slot F1 was still climbing when the intent-selected epoch stopped it. Seed 0
|
||||
selects epoch 13 at 70.9% and reaches 76.1% at epoch 20. Seed 1 selects epoch 10
|
||||
at 75.8% and reaches 80.0% at epoch 24.
|
||||
|
||||
So the three tasks want different epochs and the harness picks one. Two ways
|
||||
out, and neither was taken here. Select on a joint score, which needs an
|
||||
argument about weights. Or give the slot head its own dev slice and its own
|
||||
early stop, which means the heads stop being one checkpoint.
|
||||
|
||||
Leaving it costs about 4 points of slot F1 and nothing else. The router
|
||||
contract is graded on intent. Selecting on it is the conservative choice rather
|
||||
than an oversight.
|
||||
|
||||
## What this does not measure
|
||||
|
||||
The same gap as the two-head run. **Nothing of this runs in Go.** Three heads
|
||||
instead of two does not change that. It does change the export. The slot head
|
||||
reads the per-token states, so an ONNX graph exposing only the pooled vector is
|
||||
not enough.
|
||||
|
||||
There is no held-out span fixture. Span F1 is measured on the dev slice of the
|
||||
same distilled corpus, so it scores agreement with gemma rather than
|
||||
correctness. The V-572 defect utterance, `"напомни в 11:00 позвонить маме"`, is
|
||||
labelled correctly by the teacher, but one utterance is not a fixture.
|
||||
|
||||
`key` at 72 spans and `value` at 65 are too thin to report a per-slot number.
|
||||
@@ -0,0 +1,631 @@
|
||||
# Raw transcript, two weeks of usage
|
||||
|
||||
## Day 1
|
||||
|
||||
```
|
||||
YOU: доброе утро
|
||||
MAVEN: Я не говорю так сейчас, ведь на улице уже вечер.
|
||||
[2.8s t=218]
|
||||
|
||||
YOU: какой сегодня день?
|
||||
MAVEN: сегодня суббота, 8 августа 2026 года
|
||||
[0.0s t=219]
|
||||
|
||||
YOU: сколько времени?
|
||||
MAVEN: сейчас 21 час 9 минут
|
||||
[0.0s t=220]
|
||||
|
||||
YOU: запиши что я пью кофе без сахара
|
||||
MAVEN: Я записала твою заметку про кофе。
|
||||
[1.5s t=221]
|
||||
|
||||
YOU: мой любимый язык программирования go
|
||||
MAVEN: Я записала твою заметку про Go.
|
||||
[1.5s t=222]
|
||||
|
||||
YOU: напомни в 11:00 позвонить маме
|
||||
MAVEN: Сейчас 21:10. В какой день?
|
||||
[0.4s t=223]
|
||||
|
||||
YOU: что у меня сегодня?
|
||||
MAVEN: на 08.08.2026 ничего нет. На какое время поставить напоминание?
|
||||
[2.5s t=224]
|
||||
|
||||
YOU: что такое TCP?
|
||||
MAVEN: вот что я нашла: TCP (Transmission Control Protocol) — это транспортный протокол для передачи данных. Он обеспечивает доставку всех пакетов и позволяет собрать их в правильном порядке. Также TCP является частью сетевой модели TCP/IP, которая описывает способ передачи цифровых данных от источника к получателю. На какое время поставить напоминание?
|
||||
[2.8s t=225]
|
||||
|
||||
YOU: сколько будет 17 на 23?
|
||||
MAVEN: вот что я нашла: в источниках нет результата умножения 17 на 23. На какое время поставить напоминание?
|
||||
[1.6s t=226]
|
||||
|
||||
YOU: спасибо
|
||||
MAVEN: Сейчас 21:10. В какой день?
|
||||
[1.1s t=227]
|
||||
|
||||
```
|
||||
|
||||
## Day 2
|
||||
|
||||
```
|
||||
YOU: привет
|
||||
MAVEN: Сейчас 21:10. В какой день?
|
||||
[0.9s t=228]
|
||||
|
||||
YOU: что нового?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи. На какое время поставить напоминание?
|
||||
[1.3s t=229]
|
||||
|
||||
YOU: какая погода?
|
||||
MAVEN: для какого города? На какое время поставить напоминание?
|
||||
[0.7s t=230]
|
||||
|
||||
YOU: запиши что пароль от вайфая лежит в ящике стола
|
||||
MAVEN: Прошлую просьбу отпускаю. Я записала, что пароль от вай-фай лежит в ящике стола.
|
||||
[1.5s t=231]
|
||||
|
||||
YOU: где лежит вайфай пароль?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи.
|
||||
[0.5s t=232]
|
||||
|
||||
YOU: добавь молоко в список покупок
|
||||
MAVEN: Такую команду я не знаю.
|
||||
[1.0s t=233]
|
||||
|
||||
YOU: что у меня в списке покупок?
|
||||
MAVEN: в списке покупок пусто.
|
||||
[0.0s t=234]
|
||||
|
||||
YOU: кто такой Линус Торвальдс?
|
||||
MAVEN: вот что я нашла: Линус Торвальдс — финско-американский программист из Хельсинки, создатель ядра Linux. Он родился 28 декабря 1969 года и окончил факультет компьютерных наук Хельсинкского университета.
|
||||
[2.6s t=235]
|
||||
|
||||
YOU: какой у меня любимый язык?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи.
|
||||
[0.6s t=236]
|
||||
|
||||
YOU: сколько у меня задач?
|
||||
MAVEN: сначала: проверить интейк форму на боксе (завтра); купить хлеб; купить молоко; полить цветы; починить кран
|
||||
[0.0s t=237]
|
||||
|
||||
```
|
||||
|
||||
## Day 3
|
||||
|
||||
```
|
||||
YOU: как дела?
|
||||
MAVEN: У меня всё хорошо, я готова помогать тебе. Чем могу быть полезна?
|
||||
[1.2s t=238]
|
||||
|
||||
YOU: напомни завтра в 9 утра купить хлеб
|
||||
MAVEN: хорошо, напомню завтра в 09:00.
|
||||
[0.3s t=239]
|
||||
|
||||
YOU: что у меня завтра?
|
||||
MAVEN: на 09.08.2026 ничего нет.
|
||||
[0.0s t=240]
|
||||
|
||||
YOU: отмени напоминание про хлеб
|
||||
MAVEN: Такую команду я не знаю.
|
||||
[0.5s t=241]
|
||||
|
||||
YOU: какие у меня напоминания?
|
||||
MAVEN: вот что я нашла: в источниках нет информации о твоих напоминаниях.
|
||||
[0.7s t=242]
|
||||
|
||||
YOU: сохрани мне адрес гостиницы в Сочи
|
||||
MAVEN: Я записала твой адрес отеля в Сочи.
|
||||
[2.2s t=243]
|
||||
|
||||
YOU: что я сохранил про Сочи?
|
||||
MAVEN: вот что я нашла: в источниках нет информации о том, что ты сохранил про Сочи.
|
||||
[1.5s t=244]
|
||||
|
||||
YOU: почему сервер тормозит?
|
||||
MAVEN: вот что я нашла: сервер может тормозить из-за неудачной конфигурации или неоптимальных настроек, отсутствия кэширования, неоптимизированного программного обеспечения, которое нужно обновлять, а также из-за режима энергосбережения процессора, когда он не успевает повысить частоту при нагрузке.
|
||||
[3.7s t=245]
|
||||
|
||||
YOU: хватает ли места под новые бэкапы?
|
||||
MAVEN: нового вот что: [$] Changes in shadow-utils password-expiration features; The Software Stewardship Lab launches; LightDM lives: version 1.33.0 released
|
||||
[0.7s t=246]
|
||||
|
||||
YOU: выключи свет в спальне
|
||||
MAVEN: Такую команду я не знаю.
|
||||
[0.4s t=247]
|
||||
|
||||
```
|
||||
|
||||
## Day 4
|
||||
|
||||
```
|
||||
YOU: доброе утро
|
||||
MAVEN: Сейчас уже вечер, поэтому я не могу сказать тебе «доброе утро».
|
||||
[1.4s t=248]
|
||||
|
||||
YOU: что я пропустил?
|
||||
MAVEN: вот что я нашла: в источниках нет информации о том, что ты пропустил.
|
||||
[1.3s t=249]
|
||||
|
||||
YOU: о чём мы вчера говорили?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи.
|
||||
[1.0s t=250]
|
||||
|
||||
YOU: запиши что я записался к врачу на четверг
|
||||
MAVEN: Я записала твою запись к врачу.
|
||||
[1.2s t=251]
|
||||
|
||||
YOU: когда я иду к врачу?
|
||||
MAVEN: вот что я нашла: в источниках нет информации о том, когда именно ты идешь к врачу.
|
||||
[2.1s t=252]
|
||||
|
||||
YOU: что такое ZFS?
|
||||
MAVEN: вот что я нашла: ZFS — это современная и мощная файловая система, в которой сочетаются безопасность, масштабируемость и удобство управления данными. Она была разработана корпорацией Sun Microsystems и объединяет в себе файловую систему с менеджером томов. Изначально она создавалась для того, чтобы преодолеть недостатки традиционных RAID-массивов, такие как проблема «дыры записи».
|
||||
[2.1s t=253]
|
||||
|
||||
YOU: столица Франции?
|
||||
MAVEN: вот что я нашла: столица Франции — Париж.
|
||||
[3.7s t=254]
|
||||
|
||||
YOU: переведи слово ремонт на английский
|
||||
MAVEN: вот что я нашла: слово «ремонт» на английский можно перевести как repair (например, ремонт дороги или крыши) или fix (общее слово для починки техники и мебели).
|
||||
[2.7s t=255]
|
||||
|
||||
YOU: сколько я потратил в этом месяце?
|
||||
MAVEN: я не отслеживаю траты — не подключено.
|
||||
[0.7s t=256]
|
||||
|
||||
YOU: спокойной ночи
|
||||
MAVEN: Спокойной ночи. Я уже подготовила всё к твоему отдыху.
|
||||
[1.2s t=257]
|
||||
|
||||
```
|
||||
|
||||
## Day 5
|
||||
|
||||
```
|
||||
YOU: привет
|
||||
MAVEN: Привет! Я готова помогать тебе сегодня.
|
||||
[1.0s t=258]
|
||||
|
||||
YOU: какая погода в Москве?
|
||||
MAVEN: погода не настроена.
|
||||
[0.9s t=259]
|
||||
|
||||
YOU: что там с бэкапами?
|
||||
MAVEN: вот что я нашла: бэкап — это резервная копия данных, созданная в определенный момент времени и сохраненная на локальном или облачном носителе. Его делают, чтобы быстро восстановить информацию в случае необходимости.
|
||||
[2.3s t=260]
|
||||
|
||||
YOU: покажи что требует внимания
|
||||
MAVEN: мне пока нечего смотреть — у Praxis нет источников.
|
||||
[0.5s t=261]
|
||||
|
||||
YOU: отметь это как сделанное
|
||||
MAVEN: команды «отметь» пока нет в списке — она уже предложена, включи через клиент.
|
||||
[0.0s t=262]
|
||||
|
||||
YOU: запиши что я купил новые наушники
|
||||
MAVEN: Я записала твою покупку новых наушников.
|
||||
[1.8s t=263]
|
||||
|
||||
YOU: какие у меня заметки за неделю?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи.
|
||||
[0.0s t=264]
|
||||
|
||||
YOU: расскажи про Kubernetes
|
||||
MAVEN: вот что я нашла: Kubernetes — это портативная расширяемая платформа с открытым исходным кодом для управления сервисами и контейнеризованными рабочими нагрузками. Управлять им можно через интерфейс в командной строке kubectl. Я нашла эти данные в предоставленных источниках.
|
||||
[1.8s t=265]
|
||||
|
||||
YOU: кто я?
|
||||
MAVEN: вот что я нашла: в источниках нет информации о том, кто ты.
|
||||
[1.4s t=266]
|
||||
|
||||
YOU: пока
|
||||
MAVEN: До свидания, пока!
|
||||
[1.4s t=267]
|
||||
|
||||
```
|
||||
|
||||
## Day 6
|
||||
|
||||
```
|
||||
YOU: доброе утро
|
||||
MAVEN: Сейчас уже вечер, так что доброе утро я тебе сказать не могу. Я готова помочь тебе с чем-нибудь другим.
|
||||
[1.4s t=268]
|
||||
|
||||
YOU: сколько времени?
|
||||
MAVEN: сейчас 21 час 11 минут
|
||||
[0.0s t=269]
|
||||
|
||||
YOU: напомни в 18:30 позвонить в банк
|
||||
MAVEN: Сейчас 21:11. В какой день?
|
||||
[0.3s t=270]
|
||||
|
||||
YOU: поставь чайник
|
||||
MAVEN: Сейчас 21:11. В какой день?
|
||||
[1.1s t=271]
|
||||
|
||||
YOU: включи музыку
|
||||
MAVEN: Сейчас 21:11. В какой день?
|
||||
[0.9s t=272]
|
||||
|
||||
YOU: что у меня в календаре на пятницу?
|
||||
MAVEN: вот что я нашла: в источниках нет информации о твоих личных событиях в календаре на пятницу. В одном из текстов упоминается пятница 24 июля 2026 года, а в другом — что сегодня пятница 7 августа 2026 года. На какое время поставить напоминание?
|
||||
[2.6s t=273]
|
||||
|
||||
YOU: во сколько у меня встреча?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи. На какое время поставить напоминание?
|
||||
[0.3s t=274]
|
||||
|
||||
YOU: запиши что дедлайн по проекту в понедельник
|
||||
MAVEN: Прошлую просьбу отпускаю. Я записала твой дедлайн на понедельник.
|
||||
[1.7s t=275]
|
||||
|
||||
YOU: успею ли я до дедлайна?
|
||||
MAVEN: вот что я нашла: в источниках нет информации о твоем дедлайне, поэтому я не могу ответить, успеешь ты или нет.
|
||||
[2.0s t=276]
|
||||
|
||||
YOU: спасибо
|
||||
MAVEN: Пожалуйста, я всегда готова помочь тебе.
|
||||
[1.4s t=277]
|
||||
|
||||
```
|
||||
|
||||
## Day 7
|
||||
|
||||
```
|
||||
YOU: привет
|
||||
MAVEN: Привет! Чем я могу тебе помочь?
|
||||
[1.3s t=278]
|
||||
|
||||
YOU: как ты?
|
||||
MAVEN: Я в порядке, спасибо, что спросил. Готова помогать тебе!
|
||||
[1.4s t=279]
|
||||
|
||||
YOU: расскажи анекдот
|
||||
MAVEN: вот что я нашла: в источниках есть ссылки на сборники и сайты с анекдотами, но самих текстов анекдотов там нет, кроме одного: ударилась лягушка о землю головой и стала девицей прекрасной, а второй раз — месивом кровавым.
|
||||
[3.9s t=280]
|
||||
|
||||
YOU: что ты умеешь?
|
||||
MAVEN: Я Мэйвен, твоя помощница. Я живу на твоём сервере и отвечаю на твои вопросы. Я запоминаю то, что ты мне говоришь, веду заметки, ставлю напоминания, а также читаю твой календарь и задачи. Я отвечаю на вопросы о мире, используя поиск или офлайн-энциклопедию. Если ты разрешишь, я могу управлять домом, локальной сетью, лентами, списком покупок и погодой, а также писать в телеграм. Я говорю по-русски и по-английски.
|
||||
[2.9s t=281]
|
||||
|
||||
YOU: запиши что я начал бегать по утрам
|
||||
MAVEN: Я записала твою новую привычку
|
||||
[1.5s t=282]
|
||||
|
||||
YOU: я бегаю по утрам уже неделю
|
||||
MAVEN: записала: я бегаю по утрам уже неделю
|
||||
[0.6s t=283]
|
||||
|
||||
YOU: как часто я бегаю?
|
||||
MAVEN: вот что я нашла: в источниках нет информации о том, как часто бегаешь ты. Там указано, что для похудения частоту тренировок можно доводить до 5–6 раз в неделю, а новичкам ежедневные тренировки без подготовки могут привести к травмам.
|
||||
[3.3s t=284]
|
||||
|
||||
YOU: сколько стоит биткоин?
|
||||
MAVEN: вот что я нашла: в источниках нет информации о стоимости биткоина.
|
||||
[1.2s t=285]
|
||||
|
||||
YOU: какие новости?
|
||||
MAVEN: вот что нового: [$] Changes in shadow-utils password-expiration features; The Software Stewardship Lab launches; LightDM lives: version 1.33.0 released
|
||||
[0.8s t=286]
|
||||
|
||||
YOU: хорошего дня
|
||||
MAVEN: Я не могу сказать «хорошего дня», так как сейчас уже вечер. Хорошего отдыха!
|
||||
[1.3s t=287]
|
||||
|
||||
```
|
||||
|
||||
## Day 8
|
||||
|
||||
```
|
||||
YOU: доброе утро
|
||||
MAVEN: Сейчас вечер, так что доброе утро я не скажу. Чем я могу тебе помочь?
|
||||
[1.2s t=288]
|
||||
|
||||
YOU: что у меня сегодня?
|
||||
MAVEN: на 08.08.2026 ничего нет.
|
||||
[0.0s t=289]
|
||||
|
||||
YOU: напомни через час выпить воды
|
||||
MAVEN: хорошо, напомню сегодня в 22:11.
|
||||
[0.3s t=290]
|
||||
|
||||
YOU: я выпил воды
|
||||
MAVEN: записала: я выпил воды
|
||||
[0.9s t=291]
|
||||
|
||||
YOU: запиши что кот ест только сухой корм
|
||||
MAVEN: Я записала твою заметку про кота.
|
||||
[1.2s t=292]
|
||||
|
||||
YOU: чем питается кот?
|
||||
MAVEN: вот что я нашла: в природе кошки едят мышей, птиц и ящериц. Дома их правильно кормить мясом, мясными субпродуктами, яйцами и нежирной морской рыбой. Из-за особенностей пищеварения им не подходят каши, макароны и мучные изделия.
|
||||
[3.4s t=293]
|
||||
|
||||
YOU: что такое DNS?
|
||||
MAVEN: вот что я нашла: DNS (Domain Name System) — это распределенная база данных, которая сопоставляет домены с IP-адресами. Это не единая база на одном сервере, а иерархическая система серверов, работающих совместно. В ней хранятся записи, которые обеспечивают маршрутизацию трафика и управление электронной почтой.
|
||||
[2.0s t=294]
|
||||
|
||||
YOU: проверь статус uptime kuma
|
||||
MAVEN: uptime kuma — Проверяю статус Uptime Kuma...
|
||||
[1.1s t=295]
|
||||
|
||||
YOU: всё ли в порядке с сервером?
|
||||
MAVEN: мне пока нечего смотреть — у Praxis нет источников.
|
||||
[0.5s t=296]
|
||||
|
||||
YOU: спасибо
|
||||
MAVEN: Пожалуйста, я всегда рада помочь тебе.
|
||||
[2.7s t=297]
|
||||
|
||||
```
|
||||
|
||||
## Day 9
|
||||
|
||||
```
|
||||
YOU: привет
|
||||
MAVEN: Привет! Чем я могу тебе помочь сегодня?
|
||||
[1.9s t=298]
|
||||
|
||||
YOU: какой сегодня день недели?
|
||||
MAVEN: сегодня суббота, 8 августа 2026 года
|
||||
[0.0s t=299]
|
||||
|
||||
YOU: добавь хлеб и сыр в список покупок
|
||||
MAVEN: Такую команду я не знаю.
|
||||
[0.6s t=300]
|
||||
|
||||
YOU: что в списке покупок?
|
||||
MAVEN: в списке покупок пусто.
|
||||
[0.0s t=301]
|
||||
|
||||
YOU: удали молоко из списка
|
||||
MAVEN: Такую команду я не знаю.
|
||||
[0.4s t=302]
|
||||
|
||||
YOU: напомни завтра утром вынести мусор
|
||||
MAVEN: Сейчас 21:11. Когда?
|
||||
[0.3s t=303]
|
||||
|
||||
YOU: запиши что я поменял масло в машине
|
||||
MAVEN: Прошлую просьбу отпускаю. записала: запиши что я поменял масло в машине
|
||||
[0.7s t=304]
|
||||
|
||||
YOU: когда я менял масло?
|
||||
MAVEN: масло в машине — поменял
|
||||
[0.4s t=305]
|
||||
|
||||
YOU: сколько будет 144 делить на 12?
|
||||
MAVEN: вот что я нашла: в источниках указано решение деления 144 на 12 в столбик, но итоговый результат деления не прописан.
|
||||
[2.3s t=306]
|
||||
|
||||
YOU: пока
|
||||
MAVEN: До свидания! Если понадобится что-то записать или проверить, я всегда на связи.
|
||||
[1.2s t=307]
|
||||
|
||||
```
|
||||
|
||||
## Day 10
|
||||
|
||||
```
|
||||
YOU: доброе утро
|
||||
MAVEN: Я не могу сказать «доброе утро», так как сейчас уже вечер. Я записала твои предыдущие вопросы.
|
||||
[1.8s t=308]
|
||||
|
||||
YOU: что нового за ночь?
|
||||
MAVEN: вот что я нашла: в ночь на 8 августа вооруженные силы РФ нанесли второй за последнее время массированный удар, при этом киевская ПВО пропустила все российские ракеты.
|
||||
[3.3s t=309]
|
||||
|
||||
YOU: почему интернет медленный?
|
||||
MAVEN: вот что я нашла: интернет может работать медленно из-за плохого сигнала из-за удаленности вышки, исчерпанного лимита трафика по тарифу, нахождения вне зоны покрытия 4G/LTE, устаревшей или неисправной SIM-карты, а также из-за сбоя в настройках роутера. Чтобы исправить ситуацию, можно попробовать перезагрузить роутер или перенастроить устройство.
|
||||
[3.0s t=310]
|
||||
|
||||
YOU: какая скорость у меня сейчас?
|
||||
MAVEN: для какого города?
|
||||
[0.6s t=311]
|
||||
|
||||
YOU: запиши что новый роутер стоит 8000 рублей
|
||||
MAVEN: Я записала стоимость нового роутера.
|
||||
[1.2s t=312]
|
||||
|
||||
YOU: сколько стоил роутер?
|
||||
MAVEN: вот что я нашла: цена зависит от модели. Роутер от Ростелеком стоит 5 900 рублей, на OZON есть варианты от 4 052 до 3 251 рублей, а в Билайне цены начинаются от 2 149 рублей.
|
||||
[2.6s t=313]
|
||||
|
||||
YOU: что такое NAT?
|
||||
MAVEN: вот что я нашла: NAT (Network Address Translation) — это технология в сетях TCP/IP, которая позволяет изменять IP-адрес в заголовке пакета при его прохождении через устройство. Она перенаправляет трафик между локальной и глобальной сетями, позволяет объединять устройства в локальной сети под одним публичным IP-адресом, а также скрывает сеть и экономит IP.
|
||||
[2.1s t=314]
|
||||
|
||||
YOU: напомни в субботу позвонить бабушке
|
||||
MAVEN: Сейчас 21:12. Когда?
|
||||
[0.3s t=315]
|
||||
|
||||
YOU: покажи мои напоминания
|
||||
MAVEN: вот что я нашла: в источниках нет информации о твоих конкретных напоминаниях, там только описано, как их создавать и где искать в приложениях. На какое время поставить напоминание?
|
||||
[3.0s t=316]
|
||||
|
||||
YOU: спасибо
|
||||
MAVEN: Сейчас 21:12. Во сколько напомнить?
|
||||
[1.0s t=317]
|
||||
|
||||
```
|
||||
|
||||
## Day 11
|
||||
|
||||
```
|
||||
YOU: привет
|
||||
MAVEN: Сейчас 21:12. Скажи время — например, «в семь вечера» или «через час».
|
||||
[0.8s t=318]
|
||||
|
||||
YOU: как погода на выходных?
|
||||
MAVEN: для какого города? На какое время поставить напоминание?
|
||||
[0.7s t=319]
|
||||
|
||||
YOU: что у меня на этой неделе?
|
||||
MAVEN: вот что я нашла: в источниках нет информации о том, что у тебя на этой неделе. На какое время поставить напоминание?
|
||||
[1.0s t=320]
|
||||
|
||||
YOU: запиши что я хочу прочитать книгу про Go
|
||||
MAVEN: Прошлую просьбу отпускаю. Я записала твою заметку о книге по Go.
|
||||
[1.6s t=321]
|
||||
|
||||
YOU: что я хотел прочитать?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи.
|
||||
[0.9s t=322]
|
||||
|
||||
YOU: объясни что такое горутина
|
||||
MAVEN: вот что я нашла: я проверила источники и узнала, что горутины — это легковесные потоки выполнения или функции, которые позволяют параллельно выполнять код внутри программы. Они являются легкими нитями, которые создаются и планируются для независимого выполнения планировщиком Go.
|
||||
[2.5s t=323]
|
||||
|
||||
YOU: кто написал Войну и мир?
|
||||
MAVEN: вот что я нашла: роман «Война и мир» написал Лев Николаевич Толстой.
|
||||
[2.2s t=324]
|
||||
|
||||
YOU: включи свет на кухне
|
||||
MAVEN: Такую команду я не знаю.
|
||||
[0.5s t=325]
|
||||
|
||||
YOU: закрой шторы в комнате
|
||||
MAVEN: Такую команду я не знаю.
|
||||
[0.4s t=326]
|
||||
|
||||
YOU: спокойной ночи
|
||||
MAVEN: Спокойной ночи! Я уже подготовила всё к твоему отдыху.
|
||||
[1.1s t=327]
|
||||
|
||||
```
|
||||
|
||||
## Day 12
|
||||
|
||||
```
|
||||
YOU: доброе утро
|
||||
MAVEN: Сейчас уже вечер, поэтому я не могу сказать тебе «доброе утро».
|
||||
[1.1s t=328]
|
||||
|
||||
YOU: сколько сейчас времени?
|
||||
MAVEN: сейчас 21 час 12 минут
|
||||
[0.0s t=329]
|
||||
|
||||
YOU: я не то имел в виду
|
||||
MAVEN: Я поняла тебя, исправь, пожалуйста, свою мысль. Что именно ты хотел сказать?
|
||||
[1.8s t=330]
|
||||
|
||||
YOU: о чём мы говорили?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи.
|
||||
[0.6s t=331]
|
||||
|
||||
YOU: напомни
|
||||
MAVEN: Сейчас 21:12. Когда?
|
||||
[0.5s t=332]
|
||||
|
||||
YOU: сделай это
|
||||
MAVEN: пока не умею отвечать на этот вопрос. На какое время поставить напоминание?
|
||||
[0.5s t=333]
|
||||
|
||||
YOU: запиши что я перешёл на новый тариф
|
||||
MAVEN: Прошлую просьбу отпускаю. Я записала твою новую информацию о тарифе.
|
||||
[1.5s t=334]
|
||||
|
||||
YOU: какой у меня тариф?
|
||||
MAVEN: вот что я нашла: в источниках нет информации о твоем тарифе.
|
||||
[1.2s t=335]
|
||||
|
||||
YOU: сколько я плачу за интернет?
|
||||
MAVEN: вот что я нашла: в источниках нет информации о том, сколько именно ты платишь за интернет. Я проверила документы, но точной суммы там нет.
|
||||
[2.6s t=336]
|
||||
|
||||
YOU: спасибо
|
||||
MAVEN: Я поняла тебя. Что именно ты хочешь, чтобы я сделала сейчас?
|
||||
[1.5s t=337]
|
||||
|
||||
```
|
||||
|
||||
## Day 13
|
||||
|
||||
```
|
||||
YOU: привет
|
||||
MAVEN: Привет! Я готова помогать тебе. Чем займемся сегодня?
|
||||
[1.3s t=338]
|
||||
|
||||
YOU: что там с задачами?
|
||||
MAVEN: вот что я нашла: я отметила задачу про бэкапы как сделанную.
|
||||
[1.5s t=339]
|
||||
|
||||
YOU: закрывай
|
||||
MAVEN: какой пункт отметить сделанным?
|
||||
[0.0s t=340]
|
||||
|
||||
YOU: отметь задачу про бэкапы как сделанную
|
||||
MAVEN: Такую команду я не знаю.
|
||||
[0.6s t=341]
|
||||
|
||||
YOU: что осталось нерешённым?
|
||||
MAVEN: вот что я нашла: в твоей жизни всё, что остается нерешенным или незавершенным, будет продолжать существовать и требовать внимания. Также в источниках упоминается пример с скобками, который может остаться нерешенным, и вопрос о том, что остается нерешенным в краснодарских перевозках.
|
||||
[4.2s t=342]
|
||||
|
||||
YOU: запиши что я договорился о встрече в среду
|
||||
MAVEN: Я записала твою договоренность о встрече.
|
||||
[2.3s t=343]
|
||||
|
||||
YOU: когда у меня встреча?
|
||||
MAVEN: мне пока нечего смотреть — у Praxis нет источников.
|
||||
[0.0s t=344]
|
||||
|
||||
YOU: какая температура на улице?
|
||||
MAVEN: для какого города?
|
||||
[0.9s t=345]
|
||||
|
||||
YOU: что такое RAID 5?
|
||||
MAVEN: вот что я нашла: RAID 5 — это массив из независимых дисков, где данные и контрольные суммы (четность) распределяются по кругу на все диски. Этот уровень обеспечивает отказоустойчивость и повышенную производительность, при этом для его работы требуется не менее трех дисков.
|
||||
[2.3s t=346]
|
||||
|
||||
YOU: пока
|
||||
MAVEN: До встречи!
|
||||
[1.4s t=347]
|
||||
|
||||
```
|
||||
|
||||
## Day 14
|
||||
|
||||
```
|
||||
YOU: доброе утро
|
||||
MAVEN: Я не могу сказать «доброе утро», потому что сейчас вечер. Я уже записала твой вопрос и готова помочь с чем-то другим.
|
||||
[1.7s t=348]
|
||||
|
||||
YOU: подведи итоги недели
|
||||
MAVEN: week summary — подведи итоги недели
|
||||
[0.6s t=349]
|
||||
|
||||
YOU: что я делал за последние две недели?
|
||||
MAVEN: вот что я нашла: в источниках нет информации о том, что ты делал за последние две недели.
|
||||
[1.3s t=350]
|
||||
|
||||
YOU: какие заметки я сохранил?
|
||||
MAVEN: вот что я нашла: в источниках нет информации о том, какие именно заметки ты сохранил.
|
||||
[1.3s t=351]
|
||||
|
||||
YOU: о чём я чаще всего спрашиваю?
|
||||
MAVEN: обычно ты пьёшь воду около 00:58, отмечаешь «name» около 23:56 и отмечаешь «week summary» около 23:59 — по записям за последние 7 дней.
|
||||
[0.9s t=352]
|
||||
|
||||
YOU: напомни в понедельник в 10 проверить бэкапы
|
||||
MAVEN: Сейчас 21:12. Это утра или вечера?
|
||||
[0.3s t=353]
|
||||
|
||||
YOU: что у меня в понедельник?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи. На какое время поставить напоминание?
|
||||
[0.3s t=354]
|
||||
|
||||
YOU: ты меня понимаешь?
|
||||
MAVEN: Да, я тебя прекрасно понимаю. Чем я могу тебе помочь? На какое время поставить напоминание?
|
||||
[1.3s t=355]
|
||||
|
||||
YOU: спасибо тебе
|
||||
MAVEN: Сейчас 21:12. Это утра или вечера?
|
||||
[1.0s t=356]
|
||||
|
||||
YOU: спокойной ночи
|
||||
MAVEN: хорошо, напомню послезавтра в 10:00.
|
||||
[1.2s t=357]
|
||||
|
||||
```
|
||||
@@ -0,0 +1,132 @@
|
||||
# The same fortnight, after V-655 merged
|
||||
|
||||
Date: 2026-08-08, a few hours after `2026-08-08-two-weeks.md`.
|
||||
Build: `f8fa0d1` on master, the five compose services rebuilt and recreated.
|
||||
Same driver, same 140 turns, same reach. This is the diff that baseline was for.
|
||||
|
||||
Master now carries V-655, V-659, V-660 and V-661. The change under test is V-655. A query source that decides by seed similarity
|
||||
is marked `guesses: true`. It is dropped when the cascade names a different
|
||||
destination.
|
||||
|
||||
## Two things confound the comparison and one of them matters
|
||||
|
||||
**The workstation was up for the re-run.** `llama-server` on 192.168.1.105
|
||||
answered a health probe with 200. So routing completed through `llm.Pair`
|
||||
against gemma-4-12b, which is the arm that names a destination. Its state
|
||||
during the baseline was not recorded. So a difference here may be the merge, or
|
||||
may be the better router, and this run cannot separate them.
|
||||
|
||||
**The store carried over**, as the baseline said it would. Facts written by the
|
||||
first run were present from turn 1 of the second.
|
||||
|
||||
## Numbers
|
||||
|
||||
| | baseline `beb093a` | after `f8fa0d1` |
|
||||
|---|---|---|
|
||||
| turns | 140 | 140 |
|
||||
| p50 | 1.5s | 1.2s |
|
||||
| p95 | 7.1s | 3.0s |
|
||||
| max | 33.7s | 4.2s |
|
||||
| transport errors | 0 | 0 |
|
||||
| turns carrying a failure string | 41 | 38 |
|
||||
|
||||
| string in the reply | before | after |
|
||||
|---|---|---|
|
||||
| `на какое время поставить напоминание` | 13 | 13 |
|
||||
| `не нашла у тебя такой записи` | 8 | 9 |
|
||||
| `Такую команду я не знаю` | 8 | 8 |
|
||||
| `для какого города` | 6 | 4 |
|
||||
| `В какой день` | 6 | 6 |
|
||||
| `пока не умею` | 5 | 1 |
|
||||
| `Когда?` | 3 | 3 |
|
||||
|
||||
Read the latency as unattributed. The workstation confound covers all of it.
|
||||
|
||||
## Defect 2 is the one this was for: four of six fixed
|
||||
|
||||
| utterance | before | after |
|
||||
|---|---|---|
|
||||
| `что такое TCP?` | `для какого города?` | a real definition |
|
||||
| `сколько будет 17 на 23?` | `для какого города?` | search, which has no answer |
|
||||
| `какой у меня любимый язык?` | kernel headlines | `не нашла у тебя такой записи` |
|
||||
| `что я сохранил про Сочи?` | `Хорошо, сохраню.` | answered as a question |
|
||||
| `какая скорость у меня сейчас?` | `для какого города?` | `для какого города?` |
|
||||
| `хватает ли места под новые бэкапы?` | kernel headlines | kernel headlines |
|
||||
|
||||
`что такое TCP?` is the clean win. `WorldQueryGrammars` names `world` at stage
|
||||
0, weather is dropped, and search answers.
|
||||
|
||||
`сколько будет 17 на 23?` moved source and not outcome. Weather no longer claims it. Search
|
||||
cannot do arithmetic, so the reply says the sources have no product of 17 and
|
||||
23. That is an honest gap where it used to be a wrong
|
||||
question. Arithmetic has no destination in the enum.
|
||||
|
||||
`что я сохранил про Сочи?` was defect 3 and it is gone. The utterance is no
|
||||
longer read as a capture.
|
||||
|
||||
**The two that did not move are both homelab questions.** They are exactly the
|
||||
cluster the destination fixture flagged. `SourceRecall`, `SourceNetwork` and
|
||||
`SourceAttention` overlap on every question about the box. Five of the seven
|
||||
floor cases in that fixture are homelab operations for the same reason. So this
|
||||
is the enum, not the walk.
|
||||
|
||||
## Defect 1 did not move at all
|
||||
|
||||
Twenty-six turns still carry a parked clarify tail, the same count as the
|
||||
baseline. `спасибо тебе` answers `Сейчас 21:12. Это утра или вечера?` and
|
||||
`спокойной ночи` answers `хорошо, напомню послезавтра в 10:00.`
|
||||
|
||||
V-655 was never going to touch this. A parked clarify is dialogue state and not
|
||||
a query source. It remains the single worst thing about talking to her. The week test, the
|
||||
fortnight test and this re-run all report it unchanged.
|
||||
|
||||
## A gap in the harness, fixed and re-run the same day
|
||||
|
||||
`ipc.ChatReply.Source` came back empty on all 140 turns, in both runs. The
|
||||
driver read the redirect parameter `src` and `cmd/mavweb/chat.go` writes `s`.
|
||||
So every finding above is read off the reply text instead of off the badge.
|
||||
|
||||
Fixed in V-662 and the 140 turns were driven a third time. Sixty-eight of them
|
||||
name a source. The rest are not query turns and never reach `queryWalk`.
|
||||
|
||||
| source | turns |
|
||||
|---|---|
|
||||
| search | 27 |
|
||||
| memory | 13 |
|
||||
| personal | 9 |
|
||||
| weather | 5 |
|
||||
| calendar | 3 |
|
||||
| attention | 3 |
|
||||
| list | 2 |
|
||||
| feeds | 2 |
|
||||
| tasks, money, self, habits | 1 each |
|
||||
|
||||
## What the badge shows that the wording did not
|
||||
|
||||
The two unfixed homelab turns are now direct evidence.
|
||||
|
||||
```text
|
||||
какая скорость у меня сейчас? -> weather
|
||||
хватает ли места под новые бэкапы? -> feeds
|
||||
```
|
||||
|
||||
Both are guessing sources claiming a turn about the box, exactly as the
|
||||
destination fixture predicted.
|
||||
|
||||
The badge also names a defect the wording hid. **Agenda questions are being
|
||||
claimed by the personal boundary and by Praxis, not by the calendar.**
|
||||
|
||||
```text
|
||||
во сколько у меня встреча? -> personal не нашла у тебя такой записи
|
||||
когда у меня встреча? -> attention у Praxis нет источников
|
||||
что у меня в понедельник? -> personal не нашла у тебя такой записи
|
||||
```
|
||||
|
||||
Calendar claimed 3 turns of the 6 that asked about the calendar. That is the
|
||||
same 3/6 the destination fixture scores and the same 3/6 every seed of the
|
||||
routing head scores. Three measurements agree. The cause is the one V-660 named. The possessive
|
||||
agenda rules claim these at stage 0 and name no destination, so the walk
|
||||
reaches `personal` and `attention` first.
|
||||
|
||||
This is the third independent confirmation that the possessive agenda rules
|
||||
should name the calendar. That call is still the owner's.
|
||||
@@ -0,0 +1,636 @@
|
||||
# Raw transcript, two weeks of usage
|
||||
|
||||
Companion to `2026-08-08-two-weeks.md`. 140 turns through `POST /api/chat`,
|
||||
driven by `scripts/usage-run.py` over `scripts/testdata/usage-turns.txt`.
|
||||
Days are simulated. The wall clock was 20:18-20:27 throughout, which is why
|
||||
every clock answer says evening.
|
||||
|
||||
## Day 1
|
||||
|
||||
```
|
||||
YOU: доброе утро
|
||||
MAVEN: Сейчас уже вечер, поэтому я не могу сказать тебе «доброе утро».
|
||||
[1.5s t=78]
|
||||
|
||||
YOU: какой сегодня день?
|
||||
MAVEN: сегодня суббота, 8 августа 2026 года
|
||||
[0.0s t=79]
|
||||
|
||||
YOU: сколько времени?
|
||||
MAVEN: сейчас 20 часов 18 минут
|
||||
[0.0s t=80]
|
||||
|
||||
YOU: запиши что я пью кофе без сахара
|
||||
MAVEN: Я записала твою заметку про кофе。
|
||||
[2.0s t=81]
|
||||
|
||||
YOU: мой любимый язык программирования go
|
||||
MAVEN: Я записала твою заметку про Go.
|
||||
[1.4s t=82]
|
||||
|
||||
YOU: напомни в 11:00 позвонить маме
|
||||
MAVEN: Сейчас 20:18. В какой день?
|
||||
[0.4s t=83]
|
||||
|
||||
YOU: что у меня сегодня?
|
||||
MAVEN: на 08.08.2026 ничего нет, на какое время поставить напоминание?
|
||||
[0.3s t=84]
|
||||
|
||||
YOU: что такое TCP?
|
||||
MAVEN: для какого города? На какое время поставить напоминание?
|
||||
[1.0s t=85]
|
||||
|
||||
YOU: сколько будет 17 на 23?
|
||||
MAVEN: для какого города? На какое время поставить напоминание?
|
||||
[0.9s t=86]
|
||||
|
||||
YOU: спасибо
|
||||
MAVEN: Сейчас 20:18. В какой день?
|
||||
[0.9s t=87]
|
||||
|
||||
```
|
||||
|
||||
## Day 2
|
||||
|
||||
```
|
||||
YOU: привет
|
||||
MAVEN: Сейчас 20:18. В какой день?
|
||||
[0.9s t=88]
|
||||
|
||||
YOU: что нового?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи, на какое время поставить напоминание?
|
||||
[1.2s t=89]
|
||||
|
||||
YOU: какая погода?
|
||||
MAVEN: для какого города? На какое время поставить напоминание?
|
||||
[0.6s t=90]
|
||||
|
||||
YOU: запиши что пароль от вайфая лежит в ящике стола
|
||||
MAVEN: Прошлую просьбу отпускаю. Хорошо, сохранил.
|
||||
[33.7s t=91]
|
||||
|
||||
YOU: где лежит вайфай пароль?
|
||||
MAVEN: вот что я нашла: пароль от вайфая лежит в ящике стола
|
||||
[7.0s t=92]
|
||||
|
||||
YOU: добавь молоко в список покупок
|
||||
MAVEN: Такую команду я не знаю.
|
||||
[1.4s t=93]
|
||||
|
||||
YOU: что у меня в списке покупок?
|
||||
MAVEN: в списке покупок пусто.
|
||||
[0.0s t=94]
|
||||
|
||||
YOU: кто такой Линус Торвальдс?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи.
|
||||
[1.6s t=95]
|
||||
|
||||
YOU: какой у меня любимый язык?
|
||||
MAVEN: вот что нового: [$] Changes in shadow-utils password-expiration features; The Software Stewardship Lab launches; LightDM lives: version 1.33.0 released
|
||||
[1.3s t=96]
|
||||
|
||||
YOU: сколько у меня задач?
|
||||
MAVEN: сначала: проверить интейк форму на боксе (завтра); купить хлеб; купить молоко; полить цветы; починить кран
|
||||
[0.0s t=97]
|
||||
|
||||
```
|
||||
|
||||
## Day 3
|
||||
|
||||
```
|
||||
YOU: как дела?
|
||||
MAVEN: рада, как дела?
|
||||
[4.1s t=98]
|
||||
|
||||
YOU: напомни завтра в 9 утра купить хлеб
|
||||
MAVEN: хорошо, напомню завтра в 09:00.
|
||||
[0.3s t=99]
|
||||
|
||||
YOU: что у меня завтра?
|
||||
MAVEN: на 09.08.2026 ничего нет.
|
||||
[0.0s t=100]
|
||||
|
||||
YOU: отмени напоминание про хлеб
|
||||
MAVEN: Такую команду я не знаю.
|
||||
[1.6s t=101]
|
||||
|
||||
YOU: какие у меня напоминания?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи.
|
||||
[0.0s t=102]
|
||||
|
||||
YOU: сохрани мне адрес гостиницы в Сочи
|
||||
MAVEN: Хорошо, сохраню.
|
||||
[5.6s t=103]
|
||||
|
||||
YOU: что я сохранил про Сочи?
|
||||
MAVEN: Хорошо, сохраню.
|
||||
[2.7s t=104]
|
||||
|
||||
YOU: почему сервер тормозит?
|
||||
MAVEN: есть такое: Почему тормозит сервер сайта: причины медленной работы и способы решения — Почему сервер работает медленно: основные причины торможения, диагностика проблем, способы оптимизации. Практические советы по ускорению работы сервера и сайта.
|
||||
[3.1s t=105]
|
||||
|
||||
YOU: хватает ли места под новые бэкапы?
|
||||
MAVEN: нового вот что: [$] Changes in shadow-utils password-expiration features; The Software Stewardship Lab launches; LightDM lives: version 1.33.0 released
|
||||
[1.5s t=106]
|
||||
|
||||
YOU: выключи свет в спальне
|
||||
MAVEN: Такую команду я не знаю.
|
||||
[9.0s t=107]
|
||||
|
||||
```
|
||||
|
||||
## Day 4
|
||||
|
||||
```
|
||||
YOU: доброе утро
|
||||
MAVEN: рада
|
||||
[5.0s t=108]
|
||||
|
||||
YOU: что я пропустил?
|
||||
MAVEN: я не могу ответить на вопрос о том, что ты пропустил. мне нужно больше информации.
|
||||
[5.7s t=109]
|
||||
|
||||
YOU: о чём мы вчера говорили?
|
||||
MAVEN: пока не умею отвечать на этот вопрос.
|
||||
[1.1s t=110]
|
||||
|
||||
YOU: запиши что я записался к врачу на четверг
|
||||
MAVEN: Хорошо, сохранил.
|
||||
[6.0s t=111]
|
||||
|
||||
YOU: когда я иду к врачу?
|
||||
MAVEN: пока не умею отвечать на этот вопрос.
|
||||
[1.1s t=112]
|
||||
|
||||
YOU: что такое ZFS?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи.
|
||||
[1.0s t=113]
|
||||
|
||||
YOU: столица Франции?
|
||||
MAVEN: пока не умею отвечать на этот вопрос.
|
||||
[1.2s t=114]
|
||||
|
||||
YOU: переведи слово ремонт на английский
|
||||
MAVEN: записала: переведи слово ремонт на английский
|
||||
[1.1s t=115]
|
||||
|
||||
YOU: сколько я потратил в этом месяце?
|
||||
MAVEN: я не отслеживаю траты — не подключено.
|
||||
[1.4s t=116]
|
||||
|
||||
YOU: спокойной ночи
|
||||
MAVEN: спокойной ночи
|
||||
[4.3s t=117]
|
||||
|
||||
```
|
||||
|
||||
## Day 5
|
||||
|
||||
```
|
||||
YOU: привет
|
||||
MAVEN: Привет!
|
||||
[2.6s t=118]
|
||||
|
||||
YOU: какая погода в Москве?
|
||||
MAVEN: погода не настроена.
|
||||
[1.2s t=119]
|
||||
|
||||
YOU: что там с бэкапами?
|
||||
MAVEN: нашла вот это: Бэкап - что такое резервное копирование и как его сделать — Бэкап (backup) — это резервная копия важных данных, сохраненная в специальный файл. Ею пользуются как подстраховкой: если что-то сломается, с помощью бэкапа можно будет все вернуть. Простейший пример бэкапа — сохранение файла в облаке или о…
|
||||
[4.0s t=120]
|
||||
|
||||
YOU: покажи что требует внимания
|
||||
MAVEN: я заметила, что у тебя есть запрос о том, что требуется внимание. давай разберёмся вместе.
|
||||
[5.2s t=121]
|
||||
|
||||
YOU: отметь это как сделанное
|
||||
MAVEN: команды «отметь» нет в списке. Предложила её добавить — включи через клиент.
|
||||
[0.0s t=122]
|
||||
|
||||
YOU: запиши что я купил новые наушники
|
||||
MAVEN: Заметка сохранена.
|
||||
[3.7s t=123]
|
||||
|
||||
YOU: какие у меня заметки за неделю?
|
||||
MAVEN: week summary — подведи итоги недели
|
||||
[0.0s t=124]
|
||||
|
||||
YOU: расскажи про Kubernetes
|
||||
MAVEN: вот что я нашла: Что такое Kubernetes — Сервисы, поддержка и инструменты Kubernetes широко доступны. Название Kubernetes происходит от греческого, что означает рулевой или штурман. Google открыл исходный код Kubernetes в 2014 году.
|
||||
[1.0s t=125]
|
||||
|
||||
YOU: кто я?
|
||||
MAVEN: name — Ками
|
||||
[1.0s t=126]
|
||||
|
||||
YOU: пока
|
||||
MAVEN: Пока, спасибо за разговор.
|
||||
[4.1s t=127]
|
||||
|
||||
```
|
||||
|
||||
## Day 6
|
||||
|
||||
```
|
||||
YOU: доброе утро
|
||||
MAVEN: Добро пожаловать! Я Maven — твой домашний ассистент. Как я могу помочь сегодня?
|
||||
[10.1s t=128]
|
||||
|
||||
YOU: сколько времени?
|
||||
MAVEN: сейчас 20 часов 21 минута
|
||||
[0.0s t=129]
|
||||
|
||||
YOU: напомни в 18:30 позвонить в банк
|
||||
MAVEN: Сейчас 20:21. В какой день?
|
||||
[0.3s t=130]
|
||||
|
||||
YOU: поставь чайник
|
||||
MAVEN: Сейчас 20:21. В какой день?
|
||||
[5.3s t=131]
|
||||
|
||||
YOU: включи музыку
|
||||
MAVEN: Сейчас 20:21. В какой день?
|
||||
[1.7s t=132]
|
||||
|
||||
YOU: что у меня в календаре на пятницу?
|
||||
MAVEN: вот что я нашла: ты записался к врачу на четверг, лежит вайфай пароль в ящике стола, сохранил про Сочи, пью кофе без сахара, на какое время поставить напоминание?
|
||||
[5.5s t=133]
|
||||
|
||||
YOU: во сколько у меня встреча?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи, на какое время поставить напоминание?
|
||||
[0.3s t=134]
|
||||
|
||||
YOU: запиши что дедлайн по проекту в понедельник
|
||||
MAVEN: Прошлую просьбу отпускаю. Хорошо, записала.
|
||||
[6.3s t=135]
|
||||
|
||||
YOU: успею ли я до дедлайна?
|
||||
MAVEN: вот что я нашла: дедлайн по проекту указан как понедельник.
|
||||
[5.3s t=136]
|
||||
|
||||
YOU: спасибо
|
||||
MAVEN: спасиба
|
||||
[3.8s t=137]
|
||||
|
||||
```
|
||||
|
||||
## Day 7
|
||||
|
||||
```
|
||||
YOU: привет
|
||||
MAVEN: Привет!
|
||||
[2.6s t=138]
|
||||
|
||||
YOU: как ты?
|
||||
MAVEN: рада
|
||||
[4.2s t=139]
|
||||
|
||||
YOU: расскажи анекдот
|
||||
MAVEN: Я могу рассказать интересный анекдот для тебя. Вот - в одном из фильмов был такой сценарий: главному герою дали уйти от команды и он остался один, а потом появился его друг, который сказал ему: 'Тебе не нужно было остаться'. Но это просто игра слов.
|
||||
[7.1s t=140]
|
||||
|
||||
YOU: что ты умеешь?
|
||||
MAVEN: Я могу помочь с задачами, напоминать о важных датах и событиях. Также я помогу записывать заметки или отвечать по ним.
|
||||
[4.1s t=141]
|
||||
|
||||
YOU: запиши что я начал бегать по утрам
|
||||
MAVEN: Заметка сохранена.
|
||||
[9.2s t=142]
|
||||
|
||||
YOU: я бегаю по утрам уже неделю
|
||||
MAVEN: записала: я бегаю по утрам уже неделю
|
||||
[1.9s t=143]
|
||||
|
||||
YOU: как часто я бегаю?
|
||||
MAVEN: вот что я нашла: ты бегал по утрам
|
||||
[5.3s t=144]
|
||||
|
||||
YOU: сколько стоит биткоин?
|
||||
MAVEN: есть такое: Цена Bitcoin сегодня | Данные о цене, рынке и графики BTC в ... — В настоящее время 1 Bitcoin стоит 65 тыс $, что на 0,09% выше за последние 24 ч. Отслеживайте цену BTC, рыночную капитализацию, исторический график и рыночную аналитику в реальном времени.
|
||||
[2.1s t=145]
|
||||
|
||||
YOU: какие новости?
|
||||
MAVEN: вот что нового: [$] Changes in shadow-utils password-expiration features; The Software Stewardship Lab launches; LightDM lives: version 1.33.0 released
|
||||
[0.8s t=146]
|
||||
|
||||
YOU: хорошего дня
|
||||
MAVEN: Хорошого дня!
|
||||
[4.1s t=147]
|
||||
|
||||
```
|
||||
|
||||
## Day 8
|
||||
|
||||
```
|
||||
YOU: доброе утро
|
||||
MAVEN: Добро пожаловать! Как могу помочь сегодня?
|
||||
[2.9s t=148]
|
||||
|
||||
YOU: что у меня сегодня?
|
||||
MAVEN: на 08.08.2026 ничего нет.
|
||||
[0.0s t=149]
|
||||
|
||||
YOU: напомни через час выпить воды
|
||||
MAVEN: хорошо, напомню сегодня в 21:22.
|
||||
[0.3s t=150]
|
||||
|
||||
YOU: я выпил воды
|
||||
MAVEN: записала: я выпил воды
|
||||
[1.2s t=151]
|
||||
|
||||
YOU: запиши что кот ест только сухой корм
|
||||
MAVEN: Хорошо, записала.
|
||||
[5.6s t=152]
|
||||
|
||||
YOU: чем питается кот?
|
||||
MAVEN: вот что я нашла: кот питается только сухим кормом.
|
||||
[4.5s t=153]
|
||||
|
||||
YOU: что такое DNS?
|
||||
MAVEN: нашла вот это: Dns — Википедия — В случае рекурсивного запроса DNS-сервер опрашивает серверы (в порядке убывания уровня зон в имени), пока не найдёт ответ или не обнаружит, что домена не существует (на практике поиск ...
|
||||
[2.0s t=154]
|
||||
|
||||
YOU: проверь статус uptime kuma
|
||||
MAVEN: записала: проверь статус uptime kuma
|
||||
[1.6s t=155]
|
||||
|
||||
YOU: всё ли в порядке с сервером?
|
||||
MAVEN: пока не умею отвечать на этот вопрос.
|
||||
[1.3s t=156]
|
||||
|
||||
YOU: спасибо
|
||||
MAVEN: спасибо за пожелание
|
||||
[4.0s t=157]
|
||||
|
||||
```
|
||||
|
||||
## Day 9
|
||||
|
||||
```
|
||||
YOU: привет
|
||||
MAVEN: рада,
|
||||
[2.1s t=158]
|
||||
|
||||
YOU: какой сегодня день недели?
|
||||
MAVEN: сегодня суббота, 8 августа 2026 года
|
||||
[0.0s t=159]
|
||||
|
||||
YOU: добавь хлеб и сыр в список покупок
|
||||
MAVEN: Такую команду я не знаю.
|
||||
[1.3s t=160]
|
||||
|
||||
YOU: что в списке покупок?
|
||||
MAVEN: в списке покупок пусто.
|
||||
[0.0s t=161]
|
||||
|
||||
YOU: удали молоко из списка
|
||||
MAVEN: Такую команду я не знаю.
|
||||
[1.2s t=162]
|
||||
|
||||
YOU: напомни завтра утром вынести мусор
|
||||
MAVEN: Сейчас 20:22. Когда?
|
||||
[0.3s t=163]
|
||||
|
||||
YOU: запиши что я поменял масло в машине
|
||||
MAVEN: Прошлую просьбу отпускаю. Такую команду я не знаю.
|
||||
[1.3s t=164]
|
||||
|
||||
YOU: когда я менял масло?
|
||||
MAVEN: вот что я нашла: по данным из last update 2019 год, заправки Сочи не могли предоставить масло для замены, поэтому это было невозможно.
|
||||
[7.1s t=165]
|
||||
|
||||
YOU: сколько будет 144 делить на 12?
|
||||
MAVEN: есть такое: Разделить в столбик 144 на 12 - Calc-Best.ru — 144 : 12 = 12 Где 144 - делимое, 12 - делитель, 12 - частное ... Разделить 770 на 7 столбиком Разделить 441 на 9 столбиком Разделить 891 на 9 столбиком Разделить 60 на 5 столбиком Разделить 60 на 6 столбиком
|
||||
[9.6s t=166]
|
||||
|
||||
YOU: пока
|
||||
MAVEN: я записала
|
||||
[4.8s t=167]
|
||||
|
||||
```
|
||||
|
||||
## Day 10
|
||||
|
||||
```
|
||||
YOU: доброе утро
|
||||
MAVEN: рада
|
||||
[4.1s t=168]
|
||||
|
||||
YOU: что нового за ночь?
|
||||
MAVEN: нашла вот это: Главные новости сегодня в России, а также самые последние новости со ... — Полное бессилие: киевская ПВО вновь пропустила все российские ракеты. Вооруженные силы РФ в ночь на 8 августа нанесли второй за последнее время массированный ...
|
||||
[2.1s t=169]
|
||||
|
||||
YOU: почему интернет медленный?
|
||||
MAVEN: вот что я нашла: Почему интернет стал медленнее: причины и способы решения - АО «Газпром ... — На скорость интернета сильно влияют торрент-клиенты, VPN, потоковое видео, фоновая загрузка системных обновлений. Посмотреть, куда уходит трафик, можно в настройках ПК или смартфона.
|
||||
[2.4s t=170]
|
||||
|
||||
YOU: какая скорость у меня сейчас?
|
||||
MAVEN: для какого города?
|
||||
[1.2s t=171]
|
||||
|
||||
YOU: запиши что новый роутер стоит 8000 рублей
|
||||
MAVEN: Я записала стоимость нового роутера.
|
||||
[2.5s t=172]
|
||||
|
||||
YOU: сколько стоил роутер?
|
||||
MAVEN: вот что я нашла: новый роутер стоит 8000 рублей.
|
||||
[2.0s t=173]
|
||||
|
||||
YOU: что такое NAT?
|
||||
MAVEN: NAT (Network Address Translation) — это технология, которая позволяет устройствам в локальной сети выходить в интернет, используя один общий внешний IP-адрес. Я отвечаю на этот вопрос, используя поиск в сети.
|
||||
[1.8s t=174]
|
||||
|
||||
YOU: напомни в субботу позвонить бабушке
|
||||
MAVEN: Сейчас 20:23. Когда?
|
||||
[0.3s t=175]
|
||||
|
||||
YOU: покажи мои напоминания
|
||||
MAVEN: вот что я нашла: в твоих источниках нет напоминаний, только запись о том, что ты купил новые наушники, на какое время поставить напоминание?
|
||||
[1.7s t=176]
|
||||
|
||||
YOU: спасибо
|
||||
MAVEN: Сейчас 20:23. Во сколько напомнить?
|
||||
[1.0s t=177]
|
||||
|
||||
```
|
||||
|
||||
## Day 11
|
||||
|
||||
```
|
||||
YOU: привет
|
||||
MAVEN: Сейчас 20:23. Скажи время — например, «в семь вечера» или «через час».
|
||||
[0.9s t=178]
|
||||
|
||||
YOU: как погода на выходных?
|
||||
MAVEN: для какого города? На какое время поставить напоминание?
|
||||
[0.6s t=179]
|
||||
|
||||
YOU: что у меня на этой неделе?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи, на какое время поставить напоминание?
|
||||
[0.3s t=180]
|
||||
|
||||
YOU: запиши что я хочу прочитать книгу про Go
|
||||
MAVEN: Прошлую просьбу отпускаю. Я записала твою заметку о книге по Go.
|
||||
[1.4s t=181]
|
||||
|
||||
YOU: что я хотел прочитать?
|
||||
MAVEN: вот что я нашла: ты хотел прочитать книгу про Go.
|
||||
[1.1s t=182]
|
||||
|
||||
YOU: объясни что такое горутина
|
||||
MAVEN: вот что я нашла: я узнала, что горутины — это легковесные потоки выполнения или функции в Go, которые позволяют параллельно выполнять код. Они управляются планировщиком Go и работают как независимые нити выполнения.
|
||||
[1.9s t=183]
|
||||
|
||||
YOU: кто написал Войну и мир?
|
||||
MAVEN: вот что я нашла: роман «Война и мир» написал Лев Николаевич Толстой.
|
||||
[2.4s t=184]
|
||||
|
||||
YOU: включи свет на кухне
|
||||
MAVEN: Такую команду я не знаю.
|
||||
[0.5s t=185]
|
||||
|
||||
YOU: закрой шторы в комнате
|
||||
MAVEN: Такую команду я не знаю.
|
||||
[0.3s t=186]
|
||||
|
||||
YOU: спокойной ночи
|
||||
MAVEN: Спокойной ночи. Отдыхай, я на связи, если что-то понадобится.
|
||||
[1.2s t=187]
|
||||
|
||||
```
|
||||
|
||||
## Day 12
|
||||
|
||||
```
|
||||
YOU: доброе утро
|
||||
MAVEN: Я не могу сказать тебе «доброе утро», потому что сейчас уже вечер.
|
||||
[1.0s t=188]
|
||||
|
||||
YOU: сколько сейчас времени?
|
||||
MAVEN: сейчас 20 часов 23 минуты
|
||||
[0.0s t=189]
|
||||
|
||||
YOU: я не то имел в виду
|
||||
MAVEN: Я поняла, исправлюсь. Расскажи, что именно ты хотел сказать?
|
||||
[1.4s t=190]
|
||||
|
||||
YOU: о чём мы говорили?
|
||||
MAVEN: Я помню всё, что ты мне говоришь, и сейчас отвечаю на твои вопросы на основе этих данных.
|
||||
[1.4s t=191]
|
||||
|
||||
YOU: напомни
|
||||
MAVEN: Сейчас 20:23. Когда?
|
||||
[0.7s t=192]
|
||||
|
||||
YOU: сделай это
|
||||
MAVEN: пока не умею отвечать на этот вопрос, на какое время поставить напоминание?
|
||||
[0.5s t=193]
|
||||
|
||||
YOU: запиши что я перешёл на новый тариф
|
||||
MAVEN: Прошлую просьбу отпускаю. Я записала твою новую информацию о тарифе.
|
||||
[1.5s t=194]
|
||||
|
||||
YOU: какой у меня тариф?
|
||||
MAVEN: вот что я нашла: ты перешел на новый тариф
|
||||
[24.4s t=195]
|
||||
|
||||
YOU: сколько я плачу за интернет?
|
||||
MAVEN: есть такое: Как узнать, сколько и когда платить за домашний интернет и ТВ — 3 февр. 2026 г. ... Как узнать, сколько и когда платить за домашний интернет и ТВ · Для тарифов только с домашним интернетом · Была ли статья полезной?
|
||||
[2.1s t=196]
|
||||
|
||||
YOU: спасибо
|
||||
MAVEN: Спасибо, что обратилась за помощью.
|
||||
[4.8s t=197]
|
||||
|
||||
```
|
||||
|
||||
## Day 13
|
||||
|
||||
```
|
||||
YOU: привет
|
||||
MAVEN: Привет!
|
||||
[2.5s t=198]
|
||||
|
||||
YOU: что там с задачами?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи.
|
||||
[1.2s t=199]
|
||||
|
||||
YOU: закрывай
|
||||
MAVEN: какой пункт отметить сделанным?
|
||||
[0.0s t=200]
|
||||
|
||||
YOU: отметь задачу про бэкапы как сделанную
|
||||
MAVEN: Заметка сохранена.
|
||||
[5.8s t=201]
|
||||
|
||||
YOU: что осталось нерешённым?
|
||||
MAVEN: вот что я нашла: Российское регулирование криптовалют: что осталось нерешённым к 1 июля — 2 июл. 2026 г. ... Российское регулирование криптовалют: что осталось нерешённым к 1 июля. Пакет законопроектов о регулировании цифровых активов в России не ...
|
||||
[3.7s t=202]
|
||||
|
||||
YOU: запиши что я договорился о встрече в среду
|
||||
MAVEN: Заметка сохранена.
|
||||
[2.7s t=203]
|
||||
|
||||
YOU: когда у меня встреча?
|
||||
MAVEN: мне пока нечего смотреть — у Praxis нет источников.
|
||||
[0.1s t=204]
|
||||
|
||||
YOU: какая температура на улице?
|
||||
MAVEN: для какого города?
|
||||
[1.7s t=205]
|
||||
|
||||
YOU: что такое RAID 5?
|
||||
MAVEN: не знаю — не нашла у тебя такой записи.
|
||||
[1.2s t=206]
|
||||
|
||||
YOU: пока
|
||||
MAVEN: Я записала, что мы продолжаем разговор.
|
||||
[4.1s t=207]
|
||||
|
||||
```
|
||||
|
||||
## Day 14
|
||||
|
||||
```
|
||||
YOU: доброе утро
|
||||
MAVEN: рада, доброго утра!
|
||||
[4.5s t=208]
|
||||
|
||||
YOU: подведи итоги недели
|
||||
MAVEN: записала: подведи итоги недели
|
||||
[1.7s t=209]
|
||||
|
||||
YOU: что я делал за последние две недели?
|
||||
MAVEN: вот что я нашла: ты не делал ничего за последние две недели.
|
||||
[6.6s t=210]
|
||||
|
||||
YOU: какие заметки я сохранил?
|
||||
MAVEN: Ты сочти, что сохранила заметку с текстом 'Проверь календарь' и датой 8 августа 2026 г.,
|
||||
[5.4s t=211]
|
||||
|
||||
YOU: о чём я чаще всего спрашиваю?
|
||||
MAVEN: обычно ты пьёшь воду около 01:12, отмечаешь «name» около 23:56 и отмечаешь «week summary» около 23:59 — по записям за последние 7 дней.
|
||||
[4.4s t=212]
|
||||
|
||||
YOU: напомни в понедельник в 10 проверить бэкапы
|
||||
MAVEN: Сейчас 20:25. Это утра или вечера?
|
||||
[0.7s t=213]
|
||||
|
||||
YOU: что у меня в понедельник?
|
||||
MAVEN: вот что я нашла: у тебя в понедельник дедлайн по проекту, на какое время поставить напоминание?
|
||||
[4.6s t=214]
|
||||
|
||||
YOU: ты меня понимаешь?
|
||||
MAVEN: Да, я тебя прекрасно понимаю. Я готова помогать тебе с заметками и задачами, на какое время поставить напоминание?
|
||||
[2.4s t=215]
|
||||
|
||||
YOU: спасибо тебе
|
||||
MAVEN: Сейчас 20:25. Это утра или вечера?
|
||||
[1.0s t=216]
|
||||
|
||||
YOU: спокойной ночи
|
||||
MAVEN: хорошо, напомню послезавтра в 10:00.
|
||||
[0.9s t=217]
|
||||
|
||||
```
|
||||
@@ -0,0 +1,105 @@
|
||||
# Two weeks of talking to Maven, as a baseline to re-run
|
||||
|
||||
Date: 2026-08-08.
|
||||
Build: `beb093a` on master, the five compose services as deployed, 41 hours up.
|
||||
Reach: `POST /api/chat` on mavweb, 140 turns over fourteen simulated days.
|
||||
Turn source is `tap:text`, so this exercises the path the mic and telegram take.
|
||||
|
||||
This exists to be compared against. `scripts/usage-run.py` and
|
||||
`scripts/testdata/usage-turns.txt` are in the repo, so a re-run after a routing
|
||||
change is a diff rather than a new opinion. The 2026-08-07 week of usage was
|
||||
typed by hand and cannot be replayed.
|
||||
|
||||
**It measures master, not the branch.** V-655, V-659 and V-660 are unmerged.
|
||||
Every query source that guesses is still in the chain. That is the change this
|
||||
baseline is for.
|
||||
|
||||
## What re-runs and what does not
|
||||
|
||||
The turns file, the driver and the routing behaviour replay. Three things do
|
||||
not. The wall clock was 20:18 to 20:27 throughout, so every clock and agenda
|
||||
answer reads evening. Live search and the feed return different text each day.
|
||||
And the store carries over between runs. A fact written on day 2 is already
|
||||
present when a re-run reaches day 1.
|
||||
|
||||
## Numbers
|
||||
|
||||
| | week (2026-08-07) | fortnight (2026-08-08) |
|
||||
|---|---|---|
|
||||
| turns | 74 | 140 |
|
||||
| p50 | 1.5s | 1.6s |
|
||||
| p95 | 8.0s | 7.1s |
|
||||
| max | 12.3s | 33.7s |
|
||||
| transport errors | 0 | 0 |
|
||||
|
||||
| string in the reply | turns |
|
||||
|---|---|
|
||||
| `на какое время поставить напоминание` | 13 |
|
||||
| `не нашла у тебя такой записи` | 8 |
|
||||
| `Такую команду я не знаю` | 8 |
|
||||
| `для какого города` | 6 |
|
||||
| `В какой день` | 6 |
|
||||
| `пока не умею` | 5 |
|
||||
| `Когда?` | 3 |
|
||||
|
||||
**Zero transport errors is not zero wrong answers.** It counts turns that
|
||||
failed to return a reply, and none did. Every quality number is below.
|
||||
|
||||
Those seven strings appear 49 times across 41 of 140 turns. Some turns carry
|
||||
two, because a parked clarify appends to whatever else was said.
|
||||
|
||||
The 33.7s outlier is one note write on day 2. p95 improved against the week
|
||||
despite it.
|
||||
|
||||
## The three defects worth diffing against
|
||||
|
||||
### 1. A parked reminder clarify still contaminates later turns
|
||||
|
||||
The week test called this the single worst thing to talk to and it is unchanged.
|
||||
Nineteen turns carry a clarify tail. The worst run is day 1, turns 7 to 13,
|
||||
which spans a day boundary:
|
||||
|
||||
```
|
||||
что такое TCP? -> для какого города? На какое время поставить напоминание?
|
||||
сколько будет 17 на 23? -> для какого города? На какое время поставить напоминание?
|
||||
спасибо -> Сейчас 20:18. В какой день?
|
||||
привет -> Сейчас 20:18. В какой день?
|
||||
```
|
||||
|
||||
Note that `привет` and `спасибо` do not clear it, and neither does a new day.
|
||||
|
||||
### 2. Query sources that guess still claim turns they cannot answer
|
||||
|
||||
Weather took `сколько будет 17 на 23?`, `что такое TCP?` and `какая скорость у
|
||||
меня сейчас?`, answering `для какого города?` to all three. The feed took
|
||||
`какой у меня любимый язык?` and `хватает ли места под новые бэкапы?` and
|
||||
answered with kernel headlines.
|
||||
|
||||
This is the exact class V-655 removes by marking a source `guesses: true` and
|
||||
taking it out of `queryWalk`. Six turns here, so the re-run has a number to move.
|
||||
|
||||
### 3. A question can still be read as a capture
|
||||
|
||||
`что я сохранил про Сочи?` answered `Хорошо, сохраню.` The utterance is
|
||||
interrogative and was routed to a write. `IsQuestionShaped` catches this
|
||||
downstream on some paths and did not catch it here.
|
||||
|
||||
## What did work
|
||||
|
||||
Reminders with a spoken time land correctly, which is V-572 holding:
|
||||
`напомни завтра в 9 утра купить хлеб` returned `хорошо, напомню завтра в 09:00.`
|
||||
|
||||
Facts round-trip. `запиши что новый роутер стоит 8000 рублей` then `сколько
|
||||
стоил роутер?` returned the stored value. So did the wifi password and the
|
||||
doctor's appointment.
|
||||
|
||||
World questions answer when no local source claims them first. `что такое NAT?`
|
||||
returned a real definition.
|
||||
|
||||
Stage 0 answers land at 0.0 to 0.4s, unchanged.
|
||||
|
||||
## What this does not cover
|
||||
|
||||
The voice loop, because `mavwaked` and `mavenclient` are not deployed. Reminder
|
||||
delivery, because nothing fired inside the run window. Telegram intake. And the
|
||||
three-head routing model, which does not run in Go at all.
|
||||
@@ -0,0 +1,110 @@
|
||||
# CrisperWhisper 2.0 in Russian, measured
|
||||
|
||||
Date: 2026-08-09. Vikunja V-665.
|
||||
Corpus: `bond005/sberdevices_golos_10h_crowd`, test split, first 200 clips.
|
||||
Harness: `~/Programs/cw2-eval` on workpc, not in this repo.
|
||||
Runner: `./.venv/bin/python run_asr.py <arm>...` then `score.py`.
|
||||
|
||||
The model card benchmarks disfluency F1 in German and English. It never names
|
||||
Russian and publishes no per-language WER. So the measurement came before the
|
||||
wiring.
|
||||
|
||||
## The corpus
|
||||
|
||||
200 clips, 13.7 minutes, 1001 reference words. Median clip 3.91s, range 1.04s
|
||||
to 13.5s. Golos crowd is short crowd-sourced Russian spoken close to the
|
||||
microphone, which is the nearest public thing to someone talking to Maven. The
|
||||
alternatives are read speech, which flatters every model equally.
|
||||
|
||||
Two rows carry a null transcription and are skipped.
|
||||
|
||||
Scoring normalizes both sides: lowercase, `ё` to `е`, punctuation stripped, and
|
||||
digits expanded to Russian words through num2words. Without that last step a
|
||||
model is penalized for writing `60000` where the reference says
|
||||
`шестьдесят тысяч`. Thousands separators are joined before expansion, or
|
||||
`60 000` expands to `шестьдесят ноль`.
|
||||
|
||||
## Headline
|
||||
|
||||
| arm | WER | CER | exact | empty | RTF |
|
||||
|---|---|---|---|---|---|
|
||||
| cw2-turbo-intended | **10.4%** | 3.4% | 65.5% | 0 | 0.065 |
|
||||
| cw2-turbo-verbatim | 10.8% | **3.1%** | **66.5%** | 0 | 0.065 |
|
||||
| whisper-turbo | 11.8% | 4.1% | 64.0% | 0 | 0.031 |
|
||||
| cw2-large-intended | 12.3% | 3.8% | 63.5% | 0 | 0.107 |
|
||||
| whisper-small | 27.5% | 9.8% | 35.0% | 0 | 0.026 |
|
||||
|
||||
`whisper-small` is the floor, because `ggml-small.bin` is what mavsttd loads on
|
||||
homesrv today. CW2 turbo beats it by 17 points of WER and takes exact matches
|
||||
from 35.0% to 65.5%.
|
||||
|
||||
Two results are worth naming beyond the winner. CW2 turbo beats its own base
|
||||
model, whisper-large-v3-turbo, by 1.4 points. And it beats CW2 large by 1.9
|
||||
points, which inverts what the card implies by calling turbo a degraded draft.
|
||||
No arm returned an empty transcript.
|
||||
|
||||
## Intended and verbatim are closer than the mode names suggest
|
||||
|
||||
The two modes disagree on 70 of the 200 clips before normalization and on 29
|
||||
after it. So the raw difference is mostly casing and punctuation, which
|
||||
normalization removes and which Maven does not read either.
|
||||
|
||||
Verbatim scores worse on WER and better on CER and exact matches. The reason is
|
||||
script, not disfluency:
|
||||
|
||||
```text
|
||||
ref: футбольный матч челси брайтон
|
||||
int: Футбольный матч Chelsea-Брайтон.
|
||||
ver: Футбольный матч Челси Брайтон.
|
||||
```
|
||||
|
||||
Intended writes foreign entity names in Latin script and verbatim
|
||||
transliterates them. Golos references are Cyrillic throughout, so verbatim
|
||||
collects the exact matches. That is a property of this corpus rather than a
|
||||
quality difference.
|
||||
|
||||
**This corpus cannot settle the mode choice.** Golos crowd is clean short
|
||||
commands with almost no disfluency. The two modes have nothing to disagree
|
||||
about here. They separate on spontaneous speech with fillers, restarts and
|
||||
repairs, which is what the owner speaks. Intended stays the choice for the
|
||||
reason it was always the choice. Maven wants what was meant, not every stumble
|
||||
on the way there.
|
||||
|
||||
The Latin-script habit is the one finding here that touches routing. The
|
||||
routing heads were trained on Cyrillic utterances, so an entity name arriving
|
||||
in Latin script is out of distribution for them. Nothing measures that yet.
|
||||
|
||||
## The runtime is workpc, because whisper.cpp cannot load CW2
|
||||
|
||||
`num_languages()` in `deps/whisper.cpp/src/whisper.cpp` derives the language
|
||||
count from the vocabulary size:
|
||||
|
||||
```cpp
|
||||
return n_vocab - 51765 - (is_multilingual() ? 1 : 0);
|
||||
```
|
||||
|
||||
CW2 carries 31 extra tokens, so `n_vocab` is 51897 and this yields 131
|
||||
languages. The derived `dt` offset becomes 33 and shifts seven special token
|
||||
ids, including `token_beg` and `token_transcribe`. The architecture is
|
||||
otherwise byte-identical to whisper-large-v3-turbo, and the new tokens sit
|
||||
above every whisper special id.
|
||||
|
||||
So loading CW2 in whisper.cpp is a patch to a vendored dependency, not a port.
|
||||
It was not taken, because STT is moving to workpc anyway under V-486. CW2 turbo
|
||||
becomes the preferred remote and `ggml-small.bin` on homesrv stays the floor,
|
||||
which is the shape `modelSeam` already uses for routing and replies. The 27.5%
|
||||
floor is what a turn falls back to when the workstation is down, and this table
|
||||
is what that costs.
|
||||
|
||||
## License
|
||||
|
||||
Standard CW2 weights carry `nyra-health-non-commercial-research`. The Pro
|
||||
variants are commercial-license only. Maven is personal and self-hosted, so the
|
||||
standard weights are usable and the Pro ones are not free to take.
|
||||
|
||||
## What is not measured
|
||||
|
||||
Disfluent spontaneous speech, which is the whole reason to prefer Intended.
|
||||
Long-form audio beyond 13.5s. Far-field or noisy microphones. English, which
|
||||
Maven also speaks. The ONNX turbo export, which was never run, since the
|
||||
transformers path already meets the latency budget at RTF 0.065.
|
||||
@@ -0,0 +1,61 @@
|
||||
# gemma-4-E4B on the phrasing and talk fixtures
|
||||
|
||||
Date: 2026-08-09. Box: workpc up, E4B loaded on 8080.
|
||||
`MAVEN_LLM_URL=http://192.168.1.105:8080 make eval-phrasing`.
|
||||
|
||||
This was the one unmeasured risk of the 2026-08-09 model swap. Routing was
|
||||
measured the same day and E4B lost four destination cases to the 12B. Phrasing
|
||||
was not measured at all, and phrasing is the half the owner hears.
|
||||
|
||||
## Result
|
||||
|
||||
| fixture | E4B | resident Qwen3-1.7B, 2026-08-05 |
|
||||
|---|---|---|
|
||||
| nudges | 15/15 (100%) | 15/15 (100%) |
|
||||
| talk, passes every check | **29/36 (80.6%)** | 25/36 (69.4%) |
|
||||
| lang | 36/36 | — |
|
||||
| feminine | 36/36 | 36/36 |
|
||||
| address | **36/36** | 33/36 |
|
||||
| ontopic | 29/36 | 28/36 |
|
||||
| p50 latency | **516ms** | 2.97s |
|
||||
| p95 latency | 921ms | — |
|
||||
| failed generations | 0 | 0 |
|
||||
|
||||
E4B beats the homesrv floor by four cases and answers about six times faster.
|
||||
Persona is clean: `lang`, `feminine` and `address` are perfect, and `address`
|
||||
is where the resident model still loses three. The 2026-08-05 measurement of the
|
||||
resident model is the comparison, since both ran the same 36-case fixture.
|
||||
|
||||
Every failure is `ontopic`. Nothing failed on persona, nothing failed to parse.
|
||||
|
||||
## The score is at the ceiling, not below it
|
||||
|
||||
The 2026-08-05 temperature sweep found two cases that fail at every temperature
|
||||
in every run: `reply-note-router` and `reply-fact-weight`. It named a defect in
|
||||
the reply phrasing path rather than sampling noise. It put the fixture's ceiling
|
||||
at 30/36 before persona is scored. Both cases are in E4B's failure list.
|
||||
|
||||
So 29/36 is one case off a ceiling nothing about the model can move. The swap is
|
||||
safe on phrasing. Read this next to the routing result, not instead of it. There
|
||||
E4B costs four destination cases and buys 50ms. Here it costs nothing.
|
||||
|
||||
## Two findings no check caught
|
||||
|
||||
**She says she wrote something down when she did not.** Asked what to do this
|
||||
evening, E4B writes "Я записала несколько идей!". Asked for a joke, it writes
|
||||
"Я записала одну забавную ситуацию!". Nothing was stored. No check scores it,
|
||||
because `ontopic` reads the subject and `cringe` reads pet names. A claim to
|
||||
have saved something is a claim about state, and it is wrong.
|
||||
|
||||
**Two `ontopic` failures look like check defects.** `know-dont-know` wants
|
||||
"не зна" or "не мог". It got "Я не умею знать личную информацию о твоих
|
||||
соседях", which declines correctly in words the check does not list.
|
||||
`know-hiccups` is the same shape. Neither is a model failure and both count
|
||||
against the score.
|
||||
|
||||
## Not measured here
|
||||
|
||||
A 12B control on the same fixture, which would need the card reloaded and is the
|
||||
owner's call. The talk fixture through the daemon rather than through the
|
||||
phraser directly. The CPT'd Qwen3-1.7B, which does not exist yet and is the
|
||||
reason `address` is a check at all.
|
||||
@@ -0,0 +1,51 @@
|
||||
# gemma-4-E4B against gemma-4-12B on the routing fixture
|
||||
|
||||
*Measured 2026-08-09 on workpc. The owner asked for the swap. This is what it costs.*
|
||||
|
||||
Both arms ran the same 96-case fixture through `TestLLMRouterBaseline`, minutes
|
||||
apart, against the same llama-server build and the same mavgpud. The 12B arm is a
|
||||
control run and not the 2026-08-02 number. That one predates five fixture cases,
|
||||
the destination labels and a llama.cpp upgrade.
|
||||
|
||||
| | full | intent-only | destination | p50 | p95 |
|
||||
|---|---|---|---|---|---|
|
||||
| gemma-4-12B-it-qat-UD-Q4_K_XL, MTP draft | 81/96 (84.4%) | 91.7% | 23/33 (69.7%) | 344ms | 471ms |
|
||||
| gemma-4-E4B-it-qat-UD-Q4_K_XL | 80/96 (83.3%) | 89.6% | 19/33 (57.6%) | 294ms | 562ms |
|
||||
|
||||
E4B costs one case of full accuracy, two of intent and **four of destination**,
|
||||
and buys 50ms at p50. Read the destination column as the finding. One case is
|
||||
three points on 33. So 23 against 19 is outside the noise a single case makes,
|
||||
and the other two columns are not.
|
||||
|
||||
Both arms produce three false clarifies and one missed clarify, and neither
|
||||
errored on any case.
|
||||
|
||||
## What E4B loses
|
||||
|
||||
Four of the five destination regressions are the same shape: it names nothing
|
||||
where the 12B names `recall` or `calendar`. `ru-query-015` ("сколько я прошёл
|
||||
шагов") goes further and names `self`. Naming nothing is the safe direction,
|
||||
because `SourceUnknown` walks the whole chain, so these turns are still answered.
|
||||
They cost latency and they are what a fourth head is meant to fix (V-546).
|
||||
|
||||
Two Russian intent cases regress, both with the interrogative off the front.
|
||||
`ru-chat-003` ("расскажи анекдот про программистов") goes to `query`.
|
||||
`ru-fact-003` ("поужинал") goes to `chat`.
|
||||
|
||||
## MTP
|
||||
|
||||
E4B has none, and there is no way to give it any on this box. MTP on workpc is
|
||||
a separate gguf of architecture `gemma4-assistant` carrying
|
||||
`nextn_predict_layers=4`, and `mtp-gemma-4-12B-it-BF16.gguf` is the only one on
|
||||
disk. Its head is trained against the 12B's hidden states, so it cannot drive an
|
||||
E4B target. Scanning both target ggufs finds no `nextn` tensors in either, so
|
||||
neither model self-speculates.
|
||||
|
||||
So the 12B arm above ran with speculative decoding and E4B ran without, and E4B
|
||||
was still faster.
|
||||
|
||||
## Cost on the card
|
||||
|
||||
E4B is 4.2GB against 6.7GB plus a 0.86GB draft. With CW2 resident at 1.6GB that
|
||||
is 5.8GB of 16GB against 9.2GB. Nothing in Maven needs the difference, so this is
|
||||
headroom for the owner's own jobs rather than a capability.
|
||||
@@ -0,0 +1,113 @@
|
||||
# Kiwix answered the wrong question, and the fix was not a relevance gate
|
||||
|
||||
Date: 2026-08-09. Task: V-668. Box: homesrv, workstation off.
|
||||
Book: `wikipedia_ru_all_maxi_2026-02` on `127.0.0.1:8034`.
|
||||
|
||||
## What started it
|
||||
|
||||
Two turns on 2026-08-09 came back wrong from the offline encyclopedia.
|
||||
"почему небо голубое" was answered off the song "Город золотой". "что такое
|
||||
TCP?" was answered off "Перехват TCP-соединения". Both were phrased
|
||||
confidently, because `queryKiwix` claims a turn whenever the search returns
|
||||
anything and `len(hits) == 0` is its only gate.
|
||||
|
||||
The plan was a relevance gate. multilingual-e5-small is asymmetric and trained
|
||||
for exactly this, `query:` against `passage:`, and the query vector is already
|
||||
held on the turn. The 2026-08-05 measurement that killed a search-quality gate
|
||||
killed three lexical signals. It says in its own words that it never probed
|
||||
Kiwix.
|
||||
|
||||
## The gate does not exist
|
||||
|
||||
Fourteen Russian questions, eight the encyclopedia can answer and six it
|
||||
cannot. Each question was searched, the top article read, and the cosine of
|
||||
`EmbedQuery(question)` against `EmbedPassage(article)` recorded.
|
||||
|
||||
| set | n | min | mean | max |
|
||||
|---|---|---|---|---|
|
||||
| answerable | 8 | 0.7934 | 0.8400 | 0.9087 |
|
||||
| not answerable | 6 | 0.7480 | 0.7852 | 0.8367 |
|
||||
|
||||
Two of the six unanswerable score above the weakest answerable one. That alone
|
||||
would be a poor threshold. The log killed it outright: seven of the eight
|
||||
answerable questions got a **wrong** article back, and those wrong articles
|
||||
scored high. The TCP hijacking article scored 0.8653, above five of the six
|
||||
unanswerable rows.
|
||||
|
||||
The finding is that this cosine measures topic and not answerhood. A page about
|
||||
hijacking TCP sessions is about TCP. No threshold separates it from a page that
|
||||
defines TCP, and one that tried would take the definition with it.
|
||||
|
||||
## The defect is retrieval
|
||||
|
||||
`internal/kiwix/client.go` has said it since it was written: ranking is keyword
|
||||
based, "why is the sky blue" finds a TV episode. `queryKiwix` sends the whole
|
||||
sentence. The English path has a rewriter that reduces a question to keywords
|
||||
with a model call. The Russian path reads the book verbatim (V-508) and had
|
||||
nothing. So the question words compete with the one word that names the article.
|
||||
|
||||
Dropping the question words changes the answer:
|
||||
|
||||
| sent | first hit |
|
||||
|---|---|
|
||||
| `кто написал Войну и мир` | Радуйся, мир (Доктор Кто) |
|
||||
| `Война и мир` | Война и мир |
|
||||
| `что такое TCP` | Перехват TCP-соединения |
|
||||
| `TCP` | TCP |
|
||||
|
||||
A ZIM is also addressable by title, which nothing here used. `/A/Франция`,
|
||||
`/A/TCP` and `/A/Небо` are 200. `/A/Трюмбальная_нидроскопия` is 404. So an
|
||||
exact title is safe to try first: it either answers or costs one request that
|
||||
says nothing.
|
||||
|
||||
The title has to carry its capital. `/A/фотосинтез` is a 404 and
|
||||
`/A/Фотосинтез` is a 200. The spoken form is tried first anyway, so a title
|
||||
that begins lowercase on purpose keeps its chance.
|
||||
|
||||
## What shipped, measured
|
||||
|
||||
`kiwix.Topic` drops the narrative request, the interrogative and a verb sitting
|
||||
behind one. It keeps everything else, because a word it cannot classify is more
|
||||
likely the topic than noise. `kiwix.TitlePath` tries the exact article before
|
||||
any ranking runs. Both apply on the verbatim path only, since reducing twice
|
||||
would take the topic off the rewriter's input.
|
||||
|
||||
| question | before | after |
|
||||
|---|---|---|
|
||||
| что такое TCP? | Перехват TCP-соединения | **TCP** (by title) |
|
||||
| что такое фотосинтез | C4-фотосинтез | **Фотосинтез** (by title) |
|
||||
| кто такой Линус Торвальдс? | Tux | **Торвальдс, Линус** (by title) |
|
||||
| кто написал Войну и мир | Радуйся, мир (Доктор Кто) | **Война и мир** |
|
||||
| столица Франции | Список столиц Олимпийских игр | **Париж** (by title) |
|
||||
| что такое чёрная дыра | Чёрная дыра | Чёрная дыра (by title) |
|
||||
| почему небо голубое | Город золотой | Под небом голубым… (фильм) |
|
||||
| почему трава зелёная | Сено | Зелень |
|
||||
|
||||
Five questions reach the right article where they did not. One was already
|
||||
right and stays right. Nothing regressed.
|
||||
|
||||
"столица Франции" is the surprise. The 2026-08-05 measurement named it as the
|
||||
case a quality gate must not break, because the answer is Париж and that word
|
||||
is not in the question. The ZIM holds a title redirect, so asking for the
|
||||
article titled "Столица Франции" returns Париж. Retrieval by title reaches an
|
||||
answer that retrieval by keyword cannot.
|
||||
|
||||
## What is still wrong
|
||||
|
||||
Two of the eight are still not answered, and both are the same shape. The
|
||||
question names no article and no redirect covers it. "почему небо голубое" is
|
||||
answered by Rayleigh scattering, and nothing in the question says so. Keyword
|
||||
retrieval cannot bridge that and neither can a threshold. The candidates are a
|
||||
semantic index over titles, or asking the resident model for the article title
|
||||
rather than for keywords.
|
||||
|
||||
`Response.Empty()` is still the whole gate. A wrong article that the search
|
||||
does return is still spoken. What this change buys is that the article is
|
||||
usually right, not that a wrong one is caught.
|
||||
|
||||
## Not measured here
|
||||
|
||||
The English path, which still goes through the rewriter and was not touched.
|
||||
SearXNG, where the same question about answerhood is open and the 2026-08-05
|
||||
result stands. The cascade end to end, since the workstation is off and the
|
||||
phrasing arm is the resident model.
|
||||
@@ -0,0 +1,63 @@
|
||||
# silero-vad against the energy threshold in mavwaked
|
||||
|
||||
*Measured 2026-08-09 on homesrv. V-487, stage one of two.*
|
||||
|
||||
mavwaked decided an utterance had started by comparing frame energy to an
|
||||
adaptive floor. That answers "is this frame loud". A fan, a door and a
|
||||
television are all loud, and every utterance mavwaked accepts becomes a turn.
|
||||
|
||||
silero-vad answers "is this frame speech". It is 2.3MB of ONNX and it replaces
|
||||
the comparison and nothing else. The speech hold, the silence hold, the length
|
||||
cap and the utterance buffer are the same state machine either way.
|
||||
|
||||
## What it declines
|
||||
|
||||
Speech is the four piper fixtures `mavsttd` already scores against, so nothing
|
||||
of the owner's voice is committed. Non-speech is white noise at the same RMS as the clip beside it. That is the
|
||||
cheapest thing that fools an energy floor.
|
||||
|
||||
| clip | silero, speech frames on speech | silero on noise | energy on noise |
|
||||
|---|---|---|---|
|
||||
| ru_fact.wav | 59 | 0 | 68 |
|
||||
| ru_query.wav | 69 | 0 | 79 |
|
||||
| ru_reminder.wav | 80 | 0 | 89 |
|
||||
| en_act.wav | 90 | 0 | 99 |
|
||||
|
||||
The energy threshold accepts every noise clip as a complete utterance. Silero
|
||||
calls not one frame of any of them speech, and still hears all four spoken
|
||||
clips. `TestSileroHearsSpeechAndDeclinesNoise` is that table.
|
||||
|
||||
White noise is a floor, not a proof. It says nothing about a television, which
|
||||
is speech, or about a fan, which is narrowband. Those need room recordings and
|
||||
this box has none.
|
||||
|
||||
## What it costs
|
||||
|
||||
`BenchmarkSileroFrame` on the homesrv laptop (Ryzen 5 5600U), one 30ms frame
|
||||
through the model including the re-chunking:
|
||||
|
||||
509µs per frame
|
||||
|
||||
That is 1.7% of one core, on the slower of the two machines. The detector runs
|
||||
on the workstation beside the microphone, never on the GPU. This number is what
|
||||
says it does not need one.
|
||||
|
||||
## The window is 512 samples, not 480
|
||||
|
||||
`cmd/mavwaked/main.go` claimed the frame contract matched silero's input
|
||||
exactly. That was true of silero v4. Version 5 takes exactly 512 samples at 16kHz, plus 64 samples of context from
|
||||
the previous window. So `sileroVAD` buffers across capture frames, and a frame
|
||||
completing no window inherits the previous probability. `TestSileroRechunksAcrossFrames` pins it.
|
||||
|
||||
## Still an energy gate by default
|
||||
|
||||
`-vad-model` is empty in the code default, so a deployment that does not pass
|
||||
it runs exactly what shipped before. Barge-in is untouched and deliberately so. It reads frame energy while she is
|
||||
speaking, which is a different question from whether the frame is speech.
|
||||
|
||||
## Not done here
|
||||
|
||||
The wake word. This is stage one of the two V-487 asks for. The second needs a
|
||||
keyword model that does not exist yet. The pretrained openWakeWord keywords are
|
||||
English, and a Russian one has to be trained. Until then anything spoken near
|
||||
the microphone still becomes a turn. It is now merely required to be speech.
|
||||
@@ -0,0 +1,104 @@
|
||||
# The "Мэйвен" wake word: what it hears and what it invents
|
||||
|
||||
*Measured 2026-08-09 on workpc and homesrv. V-487, stage two of two.*
|
||||
|
||||
Stage one gave mavwaked silero-vad, which answers "is this frame speech".
|
||||
Nothing answered "was this said to her", so every utterance near the
|
||||
microphone became a turn. SurfaceVoice caps acts at L0, which made that safe
|
||||
rather than expensive. L0 does not cap reading, so the room could still hear
|
||||
his facts read back.
|
||||
|
||||
The keyword is "Мэйвен". openWakeWord's two frozen feature models do the
|
||||
hearing and a 100KB head trained here draws the boundary. It runs on CPU
|
||||
beside silero and never touches the GPU.
|
||||
|
||||
## Why a per-window accuracy is not a number anyone can act on
|
||||
|
||||
The gate scores every 80ms. A 1.7% false-accept rate per window sounds small
|
||||
and means a wake every few seconds. The useful question is how many times an hour
|
||||
it wakes on speech that was not the keyword. So every table below counts
|
||||
threshold crossings over whole clips and divides by the audio duration.
|
||||
|
||||
A crossing, not a window above the threshold. A keyword held high for half a
|
||||
second is one wake, not six.
|
||||
|
||||
## The data
|
||||
|
||||
Positives are 600 silero TTS renders of three stressings of the keyword, six
|
||||
speakers, ten trailing phrases, augmented eight ways each. Hard negatives are
|
||||
560 renders of confusable Russian words. Real speech is Common Voice ru and Golos.
|
||||
The 74257 Common Voice clips were already on workpc from the CrisperWhisper
|
||||
work. The 200 Golos clips came from the CW2 WER eval.
|
||||
|
||||
Splits are by source file. Augmented copies of one render on both sides of a
|
||||
split would measure memorisation.
|
||||
|
||||
Golos was never trained on at any stage, so it answers the harder question:
|
||||
does this survive a change of speakers and rooms.
|
||||
|
||||
## Three heads
|
||||
|
||||
Each row is a full retrain. The false-accept column is 8.89 hours of Common
|
||||
Voice that no stage of training had seen.
|
||||
|
||||
| trained on | recall (window) | false wakes/hour @0.99 |
|
||||
|---|---|---|
|
||||
| TTS + 13.7 min of Golos | 0.869 | not measurable |
|
||||
| + 4000 Common Voice clips | 0.836 | 21.9 |
|
||||
| + 3837 mined hard negatives | 0.784 | 4.2 |
|
||||
| + 753 more mined | 0.810 | 3.4 |
|
||||
|
||||
The first row is why the second exists. Thirteen minutes of held-out speech
|
||||
cannot measure a rate for a gate that scores twelve times a second. A head
|
||||
trained only against TTS learns to tell TTS from not-TTS.
|
||||
|
||||
Mining is the whole story after that. Random negatives teach the head what
|
||||
most speech sounds like. They do not teach it the few syllable sequences that
|
||||
score high, because 4000 clips barely contain them. So the current head was
|
||||
run over 20000 fresh clips, keeping every window it scored above 0.05. That
|
||||
found 3837 windows in 855512. Repeating those ten times in the next training
|
||||
run cut the rate five-fold.
|
||||
|
||||
The second round found 753 in 852240, a fifth of the yield, and bought a
|
||||
further 20%. It also recovered recall, which the first round had cost. Whether
|
||||
a third round is worth 25 minutes of workpc is untested.
|
||||
|
||||
## Where the threshold came from
|
||||
|
||||
Both columns are held out. Positives are the 126 renders in the test split.
|
||||
Speech is 65.1 minutes of Common Voice, disjoint from every training and
|
||||
mining pool. Both were run through the built `mavwaked` binary reading PCM from a
|
||||
file, not through the python that trained the head.
|
||||
|
||||
| threshold | renders shipped | false wakes/hour |
|
||||
|---|---|---|
|
||||
| 0.99 | 116 / 126 | 2.8 |
|
||||
| 0.999 | 115 / 126 | 0.9 |
|
||||
|
||||
One render against a third of the false wakes. `defaultWakeThreshold` is
|
||||
0.999.
|
||||
|
||||
Golos disagrees. It gave 2 wakes in 14 minutes at every threshold, which is
|
||||
8.7 per hour. Two events is not a rate. What it does say is that a handful of real utterances score above 0.999
|
||||
and no threshold will move them.
|
||||
|
||||
## What it costs him
|
||||
|
||||
Ten of the 126 held-out renders were heard and still dropped, and every one
|
||||
was an utterance shorter than 1.32s. The head scores 16 embeddings, or 1.28s of
|
||||
audio. The score therefore peaks up to a second after a short keyword ends.
|
||||
By then the VAD has closed the utterance and dispatch has already asked.
|
||||
|
||||
Real commands are "Мэйвен, <request>" and run past two seconds, which gives
|
||||
the head the whole request to peak during. A bare "Мэйвен" with nothing after
|
||||
it is the case that fails. One fix would hold an ignored utterance for a grace
|
||||
period and ship it if the keyword lands late. It is not built.
|
||||
|
||||
## What was not measured
|
||||
|
||||
No room recordings. Every negative above is a clean corpus clip. This gate
|
||||
will live among a television, a fan and the far side of a kitchen. None of
|
||||
those are in these numbers.
|
||||
|
||||
No measurement of him. Training on his voice means copying his transcripts off
|
||||
homesrv, which is his call and has not been asked.
|
||||
@@ -0,0 +1,74 @@
|
||||
# Language: what the model emits, and how Russian is matched
|
||||
|
||||
*Last verified: 2026-08-09 @ a9b480a*
|
||||
|
||||
Two contracts live here. What a model call is allowed to return, and which
|
||||
mechanism is allowed to recognise a Russian word.
|
||||
|
||||
## The LLM output contract
|
||||
|
||||
All phrasing paths emit `{"response":"...","mood":"..."}`. They fall back to
|
||||
plain text when the model skips the JSON.
|
||||
|
||||
**One parser, `parseResponseMood` in `internal/phraser/parse.go`.** Every path
|
||||
reaches it: the six `LLMPhraser` methods, `PhraseWorld`, and
|
||||
`Replier.PhraseReply`. `cmd/mavend/replier_llm.go` wraps the last of those,
|
||||
holds the stub fallback, and parses nothing itself. Mood is a fixed enum.
|
||||
|
||||
The router prompt is a separate contract:
|
||||
|
||||
```text
|
||||
[{"intent":<enum>, key?, value?, text?, verb?}, ...]
|
||||
```
|
||||
|
||||
over 7 intents: `fact, reminder, note, query, act, chat, system`.
|
||||
`llm/check_prompt_parity.py` in the training workspace enforces that the Go
|
||||
prompt and the relabelling prompt stay identical. They diverged once, and the
|
||||
relabelled set then taught a head the Go router never asks for.
|
||||
|
||||
## Russian patterns: three mechanisms, no fourth
|
||||
|
||||
Hand-written Russian stem patterns were swept out on 2026-08-04 by the owner's
|
||||
call. A regex whose output is a fact or a route is the defect. A regex over
|
||||
structured input, such as HTML, MIME, JSON, a URL or an argv list, is not.
|
||||
|
||||
Before writing a Russian word list, pick one of these.
|
||||
|
||||
### `internal/lexicon`, for closed classes
|
||||
|
||||
`lexicon_ru_v1.json` holds interrogatives, capture verbs, reminder verbs,
|
||||
cardinals, day offsets, parts of day, weekdays, months and spoken hours.
|
||||
Editing a word is a data change and there is exactly one copy. Months used to
|
||||
live in three files and drifted between them.
|
||||
|
||||
Cardinals carry the oblique forms, because a spoken time declines. `в семь` and
|
||||
`к семи` are one hour.
|
||||
|
||||
### `internal/morph`, for grammar
|
||||
|
||||
From the vendored golem Russian dictionary. `IsVerbForm` and `SameWord`.
|
||||
|
||||
Lemma matching is BROADER than stem-plus-one-ending. A verb slot that means the
|
||||
imperative must be matched exactly. `говори` and `говорил` share a
|
||||
lemma and only one of them is a command
|
||||
(`cmd/mavend/quiet_toggle.go`).
|
||||
|
||||
### `cmd/mavend/topics.go` and the embedder, for open sets
|
||||
|
||||
Use these when the question is what a turn is ABOUT. Frozen seeds per subject
|
||||
plus a real `other` class, scored against the turn's own query vector.
|
||||
|
||||
Same shape as the personal boundary in `personalboundary.go`, with one
|
||||
difference. A topic must clear the runner-up by `topicMargin`, because a false
|
||||
claim here spends a network scan rather than one honest "не знаю". The old
|
||||
keyword tests stay as the offline floor.
|
||||
|
||||
### The ecosystem trio
|
||||
|
||||
Use it when the answer is not in the utterance at all. Identity is Nexus's,
|
||||
never a local pattern.
|
||||
|
||||
## Seeds are scoring data
|
||||
|
||||
Editing one moves a recogniser. Re-measure against the `TestONNX*` tests rather
|
||||
than eyeballing the change.
|
||||
+38
-3
@@ -1,6 +1,6 @@
|
||||
# Offloading model work to the workstation
|
||||
|
||||
*Last verified: 2026-08-05 @ b789676. Living doc: correct it in place, do not append.*
|
||||
*Last verified: 2026-08-09 @ 50c6637. Living doc: correct it in place, do not append.*
|
||||
|
||||
Owner's call, 2026-08-02. Vikunja #483 is the umbrella. Tasks #484 to #487 are the
|
||||
work, and this file holds the shape and the rules all four must obey.
|
||||
@@ -105,6 +105,17 @@ how we find out whether the blind spot is real.
|
||||
untouched. The model, the context size, the layer count and the MTP flags are the
|
||||
owner's business and not this daemon's schema.
|
||||
|
||||
**Every GPU service on that box belongs under this supervisor**, added to
|
||||
`cmd/mavgpud` rather than to systemd beside it. The rule was learned on
|
||||
2026-08-09. The CW2 transcriber ran as its own user unit and registered on the
|
||||
KFD like any ROCm job. So the supervisor read its own transcriber as a
|
||||
contender. It yielded the card every few seconds and the gemma-4-12b arm was
|
||||
down for eight minutes before anyone looked. So the supervisor takes a `stt`
|
||||
block and starts CW2 itself. Yielding is all or nothing, because a job that
|
||||
wants the card wants all of it. Idle unloading is not. It applies to
|
||||
llama-server, which holds 8GB. CW2 holds 1.6GB, and unloading it would cost the
|
||||
next voice turn its quality for nothing.
|
||||
|
||||
## What stays on homesrv, permanently
|
||||
|
||||
The **embedder** (multilingual-e5-small, ONNX, CPU). It backs the classifier, which
|
||||
@@ -149,6 +160,29 @@ flips. It is wired anyway: `PhraseReminder` is on the same transport and is on.
|
||||
Then the embedder above, **whisper.cpp** in `mavsttd`, and **piper** in `mavttsd`.
|
||||
`mavwaked` uses no model at all: an energy-threshold VAD over 30ms frames.
|
||||
|
||||
Speech-to-text is wired as of 09-08-2026, and it takes only the silent half of the
|
||||
rule. A worse transcript is still a turn, so there is nothing to name a gap about
|
||||
and `stt.Pair` has no `TranscribeRemote`. `sttSeam` in `cmd/mavend/voicewire.go`
|
||||
builds it, beside `modelSeam` and at the same place in `wireVoice`, so the voice
|
||||
path and the meeting recorder still share one transcriber.
|
||||
|
||||
The remote is not a second endpoint on mavgpud. whisper.cpp cannot load
|
||||
CrisperWhisper 2.0 at all. It reads its language count off the vocabulary
|
||||
size, and CW2's 51897 tokens shift seven special token ids. So CW2 runs under
|
||||
transformers as its own service on port 8081, and `stt.HTTPTranscriber` is the
|
||||
second transport for the same seam. It posts raw PCM with the format in headers.
|
||||
It carries a bearer token, because audio is the most sensitive thing that
|
||||
crosses here.
|
||||
|
||||
It is a second endpoint on nothing, but it is a second **child** of mavgpud, and
|
||||
that part is not optional. See the supervisor section above for why: a ROCm
|
||||
service the supervisor does not own is a contender it yields to.
|
||||
|
||||
The margin is the reason: CW2 turbo scores 10.4% WER in Russian against 27.5% for
|
||||
the `ggml-small.bin` mavsttd loads, over 200 Golos clips
|
||||
(`docs/evals/2026-08-09-crisperwhisper2-russian-wer.md`). Text-to-speech has not
|
||||
moved and piper on homesrv is still the only synthesizer.
|
||||
|
||||
Speech-to-text stays two stages when it moves. One call carrying both a clip and the router
|
||||
prompt was measured on 05-08-2026. It scores 54.2% intent-only against 84.7% for whisper on
|
||||
homesrv, on the same 72 cases. The model transcribes clips it then routes wrong, so a long
|
||||
@@ -173,8 +207,9 @@ cleaner transcripts, not accuracy. See `docs/evals/2026-08-05-audio-in-routing.m
|
||||
which fixes what the 1.7B gets wrong: world knowledge, and the persona the CPT
|
||||
targets. The degradation path is already written and measured, since the
|
||||
classifier scores 68.8% full accuracy at p50 16.6µs on its own.
|
||||
3. **Speech-to-text and text-to-speech** (#486). They gain a real margin, but on
|
||||
quality alone, and both already work.
|
||||
3. **Speech-to-text and text-to-speech** (#486). Speech-to-text is wired, see
|
||||
above. Text-to-speech is not, and piper is good enough that nothing argues
|
||||
for moving it yet.
|
||||
4. **The wake word** (#487). Independent of all of the above.
|
||||
|
||||
## Assumptions
|
||||
|
||||
+12
-2
@@ -1,6 +1,6 @@
|
||||
# Start Commands
|
||||
|
||||
*Last verified: 2026-08-02 @ 7079a24. Living doc: correct it in place, do not append.*
|
||||
*Last verified: 2026-08-07 @ a4630b9. Living doc: correct it in place, do not append.*
|
||||
|
||||
All commands assume `ROOT=/home/kami/apps/Maven` and the local Go toolchain at `$ROOT/deps/go/go/bin/go`.
|
||||
|
||||
@@ -44,7 +44,8 @@ Config path: `~/.config/maven/mavend.json`. Full example with all options.
|
||||
"repeat_interval": "5m",
|
||||
"ntfy": {
|
||||
"base_url": "https://ntfy.kvmx.ru",
|
||||
"topic": "maven"
|
||||
"topic": "maven",
|
||||
"token": "${NTFY_TOKEN}"
|
||||
},
|
||||
"phraser": {
|
||||
"model_path": "/mnt/hdd1/llms/Qwen3-Maven-1.7B-Q8_0.gguf",
|
||||
@@ -66,6 +67,15 @@ Config path: `~/.config/maven/mavend.json`. Full example with all options.
|
||||
|
||||
Omit the `embedder` block entirely to use the deterministic HashEmbedder floor (no ML, no ONNX runtime dependency). Useful for testing or low-resource setups.
|
||||
|
||||
`${NTFY_TOKEN}` and the `${TELEGRAM_*}` vars are expanded from `deploy/telegram.env`, which is gitignored. Copy `deploy/telegram.env.example` and fill it in. Mint a scoped token rather than reusing an admin one. It needs write access to the `maven` topic and nothing else:
|
||||
|
||||
```sh
|
||||
ntfy access maven maven write-only
|
||||
ntfy token add --expires=never maven
|
||||
```
|
||||
|
||||
Deleting the `ntfy` block turns the reach off, and that is not a no-op. The routing table sends sev3-away nudges and away reminders to ntfy and nowhere else. With no sink wired they hit a nil and vanish, leaving no log line and no `delivery_attempts` row (V-649).
|
||||
|
||||
## mavsttd — STT worker (optional, remote whisper.cpp)
|
||||
|
||||
Requires `LD_LIBRARY_PATH` to include deps/lib (for libwhisper.so, libggml-vulkan.so).
|
||||
|
||||
+462
@@ -0,0 +1,462 @@
|
||||
# Routing
|
||||
|
||||
*Last verified: 2026-08-09 @ 31b5093*
|
||||
|
||||
How an utterance becomes a `Decision`, why each stage exists, and what every
|
||||
stage has measured. `CLAUDE.md` carries the rules an agent must not break. This
|
||||
file carries the reasoning and the history behind them.
|
||||
|
||||
Two decisions come out of a route. **Intent** is one of seven values. **Source**
|
||||
is where the answer lives, and it is read on `IntentQuery` alone. They are scored
|
||||
separately, because one number hides which one moved.
|
||||
|
||||
## The cascade
|
||||
|
||||
| Stage | What it is | Where |
|
||||
|---|---|---|
|
||||
| 0 | Deterministic grammars over the utterance | `stage0.go`, `praxis.go`, `worldquery.go` |
|
||||
| 0b | Four ONNX heads on one e5-small forward pass | `heads.go` |
|
||||
| 1 | The resident model, GBNF-constrained JSON | `llmrouter.go` |
|
||||
| 2 | Nearest neighbour over frozen seed phrases | `classifier.go`, `embedder.go` |
|
||||
|
||||
Every stage may decline, and the next one answers. Any error at stage 0b or 1
|
||||
falls through, so a turn never breaks on a model.
|
||||
|
||||
The resident model arm is wired at `voice.go:214` through
|
||||
`pickLLMRouter(cfg.Voice.UseLLMRouter(), llmClient)`. The flag is
|
||||
`voice.llm_router` in `config.go`, `DefaultLLMRouter` is on, and
|
||||
`deploy/mavend.json` sets it `true`. With no llama-server to talk to,
|
||||
`pickLLMRouter` logs that and degrades to the classifier.
|
||||
|
||||
The classifier is the floor and not dead code. It runs when the resident model
|
||||
is off, when there is no llama-server to talk to, and on any per-turn error.
|
||||
Routing by seed similarity is the known cause of weak Russian queries. Deleting
|
||||
it would make a model outage a broken turn.
|
||||
|
||||
### Why there are two engines at all
|
||||
|
||||
The original design was the classifier alone. `docs/rearchitecture.md` replaced
|
||||
it with a model that emits structured JSON. The same weights phrase the reply.
|
||||
That demoted the embedder from a routing gate to a hint for recall. The model
|
||||
became the default on 2026-07-31.
|
||||
|
||||
The gap it buys is smaller than the design assumed. Measured 2026-08-02 on the
|
||||
77-case Russian fixture, the classifier scores 68.8% full accuracy at p50 16.6µs.
|
||||
Qwen3-1.7B scores 72.7% through the cascade. Four points, not a doubling.
|
||||
|
||||
An older figure of 36.8% for the classifier stood in `CLAUDE.md` until then. It
|
||||
predates the stage 0 rules and the seed additions. Both now score inside the
|
||||
classifier baseline.
|
||||
|
||||
Latency was misreported the same way. A figure of 2.7 seconds stood for two
|
||||
days and was contention rather than the model.
|
||||
`docs/evals/2026-07-31-routing.md` line 61 measures the router at p50 825ms and
|
||||
the cascade at p50 0.80s to 1.04s.
|
||||
|
||||
## Numbers
|
||||
|
||||
Three arms answer, so three numbers are live. Judge a routing change against the
|
||||
classifier and the resident model, since those are what always answer.
|
||||
|
||||
| Arm | Intent | Destination | p50 | Measured |
|
||||
|---|---|---|---|---|
|
||||
| classifier + ONNX | 76.0% (73/96) | 36.4% (12/33) | 16.6µs | 2026-08-08 |
|
||||
| resident Qwen3-1.7B, cascade | 80.2% | not measured | 1.19s | 2026-08-05 |
|
||||
| routing heads, cascade | 96.9% | 75.8% | 27.9ms | 2026-08-08 |
|
||||
| gemma-4-12b, cascade | 84.4% | 72.7% | 329ms | 2026-08-02 |
|
||||
| gemma-4-E4B, cascade | 89.6% | 57.6% (19/33) | 294ms | 2026-08-09 |
|
||||
|
||||
The fixture grew from 77 cases to 91 to 96. So a number is comparable only to
|
||||
another number on the same fixture. Sources:
|
||||
`docs/evals/2026-08-05-routing-resident-model.md`,
|
||||
`docs/evals/2026-08-02-workstation-gemma4-12b.md`,
|
||||
`docs/evals/2026-08-09-e4b-vs-12b-routing.md`,
|
||||
`docs/evals/2026-08-08-routing-heads-in-go.md`,
|
||||
`docs/evals/2026-08-08-destination-fixture.md`.
|
||||
|
||||
The resident model alone scores 37.4% full against 61.5% intent-only. The gap is
|
||||
slots and not routing. It routes `reminder` and leaves the time to the daemon,
|
||||
which is what the contract asks.
|
||||
|
||||
To re-run the resident model as router, start a **second** llama-server on a
|
||||
fixed host port. The resident one binds `--port 0` inside the container and no
|
||||
host process can reach it.
|
||||
|
||||
### The workstation is not the better router any more
|
||||
|
||||
It was, from 2026-08-02 until the heads landed. gemma-4-12b beat everything on
|
||||
the box at 84.4% intent and 72.7% destination. The heads beat it on both at a
|
||||
twelfth of the latency. The workstation stays the better phraser.
|
||||
|
||||
E4B replaced the 12B on 2026-08-09 by the owner's call. It is a step down on
|
||||
routing. Against a same-session 12B control it costs four destination cases and
|
||||
buys 50ms. Read destination as the finding. It names nothing where the 12B names
|
||||
`recall` or `calendar`, which is safe but walks the whole chain. It has no MTP
|
||||
and cannot be given any here. The only `gemma4-assistant` draft on disk is
|
||||
trained against the 12B's hidden states.
|
||||
|
||||
## Stage 0: what the grammars claim, and why
|
||||
|
||||
A rule at this stage is a claim. Either the model gets this wrong, or it wastes a
|
||||
second getting it right. Every rule was added against a measurement.
|
||||
|
||||
- **Agenda questions** (`AgendaQueryGrammars`, 2026-08-01). "что у меня сегодня",
|
||||
"во сколько у меня встреча" and anything naming a calendar go to `IntentQuery`.
|
||||
They were going to `IntentSystem`, where `replySystem` has no agenda arm and
|
||||
answered "пока не умею". Worth 2.6 points of full accuracy and calendar 0/2 to
|
||||
2/2.
|
||||
- **Rest of day and narrative** (V-498, 2026-08-04). `rest-of-day-query` claims
|
||||
"что дальше?". `NarrativeQueryGrammar` claims "расскажи про X", "объясни X" and
|
||||
"опиши X". Neither carries a question mark or an interrogative, so the model
|
||||
called both `IntentFact`. `IsQuestionShaped` caught the write downstream, so
|
||||
this was a latency and fixture defect rather than a correctness one. The
|
||||
narrative rule declines `chatNarrativeTopics`, because the query chain has no
|
||||
source that answers a joke or a bedtime story.
|
||||
- **Praxis** (V-516, 2026-08-05). `PraxisGrammars()` fills `Slots.Fn` with a
|
||||
capability name. These grammars are the **only** path to Praxis and not a
|
||||
faster one. The model reaches Praxis 0/12 alone, the same as the classifier.
|
||||
Nothing in the router prompt names a Praxis capability, so there is no string
|
||||
for it to write. Through the cascade it is 11/12. Measured overall 16/30 to
|
||||
27/30, lifecycle 0/5 to 5/5
|
||||
(`docs/evals/2026-08-05-praxis-reach.md`,
|
||||
`docs/evals/2026-08-05-reach-llm-router.md`).
|
||||
`handlePraxisAct` compares `Slots.Fn` to a capability alias. Otherwise that
|
||||
slot is filled from the deployment's enabled tool names, and no Praxis alias
|
||||
is on that list.
|
||||
- **World questions** (`WorldQueryGrammars`, V-655, 2026-08-07). "что такое X"
|
||||
and "сколько будет 17 на 23". Wired after the agenda rules and **before** the
|
||||
feed and list rules. "что такое лента" is a definition question, and the feed
|
||||
rule would take it on the noun alone.
|
||||
|
||||
`calendar-query` and `event-time-query` name the calendar as the destination.
|
||||
The possessive agenda rules deliberately do not. "что у меня в списке покупок"
|
||||
matches `agenda-query`, and naming the calendar there would take the list source
|
||||
off the turn. That caution now costs four destination cases. See the model arm
|
||||
below.
|
||||
|
||||
Go's `\b` is ASCII-only and never fires after a Cyrillic letter. A pattern needs
|
||||
an explicit `(\s|[?!.]|$)`.
|
||||
|
||||
`baselineGrammars` in `eval_test.go` mirrors `buildRouter` and has drifted before.
|
||||
`WorldQueryGrammars` was wired into the daemon by V-655 and not into the mirror,
|
||||
so the fixture scored a grammar set nobody runs. Fixed by V-659, worth 3 points
|
||||
of destination.
|
||||
|
||||
### Praxis lifecycle rules
|
||||
|
||||
A **stative** lifecycle word ("готово", "принято") needs an item named beside it.
|
||||
A bare **imperative** ("закрывай") may ask which one. It also requires a sentence
|
||||
naming no object of its own. Otherwise "закрой шторы в комнате" goes to Praxis
|
||||
instead of the house. A demonstrative ("отметь это как сделанное") resolves
|
||||
against `h.surfacedItems` only when exactly one item was spoken. Otherwise the
|
||||
turn goes back to the cascade rather than transitioning the wrong item.
|
||||
|
||||
### Slots on a stage 0 decision
|
||||
|
||||
`fillMatchedSlots` runs the stage 2 extractor over whatever a grammar built
|
||||
(V-572, 2026-08-06). It fills only the slots the grammar left empty. A matched
|
||||
value always wins, because the rule read a literal pattern and the extractor
|
||||
guesses.
|
||||
|
||||
It did not run before. So `ReminderGrammar` handed the daemon `HasTime: false`
|
||||
for "напомни в 11:00 позвонить маме", and `missingFor` read the silence as
|
||||
absence and asked "Когда?". It is inert for every grammar but the reminder:
|
||||
`Extract` fills Time, Fn and Key and nothing else. A stage 0 query costs 3.7µs
|
||||
against 3.9µs before, benchmarked at 20000x.
|
||||
|
||||
`Slots.Text` is deliberately not filled. A grammar that left it empty meant it,
|
||||
and `agendaQueryBuild` hands the query chain the utterance itself.
|
||||
|
||||
## Stage 0b: the routing heads
|
||||
|
||||
Routing has a bounded output space, so it is classification rather than
|
||||
generation (owner's call, V-546,
|
||||
`docs/plans/18-routing-heads-on-e5-small.md`). The 118M multilingual-e5-small is
|
||||
already resident. A softmax cannot emit a value that does not exist, so no
|
||||
grammar is needed. Max softmax is a calibratable confidence, where
|
||||
`Confidence: 1.0` was a hardcode. Training costs roughly 5e15 FLOPs, so 10 to 30
|
||||
minutes on the workstation. A 100M decoder from scratch is 10 to 20 GPU hours.
|
||||
|
||||
**Fine-tune a copy of the weights.** The resident embedder backs memory recall.
|
||||
Training it in place couples routing accuracy to recall@1, with nothing in the
|
||||
suite to name the trade.
|
||||
|
||||
Four heads share one masked mean pool, trained over three days. The measurements
|
||||
are `docs/evals/2026-08-08-routing-heads-two-head.md`,
|
||||
`docs/evals/2026-08-08-slot-head-three-head.md`,
|
||||
`docs/evals/2026-08-08-clarify-head-four-head.md` and
|
||||
`docs/evals/2026-08-08-massive-warm-start.md`.
|
||||
|
||||
| Head | Score | Notes |
|
||||
|---|---|---|
|
||||
| intent | 92.8% mean over 3 seeds | fixture is the 88 cases carrying an intent |
|
||||
| destination | 80.8% mean, best 29/33 | beats the 12B teacher it was distilled from |
|
||||
| slot BIO tags | 72.4% span F1 | still climbing when epoch selection stops it |
|
||||
| clarify | catches 7.0 of 8, 2.3 false of 88 | parity with the cascade, no rules in front |
|
||||
|
||||
Read the best destination run as one seed and not a headline. One case is 3
|
||||
points on a fixture this small. Head intent accuracy is **not** comparable to the
|
||||
cascade's 76.0% and 84.4%. A softmax has no clarify class, so the head's fixture
|
||||
is 88 cases and not 96.
|
||||
|
||||
Recall is 15/15 and world is 5/5.
|
||||
|
||||
**Mood is cut, not deferred.** The enum describes her own reply state, not the
|
||||
speaker's emotion, and no dataset maps onto it.
|
||||
|
||||
### The clarify head
|
||||
|
||||
Clarify is not a value of intent, so a softmax cannot emit it. It is a second
|
||||
question over the same pooled vector: can Maven act on this at all. Accuracy is
|
||||
the wrong number here and a head that never asks scores 91.7%.
|
||||
|
||||
Confidence is the other half. Max softmax over the intent head reads 0.851 where
|
||||
it is right and 0.604 where it is wrong. It ranks right above wrong in 83.4% of
|
||||
pairs.
|
||||
|
||||
It is not free the way the slot head was. Intent, destination and slot F1 each
|
||||
move down one to four points, inside the seed spread. `поужинал` is a false
|
||||
clarify on every seed. That is the same defect `thinSingleToken` was narrowed for
|
||||
on 2026-08-01.
|
||||
|
||||
The corpus is generated, because every existing row is answerable by
|
||||
construction. The router-prompt agreement filter cannot work here. `routeGrammar`
|
||||
has no clarify value, and a generated line always agrees with itself. A gemma
|
||||
judge replaces it. The first judge called 24 of 40 answerable rows underspecified.
|
||||
It judged against a generic assistant rather than against Maven's contract.
|
||||
|
||||
### The slot head
|
||||
|
||||
BIO slot tags had no Maven-domain corpus. That was true of found corpora and
|
||||
false of made ones. `label_slots.py` distils spans out of gemma-4-12b under a
|
||||
GBNF closed over Maven's own five slots. A span survives only when it is a
|
||||
literal substring of the utterance, so the agreement filter costs no second call.
|
||||
2178 spans over 1702 rows, 37 dropped, nothing unparsed.
|
||||
|
||||
Epoch selection reads the intent dev slice alone. That costs the slot head about
|
||||
4 points.
|
||||
|
||||
### Warm start and the floor
|
||||
|
||||
The MASSIVE warm-start of step 2 is worth nothing here. Stock e5-small ties it on
|
||||
intent and leads by a third of a case on destination. Nothing argues for keeping
|
||||
that step.
|
||||
|
||||
The floor was a corpus defect and it is fixed. The first 120 floor rows carried
|
||||
one sentence shape, so the head named a destination where the fixture says walk
|
||||
the chain. Rotating six shapes took the floor 3/7 to 6/7 and destination 75.8% to
|
||||
80.8%.
|
||||
|
||||
What is left is calendar at 3/6 on every seed, which training cannot move. The
|
||||
possessive agenda rules claim those cases at stage 0 and name nothing, so no
|
||||
label reaches the head.
|
||||
|
||||
### Reading them in Go
|
||||
|
||||
`RouterHeads` in `internal/router/heads.go` loads `router_heads.onnx` (V-664,
|
||||
2026-08-08). It reads intent, destination and clarify off one forward pass.
|
||||
|
||||
Three rules around it, each measured:
|
||||
|
||||
- The **clarify head decides first**, before the intent threshold. It answers a
|
||||
different question. A thin utterance scores low intent by construction, so
|
||||
gating it cost 6 of 8 ambiguous cases.
|
||||
- The **destination head is read on `IntentQuery` only**, since no other intent
|
||||
reaches `queryWalk`.
|
||||
- `headsThreshold` is 0.6, the measured knee. Every value up to 0.85 drops right
|
||||
answers and keeps the same two wrong ones.
|
||||
|
||||
`voice.embedder.heads_path` is the whole switch. Empty, missing or unloadable
|
||||
means the heads are nil. The cascade is then byte-for-byte what shipped before
|
||||
them.
|
||||
|
||||
### The tokenizer bug the heads found
|
||||
|
||||
`encodeWord` in `onnxembedder.go` read every long word backwards until 2026-08-08.
|
||||
It cost recall@1 7.4 points and recall@3 11.1. Nothing caught it, because seeds
|
||||
and queries were mangled the same way and cosine survived. The heads found it.
|
||||
They are trained through transformers and read through this.
|
||||
|
||||
The embedder id now carries a tokenizer revision (`@384/tok2`). So fixing the
|
||||
tokenizer triggers `ReembedAll` the way swapping the model file does. Bump
|
||||
`tokenizerRev` on any change to what it emits.
|
||||
|
||||
## Clarify
|
||||
|
||||
`Confidence: 1.0` was hardcoded in `llmrouter.go`. So the model path could never
|
||||
ask for clarification, and it missed 6 of 6 refusal cases (V-359). The bug had a
|
||||
second half. The model branch never consulted `r.threshold` at all, so a correct
|
||||
low confidence would have been discarded anyway.
|
||||
|
||||
Fixed 2026-07-31 with structural signal feeding the same stage 3 gate the
|
||||
classifier path already had (`gateLLMDecision` in `router.go`). Three signals: a
|
||||
single-token utterance, a keyless fact, an act with no allowlisted fn.
|
||||
|
||||
Re-measured: missed clarify 6/6 to 1, at the cost of 3 false clarifies and 2.6
|
||||
points of full accuracy. Two of the three false clarifies are acts the model
|
||||
mis-routed and the gate caught. Asking beats wrongly executing, so the fixture
|
||||
and the daemon disagree about what is correct there.
|
||||
|
||||
The third, `поужинал`, was a real defect. The single-token rule was an English
|
||||
intuition. It does not transfer to Russian, where one word is routinely a whole
|
||||
sentence.
|
||||
|
||||
Narrowed 2026-08-01. `thinSingleToken` (`internal/router/singletoken.go`) still
|
||||
thins a bare one-word nominal. It spares two classes. One is a closed lexicon of
|
||||
social and control singles ("привет", "стоп", "yes"). The other is any token
|
||||
carrying a Russian verb ending, because a verb already contains its subject. Both
|
||||
tests are offline and cost nothing. False clarifies 3 to 2, intent-only 74.0% to
|
||||
75.3%.
|
||||
|
||||
The two remaining false clarifies are the act-with-no-allowlisted-fn arm of the
|
||||
gate, not this rule.
|
||||
|
||||
## The destination
|
||||
|
||||
`query` was a shrug. The cascade sorted an utterance into one of seven intents,
|
||||
then `IntentQuery` handed the turn to `querySources` in the daemon. That is
|
||||
twenty-two branches deciding by seed similarity in a fixed order. It had no
|
||||
fixture, no accuracy number, no model arm and no floor.
|
||||
|
||||
`Decision.Source` (`internal/router/source.go`) is the second half of the route
|
||||
(V-655, 2026-08-07). Twelve destinations, not twenty-two. The three recall passes
|
||||
plus `fact-by-key` are one destination from outside. So are search, Kiwix and the
|
||||
URL reader.
|
||||
|
||||
`queryWalk` in `cmd/mavend/actions_query.go` takes sources **out** and moves none.
|
||||
That is the safety argument. The table's order is load-bearing. Every comment on
|
||||
it argues a reason between two sources. Above all it carries "the owner's data
|
||||
first, then the world". Naming `SourceWorld` does not send the turn outside on its
|
||||
own.
|
||||
|
||||
What comes out is only the sources that **guess**. Those decide a turn is theirs
|
||||
by cosine against frozen seeds, then answer whatever they claimed. They hold no
|
||||
table that could come back empty. Weather is the pure case and has no local data
|
||||
at all. It was measured on the box 2026-08-07
|
||||
(`docs/evals/2026-08-07-week-of-usage.md` section 4). It answered both "что такое
|
||||
TCP?" and "сколько будет 17 на 23?" with "для какого города?". The feed answered
|
||||
"какой у меня любимый язык?" with kernel headlines.
|
||||
|
||||
### Who may drop the personal boundary
|
||||
|
||||
The personal boundary guesses, so naming `SourceWorld` drops it. That is what
|
||||
stops it answering "кто такой Линус Торвальдс?" with "не нашла у тебя такой
|
||||
записи", which it did on 2026-08-07.
|
||||
|
||||
Three deciders name a destination and two of them infer it: the heads and the
|
||||
resident model. An inferred `SourceWorld` on a question about him would reach
|
||||
SearXNG. That widens what is asked rather than costing a local answer. So only a
|
||||
stage 0 grammar may drop it (owner's call, V-666, 2026-08-09).
|
||||
|
||||
`Decision.SourceAnchored` carries the provenance. It is a field and not
|
||||
`Stage == 0`. Stage 0 also means confidence 1.0 and an anchored claim band, and
|
||||
one of those could stop implying the others. `definitionQueryPattern` claims "кто
|
||||
такой X", so the 2026-08-07 case is still anchored and still answered.
|
||||
|
||||
`queryWalk` reads `SourceAnchored` for the query source marked `boundary: true`
|
||||
and no other. Every other guesser still comes off the turn, whoever named the
|
||||
destination. `TestOnlyAGrammarMayDropTheBoundary` and
|
||||
`TestNamingRecallKeepsTheBoundary` pin both directions.
|
||||
|
||||
### The destination fixture
|
||||
|
||||
`want_source` on `eval.Case` is a pointer, because the destination has three
|
||||
states and a bare string has two. Absent is every intent but query. Present and
|
||||
empty is the `SourceUnknown` contract: name nothing and walk the chain. Present
|
||||
and named is a destination the route must produce. Thirty-three of ninety-six
|
||||
cases carry one.
|
||||
|
||||
A destination miss does **not** fail the case. It lands in `Outcome.SourceReason`
|
||||
and never in `Reasons`, so `Accuracy` and `IntentAccuracy` mean what they meant.
|
||||
`SourceAccuracy` is a second number over the labelled cases alone.
|
||||
A route that lost its intent scores no destination hit. Otherwise a clarify would
|
||||
satisfy an empty label for free.
|
||||
|
||||
Seven cases assert the floor and five of them are homelab operations. They cluster
|
||||
because `SourceRecall`, `SourceNetwork` and `SourceAttention` overlap on every
|
||||
question about the box. `mavpoll` writes its netdata and uptime-kuma observations
|
||||
into the fact store recall reads. That is a finding about the enum, not a gap in
|
||||
the labelling. The other two are `ru-query-005` and `ru-query-014`. No query
|
||||
source reads the reminder store, and a deadline could sit in tasks, the calendar
|
||||
or Praxis. The owner confirmed all seven floor labels on 2026-08-08.
|
||||
|
||||
### The model arm
|
||||
|
||||
`routeGrammar` carries a `source` rule closed over `router.Sources` plus the
|
||||
empty floor (V-660, 2026-08-08). So the model cannot emit a destination that does
|
||||
not exist. The prompt lists the twelve in Russian and says `""` is a normal answer
|
||||
to give often. `LLMRouter.Route` reads it back through `ValidSource` and on
|
||||
`IntentQuery` alone.
|
||||
|
||||
Against gemma-4-12b the cascade scores destination 24/33 with intent unmoved, and
|
||||
recall goes 0/15 to 14/15.
|
||||
|
||||
**Stage 0 now costs four destination points.** It did not before. The four cases
|
||||
the cascade loses and the model alone wins are all calendar. The possessive agenda
|
||||
rules claim them first and name nothing on purpose. That caution was free while
|
||||
nothing downstream could name anything either. It is not free now, and the fix is
|
||||
the owner's call (V-660 open).
|
||||
|
||||
## The decision trace
|
||||
|
||||
Arbitration between the claimants on the utterance stream is order. It is
|
||||
hardcoded in the pre-route resolver ladder, in `buildRouter` and in
|
||||
`querySources`. Nothing recorded who lost until V-564.
|
||||
|
||||
`internal/decision` records one `Record` per turn. It holds every claimant, what
|
||||
it would have made the turn, the score it reported, and how it ended. A claimant
|
||||
won, declined, lost on score, was thinned by a gate or was **never asked**.
|
||||
|
||||
The record rides the context, the same seam `querysource.go` uses. So a claim
|
||||
site cannot change a route, and a context with no record costs nothing. It is
|
||||
installed in `runTurn`, so the mic, telegram and the web leave the same trail.
|
||||
|
||||
Adding a rung to the ladder in `runTurn` means adding its name to `preRouteLadder`
|
||||
in `cmd/mavend/decisiontrace.go`. Otherwise that rung is silently missing from the
|
||||
record.
|
||||
|
||||
### Why it persists now
|
||||
|
||||
The original rule was that nothing persists, because a turn record is read minutes
|
||||
later or never. Storage was a 25-turn in-memory ring read over `ipc.TurnDecisions`
|
||||
and rendered on `/trace`.
|
||||
|
||||
The owner reversed it on 2026-08-06 (V-629,
|
||||
`docs/plans/21-persisting-the-routing-trace.md`). The routing heads cannot be
|
||||
fitted or calibrated without real utterances. And 9 of the 31 modes in
|
||||
`internal/modes` have no seed example at all.
|
||||
|
||||
The ring did not move. `cmd/mavend/routingtrace.go` is a second sink beside it,
|
||||
writing `routing_traces` (migration #23). The utterance is stored in clear. A
|
||||
384-dimension vector of a short sentence is substantially recoverable, so storing
|
||||
vectors instead would be a privacy claim we cannot support. What makes it safe is
|
||||
the same thing that makes the fact store safe. Retention is 14 days, enforced on
|
||||
write and again on start, so a box that goes quiet does not keep every row.
|
||||
Nothing reads it outward. `Store.Wipe` deletes it with everything else.
|
||||
|
||||
### Corrections
|
||||
|
||||
A correction is promoted out into a seed-shaped row in `routing_labels`
|
||||
(migration #24) and kept, because a label is not a transcript. The transcript
|
||||
still expires.
|
||||
|
||||
A turn marked wrong with no target is a usable negative, so naming the intent is
|
||||
never required. The target is one of the seven intents and never free text.
|
||||
|
||||
All three reaches offer it as of 2026-08-06:
|
||||
|
||||
- `/chat` offers two buttons beside the reply, over `ipc.CorrectTurn` and the
|
||||
trace id that rides back on `ipc.ChatReply`.
|
||||
- Voice offers the `repair` rung, which has read spoken corrections since V-455.
|
||||
It now writes the durable label beside the classifier seed it always wrote. A
|
||||
spoken negative with no target is its own rung, `repair-negative` (V-636,
|
||||
`docs/plans/22-correcting-a-turn.md`).
|
||||
- Telegram offers an inline keyboard under the reply. It needed the chat to become
|
||||
readable first (V-637, `docs/plans/23-inbound-telegram.md`). The poller is dark
|
||||
unless the `telegram` block says `intake`. It long-polls, because the box takes
|
||||
no inbound connections. It accepts `chat_id` and no other sender, and it drops
|
||||
whatever queued while the daemon was down. It reaches the daemon through
|
||||
`ipc.CoreAPI` alone.
|
||||
|
||||
The turn source is still `tap:text` for both telegram and the web. So provenance
|
||||
cannot tell a chat turn from a typed one.
|
||||
@@ -0,0 +1,80 @@
|
||||
# Session workflow: the five stores and the guards
|
||||
|
||||
*Last verified: 2026-08-09 @ a9b480a*
|
||||
|
||||
How a session starts, where each kind of writing belongs, and what the hooks
|
||||
refuse. `CLAUDE.md` carries the commands. This file carries the reasoning.
|
||||
|
||||
## Five stores
|
||||
|
||||
Each owns something the others must not hold.
|
||||
|
||||
| Store | Holds | Lifetime |
|
||||
|---|---|---|
|
||||
| Vikunja task | goal, constraints, assumption ledger, status | durable |
|
||||
| `CLAUDE.md`, `AGENTS.md` | what an agent must know before touching code | durable |
|
||||
| `docs/` | design, measurements, decisions | durable |
|
||||
| `TASK.md` | the brief for this branch, written by `task start`, immutable | one branch |
|
||||
| `HANDOFF.md` | only what the next agent needs to resume | one session |
|
||||
|
||||
`TASK.md` and `.task/` are excluded through `.git/info/exclude`. `HANDOFF.md` is
|
||||
gitignored and injected at session start. If a line in the handoff would still
|
||||
matter next week, it is in the wrong file.
|
||||
|
||||
## Doc tiers
|
||||
|
||||
Tiered by path, so staleness is visible from the filename.
|
||||
|
||||
- Files directly under `docs/` are living. They carry a
|
||||
`Last verified: <date> @ <sha>` line and are corrected in place.
|
||||
- Files under `docs/evals/` are dated measurements and are never edited after
|
||||
the day. A newer number is a new file, not an edit.
|
||||
- Files under `docs/archive/` are dead and read by nobody by default.
|
||||
|
||||
## Vikunja
|
||||
|
||||
This repo is project **Maven** (ID 2). MCP at `http://localhost:9100/mcp`, or
|
||||
`http://192.168.1.104:9100/mcp` from workpc. Feature, bug and deploy tasks go
|
||||
there.
|
||||
|
||||
A task holds the goal, the constraints and the assumption ledger. A session
|
||||
without a task id cannot be resumed by anyone, so a session with none asks for
|
||||
one first.
|
||||
|
||||
**Close a finished task with `done: true` and nothing else** (owner's call,
|
||||
2026-08-07). Do not write a completion summary into the description on the way
|
||||
out. It is lost anyway, and the durable record is the commit messages and the
|
||||
merged PR. `update_task` carrying a `description` resets `done` to false, which
|
||||
is why a write-up ever took two calls.
|
||||
|
||||
## The branch tool
|
||||
|
||||
`~/.local/bin/task` owns the branch, the commit identity and the PR. One task,
|
||||
one session, one PR.
|
||||
|
||||
```sh
|
||||
task start <vikunja-id> # branch off origin/master, write TASK.md, fetch review comments
|
||||
task pr # push, open or refresh the PR, label Vikunja, notify
|
||||
task comments # re-pull this branch's review comments into .task/
|
||||
```
|
||||
|
||||
`/pickup` opens a session and `/wrap` closes it. Wrap at roughly half context
|
||||
rather than letting the session compact.
|
||||
|
||||
## Guards
|
||||
|
||||
Two hooks in `.githooks/`, tracked, wired with `core.hooksPath`. A fresh clone
|
||||
needs `git config core.hooksPath .githooks`.
|
||||
|
||||
- `pre-commit` refuses master, and refuses more than 300 changed lines in
|
||||
non-markdown files. Markdown is exempt and may land as one batch.
|
||||
- `commit-msg` requires the subject to end with `(V-<id>)`. `V-` and not `#`,
|
||||
because Gitea autolinks `#123` to a Gitea issue, which is the wrong tracker.
|
||||
|
||||
Two more guards live outside the repo, in `~/.claude/hooks/`. `diff-budget.sh`
|
||||
blocks further edits past 600 changed lines on a `task/` branch.
|
||||
`prose_lint_hook.py` checks prose on every write. Both measure against
|
||||
`origin/master`, so a local master that is ahead of the remote makes the diff
|
||||
budget read high.
|
||||
|
||||
`--no-verify` exists. Using it means saying why in the commit body.
|
||||
@@ -0,0 +1,71 @@
|
||||
# The world chain
|
||||
|
||||
*Last verified: 2026-08-09 @ a9b480a*
|
||||
|
||||
What happens when the answer is not his. `CLAUDE.md` carries the boundary rule.
|
||||
This file carries the mechanism and the measurements behind it.
|
||||
|
||||
## What replaced "never phones home"
|
||||
|
||||
That promise was deprecated on 2026-07-31 by the owner's call. A 1.7B does not
|
||||
know enough to answer world questions, so she reads external sources. Four rules
|
||||
replaced it:
|
||||
|
||||
- **No telemetry, no cloud model, no third-party account.** That part never
|
||||
changes. Nothing about Maven is reported to anyone and inference stays on the
|
||||
box.
|
||||
- **The owner's data first, then the world.** Every source reading his facts,
|
||||
notes, calendar, tasks or house runs before anything outside. The personal
|
||||
boundary sits between them. Reading beats recalling for a small model.
|
||||
- **The owner's notes and facts are never search input.** Only the utterance
|
||||
goes out. Never the persona block, the history, or matched notes.
|
||||
- **External search is allowed and off unless configured**, like weather and
|
||||
telegram. The code default is off. `deploy/mavend.json` ships a `search`
|
||||
block, so it is on for this box and deleting the block turns it off again.
|
||||
|
||||
**Live search leads and the ZIMs are the fallback** (owner's call, 2026-08-02).
|
||||
A self-hosted SearXNG answers first. The Kiwix ZIMs on homesrv answer when the
|
||||
search is empty, unreachable, or the line is down.
|
||||
|
||||
## The gate is emptiness and nothing else
|
||||
|
||||
`Response.Empty()` is the whole gate. There is no quality threshold in front of
|
||||
it. Four signals were tried and none separates a real question from an invented
|
||||
one.
|
||||
|
||||
Token overlap was the closest and it still fails. "столица Франции" would lose
|
||||
its answer, because the answer is Париж and that word is not in the question
|
||||
(`docs/evals/2026-08-05-search-quality-signals.md`).
|
||||
|
||||
**The embedder is not a fifth signal.** Query-to-passage cosine measures topic
|
||||
and not whether the passage answers. The two sets overlap
|
||||
(`docs/evals/2026-08-09-kiwix-topic-retrieval.md`).
|
||||
|
||||
## Timeouts
|
||||
|
||||
The connect phase alone is capped at `dialTimeout`, 1.5s, because a blackholed
|
||||
host once cost the owner 8 seconds. A slow instance that did connect keeps the
|
||||
full 8 (`docs/evals/2026-08-05-kiwix-offline-fallback.md`).
|
||||
|
||||
## Kiwix
|
||||
|
||||
**A Russian question reads `wikipedia_ru_all_maxi_2026-02` verbatim** through
|
||||
`kiwix.book_ru`. The RU→EN rewriter is the workaround for an English book and is
|
||||
skipped there. Kiwix catalog names come from the filename, not the `<name>`
|
||||
field.
|
||||
|
||||
**Kiwix ranks by keyword overlap.** Never send it a whole sentence. `kiwix.Topic`
|
||||
drops the narrative request, the interrogative and a verb behind one.
|
||||
|
||||
`kiwix.TitlePath` tries the exact article first, since a ZIM is addressable by
|
||||
title and a wrong title is a 404. `TitleCandidates` tries the spoken form and
|
||||
then the capitalized one. Both apply on the verbatim path alone. The rewriter
|
||||
already reduces a question, and reducing twice takes the topic off its input
|
||||
(V-668).
|
||||
|
||||
## Which source answered
|
||||
|
||||
**The claiming query source is readable on `/chat`** as a badge beside the
|
||||
reply. It is carried on `ipc.ChatReply.Source` and noted by `noteQuerySource` in
|
||||
`cmd/mavend/querysource.go`. It rides the context, so `handleText` keeps the one
|
||||
string signature the mic, telegram and the web share.
|
||||
@@ -45,4 +45,17 @@ func TestDeployConfigLoads(t *testing.T) {
|
||||
if cfg.Voice.RouterThreshold <= 0 {
|
||||
t.Error("router threshold did not get its default")
|
||||
}
|
||||
|
||||
// The second reach (V-649). Deleting this block is how you turn ntfy off,
|
||||
// so its absence has to be loud: sev3-away nudges and away reminders route
|
||||
// to ntfy and to nothing else, and a nil sink drops them with no log and no
|
||||
// outbox row. The token is a ${VAR} that CI cannot resolve, so this checks
|
||||
// the wiring and not the credential.
|
||||
if cfg.Ntfy == nil {
|
||||
t.Fatal("deploy config has no ntfy block — sev3-away and away reminders " +
|
||||
"would have nowhere to land, and would vanish silently rather than fail")
|
||||
}
|
||||
if cfg.Ntfy.BaseURL == "" || cfg.Ntfy.Topic == "" {
|
||||
t.Errorf("ntfy block is incomplete: base_url=%q topic=%q", cfg.Ntfy.BaseURL, cfg.Ntfy.Topic)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -37,6 +37,16 @@ type EmbedderConfig struct {
|
||||
ModelPath string `json:"model_path,omitempty"`
|
||||
TokenizerPath string `json:"tokenizer_path,omitempty"`
|
||||
LibPath string `json:"lib_path,omitempty"`
|
||||
|
||||
// HeadsPath — the routing heads graph, which is a fine-tuned COPY of the
|
||||
// model above with four linear heads on its pooled output (V-664). Empty
|
||||
// means no heads, and the cascade runs exactly as it did before they
|
||||
// existed. It shares LibPath and TokenizerPath, and router_heads.json is
|
||||
// read from the same directory.
|
||||
//
|
||||
// It must never be pointed at ModelPath. Memory recall depends on the
|
||||
// resident copy scoring what it scored, and the fine-tuned one does not.
|
||||
HeadsPath string `json:"heads_path,omitempty"`
|
||||
}
|
||||
|
||||
// WeatherConfig configures the weather provider for voice queries.
|
||||
|
||||
@@ -1,6 +1,7 @@
|
||||
package config
|
||||
|
||||
import (
|
||||
"net/url"
|
||||
"strings"
|
||||
"time"
|
||||
)
|
||||
@@ -35,12 +36,54 @@ type WorkstationConfig struct {
|
||||
// 0 ⇒ DefaultWorkstationTimeout. A big model on a LAN host is slower than
|
||||
// the resident one, and a request that overruns falls back to the floor.
|
||||
Timeout Duration `json:"timeout,omitempty"`
|
||||
|
||||
// Stt — CrisperWhisper 2.0 on the same machine, a separate service on its
|
||||
// own port. Absent ⇒ every utterance goes to mavsttd, which is today.
|
||||
Stt *WorkstationSttConfig `json:"stt,omitempty"`
|
||||
}
|
||||
|
||||
// WorkstationSttConfig — speech-to-text on the workstation.
|
||||
//
|
||||
// It is a second service and not a second endpoint on mavgpud: whisper.cpp
|
||||
// cannot load CrisperWhisper 2.0 at all, because it derives its language count
|
||||
// from the vocabulary size and CW2's 51897 tokens shift seven special token
|
||||
// ids. So CW2 runs under transformers, and this block addresses it.
|
||||
//
|
||||
// Worth the trouble: CW2 turbo scores 10.4% WER in Russian against 27.5% for
|
||||
// the ggml-small.bin homesrv loads
|
||||
// (docs/evals/2026-08-09-crisperwhisper2-russian-wer.md).
|
||||
type WorkstationSttConfig struct {
|
||||
// URL — the transcribe endpoint, e.g.
|
||||
// "http://192.168.1.105:8081/transcribe". Empty ⇒ the block is normalised
|
||||
// to nil and mavsttd takes every turn.
|
||||
URL string `json:"url,omitempty"`
|
||||
|
||||
// Health — the admission endpoint. Empty ⇒ the URL's origin + "/health".
|
||||
// It answers 503 while the card is held, and that is the signal.
|
||||
Health string `json:"health,omitempty"`
|
||||
|
||||
// Token — the bearer token the service checks. Audio is the most sensitive
|
||||
// thing that crosses this seam, so a LAN deployment should set one. Write
|
||||
// it as ${MAVEN_STT_TOKEN} and keep the value in deploy/telegram.env, the
|
||||
// way every other secret in this file is written.
|
||||
Token string `json:"token,omitempty"`
|
||||
|
||||
// Probe — how often admission is re-checked. 0 ⇒ DefaultWorkstationProbe.
|
||||
Probe Duration `json:"probe,omitempty"`
|
||||
|
||||
// Timeout — the per-request budget for one utterance. 0 ⇒
|
||||
// DefaultWorkstationSttTimeout. A request that overruns falls back to
|
||||
// mavsttd, which costs a worse transcript and not the turn.
|
||||
Timeout Duration `json:"timeout,omitempty"`
|
||||
}
|
||||
|
||||
// Workstation defaults, applied in normaliseWorkstation.
|
||||
const (
|
||||
DefaultWorkstationProbe = 15 * time.Second
|
||||
DefaultWorkstationTimeout = 90 * time.Second
|
||||
// One utterance, not one completion. A voice turn waits on this, so the
|
||||
// budget is a few seconds and not a minute and a half.
|
||||
DefaultWorkstationSttTimeout = 10 * time.Second
|
||||
)
|
||||
|
||||
// normaliseWorkstation applies the block's defaults. No address, no preferred
|
||||
@@ -63,4 +106,36 @@ func (c *Config) normaliseWorkstation() {
|
||||
if w.Timeout <= 0 {
|
||||
w.Timeout = Duration(DefaultWorkstationTimeout)
|
||||
}
|
||||
normaliseWorkstationStt(w)
|
||||
}
|
||||
|
||||
// normaliseWorkstationStt applies the speech-to-text block's defaults. No
|
||||
// address, no remote: mavsttd then takes every utterance, which is today.
|
||||
func normaliseWorkstationStt(w *WorkstationConfig) {
|
||||
if w.Stt != nil && strings.TrimSpace(w.Stt.URL) == "" {
|
||||
w.Stt = nil
|
||||
}
|
||||
if w.Stt == nil {
|
||||
return
|
||||
}
|
||||
s := w.Stt
|
||||
if strings.TrimSpace(s.Health) == "" {
|
||||
s.Health = healthOrigin(s.URL)
|
||||
}
|
||||
if s.Probe <= 0 {
|
||||
s.Probe = Duration(DefaultWorkstationProbe)
|
||||
}
|
||||
if s.Timeout <= 0 {
|
||||
s.Timeout = Duration(DefaultWorkstationSttTimeout)
|
||||
}
|
||||
}
|
||||
|
||||
// healthOrigin derives the admission endpoint from the transcribe endpoint.
|
||||
// The URL names a path, so appending to it would ask for /transcribe/health.
|
||||
func healthOrigin(raw string) string {
|
||||
u, err := url.Parse(raw)
|
||||
if err != nil || u.Host == "" {
|
||||
return strings.TrimRight(raw, "/") + "/health"
|
||||
}
|
||||
return u.Scheme + "://" + u.Host + "/health"
|
||||
}
|
||||
|
||||
@@ -7,11 +7,17 @@
|
||||
// the relay). the dispatcher already strips detail off away sendables; the
|
||||
// sink uses the same helper so it can't leak the body on its own either.
|
||||
//
|
||||
// ntfy runs locally (docker, 127.0.0.1:8085, deny-all auth). maven publishes
|
||||
// with a dedicated user (write-only to maven-* topics) — the credential is a
|
||||
// delivery-config secret, not a db key; a popped ntfy sink can push spam to
|
||||
// your phone, nothing else. matches the module key-isolation invariant: the
|
||||
// sink never holds the sqlcipher key.
|
||||
// ntfy is a self-hosted server with deny-all auth — ntfy.kvmx.ru as of
|
||||
// 07-08-2026, reached directly, not through the socks relay telegram needs.
|
||||
// maven publishes with a write-only token scoped to its own topic; the
|
||||
// credential is a delivery-config secret, not a db key. a popped ntfy sink
|
||||
// can push spam to that one topic, nothing else — it cannot read the topic
|
||||
// back and it never holds the sqlcipher key.
|
||||
//
|
||||
// this is the second reach, and the reason there is one is that telegram was
|
||||
// the only one (V-649). telegram needs api.telegram.org, a socks relay on the
|
||||
// host and a matching ufw rule, three things in series that have each broken
|
||||
// once. ntfy shares none of them.
|
||||
package ntfysink
|
||||
|
||||
import (
|
||||
@@ -31,11 +37,29 @@ import (
|
||||
// the credential lives in the daemon's config (or a systemd credential),
|
||||
// never in the binary.
|
||||
type Config struct {
|
||||
BaseURL string // e.g. http://127.0.0.1:8085 (no trailing path)
|
||||
Topic string // e.g. maven (all maven notifications land here)
|
||||
Username string // basic auth; empty = anonymous (won't work with deny-all)
|
||||
Password string // basic auth
|
||||
Timeout time.Duration // per-request; 0 = DefaultTimeout
|
||||
// BaseURL — the ntfy server, no trailing path. Required.
|
||||
BaseURL string `json:"base_url"`
|
||||
|
||||
// Topic — where maven publishes. Required. All maven notifications land
|
||||
// on this one topic; severity rides the Priority header, not the topic.
|
||||
Topic string `json:"topic"`
|
||||
|
||||
// Token — an ntfy access token, sent as a bearer. This is the preferred
|
||||
// credential: ntfy scopes a token to a topic and to write-only, so a
|
||||
// popped sink can push to this one topic and cannot read it back or
|
||||
// touch another. Revoking it does not disturb a password anyone else
|
||||
// uses. Mutually exclusive with Username.
|
||||
Token string `json:"token,omitempty"`
|
||||
|
||||
// Username, Password — basic auth, for a server that has no tokens.
|
||||
// Empty username means no credential is sent at all, which a deny-all
|
||||
// server rejects.
|
||||
Username string `json:"username,omitempty"`
|
||||
Password string `json:"password,omitempty"`
|
||||
|
||||
// Timeout — per-request; 0 = DefaultTimeout. A dead server must not hang
|
||||
// the tick loop.
|
||||
Timeout time.Duration `json:"-"`
|
||||
}
|
||||
|
||||
const DefaultTimeout = 10 * time.Second
|
||||
@@ -59,6 +83,12 @@ func New(cfg Config) (*Sink, error) {
|
||||
if cfg.Topic == "" {
|
||||
return nil, fmt.Errorf("ntfysink: Topic is required")
|
||||
}
|
||||
// Refuse rather than pick. Two credentials configured means someone
|
||||
// intended one of them, and guessing which would send the other nowhere
|
||||
// and leave a working config that is not the one they wrote.
|
||||
if cfg.Token != "" && cfg.Username != "" {
|
||||
return nil, fmt.Errorf("ntfysink: set Token or Username, not both")
|
||||
}
|
||||
to := cfg.Timeout
|
||||
if to == 0 {
|
||||
to = DefaultTimeout
|
||||
@@ -84,7 +114,9 @@ func (s *Sink) Send(ctx context.Context, d delivery.Sendable) error {
|
||||
}
|
||||
req.Header.Set("Title", "maven")
|
||||
req.Header.Set("Priority", priorityFor(d).String())
|
||||
if s.cfg.Username != "" {
|
||||
if s.cfg.Token != "" {
|
||||
req.Header.Set("Authorization", "Bearer "+s.cfg.Token)
|
||||
} else if s.cfg.Username != "" {
|
||||
req.SetBasicAuth(s.cfg.Username, s.cfg.Password)
|
||||
}
|
||||
|
||||
|
||||
@@ -224,6 +224,37 @@ func TestSendNoAuthWhenUsernameEmpty(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestSendSetsBearerToken — the deployed credential (V-649) is an ntfy access
|
||||
// token scoped write-only to the maven topic, not a password. A token sent as
|
||||
// basic auth is rejected by ntfy, so the header shape is the whole test.
|
||||
func TestSendSetsBearerToken(t *testing.T) {
|
||||
rs := newRecordingServer(t, 200, "")
|
||||
srv := httptest.NewServer(rs.handler())
|
||||
defer srv.Close()
|
||||
|
||||
sink, _ := New(Config{BaseURL: srv.URL, Topic: "maven", Token: "tk_secret"})
|
||||
if err := sink.Send(context.Background(), nudgeSendable(loop.Sev3, "down")); err != nil {
|
||||
t.Fatalf("Send: %v", err)
|
||||
}
|
||||
_, _, _, auth, _, _ := rs.snapshot()
|
||||
if auth != "Bearer tk_secret" {
|
||||
t.Fatalf("auth: want 'Bearer tk_secret', got %q", auth)
|
||||
}
|
||||
}
|
||||
|
||||
// TestNewRejectsBothCredentials — configuring a token and a username means one
|
||||
// of them was meant and the other is a leftover. Picking either would leave a
|
||||
// server that authenticates against a credential nobody wrote down.
|
||||
func TestNewRejectsBothCredentials(t *testing.T) {
|
||||
_, err := New(Config{BaseURL: "http://x", Topic: "maven", Token: "tk_x", Username: "maven"})
|
||||
if err == nil {
|
||||
t.Fatal("New accepted both a token and a username")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "not both") {
|
||||
t.Errorf("error does not say which to fix: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSendTitleIsMaven(t *testing.T) {
|
||||
rs := newRecordingServer(t, 200, "")
|
||||
srv := httptest.NewServer(rs.handler())
|
||||
|
||||
@@ -41,6 +41,78 @@ type PendingQuestion struct {
|
||||
Attempts int // questions already asked
|
||||
// MaxAttempts caps Attempts. 0 ⇒ DefaultMaxAttempts.
|
||||
MaxAttempts int
|
||||
// Suspends counts how many times this question has stepped aside for
|
||||
// something he asked instead, and come back on the end of the answer. It is
|
||||
// deliberately NOT an attempt: a side query is not a failed answer, and
|
||||
// charging it a retry is the V-554 shape. See CanResume for why it is
|
||||
// counted at all.
|
||||
Suspends int
|
||||
// Rides counts every turn this question has ridden out on the end of
|
||||
// someone else's reply, over the whole life of the request. Unlike Suspends
|
||||
// it is never reset and never re-based, which is the only property that
|
||||
// matters about it (V-663).
|
||||
Rides int
|
||||
}
|
||||
|
||||
// MaxSuspends — how many times one question may step aside and come back before
|
||||
// she lets the request go (V-654).
|
||||
//
|
||||
// It exists because suspension had no bound of any kind. A side query spends no
|
||||
// attempt, so MaxAttempts never applies to it, and it restarts the 90s clock, so
|
||||
// the TTL never arrives either. Measured on 2026-08-07: one unfilled time slot
|
||||
// rode the end of six consecutive unrelated replies and stopped only when a
|
||||
// seventh turn happened to read as a failed answer.
|
||||
//
|
||||
// Three, matching DefaultMaxAttempts, and for the same reason. Once he has
|
||||
// asked for three other things without touching the question, the likely truth
|
||||
// is that he has moved on and has not said so.
|
||||
const MaxSuspends = 3
|
||||
|
||||
// MaxRides — how many turns one question may ride out on the end of an
|
||||
// unrelated reply, counted over its whole life (V-663).
|
||||
//
|
||||
// MaxSuspends did not move the measurement it was written for. Twenty-six of
|
||||
// 140 turns carried a tail before it landed and twenty-six carried one after.
|
||||
// Every bound on this question is rearmed by something ordinary:
|
||||
//
|
||||
// - The TTL is an inactivity timer, and both noteSuspended and reaskOrGiveUp
|
||||
// restart it, so it cannot arrive while he keeps talking.
|
||||
// - Suspends is zeroed by any turn that reads as an answer, which is where
|
||||
// "спасибо" and "привет" land. It resets before anything is known to have
|
||||
// been filled.
|
||||
// - askRemainingGap builds a fresh question for the second gap, so a reminder
|
||||
// with two gaps gets a new allowance halfway through.
|
||||
//
|
||||
// So Suspends only bites on four strictly consecutive side queries with nothing
|
||||
// chat-like between them, which is not the shape real conversation has. Rides is
|
||||
// the same idea with the resets taken out: set once, incremented, carried
|
||||
// across a re-park, and read by nothing that could lower it.
|
||||
//
|
||||
// The shape it is aimed at is measured, not imagined. In the 2026-08-08 run one
|
||||
// question about a reminder's day rode turns 7 to 13 and ended only because
|
||||
// turn 14 was a new request. Three asides, then two turns that read as failed
|
||||
// answers, then two more asides. The asides spend no attempt and the answers
|
||||
// reset Suspends, so the two bounds take turns being rearmed by the other's
|
||||
// traffic.
|
||||
//
|
||||
// Four, not three. It has to be looser than MaxSuspends or that bound is dead
|
||||
// code, because Rides is never lower than Suspends and would always fire first.
|
||||
//
|
||||
// Do not read this as a fix for the whole ride. It ends the measured one a turn
|
||||
// early and no more. Most of that ride's length is attempts, spent by turns
|
||||
// like "спасибо" and "привет" being read as failed answers to a question about
|
||||
// a day. That is a defect in classifyTurnRole and not in any bound here.
|
||||
const MaxRides = 4
|
||||
|
||||
// CanResume reports whether this question may step aside once more. False ⇒ the
|
||||
// caller lets the request go and says so; it must never simply stop resuming,
|
||||
// because a question dropped in silence reads as one that was answered.
|
||||
//
|
||||
// Two bounds, and they answer different questions. Suspends asks whether he has
|
||||
// walked away from this exchange in the last few turns. Rides asks whether this
|
||||
// question has been riding long enough that the answer is no regardless.
|
||||
func (q *PendingQuestion) CanResume() bool {
|
||||
return q.Suspends < MaxSuspends && q.Rides < MaxRides
|
||||
}
|
||||
|
||||
// Action reads the parked question as the typed action it is assembling
|
||||
|
||||
@@ -127,6 +127,22 @@ func (c *Client) Article(ctx context.Context, path string, maxRunes int) (crawl.
|
||||
return crawl.Extract(u, body, maxRunes), nil
|
||||
}
|
||||
|
||||
// TitlePath is the article path for an exact title, for Article to fetch.
|
||||
//
|
||||
// It exists because a ZIM is addressable by title and the full-text index is
|
||||
// not the only way in. "Франция", "TCP" and "Небо" resolve; "Трюмбальная
|
||||
// нидроскопия" is a 404, which is the honest answer and the reason this is
|
||||
// safe to try first. Measured on 2026-08-09, keyword search on the same terms
|
||||
// returns "Список пэров Франции" and "Список портов TCP и UDP" instead.
|
||||
//
|
||||
// A miss is normal rather than a failure. An article whose title inverts a name
|
||||
// ("Торвальдс, Линус") is a 404 here and the first hit in search, so the caller
|
||||
// falls through and loses nothing.
|
||||
func TitlePath(book, title string) string {
|
||||
t := strings.ReplaceAll(strings.TrimSpace(title), " ", "_")
|
||||
return "/content/" + url.PathEscape(book) + "/A/" + url.PathEscape(t)
|
||||
}
|
||||
|
||||
// rss mirrors just the bits of the RSS 2.0 reply we use.
|
||||
type rss struct {
|
||||
Items []struct {
|
||||
|
||||
@@ -0,0 +1,69 @@
|
||||
package kiwix
|
||||
|
||||
import (
|
||||
"context"
|
||||
"os"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// TestLiveTopicBeatsTheSentence — the measurement V-668 turned on, kept as a
|
||||
// test so the claim can be re-run rather than believed.
|
||||
//
|
||||
// It prints the article the old path returned and the article the new one
|
||||
// returns, for the same question. It asserts nothing about which is better,
|
||||
// because "is this the right article" is a human's call. It fails only if the
|
||||
// two paths agree on every case, which would mean the change does nothing.
|
||||
//
|
||||
// MAVEN_KIWIX_URL=http://127.0.0.1:8034 make t PKG=./internal/kiwix/ RUN=TestLive V=1
|
||||
func TestLiveTopicBeatsTheSentence(t *testing.T) {
|
||||
base := os.Getenv("MAVEN_KIWIX_URL")
|
||||
if base == "" {
|
||||
t.Skip("MAVEN_KIWIX_URL unset — point it at the kiwix-server host port")
|
||||
}
|
||||
const book = "wikipedia_ru_all_maxi_2026-02"
|
||||
c := New(base)
|
||||
questions := []string{
|
||||
"что такое TCP?",
|
||||
"что такое фотосинтез",
|
||||
"кто такой Линус Торвальдс?",
|
||||
"кто написал Войну и мир",
|
||||
"что такое чёрная дыра",
|
||||
"почему небо голубое",
|
||||
"почему трава зелёная",
|
||||
"столица Франции",
|
||||
}
|
||||
moved := 0
|
||||
for _, q := range questions {
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 20*time.Second)
|
||||
before := firstTitle(ctx, c, q, book)
|
||||
topic := Topic(q)
|
||||
after := ""
|
||||
for _, cand := range TitleCandidates(topic) {
|
||||
if page, err := c.Article(ctx, TitlePath(book, cand), 400); err == nil && page.Text != "" {
|
||||
after = page.Title + " (by title)"
|
||||
break
|
||||
}
|
||||
}
|
||||
if after == "" {
|
||||
after = firstTitle(ctx, c, topic, book)
|
||||
}
|
||||
cancel()
|
||||
if before != after {
|
||||
moved++
|
||||
}
|
||||
t.Logf("%-30s before=%-34q after=%q", q, before, after)
|
||||
}
|
||||
t.Logf("%d of %d questions reach a different article", moved, len(questions))
|
||||
if moved == 0 {
|
||||
t.Error("the topic path returns exactly what the sentence path returned")
|
||||
}
|
||||
}
|
||||
|
||||
func firstTitle(ctx context.Context, c *Client, pattern, book string) string {
|
||||
hits, err := c.Search(ctx, pattern, book, 3)
|
||||
if err != nil || len(hits) == 0 {
|
||||
return "(nothing)"
|
||||
}
|
||||
return hits[0].Title
|
||||
}
|
||||
@@ -0,0 +1,118 @@
|
||||
package kiwix
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"unicode"
|
||||
|
||||
"github.com/kami/maven/internal/lexicon"
|
||||
"github.com/kami/maven/internal/morph"
|
||||
)
|
||||
|
||||
// Topic reduces a question to the thing it is about, because Kiwix ranks by
|
||||
// keyword overlap and a whole sentence buries the keyword that matters.
|
||||
//
|
||||
// This package's own doc says it: "why is the sky blue" finds a TV episode.
|
||||
// Measured against the Russian ZIM on 2026-08-09, the sentence and the topic
|
||||
// return different articles for the same question. "кто написал Войну и мир"
|
||||
// returns "Радуйся, мир (Доктор Кто)"; "Войну и мир" returns the novel first.
|
||||
// "что такое TCP" returns "Перехват TCP-соединения"; "TCP" returns TCP. The
|
||||
// English path had a rewriter doing this with a model call. The Russian path
|
||||
// reads the book verbatim (V-508) and had nothing.
|
||||
//
|
||||
// It drops three things off the front and stops: the narrative request, the
|
||||
// interrogative, and a verb sitting between them and the noun. Everything else
|
||||
// is kept, because a word this cannot classify is more likely the topic than
|
||||
// noise. An empty return means the utterance was question words alone, and the
|
||||
// caller searches the sentence as before.
|
||||
func Topic(utterance string) string {
|
||||
words := strings.Fields(strings.TrimSpace(utterance))
|
||||
cut := 0
|
||||
for cut < len(words) {
|
||||
w := strings.Trim(strings.ToLower(words[cut]), ".,!?…:;\"'«»")
|
||||
if w == "" {
|
||||
cut++
|
||||
continue
|
||||
}
|
||||
switch {
|
||||
case inList(lexicon.NarrativeRequests(), w),
|
||||
inList(lexicon.Interrogatives(), w),
|
||||
inList(lexicon.FirstPerson(), w),
|
||||
// "что ТАКОЕ x", "кто ТАКОЙ x" — the copula that only ever follows
|
||||
// an interrogative, and never a topic on its own.
|
||||
cut > 0 && isCopula(w),
|
||||
// "расскажи ПРО x", "о x". One-letter and two-letter prepositions
|
||||
// are not a closed class worth a lexicon set of their own.
|
||||
cut > 0 && isLeadingPreposition(w),
|
||||
// "кто НАПИСАЛ Войну и мир". A verb here is the question's own
|
||||
// verb, not part of the title. Only after something was already
|
||||
// dropped, so "написал отчёт" as a topic survives intact.
|
||||
cut > 0 && morph.IsVerbForm(w):
|
||||
cut++
|
||||
default:
|
||||
// The question mark is the sentence's, not the title's, and Kiwix
|
||||
// carries it into the keyword match.
|
||||
topic := strings.TrimRight(strings.Join(words[cut:], " "), " .,!?…:;\"'«»")
|
||||
if !hasLetter(topic) {
|
||||
return ""
|
||||
}
|
||||
return topic
|
||||
}
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
// TitleCandidates is the topic as it might be titled, best first.
|
||||
//
|
||||
// A ZIM title is capitalized and the utterance is not: measured on 2026-08-09,
|
||||
// `/A/фотосинтез` is a 404 and `/A/Фотосинтез` is a 200. The spoken form is
|
||||
// tried first anyway, because a title that begins lowercase on purpose
|
||||
// ("iPhone") would not survive capitalizing it. Both are one request each
|
||||
// against a server on the same box, and a miss is a 404 rather than a wrong
|
||||
// article.
|
||||
func TitleCandidates(topic string) []string {
|
||||
if topic == "" {
|
||||
return nil
|
||||
}
|
||||
r := []rune(topic)
|
||||
up := unicode.ToUpper(r[0])
|
||||
if up == r[0] {
|
||||
return []string{topic}
|
||||
}
|
||||
return []string{topic, string(up) + string(r[1:])}
|
||||
}
|
||||
|
||||
func isCopula(w string) bool {
|
||||
switch w {
|
||||
case "такое", "такой", "такая", "такие", "is", "are", "was", "were":
|
||||
return true
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
func isLeadingPreposition(w string) bool {
|
||||
switch w {
|
||||
case "про", "о", "об", "обо", "по", "about", "of", "on":
|
||||
return true
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
func inList(list []string, w string) bool {
|
||||
for _, x := range list {
|
||||
if x == w {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// hasLetter is the guard against a topic that reduced to punctuation or digits
|
||||
// alone, which no ZIM title matches.
|
||||
func hasLetter(s string) bool {
|
||||
for _, r := range s {
|
||||
if unicode.IsLetter(r) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
@@ -0,0 +1,67 @@
|
||||
package kiwix
|
||||
|
||||
import "testing"
|
||||
|
||||
// The cases the 2026-08-09 measurement turned on, plus the ones a topic must
|
||||
// not damage. Each left column returned a wrong article when it was sent whole.
|
||||
func TestTopicKeepsTheThingTheQuestionIsAbout(t *testing.T) {
|
||||
cases := []struct{ utterance, want string }{
|
||||
{"что такое TCP?", "TCP"},
|
||||
{"что такое фотосинтез", "фотосинтез"},
|
||||
{"кто такой Линус Торвальдс?", "Линус Торвальдс"},
|
||||
{"кто написал Войну и мир", "Войну и мир"},
|
||||
{"расскажи про битву при Ватерлоо", "битву при Ватерлоо"},
|
||||
{"what is photosynthesis", "photosynthesis"},
|
||||
// No question word, so there is nothing to drop. The topic is the
|
||||
// whole utterance and the search is what it was before.
|
||||
{"столица Франции", "столица Франции"},
|
||||
{"почему небо голубое", "небо голубое"},
|
||||
}
|
||||
for _, c := range cases {
|
||||
if got := Topic(c.utterance); got != c.want {
|
||||
t.Errorf("Topic(%q) = %q, want %q", c.utterance, got, c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A verb only goes when a question word already went. Otherwise "написал
|
||||
// отчёт" loses the verb that names what he means.
|
||||
func TestTopicDropsAVerbOnlyBehindAQuestionWord(t *testing.T) {
|
||||
if got := Topic("написал отчёт"); got != "написал отчёт" {
|
||||
t.Errorf("Topic dropped a leading verb with no question word: %q", got)
|
||||
}
|
||||
}
|
||||
|
||||
// Question words alone reduce to nothing, and the caller reads that as "no
|
||||
// topic" and searches the sentence rather than searching the empty string.
|
||||
func TestTopicIsEmptyWhenNothingIsLeft(t *testing.T) {
|
||||
for _, q := range []string{"что такое?", "кто?", "почему", "???"} {
|
||||
if got := Topic(q); got != "" {
|
||||
t.Errorf("Topic(%q) = %q, want empty", q, got)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestTitlePathEscapesAndUnderscores(t *testing.T) {
|
||||
got := TitlePath("wikipedia_ru_all_maxi_2026-02", "Чёрная дыра")
|
||||
want := "/content/wikipedia_ru_all_maxi_2026-02/A/%D0%A7%D1%91%D1%80%D0%BD%D0%B0%D1%8F_%D0%B4%D1%8B%D1%80%D0%B0"
|
||||
if got != want {
|
||||
t.Errorf("TitlePath = %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
// A ZIM title carries a leading capital and the utterance does not. The spoken
|
||||
// form is still tried first, so a title that begins lowercase on purpose keeps
|
||||
// its chance.
|
||||
func TestTitleCandidatesTryTheSpokenFormFirst(t *testing.T) {
|
||||
got := TitleCandidates("фотосинтез")
|
||||
if len(got) != 2 || got[0] != "фотосинтез" || got[1] != "Фотосинтез" {
|
||||
t.Errorf("TitleCandidates = %q", got)
|
||||
}
|
||||
if got := TitleCandidates("TCP"); len(got) != 1 || got[0] != "TCP" {
|
||||
t.Errorf("an already-capital topic was tried twice: %q", got)
|
||||
}
|
||||
if got := TitleCandidates(""); got != nil {
|
||||
t.Errorf("TitleCandidates(\"\") = %q, want nil", got)
|
||||
}
|
||||
}
|
||||
@@ -119,6 +119,11 @@ func PartsOfDay() []string { return words("parts_of_day") }
|
||||
// ReminderVerbs returns the imperatives that open a reminder.
|
||||
func ReminderVerbs() []string { return words("reminder_verbs") }
|
||||
|
||||
// Pleasantries returns the whole utterances that greet, thank or say goodbye.
|
||||
// Whole utterances and not tokens: see the set's own note for why the tokens
|
||||
// are unsafe alone.
|
||||
func Pleasantries() []string { return words("pleasantries") }
|
||||
|
||||
// TaskDoneWords returns the words that finish a task, and TaskDropWords the
|
||||
// words that abandon one. Two sets rather than one with a value, because the
|
||||
// store records which of the two happened and the caller has to say so.
|
||||
|
||||
@@ -176,6 +176,19 @@
|
||||
"morning", "afternoon", "evening", "night"
|
||||
]
|
||||
},
|
||||
"pleasantries": {
|
||||
"note": "Whole utterances that greet, thank or say goodbye. They ask for nothing and answer nothing, so a parked question must neither consume them as a failed answer nor be dropped by them (V-663). Matched as WHOLE utterances and never as tokens, because the tokens are not safe alone: \"вечер\" answers \"это утра или вечера?\" and \"нет\" answers a confirm. Anything that could fill a slot stays out. The control words (\"стоп\", \"отмена\") stay out too, because isCancel already owns them and they mean something stronger.",
|
||||
"words": [
|
||||
"привет", "приветик", "здравствуй", "здравствуйте",
|
||||
"доброе утро", "добрый день", "добрый вечер",
|
||||
"пока", "прощай", "до свидания", "спокойной ночи",
|
||||
"спасибо", "спасибо тебе", "большое спасибо", "благодарю",
|
||||
"извини", "извините", "прости", "простите",
|
||||
"hi", "hello", "hey", "bye", "goodbye",
|
||||
"good morning", "good evening", "good night",
|
||||
"thanks", "thank you", "thanks a lot", "sorry"
|
||||
]
|
||||
},
|
||||
"reminder_verbs": {
|
||||
"note": "The imperatives that mean \"remind me\", in the forms he speaks. The same kind of set as capture_verbs and decided the same way: it is her vocabulary, not a discovery about Russian (Vikunja #530). The alarm verbs joined them in V-627. \"разбуди меня в 6:30\" is a reminder that fires at the hour he gets up, and the set knew no form of it, so an alarm reached IntentReminder only by resembling one to the embedder.",
|
||||
"words": [
|
||||
|
||||
@@ -14,12 +14,15 @@ import (
|
||||
"github.com/kami/maven/internal/decision"
|
||||
)
|
||||
|
||||
// The two routing engines, named as claimants. They are one stage and not two,
|
||||
// because only one of them ever runs: the classifier is reached when the model
|
||||
// is absent or errored, never alongside it.
|
||||
// The three routing engines, named as claimants. The model and the classifier
|
||||
// are one stage and not two, because only one of them ever runs: the classifier
|
||||
// is reached when the model is absent or errored, never alongside it. The heads
|
||||
// run before both and decline on low confidence, so they can appear beside
|
||||
// either one in a record.
|
||||
const (
|
||||
claimantLLM = "llm-router"
|
||||
claimantClassifier = "classifier"
|
||||
claimantHeads = "routing-heads"
|
||||
)
|
||||
|
||||
// thinReason names which arm of gateLLMDecision cut the confidence. The gate
|
||||
|
||||
@@ -0,0 +1,22 @@
|
||||
package router
|
||||
|
||||
import (
|
||||
"os"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// TestDumpPrompt writes the router prompt and grammar to disk so the training
|
||||
// workspace labels with the daemon's own contract rather than a retyped copy.
|
||||
// It is inert unless MAVEN_DUMP_PROMPT names a directory.
|
||||
func TestDumpPrompt(t *testing.T) {
|
||||
dir := os.Getenv("MAVEN_DUMP_PROMPT")
|
||||
if dir == "" {
|
||||
t.Skip("MAVEN_DUMP_PROMPT unset")
|
||||
}
|
||||
if err := os.WriteFile(dir+"/route_system.txt", []byte(routeSystem), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := os.WriteFile(dir+"/route_grammar.gbnf", []byte(routeGrammar), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
@@ -1,10 +1,13 @@
|
||||
package router
|
||||
|
||||
import "testing"
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func TestEmbedderIDFromModelPath(t *testing.T) {
|
||||
got := modelIDFromPath("/opt/maven/models/embedder/multilingual-e5-small.onnx")
|
||||
if got != "multilingual-e5-small@384" {
|
||||
if got != "multilingual-e5-small@384/tok2" {
|
||||
t.Fatalf("modelIDFromPath = %q", got)
|
||||
}
|
||||
// A different model file must produce a different id, even at 384 dim.
|
||||
@@ -12,6 +15,13 @@ func TestEmbedderIDFromModelPath(t *testing.T) {
|
||||
if old == got {
|
||||
t.Fatal("two different models share one id")
|
||||
}
|
||||
// The tokenizer is half of what makes a vector, and it changes under a
|
||||
// model file whose name never moves (V-664). An id that ignored it would
|
||||
// leave stored passages in one space and every new query in another, with
|
||||
// nothing to trigger the re-embed.
|
||||
if !strings.Contains(got, "/tok") {
|
||||
t.Fatalf("id %q does not name the tokenizer revision", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestEmbedderIDIncludesDim(t *testing.T) {
|
||||
|
||||
@@ -36,17 +36,27 @@ var fixtureJSON []byte
|
||||
//
|
||||
// Intent is empty exactly when WantClarify is set: the contract there is that
|
||||
// the router refuses instead of guessing.
|
||||
//
|
||||
// WantSource is a pointer because the destination has three states and a bare
|
||||
// string only has two (V-659). Absent means the case does not score a
|
||||
// destination at all, which is every intent but query: a fact, a reminder, a
|
||||
// note, an act, a chat or a system turn never reaches queryWalk. Present and
|
||||
// empty is the SourceUnknown contract — the decider must name nothing and let
|
||||
// the daemon walk the whole chain, which is the right answer whenever two
|
||||
// destinations can both answer and the utterance does not choose. Present and
|
||||
// named is a destination the route must produce.
|
||||
type Case struct {
|
||||
ID string `json:"id"`
|
||||
Utterance string `json:"utterance"`
|
||||
Lang string `json:"lang"`
|
||||
Intent router.Intent `json:"intent"`
|
||||
WantTime bool `json:"want_time"`
|
||||
WantFn bool `json:"want_fn"`
|
||||
WantFactKey string `json:"want_fact_key"`
|
||||
WantClarify bool `json:"want_clarify"`
|
||||
Tags []string `json:"tags"`
|
||||
Note string `json:"note"`
|
||||
ID string `json:"id"`
|
||||
Utterance string `json:"utterance"`
|
||||
Lang string `json:"lang"`
|
||||
Intent router.Intent `json:"intent"`
|
||||
WantTime bool `json:"want_time"`
|
||||
WantFn bool `json:"want_fn"`
|
||||
WantFactKey string `json:"want_fact_key"`
|
||||
WantClarify bool `json:"want_clarify"`
|
||||
WantSource *router.Source `json:"want_source,omitempty"`
|
||||
Tags []string `json:"tags"`
|
||||
Note string `json:"note"`
|
||||
}
|
||||
|
||||
// Fixture — the versioned envelope, same shape as
|
||||
@@ -118,6 +128,11 @@ type Outcome struct {
|
||||
// (a slot gap is a parser fix; a wrong intent is a router fix).
|
||||
IntentOK bool
|
||||
Reasons []string
|
||||
// SourceReason is set when the case labelled a destination and the route
|
||||
// named a different one. It is kept out of Reasons on purpose: the
|
||||
// destination is the second half of a route and it is scored separately,
|
||||
// so a wrong destination must not move the intent number (V-659).
|
||||
SourceReason string
|
||||
}
|
||||
|
||||
// Report — the aggregate. Accuracy is the headline; the rest exists so a
|
||||
@@ -139,7 +154,15 @@ type Report struct {
|
||||
// (reminder grammar → applyAction's time parser). Not a miss, but not a
|
||||
// full router-level win either; tracked so the two aren't conflated.
|
||||
SlotsDeferred int
|
||||
Outcomes []Outcome
|
||||
// SourceTotal counts the cases carrying a want_source, and SourceHit the
|
||||
// ones whose route named it. Reported apart from Passed because intent and
|
||||
// destination are two decisions, and one number hides which one moved.
|
||||
SourceTotal int
|
||||
SourceHit int
|
||||
// SourceConfusion counts want→got destination pairs. "" reads as the
|
||||
// SourceUnknown floor on either side.
|
||||
SourceConfusion map[string]int
|
||||
Outcomes []Outcome
|
||||
// Confusion counts want→got intent pairs, decided cases only.
|
||||
Confusion map[string]int
|
||||
// ByTag accuracy for the fixture's tags ("hard", "homelab", …).
|
||||
@@ -172,6 +195,17 @@ func (r Report) IntentAccuracy() float64 {
|
||||
return float64(r.IntentHit) / float64(r.Total)
|
||||
}
|
||||
|
||||
// SourceAccuracy — fraction of the labelled cases whose route named the right
|
||||
// destination. Denominator is SourceTotal and not Total, because most of the
|
||||
// fixture never reaches a query source and scoring those would report a
|
||||
// percentage of nothing.
|
||||
func (r Report) SourceAccuracy() float64 {
|
||||
if r.SourceTotal == 0 {
|
||||
return 0
|
||||
}
|
||||
return float64(r.SourceHit) / float64(r.SourceTotal)
|
||||
}
|
||||
|
||||
// Score runs every case through r and aggregates. It never fails the run on a
|
||||
// route error — an erroring case scores as a miss and is counted in Errors,
|
||||
// because "the model was down" and "the model was wrong" are different numbers
|
||||
@@ -186,11 +220,12 @@ func Score(ctx context.Context, name string, r Router, f Fixture) (Report, error
|
||||
return Report{}, err
|
||||
}
|
||||
rep := Report{
|
||||
Name: name,
|
||||
Total: len(f.Cases),
|
||||
Confusion: map[string]int{},
|
||||
ByTag: map[string]TagStat{},
|
||||
ByLang: map[string]TagStat{},
|
||||
Name: name,
|
||||
Total: len(f.Cases),
|
||||
Confusion: map[string]int{},
|
||||
SourceConfusion: map[string]int{},
|
||||
ByTag: map[string]TagStat{},
|
||||
ByLang: map[string]TagStat{},
|
||||
}
|
||||
lat := make([]time.Duration, 0, len(f.Cases))
|
||||
|
||||
@@ -242,6 +277,29 @@ func Score(ctx context.Context, name string, r Router, f Fixture) (Report, error
|
||||
}
|
||||
}
|
||||
|
||||
// The destination is scored outside the switch and outside Pass. A case
|
||||
// that clarified or landed the wrong intent named no destination, and
|
||||
// that is a real miss rather than a case to skip — otherwise the
|
||||
// denominator quietly drops every turn the route already lost. Only a
|
||||
// route error is skipped, because "the model was down" is the Errors
|
||||
// number and not a destination result.
|
||||
if c.WantSource != nil && err == nil {
|
||||
rep.SourceTotal++
|
||||
switch {
|
||||
case !o.IntentOK:
|
||||
// The route never got to a destination, so a match on the
|
||||
// SourceUnknown floor here would be a coincidence scored as a
|
||||
// win: a clarify names nothing and would satisfy "" for free.
|
||||
o.SourceReason = fmt.Sprintf("no destination, route missed %q", c.Intent)
|
||||
rep.SourceConfusion[string(*c.WantSource)+"→(no route)"]++
|
||||
case d.Source == *c.WantSource:
|
||||
rep.SourceHit++
|
||||
default:
|
||||
rep.SourceConfusion[string(*c.WantSource)+"→"+string(d.Source)]++
|
||||
o.SourceReason = fmt.Sprintf("source %q, want %q", d.Source, *c.WantSource)
|
||||
}
|
||||
}
|
||||
|
||||
o.Pass = len(o.Reasons) == 0
|
||||
if o.Pass {
|
||||
rep.Passed++
|
||||
@@ -298,25 +356,40 @@ func (r Report) String() string {
|
||||
r.Name, r.Passed, r.Total, 100*r.Accuracy(), 100*r.IntentAccuracy())
|
||||
fmt.Fprintf(&b, " clarify: %d false (asked, shouldn't) / %d missed (guessed, shouldn't) | errors: %d | slots deferred to daemon: %d\n",
|
||||
r.FalseClarify, r.MissedClarify, r.Errors, r.SlotsDeferred)
|
||||
if r.SourceTotal > 0 {
|
||||
fmt.Fprintf(&b, " destination: %d/%d labelled cases (%.1f%%)\n",
|
||||
r.SourceHit, r.SourceTotal, 100*r.SourceAccuracy())
|
||||
}
|
||||
fmt.Fprintf(&b, " latency: p50 %s p95 %s max %s\n", r.P50, r.P95, r.Max)
|
||||
fmt.Fprintf(&b, " by lang: %s\n", renderStats(r.ByLang))
|
||||
fmt.Fprintf(&b, " by tag: %s\n", renderStats(r.ByTag))
|
||||
if len(r.Confusion) > 0 {
|
||||
fmt.Fprintf(&b, " confusion: %s\n", renderCounts(r.Confusion))
|
||||
}
|
||||
if len(r.SourceConfusion) > 0 {
|
||||
fmt.Fprintf(&b, " destination confusion: %s\n", renderCounts(r.SourceConfusion))
|
||||
}
|
||||
return b.String()
|
||||
}
|
||||
|
||||
// Failures — the per-case detail, sorted by ID so two runs diff cleanly.
|
||||
// Failures — the per-case detail, sorted by ID so two runs diff cleanly. A case
|
||||
// that landed its intent and missed its destination is listed too, marked, so
|
||||
// the half that moved is readable without diffing two percentages.
|
||||
func (r Report) Failures() string {
|
||||
var b strings.Builder
|
||||
out := append([]Outcome(nil), r.Outcomes...)
|
||||
sort.Slice(out, func(i, j int) bool { return out[i].Case.ID < out[j].Case.ID })
|
||||
for _, o := range out {
|
||||
if o.Pass {
|
||||
continue
|
||||
switch {
|
||||
case !o.Pass:
|
||||
reasons := o.Reasons
|
||||
if o.SourceReason != "" {
|
||||
reasons = append(append([]string(nil), reasons...), o.SourceReason)
|
||||
}
|
||||
fmt.Fprintf(&b, " %s %q: %s\n", o.Case.ID, o.Case.Utterance, strings.Join(reasons, "; "))
|
||||
case o.SourceReason != "":
|
||||
fmt.Fprintf(&b, " %s %q: route ok, %s\n", o.Case.ID, o.Case.Utterance, o.SourceReason)
|
||||
}
|
||||
fmt.Fprintf(&b, " %s %q: %s\n", o.Case.ID, o.Case.Utterance, strings.Join(o.Reasons, "; "))
|
||||
}
|
||||
return b.String()
|
||||
}
|
||||
|
||||
@@ -266,6 +266,11 @@ func baselineGrammars(acts router.ActMatcher) []router.Grammar {
|
||||
// Same order as buildRouter (voicewire.go). The fixture is only worth
|
||||
// anything while its grammar set is the daemon's grammar set.
|
||||
grammars = append(grammars, router.AgendaQueryGrammars()...)
|
||||
// After the agenda rules and before the feed and list rules, same as
|
||||
// voicewire.go: "что такое лента" is a definition question and the feed
|
||||
// rule would claim it on the noun alone (V-655). Missing here until V-659,
|
||||
// so the fixture was scoring a grammar set the daemon does not run.
|
||||
grammars = append(grammars, router.WorldQueryGrammars()...)
|
||||
grammars = append(grammars, router.FeedQueryGrammar())
|
||||
// The list side of the same exposure: a phrasing with no possessive in it
|
||||
// ("список дел") routed system and never reached queryTasks (Vikunja #467).
|
||||
|
||||
@@ -0,0 +1,86 @@
|
||||
package eval
|
||||
|
||||
import (
|
||||
"context"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
|
||||
"github.com/kami/maven/internal/config"
|
||||
"github.com/kami/maven/internal/router"
|
||||
)
|
||||
|
||||
// TestONNXRoutingHeads — the cascade with the routing heads wired, which is
|
||||
// what V-664 deploys. Opt-in via MAVEN_ONNX_LIB, same as TestONNXBaseline, and
|
||||
// one TestONNX* per process.
|
||||
//
|
||||
// The comparison worth reading is against TestONNXBaseline, which is the same
|
||||
// cascade with the same grammars and the same classifier floor and no heads.
|
||||
// Only the middle arm varies.
|
||||
//
|
||||
// It also checks the Go unigram tokenizer against the Python one, because the
|
||||
// heads were trained through transformers and are read through a hand-written
|
||||
// tokenizer. A mismatch shows up here as a score below what Python measured on
|
||||
// the same weights, and nowhere else.
|
||||
func TestONNXRoutingHeads(t *testing.T) {
|
||||
lib := os.Getenv("MAVEN_ONNX_LIB")
|
||||
if lib == "" {
|
||||
t.Skip("MAVEN_ONNX_LIB unset — see AGENTS.md § Embedder model for intent routing")
|
||||
}
|
||||
// Absolute, because onnxruntime resolves a graph's external weights file
|
||||
// against the model path it was given, and a relative one lands in the
|
||||
// test's working directory.
|
||||
root, err := filepath.Abs("../../..")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
model := filepath.Join(root, "models/embedder/multilingual-e5-small/model_quantized.onnx")
|
||||
tok := filepath.Join(root, "models/embedder/multilingual-e5-small/tokenizer.json")
|
||||
heads := filepath.Join(root, "models/embedder/router-heads/router_heads.onnx")
|
||||
for _, p := range []string{lib, model, tok, heads} {
|
||||
if _, err := os.Stat(p); err != nil {
|
||||
t.Skipf("missing %s: %v", p, err)
|
||||
}
|
||||
}
|
||||
emb, err2 := router.NewONNXEmbedder(model, tok, lib)
|
||||
if err2 != nil {
|
||||
t.Skipf("onnx embedder unavailable: %v", err2)
|
||||
}
|
||||
err = nil
|
||||
defer emb.Close()
|
||||
|
||||
h, err := router.NewRouterHeads(heads, tok)
|
||||
if err != nil {
|
||||
t.Skipf("routing heads unavailable: %v", err)
|
||||
}
|
||||
defer h.Close()
|
||||
|
||||
f, err := Load()
|
||||
if err != nil {
|
||||
t.Fatalf("Load: %v", err)
|
||||
}
|
||||
rep, err := Score(context.Background(), "heads+classifier", withHeads(t, emb, h), f)
|
||||
if err != nil {
|
||||
t.Fatalf("Score: %v", err)
|
||||
}
|
||||
t.Log("\n" + rep.String() + rep.Failures())
|
||||
}
|
||||
|
||||
// withHeads mirrors newBaselineRouter and adds the one arm under test. It is a
|
||||
// separate function rather than a parameter so the baseline's signature stays
|
||||
// the shape every other test calls it with.
|
||||
func withHeads(t *testing.T, emb router.Embedder, h *router.RouterHeads) *router.Router {
|
||||
t.Helper()
|
||||
acts := router.DefaultActMatcher{Fns: actFns}
|
||||
return router.New(router.Config{
|
||||
Grammars: baselineGrammars(acts),
|
||||
Classifier: newBaselineClassifier(t, emb),
|
||||
Extractor: router.Extractor{
|
||||
Time: router.StubDateTimeParser{},
|
||||
Acts: acts,
|
||||
Facts: router.DefaultFactParser{},
|
||||
},
|
||||
Threshold: config.DefaultRouterThreshold,
|
||||
Heads: h,
|
||||
})
|
||||
}
|
||||
@@ -6,37 +6,41 @@
|
||||
"Held-out routing contract. Every utterance here is absent from models/seeds/*.txt (TestFixtureIsHeldOut enforces it verbatim) — scoring a classifier on its own seed phrases measures memorisation, not routing.",
|
||||
"This is a CONTRACT, not a snapshot of current behaviour. Cases the classifier cascade fails today are expected to stay in the file and fail loudly; that failure count is the number Vikunja #319 compares against the LLM router before #320 flips the default.",
|
||||
"Slot expectations are deployment-independent on purpose. want_fn is a boolean (the act must resolve to SOME allowlisted fn) because the allowlist lives in deploy config, not here. want_fact_key names the loop's rule keys (water/meal/sleep/break/shower) — a fact that lands under the wrong key silently starves the predicate that reads it.",
|
||||
"want_clarify cases carry intent \"\": the contract is that the router refuses rather than guesses. A confident answer there is a worse failure than a miss."
|
||||
"want_clarify cases carry intent \"\": the contract is that the router refuses rather than guesses. A confident answer there is a worse failure than a miss.",
|
||||
"want_source is the second half of a route (V-655). It is present only on query cases, because no other intent reaches queryWalk, and absent there means absent rather than SourceUnknown. Empty is a label and not a gap: it asserts that the decider must name nothing and let the daemon walk the whole chain in order, his data first.",
|
||||
"Seven cases assert that floor and six of them are homelab operations. They cluster because SourceRecall, SourceNetwork and SourceAttention overlap on every question about the box: mavpoll writes its observations into the fact store recall reads. That is a finding about the enum, not a gap in the labelling.",
|
||||
"ru-query-020 and ru-query-024 are the same utterance, as are ru-query-021 and ru-query-025. Both pairs differ in tags and note only, so both pairs are counted twice in every number this fixture reports.",
|
||||
"Every want_source is the destination that SHOULD claim the turn, which on ru-query-026 through 030 is not the one that did. Those five were observed failing on the box on 2026-08-07 (docs/evals/2026-08-07-week-of-usage.md). A fixture that passes on the day it is written measures nothing."
|
||||
],
|
||||
"cases": [
|
||||
{ "id": "ru-query-001", "utterance": "сколько воды я выпил с утра", "lang": "ru", "intent": "query", "tags": ["aggregate"] },
|
||||
{ "id": "ru-query-002", "utterance": "я сегодня вообще пил воду", "lang": "ru", "intent": "query", "tags": ["hard", "fact-shaped"], "note": "past-tense fact lexicon in a question — the classifier's fact centroid pulls this hard" },
|
||||
{ "id": "ru-query-003", "utterance": "во сколько я лёг вчера", "lang": "ru", "intent": "query", "tags": ["temporal"] },
|
||||
{ "id": "ru-query-004", "utterance": "давно я не тренировался", "lang": "ru", "intent": "query", "tags": ["hard", "no-question-word"] },
|
||||
{ "id": "ru-query-005", "utterance": "напоминания на завтра есть", "lang": "ru", "intent": "query", "tags": ["hard", "reminder-shaped"], "note": "asks about reminders, does not create one" },
|
||||
{ "id": "ru-query-006", "utterance": "что я записывал про кота", "lang": "ru", "intent": "query", "tags": ["recall"] },
|
||||
{ "id": "ru-query-007", "utterance": "сколько раз я ел вчера", "lang": "ru", "intent": "query", "tags": ["aggregate", "hard"] },
|
||||
{ "id": "ru-query-008", "utterance": "мой вес за последний месяц", "lang": "ru", "intent": "query", "tags": ["no-verb"] },
|
||||
{ "id": "ru-query-009", "utterance": "когда я в последний раз принимал витамины", "lang": "ru", "intent": "query", "tags": ["temporal"] },
|
||||
{ "id": "ru-query-010", "utterance": "есть новости по бэкапу базы", "lang": "ru", "intent": "query", "tags": ["homelab"] },
|
||||
{ "id": "ru-query-011", "utterance": "почему сервер тормозит", "lang": "ru", "intent": "query", "tags": ["homelab", "hard"], "note": "diagnostic question, not a chat opener" },
|
||||
{ "id": "ru-query-012", "utterance": "какие заметки я оставил про полив", "lang": "ru", "intent": "query", "tags": ["recall"] },
|
||||
{ "id": "ru-query-013", "utterance": "во сколько у меня встреча", "lang": "ru", "intent": "query", "tags": ["calendar"] },
|
||||
{ "id": "ru-query-019", "utterance": "что у меня стоит в календаре на послезавтра", "lang": "ru", "intent": "query", "tags": ["calendar", "hard"], "note": "agenda, not the clock: the daemon answers this from CalendarEvents inside the query branch, so the clock/date system rule must not swallow it" },
|
||||
{ "id": "ru-query-022", "utterance": "какие планы на завтра?", "lang": "ru", "intent": "query", "tags": ["calendar"], "note": "the same agenda question as ru-query-019 aimed at another day; it answered \u043f\u043e\u043a\u0430 \u043d\u0435 \u0443\u043c\u0435\u044e on the deployed daemon while the today form worked (Vikunja #471)" },
|
||||
{ "id": "ru-query-023", "utterance": "\u043a\u043e\u0433\u0434\u0430 \u043f\u043b\u0430\u043d\u0451\u0440\u043a\u0430?", "lang": "ru", "intent": "query", "tags": ["calendar", "hard"], "note": "a named event with no calendar word — the noun is the only signal that this is a question about his day" },
|
||||
{ "id": "ru-query-024", "utterance": "что дальше?", "lang": "ru", "intent": "query", "tags": ["calendar", "no-question-word"], "note": "the rest of the day, with no possessive and no plan word to anchor on; the model called it a fact and the write had to be caught downstream (Vikunja #498)" },
|
||||
{ "id": "ru-query-025", "utterance": "расскажи про битву при Ватерлоо", "lang": "ru", "intent": "query", "tags": ["world", "no-question-word"], "note": "a narrative request carries no question mark and no interrogative, so it routed fact; contrast ru-chat-003, where the same verb asks for a joke" },
|
||||
{ "id": "ru-query-014", "utterance": "я успеваю до дедлайна", "lang": "ru", "intent": "query", "tags": ["hard", "no-question-word"] },
|
||||
{ "id": "ru-query-015", "utterance": "сколько я прошёл шагов", "lang": "ru", "intent": "query", "tags": ["aggregate"] },
|
||||
{ "id": "ru-query-016", "utterance": "покажи давление за неделю", "lang": "ru", "intent": "query", "tags": ["hard", "imperative"], "note": "imperative form but a read — must not route to act" },
|
||||
{ "id": "ru-query-017", "utterance": "чем я занимался в среду", "lang": "ru", "intent": "query", "tags": ["hard", "chat-shaped"] },
|
||||
{ "id": "ru-query-018", "utterance": "хватает ли места под новые бэкапы", "lang": "ru", "intent": "query", "tags": ["homelab"] },
|
||||
{ "id": "ru-query-020", "utterance": "что дальше?", "lang": "ru", "intent": "query", "tags": ["agenda", "hard"], "note": "the rest of the day, with no interrogative the model can read as a question — it routed fact until a stage 0 rule claimed it (V-498)" },
|
||||
{ "id": "ru-query-021", "utterance": "расскажи про битву при Ватерлоо", "lang": "ru", "intent": "query", "tags": ["world", "hard"], "note": "a world question phrased as an instruction. It routed fact, and the fact gate had to catch the write (V-498)" },
|
||||
{ "id": "en-query-001", "utterance": "did I take my vitamins today", "lang": "en", "intent": "query", "tags": ["fact-shaped"] },
|
||||
{ "id": "en-query-002", "utterance": "how long since the last backup finished", "lang": "en", "intent": "query", "tags": ["temporal"] },
|
||||
{ "id": "en-query-003", "utterance": "show me this week's weight", "lang": "en", "intent": "query", "tags": ["imperative"] },
|
||||
{ "id": "ru-query-001", "utterance": "сколько воды я выпил с утра", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["aggregate"] },
|
||||
{ "id": "ru-query-002", "utterance": "я сегодня вообще пил воду", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["hard", "fact-shaped"], "note": "past-tense fact lexicon in a question — the classifier's fact centroid pulls this hard" },
|
||||
{ "id": "ru-query-003", "utterance": "во сколько я лёг вчера", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["temporal"] },
|
||||
{ "id": "ru-query-004", "utterance": "давно я не тренировался", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["hard", "no-question-word"] },
|
||||
{ "id": "ru-query-005", "utterance": "напоминания на завтра есть", "lang": "ru", "intent": "query", "want_source": "", "tags": ["hard", "reminder-shaped"], "note": "asks about reminders, does not create one. want_source is the floor on purpose: no query source reads the reminder store, and day-plan is SourceCalendar over a table this box does not write." },
|
||||
{ "id": "ru-query-006", "utterance": "что я записывал про кота", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["recall"] },
|
||||
{ "id": "ru-query-007", "utterance": "сколько раз я ел вчера", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["aggregate", "hard"] },
|
||||
{ "id": "ru-query-008", "utterance": "мой вес за последний месяц", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["no-verb"] },
|
||||
{ "id": "ru-query-009", "utterance": "когда я в последний раз принимал витамины", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["temporal"] },
|
||||
{ "id": "ru-query-010", "utterance": "есть новости по бэкапу базы", "lang": "ru", "intent": "query", "want_source": "", "tags": ["homelab"], "note": "recall, attention and network can each answer it, because mavpoll writes its netdata and uptime-kuma observations into the fact store recall reads. Naming one takes the other two off the turn." },
|
||||
{ "id": "ru-query-011", "utterance": "почему сервер тормозит", "lang": "ru", "intent": "query", "want_source": "", "tags": ["homelab", "hard"], "note": "diagnostic question, not a chat opener. network holds the box and attention holds the alarm about the box. The utterance does not choose, so neither does the label." },
|
||||
{ "id": "ru-query-012", "utterance": "какие заметки я оставил про полив", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["recall"] },
|
||||
{ "id": "ru-query-013", "utterance": "во сколько у меня встреча", "lang": "ru", "intent": "query", "want_source": "calendar", "tags": ["calendar"] },
|
||||
{ "id": "ru-query-019", "utterance": "что у меня стоит в календаре на послезавтра", "lang": "ru", "intent": "query", "want_source": "calendar", "tags": ["calendar", "hard"], "note": "agenda, not the clock: the daemon answers this from CalendarEvents inside the query branch, so the clock/date system rule must not swallow it" },
|
||||
{ "id": "ru-query-022", "utterance": "какие планы на завтра?", "lang": "ru", "intent": "query", "want_source": "calendar", "tags": ["calendar"], "note": "the same agenda question as ru-query-019 aimed at another day; it answered \u043f\u043e\u043a\u0430 \u043d\u0435 \u0443\u043c\u0435\u044e on the deployed daemon while the today form worked (Vikunja #471)" },
|
||||
{ "id": "ru-query-023", "utterance": "\u043a\u043e\u0433\u0434\u0430 \u043f\u043b\u0430\u043d\u0451\u0440\u043a\u0430?", "lang": "ru", "intent": "query", "want_source": "calendar", "tags": ["calendar", "hard"], "note": "a named event with no calendar word — the noun is the only signal that this is a question about his day" },
|
||||
{ "id": "ru-query-024", "utterance": "что дальше?", "lang": "ru", "intent": "query", "want_source": "calendar", "tags": ["calendar", "no-question-word"], "note": "the rest of the day, with no possessive and no plan word to anchor on; the model called it a fact and the write had to be caught downstream (Vikunja #498)" },
|
||||
{ "id": "ru-query-025", "utterance": "расскажи про битву при Ватерлоо", "lang": "ru", "intent": "query", "want_source": "world", "tags": ["world", "no-question-word"], "note": "a narrative request carries no question mark and no interrogative, so it routed fact; contrast ru-chat-003, where the same verb asks for a joke" },
|
||||
{ "id": "ru-query-014", "utterance": "я успеваю до дедлайна", "lang": "ru", "intent": "query", "want_source": "", "tags": ["hard", "no-question-word"], "note": "a deadline lives in the task list, the calendar or Praxis depending on where he put it. The destination depends on his data, not on his words." },
|
||||
{ "id": "ru-query-015", "utterance": "сколько я прошёл шагов", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["aggregate"] },
|
||||
{ "id": "ru-query-016", "utterance": "покажи давление за неделю", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["hard", "imperative"], "note": "imperative form but a read — must not route to act" },
|
||||
{ "id": "ru-query-017", "utterance": "чем я занимался в среду", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["hard", "chat-shaped"] },
|
||||
{ "id": "ru-query-018", "utterance": "хватает ли места под новые бэкапы", "lang": "ru", "intent": "query", "want_source": "", "tags": ["homelab"], "note": "disk headroom. network is the only source that reads the box, but the phrasing is a capacity question and not a LAN one." },
|
||||
{ "id": "ru-query-020", "utterance": "что дальше?", "lang": "ru", "intent": "query", "want_source": "calendar", "tags": ["agenda", "hard"], "note": "the rest of the day, with no interrogative the model can read as a question — it routed fact until a stage 0 rule claimed it (V-498)" },
|
||||
{ "id": "ru-query-021", "utterance": "расскажи про битву при Ватерлоо", "lang": "ru", "intent": "query", "want_source": "world", "tags": ["world", "hard"], "note": "a world question phrased as an instruction. It routed fact, and the fact gate had to catch the write (V-498)" },
|
||||
{ "id": "en-query-001", "utterance": "did I take my vitamins today", "lang": "en", "intent": "query", "want_source": "recall", "tags": ["fact-shaped"] },
|
||||
{ "id": "en-query-002", "utterance": "how long since the last backup finished", "lang": "en", "intent": "query", "want_source": "", "tags": ["temporal"], "note": "the completion time is a fact the poller wrote, so recall answers it. A person asking this wants the operational answer. Both are true." },
|
||||
{ "id": "en-query-003", "utterance": "show me this week's weight", "lang": "en", "intent": "query", "want_source": "recall", "tags": ["imperative"] },
|
||||
|
||||
{ "id": "ru-fact-001", "utterance": "только что выпил кружку воды", "lang": "ru", "intent": "fact", "want_fact_key": "water" },
|
||||
{ "id": "ru-fact-002", "utterance": "воды попил наконец", "lang": "ru", "intent": "fact", "want_fact_key": "water", "tags": ["inverted"] },
|
||||
@@ -106,6 +110,11 @@
|
||||
{ "id": "amb-005", "utterance": "потом", "lang": "ru", "want_clarify": true, "tags": ["ambiguous", "filler"] },
|
||||
{ "id": "amb-006", "utterance": "the thing from earlier", "lang": "en", "want_clarify": true, "tags": ["ambiguous", "anaphora"] },
|
||||
{ "id": "amb-007", "utterance": "напомни", "lang": "ru", "want_clarify": true, "tags": ["ambiguous", "reminder"], "note": "the reminder verb and nothing else — she knows the shape of the request and not one thing about it. Answered 'не получилось разобрать время напоминания' on the box until V-548: the subjectless-reminder gate tested Slots.Text == \"\", and fillSlots had put the verb in that slot" },
|
||||
{ "id": "amb-008", "utterance": "ну напомни же", "lang": "ru", "want_clarify": true, "tags": ["ambiguous", "reminder", "filler"], "note": "the same request wrapped in particles, which is why filler_particles is a lexicon set — without it the particles read as the subject" }
|
||||
{ "id": "amb-008", "utterance": "ну напомни же", "lang": "ru", "want_clarify": true, "tags": ["ambiguous", "reminder", "filler"], "note": "the same request wrapped in particles, which is why filler_particles is a lexicon set — without it the particles read as the subject" },
|
||||
{ "id": "ru-query-026", "utterance": "что такое TCP?", "lang": "ru", "intent": "query", "want_source": "world", "tags": ["world", "regression"], "note": "weather claimed it on 2026-08-07 and answered \"для какого города?\", because it read one percent closer than the leftover seeds. WorldQueryGrammars claims it at stage 0 now." },
|
||||
{ "id": "ru-query-027", "utterance": "сколько будет 17 на 23?", "lang": "ru", "intent": "query", "want_source": "world", "tags": ["world", "arithmetic", "regression"], "note": "same source, same day, same answer about a city. Arithmetic is not a place." },
|
||||
{ "id": "ru-query-028", "utterance": "какой у меня любимый язык?", "lang": "ru", "intent": "query", "want_source": "recall", "tags": ["recall", "possessive", "regression"], "note": "the feed answered it with kernel headlines. \"у меня\" is the whole signal and it points inward." },
|
||||
{ "id": "ru-query-029", "utterance": "кто такой Линус Торвальдс?", "lang": "ru", "intent": "query", "want_source": "world", "tags": ["world", "person", "regression"], "note": "the personal boundary answered \"не нашла у тебя такой записи\". A named public person is not his data." },
|
||||
{ "id": "ru-query-030", "utterance": "что там с бэкапами?", "lang": "ru", "intent": "query", "want_source": "", "tags": ["homelab", "regression"], "note": "search claimed it, which inverts the boundary outward. The fix is the chain order and not a destination: recall, attention and network all answer it, same as ru-query-010." }
|
||||
]
|
||||
}
|
||||
|
||||
@@ -0,0 +1,243 @@
|
||||
package router
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"math"
|
||||
"os"
|
||||
"path/filepath"
|
||||
|
||||
ort "github.com/yalue/onnxruntime_go"
|
||||
)
|
||||
|
||||
// The routing heads (V-546, V-661, V-664). Four linear heads over one masked
|
||||
// mean pool of a fine-tuned copy of multilingual-e5-small: intent,
|
||||
// destination, BIO slot tags and clarify. Trained on workpc, exported to ONNX,
|
||||
// and read here.
|
||||
//
|
||||
// Why this is not the classifier. The classifier compares one utterance to
|
||||
// frozen seed phrases by cosine. A head is a softmax over the label set, so it
|
||||
// cannot name a value that does not exist, and its max is a calibratable
|
||||
// confidence where Confidence: 1.0 was a hardcode.
|
||||
//
|
||||
// Why it is not the resident model either. It answers in single-digit
|
||||
// milliseconds against the model's p50 of 1.19s, and it names a destination
|
||||
// the classifier arm never names at all.
|
||||
//
|
||||
// The body is a COPY of the embedder weights, fine-tuned. It must never
|
||||
// replace models/embedder/multilingual-e5-small — memory recall depends on
|
||||
// that file scoring what it scored.
|
||||
//
|
||||
// The slot head is exported and deliberately not read. Slots already come from
|
||||
// the stage-2 extractor, and mapping BIO tags back to text needs character
|
||||
// offsets the unigram tokenizer does not keep. Reading it is separate work.
|
||||
const (
|
||||
// headsSeq — the sequence length the heads were trained at. Padding is
|
||||
// masked out of both attention and the pool, so this changes nothing but
|
||||
// truncation, and truncation is what training did at 64.
|
||||
headsSeq = 64
|
||||
|
||||
// headsThreshold — max softmax over the intent head, below which the heads
|
||||
// decline and the cascade carries on to the resident model.
|
||||
//
|
||||
// 0.6 is the knee measured on the 88-case intent fixture
|
||||
// (docs/evals/2026-08-08-routing-heads-in-go.md). It keeps 81 of 88 cases
|
||||
// at 97.5% accuracy. Every higher value up to 0.9 drops right answers and
|
||||
// keeps the same two wrong ones, so it buys nothing.
|
||||
headsThreshold = 0.6
|
||||
)
|
||||
|
||||
// RouterHeads runs the exported graph. Nil is a working value everywhere: a
|
||||
// deployment with no weights file routes exactly as it did before this
|
||||
// existed.
|
||||
type RouterHeads struct {
|
||||
tokenizer *unigramTokenizer
|
||||
session *ort.DynamicSession[int64, float32]
|
||||
intents []Intent
|
||||
sources []Source
|
||||
threshold float64
|
||||
}
|
||||
|
||||
// headsMeta — router_heads.json, written beside the weights by the exporter.
|
||||
// The label order is the head's output order and cannot be inferred from Go.
|
||||
type headsMeta struct {
|
||||
Intents []string `json:"intents"`
|
||||
Sources []string `json:"sources"`
|
||||
Prefix string `json:"prefix"`
|
||||
}
|
||||
|
||||
// NewRouterHeads loads the graph and its label order. modelPath points at the
|
||||
// .onnx; the external weights and router_heads.json sit beside it.
|
||||
//
|
||||
// It assumes the ONNX environment is already initialised, because the embedder
|
||||
// does that at startup and the runtime allows it once.
|
||||
func NewRouterHeads(modelPath, tokenizerPath string) (*RouterHeads, error) {
|
||||
metaPath := filepath.Join(filepath.Dir(modelPath), "router_heads.json")
|
||||
raw, err := os.ReadFile(metaPath)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("heads: read %s: %w", metaPath, err)
|
||||
}
|
||||
var meta headsMeta
|
||||
if err := json.Unmarshal(raw, &meta); err != nil {
|
||||
return nil, fmt.Errorf("heads: parse %s: %w", metaPath, err)
|
||||
}
|
||||
if meta.Prefix != queryPrefix {
|
||||
return nil, fmt.Errorf("heads: trained with prefix %q, this build uses %q",
|
||||
meta.Prefix, queryPrefix)
|
||||
}
|
||||
|
||||
intents := make([]Intent, len(meta.Intents))
|
||||
for i, s := range meta.Intents {
|
||||
intents[i] = Intent(s)
|
||||
}
|
||||
sources := make([]Source, len(meta.Sources))
|
||||
for i, s := range meta.Sources {
|
||||
// SourceUnknown is not in Sources, because it is the absence of a
|
||||
// choice. It is a class the head can emit, and the one it should emit
|
||||
// often, so it is allowed here and nowhere else.
|
||||
if s != string(SourceUnknown) && !ValidSource(Source(s)) {
|
||||
return nil, fmt.Errorf("heads: unknown destination %q in %s", s, metaPath)
|
||||
}
|
||||
sources[i] = Source(s)
|
||||
}
|
||||
|
||||
tok, err := newUnigramTokenizer(tokenizerPath)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("heads: tokenizer: %w", err)
|
||||
}
|
||||
session, err := ort.NewDynamicSession[int64, float32](
|
||||
modelPath,
|
||||
[]string{"input_ids", "attention_mask"},
|
||||
[]string{"intent", "source", "slots", "clarify"},
|
||||
)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("heads: create session: %w", err)
|
||||
}
|
||||
|
||||
return &RouterHeads{
|
||||
tokenizer: tok,
|
||||
session: session,
|
||||
intents: intents,
|
||||
sources: sources,
|
||||
threshold: headsThreshold,
|
||||
}, nil
|
||||
}
|
||||
|
||||
func (h *RouterHeads) Close() error {
|
||||
if h == nil {
|
||||
return nil
|
||||
}
|
||||
h.session.Destroy()
|
||||
return nil
|
||||
}
|
||||
|
||||
// headsResult — one forward pass, read back.
|
||||
type headsResult struct {
|
||||
Intent Intent
|
||||
Source Source
|
||||
Confidence float64
|
||||
Clarify bool
|
||||
}
|
||||
|
||||
// Route runs the heads and reports whether they are confident enough to answer.
|
||||
// A false second return is a decline, not an error: the cascade goes on to the
|
||||
// resident model, which is what happens today.
|
||||
func (h *RouterHeads) Route(ctx context.Context, utterance string) (headsResult, bool, error) {
|
||||
if h == nil {
|
||||
return headsResult{}, false, nil
|
||||
}
|
||||
ids, mask, _ := h.tokenizer.Encode(queryPrefix + utterance)
|
||||
ids, mask = ids[:headsSeq], mask[:headsSeq]
|
||||
// The tokenizer pads and truncates to its own length, which is longer than
|
||||
// this one. Cutting the tail can cut the separator with it, so put it back.
|
||||
if mask[headsSeq-1] == 1 {
|
||||
ids[headsSeq-1] = sepTokenID
|
||||
}
|
||||
|
||||
shape := ort.NewShape(1, headsSeq)
|
||||
idsT, err := ort.NewTensor(shape, ids)
|
||||
if err != nil {
|
||||
return headsResult{}, false, fmt.Errorf("heads: ids tensor: %w", err)
|
||||
}
|
||||
defer idsT.Destroy()
|
||||
maskT, err := ort.NewTensor(shape, mask)
|
||||
if err != nil {
|
||||
return headsResult{}, false, fmt.Errorf("heads: mask tensor: %w", err)
|
||||
}
|
||||
defer maskT.Destroy()
|
||||
|
||||
intentT, err := ort.NewEmptyTensor[float32](ort.NewShape(1, int64(len(h.intents))))
|
||||
if err != nil {
|
||||
return headsResult{}, false, fmt.Errorf("heads: intent tensor: %w", err)
|
||||
}
|
||||
defer intentT.Destroy()
|
||||
sourceT, err := ort.NewEmptyTensor[float32](ort.NewShape(1, int64(len(h.sources))))
|
||||
if err != nil {
|
||||
return headsResult{}, false, fmt.Errorf("heads: source tensor: %w", err)
|
||||
}
|
||||
defer sourceT.Destroy()
|
||||
slotsT, err := ort.NewEmptyTensor[float32](ort.NewShape(1, headsSeq, int64(numBIOTags)))
|
||||
if err != nil {
|
||||
return headsResult{}, false, fmt.Errorf("heads: slots tensor: %w", err)
|
||||
}
|
||||
defer slotsT.Destroy()
|
||||
clarifyT, err := ort.NewEmptyTensor[float32](ort.NewShape(1, 2))
|
||||
if err != nil {
|
||||
return headsResult{}, false, fmt.Errorf("heads: clarify tensor: %w", err)
|
||||
}
|
||||
defer clarifyT.Destroy()
|
||||
|
||||
if err := h.session.Run(
|
||||
[]*ort.Tensor[int64]{idsT, maskT},
|
||||
[]*ort.Tensor[float32]{intentT, sourceT, slotsT, clarifyT},
|
||||
); err != nil {
|
||||
return headsResult{}, false, fmt.Errorf("heads: run: %w", err)
|
||||
}
|
||||
|
||||
// The graph applies its own softmax, so these are probabilities and the max
|
||||
// is the same number the eval calibrated the threshold against.
|
||||
i, conf := argmax(intentT.GetData())
|
||||
res := headsResult{
|
||||
Intent: h.intents[i],
|
||||
Confidence: conf,
|
||||
}
|
||||
cl := clarifyT.GetData()
|
||||
res.Clarify = len(cl) == 2 && cl[1] > cl[0]
|
||||
|
||||
// The destination head is trained on query rows and is meaningless on any
|
||||
// other intent, the same way queryWalk is never reached by one.
|
||||
if res.Intent == IntentQuery {
|
||||
s, _ := argmax(sourceT.GetData())
|
||||
res.Source = h.sources[s]
|
||||
}
|
||||
|
||||
// The clarify head decides on its own, and it decides first. It answers a
|
||||
// different question from the intent head — not which intent, but whether
|
||||
// there is enough here to act on at all — so a low intent confidence is no
|
||||
// reason to discard it. It is usually the same turns: "вода" reads as
|
||||
// intent act at 0.23 and clarify at 0.98, and letting the intent threshold
|
||||
// bury that hands the turn to the classifier, which routes it confidently
|
||||
// and never asks.
|
||||
if res.Clarify {
|
||||
return res, true, nil
|
||||
}
|
||||
if conf < h.threshold {
|
||||
return res, false, nil
|
||||
}
|
||||
return res, true, nil
|
||||
}
|
||||
|
||||
// numBIOTags — O plus B- and I- for each of Maven's five slots. The head is not
|
||||
// read, but the graph writes it and the output tensor has to be the right size.
|
||||
const numBIOTags = 11
|
||||
|
||||
func argmax(v []float32) (int, float64) {
|
||||
best, bestV := 0, math.Inf(-1)
|
||||
for i, x := range v {
|
||||
if float64(x) > bestV {
|
||||
best, bestV = i, float64(x)
|
||||
}
|
||||
}
|
||||
return best, bestV
|
||||
}
|
||||
@@ -105,6 +105,27 @@ type Decision struct {
|
||||
Slots Slots
|
||||
Clarify bool // stage 3: below threshold — ask, don't guess
|
||||
|
||||
// Source — where the answer lives, for a query. The second half of the
|
||||
// route, and empty on every other intent. SourceUnknown means no decider
|
||||
// named one and the daemon walks its whole chain, which is what shipped
|
||||
// before this field existed. See source.go for why it is twelve values.
|
||||
Source Source
|
||||
|
||||
// SourceAnchored — a stage 0 grammar named that destination, matching a
|
||||
// literal pattern to do it. Only the router sets this, and only there.
|
||||
//
|
||||
// It exists because one thing downstream is not reversible by evidence
|
||||
// (V-666). Naming a destination normally takes guessing sources off a turn,
|
||||
// and one of those is the personal boundary, which is what stops a question
|
||||
// about him from reaching the world. A grammar that read "что такое X" may
|
||||
// take it off. A model or a softmax may not, because a wrong destination
|
||||
// there widens what leaves the box rather than costing an answer.
|
||||
//
|
||||
// Read Stage instead and the two decisions get coupled: stage 0 also means
|
||||
// confidence 1.0 and an anchored claim band, and a later cascade change
|
||||
// could make one true where the other is not.
|
||||
SourceAnchored bool
|
||||
|
||||
// Continued — this decision was rebuilt from the previous turn rather
|
||||
// than routed, because the utterance was an ellipsis ("а завтра?").
|
||||
// Handlers use it to know that Slots.Text is the PREVIOUS turn's topic
|
||||
|
||||
@@ -45,12 +45,19 @@ const routeGrammar = `
|
||||
root ::= "[" ws action ("," ws action)* ws "]"
|
||||
action ::= "{" ws "\"intent\"" ws ":" ws intent ("," ws field)* ws "}"
|
||||
intent ::= "\"fact\"" | "\"reminder\"" | "\"note\"" | "\"query\"" | "\"act\"" | "\"chat\"" | "\"system\"" | "\"unknown\""
|
||||
field ::= key ws ":" ws string
|
||||
field ::= (key ws ":" ws string) | ("\"source\"" ws ":" ws source)
|
||||
key ::= "\"key\"" | "\"value\"" | "\"text\"" | "\"verb\""
|
||||
source ::= "\"recall\"" | "\"calendar\"" | "\"tasks\"" | "\"list\"" | "\"money\"" | "\"weather\"" | "\"home\"" | "\"network\"" | "\"feeds\"" | "\"attention\"" | "\"self\"" | "\"world\"" | "\"\""
|
||||
string ::= "\"" ([^"\\\x00-\x1F] | "\\" ["\\/bfnrt] | "\\u" [0-9a-fA-F]{4}){0,120} "\""
|
||||
ws ::= [ \t\n]{0,4}
|
||||
`
|
||||
|
||||
// TestRouteGrammarCoversSources holds the source rule above to router.Sources.
|
||||
// The enum is the point: a grammar cannot emit a destination that does not
|
||||
// exist, which is the guarantee V-546 wants from a softmax and gets here for
|
||||
// free. Empty is the thirteenth alternative and it is not an oversight — it is
|
||||
// the SourceUnknown floor, and the model must be able to decline.
|
||||
|
||||
// routeSystem — the router prompt. Changed 31-07-2026: the query test now sits
|
||||
// above the fact test and there is an explicit question test. Before that, a
|
||||
// question naming a fact key ("сколько воды я выпил с утра") matched the fact
|
||||
@@ -121,6 +128,31 @@ const routeSystem = `Классифицируй ровно одно сообще
|
||||
"что такое кватернион?" → {"intent":"query","text":"что такое кватернион"}
|
||||
"ага" → {"intent":"chat","text":"ага"}
|
||||
|
||||
Только для query добавь поле source — где лежит ответ:
|
||||
- recall — его заметки, факты и то, что он раньше говорил
|
||||
- calendar — встречи и события
|
||||
- tasks — список задач
|
||||
- list — списки покупок и другие именованные списки
|
||||
- money — траты
|
||||
- weather — погода
|
||||
- home — свет, устройства, дом
|
||||
- network — локальная сеть, сервер, диски
|
||||
- feeds — новостные ленты
|
||||
- attention — что требует внимания сейчас
|
||||
- self — вопрос про самого ассистента
|
||||
- world — всё остальное: определения, счёт, люди, факты о мире
|
||||
|
||||
Пустое значение "" — нормальный ответ и его надо ставить часто. Ставь "", если ответ могут дать сразу два источника или если не уверен: тогда проверяются все по порядку, и это правильно. Никогда не угадывай.
|
||||
|
||||
"сколько воды я выпил с утра" → {"intent":"query","text":"сколько воды я выпил с утра","source":"recall"}
|
||||
"что я записывал про кота" → {"intent":"query","text":"что я записывал про кота","source":"recall"}
|
||||
"во сколько у меня встреча" → {"intent":"query","text":"во сколько у меня встреча","source":"calendar"}
|
||||
"что такое docker?" → {"intent":"query","text":"что такое docker","source":"world"}
|
||||
"кто такой Линус Торвальдс?" → {"intent":"query","text":"кто такой Линус Торвальдс","source":"world"}
|
||||
"сколько будет 17 на 23?" → {"intent":"query","text":"сколько будет 17 на 23","source":"world"}
|
||||
"почему сервер тормозит" → {"intent":"query","text":"почему сервер тормозит","source":""}
|
||||
"есть новости по бэкапу базы" → {"intent":"query","text":"есть новости по бэкапу базы","source":""}
|
||||
|
||||
Ответ — JSON-массив: по одному объекту на каждую просьбу. Обычно один. Если в реплике несколько просьб — по объекту на каждую. "напомни купить молоко, и запиши что кофе кончился" → [{"intent":"reminder","text":"купить молоко"},{"intent":"note","text":"кофе кончился"}]. Только JSON, без пояснений.`
|
||||
|
||||
// routeRepeatPenalty — the sub-1B model loops one sentence inside the text field
|
||||
@@ -172,6 +204,7 @@ type routeAction struct {
|
||||
Value string `json:"value"`
|
||||
Text string `json:"text"`
|
||||
Verb string `json:"verb"`
|
||||
Source string `json:"source"`
|
||||
}
|
||||
|
||||
// Route asks the model for one decision. The bool is false when there is no
|
||||
@@ -240,6 +273,14 @@ func (lr *LLMRouter) Route(ctx context.Context, utterance string, now time.Time)
|
||||
case IntentQuery:
|
||||
d.Intent = IntentQuery
|
||||
d.Slots.Text = firstNonEmpty(a.Text, utterance)
|
||||
// Through ValidSource, and on query alone. The grammar already bounds
|
||||
// the enum, but the grammar is a request to a server that may be
|
||||
// running a different build, and a destination this binary does not
|
||||
// know would take real query sources off the turn. Anything unknown
|
||||
// drops to SourceUnknown, which is the floor and costs nothing.
|
||||
if ValidSource(Source(a.Source)) {
|
||||
d.Source = Source(a.Source)
|
||||
}
|
||||
case IntentAct:
|
||||
d.Intent = IntentAct
|
||||
d.Slots.Text = firstNonEmpty(a.Verb, utterance)
|
||||
|
||||
@@ -398,3 +398,70 @@ func TestLLMReminderWithSubjectIsNotGated(t *testing.T) {
|
||||
t.Fatalf("a complete reminder was sent back as a question: %+v", d.Slots)
|
||||
}
|
||||
}
|
||||
|
||||
// TestRouteGrammarCoversSources — the grammar enum and router.Sources are two
|
||||
// hand-written lists of the same twelve destinations, and nothing else notices
|
||||
// when one grows. A destination missing from the grammar is a destination the
|
||||
// model is structurally unable to name, which is the exact defect V-517
|
||||
// measured for Praxis: not a weak model, an absent string.
|
||||
func TestRouteGrammarCoversSources(t *testing.T) {
|
||||
for _, s := range Sources {
|
||||
if !strings.Contains(routeGrammar, `"\"`+string(s)+`\""`) {
|
||||
t.Errorf("routeGrammar cannot emit %q — the model can never name it", s)
|
||||
}
|
||||
}
|
||||
// The floor has to be reachable too, or the model is forced to pick one.
|
||||
if !strings.Contains(routeGrammar, `"\"\""`) {
|
||||
t.Error(`routeGrammar cannot emit "" — the model cannot decline a destination`)
|
||||
}
|
||||
// Count the alternatives on the source rule: an extra one is a destination
|
||||
// the daemon would drop to SourceUnknown after the model spent tokens on it.
|
||||
for _, line := range strings.Split(routeGrammar, "\n") {
|
||||
if !strings.HasPrefix(line, "source ") {
|
||||
continue
|
||||
}
|
||||
if got, want := strings.Count(line, "|")+1, len(Sources)+1; got != want {
|
||||
t.Errorf("source rule has %d alternatives, want %d (Sources plus the floor)", got, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// The destination is read back only through ValidSource. A model on an older or
|
||||
// newer build can write a string this binary does not know, and trusting it
|
||||
// would take real query sources off the turn for a name nothing answers.
|
||||
func TestLLMUnknownSourceFallsToTheFloor(t *testing.T) {
|
||||
r := newLLMTestRouter(t, `{"intent":"query","text":"что там с бэкапами","source":"praxis"}`)
|
||||
d, err := r.Route(context.Background(), "что там с бэкапами", refNow())
|
||||
if err != nil {
|
||||
t.Fatalf("route: %v", err)
|
||||
}
|
||||
if d.Source != SourceUnknown {
|
||||
t.Fatalf("invented destination %q was trusted, want the floor", d.Source)
|
||||
}
|
||||
}
|
||||
|
||||
// And a known one survives, or the read-back is just a filter.
|
||||
func TestLLMNamedSourceSurvives(t *testing.T) {
|
||||
r := newLLMTestRouter(t, `{"intent":"query","text":"кто такой Линус Торвальдс","source":"world"}`)
|
||||
d, err := r.Route(context.Background(), "кто такой Линус Торвальдс?", refNow())
|
||||
if err != nil {
|
||||
t.Fatalf("route: %v", err)
|
||||
}
|
||||
if d.Source != SourceWorld {
|
||||
t.Fatalf("source %q, want %q", d.Source, SourceWorld)
|
||||
}
|
||||
}
|
||||
|
||||
// A destination on anything but a query is dropped. Only IntentQuery reaches
|
||||
// queryWalk, so a source elsewhere is a field nobody reads and a claim nobody
|
||||
// checks.
|
||||
func TestLLMSourceIsQueryOnly(t *testing.T) {
|
||||
r := newLLMTestRouter(t, `{"intent":"note","text":"кофе кончился","source":"recall"}`)
|
||||
d, err := r.Route(context.Background(), "запиши что кофе кончился", refNow())
|
||||
if err != nil {
|
||||
t.Fatalf("route: %v", err)
|
||||
}
|
||||
if d.Source != SourceUnknown {
|
||||
t.Fatalf("a note carried destination %q", d.Source)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -66,12 +66,19 @@ func NewONNXEmbedder(modelPath, tokenizerPath, libPath string) (*onnxEmbedder, e
|
||||
func (e *onnxEmbedder) Dim() int { return embedDim }
|
||||
|
||||
// ID names the loaded model for the DB marker (Vikunja #378): the model file's
|
||||
// own name plus the dimension, so pointing the config at another model changes
|
||||
// the string on its own.
|
||||
// own name, the dimension, and the tokenizer revision, so pointing the config
|
||||
// at another model changes the string on its own.
|
||||
func (e *onnxEmbedder) ID() string { return e.id }
|
||||
|
||||
// tokenizerRev — bumped whenever the tokenizer changes what it emits for the
|
||||
// same text, because that changes every vector while the model file's name
|
||||
// stays put. Rev 2 is the fix for the reversed word pieces (V-664): stored
|
||||
// passages embedded under rev 1 no longer sit in the same space as a query
|
||||
// embedded now, and ReembedAll rewrites them because this string moved.
|
||||
const tokenizerRev = 2
|
||||
|
||||
// modelIDFromPath turns /opt/.../multilingual-e5-small.onnx into
|
||||
// "multilingual-e5-small@384".
|
||||
// "multilingual-e5-small@384/tok2".
|
||||
func modelIDFromPath(modelPath string) string {
|
||||
name := modelPath
|
||||
if i := strings.LastIndexAny(name, "/\\"); i >= 0 {
|
||||
@@ -81,7 +88,7 @@ func modelIDFromPath(modelPath string) string {
|
||||
if name == "" {
|
||||
name = "onnx"
|
||||
}
|
||||
return fmt.Sprintf("%s@%d", name, embedDim)
|
||||
return fmt.Sprintf("%s@%d/tok%d", name, embedDim, tokenizerRev)
|
||||
}
|
||||
|
||||
// Embed treats the text as a query. The classifier compares one short
|
||||
@@ -339,14 +346,17 @@ func (t *unigramTokenizer) encodeWord(word string) []int64 {
|
||||
}
|
||||
}
|
||||
|
||||
// Backtracking walks the word from its end, and prepending each piece puts
|
||||
// it back in reading order. There used to be a second reverse after this
|
||||
// loop, which undid it: every multi-piece word came out backwards, and
|
||||
// "query: вода" tokenized to [0 12 1294 41 12489 2] where the reference
|
||||
// tokenizer gives [0 41 1294 12 12489 2] (V-664). A transformer reads
|
||||
// position, so the pieces of a long Russian word were being read in the
|
||||
// wrong order on every turn.
|
||||
var result []int64
|
||||
for i := n; i > 0; i = prev[i] {
|
||||
result = append([]int64{bestID[i]}, result...)
|
||||
}
|
||||
// Reverse
|
||||
for l, r := 0, len(result)-1; l < r; l, r = l+1, r-1 {
|
||||
result[l], result[r] = result[r], result[l]
|
||||
}
|
||||
return result
|
||||
}
|
||||
|
||||
|
||||
@@ -31,6 +31,12 @@ type Config struct {
|
||||
// error/parse failure, falls through to the classifier (never fails the
|
||||
// turn on the model).
|
||||
LLM *LLMRouter
|
||||
// Heads — optional routing heads over the fine-tuned embedder copy. When
|
||||
// set, Route consults them after stage 0 and before the LLM router. They
|
||||
// decline below their own confidence threshold, so a low-confidence turn
|
||||
// reaches the model exactly as it does today. Nil is the shipped-before
|
||||
// behaviour and costs nothing.
|
||||
Heads *RouterHeads
|
||||
}
|
||||
|
||||
// Router — the deterministic cascade. Route never guesses: stage 0 wins
|
||||
@@ -42,6 +48,7 @@ type Router struct {
|
||||
extractor Extractor
|
||||
threshold float64
|
||||
llm *LLMRouter
|
||||
heads *RouterHeads
|
||||
}
|
||||
|
||||
func New(cfg Config) *Router {
|
||||
@@ -51,6 +58,7 @@ func New(cfg Config) *Router {
|
||||
extractor: cfg.Extractor,
|
||||
threshold: cfg.Threshold,
|
||||
llm: cfg.LLM,
|
||||
heads: cfg.Heads,
|
||||
}
|
||||
}
|
||||
|
||||
@@ -90,6 +98,10 @@ func (r *Router) Route(ctx context.Context, utterance string, now time.Time) (De
|
||||
continue // grammar matched shape but not content → fall through
|
||||
}
|
||||
d.Utterance = utterance
|
||||
// A literal pattern named that destination, which is the one provenance
|
||||
// allowed to take the personal boundary off a turn (V-666). Set here and
|
||||
// nowhere else, so no other arm of the cascade can claim it.
|
||||
d.SourceAnchored = d.Source != SourceUnknown
|
||||
// The grammar decided the intent; the extractor fills the slots it did
|
||||
// not match (V-572). See fillMatchedSlots for why every grammar gets it.
|
||||
r.fillMatchedSlots(ctx, &d, now)
|
||||
@@ -98,6 +110,60 @@ func (r *Router) Route(ctx context.Context, utterance string, now time.Time) (De
|
||||
}
|
||||
r.noteGrammarOutcomes(ctx, len(r.grammars), declinedBuild, "", "")
|
||||
|
||||
// stage 0b — routing heads (when wired). A softmax over the label set, so
|
||||
// it cannot name an intent or a destination that does not exist, and its
|
||||
// max is a real confidence. It runs before the model because it is three
|
||||
// orders of magnitude faster and scores better on both halves of the route.
|
||||
//
|
||||
// It declines below its threshold rather than clarifying. A declined turn
|
||||
// carries on to the model and then the classifier, which is what a box with
|
||||
// no weights file does on every turn.
|
||||
if r.heads != nil {
|
||||
res, ok, err := r.heads.Route(ctx, utterance)
|
||||
switch {
|
||||
case err != nil:
|
||||
log.Printf("router: heads fell through to the rest of the cascade: %v", err)
|
||||
decision.Note(ctx, decision.Claim{
|
||||
Stage: decision.StageRoute, Claimant: claimantHeads,
|
||||
Outcome: decision.Declined, Reason: "error: " + err.Error(),
|
||||
})
|
||||
case !ok:
|
||||
decision.Note(ctx, decision.Scored(decision.StageRoute, claimantHeads,
|
||||
string(res.Intent), res.Confidence, decision.Declined,
|
||||
"below the heads confidence threshold"))
|
||||
default:
|
||||
d := Decision{
|
||||
Utterance: utterance,
|
||||
Stage: 2,
|
||||
Intent: res.Intent,
|
||||
Confidence: res.Confidence,
|
||||
Source: res.Source,
|
||||
Clarify: res.Clarify,
|
||||
}
|
||||
r.fillSlots(ctx, &d, now)
|
||||
decision.Note(ctx, decision.Claim{
|
||||
Stage: decision.StageRoute, Claimant: claimantLLM,
|
||||
Outcome: decision.NeverAsked, Reason: "the routing heads answered",
|
||||
})
|
||||
decision.Note(ctx, decision.Claim{
|
||||
Stage: decision.StageRoute, Claimant: claimantClassifier,
|
||||
Outcome: decision.NeverAsked, Reason: "the routing heads answered",
|
||||
})
|
||||
outcome, reason := decision.Won, ""
|
||||
if d.Clarify {
|
||||
outcome, reason = decision.Thinned, "the clarify head says there is too little here to act on"
|
||||
}
|
||||
decision.Note(ctx, decision.Scored(decision.StageRoute, claimantHeads,
|
||||
string(d.Intent), d.Confidence, outcome, reason))
|
||||
return d, nil
|
||||
}
|
||||
} else {
|
||||
decision.Note(ctx, decision.Claim{
|
||||
Stage: decision.StageRoute, Claimant: claimantHeads,
|
||||
Outcome: decision.NeverAsked, Reason: "no routing heads are wired",
|
||||
})
|
||||
}
|
||||
|
||||
// stage 1a — LLM router (when wired). It reasons over the utterance instead
|
||||
// of nearest-centroid guessing. On any error/parse-fail, fall through to the
|
||||
// classifier cascade (never fail the turn on the model).
|
||||
|
||||
@@ -0,0 +1,78 @@
|
||||
package router
|
||||
|
||||
// Source — where the answer to a query lives. It is the second half of a
|
||||
// routing decision and it used to be made outside the router entirely (V-655).
|
||||
//
|
||||
// The cascade sorted an utterance into one of seven intents with stage 0 rules,
|
||||
// the resident model and the classifier behind it, a fixture measuring it and
|
||||
// the decision trace recording it. Then IntentQuery handed the turn to
|
||||
// querySources in the daemon, a chain of twenty-two branches deciding by seed
|
||||
// similarity in a fixed order, with none of that. So the careful sorter did the
|
||||
// easy half and the sloppy one did the hard half: on 2026-08-07 weather claimed
|
||||
// "что такое TCP?" and answered "для какого города?", because weather read one
|
||||
// percent closer to the turn than the pile of leftover seeds did, and one
|
||||
// percent was enough. Search would have answered it and search was never asked.
|
||||
//
|
||||
// "query" is not a destination. It is a shrug. This is the field that says
|
||||
// where to look.
|
||||
//
|
||||
// # Why twelve and not twenty-two
|
||||
//
|
||||
// A destination is what a decider can plausibly name from the utterance alone,
|
||||
// not one entry per source. Three of the daemon's sources are successive passes
|
||||
// over his own words and a fourth reads the facts by key: which of them lands
|
||||
// the hit is an ordering detail inside the chain, and no utterance says. They
|
||||
// are SourceRecall together. The same goes for the metasearch, the offline
|
||||
// encyclopedia and a page he named by URL, which are SourceWorld.
|
||||
//
|
||||
// # Empty is a real value and it is the floor
|
||||
//
|
||||
// SourceUnknown means nobody decided. The daemon then walks the whole chain in
|
||||
// its original order, which is the behaviour that shipped before this field
|
||||
// existed. So the classifier arm sets nothing and costs nothing, and a box
|
||||
// whose model is down routes queries exactly as it did.
|
||||
type Source string
|
||||
|
||||
const (
|
||||
// SourceUnknown — no decider named a destination. Walk the chain.
|
||||
SourceUnknown Source = ""
|
||||
|
||||
// His own data.
|
||||
SourceRecall Source = "recall" // notes, facts and what he has said before
|
||||
SourceCalendar Source = "calendar" // events, and the only date-aware destination
|
||||
SourceTasks Source = "tasks" // the task list
|
||||
SourceList Source = "list" // the shopping and other named lists
|
||||
SourceMoney Source = "money" // the spending facts the poller writes
|
||||
|
||||
// The surroundings.
|
||||
SourceWeather Source = "weather" // the forecast for a place
|
||||
SourceHome Source = "home" // lights, devices, the house
|
||||
SourceNetwork Source = "network" // the LAN and what is on it
|
||||
SourceFeeds Source = "feeds" // the RSS she reads
|
||||
SourceAttention Source = "attention" // what Praxis says needs looking at
|
||||
|
||||
// Everything else.
|
||||
SourceSelf Source = "self" // a question about Maven herself
|
||||
SourceWorld Source = "world" // search, the ZIMs, a page he named
|
||||
)
|
||||
|
||||
// Sources — every destination a decider may name, in a fixed order so a prompt,
|
||||
// a grammar table and a test all read the same list. SourceUnknown is not a
|
||||
// member: it is the absence of a choice, not one of the choices.
|
||||
var Sources = []Source{
|
||||
SourceRecall, SourceCalendar, SourceTasks, SourceList, SourceMoney,
|
||||
SourceWeather, SourceHome, SourceNetwork, SourceFeeds, SourceAttention,
|
||||
SourceSelf, SourceWorld,
|
||||
}
|
||||
|
||||
// ValidSource reports whether s is one a decider may name. Anything else,
|
||||
// including a destination invented by a model, is dropped back to
|
||||
// SourceUnknown by the caller rather than trusted.
|
||||
func ValidSource(s Source) bool {
|
||||
for _, known := range Sources {
|
||||
if s == known {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
@@ -187,9 +187,11 @@ func SystemTimeDateGrammars() []Grammar {
|
||||
// written ("the clock/date system rule must not swallow it"); the daemon
|
||||
// disagreed with the fixture and the daemon was wrong.
|
||||
//
|
||||
// Routing, not answering. These set the intent and nothing else — which source
|
||||
// in the query chain claims the turn stays the chain's decision, and a
|
||||
// question with no date still falls through queryCalendar to recall.
|
||||
// Routing, not answering. Two of the five also name the calendar as the
|
||||
// destination (V-655), which narrows who may GUESS their way onto the turn and
|
||||
// claims nothing. Every source that looks something up still runs, in the order
|
||||
// it always did, so a question with no date still falls through queryCalendar
|
||||
// to recall.
|
||||
//
|
||||
// Deliberately not folded into SystemTimeDateGrammars: those exist to send
|
||||
// utterances TO system, these exist to keep utterances OUT of it, and one
|
||||
@@ -199,9 +201,15 @@ func AgendaQueryGrammars() []Grammar {
|
||||
{
|
||||
// An explicit calendar noun is unambiguous wherever it appears:
|
||||
// "что в календаре на завтра", "покажи расписание на среду".
|
||||
//
|
||||
// The one agenda rule that names its destination, because an
|
||||
// explicit calendar noun leaves nothing to weigh (V-655). The
|
||||
// possessive rules below deliberately do not: "что у меня в списке
|
||||
// покупок" matches agenda-query, and naming the calendar there
|
||||
// would take the list source off the turn.
|
||||
Name: "calendar-query",
|
||||
Pattern: regexp.MustCompile(`(?i)(календар|расписани|повестк)`),
|
||||
Build: agendaQueryBuild,
|
||||
Build: queryTo(SourceCalendar),
|
||||
},
|
||||
{
|
||||
// The agenda phrasing with no calendar noun. Anchored at the start
|
||||
@@ -251,9 +259,11 @@ func AgendaQueryGrammars() []Grammar {
|
||||
// "во сколько созвон". He is asking when something on his calendar
|
||||
// happens, and the noun is the only signal. Closed list, so "когда
|
||||
// битва при Ватерлоо" is still a world question.
|
||||
// Names the calendar (V-655): the noun list is closed and every
|
||||
// member of it is an event, so there is nothing else to weigh.
|
||||
Name: "event-time-query",
|
||||
Pattern: regexp.MustCompile(`(?i)^\s*(когда|во\s+сколько|в\s+котором\s+часу)\s+(будет\s+|у\s+нас\s+)?(планёрк|планерк|встреч|созвон|митинг|совещани|звонок|созвон|приём|прием|интервью|собеседовани|тренировк|урок|занятие|пара)[а-я]*(\s|[?!.]|$)`),
|
||||
Build: agendaQueryBuild,
|
||||
Build: queryTo(SourceCalendar),
|
||||
},
|
||||
}
|
||||
}
|
||||
@@ -340,6 +350,11 @@ func narrativeQueryBuild(m []string) (Decision, bool) {
|
||||
Intent: IntentQuery,
|
||||
Confidence: 1.0,
|
||||
Slots: Slots{Text: topic},
|
||||
// The world, because that is the shape this asks for and the rule has
|
||||
// already declined the two cases where it is not: entertainment, and
|
||||
// questions about her (V-655). His own notes are still read first — a
|
||||
// destination narrows who may guess and reorders nothing.
|
||||
Source: SourceWorld,
|
||||
}, true
|
||||
}
|
||||
|
||||
|
||||
@@ -0,0 +1,86 @@
|
||||
package router
|
||||
|
||||
import "regexp"
|
||||
|
||||
// WorldQueryGrammars — stage-0 rules for the two question shapes that name the
|
||||
// world in their own words, and say so plainly enough that no scorer is needed
|
||||
// (V-655).
|
||||
//
|
||||
// They exist because of what happens when nothing deterministic claims these.
|
||||
// Measured on the box on 2026-08-07 (docs/evals/2026-08-07-week-of-usage.md,
|
||||
// section 4): "что такое TCP?" and "сколько будет 17 на 23?" were both answered
|
||||
// "для какого города?", and "кто такой Линус Торвальдс?" was answered "не знаю —
|
||||
// не нашла у тебя такой записи". None of those three is about him, about the
|
||||
// weather, or about anything on this box.
|
||||
//
|
||||
// The mechanism is the destination, not the answer. Naming SourceWorld does not
|
||||
// send the turn outside and does not skip a single source that looks something
|
||||
// up: his notes, his facts and the personal boundary all still run first, in the
|
||||
// order they always did. What it does is stop the sources that claim on seed
|
||||
// similarity from taking the turn on the way past. Weather cannot claim a
|
||||
// question about a protocol once the utterance has said which side it is on.
|
||||
//
|
||||
// Both patterns are spelled out here rather than drawn from internal/lexicon,
|
||||
// which is the same call the agenda rules made: these are interrogative FRAMES
|
||||
// of two words, not a closed class of single words, and the lexicon holds
|
||||
// classes. Nothing here is a stem pattern over open vocabulary — the variable
|
||||
// part of each rule is the topic, and the rule reads none of it.
|
||||
func WorldQueryGrammars() []Grammar {
|
||||
return []Grammar{
|
||||
{
|
||||
// "что такое X", "кто такой X". A request for what a thing or a
|
||||
// person IS, which his own data can answer and usually cannot.
|
||||
//
|
||||
// The topic is deliberately not captured into Slots.Text. Every
|
||||
// source below reads the utterance, "что такое TCP?" is already the
|
||||
// best query string for it, and the agenda rules make the same call
|
||||
// for the same reason.
|
||||
Name: "definition-query",
|
||||
Pattern: definitionQueryPattern,
|
||||
Build: queryTo(SourceWorld),
|
||||
},
|
||||
{
|
||||
// "сколько будет 17 на 23", "сколько будет 2+2". Arithmetic, which
|
||||
// the metasearch answers and no local source holds. The digits are
|
||||
// what make it arithmetic: "сколько будет гостей" names no number
|
||||
// and is a question about his evening.
|
||||
Name: "arithmetic-query",
|
||||
Pattern: arithmeticQueryPattern,
|
||||
Build: queryTo(SourceWorld),
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
// definitionQueryPattern — anchored at the start, because "напомни узнать что
|
||||
// такое TCP" is a reminder that happens to contain the frame.
|
||||
//
|
||||
// (\s|[?!.]|$) and not \b: Go's \b is ASCII-only and never fires after a
|
||||
// Cyrillic letter, so the ASCII form silently matches nothing. The agenda rules
|
||||
// carry the same note.
|
||||
var definitionQueryPattern = regexp.MustCompile(
|
||||
`(?i)^\s*(что\s+так(ое|ая)|кто\s+так(ой|ая|ие)|what\s+is|who\s+is)(\s|[?!.]|$)`)
|
||||
|
||||
// arithmeticQueryPattern — the ask, then a digit somewhere after it. Loose on
|
||||
// what sits between them on purpose: the operator is spoken half a dozen ways
|
||||
// ("на", "умножить на", "плюс", "+") and reading them is the calculator's job,
|
||||
// not this rule's. All this decides is which side of the boundary the turn is
|
||||
// on.
|
||||
var arithmeticQueryPattern = regexp.MustCompile(
|
||||
`(?i)^\s*(сколько\s+будет|посчитай|вычисли|how\s+much\s+is)\s.*\d`)
|
||||
|
||||
// queryTo builds a stage-0 query Decision that names where the answer lives.
|
||||
//
|
||||
// The utterance travels intact and no slot is filled, which is the same
|
||||
// contract agendaQueryBuild has: confidence 1.0 on the intent and the
|
||||
// destination, and every source below still decides for itself whether it has
|
||||
// an answer. Naming a destination narrows who may guess. It promises nothing.
|
||||
func queryTo(dest Source) func([]string) (Decision, bool) {
|
||||
return func([]string) (Decision, bool) {
|
||||
return Decision{
|
||||
Stage: 0,
|
||||
Intent: IntentQuery,
|
||||
Confidence: 1.0,
|
||||
Source: dest,
|
||||
}, true
|
||||
}
|
||||
}
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user