Compare commits
81 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 0ca5748699 | |||
| b43bb265b5 | |||
| b9a24334ea | |||
| c97aebf55a | |||
| 891136c65d | |||
| 41c7c13f42 | |||
| a324e8f624 | |||
| 51805e7f35 | |||
| 533f0acda8 | |||
| 4f59ba78c6 | |||
| d0afd9d4f6 | |||
| 6b67e6f3c2 | |||
| 742b2ad1d7 | |||
| b300ac5c70 | |||
| 13e5170e9e | |||
| 0b90952e55 | |||
| aa8f5b2ee2 | |||
| d7cdcb63bd | |||
| ddb658ffbb | |||
| c7dadc97d9 | |||
| c9d88c152e | |||
| 1890ff5d5d | |||
| 0110e9bc8c | |||
| 50ca8c8b5a | |||
| de09471421 | |||
| d65c16a567 | |||
| 062d4252ef | |||
| 2c27e2ce1f | |||
| ccc5cba2a3 | |||
| 89d83c0b11 | |||
| a97f554802 | |||
| f4de2fc5e1 | |||
| eef5d4da4f | |||
| 09f1696fce | |||
| 80f7322294 | |||
| fa5aebfbe4 | |||
| 59cec63da1 | |||
| 02e8786695 | |||
| 0272dc9d89 | |||
| 2ad7635501 | |||
| 9949b309b1 | |||
| a788ca3915 | |||
| 62d47d28ac | |||
| e9ff2c4912 | |||
| 3dbf67f8f9 | |||
| 84ba217892 | |||
| f179ae2fde | |||
| d00929ac0b | |||
| b6f47fbeb6 | |||
| 2e9b9ec1cf | |||
| bfb57c3148 | |||
| d1f6f6355f | |||
| 92ecb691de | |||
| 4282f6b9a9 | |||
| 7bb9f9be06 | |||
| 1e47eaca5a | |||
| 892330eb84 | |||
| 9a3bcd7c46 | |||
| 98ee701e03 | |||
| 04c1088088 | |||
| 07c191d8b8 | |||
| c668310b3e | |||
| 1bd2acdc2a | |||
| 15e5dd8eaa | |||
| 10cf6f525c | |||
| e2210f6844 | |||
| 214a4032cf | |||
| dc70a5a7ab | |||
| 06aded6ab0 | |||
| d2be98ee2a | |||
| 62d320f93a | |||
| 796e6af3cf | |||
| 74a70880a8 | |||
| a40bc559d5 | |||
| 1db0fcfcd0 | |||
| 4ca68d2f3f | |||
| ee3e6a9eaf | |||
| 8acb8a97c6 | |||
| 0b3b8d0a9e | |||
| fe0e654ab1 | |||
| a54ebac0cb |
@@ -7,9 +7,20 @@ talking over unix sockets; one resident small model for routing + phrasing; whis
|
|||||||
Deploy target is a Ryzen laptop (homesrv) with Vulkan offload to the Vega iGPU (`n_gpu_layers: 99`,
|
Deploy target is a Ryzen laptop (homesrv) with Vulkan offload to the Vega iGPU (`n_gpu_layers: 99`,
|
||||||
compose passes `/dev/dri` + the render gid) — the resident model stays ≤1.7B either way.
|
compose passes `/dev/dri` + the render gid) — the resident model stays ≤1.7B either way.
|
||||||
|
|
||||||
**Resident model:** currently **Qwen3.5-0.8B** (`Q4_K_M`), the smallest checkpoint in the gguf
|
**Resident model:** currently **Qwen3-1.7B** (`UD-Q4_K_XL`), stock — not yet the CPT'd one.
|
||||||
library, picked for CPU/iGPU latency. The **target** is the locally CPT'd **Qwen3-1.7B**; that
|
It replaced Qwen3.5-0.8B on 2026-07-31 because it measured better on both fixtures we have:
|
||||||
training is still in flight (Vikunja #122), so no such gguf exists yet. Model files live in
|
67.5% vs 59.7% intent-only on the 77-case RU routing fixture, and 20/27 vs 11-17/27 on the
|
||||||
|
talk fixture. See `MODEL-BAKEOFF-31-07-2026.md`. It is a Thinking variant, so `n_ctx` is 4096
|
||||||
|
— reasoning tokens need the room, and 4096 is what the scores above were measured at.
|
||||||
|
|
||||||
|
The **target** is still the locally CPT'd **Qwen3-1.7B** (Vikunja #122, training in flight).
|
||||||
|
Stock already speaks good Russian; what it gets wrong is the persona — it writes `я рад`,
|
||||||
|
masculine, where Maven needs `рада`. That is what the CPT is for.
|
||||||
|
|
||||||
|
**Do not bother with sub-500M models.** LFM2.5-230M and 350M were measured on 2026-07-31 and
|
||||||
|
both are unusable in Russian: the 350M routes at 5.2% (worse than guessing) and answers
|
||||||
|
"столица Франции?" with the invented non-word "Сторзит"; the 230M replies to Russian in
|
||||||
|
Spanish. Their strong published IFEval/BFCL numbers are English-only. Model files live in
|
||||||
`/mnt/hdd1/llms`, bind-mounted to `/opt/maven/models/llm` — which **shadows** the repo's
|
`/mnt/hdd1/llms`, bind-mounted to `/opt/maven/models/llm` — which **shadows** the repo's
|
||||||
`models/llm/`, so the LFM2.5 gguf sitting there is not loaded by anything. Swapping the resident
|
`models/llm/`, so the LFM2.5 gguf sitting there is not loaded by anything. Swapping the resident
|
||||||
model is a one-line change to `phraser.model_path` in `deploy/mavend.json`.
|
model is a one-line change to `phraser.model_path` in `deploy/mavend.json`.
|
||||||
@@ -58,19 +69,30 @@ protocol; the config in `deploy/mavend.json` (with `${VAR}` env expansion from g
|
|||||||
|
|
||||||
## Routing — read this before touching the router
|
## Routing — read this before touching the router
|
||||||
|
|
||||||
`internal/router/` has TWO layered engines and the committed default is an **interim
|
`internal/router/` has TWO layered engines. **The LLM router is now the default and it is
|
||||||
stopgap, not the intended design** (see memory `routing-architecture-target`):
|
on in deploy** — this section used to say it was wired `nil`, which stopped being true on
|
||||||
|
2026-07-31.
|
||||||
|
|
||||||
- **Target (REARCH.md):** LLM-as-router. One resident Qwen3-1.7B (`llmrouter.go`) emits
|
- **LLM router (the intended design, REARCH.md):** the resident Qwen3-1.7B (`llmrouter.go`)
|
||||||
GBNF-constrained structured JSON, and the SAME model phrases replies. Embedder is demoted
|
emits GBNF-constrained structured JSON, and the SAME model phrases replies. Embedder is
|
||||||
from a routing gate to a RAG hint.
|
demoted from a routing gate to a RAG hint. Wired at `voice.go:214` via
|
||||||
- **Current stopgap:** `llmrouter` is wired `nil` (around `voice.go`), so the
|
`pickLLMRouter(cfg.Voice.UseLLMRouter(), llmClient)`; the flag is `voice.llm_router`
|
||||||
`classifier.go` + `embedder.go` nearest-neighbour cascade actually runs. It routes by
|
(`config.go`), `DefaultLLMRouter` is **on**, and `deploy/mavend.json` sets it `true`.
|
||||||
similarity to frozen seed phrases — the known cause of weak RU query handling.
|
- **Classifier cascade (the failure floor, not dead code):** `classifier.go` +
|
||||||
|
`embedder.go` nearest-neighbour over frozen seed phrases. It runs when the LLM router is
|
||||||
|
off, when there is no llama-server to talk to (`pickLLMRouter` logs that and degrades),
|
||||||
|
and on any per-turn LLM error. Do not delete it — routing by seed similarity is the known
|
||||||
|
cause of weak RU query handling, but a turn must never break on the model.
|
||||||
|
|
||||||
Cascade order: `stage0.go` exact-match fast-path → LLM router (when non-nil) → classifier
|
Cascade order: `stage0.go` exact-match fast-path → LLM router (when non-nil) → classifier
|
||||||
fallback. Any LLM error falls through to the classifier so a turn never breaks on the model.
|
fallback. Any LLM error falls through to the classifier so a turn never breaks on the model.
|
||||||
|
|
||||||
|
Measured on the 77-case RU fixture (`MODEL-BAKEOFF-31-07-2026.md`): the classifier scores
|
||||||
|
36.8% full accuracy at p50 31ms; Qwen3-1.7B scores 67.5% intent-only / 72.7% through the
|
||||||
|
cascade at p50 ≈2.7s. Accuracy roughly doubled, latency is ~90× worse, and that trade was
|
||||||
|
accepted deliberately. Still open: `Confidence: 1.0` is hardcoded in `llmrouter.go`, so the
|
||||||
|
LLM path never asks for clarification (6/6 refusal cases missed) — Vikunja #359.
|
||||||
|
|
||||||
## LLM output contract
|
## LLM output contract
|
||||||
|
|
||||||
All phrasing paths emit `{"response":"...","mood":"..."}` (parsed in `replier_llm.go` and
|
All phrasing paths emit `{"response":"...","mood":"..."}` (parsed in `replier_llm.go` and
|
||||||
@@ -82,8 +104,26 @@ workspace enforces that the Go and relabelling prompts remain identical.
|
|||||||
|
|
||||||
## Non-goals (hard constraints)
|
## Non-goals (hard constraints)
|
||||||
|
|
||||||
Never phones home. Not a nag, not autonomous. Maven's persona is **feminine** — Russian
|
Not a nag, not autonomous. Maven's persona is **feminine** — Russian
|
||||||
self-reference must use feminine forms (the user is male; see memory `maven-persona-gender`).
|
self-reference must use feminine forms — `рада`, not `рад`; `поняла`, not `понял`. The owner
|
||||||
|
is male and is addressed informally: "ты", singular, never "вы"/"ваш" and never "он"/"его"
|
||||||
|
(she talks TO him, not about him). Pet names ("милый", "дорогой") are forbidden; his name
|
||||||
|
("Ками") is not. The eval enforces this: `CheckAddress`, `CheckFeminine` and `CheckCringe` in
|
||||||
|
`internal/phraser/eval/checks.go`, scored by `make eval-phrasing`.
|
||||||
|
|
||||||
|
**"Never phones home" is DEPRECATED** (owner's call, 2026-07-31). It used to be a hard
|
||||||
|
constraint and it is not one any more: a 0.8B — and a 1.7B — does not know enough to answer
|
||||||
|
world questions, so she needs to read external sources. What replaces it:
|
||||||
|
|
||||||
|
- **No telemetry, no cloud model, no third-party account.** That part never changes. Nothing
|
||||||
|
about Maven is reported to anyone, and inference stays on the box.
|
||||||
|
- **Local sources first.** Kiwix ZIMs on homesrv (Wikipedia, ifixit) before anything on the
|
||||||
|
network. Reading beats recalling for a small model, and a local read costs nothing.
|
||||||
|
- **External search is allowed and off unless configured**, like the weather and telegram
|
||||||
|
capabilities.
|
||||||
|
- **His notes and facts are never search input.** Looking up why the sky is blue and sending
|
||||||
|
his stored personal notes to an upstream engine are different acts. Only the utterance goes
|
||||||
|
out, never the persona block, history, or matched notes.
|
||||||
|
|
||||||
## Web UI conventions
|
## Web UI conventions
|
||||||
|
|
||||||
|
|||||||
@@ -15,7 +15,8 @@
|
|||||||
|
|
||||||
**Maven** — self-hosted personal assistant. Manages your day, acts on your
|
**Maven** — self-hosted personal assistant. Manages your day, acts on your
|
||||||
homelab. One daemon on homesrv (always-on, not the workstation), multiple
|
homelab. One daemon on homesrv (always-on, not the workstation), multiple
|
||||||
client surfaces. All local, never phones home.
|
client surfaces. Inference and data stay on the box; she may READ external
|
||||||
|
sources (see Non-goals — "never phones home" is deprecated).
|
||||||
|
|
||||||
Primary name is "Maven", with feminine-gendered Russian self-reference
|
Primary name is "Maven", with feminine-gendered Russian self-reference
|
||||||
("она", "меня", "помогла"). Clients may choose their own UI label. Consistent
|
("она", "меня", "помогла"). Clients may choose their own UI label. Consistent
|
||||||
@@ -35,8 +36,13 @@ Inside boundary — the ones that actually constrain the build:
|
|||||||
she records. A confident wrong fact is worse than a known gap.
|
she records. A confident wrong fact is worse than a known gap.
|
||||||
- **Not a nag** — she'd rather miss a nudge than be mutable. Shuts up when
|
- **Not a nag** — she'd rather miss a nudge than be mutable. Shuts up when
|
||||||
uncertain. Load-bearing.
|
uncertain. Load-bearing.
|
||||||
- **Not a stranger** — runs on your stuff, your model, your data. Never
|
- **Not a stranger** — runs on your stuff, your model, your data. No
|
||||||
phones home.
|
telemetry, no cloud model, no third-party account. She may READ external
|
||||||
|
sources to answer world questions (Kiwix first, then optional search); she
|
||||||
|
never reports anything about you to anyone, and your notes and facts are
|
||||||
|
never used as search input. **"Never phones home" as an absolute is
|
||||||
|
deprecated** — owner's call, 2026-07-31: a small model does not know enough
|
||||||
|
to be useful without reading.
|
||||||
- **Not a relationship** — mom-tone is a function that makes nudges land, not
|
- **Not a relationship** — mom-tone is a function that makes nudges land, not
|
||||||
emotional company. Names the drift a warm small model falls into.
|
emotional company. Names the drift a warm small model falls into.
|
||||||
|
|
||||||
@@ -458,7 +464,7 @@ decides *insistence*. Both are needed.
|
|||||||
|
|
||||||
sev ≤ 2 drops on away, sev ≥ 3 holds: a missed water nudge is noise, a missed
|
sev ≤ 2 drops on away, sev ≥ 3 holds: a missed water nudge is noise, a missed
|
||||||
backup failure isn't. Away-channels (ntfy/telegram) leave the box — the one
|
backup failure isn't. Away-channels (ntfy/telegram) leave the box — the one
|
||||||
path that crosses "never phones home," through your own relay. **Minimal
|
path that leaves the box for a person to see, through your own relay. **Minimal
|
||||||
body** — "disk low on homesrv," not detail; don't make notifications a
|
body** — "disk low on homesrv," not detail; don't make notifications a
|
||||||
shoulder-surf exfil surface.
|
shoulder-surf exfil surface.
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,223 @@
|
|||||||
|
# Resident model bake-off — 31-07-2026
|
||||||
|
|
||||||
|
**Outcome: the resident model is stock Qwen3-1.7B** (`UD-Q4_K_XL`). Two sweeps ran this
|
||||||
|
evening and the second one changed the answer — read to the end before acting on any table
|
||||||
|
here. [Second sweep](#second-sweep-same-evening--five-models-and-a-resident-model-change)
|
||||||
|
is the one that holds.
|
||||||
|
|
||||||
|
## First sweep — LFM2.5-1.2B vs Qwen3.5-0.8B
|
||||||
|
|
||||||
|
**Verdict, scoped to this pair: keep Qwen3.5-0.8B over LFM2.5-1.2B.** LFM2.5-1.2B is worse
|
||||||
|
at routing (52.6% vs 60.5% intent accuracy), and the loss is almost entirely Russian
|
||||||
|
(18/61 vs 22/61 RU, while EN is a wash). It is also 2.4× slower. The Thinking variant is
|
||||||
|
far worse again. This verdict still stands as written — it rejects LFM2.5-1.2B. It is
|
||||||
|
**not** a recommendation to keep 0.8B as the resident model; the second sweep replaced it
|
||||||
|
with Qwen3-1.7B.
|
||||||
|
|
||||||
|
Settles Vikunja **#278 / #250**.
|
||||||
|
|
||||||
|
- Same fixture and scorer as `ROUTING-EVAL-31-07-2026.md`: `internal/router/eval/`
|
||||||
|
(`ru_routing_v1.json`, 76 held-out cases).
|
||||||
|
- Reproduce: `MAVEN_LLM_URL=http://127.0.0.1:<port> make eval-router`
|
||||||
|
(`TestLLMRouterBaseline`). (This line used to say there is no `make eval-models` target.
|
||||||
|
There is one now — start a server with the gguf you want, then
|
||||||
|
`make eval-models MAVEN_LLM_URL=http://127.0.0.1:<port>`. It runs only the LLM test, since
|
||||||
|
the classifier baselines do not depend on the model.)
|
||||||
|
- All three models served by the same `llama-server` flags — `-c 2048 -ngl 99 -t 6`, only
|
||||||
|
`-m` and `--port` differ. One server at a time on an otherwise idle box, so latencies are
|
||||||
|
real and not contention.
|
||||||
|
- Measured on top of the router prompt fix (`origin/overnight/router-prompt` merged in), so
|
||||||
|
the Qwen column is directly comparable to the numbers already recorded.
|
||||||
|
|
||||||
|
## Results
|
||||||
|
|
||||||
|
`llm-only` — the model alone. This is the column that measures the model.
|
||||||
|
|
||||||
|
| | Qwen3.5-0.8B | LFM2.5-1.2B Instruct | LFM2.5-1.2B Thinking |
|
||||||
|
|---|---|---|---|
|
||||||
|
| **intent-only accuracy** | **60.5%** | 52.6% | 36.8% |
|
||||||
|
| full accuracy (intent+slots+gate) | **36.8%** | 32.9% | 21.1% |
|
||||||
|
| **RU** | **22/61** | 18/61 | 10/61 |
|
||||||
|
| EN | 6/15 | **7/15** | 6/15 |
|
||||||
|
| route errors | 0 | 0 | 0 |
|
||||||
|
| **p50 / p95 latency** | **1.05s / 1.71s** | 2.47s / 3.62s | 2.42s / 3.24s |
|
||||||
|
| missed clarify | 6 / 6 | 6 / 6 | 6 / 6 |
|
||||||
|
|
||||||
|
`cascade+llm` — stage-0 → model → classifier floor, what #320 would actually ship. Same
|
||||||
|
ordering.
|
||||||
|
|
||||||
|
| | Qwen3.5-0.8B | LFM2.5-1.2B Instruct | LFM2.5-1.2B Thinking |
|
||||||
|
|---|---|---|---|
|
||||||
|
| intent-only accuracy | **61.8%** | 55.3% | 38.2% |
|
||||||
|
| full accuracy | **46.1%** | 42.1% | 30.3% |
|
||||||
|
| RU / EN | **27/61** / 8/15 | 23/61 / **9/15** | 15/61 / 8/15 |
|
||||||
|
| route errors | 0 | 0 | 0 |
|
||||||
|
| p50 / p95 latency | **1.28s / 1.94s** | 2.18s / 2.72s | 2.27s / 3.19s |
|
||||||
|
|
||||||
|
Full logs: the three runs are archived in the session scratchpad
|
||||||
|
(`qwen08.txt`, `lfm-instruct.txt`, `lfm-thinking.txt`).
|
||||||
|
|
||||||
|
## Russian-specific failures — the owner's worry is confirmed
|
||||||
|
|
||||||
|
LFM2.5's Russian loss is not spread out. It has one large, specific failure: **it hears
|
||||||
|
almost any Russian imperative or short phrase as `reminder`.**
|
||||||
|
|
||||||
|
- `перезапусти докер` → reminder (want act)
|
||||||
|
- `включи вытяжку` → reminder (want act)
|
||||||
|
- `закрой жалюзи` → reminder (want act)
|
||||||
|
- `заметка: продлить домен в августе` → reminder (want note)
|
||||||
|
- `запиши что кран на кухне снова капает` → reminder (want note)
|
||||||
|
- `доброе утро` → reminder (want chat)
|
||||||
|
- `спасибо тебе` → reminder (want note/chat)
|
||||||
|
- `переходи в тихий режим` → reminder (want system)
|
||||||
|
|
||||||
|
That is `note→reminder ×4`, `act→reminder ×4`, `chat→reminder ×2` in one run. Qwen's
|
||||||
|
equivalent failure axis is `query→fact ×8`, which is a narrower and already-understood bug.
|
||||||
|
|
||||||
|
Two more Russian-side problems worth naming:
|
||||||
|
|
||||||
|
1. **Fact keys come back empty or wrong in Russian.** `воды попил наконец`, `поужинал`,
|
||||||
|
`поспал часов пять` and `отметь что я позавтракал овсянкой` all returned an empty key.
|
||||||
|
`сходил в душ` and `отдохнул минут двадцать` both returned `water`. Qwen does not do this.
|
||||||
|
2. **It leaked German.** `slept about seven hours` produced the fact key
|
||||||
|
`"7 Stunden geschlafen"`. Grammar-valid, semantically garbage — a sign the multilingual
|
||||||
|
mix is not anchored where Maven needs it.
|
||||||
|
|
||||||
|
The claimed tool-calling advantage did not show up here. `act` is the closest thing this
|
||||||
|
fixture has to a tool call, and LFM2.5 got it wrong more often than Qwen, mostly by calling
|
||||||
|
it a reminder. It also produced no `fn` slot on any act, same as Qwen.
|
||||||
|
|
||||||
|
## The Thinking variant
|
||||||
|
|
||||||
|
Not viable. 36.8% intent accuracy, 10/61 Russian, and no latency saving over Instruct — the
|
||||||
|
thinking trace costs time without buying accuracy on a short enum classification. With the
|
||||||
|
`enable_thinking=false` diagnostic it collapsed further to 28.9% with 2 route errors
|
||||||
|
(`query→reminder ×12`). Do not pursue.
|
||||||
|
|
||||||
|
## Notes
|
||||||
|
|
||||||
|
- Nothing crashed, nothing ignored the GBNF grammar, and no model produced unparseable JSON
|
||||||
|
in the shippable configurations. Zero route errors for both Instruct and Thinking in
|
||||||
|
`llm-only` and `cascade+llm`. The problem with LFM2.5 is what it decides, not whether it
|
||||||
|
can emit the contract.
|
||||||
|
- The `6 / 6` missed clarify is unchanged across all three models. No model fixes the missing
|
||||||
|
refusal lane — that is `Confidence: 1.0` hardcoded in `llmrouter.go` (Vikunja #359), not a
|
||||||
|
model property.
|
||||||
|
- The report labels every configuration `(0.8B)`; that string is hardcoded in the test, not a
|
||||||
|
reflection of which gguf was loaded. Model identity was confirmed per run via `/v1/models`.
|
||||||
|
- No Go code was changed for this measurement, and no bug was found that needed one.
|
||||||
|
|
||||||
|
## What this does not settle
|
||||||
|
|
||||||
|
Routing only. LFM2.5 might still phrase better, and phrasing is the resident model's other
|
||||||
|
job — that needs its own fixture. But routing is the load-bearing path and Maven is
|
||||||
|
Russian-first, so on the evidence here the switch is not worth making.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# Second sweep, same evening — five models, and a resident-model change
|
||||||
|
|
||||||
|
The sections above compared LFM2.5-1.2B against Qwen3.5-0.8B on routing and concluded
|
||||||
|
"the switch is not worth making". That still holds. This sweep asked a different
|
||||||
|
question — whether a *smaller* model could work, since LFM2.5's published
|
||||||
|
instruction-following scores beat Qwen3.5-0.8B badly — and answered it, plus found a
|
||||||
|
better resident model by accident.
|
||||||
|
|
||||||
|
**Outcome: the resident model is now stock Qwen3-1.7B.** Sub-500M is a dead end.
|
||||||
|
|
||||||
|
## Routing — 77 Russian cases, one run each
|
||||||
|
|
||||||
|
| model | on disk | llm-only (full) | llm-only (intent) | cascade + fallback |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| LFM2.5-230M-Q8_0 | 246 MB | 23.4% | 33.8% | 36.4% |
|
||||||
|
| LFM2.5-350M-Q8_0 | 379 MB | 2.6% | **5.2%** | 20.8% |
|
||||||
|
| Qwen3.5-0.8B-Q4_K_M | 527 MB | 36.4% | 59.7% | 61.0% |
|
||||||
|
| Qwen3.5-2B-UD-Q4_K_XL | 1.34 GB | 42.9% | 62.3% | 63.6% |
|
||||||
|
| **Qwen3-1.7B-UD-Q4_K_XL (stock)** | 1.13 GB | **44.2%** | **67.5%** | **72.7%** |
|
||||||
|
|
||||||
|
Qwen3-1.7B wins every column, including against a model 20% larger than it.
|
||||||
|
|
||||||
|
## Talk fixture — 27 cases, three runs each, idle box
|
||||||
|
|
||||||
|
| | Qwen3.5-0.8B | Qwen3-1.7B stock |
|
||||||
|
|---|---|---|
|
||||||
|
| composite | 13, 11, 8 | **20, 21, 18** |
|
||||||
|
| address | 21, 18, 18 | **26, 25, 23** |
|
||||||
|
| feminine | 27, 25, 26 | 26, 27, 26 |
|
||||||
|
| lang | 27, 27, 26 | 26, 27, 27 |
|
||||||
|
| ontopic | 16, 19, 19 | **22, 23, 23** |
|
||||||
|
| canned fallbacks | 8, 5, 6 | **0, 2, 0** |
|
||||||
|
|
||||||
|
This also fills the row `TALK-EVAL-31-07-2026.md` had to void for contamination:
|
||||||
|
**600ch/1024tok on Qwen3.5-0.8B scores 13, 11, 8.**
|
||||||
|
|
||||||
|
`address` is the headline. It sat at 18-22 of 27 on the 0.8B no matter how the prompt
|
||||||
|
was worded — the prompt explicitly forbids "вы" and the model writes `вашей`,
|
||||||
|
`подождите`, `делаете` anyway. That was read as "prompting is out of levers", and it
|
||||||
|
was really "0.8B is out of capacity". The 1.7B mostly holds the constraint.
|
||||||
|
|
||||||
|
The fallback column matters too: 5-8 of 27 turns on the 0.8B end in a hardcoded
|
||||||
|
`"не знаю."`, meaning it failed to emit parseable JSON about a quarter of the time.
|
||||||
|
The 1.7B does that 0-2 times.
|
||||||
|
|
||||||
|
## Latency — the long tail is not the Thinking block
|
||||||
|
|
||||||
|
| | p50 | p95 |
|
||||||
|
|---|---|---|
|
||||||
|
| Qwen3.5-0.8B | 2.4s, 2.9s, 2.0s | 17.4s, 17.6s, 17.4s |
|
||||||
|
| Qwen3-1.7B stock | 2.7s, 2.6s, 2.8s | 16.4s, 6.6s, 3.9s |
|
||||||
|
|
||||||
|
p50 is flat across a 2× size difference. The first instinct on seeing the 1.7B's
|
||||||
|
16s p95 was "that is the reasoning trace, cap it" — wrong. The 0.8B's p95 is a
|
||||||
|
consistent 17s and the 1.7B beat it in two of three runs. The tail is shared and
|
||||||
|
lives somewhere else. Do not spend time on `/no_think` on this evidence.
|
||||||
|
|
||||||
|
## Sub-500M: not close, and the benchmarks say otherwise for a reason
|
||||||
|
|
||||||
|
LFM2.5-350M publishes IFEval 76.96 against Qwen3.5-0.8B's 59.94, and BFCLv3 44.11
|
||||||
|
against 35.08 — better at instruction-following and structured output, at 2/3 the
|
||||||
|
size. Those numbers are real and they are **English**. Every benchmark in that
|
||||||
|
table except Multi-IF is English-only.
|
||||||
|
|
||||||
|
In Russian, with a 300-token budget and temperature 0:
|
||||||
|
|
||||||
|
- **350M**, «Столица Франции? Ответь кратко.» → *«Сторзит в Париже.»* — `Сторзит` is
|
||||||
|
not a word; it is invented morphology.
|
||||||
|
- **350M**, asked to read back a reminder → a fortune cookie about being attentive
|
||||||
|
and confident. No reminder in it.
|
||||||
|
- **230M**, «Привет, как дела?» → answered **in Spanish**.
|
||||||
|
|
||||||
|
The 230M beating the 350M six-fold on routing (33.8% vs 5.2%) is the other tell:
|
||||||
|
when the larger sibling collapses like that it is format compliance failing, not
|
||||||
|
reasoning.
|
||||||
|
|
||||||
|
This is a pretraining gap, not a fine-tuning gap. Teaching Russian to a 350M from
|
||||||
|
near-zero is not an afternoon on a Colab, which was the premise worth checking.
|
||||||
|
|
||||||
|
## Why this vindicates the 1.7B CPT
|
||||||
|
|
||||||
|
Stock Qwen3-1.7B, untrained and unprompted, answers all three probes in fluent
|
||||||
|
correct Russian. What it gets wrong is the persona: *«Привет! Я рад, что ты здесь»*
|
||||||
|
— `рад` is masculine and Maven needs `рада`. That is the right kind of remaining
|
||||||
|
problem, and it is exactly what the CPT (Vikunja #122) is for.
|
||||||
|
|
||||||
|
The 1.7B was the correct model choice. What was wrong was treating it as a
|
||||||
|
**blocker**: stock already beats what was deployed, so it ships now and gets
|
||||||
|
swapped again when the CPT lands.
|
||||||
|
|
||||||
|
## Caveats
|
||||||
|
|
||||||
|
- Routing is one run per model, not three. The gaps between families are far larger
|
||||||
|
than the run-to-run spread seen on the talk fixture, but the 2B-vs-1.7B gap (62.3
|
||||||
|
vs 67.5) is not safe to call on one run.
|
||||||
|
- ~~The routing numbers only reach production once the LLM router is wired on. It is
|
||||||
|
still `nil`.~~ **Resolved the same evening:** the LLM router is wired at `voice.go:214`
|
||||||
|
behind `voice.llm_router`, the default is on, and `deploy/mavend.json` sets it `true`.
|
||||||
|
These numbers are the production path now, so the p50 ≈2.7s is a real per-turn cost and
|
||||||
|
not a bench artifact.
|
||||||
|
- ~~`/mnt/hdd1/llms/LFM2.5/Qwen3-1.7B-UD-Q4_K_XL.gguf` is a 293 MB truncated download
|
||||||
|
in the wrong directory.~~ **Deleted 2026-07-31.** The good 1.13 GB copy in `qwen3/` is
|
||||||
|
what `deploy/mavend.json` loads.
|
||||||
|
- Harness: `scratchpad/bakeoff.sh`, one server at a time, health-checked before each
|
||||||
|
run, `/v1/models` recorded per run. Never run two LLM consumers at once — see the
|
||||||
|
contamination note in `TALK-EVAL-31-07-2026.md`.
|
||||||
@@ -103,14 +103,17 @@ eval-router:
|
|||||||
eval-recall:
|
eval-recall:
|
||||||
MAVEN_ONNX_LIB="$(MAVEN_ONNX_LIB)" $(GO) test -v -count=1 ./internal/memory/recalleval/
|
MAVEN_ONNX_LIB="$(MAVEN_ONNX_LIB)" $(GO) test -v -count=1 ./internal/memory/recalleval/
|
||||||
|
|
||||||
# eval-phrasing -- score nudge phrasing (internal/phraser/eval). Verbose so the
|
# eval-phrasing -- score nudge phrasing AND the conversational paths (chat,
|
||||||
|
# query, general knowledge) in internal/phraser/eval. Verbose so the
|
||||||
# report and every generated message land in the terminal. With no environment
|
# report and every generated message land in the terminal. With no environment
|
||||||
# it scores the deterministic Stub only, which is what CI runs. Set
|
# it scores the deterministic Stub only, which is what CI runs. Set
|
||||||
# MAVEN_LLM_URL to add the resident model:
|
# MAVEN_LLM_URL to add the resident model:
|
||||||
# MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-phrasing
|
# MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-phrasing
|
||||||
# The model run is slow (minutes) -- the timeout is raised to match.
|
# The model run is slow (minutes) -- the timeout is raised to match. It covers
|
||||||
|
# two fixtures now (15 nudges + 27 conversational cases, and the chat replies are
|
||||||
|
# the long ones), hence 90m rather than 40m.
|
||||||
eval-phrasing:
|
eval-phrasing:
|
||||||
$(GO) test -v -count=1 -timeout 40m ./internal/phraser/eval/
|
$(GO) test -v -count=1 -timeout 90m ./internal/phraser/eval/
|
||||||
|
|
||||||
# eval-models — score ONE llama-server against the same fixture, for the
|
# eval-models — score ONE llama-server against the same fixture, for the
|
||||||
# resident-model bake-off (#278, #250). Start a server with the gguf you want,
|
# resident-model bake-off (#278, #250). Start a server with the gguf you want,
|
||||||
|
|||||||
@@ -0,0 +1,169 @@
|
|||||||
|
# Phrasing evaluation — 31-07-2026
|
||||||
|
|
||||||
|
How Maven words a nudge, measured instead of argued. Counterpart to
|
||||||
|
`ROUTING-EVAL-31-07-2026.md`.
|
||||||
|
|
||||||
|
- Fixture + scorer: `internal/phraser/eval/` (`nudges_v1.json`, 15 cases; `eval.go`, `checks.go`)
|
||||||
|
- Reproduce: `MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-phrasing`
|
||||||
|
- Model: Qwen3.5-0.8B Q4_K_M, the resident model. Not swapped.
|
||||||
|
- Commit: `a40bc55` (prompt fix)
|
||||||
|
|
||||||
|
Every check is a string or length test a human can read and disagree with. No model
|
||||||
|
grades another model here.
|
||||||
|
|
||||||
|
## Result
|
||||||
|
|
||||||
|
| | before | after |
|
||||||
|
|---|---|---|
|
||||||
|
| **cases passing every check** | **0/15** | **13/15** |
|
||||||
|
| mood in enum | 6/15 | 15/15 |
|
||||||
|
| Russian | 2/15 | 14/15 |
|
||||||
|
| length (≤120 chars, ≤16 words) | 13/15 | 15/15 |
|
||||||
|
| feminine self-reference | 15/15 | 15/15 |
|
||||||
|
| no cringe | 13/15 | 15/15 |
|
||||||
|
| on topic | 6/15 | 13/15 |
|
||||||
|
| p50 latency | 11.4s | 11.4s |
|
||||||
|
|
||||||
|
Latency did not move and is not good. 11s to word one nudge on this box.
|
||||||
|
|
||||||
|
## The bug reproduced
|
||||||
|
|
||||||
|
Yes, exactly as reported. 7 of 15 messages were the literal string `"..."`, and one was
|
||||||
|
`"full voice message"`. Both are text copied straight out of the prompt.
|
||||||
|
|
||||||
|
The system prompt said:
|
||||||
|
|
||||||
|
```
|
||||||
|
Respond ONLY with valid JSON: {"response": "full voice message", "mood": "neutral"}
|
||||||
|
```
|
||||||
|
|
||||||
|
and the user prompt said:
|
||||||
|
|
||||||
|
```
|
||||||
|
Respond as JSON: {"response": "...", "mood": "..."}
|
||||||
|
```
|
||||||
|
|
||||||
|
A 0.8B does not read `"..."` as "put your answer here". It reads it as the answer. The
|
||||||
|
prompt was a worked example whose worked part was blank, so the model filled the slot by
|
||||||
|
copying. This is the whole of finding 1.
|
||||||
|
|
||||||
|
## What else was wrong
|
||||||
|
|
||||||
|
Four separate faults, all prompt-side:
|
||||||
|
|
||||||
|
1. **Placeholder echo** (7 cases) — above.
|
||||||
|
2. **Wrong language** (13/15 failed the language check). The prompt was entirely English
|
||||||
|
and said "in the user's language (Russian or English)". The model picked English. It is
|
||||||
|
never English: the nudge is spoken by a Russian piper voice.
|
||||||
|
3. **Rule names are English identifiers.** `netdata_critical`, `service_down`, `break` went
|
||||||
|
into the prompt raw. The model cannot nudge about a topic it has not been told in words,
|
||||||
|
so 9/15 were off topic. The daemon knows what its own rules mean; now it says so.
|
||||||
|
4. **Mood invented** (`"warm"`, twice). The enum was listed in a parenthesis at the end of
|
||||||
|
an English sentence. Now it is its own line: "ровно одно из: neutral, happy, thinking,
|
||||||
|
tired, confused."
|
||||||
|
|
||||||
|
Plus two non-prompt faults the run exposed:
|
||||||
|
|
||||||
|
- **The no-parse fallback was English.** When the model returned nothing usable, the body
|
||||||
|
became `fmt.Sprintf("%s — %s", rule, sev)` — `"water — care"` — and that string went to
|
||||||
|
a Russian TTS. Now it falls back to plain Russian.
|
||||||
|
- **Durations were English.** `humanDur` returns "3 hours"; it was landing verbatim inside
|
||||||
|
Russian sentences. Nudges now use a Russian formatter.
|
||||||
|
|
||||||
|
## Three iterations, and what each taught
|
||||||
|
|
||||||
|
| | score | change |
|
||||||
|
|---|---|---|
|
||||||
|
| baseline | 0/15 | — |
|
||||||
|
| iter 1 | 2/15 | Russian prompt, filled-in examples, Russian durations |
|
||||||
|
| iter 2 | 11/15 | required keyword per rule, one example instead of five, Russian fallback |
|
||||||
|
| iter 3 | **13/15** | examples moved to topics that are not rules |
|
||||||
|
|
||||||
|
The interesting step is 1 → 2. Fixing the placeholder did not fix the disease, it moved it:
|
||||||
|
the model stopped copying `"..."` and started copying my first example instead. Five nudges
|
||||||
|
in a row came back as `"Ты не пил воду три часа. Налей стакан."` regardless of the rule.
|
||||||
|
|
||||||
|
**A small model copies the nearest concrete text in its prompt.** That is one failure mode
|
||||||
|
with two symptoms. The fix that stuck was making the examples about laundry and a laptop
|
||||||
|
battery — topics no rule ever produces, so copying them is visible in the score rather than
|
||||||
|
invisibly passing the water cases.
|
||||||
|
|
||||||
|
## Do not oversell 13/15
|
||||||
|
|
||||||
|
Seven of the thirteen passes are the **deterministic fallback**, not the model:
|
||||||
|
`"Напоминаю: таблетки."`, `"Сервис не отвечает."`, `"Критический алярм: проверь диск."`,
|
||||||
|
`"Ты давно не пил воду."`. Those are strings this commit added to Go. The model returned
|
||||||
|
nothing parseable and the fallback scored.
|
||||||
|
|
||||||
|
So the honest reading is roughly **6/15 from the model, 7/15 from a fallback, 2/15 failing**.
|
||||||
|
The prompt fix is real — `"..."` is nearly gone and the language and mood checks are clean —
|
||||||
|
but a large part of the jump is that failure now degrades into Russian instead of into
|
||||||
|
`"water — care"`. That is a genuine improvement for the operator and a weak one for the model.
|
||||||
|
|
||||||
|
The two remaining failures: one `"..."` recurrence (`routine-stretch`) and one meal nudge
|
||||||
|
that never says food.
|
||||||
|
|
||||||
|
## Tried and reverted: an example-led nudge prompt (#393)
|
||||||
|
|
||||||
|
The idea was that a 0.8B copies examples better than it follows rules, so the nudge prompt
|
||||||
|
was rewritten to lead with five on-topic examples (water, break, pills, morning, service) and
|
||||||
|
the prose rules were compressed to pay for the tokens: 1190 chars down to 986.
|
||||||
|
|
||||||
|
It measured **worse**, three runs each side, same llama-server, same fixture:
|
||||||
|
|
||||||
|
| run | before | after |
|
||||||
|
|---|---|---|
|
||||||
|
| 1 | 12/15 (address 14) | 11/15 (address 13) |
|
||||||
|
| 2 | 13/15 (address 15) | 12/15 (address 15) |
|
||||||
|
| 3 | 14/15 (address 15) | 11/15 (address 12) |
|
||||||
|
|
||||||
|
`feminine` and `hisgender` were 15/15 on all six runs, so they measure nothing here. The
|
||||||
|
regression is all in `address`: 44/45 before, 40/45 after. Formal "вы"/"ваше" and plural
|
||||||
|
imperatives came back, and so did `"..."`.
|
||||||
|
|
||||||
|
Two likely causes, both about the same thing — **examples do not carry a prohibition**. The
|
||||||
|
old prompt spent a whole sentence on «говоришь на "ты", в единственном числе»; the new one
|
||||||
|
demoted that to one item in a long "никогда" list, and the model stopped obeying it. And
|
||||||
|
making the examples on-topic let their *wording* leak: a break case came back as
|
||||||
|
«Вы давно не пили воду. Выпей стакан.» — the water example, verbatim, in the wrong slot.
|
||||||
|
That is exactly the failure the laundry/laptop examples were chosen to avoid.
|
||||||
|
|
||||||
|
Change reverted. What survives is the measurement: a rule the model must obey needs its own
|
||||||
|
sentence, and examples must stay off-topic. Also note the before side alone spans 12–14 of
|
||||||
|
15 — this fixture cannot resolve anything smaller than about three cases.
|
||||||
|
|
||||||
|
## Broken, found, not fixed
|
||||||
|
|
||||||
|
1. ~~**`checkFeminine` only catches half the constraint.**~~ **Fixed** (#381). It scanned for
|
||||||
|
masculine self-reference only, so three messages that addressed the *owner* in the feminine
|
||||||
|
("ты давно не отдыхал**а**") scored clean. There is now a second check, `hisgender`: a
|
||||||
|
feminine past-tense verb (-ла/-лась) in a sentence addressed to him ("ты", "тебе", "твой")
|
||||||
|
fails, unless the verb is hers ("я заметила", "напомнила тебе"). It is a suffix rule, not a
|
||||||
|
parser — see the comment in `checks.go` for what it misses. A fresh 15-case run after adding
|
||||||
|
it scored **12/15** with `hisgender` 15/15; the model did not repeat the feminine address in
|
||||||
|
that sample, and the check is pinned by unit tests on the recorded bad strings instead.
|
||||||
|
2. **Grammar is not checked at all, and it is bad.** `"Он не ел 11 дней"` (it was 11 hours),
|
||||||
|
`"Сонуждились 7 дней"` (not a word), `"Они забыли воду"` (wrong person entirely). Every
|
||||||
|
one of these passes all six checks. The fixture measures properties, not fluency, and at
|
||||||
|
0.8B fluency is the binding constraint.
|
||||||
|
3. **Unit confusion.** The model turns hours into days about a third of the time. The
|
||||||
|
prompt now says "11 ч"; it reads it as days.
|
||||||
|
4. **11s p50.** Unchanged and untouched here. A nudge the model takes eleven seconds to
|
||||||
|
word has missed its moment. Worth its own task.
|
||||||
|
5. **The keyword hint is close to teaching to the test.** `ruleKeywords` names the word the
|
||||||
|
on-topic check looks for. It is defensible — the daemon genuinely knows its rule topics
|
||||||
|
and the model genuinely cannot infer them from `netdata_critical` — but the on-topic
|
||||||
|
number is softer than the others because of it.
|
||||||
|
|
||||||
|
## Next steps
|
||||||
|
|
||||||
|
1. ~~**Add a second-person gender check**~~ — done, `hisgender` in `checks.go` (#381).
|
||||||
|
2. **Decide whether the fallback should count as a pass.** Right now `Score` cannot tell a
|
||||||
|
model answer from a fallback. Either mark fallback bodies in `PhrasedNudge` or count them
|
||||||
|
in their own column. Without that, any future prompt change can score well by failing
|
||||||
|
more.
|
||||||
|
3. **Attack the 11s.** Nudge phrasing is short and non-interactive; thinking off is the first
|
||||||
|
thing to try, as it was for routing (#376).
|
||||||
|
4. **Re-measure when #122 lands.** The CPT'd Qwen3-1.7B is the target. 13/15 with seven
|
||||||
|
fallbacks is the floor it has to beat, and the fluency problems above are the ones a
|
||||||
|
bigger, Russian-trained checkpoint should actually fix.
|
||||||
@@ -90,4 +90,7 @@ later* is the worker + RAG.
|
|||||||
4. **Deferred work** — larger reasoner, custom Piper voice and other expansions.
|
4. **Deferred work** — larger reasoner, custom Piper voice and other expansions.
|
||||||
|
|
||||||
## Non-goals (unchanged)
|
## Non-goals (unchanged)
|
||||||
Never phones home. Not a nag. Not autonomous. Feminine-gendered RU self-ref.
|
Not a nag. Not autonomous. Feminine-gendered RU self-ref. No telemetry, no
|
||||||
|
cloud model, no third-party account — but she MAY read external sources to
|
||||||
|
answer world questions (Kiwix first, search optional). "Never phones home" as
|
||||||
|
an absolute is deprecated, owner's call 2026-07-31; see CLAUDE.md § Non-goals.
|
||||||
|
|||||||
@@ -75,6 +75,14 @@ same vector. A note is indexed in both places with the same embedding, so if it
|
|||||||
`QueryNotes` it fails again here — the branch can only ever return a **fact**. Its comment calls it
|
`QueryNotes` it fails again here — the branch can only ever return a **fact**. Its comment calls it
|
||||||
"additive"; for notes it is not.
|
"additive"; for notes it is not.
|
||||||
|
|
||||||
|
**Fixed (Vikunja #373).** The memory pass now runs *first*, as one search over notes and facts with
|
||||||
|
one gate, so whichever memory is clearly the best match answers — note or fact. The notes-only pass
|
||||||
|
stays behind it for notes the vector index does not hold. No threshold changed, so the set of
|
||||||
|
questions Maven answers is the same; only which memory answers them. The fixture gained two mixed
|
||||||
|
note+fact cases (`ru-mixed-031`, `ru-mixed-032`), which is why the counts below are out of 27
|
||||||
|
answerable cases and not 25: hash recall@1 36.0% (9/25) → 37.0% (10/27), e5 recall@1 72.0% (18/25) →
|
||||||
|
70.4% (19/27) with answered-after-gate 68.0% → 66.7% and false recall unchanged at 1/5.
|
||||||
|
|
||||||
### 5. Ranking has no recency or type signal, and the store is not the bottleneck
|
### 5. Ranking has no recency or type signal, and the store is not the bottleneck
|
||||||
|
|
||||||
`internal/store/notes.go:67` sorts by cosine and uses `ts` only to break an exact float tie, which
|
`internal/store/notes.go:67` sorts by cosine and uses `ts` only to break an exact float tie, which
|
||||||
|
|||||||
+108
-10
@@ -66,15 +66,113 @@ Three things this run settles:
|
|||||||
`запиши что…` phrasings toward fact, and that suspicion stands — all five `ru-note-*`
|
`запиши что…` phrasings toward fact, and that suspicion stands — all five `ru-note-*`
|
||||||
cases now land on fact. Tracked as Vikunja #375.
|
cases now land on fact. Tracked as Vikunja #375.
|
||||||
|
|
||||||
**Thinking off is the best configuration measured so far**, on both accuracy and latency
|
The `thinking off` column above read as the best configuration measured so far (Vikunja #376).
|
||||||
(Vikunja #376). That is worth understanding before flipping: routing is a short
|
**It was wrong** — see the controlled re-run below. Ignore that column.
|
||||||
classification into a fixed enum with grammar-constrained output, so there is little to
|
|
||||||
reason about, and the thinking trace mostly gives a small model room to talk itself out of
|
|
||||||
the right answer. Phrasing is a different job and needs measuring separately.
|
|
||||||
|
|
||||||
Still `6 / 6` missed clarify — the router has no way to say "I don't know" (Vikunja #359).
|
Still `6 / 6` missed clarify — the router has no way to say "I don't know" (Vikunja #359).
|
||||||
That is unchanged by anything here.
|
That is unchanged by anything here.
|
||||||
|
|
||||||
|
## Thinking off — 31-07-2026, controlled re-run (Vikunja #376)
|
||||||
|
|
||||||
|
The "thinking off wins by 6 points" observation above **does not hold**. It was a measurement
|
||||||
|
artefact, and the earlier table's `thinking off` column should be ignored.
|
||||||
|
|
||||||
|
The thinking-off variant was scored by a hand-rolled HTTP client living in the test file
|
||||||
|
instead of `llm.Client`. That copy did not send `repeat_penalty`, which the real router does
|
||||||
|
send (`routeRepeatPenalty = 1.15`). So the two columns differed on two axes at once, and the
|
||||||
|
one that mattered was the penalty, not the thinking mode.
|
||||||
|
|
||||||
|
Re-measured with everything else held equal — same fixture, same prompt, same grammar, same
|
||||||
|
sampling, same idle box, the three configurations run back to back and never concurrently:
|
||||||
|
|
||||||
|
| | llm-only, thinking on | llm-only, thinking off | cascade+llm |
|
||||||
|
|---|---|---|---|
|
||||||
|
| intent-only accuracy | 59.2% (45/76) | 59.2% (45/76) | 61.8% (47/76) |
|
||||||
|
| full accuracy (intent+slots+gate) | 38.2% (29/76) | 38.2% (29/76) | 57.9% (44/76) |
|
||||||
|
| route errors | 3 | 3 | 0 |
|
||||||
|
| grammar violations | 3 (all 3 route errors) | 3 (same 3 cases) | 0 |
|
||||||
|
| missed clarify | 5 / 6 | 5 / 6 | 5 / 6 |
|
||||||
|
| p50 latency | 836ms | 920ms | 810ms |
|
||||||
|
| p95 latency | 1.41s | 2.00s | 1.31s |
|
||||||
|
|
||||||
|
Thinking off is not just a tie on the headline numbers — it is identical case for case, with
|
||||||
|
the same confusion matrix and the same three unparseable replies. The latency difference is
|
||||||
|
run-to-run noise on one box, and it points the wrong way here.
|
||||||
|
|
||||||
|
The reason is simpler than any accuracy argument: **this llama-server build ignores the
|
||||||
|
request-level thinking switch for this model.** Probed directly against the running server
|
||||||
|
with `chat_template_kwargs.enable_thinking = false`, `chat_template_kwargs.thinking = false`
|
||||||
|
and top-level `reasoning_budget = 0` — all three return a byte-identical answer with the
|
||||||
|
thinking trace still in `reasoning_content`, and the server reports the prompt prefix as
|
||||||
|
cached, meaning the rendered template did not change. There was never anything being turned
|
||||||
|
off, which is also why the numbers match exactly.
|
||||||
|
|
||||||
|
Nothing was defaulted. `internal/llm` still has no `chat_template_kwargs` field, `VoiceConfig`
|
||||||
|
has no thinking flag, and `deploy/mavend.json` is unchanged. The misleading third
|
||||||
|
configuration is removed from `internal/router/eval` so the table it produced cannot be quoted
|
||||||
|
again.
|
||||||
|
|
||||||
|
Two caveats worth saying out loud:
|
||||||
|
|
||||||
|
- **The fixture is 76 cases.** A 6-point difference on 76 cases is roughly 4-5 cases and would
|
||||||
|
not have been worth trusting even if it had reproduced. This one was exactly 0 cases, which
|
||||||
|
is a much easier call.
|
||||||
|
- **This is one server build and one checkpoint** (`b9351`, Qwen3.5-0.8B Q4_K_M). If the
|
||||||
|
#122 checkpoint or a newer llama.cpp does honour the switch, the question reopens — but it
|
||||||
|
reopens as an unmeasured question, not as a 6-point win.
|
||||||
|
|
||||||
|
Phrasing was **not** measured. Whether thinking helps there is still open, and now also blocked
|
||||||
|
on the same "can we even turn it off" question.
|
||||||
|
|
||||||
|
## Clock and calendar rule — 31-07-2026 (Vikunja #374)
|
||||||
|
|
||||||
|
`routeSystem` never said whether "который час" or "какое число завтра" are `system` or
|
||||||
|
`query`, and `system→query ×4` showed up in every run. The rule added says: the clock and the
|
||||||
|
calendar date themselves are `system`; what is *written in* the calendar or in memory
|
||||||
|
("что у меня завтра", "какие есть напоминания") stays `query`; and a time named inside a
|
||||||
|
request ("напомни завтра…") is just a detail of the request, not a reason for `system`.
|
||||||
|
|
||||||
|
That split is not a preference. In `cmd/mavend/voice.go` only `replySystem` owns the clock and
|
||||||
|
the date formatter, so a clock question routed to `query` falls into the embedder + note RAG
|
||||||
|
and answers "не знаю". The agenda, on the other hand, is answered by `ParseCalendarDate` +
|
||||||
|
`CalendarEvents` *inside* the `query` branch, so that side has to stay `query`. The rule sits
|
||||||
|
above the question test because every one of these utterances carries a question word and a
|
||||||
|
later rule would never be reached.
|
||||||
|
|
||||||
|
The fixture is now 77 cases: one calendar-agenda case was added
|
||||||
|
(`ru-query-019` "что у меня стоит в календаре на послезавтра", intent `query`) specifically so
|
||||||
|
an over-broad system rule cannot pass unnoticed. The clock/date cases (`ru-sys-001/002/005`,
|
||||||
|
`en-sys-001`) already existed.
|
||||||
|
|
||||||
|
Three runs, same box, back to back, never concurrently:
|
||||||
|
|
||||||
|
| | baseline | first rule (too broad) | rule as committed |
|
||||||
|
|---|---|---|---|
|
||||||
|
| llm-only intent-only | 59.2% (45/76) | 54.5% (42/77) | 59.7% (46/77) |
|
||||||
|
| llm-only full | 38.2% | 35.1% | 39.0% |
|
||||||
|
| llm-only route errors | 3 | 4 | 5 |
|
||||||
|
| llm-only p50 | 1.09s | 0.91s | 0.93s |
|
||||||
|
| cascade+llm intent-only | 61.8% (47/76) | 58.4% | 62.3% (48/77) |
|
||||||
|
| cascade+llm full | 57.9% | 54.5% | 59.7% |
|
||||||
|
| cascade+llm route errors | 0 | 0 | 0 |
|
||||||
|
| cascade+llm p50 | 0.91s | 0.80s | 1.04s |
|
||||||
|
|
||||||
|
**The targeted bug is fixed and the headline number did not move.** `system→query ×4` is gone
|
||||||
|
in both LLM configurations — the `time` and `date` tags go from 0/2 and 0/2 to 2/2 and 2/2 —
|
||||||
|
but the model then over-applies the rule, and `query→system ×5` plus `reminder→system ×2`
|
||||||
|
appear where they did not exist before. Net accuracy is a wash, inside the noise of a 77-case
|
||||||
|
fixture.
|
||||||
|
|
||||||
|
The first attempt is shown because it is the honest history: it said "спрашивает время, дату
|
||||||
|
или день недели → system" with no scope, which swept up reminders, and it cost 3-5 points. It
|
||||||
|
was tightened once, on the reasoning that a rule capturing "напомни завтра в 7" is simply
|
||||||
|
wrong, and not tuned further. The remaining `query/reminder → system` over-trigger is a new,
|
||||||
|
separate weakness of the sub-1B model and deserves its own task rather than more prompt
|
||||||
|
kneading against a held-out fixture.
|
||||||
|
|
||||||
|
The rule is kept. It is correct about what the daemon can answer, and the failure it replaces
|
||||||
|
was silent ("не знаю" to "который час") while the one it introduces is loud.
|
||||||
|
|
||||||
## Findings
|
## Findings
|
||||||
|
|
||||||
### 1. The resident model does route better — 50.0% vs 36.8%
|
### 1. The resident model does route better — 50.0% vs 36.8%
|
||||||
@@ -145,11 +243,11 @@ Note the grammar's `string ::= "\"" ([^"\\] | "\\" .)* "\""` is unbounded, so no
|
|||||||
|
|
||||||
### 7. Two hypotheses tested and closed
|
### 7. Two hypotheses tested and closed
|
||||||
|
|
||||||
- **Thinking mode is a non-issue.** Qwen3.5's template defaults `thinking = 1`, so
|
- **Thinking mode is a non-issue.** Confirmed twice now, the second time properly — see the
|
||||||
grammar-constrained JSON lands in `reasoning_content` with `content` empty —
|
controlled re-run section. Grammar-constrained JSON lands in `reasoning_content` with
|
||||||
`llm.Client`'s fallback handles it. A `thinking off` run scored *identically* (18/76,
|
`content` empty and `llm.Client`'s fallback handles it; the request-level switch does
|
||||||
48.7%, same p50). `internal/llm` deliberately does **not** grow a `chat_template_kwargs`
|
nothing on this build. `internal/llm` deliberately does **not** grow a
|
||||||
field.
|
`chat_template_kwargs` field.
|
||||||
- **Runaway array repetition does not reproduce.** An isolated smoke test with a stripped
|
- **Runaway array repetition does not reproduce.** An isolated smoke test with a stripped
|
||||||
grammar emitted `{"intent":"reminder"}` until `MaxTokens`; under the real `routeSystem`
|
grammar emitted `{"intent":"reminder"}` until `MaxTokens`; under the real `routeSystem`
|
||||||
prompt the few-shot examples anchor it to one object. 2 errors in 76, not 76.
|
prompt the few-shot examples anchor it to one object. 2 errors in 76, not 76.
|
||||||
|
|||||||
@@ -0,0 +1,150 @@
|
|||||||
|
# Conversational phrasing eval — 31-07-2026
|
||||||
|
|
||||||
|
Every score measured tonight, on the three paths the nudge eval never touched:
|
||||||
|
chat, query-with-notes, and general knowledge.
|
||||||
|
|
||||||
|
**Short version: the plumbing got fixed and the score barely moved.** Grammar and
|
||||||
|
Russian prompts together took the composite from ~9 to ~14 of 27. Everything
|
||||||
|
still failing is the model not knowing things or not holding a constraint, and
|
||||||
|
prompting is out of levers. Settles the measurement half of Vikunja #395 / #398 /
|
||||||
|
#400.
|
||||||
|
|
||||||
|
## How to reproduce
|
||||||
|
|
||||||
|
```sh
|
||||||
|
# llama-server: -c 4096 -ngl 99 -t 6, model /mnt/hdd1/llms/qwen3.5/Qwen3.5-0.8B.Q4_K_M.gguf
|
||||||
|
MAVEN_LLM_URL=http://127.0.0.1:18099 no_proxy=127.0.0.1,localhost \
|
||||||
|
deps/go/go/bin/go test -count=1 -timeout 40m \
|
||||||
|
-run TestLLMTalkBaseline ./internal/phraser/eval/ -v
|
||||||
|
```
|
||||||
|
|
||||||
|
Three runs per configuration, always. The fixture is 27 cases, so one reply
|
||||||
|
changing moves the composite by 3.7 points — a single run cannot tell a real
|
||||||
|
change from sampling noise. This was learned the expensive way: an earlier claim
|
||||||
|
that "one nudge case fails every run" turned out to be three different cases
|
||||||
|
across three runs.
|
||||||
|
|
||||||
|
**Run the box otherwise idle.** See the contamination note at the bottom.
|
||||||
|
|
||||||
|
## Composite, per configuration
|
||||||
|
|
||||||
|
| config | overall /27 | chat /9 | query /9 | knowledge /9 | canned fallbacks |
|
||||||
|
|---|---|---|---|---|---|
|
||||||
|
| baseline, no grammar | 7, 12, 7 | 1, 1, 0 | 2, 4, 2 | 4, 7, 5 | 0, 0, 0 |
|
||||||
|
| + GBNF grammar (#398) | 14, 15, 8 | 1, 3, 0 | 5, 6, 3 | 8, 6, 5 | 0, 0, 0 |
|
||||||
|
| + Russian prompts (#400) | 11, 17, 15 | 1, 5, 3 | 5, 6, 8 | 5, 6, 4 | 0, 0, 0 |
|
||||||
|
| + truncation fix, 1000ch/768tok | 12, 13, 10 | 2, 2, 1 | 7, 7, 5 | 3, 4, 4 | 3, 3, 6 |
|
||||||
|
| + rebalanced, 600ch/1024tok | **void — contaminated** | | | | |
|
||||||
|
|
||||||
|
"Canned fallbacks" counts replies that came back as the hardcoded `"не знаю."`
|
||||||
|
or `"поговорили."`. It is not a check, it is a health signal: those strings mean
|
||||||
|
the phraser gave up, and the eval scores them as ordinary bad replies.
|
||||||
|
|
||||||
|
## Per-check
|
||||||
|
|
||||||
|
| check | no grammar | + grammar | + RU prompts | + truncation fix |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| nonempty | 27, 27, 27 | 27, 27, 27 | 27, 27, 27 | 27, 27, 27 |
|
||||||
|
| ellipsis | 20, 19, 23 | 27, 27, 27 | 27, 27, 27 | 27, 27, 27 |
|
||||||
|
| lang | 13, 16, 15 | 23, 26, 26 | 25, 26, 25 | 26, 27, 27 |
|
||||||
|
| feminine | — | — | 25, 24, 26 | 25, 25, 27 |
|
||||||
|
| address | — | — | 21, 22, 22 | 22, 21, 22 |
|
||||||
|
| ontopic | — | — | 17, 24, 18 | 17, 19, 14 |
|
||||||
|
|
||||||
|
`nonempty` reading 27/27 everywhere is not good news — it was a broken check.
|
||||||
|
It tested for a non-blank string, so replies of literally `{` and `"15-16"`
|
||||||
|
passed it. Fixed on `overnight/fix-truncation`; it needs a letter now.
|
||||||
|
|
||||||
|
## What each change actually bought
|
||||||
|
|
||||||
|
**GBNF grammar (#398) — the biggest single win.** Qwen3.5-0.8B writes
|
||||||
|
`Thinking Process:` as plain text with no tags, `stripThink` only handles
|
||||||
|
`</think>`, so the JSON never closed and the plain-text fallback shipped the
|
||||||
|
literal reasoning. `ellipsis` went 20→27 and `lang` 13→26. The router had been
|
||||||
|
using a grammar for ages; the phraser asking nicely in the prompt was the
|
||||||
|
oversight.
|
||||||
|
|
||||||
|
**Russian prompts (#400) — modest, plus a large latency win.** Chat 1.3→3.0
|
||||||
|
average, query 4.7→6.3, knowledge 6.3→5.0. All inside the run-to-run spread, so
|
||||||
|
"probably better on the paths it targeted, not provable in three runs". p50
|
||||||
|
latency dropped from ~11.5s to ~2.3s and that part is consistent across all
|
||||||
|
three runs — shorter prompts, and she stopped emitting English reasoning first.
|
||||||
|
|
||||||
|
**Truncation fix — necessary, and did not help the score.** Two real bugs
|
||||||
|
(replies of `{`, and a `nonempty` check that passed them), both fixed, and the
|
||||||
|
composite went nowhere. A complete rambling wrong answer fails the same checks a
|
||||||
|
truncated one did. Worth doing anyway: the daemon was shipping `{` to a
|
||||||
|
text-to-speech voice.
|
||||||
|
|
||||||
|
## The truncation bug, since the cause was counter-intuitive
|
||||||
|
|
||||||
|
The grammar's `string ::= ... {0,400}` rule was the cause, not the token cap.
|
||||||
|
Measured against Qwen3.5-0.8B at three caps — 256, 768 and 2048 — the reply came
|
||||||
|
back **exactly 400 characters every time, cut mid-word** (`"Нужно записать и,"`).
|
||||||
|
|
||||||
|
Then I raised the bound to 1000 while the cap was 768 tokens and made it worse:
|
||||||
|
Russian runs ~1.5 characters per token here, so generation died on the *token*
|
||||||
|
cap instead, mid-object, and the new guard correctly refused it and shipped
|
||||||
|
`"не знаю."` — 3, 3 and 6 fallbacks per run, from zero. **The two limits have to
|
||||||
|
agree.** 600 characters needs ~400 tokens; the cap is 1024.
|
||||||
|
|
||||||
|
## Where the remaining failures live
|
||||||
|
|
||||||
|
`address` is stuck at 21-22 of 27 and `ontopic` at 14-19. Both resist prompting.
|
||||||
|
|
||||||
|
**The prompt now explicitly forbids exactly what she does.** It says never "вы",
|
||||||
|
use the singular — and she writes `вашей`, `подождите`, `делаете`, `хотите`,
|
||||||
|
`напишите`. Telling a 0.8B "never do X" does not work. Same for
|
||||||
|
`feminine`: `я готов`, `я понял`, `я нашел`, `я заметил`, `я сказал`.
|
||||||
|
|
||||||
|
**Some of `ontopic` is the fixture, not the model.** `chat-how-are-you` got
|
||||||
|
`"Привет! Я здесь, чтобы поговорить. Как дела сегодня?"` — a fine reply that
|
||||||
|
fails because `want_any` is `[норм, хорош, порядк, тут, работ]`. It fails in
|
||||||
|
every run, so it inflates the count. The `ontopic` column currently measures the
|
||||||
|
fixture as much as the model. Not fixed yet, deliberately: changing it would
|
||||||
|
break comparability with the runs above.
|
||||||
|
|
||||||
|
**Two replies worth reading, because they are not fixable by prompting:**
|
||||||
|
|
||||||
|
- Thunder and lightning: *"Скорость молнии — 8-10 тысяч километров в секунду, но
|
||||||
|
звук — 300 метров в секунду, что делает молнию громче."* Confidently wrong,
|
||||||
|
and it concludes lightning is *louder* rather than sound being *slower*.
|
||||||
|
- "расскажи обо мне": *"Ты — прекрасное существо, с душой и вниманием… Спасибо за
|
||||||
|
твою улыбку… О тебе — заповедь любви."* Sycophantic filler, zero information,
|
||||||
|
and precisely the "not a relationship" non-goal.
|
||||||
|
- Boiling an egg: `"15-16"` one run, `"1"` another. No unit, wrong number.
|
||||||
|
|
||||||
|
The first argues for reading instead of recalling (#403 — Kiwix retrieval scores
|
||||||
|
8/8 on the same questions given English keywords). The second and third argue
|
||||||
|
for templates on the paths where correctness matters (#392).
|
||||||
|
|
||||||
|
## Contamination note — how the last row got voided
|
||||||
|
|
||||||
|
I started the query-rewrite agent against the same llama-server the sweep was
|
||||||
|
using, and assumed contention would only affect latency. It did not. The
|
||||||
|
knowledge path collapsed to 0 of 9 with eight canned `"не знаю."` replies, p95
|
||||||
|
tripled to 23.7s, and **the report still said "0 errors"**.
|
||||||
|
|
||||||
|
That is Vikunja #397, and it is worse than filed: a merely *busy* server
|
||||||
|
produces a clean-looking report with a third of the fixture silently answering
|
||||||
|
`"не знаю."`. `PhraseChat` and `PhraseQuery` swallow every failure and return a
|
||||||
|
hardcoded string, so infrastructure trouble is indistinguishable from bad
|
||||||
|
phrasing in the score. The talk test guards the *start* and *end* of a run with
|
||||||
|
a model check, which catches a dead server but not a loaded one.
|
||||||
|
|
||||||
|
**Until #397 is fixed, treat any run made on a busy box as void.**
|
||||||
|
|
||||||
|
## Next
|
||||||
|
|
||||||
|
- Re-run 600ch/1024tok clean, to fill the void row.
|
||||||
|
- Score `Qwen3.5-2B-UD-Q4_K_XL` (already at `/mnt/hdd1/llms/qwen3.5/`, never
|
||||||
|
measured) on this fixture and the router fixture. Not the 4B — too big for
|
||||||
|
this box, owner's call.
|
||||||
|
- Newer sub-500M candidates (LFM2.5 200M/300M) are worth a run for routing.
|
||||||
|
Note `MODEL-BAKEOFF-31-07-2026.md` found LFM2.5-**1.2B** worse than
|
||||||
|
Qwen3.5-0.8B at Russian routing and 2.4× slower — but those are a different,
|
||||||
|
older generation, so that result does not predict the small ones.
|
||||||
|
- Fix `chat-how-are-you`'s `want_any`, and re-baseline once, so `ontopic`
|
||||||
|
measures the model.
|
||||||
|
- #397 first if anything, since it decides whether any of the above is
|
||||||
|
trustworthy.
|
||||||
+131
-19
@@ -3,6 +3,8 @@ package main
|
|||||||
import (
|
import (
|
||||||
"context"
|
"context"
|
||||||
"log"
|
"log"
|
||||||
|
"math/rand"
|
||||||
|
"strings"
|
||||||
"time"
|
"time"
|
||||||
|
|
||||||
"github.com/kami/maven/internal/dialogue"
|
"github.com/kami/maven/internal/dialogue"
|
||||||
@@ -21,8 +23,12 @@ const clarifyTTL = 90 * time.Second
|
|||||||
// raw utterance, chat and system have nothing to fill in. For those a clarify
|
// raw utterance, chat and system have nothing to fill in. For those a clarify
|
||||||
// decision keeps the canned "не поняла" reply — inventing a question for noise
|
// decision keeps the canned "не поняла" reply — inventing a question for noise
|
||||||
// is worse than admitting she missed it.
|
// is worse than admitting she missed it.
|
||||||
|
// A reminder wants BOTH what to remind about and when. Subject first: "напомни
|
||||||
|
// в 11" has a time and nothing to say at 11, and a reminder with no subject is
|
||||||
|
// not worth setting. Order here is the order she asks in — she still only asks
|
||||||
|
// about the first one missing.
|
||||||
var wantedSlots = map[router.Intent][]dialogue.Slot{
|
var wantedSlots = map[router.Intent][]dialogue.Slot{
|
||||||
router.IntentReminder: {dialogue.SlotTime},
|
router.IntentReminder: {dialogue.SlotText, dialogue.SlotTime},
|
||||||
router.IntentFact: {dialogue.SlotKey},
|
router.IntentFact: {dialogue.SlotKey},
|
||||||
router.IntentAct: {dialogue.SlotFn},
|
router.IntentAct: {dialogue.SlotFn},
|
||||||
}
|
}
|
||||||
@@ -35,14 +41,94 @@ var wantedSlots = map[router.Intent][]dialogue.Slot{
|
|||||||
// questions, so there is no gender agreement to get wrong; the feminine
|
// questions, so there is no gender agreement to get wrong; the feminine
|
||||||
// self-reference lives in the reply she gives when she drops the request.
|
// self-reference lives in the reply she gives when she drops the request.
|
||||||
var clarifyQuestions = map[dialogue.Slot]string{
|
var clarifyQuestions = map[dialogue.Slot]string{
|
||||||
dialogue.SlotTime: "На когда напомнить?",
|
dialogue.SlotTime: "Когда?",
|
||||||
|
dialogue.SlotText: "О чём напомнить?",
|
||||||
dialogue.SlotKey: "Что записать?",
|
dialogue.SlotKey: "Что записать?",
|
||||||
dialogue.SlotFn: "Что сделать?",
|
dialogue.SlotFn: "Что сделать?",
|
||||||
}
|
}
|
||||||
|
|
||||||
// clarifyDropped — she asked once, the answer still did not fill the gap, so
|
// clarifyGaveUp — she is out of questions and still does not have the slot. She
|
||||||
// the request is gone. Said plainly, once, with no second question.
|
// says so out loud: dropping the request in silence would leave him thinking it
|
||||||
const clarifyDropped = "Не разобрала — скажи целиком, пожалуйста."
|
// landed. Feminine self-reference ("поняла"), as everywhere.
|
||||||
|
const clarifyGaveUp = "Прости, я не поняла. Скажи, пожалуйста, по-другому."
|
||||||
|
|
||||||
|
// clarifyExpiredVariants — his answer came after the TTL, so the parked request
|
||||||
|
// is already gone. Same tone as clarifyGaveUp, different reason: too much time
|
||||||
|
// passed, not "I did not understand". Feminine self-reference ("ждала",
|
||||||
|
// "отпустила"); he is addressed with a plain imperative.
|
||||||
|
//
|
||||||
|
// Five phrasings, not one. This is the line he hears whenever he walks off
|
||||||
|
// mid-request, so it is the line that repeats most — and the same sentence every
|
||||||
|
// time is what makes a house assistant sound like a kiosk. They all carry the
|
||||||
|
// same two facts (the old request is gone; say it again if it still matters),
|
||||||
|
// because the wording may vary and the meaning may not.
|
||||||
|
//
|
||||||
|
// Fixed templates rather than model output, for the same reason as
|
||||||
|
// clarifyQuestions: this text has to be right every time, and it is not worth a
|
||||||
|
// generation to say something this small.
|
||||||
|
var clarifyExpiredVariants = []string{
|
||||||
|
"Прости, я слишком долго ждала ответа и отпустила прошлую просьбу. Если она ещё нужна, скажи заново.",
|
||||||
|
"Кажется, прошлая просьба уже не важна — я её отпустила. Если я ошибаюсь, повтори.",
|
||||||
|
"Ты как-то резко замолчал, и я не стала ждать дальше. Если та просьба ещё нужна, скажи заново.",
|
||||||
|
"Я не дождалась ответа и убрала прошлую просьбу. Повтори, если она всё ещё нужна.",
|
||||||
|
"Столько времени прошло, что я отпустила прошлую просьбу. Скажи заново, если она в силе.",
|
||||||
|
}
|
||||||
|
|
||||||
|
// clarifyExpiredLine picks one of them at random.
|
||||||
|
func clarifyExpiredLine() string {
|
||||||
|
return clarifyExpiredVariants[rand.Intn(len(clarifyExpiredVariants))]
|
||||||
|
}
|
||||||
|
|
||||||
|
// isClarifyExpired reports whether s opens with any of the expiry lines. The
|
||||||
|
// notice is glued in front of this turn's reply (see withNotice), so a caller
|
||||||
|
// checking for it has to match a prefix, not the whole string.
|
||||||
|
func isClarifyExpired(s string) bool {
|
||||||
|
for _, v := range clarifyExpiredVariants {
|
||||||
|
if strings.HasPrefix(s, v) {
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
|
||||||
|
// trimClarifyExpired strips a leading expiry notice, leaving this turn's actual
|
||||||
|
// reply. "" ⇒ the notice was the whole thing.
|
||||||
|
func trimClarifyExpired(s string) string {
|
||||||
|
for _, v := range clarifyExpiredVariants {
|
||||||
|
if strings.HasPrefix(s, v) {
|
||||||
|
return strings.TrimSpace(strings.TrimPrefix(s, v))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return strings.TrimSpace(s)
|
||||||
|
}
|
||||||
|
|
||||||
|
// clarifyExpiredNotice returns that line when a parked question had just timed
|
||||||
|
// out, and "" when nothing was parked. Call it right after
|
||||||
|
// resolveClarifyAnswer: a live question is answered there, an expired one is
|
||||||
|
// only reported here — the words themselves still go on to be routed fresh.
|
||||||
|
func (h *reactiveHandler) clarifyExpiredNotice() string {
|
||||||
|
if h.clarifyStore == nil {
|
||||||
|
return ""
|
||||||
|
}
|
||||||
|
if !h.clarifyStore.TakeExpired(voiceDialogueID, h.now()) {
|
||||||
|
return ""
|
||||||
|
}
|
||||||
|
log.Printf("voice: clarify — parked question expired, telling him and routing the words fresh")
|
||||||
|
return clarifyExpiredLine()
|
||||||
|
}
|
||||||
|
|
||||||
|
// withNotice glues the expiry notice in front of this turn's reply. One turn
|
||||||
|
// carries one reply on the wire, so the notice cannot be a message of its own —
|
||||||
|
// but neither the notice nor the fresh answer may be dropped.
|
||||||
|
func withNotice(notice, reply string) string {
|
||||||
|
if notice == "" {
|
||||||
|
return reply
|
||||||
|
}
|
||||||
|
if reply == "" {
|
||||||
|
return notice
|
||||||
|
}
|
||||||
|
return notice + " " + reply
|
||||||
|
}
|
||||||
|
|
||||||
// missingFor returns the slots a decision still needs, most important first.
|
// missingFor returns the slots a decision still needs, most important first.
|
||||||
// Empty ⇒ there is nothing identifiable to ask about.
|
// Empty ⇒ there is nothing identifiable to ask about.
|
||||||
@@ -79,13 +165,14 @@ func (h *reactiveHandler) askClarify(dec router.Decision) (string, bool) {
|
|||||||
return "", false
|
return "", false
|
||||||
}
|
}
|
||||||
h.clarifyStore.Put(voiceDialogueID, &dialogue.PendingQuestion{
|
h.clarifyStore.Put(voiceDialogueID, &dialogue.PendingQuestion{
|
||||||
Intent: dialogue.Intent(dec.Intent),
|
Intent: dialogue.Intent(dec.Intent),
|
||||||
Slots: toDialogueSlots(dec.Slots),
|
Slots: toDialogueSlots(dec.Slots),
|
||||||
Missing: []dialogue.Slot{slot},
|
Missing: []dialogue.Slot{slot},
|
||||||
Utterance: dec.Utterance,
|
Utterance: dec.Utterance,
|
||||||
Asked: h.now(),
|
Asked: h.now(),
|
||||||
TTL: clarifyTTL,
|
TTL: clarifyTTL,
|
||||||
Attempts: 1, // asked once; MaxAttempts is 1, so there is no second ask
|
Attempts: 1, // this ask
|
||||||
|
MaxAttempts: h.clarifyMaxAttempts,
|
||||||
})
|
})
|
||||||
log.Printf("voice: clarify — asked about %s for intent=%s", slot, dec.Intent)
|
log.Printf("voice: clarify — asked about %s for intent=%s", slot, dec.Intent)
|
||||||
return question, true
|
return question, true
|
||||||
@@ -97,8 +184,9 @@ func (h *reactiveHandler) askClarify(dec router.Decision) (string, bool) {
|
|||||||
// resolveConfirm and checked in the same place.
|
// resolveConfirm and checked in the same place.
|
||||||
//
|
//
|
||||||
// The answer is parsed with the same extractor the router uses, for the intent
|
// The answer is parsed with the same extractor the router uses, for the intent
|
||||||
// she parked — no second parser. If it still does not fill the gap the request
|
// she parked — no second parser. If it still does not fill the gap she asks
|
||||||
// is dropped: she does not ask again.
|
// again, up to MaxAttempts; after that she says out loud that she did not
|
||||||
|
// understand. She never drops the request in silence.
|
||||||
func (h *reactiveHandler) resolveClarifyAnswer(ctx context.Context, text string) (string, bool) {
|
func (h *reactiveHandler) resolveClarifyAnswer(ctx context.Context, text string) (string, bool) {
|
||||||
if h.clarifyStore == nil {
|
if h.clarifyStore == nil {
|
||||||
return "", false
|
return "", false
|
||||||
@@ -107,17 +195,14 @@ func (h *reactiveHandler) resolveClarifyAnswer(ctx context.Context, text string)
|
|||||||
if q == nil {
|
if q == nil {
|
||||||
return "", false
|
return "", false
|
||||||
}
|
}
|
||||||
// One shot either way: the question is consumed whether or not the answer
|
|
||||||
// works, so a failed answer can't leave the question armed.
|
|
||||||
h.clarifyStore.Delete(voiceDialogueID)
|
|
||||||
|
|
||||||
intent := router.Intent(q.Intent)
|
intent := router.Intent(q.Intent)
|
||||||
answer := h.extractor.Extract(ctx, intent, text, h.now())
|
answer := h.extractor.Extract(ctx, intent, text, h.now())
|
||||||
merged := q.Answer(text, toDialogueSlots(answer))
|
merged := q.Answer(text, toDialogueSlots(answer))
|
||||||
if len(dialogue.StillMissing(q.Missing, merged)) > 0 {
|
if len(dialogue.StillMissing(q.Missing, merged)) > 0 {
|
||||||
log.Printf("voice: clarify — answer %q did not fill %v, dropping", text, q.Missing)
|
return h.reaskOrGiveUp(q, merged, text), true
|
||||||
return clarifyDropped, true
|
|
||||||
}
|
}
|
||||||
|
h.clarifyStore.Delete(voiceDialogueID)
|
||||||
|
|
||||||
// Rebuild the decision as if it had routed cleanly, then run it down the
|
// Rebuild the decision as if it had routed cleanly, then run it down the
|
||||||
// normal path. Clarify is deliberately false and the intent is unchanged:
|
// normal path. Clarify is deliberately false and the intent is unchanged:
|
||||||
@@ -133,6 +218,29 @@ func (h *reactiveHandler) resolveClarifyAnswer(ctx context.Context, text string)
|
|||||||
return h.finishClarified(ctx, dec), true
|
return h.finishClarified(ctx, dec), true
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// reaskOrGiveUp handles an answer that left the gap open: ask the same question
|
||||||
|
// again while she has attempts left, otherwise say she did not understand and
|
||||||
|
// let the request go. Never returns "" — a mute give-up reads as "done".
|
||||||
|
func (h *reactiveHandler) reaskOrGiveUp(q *dialogue.PendingQuestion, merged dialogue.Slots, text string) string {
|
||||||
|
question := ""
|
||||||
|
if len(q.Missing) > 0 {
|
||||||
|
question = clarifyQuestions[q.Missing[0]]
|
||||||
|
}
|
||||||
|
if question == "" || !q.CanAsk() {
|
||||||
|
h.clarifyStore.Delete(voiceDialogueID)
|
||||||
|
log.Printf("voice: clarify — gave up on %v after %d question(s), answer was %q", q.Missing, q.Attempts, text)
|
||||||
|
return clarifyGaveUp
|
||||||
|
}
|
||||||
|
// Re-park with whatever the answer DID give, the clock restarted and one
|
||||||
|
// more question spent.
|
||||||
|
q.Slots = merged
|
||||||
|
q.Attempts++
|
||||||
|
q.Asked = h.now()
|
||||||
|
h.clarifyStore.Put(voiceDialogueID, q)
|
||||||
|
log.Printf("voice: clarify — answer %q did not fill %v, asking again (attempt %d)", text, q.Missing, q.Attempts)
|
||||||
|
return question
|
||||||
|
}
|
||||||
|
|
||||||
// finishClarified runs a completed decision through the same steps a freshly
|
// finishClarified runs a completed decision through the same steps a freshly
|
||||||
// routed one takes: remember the turn, act, then phrase.
|
// routed one takes: remember the turn, act, then phrase.
|
||||||
func (h *reactiveHandler) finishClarified(ctx context.Context, dec router.Decision) string {
|
func (h *reactiveHandler) finishClarified(ctx context.Context, dec router.Decision) string {
|
||||||
@@ -146,6 +254,10 @@ func (h *reactiveHandler) finishClarified(ctx context.Context, dec router.Decisi
|
|||||||
if reply == "" {
|
if reply == "" {
|
||||||
reply = h.replier.Reply(dec)
|
reply = h.replier.Reply(dec)
|
||||||
}
|
}
|
||||||
|
if reply == "" {
|
||||||
|
// Belt: an empty reply here would be a silent drop.
|
||||||
|
reply = clarifyGaveUp
|
||||||
|
}
|
||||||
return reply
|
return reply
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
+108
-14
@@ -56,10 +56,13 @@ func TestClarifyQuestionForMissingSlot(t *testing.T) {
|
|||||||
want string
|
want string
|
||||||
asked bool
|
asked bool
|
||||||
}{
|
}{
|
||||||
{"reminder without a time", clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме"), "На когда напомнить?", true},
|
{"reminder without a time", clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме"), "Когда?", true},
|
||||||
{"fact without a key", clarifyDec(router.IntentFact, router.Slots{Text: "запиши"}, "запиши"), "Что записать?", true},
|
{"fact without a key", clarifyDec(router.IntentFact, router.Slots{Text: "запиши"}, "запиши"), "Что записать?", true},
|
||||||
{"act without a fn", clarifyDec(router.IntentAct, router.Slots{Text: "сделай это"}, "сделай это"), "Что сделать?", true},
|
{"act without a fn", clarifyDec(router.IntentAct, router.Slots{Text: "сделай это"}, "сделай это"), "Что сделать?", true},
|
||||||
{"reminder that already has a time", clarifyDec(router.IntentReminder, router.Slots{HasTime: true}, "напомни в 11"), "", false},
|
// A time with nothing to say at that time is still half a reminder, so
|
||||||
|
// the subject is what she asks about — not silence.
|
||||||
|
{"reminder that has a time but no subject", clarifyDec(router.IntentReminder, router.Slots{HasTime: true}, "напомни в 11"), "О чём напомнить?", true},
|
||||||
|
{"reminder that has both", clarifyDec(router.IntentReminder, router.Slots{Text: "позвонить маме", HasTime: true}, "напомни в 11 позвонить маме"), "", false},
|
||||||
{"chat is never worth a question", clarifyDec(router.IntentChat, router.Slots{Text: "мгм"}, "мгм"), "", false},
|
{"chat is never worth a question", clarifyDec(router.IntentChat, router.Slots{Text: "мгм"}, "мгм"), "", false},
|
||||||
{"query is never worth a question", clarifyDec(router.IntentQuery, router.Slots{Text: "а"}, "а"), "", false},
|
{"query is never worth a question", clarifyDec(router.IntentQuery, router.Slots{Text: "а"}, "а"), "", false},
|
||||||
}
|
}
|
||||||
@@ -78,7 +81,7 @@ func TestClarifyReminderCompletesOnAnswer(t *testing.T) {
|
|||||||
h, st, _ := newClarifyHandler(t)
|
h, st, _ := newClarifyHandler(t)
|
||||||
|
|
||||||
question, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме"))
|
question, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме"))
|
||||||
if !asked || question != "На когда напомнить?" {
|
if !asked || question != "Когда?" {
|
||||||
t.Fatalf("expected the time question, got %q asked=%v", question, asked)
|
t.Fatalf("expected the time question, got %q asked=%v", question, asked)
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -86,7 +89,7 @@ func TestClarifyReminderCompletesOnAnswer(t *testing.T) {
|
|||||||
if !handled {
|
if !handled {
|
||||||
t.Fatal("the answer to an open question must be consumed as an answer")
|
t.Fatal("the answer to an open question must be consumed as an answer")
|
||||||
}
|
}
|
||||||
if reply == clarifyDropped {
|
if reply == clarifyGaveUp {
|
||||||
t.Fatalf("a good answer must not drop the request: %q", reply)
|
t.Fatalf("a good answer must not drop the request: %q", reply)
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -111,7 +114,7 @@ func TestClarifyFactCompletesOnAnswer(t *testing.T) {
|
|||||||
if _, asked := h.askClarify(clarifyDec(router.IntentFact, router.Slots{Text: "запиши"}, "запиши")); !asked {
|
if _, asked := h.askClarify(clarifyDec(router.IntentFact, router.Slots{Text: "запиши"}, "запиши")); !asked {
|
||||||
t.Fatal("a fact with no key should be asked about")
|
t.Fatal("a fact with no key should be asked about")
|
||||||
}
|
}
|
||||||
if reply, handled := h.resolveClarifyAnswer(ctx, "пил воду"); !handled || reply == clarifyDropped {
|
if reply, handled := h.resolveClarifyAnswer(ctx, "пил воду"); !handled || reply == clarifyGaveUp {
|
||||||
t.Fatalf("answer should complete the fact, handled=%v reply=%q", handled, reply)
|
t.Fatalf("answer should complete the fact, handled=%v reply=%q", handled, reply)
|
||||||
}
|
}
|
||||||
if fact, err := st.LatestFact(ctx, "water"); err != nil || fact.Key != "water" {
|
if fact, err := st.LatestFact(ctx, "water"); err != nil || fact.Key != "water" {
|
||||||
@@ -137,26 +140,87 @@ func TestClarifyAnswerAfterTTLIsANewRequest(t *testing.T) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// TestClarifyUnclearAnswerDropsWithoutAskingAgain — MaxAttempts is 1.
|
// TestClarifyAsksThreeTimesThenSaysSo — three questions are allowed, the fourth
|
||||||
func TestClarifyUnclearAnswerDropsWithoutAskingAgain(t *testing.T) {
|
// is not, and running out is SPOKEN. Silence would read as "handled".
|
||||||
|
func TestClarifyAsksThreeTimesThenSaysSo(t *testing.T) {
|
||||||
ctx := context.Background()
|
ctx := context.Background()
|
||||||
h, st, _ := newClarifyHandler(t)
|
h, st, _ := newClarifyHandler(t)
|
||||||
|
|
||||||
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
|
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
|
||||||
t.Fatal("expected a question")
|
t.Fatal("expected a first question")
|
||||||
}
|
}
|
||||||
|
// Two more unclear answers ⇒ two more questions (3 asks in total).
|
||||||
|
for i := 2; i <= 3; i++ {
|
||||||
|
reply, handled := h.resolveClarifyAnswer(ctx, "ну не знаю")
|
||||||
|
if !handled {
|
||||||
|
t.Fatalf("answer %d must be consumed as an answer", i)
|
||||||
|
}
|
||||||
|
if reply != "Когда?" {
|
||||||
|
t.Fatalf("attempt %d should ask again, got %q", i, reply)
|
||||||
|
}
|
||||||
|
if h.clarifyStore.Get(voiceDialogueID, h.now()) == nil {
|
||||||
|
t.Fatalf("attempt %d must leave the question armed", i)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
reply, handled := h.resolveClarifyAnswer(ctx, "ну не знаю")
|
reply, handled := h.resolveClarifyAnswer(ctx, "ну не знаю")
|
||||||
if !handled || reply != clarifyDropped {
|
if !handled || reply != clarifyGaveUp {
|
||||||
t.Fatalf("an unclear answer should drop the request, handled=%v reply=%q", handled, reply)
|
t.Fatalf("the fourth try must give up out loud, handled=%v reply=%q", handled, reply)
|
||||||
}
|
}
|
||||||
if strings.Contains(reply, "?") {
|
if reply == "" || strings.Contains(reply, "?") {
|
||||||
t.Fatalf("she must not ask a second question: %q", reply)
|
t.Fatalf("giving up must be spoken and must not be another question: %q", reply)
|
||||||
}
|
}
|
||||||
if h.clarifyStore.Get(voiceDialogueID, h.now()) != nil {
|
if h.clarifyStore.Get(voiceDialogueID, h.now()) != nil {
|
||||||
t.Fatal("a dropped request must leave no armed question")
|
t.Fatal("a given-up request must leave no armed question")
|
||||||
}
|
}
|
||||||
if reminders, err := st.DueReminders(ctx, h.now().Add(48*time.Hour)); err != nil || len(reminders) != 0 {
|
if reminders, err := st.DueReminders(ctx, h.now().Add(48*time.Hour)); err != nil || len(reminders) != 0 {
|
||||||
t.Fatalf("a dropped request must not create anything: reminders=%v err=%v", reminders, err)
|
t.Fatalf("a given-up request must not create anything: reminders=%v err=%v", reminders, err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestClarifyMaxAttemptsIsConfigurable — one question when the config says one.
|
||||||
|
func TestClarifyMaxAttemptsIsConfigurable(t *testing.T) {
|
||||||
|
ctx := context.Background()
|
||||||
|
h, _, _ := newClarifyHandler(t)
|
||||||
|
h.clarifyMaxAttempts = 1
|
||||||
|
|
||||||
|
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
|
||||||
|
t.Fatal("expected a question")
|
||||||
|
}
|
||||||
|
if reply, handled := h.resolveClarifyAnswer(ctx, "ну не знаю"); !handled || reply != clarifyGaveUp {
|
||||||
|
t.Fatalf("with max 1 she must give up at once, handled=%v reply=%q", handled, reply)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestClarifyRestatedAnswerWins — «в 11:00», then «нет, в 15:00». The second
|
||||||
|
// value is the one that lands.
|
||||||
|
func TestClarifyRestatedAnswerWins(t *testing.T) {
|
||||||
|
ctx := context.Background()
|
||||||
|
h, st, _ := newClarifyHandler(t)
|
||||||
|
|
||||||
|
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме")); !asked {
|
||||||
|
t.Fatal("expected a question")
|
||||||
|
}
|
||||||
|
// First answer parses, but re-park it by hand as if she had asked again:
|
||||||
|
// what matters here is that Answer prefers the newer value over the parked
|
||||||
|
// one, which is the case the daemon hits on a re-ask.
|
||||||
|
q := h.clarifyStore.Get(voiceDialogueID, h.now())
|
||||||
|
if q == nil {
|
||||||
|
t.Fatal("expected an armed question")
|
||||||
|
}
|
||||||
|
first := h.extractor.Extract(ctx, router.IntentReminder, "в 11:00", h.now())
|
||||||
|
q.Slots = q.Answer("в 11:00", toDialogueSlots(first))
|
||||||
|
|
||||||
|
if reply, handled := h.resolveClarifyAnswer(ctx, "нет, в 15:00"); !handled || reply == clarifyGaveUp {
|
||||||
|
t.Fatalf("the restated answer should complete the request, handled=%v reply=%q", handled, reply)
|
||||||
|
}
|
||||||
|
reminders, err := st.DueReminders(ctx, h.now().Add(48*time.Hour))
|
||||||
|
if err != nil || len(reminders) != 1 {
|
||||||
|
t.Fatalf("expected one reminder: %v err=%v", reminders, err)
|
||||||
|
}
|
||||||
|
want := h.extractor.Extract(ctx, router.IntentReminder, "в 15:00", h.now())
|
||||||
|
if !reminders[0].FireTs.Equal(want.Time) {
|
||||||
|
t.Fatalf("reminder at %v, want the restated %v", reminders[0].FireTs, want.Time)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -228,6 +292,36 @@ func TestNoQuestionWhenNothingIsMissing(t *testing.T) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// TestClarifyExpiryIsAnnouncedAndWordsStillRoute — his answer lands after the
|
||||||
|
// TTL: she must say the old request is gone AND still answer the new words.
|
||||||
|
func TestClarifyExpiryIsAnnouncedAndWordsStillRoute(t *testing.T) {
|
||||||
|
ctx := context.Background()
|
||||||
|
h, _, now := newClarifyHandler(t)
|
||||||
|
emb := router.NewHashEmbedder(1024)
|
||||||
|
h.embedder = emb
|
||||||
|
h.router = buildRouter(emb, h.matcher, 0.55, nil)
|
||||||
|
|
||||||
|
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
|
||||||
|
t.Fatal("expected a question")
|
||||||
|
}
|
||||||
|
*now = now.Add(clarifyTTL + time.Second)
|
||||||
|
|
||||||
|
reply := h.handleText(ctx, "как дела")
|
||||||
|
if !isClarifyExpired(reply) {
|
||||||
|
t.Fatalf("expired question must be announced first, got %q", reply)
|
||||||
|
}
|
||||||
|
if trimClarifyExpired(reply) == "" {
|
||||||
|
t.Fatalf("the new words must still be answered, got only the notice: %q", reply)
|
||||||
|
}
|
||||||
|
if h.clarifyStore.Get(voiceDialogueID, h.now()) != nil {
|
||||||
|
t.Fatal("the expired question must be gone")
|
||||||
|
}
|
||||||
|
// The notice is said once, not on every later utterance.
|
||||||
|
if reply := h.handleText(ctx, "как дела"); isClarifyExpired(reply) {
|
||||||
|
t.Fatalf("notice repeated on a later turn: %q", reply)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
// TestNoPendingQuestionFallsThrough — with nothing parked, an utterance routes
|
// TestNoPendingQuestionFallsThrough — with nothing parked, an utterance routes
|
||||||
// normally.
|
// normally.
|
||||||
func TestNoPendingQuestionFallsThrough(t *testing.T) {
|
func TestNoPendingQuestionFallsThrough(t *testing.T) {
|
||||||
|
|||||||
+43
-21
@@ -58,6 +58,7 @@ import (
|
|||||||
"github.com/kami/maven/internal/delivery/telegramsink"
|
"github.com/kami/maven/internal/delivery/telegramsink"
|
||||||
"github.com/kami/maven/internal/ipc"
|
"github.com/kami/maven/internal/ipc"
|
||||||
"github.com/kami/maven/internal/loop"
|
"github.com/kami/maven/internal/loop"
|
||||||
|
"github.com/kami/maven/internal/persona"
|
||||||
"github.com/kami/maven/internal/phraser"
|
"github.com/kami/maven/internal/phraser"
|
||||||
"github.com/kami/maven/internal/store"
|
"github.com/kami/maven/internal/store"
|
||||||
"github.com/kami/maven/internal/webauthn"
|
"github.com/kami/maven/internal/webauthn"
|
||||||
@@ -189,7 +190,9 @@ func (l *lockedAPI) MorningStatus(ctx context.Context) ([]ipc.MorningRoutineStat
|
|||||||
func run(args []string) error {
|
func run(args []string) error {
|
||||||
cfgPath := flag.String("config", defaultConfigPath(), "path to mavend JSON config")
|
cfgPath := flag.String("config", defaultConfigPath(), "path to mavend JSON config")
|
||||||
wrappedKeyPath := flag.String("wrapped-key-file", "", "path to wrapped encryption key blob (enables cold-start unlock)")
|
wrappedKeyPath := flag.String("wrapped-key-file", "", "path to wrapped encryption key blob (enables cold-start unlock)")
|
||||||
|
reembed := flag.Bool("reembed", false, "re-embed every stored note and fact with the configured embedder, then serve normally (run once after an embedder swap)")
|
||||||
flag.CommandLine.Parse(args)
|
flag.CommandLine.Parse(args)
|
||||||
|
reembedOnStart = *reembed
|
||||||
cfg, err := config.Load(*cfgPath)
|
cfg, err := config.Load(*cfgPath)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return err
|
return err
|
||||||
@@ -269,13 +272,14 @@ func run(args []string) error {
|
|||||||
phr = phraser.NewStub()
|
phr = phraser.NewStub()
|
||||||
if cfg.Phraser != nil {
|
if cfg.Phraser != nil {
|
||||||
pc := phraser.Config{
|
pc := phraser.Config{
|
||||||
ModelPath: cfg.Phraser.ModelPath,
|
ModelPath: cfg.Phraser.ModelPath,
|
||||||
BinPath: cfg.Phraser.BinPath,
|
BinPath: cfg.Phraser.BinPath,
|
||||||
Listen: cfg.Phraser.Listen,
|
Listen: cfg.Phraser.Listen,
|
||||||
NGpuLayers: cfg.Phraser.NGpuLayers,
|
NGpuLayers: cfg.Phraser.NGpuLayers,
|
||||||
NCtx: cfg.Phraser.NCtx,
|
NCtx: cfg.Phraser.NCtx,
|
||||||
Timeout: time.Duration(cfg.Phraser.Timeout),
|
Timeout: time.Duration(cfg.Phraser.Timeout),
|
||||||
Persona: personaFromCfg(cfg),
|
LLMNudges: cfg.Phraser.LLMNudges,
|
||||||
|
ContextBlock: contextBlockFn(cfg, time.Now),
|
||||||
}
|
}
|
||||||
if pc.BinPath == "" {
|
if pc.BinPath == "" {
|
||||||
pc.BinPath = "llama-server"
|
pc.BinPath = "llama-server"
|
||||||
@@ -439,13 +443,14 @@ func run(args []string) error {
|
|||||||
phr = phraser.NewStub()
|
phr = phraser.NewStub()
|
||||||
if cfg.Phraser != nil {
|
if cfg.Phraser != nil {
|
||||||
pc := phraser.Config{
|
pc := phraser.Config{
|
||||||
ModelPath: cfg.Phraser.ModelPath,
|
ModelPath: cfg.Phraser.ModelPath,
|
||||||
BinPath: cfg.Phraser.BinPath,
|
BinPath: cfg.Phraser.BinPath,
|
||||||
Listen: cfg.Phraser.Listen,
|
Listen: cfg.Phraser.Listen,
|
||||||
NGpuLayers: cfg.Phraser.NGpuLayers,
|
NGpuLayers: cfg.Phraser.NGpuLayers,
|
||||||
NCtx: cfg.Phraser.NCtx,
|
NCtx: cfg.Phraser.NCtx,
|
||||||
Timeout: time.Duration(cfg.Phraser.Timeout),
|
Timeout: time.Duration(cfg.Phraser.Timeout),
|
||||||
Persona: personaFromCfg(cfg),
|
LLMNudges: cfg.Phraser.LLMNudges,
|
||||||
|
ContextBlock: contextBlockFn(cfg, time.Now),
|
||||||
}
|
}
|
||||||
if pc.BinPath == "" {
|
if pc.BinPath == "" {
|
||||||
pc.BinPath = "llama-server"
|
pc.BinPath = "llama-server"
|
||||||
@@ -599,12 +604,29 @@ func run(args []string) error {
|
|||||||
return nil
|
return nil
|
||||||
}
|
}
|
||||||
|
|
||||||
// personaFromCfg extracts the voice persona from the config, or returns ""
|
// personaFacts reads the optional, deployment-specific facts (his name, his
|
||||||
// when voice isn't configured. Used to pass a character prompt into the
|
// city, the free-text persona string) out of the config. Everything here may
|
||||||
// LLM phraser without requiring voice to be enabled.
|
// be empty — the context block is correct without any of it.
|
||||||
func personaFromCfg(cfg *config.Config) string {
|
func personaFacts(cfg *config.Config) persona.Facts {
|
||||||
if cfg.Voice != nil {
|
f := persona.Facts{
|
||||||
return cfg.Voice.Persona
|
// Telegram lives outside the voice block, so it counts either way.
|
||||||
|
Telegram: cfg.Telegram != nil && cfg.Telegram.BotToken != "" && cfg.Telegram.ChatID != "",
|
||||||
}
|
}
|
||||||
return ""
|
if cfg.Voice == nil {
|
||||||
|
return f
|
||||||
|
}
|
||||||
|
f.OwnerName = cfg.Voice.OwnerName
|
||||||
|
f.City = cfg.Voice.City
|
||||||
|
f.Static = cfg.Voice.Persona
|
||||||
|
// Same test wireVoice uses to pick the real provider over the stub.
|
||||||
|
f.Weather = cfg.Voice.Weather != nil && cfg.Voice.Weather.Provider == "open-meteo"
|
||||||
|
f.Tools = len(cfg.Voice.Tools) > 0
|
||||||
|
return f
|
||||||
|
}
|
||||||
|
|
||||||
|
// contextBlockFn returns the per-turn renderer of the shared context block.
|
||||||
|
// Per turn, not once at startup, because the block states the current time.
|
||||||
|
func contextBlockFn(cfg *config.Config, now func() time.Time) func() string {
|
||||||
|
f := personaFacts(cfg)
|
||||||
|
return func() string { return f.Block(now()) }
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,158 @@
|
|||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"fmt"
|
||||||
|
"math"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"github.com/kami/maven/internal/ipc"
|
||||||
|
"github.com/kami/maven/internal/memory"
|
||||||
|
"github.com/kami/maven/internal/phraser"
|
||||||
|
"github.com/kami/maven/internal/router"
|
||||||
|
"github.com/kami/maven/internal/voice"
|
||||||
|
)
|
||||||
|
|
||||||
|
// fixedEmbedder hands back a vector chosen per text, so a test can say exactly
|
||||||
|
// how close each stored memory is to the question. The real embedders make
|
||||||
|
// scores that are realistic but not controllable, and this test is about the
|
||||||
|
// gate, not about the embedder.
|
||||||
|
type fixedEmbedder struct{ vecs map[string][]float32 }
|
||||||
|
|
||||||
|
func (f *fixedEmbedder) Dim() int { return 4 }
|
||||||
|
func (f *fixedEmbedder) Close() error { return nil }
|
||||||
|
|
||||||
|
func (f *fixedEmbedder) Embed(_ context.Context, text string) ([]float32, error) {
|
||||||
|
v, ok := f.vecs[text]
|
||||||
|
if !ok {
|
||||||
|
return nil, fmt.Errorf("fixedEmbedder: no vector for %q", text)
|
||||||
|
}
|
||||||
|
return v, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// scoreVec builds a unit vector whose cosine against the query vector
|
||||||
|
// (1,0,0,0) is exactly score.
|
||||||
|
func scoreVec(score float64) []float32 {
|
||||||
|
rest := math.Sqrt(1 - score*score)
|
||||||
|
return []float32{float32(score), float32(rest), 0, 0}
|
||||||
|
}
|
||||||
|
|
||||||
|
// recordingPhraser remembers what the query path handed it to phrase, which is
|
||||||
|
// how the test can tell which pass produced the answer.
|
||||||
|
type recordingPhraser struct {
|
||||||
|
*phraser.Stub
|
||||||
|
notes []string
|
||||||
|
}
|
||||||
|
|
||||||
|
func (r *recordingPhraser) PhraseQuery(ctx context.Context, utterance string, notes []string) (string, error) {
|
||||||
|
r.notes = notes
|
||||||
|
return r.Stub.PhraseQuery(ctx, utterance, notes)
|
||||||
|
}
|
||||||
|
|
||||||
|
// recallCase — one stored memory: its text, how close it is to the question,
|
||||||
|
// whether it is a note or a fact, and whether the notes table holds it too.
|
||||||
|
type recallCase struct {
|
||||||
|
text string
|
||||||
|
score float64
|
||||||
|
kind string
|
||||||
|
}
|
||||||
|
|
||||||
|
// buildRecallHandler stores the given memories and returns a handler whose
|
||||||
|
// query path can be run directly. Notes go into BOTH the notes table and the
|
||||||
|
// vector index, which is what the daemon does (voice.go's IntentNote).
|
||||||
|
func buildRecallHandler(t *testing.T, question string, mems []recallCase) (*reactiveHandler, *recordingPhraser) {
|
||||||
|
t.Helper()
|
||||||
|
ctx := context.Background()
|
||||||
|
st := newTestStore(t)
|
||||||
|
emb := &fixedEmbedder{vecs: map[string][]float32{question: {1, 0, 0, 0}}}
|
||||||
|
mem := memory.NewInMemoryStore()
|
||||||
|
now := time.Now()
|
||||||
|
|
||||||
|
for i, m := range mems {
|
||||||
|
vec := scoreVec(m.score)
|
||||||
|
emb.vecs[m.text] = vec
|
||||||
|
id := fmt.Sprintf("%s:%d", m.kind, i)
|
||||||
|
if m.kind == "note" {
|
||||||
|
if _, err := st.WriteNote(ctx, now, m.text, vec, "tap:voice"); err != nil {
|
||||||
|
t.Fatalf("WriteNote: %v", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if err := mem.Insert(ctx, id, vec, map[string]string{"text": m.text, "type": m.kind}); err != nil {
|
||||||
|
t.Fatalf("memory insert: %v", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
phr := &recordingPhraser{Stub: phraser.NewStub()}
|
||||||
|
h := &reactiveHandler{
|
||||||
|
api: ipc.NewStoreAPI(st),
|
||||||
|
embedder: emb,
|
||||||
|
replier: voice.NewStubReplier(),
|
||||||
|
phraser: phr,
|
||||||
|
now: func() time.Time { return now },
|
||||||
|
memStore: mem,
|
||||||
|
dataStore: st,
|
||||||
|
queryMinScore: 0.55,
|
||||||
|
queryMinMargin: 0.008,
|
||||||
|
weatherProvider: nil,
|
||||||
|
}
|
||||||
|
return h, phr
|
||||||
|
}
|
||||||
|
|
||||||
|
func askQuery(t *testing.T, h *reactiveHandler, question string) string {
|
||||||
|
t.Helper()
|
||||||
|
return h.applyAction(context.Background(), router.Decision{
|
||||||
|
Intent: router.IntentQuery,
|
||||||
|
Utterance: question,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestQueryRecallNoteCanWin — the note-recall regression (Vikunja #373). Notes
|
||||||
|
// and facts share one vector index, and a note that clearly beats everything
|
||||||
|
// else must be the answer. Before the fix the memory pass only ran after the
|
||||||
|
// notes-only gate had already rejected the same note at the same score, so only
|
||||||
|
// a fact could ever come back from it.
|
||||||
|
func TestQueryRecallNoteCanWin(t *testing.T) {
|
||||||
|
const q = "где молоко"
|
||||||
|
|
||||||
|
t.Run("a clearly best note answers", func(t *testing.T) {
|
||||||
|
h, phr := buildRecallHandler(t, q, []recallCase{
|
||||||
|
{text: "молоко стоит в холодильнике", score: 0.90, kind: "note"},
|
||||||
|
{text: "выучил пару аккордов", score: 0.50, kind: "note"},
|
||||||
|
})
|
||||||
|
reply := askQuery(t, h, q)
|
||||||
|
if want := "вот что я нашла: молоко стоит в холодильнике"; reply != want {
|
||||||
|
t.Errorf("reply %q, want %q", reply, want)
|
||||||
|
}
|
||||||
|
// One text, the winning memory's — the answer came from the memory
|
||||||
|
// pass, not from handing the phraser every note in the table.
|
||||||
|
if len(phr.notes) != 1 || phr.notes[0] != "молоко стоит в холодильнике" {
|
||||||
|
t.Errorf("phraser got %q, want just the recalled note", phr.notes)
|
||||||
|
}
|
||||||
|
})
|
||||||
|
|
||||||
|
// The other half of "one gate over everything": a fact that matches better
|
||||||
|
// than the best note now answers, instead of losing to a note that only had
|
||||||
|
// to beat other notes.
|
||||||
|
t.Run("the better-matching fact answers", func(t *testing.T) {
|
||||||
|
h, _ := buildRecallHandler(t, q, []recallCase{
|
||||||
|
{text: "молоко стоит в холодильнике", score: 0.80, kind: "note"},
|
||||||
|
{text: "купил молоко в среду", score: 0.95, kind: "fact"},
|
||||||
|
})
|
||||||
|
if reply := askQuery(t, h, q); reply != "купил молоко в среду" {
|
||||||
|
t.Errorf("reply %q, want the fact read back", reply)
|
||||||
|
}
|
||||||
|
})
|
||||||
|
|
||||||
|
// The gate is untouched: two memories this close mean the embedder cannot
|
||||||
|
// tell them apart, and silence still beats a coin flip.
|
||||||
|
t.Run("no clear best stays silent", func(t *testing.T) {
|
||||||
|
h, _ := buildRecallHandler(t, q, []recallCase{
|
||||||
|
{text: "молоко стоит в холодильнике", score: 0.860, kind: "note"},
|
||||||
|
{text: "молоко закончилось", score: 0.858, kind: "note"},
|
||||||
|
})
|
||||||
|
if reply := askQuery(t, h, q); reply != "не знаю." {
|
||||||
|
t.Errorf("reply %q, want silence", reply)
|
||||||
|
}
|
||||||
|
})
|
||||||
|
}
|
||||||
+17
-14
@@ -2,21 +2,24 @@ package main
|
|||||||
|
|
||||||
import "github.com/kami/maven/internal/memory"
|
import "github.com/kami/maven/internal/memory"
|
||||||
|
|
||||||
// bestRecall is the read side of the long-term memory store: the top hit's
|
// bestRecall is the read side of the long-term memory store: the top hit when
|
||||||
// stored text when it clears the confidence gate. This recalls across BOTH
|
// it clears the confidence gate. The index holds BOTH notes and facts, and
|
||||||
// notes and facts (facts aren't in the notes table, so this is the only path
|
// either can win — the caller looks at the returned hit's meta["type"] to see
|
||||||
// that can answer "when did I last …?" from a captured fact). A note hit here
|
// which. Facts aren't in the notes table, so this is the only path that can
|
||||||
// is redundant with the notes-RAG path — by design; the two indexes can diverge
|
// answer "when did I last …?" from a captured fact.
|
||||||
// once the backend is swapped for a persistent/external store. ok=false when
|
//
|
||||||
// the hit fails the confidence gate (see memory.Confident: an absolute floor
|
// The whole hit is returned, not just its text, because "which memory answered"
|
||||||
// plus a margin over the runner-up) or carries no text.
|
// decides how the answer is said: a note gets phrased in Maven's voice, a fact
|
||||||
func bestRecall(results []memory.Result, minScore, minMargin float64) (string, bool) {
|
// is read back as stored.
|
||||||
|
//
|
||||||
|
// ok=false when the hit fails the confidence gate (see memory.Confident: an
|
||||||
|
// absolute floor plus a margin over the runner-up) or carries no text.
|
||||||
|
func bestRecall(results []memory.Result, minScore, minMargin float64) (memory.Result, bool) {
|
||||||
if !memory.Confident(results, minScore, minMargin) {
|
if !memory.Confident(results, minScore, minMargin) {
|
||||||
return "", false
|
return memory.Result{}, false
|
||||||
}
|
}
|
||||||
text := results[0].Meta["text"]
|
if results[0].Meta["text"] == "" {
|
||||||
if text == "" {
|
return memory.Result{}, false
|
||||||
return "", false
|
|
||||||
}
|
}
|
||||||
return text, true
|
return results[0], true
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -39,8 +39,27 @@ func TestBestRecall(t *testing.T) {
|
|||||||
if !ok {
|
if !ok {
|
||||||
t.Fatal("clearing hit not returned")
|
t.Fatal("clearing hit not returned")
|
||||||
}
|
}
|
||||||
if got != "выпил воды в три часа" {
|
if got.Meta["text"] != "выпил воды в три часа" {
|
||||||
t.Errorf("wrong text: %q", got)
|
t.Errorf("wrong text: %q", got.Meta["text"])
|
||||||
|
}
|
||||||
|
if got.Meta["type"] != "fact" {
|
||||||
|
t.Errorf("kind lost: %q", got.Meta["type"])
|
||||||
|
}
|
||||||
|
})
|
||||||
|
|
||||||
|
// The index holds notes and facts together, so a note has to be able to win
|
||||||
|
// it — for a long time it could not (Vikunja #373).
|
||||||
|
t.Run("a note can win", func(t *testing.T) {
|
||||||
|
res := []memory.Result{
|
||||||
|
{Score: 0.86, Meta: map[string]string{"text": "молоко в холодильнике", "type": "note"}},
|
||||||
|
{Score: 0.61, Meta: map[string]string{"text": "выпил воды", "type": "fact"}},
|
||||||
|
}
|
||||||
|
got, ok := bestRecall(res, min, margin)
|
||||||
|
if !ok {
|
||||||
|
t.Fatal("clearly-best note not returned")
|
||||||
|
}
|
||||||
|
if got.Meta["type"] != "note" || got.Meta["text"] != "молоко в холодильнике" {
|
||||||
|
t.Errorf("got %v, want the note", got.Meta)
|
||||||
}
|
}
|
||||||
})
|
})
|
||||||
|
|
||||||
|
|||||||
@@ -7,6 +7,7 @@ import (
|
|||||||
"time"
|
"time"
|
||||||
|
|
||||||
"github.com/kami/maven/internal/llm"
|
"github.com/kami/maven/internal/llm"
|
||||||
|
"github.com/kami/maven/internal/persona"
|
||||||
"github.com/kami/maven/internal/router"
|
"github.com/kami/maven/internal/router"
|
||||||
"github.com/kami/maven/internal/voice"
|
"github.com/kami/maven/internal/voice"
|
||||||
)
|
)
|
||||||
@@ -23,13 +24,19 @@ type completer interface {
|
|||||||
type llmReplier struct {
|
type llmReplier struct {
|
||||||
c completer
|
c completer
|
||||||
stub *voice.StubReplier
|
stub *voice.StubReplier
|
||||||
|
|
||||||
|
// block renders the shared context block per turn (who he is, the time).
|
||||||
|
// nil ⇒ the prompt stands alone.
|
||||||
|
block func() string
|
||||||
}
|
}
|
||||||
|
|
||||||
func newLLMReplier(c completer) *llmReplier {
|
func newLLMReplier(c completer, block func() string) *llmReplier {
|
||||||
return &llmReplier{c: c, stub: voice.NewStubReplier()}
|
return &llmReplier{c: c, stub: voice.NewStubReplier(), block: block}
|
||||||
}
|
}
|
||||||
|
|
||||||
const replySystem = `Ты — Maven, домашняя ассистентка (о себе — в женском роде). Подтверди действие РОВНО ОДНИМ коротким предложением (≤120 символов), тепло и по-русски. Не задавай вопросов, не повторяй слова, не добавляй ничего после точки. Respond ONLY with valid JSON: {"response": "...", "mood": "neutral"}.`
|
const replySystem = `Ты — Maven, домашняя ассистентка (о себе — в женском роде). Владелец — мужчина, говоришь с ним на "ты", в единственном числе; никогда не "вы"/"ваш" и не "он"/"его". Подтверди действие РОВНО ОДНИМ коротким предложением (≤120 символов), по-русски, спокойно и без официальных формулировок. Не задавай вопросов, не повторяй слова, не добавляй ничего после точки. Отвечай ТОЛЬКО одним объектом JSON с полями "response" (текст) и "mood" (ровно одно из: neutral, happy, thinking, tired, confused).
|
||||||
|
Пример: {"response": "Записала, что ты выпил стакан воды.", "mood": "neutral"}
|
||||||
|
Никогда не пиши "..." в поле response.`
|
||||||
|
|
||||||
func (r *llmReplier) Reply(d router.Decision) string {
|
func (r *llmReplier) Reply(d router.Decision) string {
|
||||||
if d.Clarify {
|
if d.Clarify {
|
||||||
@@ -38,7 +45,7 @@ func (r *llmReplier) Reply(d router.Decision) string {
|
|||||||
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
|
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
|
||||||
defer cancel()
|
defer cancel()
|
||||||
out, err := r.c.Complete(ctx, llm.Req{
|
out, err := r.c.Complete(ctx, llm.Req{
|
||||||
System: replySystem,
|
System: persona.Prepend(r.block, replySystem),
|
||||||
User: replyContext(d),
|
User: replyContext(d),
|
||||||
MaxTokens: 512,
|
MaxTokens: 512,
|
||||||
RepeatPenalty: 1.3,
|
RepeatPenalty: 1.3,
|
||||||
|
|||||||
@@ -17,7 +17,7 @@ type mockCompleter struct {
|
|||||||
func (m mockCompleter) Complete(_ context.Context, _ llm.Req) (string, error) { return m.out, m.err }
|
func (m mockCompleter) Complete(_ context.Context, _ llm.Req) (string, error) { return m.out, m.err }
|
||||||
|
|
||||||
func TestLLMReplierReturnsLLMReply(t *testing.T) {
|
func TestLLMReplierReturnsLLMReply(t *testing.T) {
|
||||||
r := newLLMReplier(mockCompleter{out: `{"response":"записала, кофе закончился","mood":"neutral"}`})
|
r := newLLMReplier(mockCompleter{out: `{"response":"записала, кофе закончился","mood":"neutral"}`}, nil)
|
||||||
got := r.Reply(router.Decision{Intent: router.IntentNote, Slots: router.Slots{Text: "кофе закончился"}})
|
got := r.Reply(router.Decision{Intent: router.IntentNote, Slots: router.Slots{Text: "кофе закончился"}})
|
||||||
if got != "записала, кофе закончился" {
|
if got != "записала, кофе закончился" {
|
||||||
t.Errorf("got %q, want %q", got, "записала, кофе закончился")
|
t.Errorf("got %q, want %q", got, "записала, кофе закончился")
|
||||||
@@ -25,7 +25,7 @@ func TestLLMReplierReturnsLLMReply(t *testing.T) {
|
|||||||
}
|
}
|
||||||
|
|
||||||
func TestLLMReplierFallsBackToPlainText(t *testing.T) {
|
func TestLLMReplierFallsBackToPlainText(t *testing.T) {
|
||||||
r := newLLMReplier(mockCompleter{out: "записала, кофе закончился"})
|
r := newLLMReplier(mockCompleter{out: "записала, кофе закончился"}, nil)
|
||||||
got := r.Reply(router.Decision{Intent: router.IntentNote, Slots: router.Slots{Text: "кофе закончился"}})
|
got := r.Reply(router.Decision{Intent: router.IntentNote, Slots: router.Slots{Text: "кофе закончился"}})
|
||||||
if got != "записала, кофе закончился" {
|
if got != "записала, кофе закончился" {
|
||||||
t.Errorf("got %q, want %q", got, "записала, кофе закончился")
|
t.Errorf("got %q, want %q", got, "записала, кофе закончился")
|
||||||
@@ -33,7 +33,7 @@ func TestLLMReplierFallsBackToPlainText(t *testing.T) {
|
|||||||
}
|
}
|
||||||
|
|
||||||
func TestLLMReplierFallsBackToStubOnError(t *testing.T) {
|
func TestLLMReplierFallsBackToStubOnError(t *testing.T) {
|
||||||
r := newLLMReplier(mockCompleter{err: errTestLLMDown})
|
r := newLLMReplier(mockCompleter{err: errTestLLMDown}, nil)
|
||||||
noteDec := router.Decision{Intent: router.IntentNote}
|
noteDec := router.Decision{Intent: router.IntentNote}
|
||||||
got := r.Reply(noteDec)
|
got := r.Reply(noteDec)
|
||||||
want := voice.NewStubReplier().Reply(noteDec)
|
want := voice.NewStubReplier().Reply(noteDec)
|
||||||
@@ -43,7 +43,7 @@ func TestLLMReplierFallsBackToStubOnError(t *testing.T) {
|
|||||||
}
|
}
|
||||||
|
|
||||||
func TestLLMReplierFallsBackToStubOnEmpty(t *testing.T) {
|
func TestLLMReplierFallsBackToStubOnEmpty(t *testing.T) {
|
||||||
r := newLLMReplier(mockCompleter{out: ""})
|
r := newLLMReplier(mockCompleter{out: ""}, nil)
|
||||||
noteDec := router.Decision{Intent: router.IntentNote}
|
noteDec := router.Decision{Intent: router.IntentNote}
|
||||||
got := r.Reply(noteDec)
|
got := r.Reply(noteDec)
|
||||||
want := voice.NewStubReplier().Reply(noteDec)
|
want := voice.NewStubReplier().Reply(noteDec)
|
||||||
@@ -53,7 +53,7 @@ func TestLLMReplierFallsBackToStubOnEmpty(t *testing.T) {
|
|||||||
}
|
}
|
||||||
|
|
||||||
func TestLLMReplierClarifyUsesStub(t *testing.T) {
|
func TestLLMReplierClarifyUsesStub(t *testing.T) {
|
||||||
r := newLLMReplier(mockCompleter{out: "я всё поняла"})
|
r := newLLMReplier(mockCompleter{out: "я всё поняла"}, nil)
|
||||||
clarifyDec := router.Decision{Clarify: true}
|
clarifyDec := router.Decision{Clarify: true}
|
||||||
got := r.Reply(clarifyDec)
|
got := r.Reply(clarifyDec)
|
||||||
want := voice.NewStubReplier().Reply(clarifyDec)
|
want := voice.NewStubReplier().Reply(clarifyDec)
|
||||||
|
|||||||
@@ -0,0 +1,78 @@
|
|||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"github.com/kami/maven/internal/router"
|
||||||
|
)
|
||||||
|
|
||||||
|
// systemHandler — a handler with nothing but a fixed clock, which is all
|
||||||
|
// replySystem needs.
|
||||||
|
func systemHandler(now time.Time) *reactiveHandler {
|
||||||
|
return &reactiveHandler{now: func() time.Time { return now }}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestReplySystemDateOffset — "какое число завтра" must answer tomorrow's
|
||||||
|
// date, not today's (Vikunja #388).
|
||||||
|
func TestReplySystemDateOffset(t *testing.T) {
|
||||||
|
// Thursday, 30 July 2026.
|
||||||
|
now := time.Date(2026, 7, 30, 14, 5, 0, 0, time.UTC)
|
||||||
|
h := systemHandler(now)
|
||||||
|
cases := []struct{ utterance, want string }{
|
||||||
|
{"какое сегодня число", "сегодня четверг, 30 июля 2026 года"},
|
||||||
|
{"какое число", "сегодня четверг, 30 июля 2026 года"},
|
||||||
|
{"какое число завтра", "завтра пятница, 31 июля 2026 года"},
|
||||||
|
{"какое число послезавтра", "послезавтра суббота, 1 августа 2026 года"},
|
||||||
|
{"какое было число вчера", "вчера среда, 29 июля 2026 года"},
|
||||||
|
}
|
||||||
|
for _, c := range cases {
|
||||||
|
got := h.replySystem(context.Background(), router.Decision{Utterance: c.utterance})
|
||||||
|
if got != c.want {
|
||||||
|
t.Errorf("replySystem(%q) = %q, want %q", c.utterance, got, c.want)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// A day she cannot work out must not come back as today's date — that is the
|
||||||
|
// same silent wrong answer #388 was about, one step further out.
|
||||||
|
func TestReplySystemUnknownDayIsHonest(t *testing.T) {
|
||||||
|
now := time.Date(2026, 7, 30, 14, 5, 0, 0, time.UTC)
|
||||||
|
h := systemHandler(now)
|
||||||
|
for _, u := range []string{
|
||||||
|
"какое число в пятницу",
|
||||||
|
"какое число через неделю",
|
||||||
|
"какое число в понедельник",
|
||||||
|
} {
|
||||||
|
got := h.replySystem(context.Background(), router.Decision{Utterance: u})
|
||||||
|
if got != onlyNearDaysReply {
|
||||||
|
t.Errorf("replySystem(%q) = %q, want the honest reply", u, got)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
// The days she does know must not be caught by the same guard.
|
||||||
|
if got := h.replySystem(context.Background(), router.Decision{Utterance: "какое число завтра"}); got == onlyNearDaysReply {
|
||||||
|
t.Error("завтра was treated as an unknown day")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestReplySystemClockCity — the clock arm must not answer local time for a
|
||||||
|
// question about another city (Vikunja #388). She keeps one clock, so every
|
||||||
|
// named place gets the honest "local time only" answer.
|
||||||
|
func TestReplySystemClockCity(t *testing.T) {
|
||||||
|
now := time.Date(2026, 7, 30, 12, 0, 0, 0, time.UTC)
|
||||||
|
h := systemHandler(now)
|
||||||
|
cases := []struct{ utterance, want string }{
|
||||||
|
{"который час", "сейчас 12 часов ровно"},
|
||||||
|
{"который час в киеве", onlyLocalTimeReply},
|
||||||
|
{"сколько времени в москве", onlyLocalTimeReply},
|
||||||
|
{"который час в лондоне", onlyLocalTimeReply},
|
||||||
|
{"который час в бишкеке", onlyLocalTimeReply},
|
||||||
|
}
|
||||||
|
for _, c := range cases {
|
||||||
|
got := h.replySystem(context.Background(), router.Decision{Utterance: c.utterance})
|
||||||
|
if got != c.want {
|
||||||
|
t.Errorf("replySystem(%q) = %q, want %q", c.utterance, got, c.want)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
+267
-41
@@ -171,6 +171,7 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
|
|||||||
emb = router.NewHashEmbedder(1024)
|
emb = router.NewHashEmbedder(1024)
|
||||||
}
|
}
|
||||||
w.embedder = emb
|
w.embedder = emb
|
||||||
|
checkStoredEmbedder(dataStore, emb)
|
||||||
|
|
||||||
// ----- tool executor (the enabled act allowlist, store-backed) -----
|
// ----- tool executor (the enabled act allowlist, store-backed) -----
|
||||||
// Config tools are the declarative bootstrap: seed them into the store as
|
// Config tools are the declarative bootstrap: seed them into the store as
|
||||||
@@ -227,14 +228,25 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
|
|||||||
}
|
}
|
||||||
|
|
||||||
// ----- dialogue (multi-turn slot carry-over; 2-min follow-up window) -----
|
// ----- dialogue (multi-turn slot carry-over; 2-min follow-up window) -----
|
||||||
dialogueSessions := dialogue.NewSessionStore(2 * time.Minute)
|
// Store-backed when the daemon passes a store, so a restart mid-conversation
|
||||||
|
// keeps the thread (Vikunja #363). Sessions past their TTL are dropped on
|
||||||
|
// load, never revived. Clarify's parked question stays in memory only.
|
||||||
|
var dialogueSessions *dialogue.SessionStore
|
||||||
|
if dataStore != nil {
|
||||||
|
dialogueSessions = dialogue.NewPersistentSessionStore(2*time.Minute, dataStore)
|
||||||
|
if err := dialogueSessions.Load(context.Background(), time.Now()); err != nil {
|
||||||
|
log.Printf("dialogue: load saved sessions: %v", err)
|
||||||
|
}
|
||||||
|
} else {
|
||||||
|
dialogueSessions = dialogue.NewSessionStore(2 * time.Minute)
|
||||||
|
}
|
||||||
clarifyStore := dialogue.NewClarifyStore(clarifyTTL)
|
clarifyStore := dialogue.NewClarifyStore(clarifyTTL)
|
||||||
timeParser := router.NewPythonDateParser()
|
timeParser := router.NewPythonDateParser()
|
||||||
|
|
||||||
// ----- replier (LLM-backed when the engine is on, Stub floor otherwise) -----
|
// ----- replier (LLM-backed when the engine is on, Stub floor otherwise) -----
|
||||||
replier := voice.Replier(voice.NewStubReplier())
|
replier := voice.Replier(voice.NewStubReplier())
|
||||||
if llmClient != nil {
|
if llmClient != nil {
|
||||||
replier = newLLMReplier(llmClient)
|
replier = newLLMReplier(llmClient, contextBlockFn(cfg, time.Now))
|
||||||
}
|
}
|
||||||
|
|
||||||
// ----- the handler (the reactive path; closes over stt / tts / router / coreAPI / memory) -----
|
// ----- the handler (the reactive path; closes over stt / tts / router / coreAPI / memory) -----
|
||||||
@@ -255,11 +267,13 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
|
|||||||
dataStore: dataStore,
|
dataStore: dataStore,
|
||||||
dialogueSessions: dialogueSessions,
|
dialogueSessions: dialogueSessions,
|
||||||
clarifyStore: clarifyStore,
|
clarifyStore: clarifyStore,
|
||||||
extractor: router.Extractor{Time: timeParser, Acts: matcher, Facts: router.DefaultFactParser{}},
|
// 0 here (unset config) ⇒ the dialogue default.
|
||||||
queryMinScore: cfg.Voice.QueryMinScore,
|
clarifyMaxAttempts: cfg.Voice.ClarifyMaxAttempts,
|
||||||
queryMinMargin: cfg.Voice.QueryMinMargin,
|
extractor: router.Extractor{Time: timeParser, Acts: matcher, Facts: router.DefaultFactParser{}},
|
||||||
timeParser: timeParser,
|
queryMinScore: cfg.Voice.QueryMinScore,
|
||||||
ecosystem: eco,
|
queryMinMargin: cfg.Voice.QueryMinMargin,
|
||||||
|
timeParser: timeParser,
|
||||||
|
ecosystem: eco,
|
||||||
}
|
}
|
||||||
|
|
||||||
// ----- the server (TCP listener) -----
|
// ----- the server (TCP listener) -----
|
||||||
@@ -318,6 +332,10 @@ type reactiveHandler struct {
|
|||||||
// clarify.go). nil ⇒ she falls back to the canned "не поняла" reply.
|
// clarify.go). nil ⇒ she falls back to the canned "не поняла" reply.
|
||||||
clarifyStore *dialogue.ClarifyStore
|
clarifyStore *dialogue.ClarifyStore
|
||||||
|
|
||||||
|
// clarifyMaxAttempts — questions per request before she gives up out loud.
|
||||||
|
// 0 ⇒ dialogue.DefaultMaxAttempts (3). Set from VoiceConfig.
|
||||||
|
clarifyMaxAttempts int
|
||||||
|
|
||||||
// extractor parses the answer to an open question, with the same parsers
|
// extractor parses the answer to an open question, with the same parsers
|
||||||
// the router's own stage-2 uses.
|
// the router's own stage-2 uses.
|
||||||
extractor router.Extractor
|
extractor router.Extractor
|
||||||
@@ -396,9 +414,17 @@ func (h *reactiveHandler) HandlePushToTalk(ctx context.Context, req voice.PushTo
|
|||||||
return h.reply(ctx, reply, nil)
|
return h.reply(ctx, reply, nil)
|
||||||
}
|
}
|
||||||
|
|
||||||
// 1b2. clarify answer — if she asked a question last turn, this utterance is
|
// 1b2. expired clarify — a question was parked but its TTL ran out, so the
|
||||||
// its answer, not a fresh command. After the confirm check: a y/n gate is
|
// request behind it is gone. Say that out loud (see clarify.go) and carry
|
||||||
// armed by her own prompt and is the narrower claim on the utterance.
|
// on: these words are still routed as a fresh utterance below, with the
|
||||||
|
// notice glued in front of whatever the fresh routing answers. Checked
|
||||||
|
// BEFORE the answer path: reading a parked question drops an expired one.
|
||||||
|
expiredNotice := h.clarifyExpiredNotice()
|
||||||
|
|
||||||
|
// 1b3. clarify answer — if she asked a live question last turn, this
|
||||||
|
// utterance is its answer, not a fresh command. After the confirm check: a
|
||||||
|
// y/n gate is armed by her own prompt and is the narrower claim on the
|
||||||
|
// utterance.
|
||||||
if reply, handled := h.resolveClarifyAnswer(ctx, text); handled {
|
if reply, handled := h.resolveClarifyAnswer(ctx, text); handled {
|
||||||
return h.reply(ctx, reply, nil)
|
return h.reply(ctx, reply, nil)
|
||||||
}
|
}
|
||||||
@@ -408,7 +434,7 @@ func (h *reactiveHandler) HandlePushToTalk(ctx context.Context, req voice.PushTo
|
|||||||
// unreliably (it's a command, not a free-form query), so we match it
|
// unreliably (it's a command, not a free-form query), so we match it
|
||||||
// before routing. Same pattern as the confirm turn above.
|
// before routing. Same pattern as the confirm turn above.
|
||||||
if reply, handled := h.resolveQuietToggle(ctx, text); handled {
|
if reply, handled := h.resolveQuietToggle(ctx, text); handled {
|
||||||
return h.reply(ctx, reply, nil)
|
return h.reply(ctx, withNotice(expiredNotice, reply), nil)
|
||||||
}
|
}
|
||||||
|
|
||||||
// 2. router — classify the utterance.
|
// 2. router — classify the utterance.
|
||||||
@@ -417,10 +443,10 @@ func (h *reactiveHandler) HandlePushToTalk(ctx context.Context, req voice.PushTo
|
|||||||
// ErrNoIntents ⇒ classifier unseeded (cold boot). reply with a
|
// ErrNoIntents ⇒ classifier unseeded (cold boot). reply with a
|
||||||
// "still warming up" rather than a wire error.
|
// "still warming up" rather than a wire error.
|
||||||
if errors.Is(err, router.ErrNoIntents) {
|
if errors.Is(err, router.ErrNoIntents) {
|
||||||
return h.reply(ctx, "я ещё не понимаю свободную речь — скоро научусь.", nil)
|
return h.reply(ctx, withNotice(expiredNotice, "я ещё не понимаю свободную речь — скоро научусь."), nil)
|
||||||
}
|
}
|
||||||
log.Printf("voice: router error: %v", err)
|
log.Printf("voice: router error: %v", err)
|
||||||
return h.reply(ctx, "не получилось разобрать команду.", nil)
|
return h.reply(ctx, withNotice(expiredNotice, "не получилось разобрать команду."), nil)
|
||||||
}
|
}
|
||||||
|
|
||||||
// 2b. dialogue — fill this turn's missing slots from a prior same-intent
|
// 2b. dialogue — fill this turn's missing slots from a prior same-intent
|
||||||
@@ -441,7 +467,7 @@ func (h *reactiveHandler) HandlePushToTalk(ctx context.Context, req voice.PushTo
|
|||||||
// stands.
|
// stands.
|
||||||
if dec.Clarify {
|
if dec.Clarify {
|
||||||
if question, asked := h.askClarify(dec); asked {
|
if question, asked := h.askClarify(dec); asked {
|
||||||
return h.reply(ctx, question, nil)
|
return h.reply(ctx, withNotice(expiredNotice, question), nil)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -457,7 +483,7 @@ func (h *reactiveHandler) HandlePushToTalk(ctx context.Context, req voice.PushTo
|
|||||||
|
|
||||||
// 5. tts — synthesise the reply text; return to the voice server which
|
// 5. tts — synthesise the reply text; return to the voice server which
|
||||||
// ships it back on the conn.
|
// ships it back on the conn.
|
||||||
return h.reply(ctx, replyText, nil)
|
return h.reply(ctx, withNotice(expiredNotice, replyText), nil)
|
||||||
}
|
}
|
||||||
|
|
||||||
// handleText — the core reactive path without stt/tts: confirm check →
|
// handleText — the core reactive path without stt/tts: confirm check →
|
||||||
@@ -472,7 +498,9 @@ func (h *reactiveHandler) handleText(ctx context.Context, text string) string {
|
|||||||
return reply
|
return reply
|
||||||
}
|
}
|
||||||
|
|
||||||
// 1b2. clarify answer — same check as HandlePushToTalk.
|
// 1b2/1b3. expired clarify then clarify answer — same order and reasons as
|
||||||
|
// HandlePushToTalk.
|
||||||
|
expiredNotice := h.clarifyExpiredNotice()
|
||||||
if reply, handled := h.resolveClarifyAnswer(ctx, text); handled {
|
if reply, handled := h.resolveClarifyAnswer(ctx, text); handled {
|
||||||
return reply
|
return reply
|
||||||
}
|
}
|
||||||
@@ -481,10 +509,10 @@ func (h *reactiveHandler) handleText(ctx context.Context, text string) string {
|
|||||||
dec, err := h.router.Route(ctx, text, h.now())
|
dec, err := h.router.Route(ctx, text, h.now())
|
||||||
if err != nil {
|
if err != nil {
|
||||||
if errors.Is(err, router.ErrNoIntents) {
|
if errors.Is(err, router.ErrNoIntents) {
|
||||||
return "я ещё не понимаю свободную речь — скоро научусь."
|
return withNotice(expiredNotice, "я ещё не понимаю свободную речь — скоро научусь.")
|
||||||
}
|
}
|
||||||
log.Printf("voice: handleText router error: %v", err)
|
log.Printf("voice: handleText router error: %v", err)
|
||||||
return "не получилось разобрать команду."
|
return withNotice(expiredNotice, "не получилось разобрать команду.")
|
||||||
}
|
}
|
||||||
log.Printf("voice: route result: intent=%s slots=%+v", dec.Intent, dec.Slots)
|
log.Printf("voice: route result: intent=%s slots=%+v", dec.Intent, dec.Slots)
|
||||||
|
|
||||||
@@ -501,7 +529,7 @@ func (h *reactiveHandler) handleText(ctx context.Context, text string) string {
|
|||||||
// 2c. clarify — same as HandlePushToTalk: ask about the one missing thing.
|
// 2c. clarify — same as HandlePushToTalk: ask about the one missing thing.
|
||||||
if dec.Clarify {
|
if dec.Clarify {
|
||||||
if question, asked := h.askClarify(dec); asked {
|
if question, asked := h.askClarify(dec); asked {
|
||||||
return question
|
return withNotice(expiredNotice, question)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -513,7 +541,7 @@ func (h *reactiveHandler) handleText(ctx context.Context, text string) string {
|
|||||||
if replyText == "" {
|
if replyText == "" {
|
||||||
replyText = h.replier.Reply(dec)
|
replyText = h.replier.Reply(dec)
|
||||||
}
|
}
|
||||||
return replyText
|
return withNotice(expiredNotice, replyText)
|
||||||
}
|
}
|
||||||
|
|
||||||
// applyAction — executes the router's Decision. Intent-by-intent:
|
// applyAction — executes the router's Decision. Intent-by-intent:
|
||||||
@@ -724,7 +752,9 @@ func (h *reactiveHandler) applyAction(ctx context.Context, dec router.Decision)
|
|||||||
}
|
}
|
||||||
|
|
||||||
// Calendar questions: "что у меня сегодня?", "планы на завтра?"
|
// Calendar questions: "что у меня сегодня?", "планы на завтра?"
|
||||||
if date, ok := router.ParseCalendarDate(dec.Utterance, time.Now()); ok {
|
// h.now(), not time.Now(): the handler's clock is the injected one, so
|
||||||
|
// this arm can be tested at a fixed time like the rest.
|
||||||
|
if date, ok := router.ParseCalendarDate(dec.Utterance, h.now()); ok {
|
||||||
events, err := h.api.CalendarEvents(ctx, date, date.Add(24*time.Hour))
|
events, err := h.api.CalendarEvents(ctx, date, date.Add(24*time.Hour))
|
||||||
if err != nil {
|
if err != nil {
|
||||||
log.Printf("voice: calendar events: %v", err)
|
log.Printf("voice: calendar events: %v", err)
|
||||||
@@ -759,11 +789,43 @@ func (h *reactiveHandler) applyAction(ctx context.Context, dec router.Decision)
|
|||||||
log.Printf("voice: embed query: %v", err)
|
log.Printf("voice: embed query: %v", err)
|
||||||
return "не получилось найти ответ."
|
return "не получилось найти ответ."
|
||||||
}
|
}
|
||||||
|
// Long-term memory first: ONE search over everything Maven remembers
|
||||||
|
// (notes and facts share this index) and ONE confidence gate, so the
|
||||||
|
// memory that is clearly the best match answers — a note just as much
|
||||||
|
// as a fact.
|
||||||
|
//
|
||||||
|
// This used to run only after the notes-only gate below had already
|
||||||
|
// rejected the same note at the same score, which no note could ever
|
||||||
|
// survive a second time: the branch could only return a fact (#373).
|
||||||
|
// Order, not the gate, was the bug — the set of questions Maven answers
|
||||||
|
// is unchanged, only which memory gets to answer them.
|
||||||
|
if h.memStore != nil {
|
||||||
|
if hits, herr := h.memStore.Search(ctx, vec, 3); herr == nil {
|
||||||
|
if hit, ok := bestRecall(hits, h.queryMinScore, h.queryMinMargin); ok {
|
||||||
|
text := hit.Meta["text"]
|
||||||
|
// A note is phrased in Maven's voice; a fact is read back
|
||||||
|
// as it was stored.
|
||||||
|
if hit.Meta["type"] == "note" {
|
||||||
|
if reply, perr := h.phraser.PhraseQuery(ctx, dec.Utterance, []string{text}); perr == nil && reply != "" {
|
||||||
|
return reply
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return text
|
||||||
|
}
|
||||||
|
} else {
|
||||||
|
log.Printf("voice: memory search: %v", herr)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
notes, err := h.api.QueryNotes(ctx, vec, 5)
|
notes, err := h.api.QueryNotes(ctx, vec, 5)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
log.Printf("voice: query notes: %v", err)
|
log.Printf("voice: query notes: %v", err)
|
||||||
return "не получилось найти ответ."
|
return "не получилось найти ответ."
|
||||||
}
|
}
|
||||||
|
// Notes-only pass, for notes the vector index above does not hold (an
|
||||||
|
// older note written before it existed). Same gate, notes-only
|
||||||
|
// candidates.
|
||||||
|
//
|
||||||
// Confidence gate: below it, say "I don't know" rather than read back
|
// Confidence gate: below it, say "I don't know" rather than read back
|
||||||
// the least-unrelated note — a confident wrong recall is worse than a
|
// the least-unrelated note — a confident wrong recall is worse than a
|
||||||
// gap (spec's "not a guesser-of-truth"). Same instinct as the loop's
|
// gap (spec's "not a guesser-of-truth"). Same instinct as the loop's
|
||||||
@@ -775,16 +837,6 @@ func (h *reactiveHandler) applyAction(ctx context.Context, dec router.Decision)
|
|||||||
noteScores[i] = n.Score
|
noteScores[i] = n.Score
|
||||||
}
|
}
|
||||||
if !memory.ConfidentScores(noteScores, h.queryMinScore, h.queryMinMargin) {
|
if !memory.ConfidentScores(noteScores, h.queryMinScore, h.queryMinMargin) {
|
||||||
// Long-term memory recall (notes + facts) before general knowledge:
|
|
||||||
// the notes table can't answer fact questions, but the memory store
|
|
||||||
// indexes both. Only runs when notes-RAG already gave up → additive.
|
|
||||||
if h.memStore != nil {
|
|
||||||
if hits, herr := h.memStore.Search(ctx, vec, 3); herr == nil {
|
|
||||||
if text, ok := bestRecall(hits, h.queryMinScore, h.queryMinMargin); ok {
|
|
||||||
return text
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
// Try general knowledge from the phraser before giving up
|
// Try general knowledge from the phraser before giving up
|
||||||
reply, err := h.phraser.PhraseQuery(ctx, dec.Utterance, nil)
|
reply, err := h.phraser.PhraseQuery(ctx, dec.Utterance, nil)
|
||||||
if err != nil || reply == "" {
|
if err != nil || reply == "" {
|
||||||
@@ -886,6 +938,101 @@ var ruMonths = []string{
|
|||||||
"июля", "августа", "сентября", "октября", "ноября", "декабря",
|
"июля", "августа", "сентября", "октября", "ноября", "декабря",
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// onlyLocalTimeReply — the honest answer when the user asks the time somewhere
|
||||||
|
// other than here. She only keeps one clock, and saying so is better than
|
||||||
|
// naming the wrong city's time.
|
||||||
|
//
|
||||||
|
// There used to be a city→time-zone table here. It was removed on purpose: the
|
||||||
|
// user only ever asks for local time, so the table was a second list of cities
|
||||||
|
// to keep in step with the weather one for no gain.
|
||||||
|
const onlyLocalTimeReply = "я знаю только местное время, про другие города пока не скажу."
|
||||||
|
|
||||||
|
// notPlaceAfterV — words that follow "в" without naming a place, so
|
||||||
|
// mentionsUnknownPlace does not mistake them for a city.
|
||||||
|
var notPlaceAfterV = map[string]bool{
|
||||||
|
"данный": true, "данную": true, "этот": true, "эту": true,
|
||||||
|
"котором": true, "какое": true, "какой": true, "который": true,
|
||||||
|
"общем": true, "точности": true, "курсе": true, "сутках": true,
|
||||||
|
"часах": true, "минутах": true, "секундах": true, "неделе": true,
|
||||||
|
}
|
||||||
|
|
||||||
|
// mentionsUnknownPlace reports whether the question has a "в <слово>" phrase
|
||||||
|
// that looks like a place we do not know ("который час в киеве"). Used only to
|
||||||
|
// pick the honest "local time only" reply instead of answering local time as
|
||||||
|
// if it were the city's.
|
||||||
|
func mentionsUnknownPlace(u string) bool {
|
||||||
|
toks := strings.Fields(u)
|
||||||
|
for i := 0; i+1 < len(toks); i++ {
|
||||||
|
if toks[i] != "в" && toks[i] != "во" {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
next := strings.Trim(toks[i+1], ".,?!")
|
||||||
|
if next == "" || notPlaceAfterV[next] {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
// A number after "в" is a clock ("в 5 часов"), not a place.
|
||||||
|
if _, err := strconv.Atoi(strings.SplitN(next, ":", 2)[0]); err == nil {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
|
||||||
|
// onlyNearDaysReply — she can work out today, tomorrow, the day after and
|
||||||
|
// yesterday, and nothing further. Said out loud instead of answering today's
|
||||||
|
// date for a day she did not understand.
|
||||||
|
const onlyNearDaysReply = "я считаю только сегодня, завтра, послезавтра и вчера — про другие дни пока не скажу."
|
||||||
|
|
||||||
|
// dayWords — day references the calendar parser cannot resolve. A weekday name
|
||||||
|
// or a "через …" phrase means he asked about a specific other day.
|
||||||
|
var dayWords = []string{
|
||||||
|
"понедельник", "вторник", "сред", "четверг", "пятниц", "суббот", "воскресен",
|
||||||
|
"через", "monday", "tuesday", "wednesday", "thursday", "friday", "saturday", "sunday",
|
||||||
|
}
|
||||||
|
|
||||||
|
// mentionsUnknownDay reports whether the question names a day the calendar
|
||||||
|
// parser could not resolve. Mirror of mentionsUnknownPlace: it exists only to
|
||||||
|
// pick an honest reply over a confidently wrong one.
|
||||||
|
//
|
||||||
|
// Only called after ParseCalendarDate has already failed, so "завтра" and the
|
||||||
|
// other words it does know never reach here.
|
||||||
|
func mentionsUnknownDay(u string) bool {
|
||||||
|
for _, w := range dayWords {
|
||||||
|
if strings.Contains(u, w) {
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
|
||||||
|
// ruClock renders the clock part of the time reply: "15 часов 4 минуты".
|
||||||
|
func ruClock(t time.Time) string {
|
||||||
|
h, m := t.Hour(), t.Minute()
|
||||||
|
hourWord := ruPlural(h, "час", "часа", "часов")
|
||||||
|
if m == 0 {
|
||||||
|
return fmt.Sprintf("%d %s ровно", h, hourWord)
|
||||||
|
}
|
||||||
|
return fmt.Sprintf("%d %s %d %s", h, hourWord, m, ruPlural(m, "минута", "минуты", "минут"))
|
||||||
|
}
|
||||||
|
|
||||||
|
// dayPrefix names the day relative to now ("завтра", "вчера", …) so the date
|
||||||
|
// reply opens the way a person would say it.
|
||||||
|
func dayPrefix(now, day time.Time) string {
|
||||||
|
base := time.Date(now.Year(), now.Month(), now.Day(), 0, 0, 0, 0, now.Location())
|
||||||
|
switch int(day.Sub(base).Hours() / 24) {
|
||||||
|
case -1:
|
||||||
|
return "вчера"
|
||||||
|
case 0:
|
||||||
|
return "сегодня"
|
||||||
|
case 1:
|
||||||
|
return "завтра"
|
||||||
|
case 2:
|
||||||
|
return "послезавтра"
|
||||||
|
}
|
||||||
|
return "это"
|
||||||
|
}
|
||||||
|
|
||||||
func ruPlural(n int, one, two, many string) string {
|
func ruPlural(n int, one, two, many string) string {
|
||||||
n = n % 100
|
n = n % 100
|
||||||
if n > 10 && n < 20 {
|
if n > 10 && n < 20 {
|
||||||
@@ -964,18 +1111,29 @@ func (h *reactiveHandler) replySystem(ctx context.Context, dec router.Decision)
|
|||||||
|
|
||||||
switch {
|
switch {
|
||||||
case strings.Contains(u, "час") || strings.Contains(u, "врем"):
|
case strings.Contains(u, "час") || strings.Contains(u, "врем"):
|
||||||
h := now.Hour()
|
// "который час в киеве" — she keeps one clock, so any named place gets
|
||||||
m := now.Minute()
|
// the honest answer. Never local time dressed up as the city's.
|
||||||
hourWord := ruPlural(h, "час", "часа", "часов")
|
if mentionsUnknownPlace(u) {
|
||||||
if m == 0 {
|
return onlyLocalTimeReply
|
||||||
return fmt.Sprintf("сейчас %d %s ровно", h, hourWord)
|
|
||||||
}
|
}
|
||||||
minWord := ruPlural(m, "минута", "минуты", "минут")
|
return "сейчас " + ruClock(now)
|
||||||
return fmt.Sprintf("сейчас %d %s %d %s", h, hourWord, m, minWord)
|
|
||||||
case strings.Contains(u, "день") || strings.Contains(u, "числ"):
|
case strings.Contains(u, "день") || strings.Contains(u, "числ"):
|
||||||
dow := ruWeekdays[now.Weekday()]
|
// "какое число завтра" — answer for the day the user asked about,
|
||||||
month := ruMonths[now.Month()-1]
|
// not today. Reuses the router's calendar day-word parser.
|
||||||
return fmt.Sprintf("сегодня %s, %d %s %d года", dow, now.Day(), month, now.Year())
|
day := now
|
||||||
|
prefix := "сегодня"
|
||||||
|
if d, ok := router.ParseCalendarDate(u, now); ok {
|
||||||
|
day = d
|
||||||
|
prefix = dayPrefix(now, d)
|
||||||
|
} else if mentionsUnknownDay(u) {
|
||||||
|
// He named a day she cannot work out ("в пятницу", "через неделю").
|
||||||
|
// Answering today's date here would be the same silent wrong answer
|
||||||
|
// this arm was fixed for, so say what she can do instead.
|
||||||
|
return onlyNearDaysReply
|
||||||
|
}
|
||||||
|
dow := ruWeekdays[day.Weekday()]
|
||||||
|
month := ruMonths[day.Month()-1]
|
||||||
|
return fmt.Sprintf("%s %s, %d %s %d года", prefix, dow, day.Day(), month, day.Year())
|
||||||
case strings.Contains(u, "кто дома") || strings.Contains(u, "человек дома"):
|
case strings.Contains(u, "кто дома") || strings.Contains(u, "человек дома"):
|
||||||
return "присутствие пока не подключено к голосовому запросу."
|
return "присутствие пока не подключено к голосовому запросу."
|
||||||
case strings.Contains(u, "памят") || strings.Contains(u, "процессор") || strings.Contains(u, "загрузк") || strings.Contains(u, "статус") || strings.Contains(u, "работа") || strings.Contains(u, "сервис") || strings.Contains(u, "диск") || strings.Contains(u, "ip") || strings.Contains(u, "аптайм") || strings.Contains(u, "трафик") || strings.Contains(u, "интернет"):
|
case strings.Contains(u, "памят") || strings.Contains(u, "процессор") || strings.Contains(u, "загрузк") || strings.Contains(u, "статус") || strings.Contains(u, "работа") || strings.Contains(u, "сервис") || strings.Contains(u, "диск") || strings.Contains(u, "ip") || strings.Contains(u, "аптайм") || strings.Contains(u, "трафик") || strings.Contains(u, "интернет"):
|
||||||
@@ -1718,3 +1876,71 @@ func jsonStringImpl(s string) string {
|
|||||||
b = append(b, '"')
|
b = append(b, '"')
|
||||||
return string(b)
|
return string(b)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// reembedOnStart is the -reembed flag (set in run()). Opt-in on purpose: see
|
||||||
|
// runReembed.
|
||||||
|
var reembedOnStart bool
|
||||||
|
|
||||||
|
// checkStoredEmbedder compares the embedder we just loaded with the one that
|
||||||
|
// wrote the vectors already in the DB (Vikunja #378).
|
||||||
|
//
|
||||||
|
// The two models we have both make 384-dim vectors, so a size check catches
|
||||||
|
// nothing: after a swap, recall silently compares vectors from different
|
||||||
|
// spaces and the scores are noise. So we say it out loud. Recall itself is not
|
||||||
|
// changed here — the fix is `mavend -reembed`.
|
||||||
|
func checkStoredEmbedder(dataStore *store.Store, emb router.Embedder) {
|
||||||
|
if dataStore == nil {
|
||||||
|
return
|
||||||
|
}
|
||||||
|
current := router.EmbedderID(emb)
|
||||||
|
if reembedOnStart {
|
||||||
|
runReembed(dataStore, emb, current)
|
||||||
|
return
|
||||||
|
}
|
||||||
|
stored, mismatch, err := dataStore.CheckEmbedder(context.Background(), current)
|
||||||
|
if err != nil {
|
||||||
|
log.Printf("voice: embedder marker check failed: %v", err)
|
||||||
|
return
|
||||||
|
}
|
||||||
|
if mismatch {
|
||||||
|
log.Printf("voice: WARNING embedder MISMATCH — stored vectors were written by %q but the configured embedder is %q; recall scores are noise until the notes and facts are re-embedded — run `mavend -reembed` once (Vikunja #378)", stored, current)
|
||||||
|
return
|
||||||
|
}
|
||||||
|
log.Printf("voice: embedder marker ok (%s)", current)
|
||||||
|
}
|
||||||
|
|
||||||
|
// runReembed is the one-shot backfill behind -reembed.
|
||||||
|
//
|
||||||
|
// Why a flag and not automatic on mismatch: the embedder is ONNX on the
|
||||||
|
// laptop's CPU, so a few thousand notes is minutes of work. Doing that silently
|
||||||
|
// inside a normal start would look like the daemon hanging on boot. So the user
|
||||||
|
// runs it once, deliberately, after an embedder swap; the mismatch warning
|
||||||
|
// above tells them to. It re-embeds, logs what it did, and then the daemon
|
||||||
|
// carries on serving as usual — no separate binary, no second start needed.
|
||||||
|
func runReembed(dataStore *store.Store, emb router.Embedder, current string) {
|
||||||
|
log.Printf("voice: re-embedding stored notes and facts with %s — this can take a few minutes, do not interrupt", current)
|
||||||
|
res, err := dataStore.ReembedAll(context.Background(), current,
|
||||||
|
// EmbedPassage, not EmbedQuery: these are stored texts being searched
|
||||||
|
// FOR, which is the side they were written with.
|
||||||
|
func(ctx context.Context, text string) ([]float32, error) {
|
||||||
|
return router.EmbedPassage(ctx, emb, text)
|
||||||
|
})
|
||||||
|
if err != nil {
|
||||||
|
log.Printf("voice: re-embed FAILED, nothing was changed and no marker was written — safe to run again: %v", err)
|
||||||
|
return
|
||||||
|
}
|
||||||
|
if res.Skipped {
|
||||||
|
log.Printf("voice: re-embed skipped — the stored vectors were already written by %s", current)
|
||||||
|
return
|
||||||
|
}
|
||||||
|
log.Printf("voice: re-embed done — %d notes in the notes table, %d notes and %d facts in the memory index, took %s; stored vectors now belong to %s",
|
||||||
|
res.Notes, res.MemNotes, res.Facts, res.Took.Round(time.Second), current)
|
||||||
|
|
||||||
|
// A row with no text cannot be re-embedded, so its vector is still the old
|
||||||
|
// model's noise while the marker now says everything is current. Both write
|
||||||
|
// paths always store the text, so this should be zero — say it loudly
|
||||||
|
// rather than bury it in the line above if it ever isn't.
|
||||||
|
if res.NoText > 0 {
|
||||||
|
log.Printf("voice: WARNING %d stored rows had no text, so their vectors could not be re-embedded and are still noise; they will never match anything useful (Vikunja #378)", res.NoText)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|||||||
+5
-3
@@ -6,11 +6,12 @@
|
|||||||
"state_dir": "/var/lib/maven",
|
"state_dir": "/var/lib/maven",
|
||||||
|
|
||||||
"phraser": {
|
"phraser": {
|
||||||
"model_path": "/opt/maven/models/llm/qwen3.5/Qwen3.5-0.8B.Q4_K_M.gguf",
|
"model_path": "/opt/maven/models/llm/qwen3/Qwen3-1.7B-UD-Q4_K_XL.gguf",
|
||||||
"bin_path": "llama-server",
|
"bin_path": "llama-server",
|
||||||
"n_gpu_layers": 99,
|
"n_gpu_layers": 99,
|
||||||
"n_ctx": 2048,
|
"n_ctx": 4096,
|
||||||
"timeout": "60s"
|
"timeout": "60s",
|
||||||
|
"llm_nudges": false
|
||||||
},
|
},
|
||||||
|
|
||||||
"telegram": {
|
"telegram": {
|
||||||
@@ -43,6 +44,7 @@
|
|||||||
"llm_router": true,
|
"llm_router": true,
|
||||||
"query_min_score": 0.55,
|
"query_min_score": 0.55,
|
||||||
"query_min_margin": 0.008,
|
"query_min_margin": 0.008,
|
||||||
|
"clarify_max_attempts": 3,
|
||||||
"tool_timeout": "30s",
|
"tool_timeout": "30s",
|
||||||
"tools": [
|
"tools": [
|
||||||
{ "name": "status", "cmd": ["systemctl", "status"], "scope": "homelab", "destructive": false },
|
{ "name": "status", "cmd": ["systemctl", "status"], "scope": "homelab", "destructive": false },
|
||||||
|
|||||||
@@ -292,12 +292,23 @@ type VoiceConfig struct {
|
|||||||
// Negative ⇒ off. 0 ⇒ the default below.
|
// Negative ⇒ off. 0 ⇒ the default below.
|
||||||
QueryMinMargin float64 `json:"query_min_margin,omitempty"`
|
QueryMinMargin float64 `json:"query_min_margin,omitempty"`
|
||||||
|
|
||||||
|
// ClarifyMaxAttempts — how many clarifying questions she may ask about one
|
||||||
|
// request before she gives up and says she did not understand. Default 3.
|
||||||
|
ClarifyMaxAttempts int `json:"clarify_max_attempts,omitempty"`
|
||||||
|
|
||||||
// Persona — optional prompt prefix that tunes maven's character. Prepended
|
// Persona — optional prompt prefix that tunes maven's character. Prepended
|
||||||
// to every LLM system prompt (nudge phrasing, note queries, general
|
// to every LLM system prompt (nudge phrasing, note queries, general
|
||||||
// knowledge). Empty string ⇒ current hardcoded persona (feminine-gendered
|
// knowledge). Empty string ⇒ current hardcoded persona (feminine-gendered
|
||||||
// Russian self-reference). Example: "Be formal and answer in English only."
|
// Russian self-reference). Example: "Be formal and answer in English only."
|
||||||
Persona string `json:"persona,omitempty"`
|
Persona string `json:"persona,omitempty"`
|
||||||
|
|
||||||
|
// OwnerName / City — optional facts about the owner, added to the shared
|
||||||
|
// context block (internal/persona). Empty is fine: the block still states
|
||||||
|
// who he is grammatically (a man, addressed as "ты") and the current time.
|
||||||
|
// Nothing about correct behaviour may depend on these being filled in.
|
||||||
|
OwnerName string `json:"owner_name,omitempty"`
|
||||||
|
City string `json:"city,omitempty"`
|
||||||
|
|
||||||
// Weather — the weather provider config. nil ⇒ the daemon wires
|
// Weather — the weather provider config. nil ⇒ the daemon wires
|
||||||
// the stub provider (returns ErrNotConfigured — "погода не настроена").
|
// the stub provider (returns ErrNotConfigured — "погода не настроена").
|
||||||
// Set provider to "open-meteo" to use the keyless Open-Meteo API.
|
// Set provider to "open-meteo" to use the keyless Open-Meteo API.
|
||||||
@@ -358,6 +369,12 @@ type PhraserConfig struct {
|
|||||||
NGpuLayers int `json:"n_gpu_layers,omitempty"`
|
NGpuLayers int `json:"n_gpu_layers,omitempty"`
|
||||||
NCtx int `json:"n_ctx,omitempty"`
|
NCtx int `json:"n_ctx,omitempty"`
|
||||||
Timeout Duration `json:"timeout,omitempty"`
|
Timeout Duration `json:"timeout,omitempty"`
|
||||||
|
|
||||||
|
// LLMNudges — let the model word nudges again. Off by default: nudges are
|
||||||
|
// worded from hand-written Russian templates now (the model broke the
|
||||||
|
// persona and invented units). Chat, query and reminder phrasing always go
|
||||||
|
// through the model regardless. See phraser.Config.LLMNudges.
|
||||||
|
LLMNudges bool `json:"llm_nudges,omitempty"`
|
||||||
}
|
}
|
||||||
|
|
||||||
// EmbedderConfig — paths for the ONNX multilingual embedder. The daemon
|
// EmbedderConfig — paths for the ONNX multilingual embedder. The daemon
|
||||||
@@ -424,7 +441,9 @@ const (
|
|||||||
// false recall from 5/5 to 1/5. Every larger delta costs real recall
|
// false recall from 5/5 to 1/5. Every larger delta costs real recall
|
||||||
// without removing that last one until 0.020, which drops recall to 44%.
|
// without removing that last one until 0.020, which drops recall to 44%.
|
||||||
DefaultQueryMinMargin = 0.008
|
DefaultQueryMinMargin = 0.008
|
||||||
DefaultToolTimeout = 30 * time.Second
|
// DefaultClarifyMaxAttempts — see dialogue.DefaultMaxAttempts.
|
||||||
|
DefaultClarifyMaxAttempts = 3
|
||||||
|
DefaultToolTimeout = 30 * time.Second
|
||||||
// DefaultLLMRouter — route with the resident model unless told otherwise.
|
// DefaultLLMRouter — route with the resident model unless told otherwise.
|
||||||
DefaultLLMRouter = true
|
DefaultLLMRouter = true
|
||||||
|
|
||||||
@@ -516,6 +535,9 @@ func (c *Config) applyDefaults() {
|
|||||||
case c.Voice.QueryMinMargin < 0:
|
case c.Voice.QueryMinMargin < 0:
|
||||||
c.Voice.QueryMinMargin = 0
|
c.Voice.QueryMinMargin = 0
|
||||||
}
|
}
|
||||||
|
if c.Voice.ClarifyMaxAttempts <= 0 {
|
||||||
|
c.Voice.ClarifyMaxAttempts = DefaultClarifyMaxAttempts
|
||||||
|
}
|
||||||
if c.Voice.ToolTimeout <= 0 {
|
if c.Voice.ToolTimeout <= 0 {
|
||||||
c.Voice.ToolTimeout = Duration(DefaultToolTimeout)
|
c.Voice.ToolTimeout = Duration(DefaultToolTimeout)
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -35,6 +35,27 @@ func TestLoadDefaults(t *testing.T) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Nudges come from templates unless the config says otherwise.
|
||||||
|
func TestPhraserLLMNudgesDefaultsOff(t *testing.T) {
|
||||||
|
p := writeConfig(t, `{"phraser":{"model_path":"/tmp/m.gguf"}}`)
|
||||||
|
c, err := Load(p)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("Load: %v", err)
|
||||||
|
}
|
||||||
|
if c.Phraser.LLMNudges {
|
||||||
|
t.Error("llm_nudges defaults on; templates must be the default")
|
||||||
|
}
|
||||||
|
|
||||||
|
p = writeConfig(t, `{"phraser":{"model_path":"/tmp/m.gguf","llm_nudges":true}}`)
|
||||||
|
c, err = Load(p)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("Load: %v", err)
|
||||||
|
}
|
||||||
|
if !c.Phraser.LLMNudges {
|
||||||
|
t.Error("llm_nudges:true did not parse")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
func TestLoadDurationsParse(t *testing.T) {
|
func TestLoadDurationsParse(t *testing.T) {
|
||||||
p := writeConfig(t, `{"tick_interval":"90s","repeat_interval":"10m"}`)
|
p := writeConfig(t, `{"tick_interval":"90s","repeat_interval":"10m"}`)
|
||||||
c, err := Load(p)
|
c, err := Load(p)
|
||||||
|
|||||||
@@ -137,9 +137,10 @@ func NewDispatcher(cfg Config) *Dispatcher {
|
|||||||
// picks for (severity, presence), sends via the matching sink, and records
|
// picks for (severity, presence), sends via the matching sink, and records
|
||||||
// one nudge row per successful send. returns the dispatches (one per channel).
|
// one nudge row per successful send. returns the dispatches (one per channel).
|
||||||
//
|
//
|
||||||
// a Drop channel = no send, no record (the nudge was suppressed by routing,
|
// a Drop channel = no send (the nudge was suppressed by routing, not by a
|
||||||
// not by a failure — "a missed water nudge is noise"). a nil sink = channel
|
// failure — "a missed water nudge is noise"), but it does leave a 'dropped'
|
||||||
// not wired, skip silently. a send error stops the dispatch and returns what
|
// outbox row so the suppression is visible. a nil sink = channel not wired,
|
||||||
|
// skip silently. a send error stops the dispatch and returns what
|
||||||
// got through — the daemon decides whether to retry.
|
// got through — the daemon decides whether to retry.
|
||||||
func (d *Dispatcher) DispatchNudge(ctx context.Context, pn PhrasedNudge, now time.Time) ([]Dispatch, error) {
|
func (d *Dispatcher) DispatchNudge(ctx context.Context, pn PhrasedNudge, now time.Time) ([]Dispatch, error) {
|
||||||
c := pn.Candidate
|
c := pn.Candidate
|
||||||
@@ -148,6 +149,16 @@ func (d *Dispatcher) DispatchNudge(ctx context.Context, pn PhrasedNudge, now tim
|
|||||||
for i := 0; i < len(channels); i++ {
|
for i := 0; i < len(channels); i++ {
|
||||||
ch := channels[i]
|
ch := channels[i]
|
||||||
if ch == ChannelDrop {
|
if ch == ChannelDrop {
|
||||||
|
// the routing table suppressed this nudge on purpose (a care nudge
|
||||||
|
// while you're away is noise). that stays — but it must not be
|
||||||
|
// invisible, or "she dropped it" and "the rule never fired" look
|
||||||
|
// the same afterwards. no nudges row: that table feeds the
|
||||||
|
// ignored_rate signal, and a nudge nobody could see must not
|
||||||
|
// count as ignored.
|
||||||
|
id := d.beginOutbox(ctx, "nudge", c.Rule.Name, 0, ch, pn.Summary, now)
|
||||||
|
d.completeOutbox(ctx, id, store.DeliveryDropped, now)
|
||||||
|
log.Printf("dispatcher: dropped %s (sev%d, presence=%s) — routing table suppressed it",
|
||||||
|
c.Rule.Name, c.Severity, c.State.Presence)
|
||||||
continue
|
continue
|
||||||
}
|
}
|
||||||
s := Sendable{
|
s := Sendable{
|
||||||
@@ -396,6 +407,13 @@ func messageForChannel(s Sendable) string {
|
|||||||
if !isAway(s.Channel) {
|
if !isAway(s.Channel) {
|
||||||
return s.Body
|
return s.Body
|
||||||
}
|
}
|
||||||
|
return AwayMessage(s)
|
||||||
|
}
|
||||||
|
|
||||||
|
// AwayMessage — the only text an off-box channel may ever carry. Exported so
|
||||||
|
// the away sinks share this one rule instead of each inventing a fallback: the
|
||||||
|
// summary if we have one, otherwise a fixed generic line. Never the body.
|
||||||
|
func AwayMessage(s Sendable) string {
|
||||||
if s.Summary != "" {
|
if s.Summary != "" {
|
||||||
return s.Summary
|
return s.Summary
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -37,10 +37,12 @@ func TestVoiceNoSessionFallthroughLeavesOutboxTrail(t *testing.T) {
|
|||||||
[]string{"voice", "ntfy"}, []string{store.DeliveryFailed, store.DeliverySent}},
|
[]string{"voice", "ntfy"}, []string{store.DeliveryFailed, store.DeliverySent}},
|
||||||
{"sev4 falls through to telegram", loop.Sev4,
|
{"sev4 falls through to telegram", loop.Sev4,
|
||||||
[]string{"voice", "telegram"}, []string{store.DeliveryFailed, store.DeliverySent}},
|
[]string{"voice", "telegram"}, []string{store.DeliveryFailed, store.DeliverySent}},
|
||||||
{"sev1 does not fall through", loop.Sev1,
|
// care severities still don't reach an away channel; since #370 the
|
||||||
[]string{"voice"}, []string{store.DeliveryFailed}},
|
// drop itself is a visible row instead of nothing.
|
||||||
{"sev2 does not fall through", loop.Sev2,
|
{"sev1 drops instead of falling through", loop.Sev1,
|
||||||
[]string{"voice"}, []string{store.DeliveryFailed}},
|
[]string{"voice", "drop"}, []string{store.DeliveryFailed, store.DeliveryDropped}},
|
||||||
|
{"sev2 drops instead of falling through", loop.Sev2,
|
||||||
|
[]string{"voice", "drop"}, []string{store.DeliveryFailed, store.DeliveryDropped}},
|
||||||
}
|
}
|
||||||
for _, c := range cases {
|
for _, c := range cases {
|
||||||
t.Run(c.name, func(t *testing.T) {
|
t.Run(c.name, func(t *testing.T) {
|
||||||
|
|||||||
@@ -2,10 +2,10 @@
|
|||||||
//
|
//
|
||||||
// ntfy is the away-channel for sev3 (ops soft) nudges, sev4 (ops hard)
|
// ntfy is the away-channel for sev3 (ops soft) nudges, sev4 (ops hard)
|
||||||
// nudges when present (alongside voice), and reminders when away. the
|
// nudges when present (alongside voice), and reminders when away. the
|
||||||
// message body is the Sendable's Summary — the minimal-body rule from the
|
// message body is delivery.AwayMessage — the minimal-body rule from the
|
||||||
// spec ("disk low on homesrv," not detail; no shoulder-surf exfil through
|
// spec ("disk low on homesrv," not detail; no shoulder-surf exfil through
|
||||||
// the relay). voice gets Body; away channels get Summary, enforced at the
|
// the relay). the dispatcher already strips detail off away sendables; the
|
||||||
// sink so a phraser bug can't exfil.
|
// sink uses the same helper so it can't leak the body on its own either.
|
||||||
//
|
//
|
||||||
// ntfy runs locally (docker, 127.0.0.1:8085, deny-all auth). maven publishes
|
// ntfy runs locally (docker, 127.0.0.1:8085, deny-all auth). maven publishes
|
||||||
// with a dedicated user (write-only to maven-* topics) — the credential is a
|
// with a dedicated user (write-only to maven-* topics) — the credential is a
|
||||||
@@ -69,18 +69,14 @@ func New(cfg Config) (*Sink, error) {
|
|||||||
}, nil
|
}, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
// Send publishes one notification to ntfy. the body is the Sendable's Summary
|
// Send publishes one notification to ntfy. the body is the minimal away
|
||||||
// (minimal body); Title is "maven" (consistent sender identity on the lock
|
// message (never the full body); Title is "maven" (consistent sender identity
|
||||||
// screen — the content is in the body). Priority maps from severity/kind so
|
// on the lock screen — the content is in the body). Priority maps from severity/kind so
|
||||||
// the phone client can ring differently for an alarm vs a soft ops nudge.
|
// the phone client can ring differently for an alarm vs a soft ops nudge.
|
||||||
func (s *Sink) Send(ctx context.Context, d delivery.Sendable) error {
|
func (s *Sink) Send(ctx context.Context, d delivery.Sendable) error {
|
||||||
body := d.Summary
|
// never fall back to d.Body: ntfy leaves the box, so an empty summary gets
|
||||||
if body == "" {
|
// a generic line instead of the full detail.
|
||||||
body = d.Body // terse full message beats no message
|
body := delivery.AwayMessage(d)
|
||||||
}
|
|
||||||
if body == "" {
|
|
||||||
return fmt.Errorf("ntfysink: empty message for %s", d.Channel)
|
|
||||||
}
|
|
||||||
|
|
||||||
req, err := http.NewRequestWithContext(ctx, http.MethodPost, s.topicURL(), strings.NewReader(body))
|
req, err := http.NewRequestWithContext(ctx, http.MethodPost, s.topicURL(), strings.NewReader(body))
|
||||||
if err != nil {
|
if err != nil {
|
||||||
|
|||||||
@@ -147,9 +147,9 @@ func TestSendBodyIsSummaryNotFullBody(t *testing.T) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
func TestSendFallsBackToBodyWhenSummaryEmpty(t *testing.T) {
|
func TestSendNeverSendsTheBodyWhenSummaryEmpty(t *testing.T) {
|
||||||
// a terse full message is better than no message; the phraser should
|
// #368: this used to fall back to the full body. ntfy leaves the box, so
|
||||||
// produce a summary for away-bound severities, but don't silently drop.
|
// an empty summary gets a fixed generic line plus the rule name instead.
|
||||||
rs := newRecordingServer(t, 200, "")
|
rs := newRecordingServer(t, 200, "")
|
||||||
srv := httptest.NewServer(rs.handler())
|
srv := httptest.NewServer(rs.handler())
|
||||||
defer srv.Close()
|
defer srv.Close()
|
||||||
@@ -160,12 +160,15 @@ func TestSendFallsBackToBodyWhenSummaryEmpty(t *testing.T) {
|
|||||||
t.Fatalf("Send: %v", err)
|
t.Fatalf("Send: %v", err)
|
||||||
}
|
}
|
||||||
_, _, body, _, _, _ := rs.snapshot()
|
_, _, body, _, _, _ := rs.snapshot()
|
||||||
if body != s.Body {
|
want := delivery.GenericAwayMessage + ": service_down"
|
||||||
t.Fatalf("fallback body: want %q, got %q", s.Body, body)
|
if body != want {
|
||||||
|
t.Fatalf("body: want %q, got %q", want, body)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
func TestSendRejectsEmptyMessage(t *testing.T) {
|
func TestSendNeverSendsAnEmptyMessage(t *testing.T) {
|
||||||
|
// with nothing at all to say we still send the generic line — an away
|
||||||
|
// channel can never carry detail, but it also never goes out blank.
|
||||||
rs := newRecordingServer(t, 200, "")
|
rs := newRecordingServer(t, 200, "")
|
||||||
srv := httptest.NewServer(rs.handler())
|
srv := httptest.NewServer(rs.handler())
|
||||||
defer srv.Close()
|
defer srv.Close()
|
||||||
@@ -173,9 +176,13 @@ func TestSendRejectsEmptyMessage(t *testing.T) {
|
|||||||
sink, _ := New(Config{BaseURL: srv.URL, Topic: "maven"})
|
sink, _ := New(Config{BaseURL: srv.URL, Topic: "maven"})
|
||||||
s := nudgeSendable(loop.Sev3, "")
|
s := nudgeSendable(loop.Sev3, "")
|
||||||
s.Body = ""
|
s.Body = ""
|
||||||
err := sink.Send(context.Background(), s)
|
s.RuleName = ""
|
||||||
if err == nil {
|
if err := sink.Send(context.Background(), s); err != nil {
|
||||||
t.Fatal("want error for empty message")
|
t.Fatalf("Send: %v", err)
|
||||||
|
}
|
||||||
|
_, _, body, _, _, _ := rs.snapshot()
|
||||||
|
if body != delivery.GenericAwayMessage {
|
||||||
|
t.Fatalf("body: want %q, got %q", delivery.GenericAwayMessage, body)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -208,10 +208,9 @@ func TestAwayChannelsGetMinimalBody(t *testing.T) {
|
|||||||
// TestCareAwayDropIsRecorded — DESIGN.md's drop is a decision ("a missed water
|
// TestCareAwayDropIsRecorded — DESIGN.md's drop is a decision ("a missed water
|
||||||
// nudge is noise, a missed backup failure isn't"), so it should be visible
|
// nudge is noise, a missed backup failure isn't"), so it should be visible
|
||||||
// rather than vanish. Today drop is a bare `continue`: no nudge row, no outbox
|
// rather than vanish. Today drop is a bare `continue`: no nudge row, no outbox
|
||||||
// attempt, no log — nothing an operator can see afterwards.
|
// attempt, no log — nothing an operator can see afterwards. now it leaves a
|
||||||
|
// 'dropped' outbox row.
|
||||||
func TestCareAwayDropIsRecorded(t *testing.T) {
|
func TestCareAwayDropIsRecorded(t *testing.T) {
|
||||||
t.Skip("not implemented: dispatcher.go:149-151 skips a Drop channel with no record; there is no 'dropped' outcome in store/delivery.go:16-21")
|
|
||||||
|
|
||||||
ob := &fakeOutbox{}
|
ob := &fakeOutbox{}
|
||||||
d := NewDispatcher(Config{Voice: &fakeSink{}, Nudges: &fakeNudgeRecorder{}, Outbox: ob})
|
d := NewDispatcher(Config{Voice: &fakeSink{}, Nudges: &fakeNudgeRecorder{}, Outbox: ob})
|
||||||
|
|
||||||
|
|||||||
@@ -2,11 +2,12 @@
|
|||||||
//
|
//
|
||||||
// telegram is the away-channel for sev4 (ops hard) nudges — "disk-fire alarm
|
// telegram is the away-channel for sev4 (ops hard) nudges — "disk-fire alarm
|
||||||
// at 2am routes to telegram, repeat til ack." the message body is the
|
// at 2am routes to telegram, repeat til ack." the message body is the
|
||||||
// Sendable's Summary — the minimal-body rule from the spec ("disk low on
|
// delivery.AwayMessage — the minimal-body rule from the spec ("disk low on
|
||||||
// homesrv," not detail; no shoulder-surf exfil through the relay). voice gets
|
// homesrv," not detail; no shoulder-surf exfil through the relay). the
|
||||||
// Body; away channels get Summary, enforced at the sink so a phraser bug can't
|
// dispatcher already strips detail off away sendables; the sink uses the same
|
||||||
// exfil. additionally, protect_content=true is passed on every send so the
|
// helper so it can't leak the body on its own either. additionally,
|
||||||
// message can't be forwarded out of the chat — locks the minimal body further.
|
// protect_content=true is passed on every send so the message can't be
|
||||||
|
// forwarded out of the chat — locks the minimal body further.
|
||||||
//
|
//
|
||||||
// telegram's bot API is region-restricted for this homesrv — direct egress to
|
// telegram's bot API is region-restricted for this homesrv — direct egress to
|
||||||
// api.telegram.org is unreliable. the spec's "away channels leave the box —
|
// api.telegram.org is unreliable. the spec's "away channels leave the box —
|
||||||
@@ -140,18 +141,13 @@ type telegramResp struct {
|
|||||||
}
|
}
|
||||||
|
|
||||||
// Send publishes one message to the configured telegram chat. the body is the
|
// Send publishes one message to the configured telegram chat. the body is the
|
||||||
// Sendable's Summary (minimal body); empty Summary falls back to Body (terse
|
// minimal away message (never the full body). protect_content=true so even
|
||||||
// full message beats no message). protect_content=true so a phraser bug (Body
|
// that can't be forwarded onward by the user or a chat observer — locks the
|
||||||
// leaking detail through Summary) can't be forwarded onward by the user or a
|
// minimal-body rule at the channel's own last mile.
|
||||||
// chat observer — locks the minimal-body rule at the channel's own last mile.
|
|
||||||
func (s *Sink) Send(ctx context.Context, d delivery.Sendable) error {
|
func (s *Sink) Send(ctx context.Context, d delivery.Sendable) error {
|
||||||
body := d.Summary
|
// never fall back to d.Body: telegram leaves the box, so an empty summary
|
||||||
if body == "" {
|
// gets a generic line instead of the full detail.
|
||||||
body = d.Body
|
body := delivery.AwayMessage(d)
|
||||||
}
|
|
||||||
if body == "" {
|
|
||||||
return fmt.Errorf("telegramsink: empty message for %s", d.Channel)
|
|
||||||
}
|
|
||||||
|
|
||||||
payload := sendMessageReq{
|
payload := sendMessageReq{
|
||||||
ChatID: s.cfg.ChatID,
|
ChatID: s.cfg.ChatID,
|
||||||
|
|||||||
@@ -173,9 +173,9 @@ func TestSendBodyIsSummaryNotFullBody(t *testing.T) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
func TestSendFallsBackToBodyWhenSummaryEmpty(t *testing.T) {
|
func TestSendNeverSendsTheBodyWhenSummaryEmpty(t *testing.T) {
|
||||||
// terse full message beats none; the phraser should produce a summary for
|
// #368: this used to fall back to the full body. telegram leaves the box,
|
||||||
// away-bound severities, but don't silently drop.
|
// so an empty summary gets a fixed generic line plus the rule name.
|
||||||
rs := newRecordingServer(t, 200, "")
|
rs := newRecordingServer(t, 200, "")
|
||||||
srv := httptest.NewServer(rs.handler())
|
srv := httptest.NewServer(rs.handler())
|
||||||
defer srv.Close()
|
defer srv.Close()
|
||||||
@@ -188,12 +188,14 @@ func TestSendFallsBackToBodyWhenSummaryEmpty(t *testing.T) {
|
|||||||
_, _, body, _, _ := rs.snapshot()
|
_, _, body, _, _ := rs.snapshot()
|
||||||
var req sendMessageReq
|
var req sendMessageReq
|
||||||
_ = json.Unmarshal([]byte(body), &req)
|
_ = json.Unmarshal([]byte(body), &req)
|
||||||
if req.Text != s.Body {
|
want := delivery.GenericAwayMessage + ": service_down"
|
||||||
t.Fatalf("fallback text: want %q, got %q", s.Body, req.Text)
|
if req.Text != want {
|
||||||
|
t.Fatalf("text: want %q, got %q", want, req.Text)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
func TestSendRejectsEmptyMessage(t *testing.T) {
|
func TestSendNeverSendsAnEmptyMessage(t *testing.T) {
|
||||||
|
// with nothing at all to say we still send the generic line.
|
||||||
rs := newRecordingServer(t, 200, "")
|
rs := newRecordingServer(t, 200, "")
|
||||||
srv := httptest.NewServer(rs.handler())
|
srv := httptest.NewServer(rs.handler())
|
||||||
defer srv.Close()
|
defer srv.Close()
|
||||||
@@ -201,9 +203,15 @@ func TestSendRejectsEmptyMessage(t *testing.T) {
|
|||||||
sink, _ := New(sinkCfg(srv.URL))
|
sink, _ := New(sinkCfg(srv.URL))
|
||||||
s := nudgeSendable(loop.Sev4, "")
|
s := nudgeSendable(loop.Sev4, "")
|
||||||
s.Body = ""
|
s.Body = ""
|
||||||
err := sink.Send(context.Background(), s)
|
s.RuleName = ""
|
||||||
if err == nil {
|
if err := sink.Send(context.Background(), s); err != nil {
|
||||||
t.Fatal("want error for empty message")
|
t.Fatalf("Send: %v", err)
|
||||||
|
}
|
||||||
|
_, _, body, _, _ := rs.snapshot()
|
||||||
|
var req sendMessageReq
|
||||||
|
_ = json.Unmarshal([]byte(body), &req)
|
||||||
|
if req.Text != delivery.GenericAwayMessage {
|
||||||
|
t.Fatalf("text: want %q, got %q", delivery.GenericAwayMessage, req.Text)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -17,10 +17,11 @@ const (
|
|||||||
SlotText Slot = "text" // Slots.Text
|
SlotText Slot = "text" // Slots.Text
|
||||||
)
|
)
|
||||||
|
|
||||||
// MaxAttempts is 1 because Maven is not a nag (DESIGN.md § Non-goals). She asks
|
// DefaultMaxAttempts — how many questions she may ask about one request.
|
||||||
// one clarifying question. If the answer still leaves the slot empty she drops
|
// Three, because after three tries the likely problem is that she misheard the
|
||||||
// the request instead of asking again.
|
// whole request, not one slot — so another question about that slot won't help.
|
||||||
const MaxAttempts = 1
|
// Configurable: voice.clarify_max_attempts.
|
||||||
|
const DefaultMaxAttempts = 3
|
||||||
|
|
||||||
// PendingQuestion is what Maven holds while she waits for an answer to an open
|
// PendingQuestion is what Maven holds while she waits for an answer to an open
|
||||||
// question. Unlike the yes/no confirms in cmd/mavend/voice.go, the answer here
|
// question. Unlike the yes/no confirms in cmd/mavend/voice.go, the answer here
|
||||||
@@ -32,7 +33,17 @@ type PendingQuestion struct {
|
|||||||
Utterance string // the user's original raw words
|
Utterance string // the user's original raw words
|
||||||
Asked time.Time
|
Asked time.Time
|
||||||
TTL time.Duration
|
TTL time.Duration
|
||||||
Attempts int // questions already asked; capped by MaxAttempts
|
Attempts int // questions already asked
|
||||||
|
// MaxAttempts caps Attempts. 0 ⇒ DefaultMaxAttempts.
|
||||||
|
MaxAttempts int
|
||||||
|
}
|
||||||
|
|
||||||
|
// maxAttempts is MaxAttempts with the default filled in.
|
||||||
|
func (q *PendingQuestion) maxAttempts() int {
|
||||||
|
if q.MaxAttempts <= 0 {
|
||||||
|
return DefaultMaxAttempts
|
||||||
|
}
|
||||||
|
return q.MaxAttempts
|
||||||
}
|
}
|
||||||
|
|
||||||
func (q *PendingQuestion) IsExpired(now time.Time) bool {
|
func (q *PendingQuestion) IsExpired(now time.Time) bool {
|
||||||
@@ -41,12 +52,9 @@ func (q *PendingQuestion) IsExpired(now time.Time) bool {
|
|||||||
|
|
||||||
// CanAsk reports whether Maven may ask another question about this request.
|
// CanAsk reports whether Maven may ask another question about this request.
|
||||||
func (q *PendingQuestion) CanAsk() bool {
|
func (q *PendingQuestion) CanAsk() bool {
|
||||||
return q.Attempts < MaxAttempts
|
return q.Attempts < q.maxAttempts()
|
||||||
}
|
}
|
||||||
|
|
||||||
// TODO: the daemon will phrase the question text from Missing (one short ru
|
|
||||||
// question per Slot, feminine self-reference) and speak it here.
|
|
||||||
|
|
||||||
// ClarifyStore holds the parked questions. Same shape and locking as
|
// ClarifyStore holds the parked questions. Same shape and locking as
|
||||||
// SessionStore: keyed by dialogue id, expired entries dropped on read.
|
// SessionStore: keyed by dialogue id, expired entries dropped on read.
|
||||||
type ClarifyStore struct {
|
type ClarifyStore struct {
|
||||||
@@ -67,8 +75,7 @@ func NewClarifyStore(defaultTTL time.Duration) *ClarifyStore {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// TODO: the daemon will Put a question here when Decision.Clarify fires, in
|
// Put parks a question. Called on a clarify decision (cmd/mavend/clarify.go).
|
||||||
// place of the flat "не разобрала" reply (cmd/mavend/voice.go).
|
|
||||||
func (s *ClarifyStore) Put(id string, q *PendingQuestion) {
|
func (s *ClarifyStore) Put(id string, q *PendingQuestion) {
|
||||||
if q.TTL <= 0 {
|
if q.TTL <= 0 {
|
||||||
q.TTL = s.defaultTTL
|
q.TTL = s.defaultTTL
|
||||||
@@ -78,8 +85,7 @@ func (s *ClarifyStore) Put(id string, q *PendingQuestion) {
|
|||||||
s.mu.Unlock()
|
s.mu.Unlock()
|
||||||
}
|
}
|
||||||
|
|
||||||
// TODO: the daemon will Get on the next turn, parse that turn into Slots, call
|
// Get returns the live parked question, or nil when there is none.
|
||||||
// Answer, and Delete — the open-question twin of resolveConfirm.
|
|
||||||
func (s *ClarifyStore) Get(id string, now time.Time) *PendingQuestion {
|
func (s *ClarifyStore) Get(id string, now time.Time) *PendingQuestion {
|
||||||
s.mu.RLock()
|
s.mu.RLock()
|
||||||
q, ok := s.questions[id]
|
q, ok := s.questions[id]
|
||||||
@@ -94,6 +100,21 @@ func (s *ClarifyStore) Get(id string, now time.Time) *PendingQuestion {
|
|||||||
return q
|
return q
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// TakeExpired reports whether a question was parked here but its TTL ran out,
|
||||||
|
// and drops it. Get drops such a question silently, which leaves the user
|
||||||
|
// thinking his request is still alive — the caller uses this to tell him it is
|
||||||
|
// gone before treating his words as a fresh utterance.
|
||||||
|
func (s *ClarifyStore) TakeExpired(id string, now time.Time) bool {
|
||||||
|
s.mu.Lock()
|
||||||
|
defer s.mu.Unlock()
|
||||||
|
q, ok := s.questions[id]
|
||||||
|
if !ok || !q.IsExpired(now) {
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
delete(s.questions, id)
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
|
||||||
func (s *ClarifyStore) Delete(id string) {
|
func (s *ClarifyStore) Delete(id string) {
|
||||||
s.mu.Lock()
|
s.mu.Lock()
|
||||||
delete(s.questions, id)
|
delete(s.questions, id)
|
||||||
@@ -101,44 +122,45 @@ func (s *ClarifyStore) Delete(id string) {
|
|||||||
}
|
}
|
||||||
|
|
||||||
// Answer merges the slots parsed from the user's answer into the parked ones.
|
// Answer merges the slots parsed from the user's answer into the parked ones.
|
||||||
// Only the slots listed in Missing are filled, and an already filled slot is
|
// Only the slots listed in Missing are touched. Within those, a value the answer
|
||||||
// never overwritten — the answer completes the original request, it does not
|
// carries WINS over what was parked: she asked about this slot, so «нет, в пять»
|
||||||
// restate it. Parsing the answer text into `answer` is the caller's job; this
|
// after «в три» must replace the time, not be thrown away.
|
||||||
// package must stay free of internal/router.
|
//
|
||||||
|
// This is the clarify answer only. A correction in a fresh turn ("вообще-то
|
||||||
|
// перенеси на пять") is a different code path (followUpMerge) — not here.
|
||||||
|
//
|
||||||
|
// Parsing the answer text into `answer` is the caller's job; this package must
|
||||||
|
// stay free of internal/router.
|
||||||
func (q *PendingQuestion) Answer(text string, answer Slots) Slots {
|
func (q *PendingQuestion) Answer(text string, answer Slots) Slots {
|
||||||
out := q.Slots
|
out := q.Slots
|
||||||
for _, slot := range q.Missing {
|
for _, slot := range q.Missing {
|
||||||
switch slot {
|
switch slot {
|
||||||
case SlotTime:
|
case SlotTime:
|
||||||
if !out.HasTime && answer.HasTime {
|
if answer.HasTime {
|
||||||
out.Time = answer.Time
|
out.Time = answer.Time
|
||||||
out.HasTime = true
|
out.HasTime = true
|
||||||
}
|
}
|
||||||
case SlotKey:
|
case SlotKey:
|
||||||
if !out.HasKey && answer.HasKey {
|
if answer.HasKey {
|
||||||
out.Key = answer.Key
|
out.Key = answer.Key
|
||||||
out.HasKey = true
|
out.HasKey = true
|
||||||
}
|
}
|
||||||
case SlotValue:
|
case SlotValue:
|
||||||
if out.Value == "" && answer.Value != "" {
|
if answer.Value != "" {
|
||||||
out.Value = answer.Value
|
out.Value = answer.Value
|
||||||
}
|
}
|
||||||
case SlotFn:
|
case SlotFn:
|
||||||
if !out.HasFn && answer.HasFn {
|
if answer.HasFn {
|
||||||
out.Fn = answer.Fn
|
out.Fn = answer.Fn
|
||||||
out.HasFn = true
|
out.HasFn = true
|
||||||
if len(out.Args) == 0 {
|
out.Args = append([]string(nil), answer.Args...)
|
||||||
out.Args = append([]string(nil), answer.Args...)
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
case SlotText:
|
case SlotText:
|
||||||
if out.Text == "" {
|
if answer.Text != "" {
|
||||||
if answer.Text != "" {
|
out.Text = answer.Text
|
||||||
out.Text = answer.Text
|
} else if out.Text == "" {
|
||||||
} else {
|
// No parse for a text slot — the raw answer IS the text.
|
||||||
// No parse for a text slot — the raw answer IS the text.
|
out.Text = text
|
||||||
out.Text = text
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -1,6 +1,7 @@
|
|||||||
package dialogue
|
package dialogue
|
||||||
|
|
||||||
import (
|
import (
|
||||||
|
"reflect"
|
||||||
"testing"
|
"testing"
|
||||||
"time"
|
"time"
|
||||||
)
|
)
|
||||||
@@ -59,6 +60,28 @@ func TestClarifyStoreGetPutDelete(t *testing.T) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// TestClarifyStoreTakeExpired — TakeExpired reports (and drops) only a question
|
||||||
|
// whose TTL ran out.
|
||||||
|
func TestClarifyStoreTakeExpired(t *testing.T) {
|
||||||
|
s := NewClarifyStore(time.Minute)
|
||||||
|
if s.TakeExpired("voice", base) {
|
||||||
|
t.Fatal("nothing parked ⇒ nothing expired")
|
||||||
|
}
|
||||||
|
s.Put("voice", &PendingQuestion{Missing: []Slot{SlotTime}, Asked: base, TTL: time.Minute})
|
||||||
|
if s.TakeExpired("voice", base.Add(30*time.Second)) {
|
||||||
|
t.Fatal("a live question must not report as expired")
|
||||||
|
}
|
||||||
|
if s.Get("voice", base.Add(30*time.Second)) == nil {
|
||||||
|
t.Fatal("a live question must survive TakeExpired")
|
||||||
|
}
|
||||||
|
if !s.TakeExpired("voice", base.Add(2*time.Minute)) {
|
||||||
|
t.Fatal("a stale question must report as expired")
|
||||||
|
}
|
||||||
|
if s.TakeExpired("voice", base.Add(2*time.Minute)) {
|
||||||
|
t.Fatal("TakeExpired must drop the question, so the second call is false")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
func TestNewClarifyStoreDefaultTTL(t *testing.T) {
|
func TestNewClarifyStoreDefaultTTL(t *testing.T) {
|
||||||
s := NewClarifyStore(0)
|
s := NewClarifyStore(0)
|
||||||
q := &PendingQuestion{Asked: base}
|
q := &PendingQuestion{Asked: base}
|
||||||
@@ -89,12 +112,13 @@ func TestAnswerFillsOnlyMissingSlots(t *testing.T) {
|
|||||||
want: Slots{Text: "напомни позвонить", Time: answerTime, HasTime: true},
|
want: Slots{Text: "напомни позвонить", Time: answerTime, HasTime: true},
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
name: "does not overwrite a filled time",
|
// He restated it: «нет, в пять». The new value wins.
|
||||||
|
name: "a restated time overwrites the parked one",
|
||||||
parked: Slots{Time: other, HasTime: true},
|
parked: Slots{Time: other, HasTime: true},
|
||||||
missing: []Slot{SlotTime},
|
missing: []Slot{SlotTime},
|
||||||
text: "в три",
|
text: "нет, в три",
|
||||||
answer: Slots{Time: answerTime, HasTime: true},
|
answer: Slots{Time: answerTime, HasTime: true},
|
||||||
want: Slots{Time: other, HasTime: true},
|
want: Slots{Time: answerTime, HasTime: true},
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
name: "ignores slots that were not missing",
|
name: "ignores slots that were not missing",
|
||||||
@@ -121,12 +145,12 @@ func TestAnswerFillsOnlyMissingSlots(t *testing.T) {
|
|||||||
want: Slots{Fn: "restart", Args: []string{"nginx"}, HasFn: true},
|
want: Slots{Fn: "restart", Args: []string{"nginx"}, HasFn: true},
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
name: "keeps existing args when fn was already known",
|
name: "a restated fn replaces the fn and its args",
|
||||||
parked: Slots{Fn: "restart", Args: []string{"nginx"}, HasFn: true},
|
parked: Slots{Fn: "restart", Args: []string{"nginx"}, HasFn: true},
|
||||||
missing: []Slot{SlotFn},
|
missing: []Slot{SlotFn},
|
||||||
text: "останови postgres",
|
text: "останови postgres",
|
||||||
answer: Slots{Fn: "stop", Args: []string{"postgres"}, HasFn: true},
|
answer: Slots{Fn: "stop", Args: []string{"postgres"}, HasFn: true},
|
||||||
want: Slots{Fn: "restart", Args: []string{"nginx"}, HasFn: true},
|
want: Slots{Fn: "stop", Args: []string{"postgres"}, HasFn: true},
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
name: "raw answer becomes the text when nothing was parsed",
|
name: "raw answer becomes the text when nothing was parsed",
|
||||||
@@ -157,36 +181,34 @@ func TestAnswerFillsOnlyMissingSlots(t *testing.T) {
|
|||||||
for _, tc := range cases {
|
for _, tc := range cases {
|
||||||
t.Run(tc.name, func(t *testing.T) {
|
t.Run(tc.name, func(t *testing.T) {
|
||||||
q := &PendingQuestion{Slots: tc.parked, Missing: tc.missing, Asked: base}
|
q := &PendingQuestion{Slots: tc.parked, Missing: tc.missing, Asked: base}
|
||||||
got := q.Answer(tc.text, tc.answer)
|
// Whole-struct compare: a new field in Slots is covered for free.
|
||||||
if got.Time != tc.want.Time || got.HasTime != tc.want.HasTime ||
|
if got := q.Answer(tc.text, tc.answer); !reflect.DeepEqual(got, tc.want) {
|
||||||
got.Key != tc.want.Key || got.HasKey != tc.want.HasKey ||
|
|
||||||
got.Value != tc.want.Value || got.Text != tc.want.Text ||
|
|
||||||
got.Fn != tc.want.Fn || got.HasFn != tc.want.HasFn {
|
|
||||||
t.Fatalf("Answer = %+v, want %+v", got, tc.want)
|
t.Fatalf("Answer = %+v, want %+v", got, tc.want)
|
||||||
}
|
}
|
||||||
if len(got.Args) != len(tc.want.Args) {
|
|
||||||
t.Fatalf("Args = %v, want %v", got.Args, tc.want.Args)
|
|
||||||
}
|
|
||||||
for i := range got.Args {
|
|
||||||
if got.Args[i] != tc.want.Args[i] {
|
|
||||||
t.Fatalf("Args = %v, want %v", got.Args, tc.want.Args)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
})
|
})
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
func TestCanAskCapsAtOneQuestion(t *testing.T) {
|
func TestCanAskAllowsThreeQuestionsByDefault(t *testing.T) {
|
||||||
if MaxAttempts != 1 {
|
if DefaultMaxAttempts != 3 {
|
||||||
t.Fatalf("MaxAttempts = %d, want 1 (Maven asks once, she is not a nag)", MaxAttempts)
|
t.Fatalf("DefaultMaxAttempts = %d, want 3", DefaultMaxAttempts)
|
||||||
}
|
}
|
||||||
q := &PendingQuestion{Asked: base}
|
q := &PendingQuestion{Asked: base} // MaxAttempts unset ⇒ the default
|
||||||
if !q.CanAsk() {
|
for i := 0; i < 3; i++ {
|
||||||
t.Fatal("a fresh question should be askable")
|
if !q.CanAsk() {
|
||||||
|
t.Fatalf("question %d should be allowed", i+1)
|
||||||
|
}
|
||||||
|
q.Attempts++
|
||||||
}
|
}
|
||||||
q.Attempts = MaxAttempts
|
|
||||||
if q.CanAsk() {
|
if q.CanAsk() {
|
||||||
t.Fatal("the question should not be asked twice")
|
t.Fatal("a fourth question must not be allowed")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestCanAskHonoursConfiguredMax(t *testing.T) {
|
||||||
|
q := &PendingQuestion{Asked: base, MaxAttempts: 1, Attempts: 1}
|
||||||
|
if q.CanAsk() {
|
||||||
|
t.Fatal("MaxAttempts 1 means one question only")
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -1,8 +1,12 @@
|
|||||||
package dialogue
|
package dialogue
|
||||||
|
|
||||||
import (
|
import (
|
||||||
|
"context"
|
||||||
|
"encoding/json"
|
||||||
"sync"
|
"sync"
|
||||||
"time"
|
"time"
|
||||||
|
|
||||||
|
"github.com/kami/maven/internal/store"
|
||||||
)
|
)
|
||||||
|
|
||||||
type Intent string
|
type Intent string
|
||||||
@@ -50,10 +54,22 @@ func (s *Session) IsExpired(now time.Time) bool {
|
|||||||
return now.After(s.Timestamp.Add(s.TTL))
|
return now.After(s.Timestamp.Add(s.TTL))
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// SessionPersister — the bit of the store the session needs, as an interface
|
||||||
|
// so tests can swap it out. Data is an opaque blob: the store never looks
|
||||||
|
// inside, we encode the session as JSON here.
|
||||||
|
type SessionPersister interface {
|
||||||
|
SaveDialogueSession(ctx context.Context, id string, data []byte, ts time.Time, ttl time.Duration) error
|
||||||
|
DeleteDialogueSession(ctx context.Context, id string) error
|
||||||
|
LoadDialogueSessions(ctx context.Context, now time.Time) ([]store.DialogueSessionRow, error)
|
||||||
|
}
|
||||||
|
|
||||||
|
// SessionStore keeps the live sessions in a map (the fast path) and mirrors
|
||||||
|
// every write to the persister, so a daemon restart can load them back.
|
||||||
type SessionStore struct {
|
type SessionStore struct {
|
||||||
mu sync.RWMutex
|
mu sync.RWMutex
|
||||||
sessions map[string]*Session
|
sessions map[string]*Session
|
||||||
defaultTTL time.Duration
|
defaultTTL time.Duration
|
||||||
|
persist SessionPersister // may be nil: memory only (tests, no-store paths)
|
||||||
}
|
}
|
||||||
|
|
||||||
func NewSessionStore(defaultTTL time.Duration) *SessionStore {
|
func NewSessionStore(defaultTTL time.Duration) *SessionStore {
|
||||||
@@ -66,6 +82,43 @@ func NewSessionStore(defaultTTL time.Duration) *SessionStore {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// NewPersistentSessionStore — same store, but writes also go to the DB.
|
||||||
|
// Call Load once after this to bring back sessions from a previous run.
|
||||||
|
func NewPersistentSessionStore(defaultTTL time.Duration, p SessionPersister) *SessionStore {
|
||||||
|
s := NewSessionStore(defaultTTL)
|
||||||
|
s.persist = p
|
||||||
|
return s
|
||||||
|
}
|
||||||
|
|
||||||
|
// Load — read the saved sessions back into memory. Anything past its TTL is
|
||||||
|
// dropped (and deleted from the DB by the store), never revived.
|
||||||
|
func (s *SessionStore) Load(ctx context.Context, now time.Time) error {
|
||||||
|
if s.persist == nil {
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
rows, err := s.persist.LoadDialogueSessions(ctx, now)
|
||||||
|
if err != nil {
|
||||||
|
return err
|
||||||
|
}
|
||||||
|
s.mu.Lock()
|
||||||
|
defer s.mu.Unlock()
|
||||||
|
for _, r := range rows {
|
||||||
|
var sess Session
|
||||||
|
if err := json.Unmarshal(r.Data, &sess); err != nil {
|
||||||
|
// A blob we can't read is not worth failing a startup over.
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if sess.TTL <= 0 {
|
||||||
|
sess.TTL = r.TTL
|
||||||
|
}
|
||||||
|
if sess.IsExpired(now) {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
s.sessions[r.ID] = &sess
|
||||||
|
}
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
|
||||||
func (s *SessionStore) Get(id string, now time.Time) *Session {
|
func (s *SessionStore) Get(id string, now time.Time) *Session {
|
||||||
s.mu.RLock()
|
s.mu.RLock()
|
||||||
sess, ok := s.sessions[id]
|
sess, ok := s.sessions[id]
|
||||||
@@ -87,12 +140,33 @@ func (s *SessionStore) Put(id string, sess *Session) {
|
|||||||
s.mu.Lock()
|
s.mu.Lock()
|
||||||
s.sessions[id] = sess
|
s.sessions[id] = sess
|
||||||
s.mu.Unlock()
|
s.mu.Unlock()
|
||||||
|
s.save(id, sess)
|
||||||
}
|
}
|
||||||
|
|
||||||
func (s *SessionStore) Delete(id string) {
|
func (s *SessionStore) Delete(id string) {
|
||||||
s.mu.Lock()
|
s.mu.Lock()
|
||||||
delete(s.sessions, id)
|
delete(s.sessions, id)
|
||||||
s.mu.Unlock()
|
s.mu.Unlock()
|
||||||
|
if s.persist != nil {
|
||||||
|
_ = s.persist.DeleteDialogueSession(context.Background(), id)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// save — mirror one session to the DB. Best effort: memory already has it, so
|
||||||
|
// a write error costs us the restart safety net, not the current turn.
|
||||||
|
func (s *SessionStore) save(id string, sess *Session) {
|
||||||
|
if s.persist == nil {
|
||||||
|
return
|
||||||
|
}
|
||||||
|
data, err := json.Marshal(sess)
|
||||||
|
if err != nil {
|
||||||
|
return
|
||||||
|
}
|
||||||
|
ts := sess.Timestamp
|
||||||
|
if ts.IsZero() {
|
||||||
|
ts = time.Now()
|
||||||
|
}
|
||||||
|
_ = s.persist.SaveDialogueSession(context.Background(), id, data, ts, sess.TTL)
|
||||||
}
|
}
|
||||||
|
|
||||||
func InheritSlots(prev, cur Slots) Slots {
|
func InheritSlots(prev, cur Slots) Slots {
|
||||||
|
|||||||
@@ -0,0 +1,118 @@
|
|||||||
|
package dialogue
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"path/filepath"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"github.com/kami/maven/internal/store"
|
||||||
|
)
|
||||||
|
|
||||||
|
// openStore — a store on disk, so a second handle can reopen the same file.
|
||||||
|
func openStore(t *testing.T, path string) *store.Store {
|
||||||
|
t.Helper()
|
||||||
|
s, err := store.Open(context.Background(), path)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("open store: %v", err)
|
||||||
|
}
|
||||||
|
t.Cleanup(func() { _ = s.Close() })
|
||||||
|
return s
|
||||||
|
}
|
||||||
|
|
||||||
|
// A session written before a restart comes back and still merges a follow-up.
|
||||||
|
func TestSessionSurvivesRestart(t *testing.T) {
|
||||||
|
ctx := context.Background()
|
||||||
|
path := filepath.Join(t.TempDir(), "maven_test.db")
|
||||||
|
now := time.Now().UTC().Truncate(time.Millisecond)
|
||||||
|
|
||||||
|
first := openStore(t, path)
|
||||||
|
before := NewPersistentSessionStore(2*time.Minute, first)
|
||||||
|
before.Put("voice", &Session{
|
||||||
|
Intent: IntentReminder,
|
||||||
|
Slots: Slots{Text: "полить цветы", Time: now.Add(time.Hour), HasTime: true},
|
||||||
|
Timestamp: now,
|
||||||
|
TTL: 2 * time.Minute,
|
||||||
|
})
|
||||||
|
if err := first.Close(); err != nil {
|
||||||
|
t.Fatalf("close: %v", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
// fresh handle, fresh in-memory map — as after a daemon restart
|
||||||
|
after := openStore(t, path)
|
||||||
|
reloaded := NewPersistentSessionStore(2*time.Minute, after)
|
||||||
|
if err := reloaded.Load(ctx, now.Add(10*time.Second)); err != nil {
|
||||||
|
t.Fatalf("Load: %v", err)
|
||||||
|
}
|
||||||
|
sess := reloaded.Get("voice", now.Add(10*time.Second))
|
||||||
|
if sess == nil {
|
||||||
|
t.Fatal("session did not survive the restart")
|
||||||
|
}
|
||||||
|
if sess.Intent != IntentReminder {
|
||||||
|
t.Fatalf("intent = %q, want reminder", sess.Intent)
|
||||||
|
}
|
||||||
|
// the follow-up carries no text of its own; it must inherit the old one
|
||||||
|
merged := InheritSlots(sess.Slots, Slots{Time: now.Add(2 * time.Hour), HasTime: true})
|
||||||
|
if merged.Text != "полить цветы" {
|
||||||
|
t.Fatalf("merged text = %q, want the earlier turn's text", merged.Text)
|
||||||
|
}
|
||||||
|
if !merged.Time.Equal(now.Add(2 * time.Hour)) {
|
||||||
|
t.Fatalf("merged time = %v, want the follow-up's time", merged.Time)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// A session past its TTL is dead: a restart must not bring it back.
|
||||||
|
func TestExpiredSessionNotResurrected(t *testing.T) {
|
||||||
|
ctx := context.Background()
|
||||||
|
path := filepath.Join(t.TempDir(), "maven_test.db")
|
||||||
|
now := time.Now().UTC().Truncate(time.Millisecond)
|
||||||
|
|
||||||
|
first := openStore(t, path)
|
||||||
|
before := NewPersistentSessionStore(2*time.Minute, first)
|
||||||
|
before.Put("voice", &Session{
|
||||||
|
Intent: IntentReminder,
|
||||||
|
Slots: Slots{Text: "полить цветы"},
|
||||||
|
Timestamp: now,
|
||||||
|
TTL: time.Minute,
|
||||||
|
})
|
||||||
|
if err := first.Close(); err != nil {
|
||||||
|
t.Fatalf("close: %v", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
after := openStore(t, path)
|
||||||
|
reloaded := NewPersistentSessionStore(2*time.Minute, after)
|
||||||
|
later := now.Add(5 * time.Minute) // well past the 1-min TTL
|
||||||
|
if err := reloaded.Load(ctx, later); err != nil {
|
||||||
|
t.Fatalf("Load: %v", err)
|
||||||
|
}
|
||||||
|
if sess := reloaded.Get("voice", later); sess != nil {
|
||||||
|
t.Fatalf("expired session came back: %+v", sess)
|
||||||
|
}
|
||||||
|
// and it is gone from the DB too, not just from memory
|
||||||
|
rows, err := after.LoadDialogueSessions(ctx, later)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("LoadDialogueSessions: %v", err)
|
||||||
|
}
|
||||||
|
if len(rows) != 0 {
|
||||||
|
t.Fatalf("expired row still in the DB: %+v", rows)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Delete removes the row as well, so an ended conversation stays ended.
|
||||||
|
func TestDeleteRemovesPersistedSession(t *testing.T) {
|
||||||
|
ctx := context.Background()
|
||||||
|
path := filepath.Join(t.TempDir(), "maven_test.db")
|
||||||
|
now := time.Now().UTC().Truncate(time.Millisecond)
|
||||||
|
|
||||||
|
s := openStore(t, path)
|
||||||
|
ss := NewPersistentSessionStore(2*time.Minute, s)
|
||||||
|
ss.Put("voice", &Session{Intent: IntentChat, Timestamp: now, TTL: time.Minute})
|
||||||
|
ss.Delete("voice")
|
||||||
|
rows, err := s.LoadDialogueSessions(ctx, now)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("LoadDialogueSessions: %v", err)
|
||||||
|
}
|
||||||
|
if len(rows) != 0 {
|
||||||
|
t.Fatalf("row survived Delete: %+v", rows)
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,116 @@
|
|||||||
|
// Package kiwix reads a local Kiwix server (offline Wikipedia and friends).
|
||||||
|
//
|
||||||
|
// Why: the resident model is a 0.8B and invents facts. Letting her read a local
|
||||||
|
// article snippet beats letting her recall. Nothing here talks to the internet;
|
||||||
|
// the Kiwix server is on the same box.
|
||||||
|
//
|
||||||
|
// This is search only. Full articles are ~100KB of HTML, far too big for a 4096
|
||||||
|
// token context, so the unit of context is the search snippet (~500 chars).
|
||||||
|
package kiwix
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"encoding/xml"
|
||||||
|
"fmt"
|
||||||
|
"html"
|
||||||
|
"io"
|
||||||
|
"net/http"
|
||||||
|
"net/url"
|
||||||
|
"regexp"
|
||||||
|
"strconv"
|
||||||
|
"strings"
|
||||||
|
"time"
|
||||||
|
)
|
||||||
|
|
||||||
|
// Result is one search hit.
|
||||||
|
type Result struct {
|
||||||
|
Title string // article title, e.g. "Rayleigh scattering"
|
||||||
|
Path string // e.g. /content/wikipedia_en_all_maxi_2026-02/Rayleigh_scattering
|
||||||
|
Snippet string // plain text, tags stripped, entities decoded
|
||||||
|
WordCount int // 0 if the server did not say
|
||||||
|
}
|
||||||
|
|
||||||
|
// Client is a Kiwix HTTP client. Boring on purpose: no retries, no cache.
|
||||||
|
type Client struct {
|
||||||
|
base string
|
||||||
|
http *http.Client
|
||||||
|
}
|
||||||
|
|
||||||
|
// New makes a client for a Kiwix base URL like http://127.0.0.1:8034.
|
||||||
|
func New(baseURL string) *Client {
|
||||||
|
return &Client{
|
||||||
|
base: strings.TrimRight(baseURL, "/"),
|
||||||
|
http: &http.Client{Timeout: 10 * time.Second},
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Search runs a keyword search in one ZIM (book) and returns up to limit hits.
|
||||||
|
//
|
||||||
|
// Ranking is keyword based, not semantic: "Rayleigh scattering" finds the right
|
||||||
|
// article, "why is the sky blue" finds a TV episode. Pass keywords, not questions.
|
||||||
|
func (c *Client) Search(ctx context.Context, pattern, book string, limit int) ([]Result, error) {
|
||||||
|
if limit <= 0 {
|
||||||
|
limit = 5
|
||||||
|
}
|
||||||
|
q := url.Values{}
|
||||||
|
q.Set("pattern", pattern)
|
||||||
|
q.Set("books.name", book)
|
||||||
|
q.Set("format", "xml")
|
||||||
|
q.Set("pageLength", strconv.Itoa(limit))
|
||||||
|
|
||||||
|
req, err := http.NewRequestWithContext(ctx, http.MethodGet, c.base+"/search?"+q.Encode(), nil)
|
||||||
|
if err != nil {
|
||||||
|
return nil, err
|
||||||
|
}
|
||||||
|
resp, err := c.http.Do(req)
|
||||||
|
if err != nil {
|
||||||
|
return nil, err
|
||||||
|
}
|
||||||
|
defer resp.Body.Close()
|
||||||
|
if resp.StatusCode != http.StatusOK {
|
||||||
|
return nil, fmt.Errorf("kiwix search: http %d", resp.StatusCode)
|
||||||
|
}
|
||||||
|
return ParseSearchRSS(resp.Body)
|
||||||
|
}
|
||||||
|
|
||||||
|
// rss mirrors just the bits of the RSS 2.0 reply we use.
|
||||||
|
type rss struct {
|
||||||
|
Items []struct {
|
||||||
|
Title string `xml:"title"`
|
||||||
|
Link string `xml:"link"`
|
||||||
|
// innerxml keeps the <b> match markers so we can strip them ourselves.
|
||||||
|
Description struct {
|
||||||
|
Inner string `xml:",innerxml"`
|
||||||
|
} `xml:"description"`
|
||||||
|
WordCount string `xml:"wordCount"`
|
||||||
|
} `xml:"channel>item"`
|
||||||
|
}
|
||||||
|
|
||||||
|
var tagRE = regexp.MustCompile(`<[^>]*>`)
|
||||||
|
|
||||||
|
// ParseSearchRSS turns a Kiwix search reply into results. Exported so the parser
|
||||||
|
// is testable from a captured response, with no server running.
|
||||||
|
func ParseSearchRSS(r io.Reader) ([]Result, error) {
|
||||||
|
var doc rss
|
||||||
|
if err := xml.NewDecoder(r).Decode(&doc); err != nil {
|
||||||
|
return nil, fmt.Errorf("kiwix search: bad xml: %w", err)
|
||||||
|
}
|
||||||
|
out := make([]Result, 0, len(doc.Items))
|
||||||
|
for _, it := range doc.Items {
|
||||||
|
n, _ := strconv.Atoi(strings.ReplaceAll(it.WordCount, ",", ""))
|
||||||
|
out = append(out, Result{
|
||||||
|
Title: strings.TrimSpace(it.Title),
|
||||||
|
Path: strings.TrimSpace(it.Link),
|
||||||
|
Snippet: plainText(it.Description.Inner),
|
||||||
|
WordCount: n,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
return out, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// plainText drops markup and decodes entities, leaving text a model can read.
|
||||||
|
func plainText(s string) string {
|
||||||
|
s = tagRE.ReplaceAllString(s, "")
|
||||||
|
s = html.UnescapeString(s)
|
||||||
|
return strings.TrimSpace(strings.Join(strings.Fields(s), " "))
|
||||||
|
}
|
||||||
@@ -0,0 +1,90 @@
|
|||||||
|
package kiwix
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"os"
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
)
|
||||||
|
|
||||||
|
// A real reply from the live server, trimmed to two items.
|
||||||
|
const sampleRSS = `<?xml version="1.0" encoding="UTF-8"?>
|
||||||
|
<rss version="2.0" xmlns:opensearch="http://a9.com/-/spec/opensearch/1.1/">
|
||||||
|
<channel>
|
||||||
|
<title>Search: Rayleigh scattering</title>
|
||||||
|
<opensearch:totalResults>800</opensearch:totalResults>
|
||||||
|
<item>
|
||||||
|
<title>Rayleigh scattering</title>
|
||||||
|
<link>/content/wikipedia_en_all_maxi_2026-02/Rayleigh_scattering</link>
|
||||||
|
<description><b>Rayleigh</b> scattering causes the blue color of the sky & yellow colors near the Sun.[1]</description>
|
||||||
|
<book><title>Wikipedia</title></book>
|
||||||
|
<wordCount>2,818</wordCount>
|
||||||
|
</item>
|
||||||
|
<item>
|
||||||
|
<title>Hyper–Rayleigh scattering</title>
|
||||||
|
<link>/content/wikipedia_en_all_maxi_2026-02/Hyper%E2%80%93Rayleigh_scattering</link>
|
||||||
|
<description>...<b>Rayleigh</b> scattering" is a nonlinear optical counterpart.</description>
|
||||||
|
<book><title>Wikipedia</title></book>
|
||||||
|
<wordCount>914</wordCount>
|
||||||
|
</item>
|
||||||
|
</channel>
|
||||||
|
</rss>`
|
||||||
|
|
||||||
|
func TestParseSearchRSS(t *testing.T) {
|
||||||
|
got, err := ParseSearchRSS(strings.NewReader(sampleRSS))
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("parse: %v", err)
|
||||||
|
}
|
||||||
|
if len(got) != 2 {
|
||||||
|
t.Fatalf("want 2 results, got %d", len(got))
|
||||||
|
}
|
||||||
|
if got[0].Title != "Rayleigh scattering" {
|
||||||
|
t.Errorf("title = %q", got[0].Title)
|
||||||
|
}
|
||||||
|
if got[0].Path != "/content/wikipedia_en_all_maxi_2026-02/Rayleigh_scattering" {
|
||||||
|
t.Errorf("path = %q", got[0].Path)
|
||||||
|
}
|
||||||
|
if got[0].WordCount != 2818 {
|
||||||
|
t.Errorf("wordCount = %d", got[0].WordCount)
|
||||||
|
}
|
||||||
|
want := "Rayleigh scattering causes the blue color of the sky & yellow colors near the Sun.[1]"
|
||||||
|
if got[0].Snippet != want {
|
||||||
|
t.Errorf("snippet = %q, want %q", got[0].Snippet, want)
|
||||||
|
}
|
||||||
|
if strings.Contains(got[1].Snippet, "<b>") {
|
||||||
|
t.Errorf("second snippet still has tags: %q", got[1].Snippet)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestParseSearchRSSBadXML(t *testing.T) {
|
||||||
|
if _, err := ParseSearchRSS(strings.NewReader("not xml at all")); err == nil {
|
||||||
|
t.Fatal("want an error on junk input")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Opt-in: needs a live Kiwix server. CI has none.
|
||||||
|
// MAVEN_KIWIX_URL=http://127.0.0.1:8034 no_proxy=127.0.0.1,localhost go test -run Retrieval -v ./internal/kiwix/
|
||||||
|
func TestRetrievalEval(t *testing.T) {
|
||||||
|
base := os.Getenv("MAVEN_KIWIX_URL")
|
||||||
|
if base == "" {
|
||||||
|
t.Skip("set MAVEN_KIWIX_URL to run the retrieval eval")
|
||||||
|
}
|
||||||
|
noProxyLoopback(t)
|
||||||
|
|
||||||
|
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute)
|
||||||
|
defer cancel()
|
||||||
|
|
||||||
|
rep, err := RunRetrievalEval(ctx, New(base), 5)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("eval: %v", err)
|
||||||
|
}
|
||||||
|
// No pass bar on purpose: the number is the finding.
|
||||||
|
t.Log("\n" + rep.String() + rep.Detail())
|
||||||
|
}
|
||||||
|
|
||||||
|
// noProxyLoopback stops the box's SOCKS bridge from eating loopback requests.
|
||||||
|
func noProxyLoopback(t *testing.T) {
|
||||||
|
t.Setenv("no_proxy", "127.0.0.1,localhost")
|
||||||
|
t.Setenv("NO_PROXY", "127.0.0.1,localhost")
|
||||||
|
}
|
||||||
@@ -0,0 +1,63 @@
|
|||||||
|
{
|
||||||
|
"name": "kiwix-knowledge-v1",
|
||||||
|
"book": "wikipedia_en_all_maxi_2026-02",
|
||||||
|
"note": "The 9 knowledge cases from internal/phraser/eval/talk_v1.json. Queries are hand-written English keywords on purpose: Kiwix ranks by keyword, not meaning, so a natural question fails. Writing them by hand separates 'retrieval is broken' from 'the model writes bad queries'.",
|
||||||
|
"cases": [
|
||||||
|
{
|
||||||
|
"id": "know-sky-blue",
|
||||||
|
"question": "почему небо синее?",
|
||||||
|
"query": "Rayleigh scattering sky blue",
|
||||||
|
"want_titles": ["Rayleigh scattering", "Diffuse sky radiation"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "know-boil-egg",
|
||||||
|
"question": "сколько варить яйцо вкрутую?",
|
||||||
|
"query": "boiled egg cooking",
|
||||||
|
"want_titles": ["Boiled egg", "Egg as food"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "know-ssd-vs-hdd",
|
||||||
|
"question": "чем ssd отличается от hdd?",
|
||||||
|
"query": "solid-state drive",
|
||||||
|
"want_titles": ["Solid-state drive", "Hard disk drive"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "know-cat-purr",
|
||||||
|
"question": "почему кошки мурчат?",
|
||||||
|
"query": "cat purr",
|
||||||
|
"want_titles": ["Purr", "Cat communication"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "know-hiccups",
|
||||||
|
"question": "как быстро избавиться от икоты?",
|
||||||
|
"query": "hiccup",
|
||||||
|
"want_titles": ["Hiccup"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "know-polite-form",
|
||||||
|
"question": "не могли бы вы объяснить, что такое vpn?",
|
||||||
|
"query": "virtual private network",
|
||||||
|
"want_titles": ["Virtual private network"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "know-dont-know",
|
||||||
|
"question": "как зовут моего соседа снизу?",
|
||||||
|
"query": "name of my downstairs neighbour",
|
||||||
|
"want_titles": [],
|
||||||
|
"expect_miss": true,
|
||||||
|
"note": "Unanswerable by design. Retrieval SHOULD find nothing useful. Counted as a hit only when nothing relevant comes back."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "know-water-per-day",
|
||||||
|
"question": "сколько воды в день надо пить?",
|
||||||
|
"query": "human daily water requirement drinking",
|
||||||
|
"want_titles": ["Drinking water", "Water", "Dehydration", "Hydration"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "know-thunder-delay",
|
||||||
|
"question": "почему гром слышно позже молнии?",
|
||||||
|
"query": "thunder speed of sound lightning",
|
||||||
|
"want_titles": ["Thunder", "Lightning"]
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
@@ -0,0 +1,135 @@
|
|||||||
|
package kiwix
|
||||||
|
|
||||||
|
// This scores retrieval alone: no LLM. For each general-knowledge question we
|
||||||
|
// hand-write English keywords and ask whether the article that would answer it
|
||||||
|
// comes back in the top N hits. If this score is low, reading Wikipedia cannot
|
||||||
|
// help the model no matter how good the prompt is.
|
||||||
|
//
|
||||||
|
// The unanswerable case (know-dont-know) is not scored. Whether the junk it
|
||||||
|
// returns is "nothing useful" is a human judgement, so the report just prints
|
||||||
|
// the titles and leaves the score to the 8 answerable cases.
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
_ "embed"
|
||||||
|
"encoding/json"
|
||||||
|
"fmt"
|
||||||
|
"strings"
|
||||||
|
)
|
||||||
|
|
||||||
|
//go:embed knowledge_v1.json
|
||||||
|
var knowledgeFixtureJSON []byte
|
||||||
|
|
||||||
|
// EvalCase — one question with hand-written keywords.
|
||||||
|
type EvalCase struct {
|
||||||
|
ID string `json:"id"`
|
||||||
|
Question string `json:"question"`
|
||||||
|
Query string `json:"query"`
|
||||||
|
WantTitles []string `json:"want_titles"`
|
||||||
|
ExpectMiss bool `json:"expect_miss"`
|
||||||
|
}
|
||||||
|
|
||||||
|
type fixture struct {
|
||||||
|
Name string `json:"name"`
|
||||||
|
Book string `json:"book"`
|
||||||
|
Cases []EvalCase `json:"cases"`
|
||||||
|
}
|
||||||
|
|
||||||
|
// Outcome — what one case retrieved.
|
||||||
|
type Outcome struct {
|
||||||
|
Case EvalCase
|
||||||
|
Titles []string // titles of the top N hits, in rank order
|
||||||
|
Rank int // 1-based rank of the first wanted title, 0 if none
|
||||||
|
Err error
|
||||||
|
}
|
||||||
|
|
||||||
|
// Hit is true when a wanted title came back.
|
||||||
|
func (o Outcome) Hit() bool { return o.Rank > 0 }
|
||||||
|
|
||||||
|
// Report — the score plus per-case detail.
|
||||||
|
type Report struct {
|
||||||
|
Name string
|
||||||
|
Book string
|
||||||
|
TopN int
|
||||||
|
Scored int // answerable cases
|
||||||
|
Hits int
|
||||||
|
Errors int
|
||||||
|
Outcomes []Outcome
|
||||||
|
}
|
||||||
|
|
||||||
|
// Accuracy over the answerable cases.
|
||||||
|
func (r Report) Accuracy() float64 {
|
||||||
|
if r.Scored == 0 {
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
return float64(r.Hits) / float64(r.Scored)
|
||||||
|
}
|
||||||
|
|
||||||
|
// RunRetrievalEval searches for every fixture case.
|
||||||
|
func RunRetrievalEval(ctx context.Context, c *Client, topN int) (Report, error) {
|
||||||
|
var f fixture
|
||||||
|
if err := json.Unmarshal(knowledgeFixtureJSON, &f); err != nil {
|
||||||
|
return Report{}, err
|
||||||
|
}
|
||||||
|
rep := Report{Name: f.Name, Book: f.Book, TopN: topN}
|
||||||
|
for _, cs := range f.Cases {
|
||||||
|
res, err := c.Search(ctx, cs.Query, f.Book, topN)
|
||||||
|
o := Outcome{Case: cs, Err: err}
|
||||||
|
if err != nil {
|
||||||
|
rep.Errors++
|
||||||
|
}
|
||||||
|
for i, hit := range res {
|
||||||
|
o.Titles = append(o.Titles, hit.Title)
|
||||||
|
if o.Rank == 0 && matches(cs.WantTitles, hit.Title) {
|
||||||
|
o.Rank = i + 1
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if !cs.ExpectMiss {
|
||||||
|
rep.Scored++
|
||||||
|
if o.Hit() {
|
||||||
|
rep.Hits++
|
||||||
|
}
|
||||||
|
}
|
||||||
|
rep.Outcomes = append(rep.Outcomes, o)
|
||||||
|
}
|
||||||
|
return rep, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
func matches(want []string, title string) bool {
|
||||||
|
for _, w := range want {
|
||||||
|
if strings.EqualFold(strings.TrimSpace(title), w) {
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
|
||||||
|
// String — the headline number.
|
||||||
|
func (r Report) String() string {
|
||||||
|
var b strings.Builder
|
||||||
|
fmt.Fprintf(&b, "%s: %d/%d answerable questions retrieve a wanted article in top %d (%.1f%%), %d errors\n",
|
||||||
|
r.Name, r.Hits, r.Scored, r.TopN, 100*r.Accuracy(), r.Errors)
|
||||||
|
fmt.Fprintf(&b, " book: %s\n", r.Book)
|
||||||
|
return b.String()
|
||||||
|
}
|
||||||
|
|
||||||
|
// Detail — per case: what was asked, what was searched, what came back.
|
||||||
|
func (r Report) Detail() string {
|
||||||
|
var b strings.Builder
|
||||||
|
for _, o := range r.Outcomes {
|
||||||
|
mark := "MISS"
|
||||||
|
switch {
|
||||||
|
case o.Case.ExpectMiss:
|
||||||
|
mark = "n/a "
|
||||||
|
case o.Hit():
|
||||||
|
mark = fmt.Sprintf("hit@%d", o.Rank)
|
||||||
|
}
|
||||||
|
fmt.Fprintf(&b, " %-6s %-20s q=%q\n", mark, o.Case.ID, o.Case.Query)
|
||||||
|
if o.Err != nil {
|
||||||
|
fmt.Fprintf(&b, " error: %v\n", o.Err)
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
fmt.Fprintf(&b, " got: %s\n", strings.Join(o.Titles, " | "))
|
||||||
|
}
|
||||||
|
return b.String()
|
||||||
|
}
|
||||||
@@ -0,0 +1,183 @@
|
|||||||
|
package kiwix
|
||||||
|
|
||||||
|
// Turning a Russian question into an English Kiwix search.
|
||||||
|
//
|
||||||
|
// Kiwix ranks by keyword, not by meaning. "why is the sky blue" returns a TV
|
||||||
|
// episode; "Rayleigh scattering sky blue" returns the right article. So the
|
||||||
|
// model's job here is NOT translation — it is naming the English article the
|
||||||
|
// answer lives in.
|
||||||
|
//
|
||||||
|
// The output space is a handful of words, so it is worth locking down hard: a
|
||||||
|
// GBNF grammar for the shape, a tiny token cap, and a cleanup pass that throws
|
||||||
|
// away anything odd rather than handing junk to Kiwix.
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"encoding/json"
|
||||||
|
"fmt"
|
||||||
|
"strings"
|
||||||
|
"unicode"
|
||||||
|
|
||||||
|
"github.com/kami/maven/internal/llm"
|
||||||
|
)
|
||||||
|
|
||||||
|
// Completer — the LLM seam, so tests can fake it. *llm.Client satisfies it.
|
||||||
|
type Completer interface {
|
||||||
|
Complete(ctx context.Context, r llm.Req) (string, error)
|
||||||
|
}
|
||||||
|
|
||||||
|
// queryGrammar — one JSON object holding 1..6 keyword words. Latin letters,
|
||||||
|
// digits and hyphens only, so the model physically cannot answer the question
|
||||||
|
// or reply in Russian.
|
||||||
|
//
|
||||||
|
// Why the JSON wrapper: this model always thinks out loud and this llama-server
|
||||||
|
// build ignores the thinking switch (see ROUTING-EVAL-31-07-2026.md). A bare
|
||||||
|
// word-list grammar just captured the reasoning — every case came back as
|
||||||
|
// "Let me analyze this request carefully". Demanding JSON, like routeGrammar and
|
||||||
|
// responseGrammar already do, gives the reasoning nowhere to go.
|
||||||
|
const queryGrammar = `
|
||||||
|
root ::= "{" ws "\"query\"" ws ":" ws "\"" word (" " word){0,5} "\"" ws "}"
|
||||||
|
word ::= [A-Za-z0-9] [A-Za-z0-9-]{0,23}
|
||||||
|
ws ::= [ \t\n]*
|
||||||
|
`
|
||||||
|
|
||||||
|
// rewriteSystem — asks for search keywords, not an answer and not a translation.
|
||||||
|
const rewriteSystem = `You turn a question into a search query for English Wikipedia.
|
||||||
|
|
||||||
|
Rules:
|
||||||
|
- Output ONLY English search keywords. Never an answer, never an explanation.
|
||||||
|
- Do NOT translate the sentence. Name the thing the answer is about.
|
||||||
|
- The output must be a noun phrase, like a Wikipedia article title.
|
||||||
|
- Never use question words: no why, how, what, when, which, "how much",
|
||||||
|
"how long", "how to", "vs", "reason", "difference".
|
||||||
|
- 2 to 4 words.
|
||||||
|
|
||||||
|
Reply with JSON: {"query":"<keywords>"}
|
||||||
|
|
||||||
|
Good:
|
||||||
|
"почему листья желтеют осенью?" -> {"query":"leaf senescence autumn"}
|
||||||
|
"как работает микроволновка?" -> {"query":"microwave oven"}
|
||||||
|
"не могли бы вы объяснить, что такое блокчейн?" -> {"query":"blockchain"}
|
||||||
|
"сколько живут собаки?" -> {"query":"dog lifespan"}
|
||||||
|
"как избавиться от комаров в квартире?" -> {"query":"mosquito control"}
|
||||||
|
"чем чай отличается от кофе?" -> {"query":"tea"}
|
||||||
|
|
||||||
|
Only JSON, no explanation.`
|
||||||
|
|
||||||
|
// maxQueryTokens — the output is a few words plus the JSON wrapper. A tight cap
|
||||||
|
// is the cheapest guard against the model rambling into an answer.
|
||||||
|
const maxQueryTokens = 32
|
||||||
|
|
||||||
|
// Rewriter asks the resident model for English search keywords.
|
||||||
|
type Rewriter struct{ c Completer }
|
||||||
|
|
||||||
|
func NewRewriter(c Completer) *Rewriter { return &Rewriter{c: c} }
|
||||||
|
|
||||||
|
// Rewrite returns English keywords for a question in any language.
|
||||||
|
// It errors rather than returning something Kiwix should not see.
|
||||||
|
func (r *Rewriter) Rewrite(ctx context.Context, question string) (string, error) {
|
||||||
|
raw, err := r.c.Complete(ctx, llm.Req{
|
||||||
|
System: rewriteSystem,
|
||||||
|
User: strings.TrimSpace(question),
|
||||||
|
Grammar: queryGrammar,
|
||||||
|
MaxTokens: maxQueryTokens,
|
||||||
|
RepeatPenalty: 1.15,
|
||||||
|
})
|
||||||
|
if err != nil {
|
||||||
|
return "", err
|
||||||
|
}
|
||||||
|
return CleanQuery(unwrapJSON(raw))
|
||||||
|
}
|
||||||
|
|
||||||
|
// unwrapJSON pulls the query out of {"query":"..."}. If the reply is not that
|
||||||
|
// shape it is returned as-is, and CleanQuery decides whether it is usable.
|
||||||
|
func unwrapJSON(raw string) string {
|
||||||
|
s := strings.TrimSpace(raw)
|
||||||
|
if !strings.HasPrefix(s, "{") {
|
||||||
|
return s
|
||||||
|
}
|
||||||
|
var got struct{ Query string }
|
||||||
|
if err := json.Unmarshal([]byte(s), &got); err != nil {
|
||||||
|
return s
|
||||||
|
}
|
||||||
|
return got.Query
|
||||||
|
}
|
||||||
|
|
||||||
|
// maxQueryWords matches the grammar's bound. Anything longer is prose.
|
||||||
|
const maxQueryWords = 6
|
||||||
|
|
||||||
|
// CleanQuery checks and tidies whatever the model produced. The grammar makes
|
||||||
|
// bad output unlikely, not impossible (a server without grammar support, a
|
||||||
|
// different model), so this is the real gate in front of Kiwix.
|
||||||
|
//
|
||||||
|
// Exported so it can be tested without a model.
|
||||||
|
func CleanQuery(raw string) (string, error) {
|
||||||
|
s := strings.TrimSpace(raw)
|
||||||
|
// Models like to wrap answers in quotes. Drop surrounding ones.
|
||||||
|
s = strings.Trim(s, "\"'`")
|
||||||
|
// Keep the first line only: everything after it is prose.
|
||||||
|
if i := strings.IndexAny(s, "\r\n"); i >= 0 {
|
||||||
|
s = s[:i]
|
||||||
|
}
|
||||||
|
// Keep letters, digits, spaces and hyphens; anything else becomes a space.
|
||||||
|
var b strings.Builder
|
||||||
|
for _, ru := range s {
|
||||||
|
switch {
|
||||||
|
case unicode.IsLetter(ru) || unicode.IsDigit(ru) || ru == '-':
|
||||||
|
b.WriteRune(ru)
|
||||||
|
default:
|
||||||
|
b.WriteRune(' ')
|
||||||
|
}
|
||||||
|
}
|
||||||
|
words := strings.Fields(b.String())
|
||||||
|
if len(words) == 0 {
|
||||||
|
return "", fmt.Errorf("kiwix rewrite: empty query")
|
||||||
|
}
|
||||||
|
if len(words) > maxQueryWords {
|
||||||
|
return "", fmt.Errorf("kiwix rewrite: %d words, want at most %d (looks like prose)", len(words), maxQueryWords)
|
||||||
|
}
|
||||||
|
words = dropStopWords(words)
|
||||||
|
out := strings.Join(words, " ")
|
||||||
|
// The ZIMs are English. Non-Latin letters mean the model ignored the ask.
|
||||||
|
for _, ru := range out {
|
||||||
|
if unicode.IsLetter(ru) && !isLatin(ru) {
|
||||||
|
return "", fmt.Errorf("kiwix rewrite: query is not English: %q", out)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return out, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// stopWords — question words and filler. The model keeps writing question-shaped
|
||||||
|
// queries ("why is the sky blue", "how much water to drink daily") no matter how
|
||||||
|
// the prompt is worded, and Kiwix ranks on every word, so those words drag in
|
||||||
|
// song and episode titles. Dropping them in code is not a style preference: a
|
||||||
|
// keyword ranker gets nothing from them.
|
||||||
|
var stopWords = map[string]bool{
|
||||||
|
"a": true, "an": true, "the": true, "is": true, "are": true, "was": true,
|
||||||
|
"do": true, "does": true, "did": true, "to": true, "of": true, "in": true,
|
||||||
|
"on": true, "for": true, "and": true, "or": true, "my": true, "me": true,
|
||||||
|
"i": true, "it": true, "its": true, "be": true, "been": true, "get": true,
|
||||||
|
"how": true, "why": true, "what": true, "when": true, "which": true,
|
||||||
|
"who": true, "where": true, "much": true, "many": true, "long": true,
|
||||||
|
"vs": true, "than": true, "rid": true, "from": true, "about": true,
|
||||||
|
}
|
||||||
|
|
||||||
|
// dropStopWords removes filler, but never everything: if the query was nothing
|
||||||
|
// but stop words there is nothing better to search, so the original is kept and
|
||||||
|
// the caller sees whatever Kiwix makes of it.
|
||||||
|
func dropStopWords(words []string) []string {
|
||||||
|
kept := make([]string, 0, len(words))
|
||||||
|
for _, w := range words {
|
||||||
|
if !stopWords[strings.ToLower(w)] {
|
||||||
|
kept = append(kept, w)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if len(kept) == 0 {
|
||||||
|
return words
|
||||||
|
}
|
||||||
|
return kept
|
||||||
|
}
|
||||||
|
|
||||||
|
func isLatin(ru rune) bool {
|
||||||
|
return (ru >= 'a' && ru <= 'z') || (ru >= 'A' && ru <= 'Z')
|
||||||
|
}
|
||||||
@@ -0,0 +1,96 @@
|
|||||||
|
package kiwix
|
||||||
|
|
||||||
|
// End-to-end score: Russian question -> model rewrite -> Kiwix search -> did a
|
||||||
|
// wanted article come back. Same 9 cases as the retrieval eval, so the two
|
||||||
|
// numbers are directly comparable: retrieval with hand-written keywords is the
|
||||||
|
// ceiling, this is what the model actually reaches.
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"encoding/json"
|
||||||
|
"fmt"
|
||||||
|
"strings"
|
||||||
|
)
|
||||||
|
|
||||||
|
// RewriteOutcome — one case, end to end.
|
||||||
|
type RewriteOutcome struct {
|
||||||
|
Outcome
|
||||||
|
ModelQuery string // what the model asked for ("" if it failed)
|
||||||
|
RewriteErr error
|
||||||
|
}
|
||||||
|
|
||||||
|
// RunRewriteEval rewrites every question with the model, then searches.
|
||||||
|
func RunRewriteEval(ctx context.Context, c *Client, rw *Rewriter, topN int) (RewriteReport, error) {
|
||||||
|
var f fixture
|
||||||
|
if err := json.Unmarshal(knowledgeFixtureJSON, &f); err != nil {
|
||||||
|
return RewriteReport{}, err
|
||||||
|
}
|
||||||
|
rep := RewriteReport{Report: Report{Name: f.Name + "-rewrite", Book: f.Book, TopN: topN}}
|
||||||
|
for _, cs := range f.Cases {
|
||||||
|
out := RewriteOutcome{Outcome: Outcome{Case: cs}}
|
||||||
|
q, err := rw.Rewrite(ctx, cs.Question)
|
||||||
|
out.ModelQuery, out.RewriteErr = q, err
|
||||||
|
if err == nil {
|
||||||
|
res, serr := c.Search(ctx, q, f.Book, topN)
|
||||||
|
out.Err = serr
|
||||||
|
for i, hit := range res {
|
||||||
|
out.Titles = append(out.Titles, hit.Title)
|
||||||
|
if out.Rank == 0 && matches(cs.WantTitles, hit.Title) {
|
||||||
|
out.Rank = i + 1
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if out.RewriteErr != nil || out.Err != nil {
|
||||||
|
rep.Errors++
|
||||||
|
}
|
||||||
|
if !cs.ExpectMiss {
|
||||||
|
rep.Scored++
|
||||||
|
if out.Hit() {
|
||||||
|
rep.Hits++
|
||||||
|
}
|
||||||
|
}
|
||||||
|
rep.Cases = append(rep.Cases, out)
|
||||||
|
}
|
||||||
|
return rep, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// RewriteReport — the score plus per-case detail.
|
||||||
|
type RewriteReport struct {
|
||||||
|
Report
|
||||||
|
Cases []RewriteOutcome
|
||||||
|
}
|
||||||
|
|
||||||
|
// String — the headline number.
|
||||||
|
func (r RewriteReport) String() string {
|
||||||
|
return fmt.Sprintf("%s: %d/%d answerable questions retrieve a wanted article in top %d (%.1f%%), %d errors\n book: %s\n",
|
||||||
|
r.Name, r.Hits, r.Scored, r.TopN, 100*r.Accuracy(), r.Errors, r.Book)
|
||||||
|
}
|
||||||
|
|
||||||
|
// Detail — per case: hand-written query next to the model's, and what came back.
|
||||||
|
// The point is seeing WHERE the model's phrasing differs, not just the score.
|
||||||
|
func (r RewriteReport) Detail() string {
|
||||||
|
var b strings.Builder
|
||||||
|
for _, o := range r.Cases {
|
||||||
|
mark := "MISS"
|
||||||
|
switch {
|
||||||
|
case o.Case.ExpectMiss:
|
||||||
|
mark = "n/a "
|
||||||
|
case o.Hit():
|
||||||
|
mark = fmt.Sprintf("hit@%d", o.Rank)
|
||||||
|
}
|
||||||
|
fmt.Fprintf(&b, " %-6s %-20s\n", mark, o.Case.ID)
|
||||||
|
fmt.Fprintf(&b, " asked: %s\n", o.Case.Question)
|
||||||
|
fmt.Fprintf(&b, " hand: %q\n", o.Case.Query)
|
||||||
|
fmt.Fprintf(&b, " model: %q\n", o.ModelQuery)
|
||||||
|
if o.RewriteErr != nil {
|
||||||
|
fmt.Fprintf(&b, " rewrite rejected: %v\n", o.RewriteErr)
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if o.Err != nil {
|
||||||
|
fmt.Fprintf(&b, " search error: %v\n", o.Err)
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
fmt.Fprintf(&b, " got: %s\n", strings.Join(o.Titles, " | "))
|
||||||
|
}
|
||||||
|
return b.String()
|
||||||
|
}
|
||||||
@@ -0,0 +1,33 @@
|
|||||||
|
package kiwix
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"os"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"github.com/kami/maven/internal/llm"
|
||||||
|
)
|
||||||
|
|
||||||
|
// Opt-in: needs a live Kiwix server AND a live llama-server.
|
||||||
|
// MAVEN_KIWIX_URL=http://127.0.0.1:8034 MAVEN_LLM_URL=http://127.0.0.1:18099 \
|
||||||
|
//
|
||||||
|
// no_proxy=127.0.0.1,localhost go test -run RewriteEval -v ./internal/kiwix/
|
||||||
|
func TestRewriteEval(t *testing.T) {
|
||||||
|
kbase, lbase := os.Getenv("MAVEN_KIWIX_URL"), os.Getenv("MAVEN_LLM_URL")
|
||||||
|
if kbase == "" || lbase == "" {
|
||||||
|
t.Skip("set MAVEN_KIWIX_URL and MAVEN_LLM_URL to run the rewrite eval")
|
||||||
|
}
|
||||||
|
noProxyLoopback(t)
|
||||||
|
|
||||||
|
ctx, cancel := context.WithTimeout(context.Background(), 15*time.Minute)
|
||||||
|
defer cancel()
|
||||||
|
|
||||||
|
rw := NewRewriter(llm.New(lbase, 3*time.Minute))
|
||||||
|
rep, err := RunRewriteEval(ctx, New(kbase), rw, 5)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("eval: %v", err)
|
||||||
|
}
|
||||||
|
// No pass bar on purpose: the number is the finding.
|
||||||
|
t.Log("\n" + rep.String() + rep.Detail())
|
||||||
|
}
|
||||||
@@ -0,0 +1,105 @@
|
|||||||
|
package kiwix
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"testing"
|
||||||
|
|
||||||
|
"github.com/kami/maven/internal/llm"
|
||||||
|
)
|
||||||
|
|
||||||
|
// Bad model output must never reach Kiwix. No model needed for this.
|
||||||
|
func TestCleanQueryRejectsJunk(t *testing.T) {
|
||||||
|
bad := []struct{ name, raw string }{
|
||||||
|
{"empty", ""},
|
||||||
|
{"blank", " \n "},
|
||||||
|
{"russian came back", "почему небо синее"},
|
||||||
|
{"mixed russian", "sky синее scattering"},
|
||||||
|
{"full sentence", "The sky looks blue because of the scattering of sunlight by air molecules"},
|
||||||
|
{"prose with quotes", `Sure! Here is a good search query: "Rayleigh scattering", which explains it.`},
|
||||||
|
}
|
||||||
|
for _, c := range bad {
|
||||||
|
if got, err := CleanQuery(c.raw); err == nil {
|
||||||
|
t.Errorf("%s: want rejection, got %q", c.name, got)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestCleanQueryCleans(t *testing.T) {
|
||||||
|
ok := []struct{ raw, want string }{
|
||||||
|
{"Rayleigh scattering sky", "Rayleigh scattering sky"},
|
||||||
|
{" boiled egg cooking \n", "boiled egg cooking"},
|
||||||
|
{`"virtual private network"`, "virtual private network"},
|
||||||
|
{"solid-state drive", "solid-state drive"},
|
||||||
|
{"cat purr.", "cat purr"},
|
||||||
|
{"hiccup\nAlso: hiccough", "hiccup"},
|
||||||
|
// Question words are filler to a keyword ranker, so they go.
|
||||||
|
{"why is the sky blue", "sky blue"},
|
||||||
|
{"how much water to drink daily", "water drink daily"},
|
||||||
|
{"SSD vs HDD comparison", "SSD HDD comparison"},
|
||||||
|
// Nothing but filler: keep it rather than return nothing.
|
||||||
|
{"what is it", "what is it"},
|
||||||
|
}
|
||||||
|
for _, c := range ok {
|
||||||
|
got, err := CleanQuery(c.raw)
|
||||||
|
if err != nil {
|
||||||
|
t.Errorf("%q: %v", c.raw, err)
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if got != c.want {
|
||||||
|
t.Errorf("%q -> %q, want %q", c.raw, got, c.want)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
type fakeCompleter struct {
|
||||||
|
out string
|
||||||
|
req llm.Req
|
||||||
|
}
|
||||||
|
|
||||||
|
func (f *fakeCompleter) Complete(_ context.Context, r llm.Req) (string, error) {
|
||||||
|
f.req = r
|
||||||
|
return f.out, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestRewriteConstrainsTheCall(t *testing.T) {
|
||||||
|
f := &fakeCompleter{out: `{"query":"Rayleigh scattering sky"}`}
|
||||||
|
got, err := NewRewriter(f).Rewrite(context.Background(), "почему небо синее?")
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("rewrite: %v", err)
|
||||||
|
}
|
||||||
|
if got != "Rayleigh scattering sky" {
|
||||||
|
t.Errorf("query = %q", got)
|
||||||
|
}
|
||||||
|
if f.req.Grammar == "" {
|
||||||
|
t.Error("no grammar sent")
|
||||||
|
}
|
||||||
|
if f.req.MaxTokens == 0 || f.req.MaxTokens > 32 {
|
||||||
|
t.Errorf("max_tokens = %d, want a small cap", f.req.MaxTokens)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestRewriteRejectsBadModelOutput(t *testing.T) {
|
||||||
|
bad := []string{
|
||||||
|
`{"query":"почему небо синее"}`, // never translated
|
||||||
|
`{"query":""}`, // empty
|
||||||
|
`{"query":"the sky is blue because sunlight is scattered by air"}`, // an answer
|
||||||
|
// Note: a SHORT English prose fragment ("Let me analyze this request")
|
||||||
|
// is under the word cap and cannot be caught here. The grammar is what
|
||||||
|
// stops that one.
|
||||||
|
}
|
||||||
|
for _, out := range bad {
|
||||||
|
f := &fakeCompleter{out: out}
|
||||||
|
if got, err := NewRewriter(f).Rewrite(context.Background(), "почему небо синее?"); err == nil {
|
||||||
|
t.Errorf("%s: want rejection, got %q", out, got)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// A reply that is not the JSON shape but is still usable keywords should pass.
|
||||||
|
func TestRewriteFallsBackToPlainText(t *testing.T) {
|
||||||
|
f := &fakeCompleter{out: "Rayleigh scattering sky"}
|
||||||
|
got, err := NewRewriter(f).Rewrite(context.Background(), "почему небо синее?")
|
||||||
|
if err != nil || got != "Rayleigh scattering sky" {
|
||||||
|
t.Errorf("got %q, %v", got, err)
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -1,4 +1,4 @@
|
|||||||
package eval
|
package llm
|
||||||
|
|
||||||
import (
|
import (
|
||||||
"context"
|
"context"
|
||||||
@@ -6,8 +6,20 @@ import (
|
|||||||
"fmt"
|
"fmt"
|
||||||
"net/http"
|
"net/http"
|
||||||
"strings"
|
"strings"
|
||||||
|
"time"
|
||||||
)
|
)
|
||||||
|
|
||||||
|
// UnknownModel is the label to print when the server would not say what it has
|
||||||
|
// loaded. Deliberately ugly: an honest "unknown" is fine, a plausible-looking
|
||||||
|
// but wrong model name is the bug this whole file exists to prevent.
|
||||||
|
const UnknownModel = "unknown-model"
|
||||||
|
|
||||||
|
// llama-server is local, so never send this through a proxy: this box's
|
||||||
|
// http_proxy answers 503 for loopback, which would look like "server won't say
|
||||||
|
// which model it has" when the server is right there and fine.
|
||||||
|
// A Transport with no Proxy set bypasses http_proxy entirely.
|
||||||
|
var modelHTTP = &http.Client{Timeout: 10 * time.Second, Transport: &http.Transport{}}
|
||||||
|
|
||||||
// ModelID asks llama-server which model it has loaded, so a scoring run can
|
// ModelID asks llama-server which model it has loaded, so a scoring run can
|
||||||
// label itself. Without this a bake-off between two models produces two tables
|
// label itself. Without this a bake-off between two models produces two tables
|
||||||
// that look identical, and the operator has to remember which server was up.
|
// that look identical, and the operator has to remember which server was up.
|
||||||
@@ -19,7 +31,7 @@ func ModelID(ctx context.Context, base string) (string, error) {
|
|||||||
if err != nil {
|
if err != nil {
|
||||||
return "", err
|
return "", err
|
||||||
}
|
}
|
||||||
resp, err := http.DefaultClient.Do(req)
|
resp, err := modelHTTP.Do(req)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return "", err
|
return "", err
|
||||||
}
|
}
|
||||||
@@ -38,14 +50,21 @@ func ModelID(ctx context.Context, base string) (string, error) {
|
|||||||
if len(out.Data) == 0 {
|
if len(out.Data) == 0 {
|
||||||
return "", fmt.Errorf("models: empty list")
|
return "", fmt.Errorf("models: empty list")
|
||||||
}
|
}
|
||||||
return shortModelID(out.Data[0].ID), nil
|
short := shortModelID(out.Data[0].ID)
|
||||||
|
if short == "" {
|
||||||
|
// Server answered but the id field was missing or blank. Say so
|
||||||
|
// instead of handing back an empty label that reads as a real name.
|
||||||
|
return "", fmt.Errorf("models: no id in response")
|
||||||
|
}
|
||||||
|
return short, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
// shortModelID trims the path and the .gguf suffix — llama-server reports the
|
// shortModelID trims the path and the .gguf suffix — llama-server reports the
|
||||||
// file name it was started with, which is too long for a table header.
|
// file name it was started with, which is too long for a table header.
|
||||||
func shortModelID(id string) string {
|
func shortModelID(id string) string {
|
||||||
|
id = strings.TrimSpace(id)
|
||||||
if i := strings.LastIndexAny(id, "/\\"); i >= 0 {
|
if i := strings.LastIndexAny(id, "/\\"); i >= 0 {
|
||||||
id = id[i+1:]
|
id = id[i+1:]
|
||||||
}
|
}
|
||||||
return strings.TrimSuffix(id, ".gguf")
|
return strings.TrimSpace(strings.TrimSuffix(id, ".gguf"))
|
||||||
}
|
}
|
||||||
@@ -0,0 +1,64 @@
|
|||||||
|
package llm
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"net/http"
|
||||||
|
"net/http/httptest"
|
||||||
|
"testing"
|
||||||
|
)
|
||||||
|
|
||||||
|
// The point of these tests: a wrong-but-plausible model label is the bug, so
|
||||||
|
// every path that cannot learn the real name must return an error instead of a
|
||||||
|
// guess. No llama-server needed — a stub server stands in.
|
||||||
|
func TestModelID(t *testing.T) {
|
||||||
|
cases := []struct {
|
||||||
|
name string
|
||||||
|
body string
|
||||||
|
code int
|
||||||
|
want string // "" ⇒ expect an error
|
||||||
|
}{
|
||||||
|
{"full path", `{"data":[{"id":"/mnt/hdd1/llms/qwen3.5/Qwen3.5-0.8B.Q4_K_M.gguf"}]}`, 200, "Qwen3.5-0.8B.Q4_K_M"},
|
||||||
|
{"bare name", `{"data":[{"id":"LFM2.5-1.2B"}]}`, 200, "LFM2.5-1.2B"},
|
||||||
|
{"empty list", `{"data":[]}`, 200, ""},
|
||||||
|
{"id missing", `{"data":[{}]}`, 200, ""},
|
||||||
|
{"id blank", `{"data":[{"id":" "}]}`, 200, ""},
|
||||||
|
{"server error", `nope`, 500, ""},
|
||||||
|
{"not json", `<html>`, 200, ""},
|
||||||
|
}
|
||||||
|
for _, c := range cases {
|
||||||
|
t.Run(c.name, func(t *testing.T) {
|
||||||
|
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||||
|
if r.URL.Path != "/v1/models" {
|
||||||
|
t.Errorf("asked for %s, want /v1/models", r.URL.Path)
|
||||||
|
}
|
||||||
|
w.WriteHeader(c.code)
|
||||||
|
_, _ = w.Write([]byte(c.body))
|
||||||
|
}))
|
||||||
|
defer srv.Close()
|
||||||
|
|
||||||
|
got, err := ModelID(context.Background(), srv.URL+"/")
|
||||||
|
if c.want == "" {
|
||||||
|
if err == nil {
|
||||||
|
t.Fatalf("want an error, got label %q", got)
|
||||||
|
}
|
||||||
|
return
|
||||||
|
}
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("ModelID: %v", err)
|
||||||
|
}
|
||||||
|
if got != c.want {
|
||||||
|
t.Errorf("got %q, want %q", got, c.want)
|
||||||
|
}
|
||||||
|
})
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestModelIDUnreachable(t *testing.T) {
|
||||||
|
srv := httptest.NewServer(http.HandlerFunc(func(http.ResponseWriter, *http.Request) {}))
|
||||||
|
url := srv.URL
|
||||||
|
srv.Close() // nothing listening now
|
||||||
|
|
||||||
|
if got, err := ModelID(context.Background(), url); err == nil {
|
||||||
|
t.Fatalf("want an error from a dead server, got label %q", got)
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -422,6 +422,8 @@ func rankNote(inTop3 bool) string {
|
|||||||
// bestRecall mirrors cmd/mavend/recall.go — the gate the daemon actually
|
// bestRecall mirrors cmd/mavend/recall.go — the gate the daemon actually
|
||||||
// applies to a memory hit. Duplicated rather than imported because package main
|
// applies to a memory hit. Duplicated rather than imported because package main
|
||||||
// is not importable; recalleval_test.go asserts the two agree in behaviour.
|
// is not importable; recalleval_test.go asserts the two agree in behaviour.
|
||||||
|
// The daemon returns the whole hit (a note and a fact are said differently);
|
||||||
|
// the harness only scores what came back, so it keeps returning the text.
|
||||||
func bestRecall(results []memory.Result, minScore, minMargin float64) string {
|
func bestRecall(results []memory.Result, minScore, minMargin float64) string {
|
||||||
if !memory.Confident(results, minScore, minMargin) {
|
if !memory.Confident(results, minScore, minMargin) {
|
||||||
return ""
|
return ""
|
||||||
|
|||||||
@@ -197,7 +197,8 @@ func TestHashRecallBaseline(t *testing.T) {
|
|||||||
t.Log("\n" + rep.String() + rep.Failures())
|
t.Log("\n" + rep.String() + rep.Failures())
|
||||||
t.Log("\ngate sweep:\n" + sweep(t, router.NewHashEmbedder(hashDim), f))
|
t.Log("\ngate sweep:\n" + sweep(t, router.NewHashEmbedder(hashDim), f))
|
||||||
|
|
||||||
// 0.32 sits under the observed 0.360 recall@1.
|
// 0.32 sits under the observed 0.370 recall@1 (was 0.360 over 25 answerable
|
||||||
|
// cases; the two mixed note+fact cases added with #373 make it 27).
|
||||||
const floorRecall1 = 0.32
|
const floorRecall1 = 0.32
|
||||||
if rep.Recall1() < floorRecall1 {
|
if rep.Recall1() < floorRecall1 {
|
||||||
t.Errorf("recall@1 %.3f below ratchet %.2f — note recall regressed", rep.Recall1(), floorRecall1)
|
t.Errorf("recall@1 %.3f below ratchet %.2f — note recall regressed", rep.Recall1(), floorRecall1)
|
||||||
|
|||||||
@@ -387,6 +387,32 @@
|
|||||||
{"id": "n2", "text": "wifi channel is 6", "kind": "note"},
|
{"id": "n2", "text": "wifi channel is 6", "kind": "note"},
|
||||||
{"id": "n3", "text": "the guest network is off", "kind": "note"}
|
{"id": "n3", "text": "the guest network is off", "kind": "note"}
|
||||||
]
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "ru-mixed-031",
|
||||||
|
"lang": "ru",
|
||||||
|
"tags": ["mixed", "paraphrase", "hard"],
|
||||||
|
"note": "notes and facts in one store and the note is the answer — the daemon indexes both (Vikunja #373)",
|
||||||
|
"query": "куда я спрятал второй ключ от квартиры",
|
||||||
|
"want": "n1",
|
||||||
|
"notes": [
|
||||||
|
{"id": "n1", "text": "запасной ключ от квартиры лежит в синей коробке на полке", "kind": "note"},
|
||||||
|
{"id": "x1", "text": "поменял замок в двери двадцатого июня", "kind": "fact"},
|
||||||
|
{"id": "x2", "text": "отдал ключ соседке в мае", "kind": "fact"}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "ru-mixed-032",
|
||||||
|
"lang": "ru",
|
||||||
|
"tags": ["mixed", "distractor"],
|
||||||
|
"note": "the mirror of ru-mixed-031: the fact answers and the notes are the distractors",
|
||||||
|
"query": "когда я в последний раз заливал бензин",
|
||||||
|
"want": "x1",
|
||||||
|
"notes": [
|
||||||
|
{"id": "x1", "text": "залил полный бак в четверг вечером", "kind": "fact"},
|
||||||
|
{"id": "n1", "text": "на заправке у моста дешевле бензин", "kind": "note"},
|
||||||
|
{"id": "n2", "text": "надо поменять зимние шины", "kind": "note"}
|
||||||
|
]
|
||||||
}
|
}
|
||||||
]
|
]
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,138 @@
|
|||||||
|
// Package persona builds the one shared context block that goes in front of
|
||||||
|
// every LLM system prompt: who the owner is, how to address him, and what
|
||||||
|
// time it is right now.
|
||||||
|
//
|
||||||
|
// Why one block and not a line pasted into each prompt: there are five
|
||||||
|
// prompts (nudges, action replies, chat, note queries, general knowledge) and
|
||||||
|
// the "address him as ты" rule had only reached two of them. Five copies drift.
|
||||||
|
// One block cannot.
|
||||||
|
//
|
||||||
|
// The rules here are defaults in code, not config. Maven is feminine and the
|
||||||
|
// owner is a man addressed informally — that is a hard constraint of the
|
||||||
|
// product, so it must hold with an empty config file. Config only ADDS
|
||||||
|
// optional facts (his name, his city).
|
||||||
|
package persona
|
||||||
|
|
||||||
|
import (
|
||||||
|
"fmt"
|
||||||
|
"strings"
|
||||||
|
"time"
|
||||||
|
)
|
||||||
|
|
||||||
|
// Facts — the optional, deployment-specific half of the block. All fields may
|
||||||
|
// be empty; the block is still correct and useful without them.
|
||||||
|
type Facts struct {
|
||||||
|
OwnerName string // his name, e.g. "Ками"
|
||||||
|
City string // where he is, e.g. "Москва"
|
||||||
|
Static string // the free-text `persona` config string, appended verbatim
|
||||||
|
|
||||||
|
// The two config-gated capabilities. They are listed only when this
|
||||||
|
// deployment actually has them, because a capability she names and cannot
|
||||||
|
// do is worse than one she never mentions.
|
||||||
|
Weather bool // an open-meteo provider is configured
|
||||||
|
Telegram bool // a telegram bot token + chat id are configured
|
||||||
|
Tools bool // at least one shell act is on the allowlist
|
||||||
|
}
|
||||||
|
|
||||||
|
var ruWeekdays = [...]string{"воскресенье", "понедельник", "вторник", "среда", "четверг", "пятница", "суббота"}
|
||||||
|
|
||||||
|
var ruMonths = [...]string{
|
||||||
|
"января", "февраля", "марта", "апреля", "мая", "июня",
|
||||||
|
"июля", "августа", "сентября", "октября", "ноября", "декабря",
|
||||||
|
}
|
||||||
|
|
||||||
|
// Block renders the context block for one turn. Russian even in front of the
|
||||||
|
// English prompts: the rules it states are Russian grammar (ты/тебя, feminine
|
||||||
|
// verbs), and a Russian rule reads best stated in Russian.
|
||||||
|
//
|
||||||
|
// Keep it short. It ships on every turn to a 0.8B on laptop CPU, so every
|
||||||
|
// line here is latency.
|
||||||
|
func (f Facts) Block(now time.Time) string {
|
||||||
|
var b strings.Builder
|
||||||
|
|
||||||
|
b.WriteString("Ты — Maven, домашняя ассистентка. О себе говоришь в женском роде: \"я записала\", \"я проверила\".\n")
|
||||||
|
|
||||||
|
// The address form gets its own line. It is the thing that kept getting
|
||||||
|
// lost when it was buried in prose.
|
||||||
|
b.WriteString("ОБРАЩЕНИЕ: владелец — мужчина, всегда на \"ты\" (ты, тебя, тебе, твой) и в единственном числе (\"выпей\", \"посмотри\"). Никогда \"вы\"/\"вас\"/\"ваш\". Никогда \"он\"/\"его\" о нём — ты говоришь ему, а не о нём. Глаголы о нём — в мужском роде (\"ты забыл\").\n")
|
||||||
|
|
||||||
|
if who := f.who(); who != "" {
|
||||||
|
b.WriteString(who + "\n")
|
||||||
|
}
|
||||||
|
|
||||||
|
b.WriteString(fmt.Sprintf("Сейчас: %s, %d %s %d, %02d:%02d (местное время).\n",
|
||||||
|
ruWeekdays[int(now.Weekday())], now.Day(), ruMonths[int(now.Month())-1], now.Year(),
|
||||||
|
now.Hour(), now.Minute()))
|
||||||
|
|
||||||
|
b.WriteString("Умеешь: " + strings.Join(f.can(), "; ") +
|
||||||
|
". Других ДЕЙСТВИЙ не умеешь — если просят такое, скажи прямо.\n")
|
||||||
|
|
||||||
|
if s := strings.TrimSpace(f.Static); s != "" {
|
||||||
|
b.WriteString(s + "\n")
|
||||||
|
}
|
||||||
|
return b.String()
|
||||||
|
}
|
||||||
|
|
||||||
|
// can lists what she can really do. Every entry here is a code path that
|
||||||
|
// exists in the daemon today:
|
||||||
|
// - reminders: IntentReminder → CoreAPI.CreateReminder, fired by the tick.
|
||||||
|
// - notes and facts: IntentNote/IntentFact write, IntentQuery reads them back.
|
||||||
|
// - calendar: IntentQuery answers "что у меня сегодня" from CalendarEvents.
|
||||||
|
// - weather / telegram / shell acts: only when configured (see Facts).
|
||||||
|
//
|
||||||
|
// Nothing speculative goes in this list. A capability she offers and cannot
|
||||||
|
// perform is worse than one she never mentions.
|
||||||
|
func (f Facts) can() []string {
|
||||||
|
c := []string{
|
||||||
|
// Talking comes first, and the closing line says "действий" rather than
|
||||||
|
// "ничего", because this same block sits in front of the chat and
|
||||||
|
// general-knowledge prompts. A flat "you can do nothing else" would
|
||||||
|
// tell her to refuse the exact thing those two prompts are for.
|
||||||
|
"разговаривать и отвечать на вопросы",
|
||||||
|
"ставить напоминания",
|
||||||
|
"записывать заметки и факты и отвечать по ним",
|
||||||
|
"смотреть календарь",
|
||||||
|
}
|
||||||
|
if f.Weather {
|
||||||
|
c = append(c, "говорить погоду")
|
||||||
|
}
|
||||||
|
if f.Telegram {
|
||||||
|
c = append(c, "писать в телеграм")
|
||||||
|
}
|
||||||
|
if f.Tools {
|
||||||
|
c = append(c, "запускать разрешённые команды на сервере")
|
||||||
|
}
|
||||||
|
return c
|
||||||
|
}
|
||||||
|
|
||||||
|
// who renders the optional name/city line, or "" when neither is configured.
|
||||||
|
//
|
||||||
|
// Written as labels ("Имя владельца: ..."), not as a sentence with pronouns:
|
||||||
|
// the block's own "ты" is Maven, so "тебя зовут" would read as her name and
|
||||||
|
// "его" would model the third-person form she must never use about him.
|
||||||
|
func (f Facts) who() string {
|
||||||
|
name := strings.TrimSpace(f.OwnerName)
|
||||||
|
city := strings.TrimSpace(f.City)
|
||||||
|
switch {
|
||||||
|
case name != "" && city != "":
|
||||||
|
return "Имя владельца: " + name + ". Город: " + city + "."
|
||||||
|
case name != "":
|
||||||
|
return "Имя владельца: " + name + "."
|
||||||
|
case city != "":
|
||||||
|
return "Город: " + city + "."
|
||||||
|
}
|
||||||
|
return ""
|
||||||
|
}
|
||||||
|
|
||||||
|
// Prepend puts the block in front of a system prompt. Nil-safe: a nil renderer
|
||||||
|
// (tests, the stub paths) returns the prompt untouched.
|
||||||
|
func Prepend(block func() string, prompt string) string {
|
||||||
|
if block == nil {
|
||||||
|
return prompt
|
||||||
|
}
|
||||||
|
s := strings.TrimSpace(block())
|
||||||
|
if s == "" {
|
||||||
|
return prompt
|
||||||
|
}
|
||||||
|
return s + "\n\n" + prompt
|
||||||
|
}
|
||||||
@@ -0,0 +1,72 @@
|
|||||||
|
package persona
|
||||||
|
|
||||||
|
import (
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
)
|
||||||
|
|
||||||
|
var ref = time.Date(2026, 7, 31, 14, 5, 0, 0, time.UTC)
|
||||||
|
|
||||||
|
// The block must be correct with an empty config: the address form and the
|
||||||
|
// gender rules are hard constraints, not preferences.
|
||||||
|
func TestBlockWorksWithZeroConfig(t *testing.T) {
|
||||||
|
b := Facts{}.Block(ref)
|
||||||
|
for _, want := range []string{"женском роде", "ОБРАЩЕНИЕ", "\"ты\"", "31 июля 2026", "пятница", "14:05"} {
|
||||||
|
if !strings.Contains(b, want) {
|
||||||
|
t.Errorf("block missing %q:\n%s", want, b)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestBlockAddsOptionalFacts(t *testing.T) {
|
||||||
|
b := Facts{OwnerName: "Ками", City: "Москва", Static: "Будь краткой."}.Block(ref)
|
||||||
|
for _, want := range []string{"Ками", "Москва", "Будь краткой."} {
|
||||||
|
if !strings.Contains(b, want) {
|
||||||
|
t.Errorf("block missing %q:\n%s", want, b)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The time changes between turns, so two renders must differ.
|
||||||
|
func TestBlockRendersTimePerTurn(t *testing.T) {
|
||||||
|
a := Facts{}.Block(ref)
|
||||||
|
c := Facts{}.Block(ref.Add(time.Hour))
|
||||||
|
if a == c {
|
||||||
|
t.Errorf("block did not change with the clock:\n%s", a)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// She may only offer what this deployment actually has.
|
||||||
|
func TestCapabilitiesAreConfigGated(t *testing.T) {
|
||||||
|
bare := Facts{}.Block(ref)
|
||||||
|
for _, want := range []string{"напоминания", "заметки", "календарь"} {
|
||||||
|
if !strings.Contains(bare, want) {
|
||||||
|
t.Errorf("block missing always-on capability %q:\n%s", want, bare)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
for _, unwanted := range []string{"погоду", "телеграм", "команды"} {
|
||||||
|
if strings.Contains(bare, unwanted) {
|
||||||
|
t.Errorf("block offers unconfigured %q:\n%s", unwanted, bare)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
full := Facts{Weather: true, Telegram: true, Tools: true}.Block(ref)
|
||||||
|
for _, want := range []string{"погоду", "телеграм", "команды"} {
|
||||||
|
if !strings.Contains(full, want) {
|
||||||
|
t.Errorf("block missing configured capability %q:\n%s", want, full)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestPrependNilIsSafe(t *testing.T) {
|
||||||
|
if got := Prepend(nil, "PROMPT"); got != "PROMPT" {
|
||||||
|
t.Errorf("Prepend(nil) = %q", got)
|
||||||
|
}
|
||||||
|
if got := Prepend(func() string { return " " }, "PROMPT"); got != "PROMPT" {
|
||||||
|
t.Errorf("Prepend(blank) = %q", got)
|
||||||
|
}
|
||||||
|
if got := Prepend(func() string { return "CTX" }, "PROMPT"); got != "CTX\n\nPROMPT" {
|
||||||
|
t.Errorf("Prepend = %q", got)
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,54 @@
|
|||||||
|
package phraser
|
||||||
|
|
||||||
|
import (
|
||||||
|
"errors"
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
)
|
||||||
|
|
||||||
|
// A reply that starts a JSON object and never finishes it is a failed
|
||||||
|
// generation, not a reply. Before this, the parser returned ("", "") for these
|
||||||
|
// and every caller then shipped the raw fragment as the thing Maven said. A
|
||||||
|
// real run produced replies of literally "{" and "{\n \"".
|
||||||
|
func TestParseResponseMoodRejectsUnfinishedJSON(t *testing.T) {
|
||||||
|
for _, raw := range []string{
|
||||||
|
`{`,
|
||||||
|
"{\n \"",
|
||||||
|
`{"response": "неполн`,
|
||||||
|
`{"response": "текст", "mood":`,
|
||||||
|
} {
|
||||||
|
text, mood, err := parseResponseMood(raw)
|
||||||
|
if !errors.Is(err, errBrokenJSON) {
|
||||||
|
t.Errorf("parseResponseMood(%q) err = %v, want errBrokenJSON", raw, err)
|
||||||
|
}
|
||||||
|
if text != "" || mood != "" {
|
||||||
|
t.Errorf("parseResponseMood(%q) leaked %q/%q — a fragment must never come back as a reply", raw, text, mood)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Bare prose is still fine. Small models sometimes answer without any JSON at
|
||||||
|
// all, and that reply is usable — so the new error must not swallow it.
|
||||||
|
func TestParseResponseMoodAllowsBareProse(t *testing.T) {
|
||||||
|
for _, raw := range []string{
|
||||||
|
"норм, а ты как?",
|
||||||
|
"вот что я нашла: ключ у соседа",
|
||||||
|
} {
|
||||||
|
text, mood, err := parseResponseMood(raw)
|
||||||
|
if err != nil {
|
||||||
|
t.Errorf("parseResponseMood(%q) err = %v, want nil", raw, err)
|
||||||
|
}
|
||||||
|
// No JSON means no fields; the caller ships raw as-is.
|
||||||
|
if text != "" || mood != "" {
|
||||||
|
t.Errorf("parseResponseMood(%q) = %q/%q, want empty", raw, text, mood)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The measured failure: the model wants more than 400 characters and the old
|
||||||
|
// grammar cut it off mid-word. Guards the bound against being tightened back.
|
||||||
|
func TestGrammarStringBoundHasRoomForARealAnswer(t *testing.T) {
|
||||||
|
if !strings.Contains(responseGrammar, "{0,1000}") {
|
||||||
|
t.Error("grammar string bound is not 1000; 400 truncated real replies mid-word (see the comment on responseGrammar)")
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,33 @@
|
|||||||
|
package phraser
|
||||||
|
|
||||||
|
import (
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
)
|
||||||
|
|
||||||
|
// Every phrasing prompt must carry the shared context block. This is the
|
||||||
|
// regression guard for the bug that started this: the "ты" rule reached only
|
||||||
|
// two of the five prompts because each prompt had its own copy of the rules.
|
||||||
|
func TestEveryPromptCarriesTheContextBlock(t *testing.T) {
|
||||||
|
block := func() string { return "CTXBLOCK" }
|
||||||
|
p := &LLMPhraser{cfg: Config{ContextBlock: block}}
|
||||||
|
|
||||||
|
prompts := map[string]string{
|
||||||
|
"nudge": p.systemPrompt(),
|
||||||
|
"query": p.querySystemPrompt(),
|
||||||
|
"chat": chatSystemPrompt(block),
|
||||||
|
}
|
||||||
|
for name, got := range prompts {
|
||||||
|
if !strings.HasPrefix(got, "CTXBLOCK\n\n") {
|
||||||
|
t.Errorf("%s prompt does not start with the context block:\n%s", name, got)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Without a block the prompts are unchanged — the stub and test paths pass nil.
|
||||||
|
func TestPromptsWithoutBlockAreUnchanged(t *testing.T) {
|
||||||
|
p := &LLMPhraser{}
|
||||||
|
if p.systemPrompt() != nudgeSystem {
|
||||||
|
t.Errorf("nudge prompt changed with no block set")
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,65 @@
|
|||||||
|
package eval
|
||||||
|
|
||||||
|
import (
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
)
|
||||||
|
|
||||||
|
// TestAddressReportsEveryBreak — the real reply from a nudge eval run broke in
|
||||||
|
// two ways at once and the check named only the plural. Both must print: a
|
||||||
|
// half-reported failure reads as a milder problem than it is.
|
||||||
|
func TestAddressReportsEveryBreak(t *testing.T) {
|
||||||
|
body := "Смотрите на его потребление воды."
|
||||||
|
res := checkAddress(body)
|
||||||
|
if res.Pass {
|
||||||
|
t.Fatalf("checkAddress passed %q", body)
|
||||||
|
}
|
||||||
|
for _, want := range []string{"смотрите", "его"} {
|
||||||
|
if !strings.Contains(res.Detail, want) {
|
||||||
|
t.Errorf("detail %q does not name %q", res.Detail, want)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// One word repeated is one problem, so the detail must not say it twice.
|
||||||
|
func TestAddressDeduplicates(t *testing.T) {
|
||||||
|
res := checkAddress("Вам стоит поесть, вам это нужно.")
|
||||||
|
if res.Pass {
|
||||||
|
t.Fatal("expected failure")
|
||||||
|
}
|
||||||
|
if n := strings.Count(res.Detail, "formal"); n != 1 {
|
||||||
|
t.Errorf("detail repeats the same break %d times: %q", n, res.Detail)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The fragments a real run produced. All of them scored as non-empty replies
|
||||||
|
// before checkNonEmpty looked for letters.
|
||||||
|
func TestNonEmptyNeedsLetters(t *testing.T) {
|
||||||
|
for _, body := range []string{
|
||||||
|
"{",
|
||||||
|
"{\n \"",
|
||||||
|
"15-16",
|
||||||
|
`{"`,
|
||||||
|
" ",
|
||||||
|
"...",
|
||||||
|
} {
|
||||||
|
if got := checkNonEmpty(body); got.Pass {
|
||||||
|
t.Errorf("checkNonEmpty(%q) passed — that is not a reply", body)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// And it must not start failing real replies. Latin counts as well as Cyrillic:
|
||||||
|
// answers about ssd or vpn are legitimately part English.
|
||||||
|
func TestNonEmptyAcceptsRealReplies(t *testing.T) {
|
||||||
|
for _, body := range []string{
|
||||||
|
"норм, а ты как?",
|
||||||
|
"вот что я нашла: ключ у соседа",
|
||||||
|
"ssd быстрее hdd.",
|
||||||
|
"9 минут.",
|
||||||
|
} {
|
||||||
|
if got := checkNonEmpty(body); !got.Pass {
|
||||||
|
t.Errorf("checkNonEmpty(%q) failed: %s", body, got.Detail)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,42 @@
|
|||||||
|
package eval
|
||||||
|
|
||||||
|
import "testing"
|
||||||
|
|
||||||
|
func TestAddressTimeWordDoesNotBlind(t *testing.T) {
|
||||||
|
// A nudge that opens with a time word must still be caught. Without the
|
||||||
|
// time words in the stoplist, "сегодня" was read as the third party.
|
||||||
|
for _, s := range []string{
|
||||||
|
"сегодня он не ел 11 дней",
|
||||||
|
"вчера он не пил воду",
|
||||||
|
"опять он забыл про таблетки",
|
||||||
|
} {
|
||||||
|
if r := checkAddress(s); r.Pass {
|
||||||
|
t.Errorf("checkAddress(%q) passed, want a third-person failure", s)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
// Still must not fire when a third party really is named.
|
||||||
|
for _, s := range []string{
|
||||||
|
"сегодня сервис упал, он не отвечает",
|
||||||
|
"ты не пил воду четыре часа",
|
||||||
|
} {
|
||||||
|
if r := checkAddress(s); !r.Pass {
|
||||||
|
t.Errorf("checkAddress(%q) failed: %s", s, r.Detail)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestAddressVerbIsNotAnAntecedent — a nudge is mostly verbs, and a verb is
|
||||||
|
// never who "он" refers to. This exact string passed the check before.
|
||||||
|
func TestAddressVerbIsNotAnAntecedent(t *testing.T) {
|
||||||
|
s := "попробуй встать и отдохнуть — у него есть перерыв"
|
||||||
|
if r := checkAddress(s); r.Pass {
|
||||||
|
t.Errorf("checkAddress(%q) passed, want a third-person failure", s)
|
||||||
|
}
|
||||||
|
// Still missed, and this is the documented hole: "выпей воды, он не пил" has
|
||||||
|
// a real noun ("воды") before the pronoun, so the scan believes somebody
|
||||||
|
// else was named. Telling that apart needs a parser, not a suffix rule.
|
||||||
|
// A named third party still wins over the verbs around it.
|
||||||
|
if r := checkAddress("сервис упал, он не отвечает"); !r.Pass {
|
||||||
|
t.Errorf("checkAddress on a real third party failed: %s", r.Detail)
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -17,10 +17,18 @@ const (
|
|||||||
CheckFeminine = "feminine" // her self-reference is feminine (hard constraint)
|
CheckFeminine = "feminine" // her self-reference is feminine (hard constraint)
|
||||||
CheckCringe = "cringe" // DESIGN.md § Non-goals, "not a relationship"
|
CheckCringe = "cringe" // DESIGN.md § Non-goals, "not a relationship"
|
||||||
CheckOnTopic = "ontopic" // says the thing the rule is about
|
CheckOnTopic = "ontopic" // says the thing the rule is about
|
||||||
|
|
||||||
|
// CheckHisGender — the other half of the persona rule: SHE is feminine, HE
|
||||||
|
// is male. "ты давно не отдыхала" addresses the operator as a woman.
|
||||||
|
CheckHisGender = "hisgender"
|
||||||
|
|
||||||
|
// CheckAddress — she talks TO him, informally, one to one. Not "вы", not
|
||||||
|
// "он". See the comment block above checkAddress.
|
||||||
|
CheckAddress = "address"
|
||||||
)
|
)
|
||||||
|
|
||||||
// CheckNames — report order.
|
// CheckNames — report order.
|
||||||
var CheckNames = []string{CheckMood, CheckLang, CheckLength, CheckFeminine, CheckCringe, CheckOnTopic}
|
var CheckNames = []string{CheckMood, CheckLang, CheckLength, CheckFeminine, CheckHisGender, CheckAddress, CheckCringe, CheckOnTopic}
|
||||||
|
|
||||||
// Result — one check on one message.
|
// Result — one check on one message.
|
||||||
type Result struct {
|
type Result struct {
|
||||||
@@ -52,6 +60,8 @@ func RunChecks(c Case, body, mood string) []Result {
|
|||||||
checkLang(body),
|
checkLang(body),
|
||||||
checkLength(body),
|
checkLength(body),
|
||||||
checkFeminine(body),
|
checkFeminine(body),
|
||||||
|
checkHisGender(body),
|
||||||
|
checkAddress(body),
|
||||||
checkCringe(body),
|
checkCringe(body),
|
||||||
checkOnTopic(c, body),
|
checkOnTopic(c, body),
|
||||||
}
|
}
|
||||||
@@ -176,6 +186,315 @@ func checkFeminine(body string) Result {
|
|||||||
return Result{CheckFeminine, true, ""}
|
return Result{CheckFeminine, true, ""}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// --- he is male ----------------------------------------------------------
|
||||||
|
//
|
||||||
|
// The mirror of checkFeminine, and the failure it was written for: the model
|
||||||
|
// wrote "ты давно не отдыхала", which addresses the operator as a woman. That
|
||||||
|
// scored clean, because checkFeminine only ever looks at how SHE speaks about
|
||||||
|
// herself.
|
||||||
|
//
|
||||||
|
// How it works: Russian past tense is gendered by suffix, -л (m) / -ла (f). So
|
||||||
|
// this looks for feminine past-tense words in a sentence that also talks TO him
|
||||||
|
// ("ты", "тебя", "тебе", "твой", …). A feminine verb that belongs to her ("я
|
||||||
|
// заметила", "напомнила тебе") is skipped — that one is correct.
|
||||||
|
//
|
||||||
|
// Honest about the limits: this is a suffix rule, not a parser.
|
||||||
|
// - False positives: a feminine noun can be the subject in the same sentence
|
||||||
|
// ("зарядка была утром, ты её пропустил"). The guard below skips a verb whose
|
||||||
|
// previous word looks like a feminine noun, which helps but will not always
|
||||||
|
// be right.
|
||||||
|
// - False negatives: gender also shows up outside the past tense (short
|
||||||
|
// adjectives, "сама"), and none of that is checked here.
|
||||||
|
//
|
||||||
|
// That is acceptable for an eval check. It is a signal to read the message, not
|
||||||
|
// a grammar verdict, and every hit prints the word it tripped on so a human can
|
||||||
|
// disagree.
|
||||||
|
|
||||||
|
// hisMarkers — words that mean the sentence is addressed to him.
|
||||||
|
var hisMarkers = map[string]bool{
|
||||||
|
"ты": true, "тебя": true, "тебе": true, "тобой": true, "тобою": true,
|
||||||
|
"твой": true, "твоя": true, "твоё": true, "твое": true, "твои": true, "твою": true,
|
||||||
|
}
|
||||||
|
|
||||||
|
// notFeminineVerb — ordinary words ending in "-ла" that are not verbs. Small on
|
||||||
|
// purpose: it only has to cover words a nudge might actually use.
|
||||||
|
//
|
||||||
|
// Words that are both a noun and a verb are deliberately NOT here. "села",
|
||||||
|
// "мыла" and "стекла" are nouns on paper, but in a nudge they are almost always
|
||||||
|
// verbs ("ты села", "ты мыла"), and listing them would make the check miss the
|
||||||
|
// exact thing it is for. Missing a real hit is worse than one false alarm.
|
||||||
|
var notFeminineVerb = map[string]bool{
|
||||||
|
"школа": true, "скала": true, "игла": true, "метла": true, "смола": true,
|
||||||
|
"дела": true, "тела": true, "масла": true, "весла": true,
|
||||||
|
"зола": true, "пчела": true, "числа": true,
|
||||||
|
}
|
||||||
|
|
||||||
|
// femininePast reports whether a word looks like a feminine past-tense verb:
|
||||||
|
// "отдыхала", "поела", "выспалась".
|
||||||
|
func femininePast(w string) bool {
|
||||||
|
if len([]rune(w)) < 3 || notFeminineVerb[w] {
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
return strings.HasSuffix(w, "ла") || strings.HasSuffix(w, "лась")
|
||||||
|
}
|
||||||
|
|
||||||
|
// looksFeminineNoun — a crude guard against "зарядка была": a word right before
|
||||||
|
// the verb that ends in "а"/"я" and is not itself a verb is probably the subject.
|
||||||
|
func looksFeminineNoun(w string) bool {
|
||||||
|
if femininePast(w) || len([]rune(w)) < 3 {
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
return strings.HasSuffix(w, "а") || strings.HasSuffix(w, "я")
|
||||||
|
}
|
||||||
|
|
||||||
|
// sentenceRE splits on sentence-ending punctuation, so a feminine verb in one
|
||||||
|
// sentence is not blamed on a "ты" in the next.
|
||||||
|
var sentenceRE = regexp.MustCompile(`[.!?;…]+`)
|
||||||
|
|
||||||
|
func checkHisGender(body string) Result {
|
||||||
|
for _, sentence := range sentenceRE.Split(strings.ToLower(body), -1) {
|
||||||
|
words := wordRE.FindAllString(sentence, -1)
|
||||||
|
addressed := false
|
||||||
|
for _, w := range words {
|
||||||
|
if hisMarkers[w] {
|
||||||
|
addressed = true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if !addressed {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
for i, w := range words {
|
||||||
|
if !femininePast(w) || hersNotHis(words, i) {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if i > 0 && looksFeminineNoun(prevWord(words, i)) {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
return Result{CheckHisGender, false,
|
||||||
|
fmt.Sprintf("feminine %q addressed to him — he is male", w)}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return Result{CheckHisGender, true, ""}
|
||||||
|
}
|
||||||
|
|
||||||
|
// hersNotHis — the verb is Maven's own if "я" comes shortly before it, or if the
|
||||||
|
// thing she did was done to him ("напомнила тебе", "проверила за тебя").
|
||||||
|
func hersNotHis(words []string, i int) bool {
|
||||||
|
for j := i - 1; j >= 0 && j >= i-3; j-- {
|
||||||
|
if words[j] == "я" {
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if i+1 < len(words) {
|
||||||
|
switch words[i+1] {
|
||||||
|
case "тебе", "тебя", "за", "тобой":
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
|
||||||
|
// prevWord — the word before i, skipping "не" and punctuation, so "не отдыхала"
|
||||||
|
// still sees the subject.
|
||||||
|
func prevWord(words []string, i int) string {
|
||||||
|
for j := i - 1; j >= 0; j-- {
|
||||||
|
w := words[j]
|
||||||
|
if w == "не" || w == "ни" || !unicode.Is(unicode.Cyrillic, []rune(w)[0]) {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
return w
|
||||||
|
}
|
||||||
|
return ""
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- how she addresses him ------------------------------------------------
|
||||||
|
//
|
||||||
|
// Persona hard constraint: Maven speaks TO him, informally, one to one. The
|
||||||
|
// phrasing eval produced two breaks of it, and both scored clean:
|
||||||
|
//
|
||||||
|
// - "Приходите… Жду вас" — the formal plural. Correct is ты/тебя/тебе and a
|
||||||
|
// singular imperative ("приходи", "жду тебя").
|
||||||
|
// - "Он не ел 11 дней" — she talks ABOUT him, in the third person, as if
|
||||||
|
// reporting to somebody else. Correct is "ты не ел 11 дней".
|
||||||
|
//
|
||||||
|
// Like checkHisGender this is a keyword + suffix heuristic, NOT a parser. Every
|
||||||
|
// hit prints the word it tripped on, so a false alarm is obvious at a glance and
|
||||||
|
// can be dismissed.
|
||||||
|
//
|
||||||
|
// Part 1, formal address. Two signals:
|
||||||
|
// - the "вы" pronoun family, matched as whole words, so there is nothing to
|
||||||
|
// exclude — "вы" and "вас" are never anything else.
|
||||||
|
// - a plural verb ending: -ите/-ете/-йте/-ьте ("приходите", "выпейте",
|
||||||
|
// "не забудьте", "хотите"). Nouns in the prepositional case share those
|
||||||
|
// endings ("в интернете", "в свете"), so a word right after a preposition is
|
||||||
|
// skipped. That is the whole exclusion list, on purpose: a bigger one would
|
||||||
|
// start swallowing real imperatives.
|
||||||
|
//
|
||||||
|
// Part 2, third person. "он" is perfectly fine when the message really is about
|
||||||
|
// somebody or something else ("сервис упал, он не отвечает"). The way to tell
|
||||||
|
// them apart: a legitimate third person has an ANTECEDENT — the thing it refers
|
||||||
|
// to was named earlier in the message. So "он" is only flagged when nothing
|
||||||
|
// before it in the message could be that thing.
|
||||||
|
//
|
||||||
|
// Where this gives up, plainly:
|
||||||
|
// - it only looks BACKWARD. "Он не отвечает, сервис упал" names the subject
|
||||||
|
// after the pronoun and is flagged wrongly.
|
||||||
|
// - any noun earlier in the message counts as an antecedent, even when it is
|
||||||
|
// not one ("после обеда он не ел", "выпей воды, он не пил" — both missed).
|
||||||
|
// Verbs and time words no longer count, which covers the usual nudge, but a
|
||||||
|
// plain noun before the pronoun still blinds it. The
|
||||||
|
// common time words are stoplisted so the usual nudge opening does not
|
||||||
|
// blind it, but a message with any other noun in front still slips through.
|
||||||
|
// This is the check's real hole; widening it further would start flagging
|
||||||
|
// legitimate third-party messages, so it stops here.
|
||||||
|
// - a message that opens with "ты" and only later slips into "он" is missed,
|
||||||
|
// because "ты" itself is skipped but the words around it are not.
|
||||||
|
// - formal address outside these endings (short adjectives, "вашими" style
|
||||||
|
// forms not listed) is missed.
|
||||||
|
|
||||||
|
// addressWordRE also takes Latin words, because "him"/"he" is the same break in
|
||||||
|
// English.
|
||||||
|
var addressWordRE = regexp.MustCompile(`[\p{Cyrillic}]+|[a-zA-Z]+|[,.;:!?…—-]`)
|
||||||
|
|
||||||
|
// formalPronouns — the "вы" family. Whole-word match, so no false hits.
|
||||||
|
var formalPronouns = map[string]bool{
|
||||||
|
"вы": true, "вас": true, "вам": true, "вами": true,
|
||||||
|
"ваш": true, "ваша": true, "ваше": true, "ваши": true,
|
||||||
|
"вашего": true, "вашей": true, "вашему": true, "вашим": true,
|
||||||
|
"вашими": true, "вашу": true,
|
||||||
|
}
|
||||||
|
|
||||||
|
// prepositions — used twice: to skip prepositional-case nouns that look like
|
||||||
|
// plural verbs, and as words that cannot be what "он" refers to.
|
||||||
|
var prepositions = map[string]bool{
|
||||||
|
"в": true, "во": true, "на": true, "о": true, "об": true, "обо": true,
|
||||||
|
"при": true, "по": true, "за": true, "из": true, "с": true, "со": true,
|
||||||
|
"к": true, "ко": true, "до": true, "от": true, "у": true, "над": true,
|
||||||
|
"под": true, "про": true, "без": true, "для": true, "через": true,
|
||||||
|
}
|
||||||
|
|
||||||
|
// pluralVerb reports whether a word looks like a plural/formal verb form:
|
||||||
|
// "приходите", "выпейте", "забудьте", "хотите".
|
||||||
|
func pluralVerb(w string) bool {
|
||||||
|
if len([]rune(w)) < 5 {
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
return strings.HasSuffix(w, "ите") || strings.HasSuffix(w, "ете") ||
|
||||||
|
strings.HasSuffix(w, "йте") || strings.HasSuffix(w, "ьте")
|
||||||
|
}
|
||||||
|
|
||||||
|
// thirdPersonHim — pronouns that would be talking about him instead of to him.
|
||||||
|
var thirdPersonHim = map[string]bool{
|
||||||
|
"он": true, "его": true, "ему": true, "него": true, "нему": true, "ним": true,
|
||||||
|
"he": true, "him": true, "his": true,
|
||||||
|
}
|
||||||
|
|
||||||
|
// notAnAntecedent — words that cannot be the thing "он" refers to: pronouns,
|
||||||
|
// particles, conjunctions, adverbs of time. If only these come before "он", the
|
||||||
|
// message never named a third party and "он" is him.
|
||||||
|
var notAnAntecedent = map[string]bool{
|
||||||
|
"не": true, "ни": true, "и": true, "а": true, "но": true, "да": true,
|
||||||
|
"же": true, "бы": true, "ли": true, "вот": true, "уже": true,
|
||||||
|
"ещё": true, "еще": true, "тоже": true, "там": true, "тут": true,
|
||||||
|
"здесь": true, "это": true, "что": true, "как": true, "когда": true,
|
||||||
|
"чтобы": true, "потому": true, "сейчас": true, "потом": true,
|
||||||
|
// Time words. A nudge almost always opens with one ("сегодня он не ел"),
|
||||||
|
// and without them the very next word is read as the person being talked
|
||||||
|
// about, so the check misses the exact break it was written for.
|
||||||
|
"сегодня": true, "вчера": true, "завтра": true, "послезавтра": true,
|
||||||
|
"утром": true, "днём": true, "днем": true, "вечером": true, "ночью": true,
|
||||||
|
"опять": true, "снова": true, "весь": true, "всю": true, "целый": true,
|
||||||
|
"я": true, "мне": true, "меня": true, "мной": true, "мы": true, "нас": true,
|
||||||
|
"ты": true, "тебя": true, "тебе": true, "тобой": true,
|
||||||
|
"твой": true, "твоя": true, "твоё": true, "твое": true, "твои": true, "твою": true,
|
||||||
|
}
|
||||||
|
|
||||||
|
// looksVerb — a verb is never the thing "он" refers to, so it must not count as
|
||||||
|
// an antecedent. Past tense keeps "сервис упал, он не отвечает" working off
|
||||||
|
// "сервис"; the infinitive and imperative endings are here because a nudge is
|
||||||
|
// mostly made of them ("попробуй встать и отдохнуть — у него есть перерыв"
|
||||||
|
// slipped through with "попробуй" taken for the person being talked about).
|
||||||
|
func looksVerb(w string) bool {
|
||||||
|
r := []rune(w)
|
||||||
|
if len(r) < 3 {
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
for _, suf := range []string{
|
||||||
|
"л", "ла", "ло", "ли", // past tense
|
||||||
|
"ть", "ться", "ти", "чь", // infinitive
|
||||||
|
"й", "йся", "йте", // imperative
|
||||||
|
} {
|
||||||
|
if strings.HasSuffix(w, suf) {
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
|
||||||
|
func checkAddress(body string) Result {
|
||||||
|
words := addressWordRE.FindAllString(strings.ToLower(body), -1)
|
||||||
|
|
||||||
|
// Every break, not just the first. A bad reply usually breaks in more than
|
||||||
|
// one way at once — "Смотрите на его потребление воды" is a plural imperative
|
||||||
|
// AND third person about him — and reporting only the first hid the second,
|
||||||
|
// which made the failure look milder than it was.
|
||||||
|
var breaks []string
|
||||||
|
seen := map[string]bool{}
|
||||||
|
add := func(msg string) {
|
||||||
|
if seen[msg] {
|
||||||
|
return // the same word twice in one message is one problem, not two
|
||||||
|
}
|
||||||
|
seen[msg] = true
|
||||||
|
breaks = append(breaks, msg)
|
||||||
|
}
|
||||||
|
|
||||||
|
for i, w := range words {
|
||||||
|
if formalPronouns[w] {
|
||||||
|
add(fmt.Sprintf("formal %q — she says ты/тебя/тебе", w))
|
||||||
|
}
|
||||||
|
if pluralVerb(w) && !(i > 0 && prepositions[words[i-1]]) {
|
||||||
|
add(fmt.Sprintf("plural imperative %q — she uses the singular", w))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
for i, w := range words {
|
||||||
|
if !thirdPersonHim[w] {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
named := false
|
||||||
|
for j := 0; j < i; j++ {
|
||||||
|
p := words[j]
|
||||||
|
if !unicode.Is(unicode.Cyrillic, []rune(p)[0]) && !isLatinWord(p) {
|
||||||
|
continue // punctuation
|
||||||
|
}
|
||||||
|
// pluralVerb as well as looksVerb: looksVerb knows the imperative in
|
||||||
|
// -й/-йте but not the -те plural ("смотрите"), so "Смотрите на его
|
||||||
|
// потребление воды" counted "смотрите" as the person being talked
|
||||||
|
// about and the "его" never printed. Third time a verb form has
|
||||||
|
// blinded this check — if a fourth turns up, the antecedent test
|
||||||
|
// wants a real morphology table, not another suffix.
|
||||||
|
if notAnAntecedent[p] || prepositions[p] || thirdPersonHim[p] || looksVerb(p) || pluralVerb(p) {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
named = true
|
||||||
|
break
|
||||||
|
}
|
||||||
|
if !named {
|
||||||
|
add(fmt.Sprintf("third person %q with nobody else named — she talks to him, not about him", w))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
if len(breaks) > 0 {
|
||||||
|
return Result{CheckAddress, false, strings.Join(breaks, " + ")}
|
||||||
|
}
|
||||||
|
return Result{CheckAddress, true, ""}
|
||||||
|
}
|
||||||
|
|
||||||
|
func isLatinWord(w string) bool {
|
||||||
|
r := []rune(w)[0]
|
||||||
|
return (r >= 'a' && r <= 'z') || (r >= 'A' && r <= 'Z')
|
||||||
|
}
|
||||||
|
|
||||||
// --- the cringe checks ---------------------------------------------------
|
// --- the cringe checks ---------------------------------------------------
|
||||||
//
|
//
|
||||||
// "Think Jarvis without the cringe part". DESIGN.md § Non-goals: "Not a
|
// "Think Jarvis without the cringe part". DESIGN.md § Non-goals: "Not a
|
||||||
@@ -276,12 +595,60 @@ func checkCringe(body string) Result {
|
|||||||
// checkOnTopic — the message must name the thing the rule is about. A nudge
|
// checkOnTopic — the message must name the thing the rule is about. A nudge
|
||||||
// that never mentions water leaves the operator with a chime and no action.
|
// that never mentions water leaves the operator with a chime and no action.
|
||||||
func checkOnTopic(c Case, body string) Result {
|
func checkOnTopic(c Case, body string) Result {
|
||||||
|
return checkOnTopicAny(c.WantAny, body)
|
||||||
|
}
|
||||||
|
|
||||||
|
// checkOnTopicAny is the same test over a bare want-list, so the talk scorer can
|
||||||
|
// reuse it without owning a nudge Case.
|
||||||
|
func checkOnTopicAny(wantAny []string, body string) Result {
|
||||||
low := strings.ToLower(body)
|
low := strings.ToLower(body)
|
||||||
for _, want := range c.WantAny {
|
for _, want := range wantAny {
|
||||||
if strings.Contains(low, strings.ToLower(want)) {
|
if strings.Contains(low, strings.ToLower(want)) {
|
||||||
return Result{CheckOnTopic, true, ""}
|
return Result{CheckOnTopic, true, ""}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
return Result{CheckOnTopic, false,
|
return Result{CheckOnTopic, false,
|
||||||
fmt.Sprintf("mentions none of %v", c.WantAny)}
|
fmt.Sprintf("mentions none of %v", wantAny)}
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- shape checks for the free-form paths --------------------------------
|
||||||
|
//
|
||||||
|
// The nudge checks assume one short sentence. Chat and query replies are longer
|
||||||
|
// by design, so the only shape worth testing there is that the model produced a
|
||||||
|
// reply at all and did not trail off. Both are failure modes the fallbacks in
|
||||||
|
// llmphraser.go hide: a truncated or empty generation still returns nil error.
|
||||||
|
|
||||||
|
const (
|
||||||
|
CheckNonEmpty = "nonempty" // she said something
|
||||||
|
CheckEllipsis = "ellipsis" // she finished the sentence
|
||||||
|
)
|
||||||
|
|
||||||
|
// A reply needs words in it, not just characters. This check used to test for a
|
||||||
|
// non-empty string, which scored 27/27 on a run where two replies were "{" and
|
||||||
|
// "{\n \"" — punctuation passed as content. Braces, quotes, digits and spaces
|
||||||
|
// are all empty in the only sense that matters.
|
||||||
|
//
|
||||||
|
// Digits alone fail too, and that is deliberate: the same run answered "сколько
|
||||||
|
// варить яйцо вкрутую?" with "15-16". No unit, no words, and it is also the
|
||||||
|
// wrong number. Whatever that is, it is not something she said.
|
||||||
|
func checkNonEmpty(body string) Result {
|
||||||
|
if strings.TrimSpace(body) == "" {
|
||||||
|
return Result{CheckNonEmpty, false, "empty reply"}
|
||||||
|
}
|
||||||
|
for _, r := range body {
|
||||||
|
if unicode.IsLetter(r) {
|
||||||
|
return Result{CheckNonEmpty, true, ""}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return Result{CheckNonEmpty, false, fmt.Sprintf("no letters in the reply %q — punctuation or digits only", strings.TrimSpace(body))}
|
||||||
|
}
|
||||||
|
|
||||||
|
// checkEllipsis — a reply ending in "…" or "..." is a generation that ran out of
|
||||||
|
// tokens, not a stylistic pause. Mid-sentence ellipses are left alone.
|
||||||
|
func checkEllipsis(body string) Result {
|
||||||
|
trimmed := strings.TrimRight(strings.TrimSpace(body), `"'»)`)
|
||||||
|
if strings.HasSuffix(trimmed, "…") || strings.HasSuffix(trimmed, "...") {
|
||||||
|
return Result{CheckEllipsis, false, "reply trails off in an ellipsis — likely truncated"}
|
||||||
|
}
|
||||||
|
return Result{CheckEllipsis, true, ""}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -281,7 +281,7 @@ func (r Report) String() string {
|
|||||||
fmt.Fprintf(&b, "%s: %d/%d cases pass every check (%.1f%%), %d errors\n",
|
fmt.Fprintf(&b, "%s: %d/%d cases pass every check (%.1f%%), %d errors\n",
|
||||||
r.Name, r.Passed, r.Total, 100*r.Accuracy(), r.Errors)
|
r.Name, r.Passed, r.Total, 100*r.Accuracy(), r.Errors)
|
||||||
for _, name := range CheckNames {
|
for _, name := range CheckNames {
|
||||||
fmt.Fprintf(&b, " %-9s %d/%d\n", name, r.ByCheck[name], r.Total)
|
fmt.Fprintf(&b, " %-10s %d/%d\n", name, r.ByCheck[name], r.Total)
|
||||||
}
|
}
|
||||||
fmt.Fprintf(&b, " latency: p50 %s p95 %s max %s\n", r.P50, r.P95, r.Max)
|
fmt.Fprintf(&b, " latency: p50 %s p95 %s max %s\n", r.P50, r.P95, r.Max)
|
||||||
fmt.Fprintf(&b, " by rule: %s\n", renderStats(r.ByRule))
|
fmt.Fprintf(&b, " by rule: %s\n", renderStats(r.ByRule))
|
||||||
|
|||||||
@@ -72,10 +72,12 @@ func TestStubBaseline(t *testing.T) {
|
|||||||
// ceiling ("you've been at your desk for 4 hours without a break — step
|
// ceiling ("you've been at your desk for 4 hours without a break — step
|
||||||
// away for a bit." is 76 chars but 16 words). Left failing rather than
|
// away for a bit." is 76 chars but 16 words). Left failing rather than
|
||||||
// raising the ceiling to hide it.
|
// raising the ceiling to hide it.
|
||||||
CheckLength: 12,
|
CheckLength: 12,
|
||||||
CheckFeminine: 15,
|
CheckFeminine: 15,
|
||||||
CheckCringe: 15,
|
CheckHisGender: 15,
|
||||||
CheckOnTopic: 12,
|
CheckAddress: 15,
|
||||||
|
CheckCringe: 15,
|
||||||
|
CheckOnTopic: 12,
|
||||||
}
|
}
|
||||||
for name, floor := range floors {
|
for name, floor := range floors {
|
||||||
if rep.ByCheck[name] < floor {
|
if rep.ByCheck[name] < floor {
|
||||||
@@ -104,6 +106,14 @@ func TestChecksCatchWhatTheyClaim(t *testing.T) {
|
|||||||
{"masculine predicative", "я должен сказать: попей воды.", CheckFeminine},
|
{"masculine predicative", "я должен сказать: попей воды.", CheckFeminine},
|
||||||
// The other direction: HE is male, so second-person masculine is right.
|
// The other direction: HE is male, so second-person masculine is right.
|
||||||
{"second person masculine ok", "ты не пил воду четыре часа.", ""},
|
{"second person masculine ok", "ты не пил воду четыре часа.", ""},
|
||||||
|
// The real observed failure: she addressed him as a woman.
|
||||||
|
{"feminine second person", "ты давно не отдыхала — попей воды.", CheckHisGender},
|
||||||
|
{"feminine second person no dash", "ты пила воду четыре часа назад.", CheckHisGender},
|
||||||
|
// Her own feminine verb next to "ты" is correct and must not be flagged.
|
||||||
|
{"her feminine verb near ты", "я заметила, что ты не пил воду.", ""},
|
||||||
|
{"her feminine verb about him", "напомнила тебе про воду.", ""},
|
||||||
|
// A feminine noun subject in the same sentence is not him.
|
||||||
|
{"feminine noun subject ok", "зарядка была утром, ты её пропустил, попей воды.", ""},
|
||||||
{"feminine self ok", "я заметила: воды не было четыре часа.", ""},
|
{"feminine self ok", "я заметила: воды не было четыре часа.", ""},
|
||||||
{"pet name", "милый, попей воды.", CheckCringe},
|
{"pet name", "милый, попей воды.", CheckCringe},
|
||||||
{"emoji", "попей воды 💧", CheckCringe},
|
{"emoji", "попей воды 💧", CheckCringe},
|
||||||
@@ -114,6 +124,10 @@ func TestChecksCatchWhatTheyClaim(t *testing.T) {
|
|||||||
{"asks how he feels", "как ты себя чувствуешь? попей воды.", CheckCringe},
|
{"asks how he feels", "как ты себя чувствуешь? попей воды.", CheckCringe},
|
||||||
{"praise", "молодец! теперь попей воды.", CheckCringe},
|
{"praise", "молодец! теперь попей воды.", CheckCringe},
|
||||||
{"off topic", "пора бы уже что-то сделать.", CheckOnTopic},
|
{"off topic", "пора бы уже что-то сделать.", CheckOnTopic},
|
||||||
|
// The two recorded persona breaks from the phrasing eval run. Pinned as
|
||||||
|
// unit tests because an eval run is sampled and may not reproduce them.
|
||||||
|
{"formal plural", "Приходите… Жду вас", CheckAddress},
|
||||||
|
{"third person about him", "Он не ел 11 дней", CheckAddress},
|
||||||
}
|
}
|
||||||
|
|
||||||
for _, tc := range cases {
|
for _, tc := range cases {
|
||||||
@@ -135,6 +149,36 @@ func TestChecksCatchWhatTheyClaim(t *testing.T) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// TestAddressCheck — the address check on its own, so the messages that must NOT
|
||||||
|
// trip it can be written without also having to satisfy the on-topic check.
|
||||||
|
func TestAddressCheck(t *testing.T) {
|
||||||
|
bad := []string{
|
||||||
|
"Приходите… Жду вас", // the recorded formal-plural break
|
||||||
|
"Он не ел 11 дней", // the recorded third-person break
|
||||||
|
"Выпейте воды, пожалуйста.", // plural imperative on its own
|
||||||
|
"Ваш обед был давно.", // formal possessive
|
||||||
|
}
|
||||||
|
for _, body := range bad {
|
||||||
|
if r := checkAddress(body); r.Pass {
|
||||||
|
t.Errorf("persona break not caught: %q", body)
|
||||||
|
} else {
|
||||||
|
t.Logf("%q -> %s", body, r.Detail)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
good := []string{
|
||||||
|
"ты не пил воду четыре часа — попей.", // correct informal address
|
||||||
|
"сервис netdata упал, он не отвечает.", // legitimately about a third party
|
||||||
|
"я заметила, что зарядка была утром.", // no address at all
|
||||||
|
"в интернете опять тихо, всё работает.", // "интернете" is a noun, not an imperative
|
||||||
|
}
|
||||||
|
for _, body := range good {
|
||||||
|
if r := checkAddress(body); !r.Pass {
|
||||||
|
t.Errorf("clean message flagged: %q -> %s", body, r.Detail)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
func TestMoodCheckUsesTheEnum(t *testing.T) {
|
func TestMoodCheckUsesTheEnum(t *testing.T) {
|
||||||
if r := checkMood("cheerful"); r.Pass {
|
if r := checkMood("cheerful"); r.Pass {
|
||||||
t.Error("mood outside the enum passed")
|
t.Error("mood outside the enum passed")
|
||||||
|
|||||||
@@ -7,6 +7,8 @@ import (
|
|||||||
"testing"
|
"testing"
|
||||||
"time"
|
"time"
|
||||||
|
|
||||||
|
"github.com/kami/maven/internal/llm"
|
||||||
|
"github.com/kami/maven/internal/persona"
|
||||||
"github.com/kami/maven/internal/phraser"
|
"github.com/kami/maven/internal/phraser"
|
||||||
)
|
)
|
||||||
|
|
||||||
@@ -32,6 +34,7 @@ func TestLLMPhrasingBaseline(t *testing.T) {
|
|||||||
// case as a phrasing error and read as "the model cannot phrase".
|
// case as a phrasing error and read as "the model cannot phrase".
|
||||||
noProxyLoopback(t)
|
noProxyLoopback(t)
|
||||||
|
|
||||||
|
ctx := context.Background()
|
||||||
f, err := Load()
|
f, err := Load()
|
||||||
if err != nil {
|
if err != nil {
|
||||||
t.Fatalf("Load: %v", err)
|
t.Fatalf("Load: %v", err)
|
||||||
@@ -41,10 +44,25 @@ func TestLLMPhrasingBaseline(t *testing.T) {
|
|||||||
// Generous: an unconstrained 0.8B can spend a minute thinking before it
|
// Generous: an unconstrained 0.8B can spend a minute thinking before it
|
||||||
// writes a word, and a timeout would be scored as a model failure.
|
// writes a word, and a timeout would be scored as a model failure.
|
||||||
cfg.Timeout = 5 * time.Minute
|
cfg.Timeout = 5 * time.Minute
|
||||||
|
// The same shared context block the daemon prepends (internal/persona),
|
||||||
|
// with an empty config — that is the deployment we actually ship.
|
||||||
|
cfg.ContextBlock = func() string { return persona.Facts{}.Block(time.Now()) }
|
||||||
p := phraser.NewLLMPhraserAt(base, cfg)
|
p := phraser.NewLLMPhraserAt(base, cfg)
|
||||||
defer p.Close()
|
defer p.Close()
|
||||||
|
|
||||||
rep, err := Score(context.Background(), "llm (0.8B, built-in persona)", p, f)
|
// Label the run with whatever gguf the server actually has loaded. It used
|
||||||
|
// to say "0.8B" no matter what, so two runs of two different models came
|
||||||
|
// out named the same and were easy to mix up when comparing.
|
||||||
|
model, err := llm.ModelID(ctx, base)
|
||||||
|
if err != nil {
|
||||||
|
// An unlabelled score is still a score, but say so loudly — a made-up
|
||||||
|
// name in a bake-off table is worse than no name.
|
||||||
|
t.Logf("could not read model id from %s: %v — report will say %q", base, err, llm.UnknownModel)
|
||||||
|
model = llm.UnknownModel
|
||||||
|
}
|
||||||
|
t.Logf("scoring model %s at %s", model, base)
|
||||||
|
|
||||||
|
rep, err := Score(ctx, "llm ("+model+", built-in persona)", p, f)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
t.Fatalf("Score: %v", err)
|
t.Fatalf("Score: %v", err)
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,265 @@
|
|||||||
|
package eval
|
||||||
|
|
||||||
|
// This file scores the CONVERSATIONAL paths, the ones the nudge fixture never
|
||||||
|
// touches: chat, query-with-notes, and general knowledge. All three now carry
|
||||||
|
// the shared persona block (internal/persona), and all three produce long
|
||||||
|
// free-form Russian — which is exactly where a persona break (formality, third
|
||||||
|
// person, masculine self-reference) is most likely and where, until this file,
|
||||||
|
// nothing could see one.
|
||||||
|
//
|
||||||
|
// Why a second fixture instead of more nudge cases: the checks differ. A nudge
|
||||||
|
// must be one short sentence with no question in it; a chat reply is allowed
|
||||||
|
// 1-3 sentences and a follow-up question is a FEATURE there. Mixing them would
|
||||||
|
// need per-case check masks, and the nudge scorer stays untouched this way.
|
||||||
|
//
|
||||||
|
// Why per-path reporting: a chat regression and a knowledge regression have
|
||||||
|
// different causes (chat prompt vs router.KnowledgePrompt), and one blended
|
||||||
|
// percentage cannot tell them apart.
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
_ "embed"
|
||||||
|
"encoding/json"
|
||||||
|
"fmt"
|
||||||
|
"sort"
|
||||||
|
"strings"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"github.com/kami/maven/internal/dialogue"
|
||||||
|
)
|
||||||
|
|
||||||
|
//go:embed talk_v1.json
|
||||||
|
var talkFixtureJSON []byte
|
||||||
|
|
||||||
|
// The three phrasing paths under test. Values match the fixture's "path" field.
|
||||||
|
const (
|
||||||
|
PathChat = "chat" // PhraseChat
|
||||||
|
PathQuery = "query" // PhraseQuery with notes
|
||||||
|
PathKnowledge = "knowledge" // PhraseQuery with no notes
|
||||||
|
)
|
||||||
|
|
||||||
|
// TalkPaths — report order.
|
||||||
|
var TalkPaths = []string{PathChat, PathQuery, PathKnowledge}
|
||||||
|
|
||||||
|
// TalkCheckNames — the checks that apply to a free-form reply, in report order.
|
||||||
|
// Deliberately a subset of CheckNames: length, mood and "no questions" are nudge
|
||||||
|
// properties and would fail a correct chat reply. These paths return no mood at
|
||||||
|
// all, so there is nothing to check there.
|
||||||
|
var TalkCheckNames = []string{
|
||||||
|
CheckNonEmpty, CheckEllipsis, CheckLang, CheckFeminine, CheckAddress, CheckOnTopic,
|
||||||
|
}
|
||||||
|
|
||||||
|
// TalkCase — one turn as the daemon would present it.
|
||||||
|
//
|
||||||
|
// History is flat text because that is all PhraseChat uses (it concatenates
|
||||||
|
// turn texts into one user message); intents and slots would be dead fields.
|
||||||
|
// Notes are what the store would have matched for a query.
|
||||||
|
//
|
||||||
|
// WantAny is the on-topic contract: at least one lowercased fragment must appear
|
||||||
|
// in the reply. Fragments are stems ("пароль" → "парол") so declension does not
|
||||||
|
// defeat them.
|
||||||
|
type TalkCase struct {
|
||||||
|
ID string `json:"id"`
|
||||||
|
Path string `json:"path"`
|
||||||
|
Utterance string `json:"utterance"`
|
||||||
|
History []string `json:"history,omitempty"`
|
||||||
|
Notes []string `json:"notes,omitempty"`
|
||||||
|
WantAny []string `json:"want_any"`
|
||||||
|
Tags []string `json:"tags,omitempty"`
|
||||||
|
Note string `json:"note,omitempty"`
|
||||||
|
}
|
||||||
|
|
||||||
|
// TalkFixture — the versioned envelope, same gating as Fixture.
|
||||||
|
type TalkFixture struct {
|
||||||
|
SchemaVersion int `json:"schema_version"`
|
||||||
|
Name string `json:"name"`
|
||||||
|
Notes []string `json:"notes"`
|
||||||
|
Cases []TalkCase `json:"cases"`
|
||||||
|
}
|
||||||
|
|
||||||
|
// LoadTalk returns the embedded conversational fixture.
|
||||||
|
func LoadTalk() (TalkFixture, error) {
|
||||||
|
var f TalkFixture
|
||||||
|
if err := json.Unmarshal(talkFixtureJSON, &f); err != nil {
|
||||||
|
return TalkFixture{}, fmt.Errorf("parse talk fixture: %w", err)
|
||||||
|
}
|
||||||
|
if f.SchemaVersion != SchemaVersion {
|
||||||
|
return TalkFixture{}, fmt.Errorf("talk fixture schema_version %d, want %d", f.SchemaVersion, SchemaVersion)
|
||||||
|
}
|
||||||
|
if len(f.Cases) == 0 {
|
||||||
|
return TalkFixture{}, fmt.Errorf("talk fixture has no cases")
|
||||||
|
}
|
||||||
|
return f, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// Talker — the two methods a conversational path must have to be scorable.
|
||||||
|
// *phraser.LLMPhraser satisfies it; same trick as Nudger.
|
||||||
|
type Talker interface {
|
||||||
|
PhraseChat(ctx context.Context, utterance string, history []dialogue.Turn) (string, error)
|
||||||
|
PhraseQuery(ctx context.Context, utterance string, notes []string) (string, error)
|
||||||
|
}
|
||||||
|
|
||||||
|
// TalkOutcome — one scored case.
|
||||||
|
type TalkOutcome struct {
|
||||||
|
Case TalkCase
|
||||||
|
Reply string
|
||||||
|
Err error
|
||||||
|
Latency time.Duration
|
||||||
|
Pass bool
|
||||||
|
Failed []string
|
||||||
|
Reasons []string
|
||||||
|
}
|
||||||
|
|
||||||
|
// TalkReport — the aggregate. ByPath is the point of this scorer.
|
||||||
|
type TalkReport struct {
|
||||||
|
Name string
|
||||||
|
Total int
|
||||||
|
Passed int
|
||||||
|
Errors int
|
||||||
|
ByCheck map[string]int
|
||||||
|
ByPath map[string]TagStat
|
||||||
|
Outcomes []TalkOutcome
|
||||||
|
P50 time.Duration
|
||||||
|
P95 time.Duration
|
||||||
|
Max time.Duration
|
||||||
|
}
|
||||||
|
|
||||||
|
// Accuracy — fraction of cases that passed every check.
|
||||||
|
func (r TalkReport) Accuracy() float64 {
|
||||||
|
if r.Total == 0 {
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
return float64(r.Passed) / float64(r.Total)
|
||||||
|
}
|
||||||
|
|
||||||
|
// ScoreTalk runs every case through t and aggregates. A phrasing error scores as
|
||||||
|
// a miss and is counted separately: "the model was down" and "the model wrote
|
||||||
|
// something bad" must not be the same number.
|
||||||
|
func ScoreTalk(ctx context.Context, name string, t Talker, f TalkFixture) (TalkReport, error) {
|
||||||
|
rep := TalkReport{
|
||||||
|
Name: name,
|
||||||
|
Total: len(f.Cases),
|
||||||
|
ByCheck: map[string]int{},
|
||||||
|
ByPath: map[string]TagStat{},
|
||||||
|
}
|
||||||
|
for _, n := range TalkCheckNames {
|
||||||
|
rep.ByCheck[n] = 0
|
||||||
|
}
|
||||||
|
lat := make([]time.Duration, 0, len(f.Cases))
|
||||||
|
|
||||||
|
for _, c := range f.Cases {
|
||||||
|
start := time.Now()
|
||||||
|
reply, err := c.run(ctx, t)
|
||||||
|
o := TalkOutcome{Case: c, Reply: reply, Err: err, Latency: time.Since(start)}
|
||||||
|
lat = append(lat, o.Latency)
|
||||||
|
|
||||||
|
if err != nil {
|
||||||
|
rep.Errors++
|
||||||
|
o.Failed = append(o.Failed, "call")
|
||||||
|
o.Reasons = append(o.Reasons, fmt.Sprintf("phrase error: %v", err))
|
||||||
|
} else {
|
||||||
|
for _, res := range RunTalkChecks(c, reply) {
|
||||||
|
if res.Pass {
|
||||||
|
rep.ByCheck[res.Name]++
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
o.Failed = append(o.Failed, res.Name)
|
||||||
|
o.Reasons = append(o.Reasons, res.Name+": "+res.Detail)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
o.Pass = len(o.Failed) == 0
|
||||||
|
if o.Pass {
|
||||||
|
rep.Passed++
|
||||||
|
}
|
||||||
|
bump(rep.ByPath, c.Path, o.Pass)
|
||||||
|
rep.Outcomes = append(rep.Outcomes, o)
|
||||||
|
}
|
||||||
|
|
||||||
|
sort.Slice(lat, func(i, j int) bool { return lat[i] < lat[j] })
|
||||||
|
rep.P50, rep.P95 = percentile(lat, 0.50), percentile(lat, 0.95)
|
||||||
|
if len(lat) > 0 {
|
||||||
|
rep.Max = lat[len(lat)-1]
|
||||||
|
}
|
||||||
|
return rep, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// run dispatches the case to its path. knowledge and query are the same method;
|
||||||
|
// the empty notes slice is what selects the no-notes branch inside PhraseQuery.
|
||||||
|
func (c TalkCase) run(ctx context.Context, t Talker) (string, error) {
|
||||||
|
switch c.Path {
|
||||||
|
case PathChat:
|
||||||
|
return t.PhraseChat(ctx, c.Utterance, c.turns())
|
||||||
|
case PathQuery:
|
||||||
|
return t.PhraseQuery(ctx, c.Utterance, c.Notes)
|
||||||
|
case PathKnowledge:
|
||||||
|
return t.PhraseQuery(ctx, c.Utterance, nil)
|
||||||
|
}
|
||||||
|
return "", fmt.Errorf("unknown path %q", c.Path)
|
||||||
|
}
|
||||||
|
|
||||||
|
func (c TalkCase) turns() []dialogue.Turn {
|
||||||
|
turns := make([]dialogue.Turn, 0, len(c.History))
|
||||||
|
for _, h := range c.History {
|
||||||
|
turns = append(turns, dialogue.Turn{Text: h})
|
||||||
|
}
|
||||||
|
return turns
|
||||||
|
}
|
||||||
|
|
||||||
|
// RunTalkChecks scores one reply. Order matches TalkCheckNames.
|
||||||
|
func RunTalkChecks(c TalkCase, reply string) []Result {
|
||||||
|
return []Result{
|
||||||
|
checkNonEmpty(reply),
|
||||||
|
checkEllipsis(reply),
|
||||||
|
checkLang(reply),
|
||||||
|
checkFeminine(reply),
|
||||||
|
checkAddress(reply),
|
||||||
|
checkOnTopicAny(c.WantAny, reply),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// String renders the comparison table — composite, then per-check so a
|
||||||
|
// regression names the property, then per-path so it names the prompt.
|
||||||
|
func (r TalkReport) String() string {
|
||||||
|
var b strings.Builder
|
||||||
|
fmt.Fprintf(&b, "%s: %d/%d cases pass every check (%.1f%%), %d errors\n",
|
||||||
|
r.Name, r.Passed, r.Total, 100*r.Accuracy(), r.Errors)
|
||||||
|
for _, name := range TalkCheckNames {
|
||||||
|
fmt.Fprintf(&b, " %-10s %d/%d\n", name, r.ByCheck[name], r.Total)
|
||||||
|
}
|
||||||
|
fmt.Fprintf(&b, " latency: p50 %s p95 %s max %s\n", r.P50, r.P95, r.Max)
|
||||||
|
fmt.Fprintf(&b, " by path: %s\n", renderStats(r.ByPath))
|
||||||
|
return b.String()
|
||||||
|
}
|
||||||
|
|
||||||
|
// Failures — per-case detail, sorted by ID so two runs diff cleanly.
|
||||||
|
func (r TalkReport) Failures() string {
|
||||||
|
var b strings.Builder
|
||||||
|
for _, o := range r.sorted() {
|
||||||
|
if o.Pass {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
fmt.Fprintf(&b, " %s %q\n %s\n", o.Case.ID, o.Reply, strings.Join(o.Reasons, "; "))
|
||||||
|
}
|
||||||
|
return b.String()
|
||||||
|
}
|
||||||
|
|
||||||
|
// Replies — every generated reply verbatim. This is what a human reads to judge
|
||||||
|
// tone; the score only says which checks fired.
|
||||||
|
func (r TalkReport) Replies() string {
|
||||||
|
var b strings.Builder
|
||||||
|
for _, o := range r.sorted() {
|
||||||
|
mark := "ok "
|
||||||
|
if !o.Pass {
|
||||||
|
mark = "FAIL"
|
||||||
|
}
|
||||||
|
fmt.Fprintf(&b, " %s %-9s %-22s %q\n", mark, o.Case.Path, o.Case.ID, o.Reply)
|
||||||
|
}
|
||||||
|
return b.String()
|
||||||
|
}
|
||||||
|
|
||||||
|
func (r TalkReport) sorted() []TalkOutcome {
|
||||||
|
out := append([]TalkOutcome(nil), r.Outcomes...)
|
||||||
|
sort.Slice(out, func(i, j int) bool { return out[i].Case.ID < out[j].Case.ID })
|
||||||
|
return out
|
||||||
|
}
|
||||||
@@ -0,0 +1,163 @@
|
|||||||
|
package eval
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"os"
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"github.com/kami/maven/internal/dialogue"
|
||||||
|
"github.com/kami/maven/internal/llm"
|
||||||
|
"github.com/kami/maven/internal/persona"
|
||||||
|
"github.com/kami/maven/internal/phraser"
|
||||||
|
)
|
||||||
|
|
||||||
|
// perPathMinimum — the resolution floor. A per-path score built on a handful of
|
||||||
|
// cases moves by 12% when a single reply changes, which cannot distinguish a
|
||||||
|
// prompt regression from noise.
|
||||||
|
const perPathMinimum = 8
|
||||||
|
|
||||||
|
// TestTalkFixture — the fixture itself has to be sound before any score off it
|
||||||
|
// means anything.
|
||||||
|
func TestTalkFixture(t *testing.T) {
|
||||||
|
f, err := LoadTalk()
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("LoadTalk: %v", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
seen := map[string]bool{}
|
||||||
|
byPath := map[string]int{}
|
||||||
|
for _, c := range f.Cases {
|
||||||
|
if seen[c.ID] {
|
||||||
|
t.Errorf("duplicate case id %q", c.ID)
|
||||||
|
}
|
||||||
|
seen[c.ID] = true
|
||||||
|
|
||||||
|
switch c.Path {
|
||||||
|
case PathChat, PathQuery, PathKnowledge:
|
||||||
|
default:
|
||||||
|
t.Errorf("%s: unknown path %q", c.ID, c.Path)
|
||||||
|
}
|
||||||
|
byPath[c.Path]++
|
||||||
|
|
||||||
|
if strings.TrimSpace(c.Utterance) == "" {
|
||||||
|
t.Errorf("%s: empty utterance", c.ID)
|
||||||
|
}
|
||||||
|
if len(c.WantAny) == 0 {
|
||||||
|
t.Errorf("%s: no want_any — the reply cannot be checked for topic", c.ID)
|
||||||
|
}
|
||||||
|
// A query case with no notes would silently score the knowledge path.
|
||||||
|
if c.Path == PathQuery && len(c.Notes) == 0 {
|
||||||
|
t.Errorf("%s: query case has no notes", c.ID)
|
||||||
|
}
|
||||||
|
if c.Path == PathKnowledge && len(c.Notes) > 0 {
|
||||||
|
t.Errorf("%s: knowledge case must have no notes", c.ID)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
for _, p := range TalkPaths {
|
||||||
|
if byPath[p] < perPathMinimum {
|
||||||
|
t.Errorf("path %s has %d cases, want at least %d", p, byPath[p], perPathMinimum)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// fakeTalker — a scripted Talker, so the scorer is testable without a model.
|
||||||
|
type fakeTalker struct{ reply string }
|
||||||
|
|
||||||
|
func (f fakeTalker) PhraseChat(context.Context, string, []dialogue.Turn) (string, error) {
|
||||||
|
return f.reply, nil
|
||||||
|
}
|
||||||
|
func (f fakeTalker) PhraseQuery(context.Context, string, []string) (string, error) {
|
||||||
|
return f.reply, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestScoreTalkCounts — a reply that fails on purpose must be counted on every
|
||||||
|
// path, so a real run cannot report a hidden zero.
|
||||||
|
func TestScoreTalkCounts(t *testing.T) {
|
||||||
|
f, err := LoadTalk()
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("LoadTalk: %v", err)
|
||||||
|
}
|
||||||
|
// Formal address, off-topic, trailing ellipsis: three checks fail at once.
|
||||||
|
rep, err := ScoreTalk(context.Background(), "fake", fakeTalker{"Приходите, я вас жду…"}, f)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("ScoreTalk: %v", err)
|
||||||
|
}
|
||||||
|
if rep.Total != len(f.Cases) || rep.Passed != 0 {
|
||||||
|
t.Errorf("got %d/%d passing, want 0/%d", rep.Passed, rep.Total, len(f.Cases))
|
||||||
|
}
|
||||||
|
if rep.ByCheck[CheckAddress] != 0 {
|
||||||
|
t.Errorf("formal reply passed the address check %d times", rep.ByCheck[CheckAddress])
|
||||||
|
}
|
||||||
|
if rep.ByCheck[CheckEllipsis] != 0 {
|
||||||
|
t.Errorf("truncated reply passed the ellipsis check %d times", rep.ByCheck[CheckEllipsis])
|
||||||
|
}
|
||||||
|
for _, p := range TalkPaths {
|
||||||
|
if rep.ByPath[p].Total == 0 {
|
||||||
|
t.Errorf("path %s missing from the report", p)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if !strings.Contains(rep.String(), "by path") {
|
||||||
|
t.Error("report does not break down by path")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestLLMTalkBaseline — the resident model on the three conversational paths.
|
||||||
|
// Opt-in exactly like TestLLMPhrasingBaseline: CI has no model and a run costs
|
||||||
|
// minutes on the CPU target.
|
||||||
|
//
|
||||||
|
// MAVEN_LLM_URL=http://127.0.0.1:18099 \
|
||||||
|
// go test -run TestLLMTalkBaseline ./internal/phraser/eval/
|
||||||
|
//
|
||||||
|
// Reports, does not assert a quality bar — the numbers are the input to tuning
|
||||||
|
// the persona prompt. The one thing worth failing on is a harness fault.
|
||||||
|
func TestLLMTalkBaseline(t *testing.T) {
|
||||||
|
base := os.Getenv("MAVEN_LLM_URL")
|
||||||
|
if base == "" {
|
||||||
|
t.Skip("MAVEN_LLM_URL unset — point it at a running llama-server (see doc comment)")
|
||||||
|
}
|
||||||
|
noProxyLoopback(t)
|
||||||
|
|
||||||
|
ctx := context.Background()
|
||||||
|
f, err := LoadTalk()
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("LoadTalk: %v", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
cfg := phraser.DefaultConfig("")
|
||||||
|
cfg.Timeout = 5 * time.Minute
|
||||||
|
cfg.ContextBlock = func() string { return persona.Facts{}.Block(time.Now()) }
|
||||||
|
p := phraser.NewLLMPhraserAt(base, cfg)
|
||||||
|
defer p.Close()
|
||||||
|
|
||||||
|
// Unreachable server is fatal here, not a logged warning, and that differs
|
||||||
|
// from the nudge test on purpose. PhraseNudge returns its errors, so a dead
|
||||||
|
// server there shows up honestly in the Errors column. PhraseChat and
|
||||||
|
// PhraseQuery do NOT: they swallow every failure and return a canned string
|
||||||
|
// ("поговорили.", "не знаю.", "вот что я нашла: …"). So on these three paths
|
||||||
|
// a dead server produces a full report with 0 errors and a terrible score —
|
||||||
|
// a number that looks like bad phrasing and is really no phrasing at all.
|
||||||
|
// Refusing to score without a confirmed model is the only guard available
|
||||||
|
// until the phraser reports its failures (Vikunja #397).
|
||||||
|
model, err := llm.ModelID(ctx, base)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("no model at %s: %v — refusing to score, these paths hide their errors "+
|
||||||
|
"and would report a plausible-looking result off a dead server", base, err)
|
||||||
|
}
|
||||||
|
t.Logf("scoring model %s at %s", model, base)
|
||||||
|
|
||||||
|
rep, err := ScoreTalk(ctx, "llm ("+model+", built-in persona)", p, f)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("ScoreTalk: %v", err)
|
||||||
|
}
|
||||||
|
t.Log("\n" + rep.String() + "\nreplies:\n" + rep.Replies() + "\nfailures:\n" + rep.Failures())
|
||||||
|
|
||||||
|
// And again afterwards: the run takes minutes, and a server that died or got
|
||||||
|
// OOM-killed halfway through would leave the first cases scored and the rest
|
||||||
|
// silently canned. Checking only at the start would not catch that.
|
||||||
|
if _, err := llm.ModelID(ctx, base); err != nil {
|
||||||
|
t.Fatalf("model at %s went away during the run: %v — the score above is not trustworthy", base, err)
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,227 @@
|
|||||||
|
{
|
||||||
|
"schema_version": 1,
|
||||||
|
"name": "ru-talk-v1",
|
||||||
|
"notes": [
|
||||||
|
"Scores the three conversational phrasing paths: chat (PhraseChat), query (PhraseQuery with notes) and knowledge (PhraseQuery with no notes). The nudge fixture does not cover any of them.",
|
||||||
|
"Nine cases per path, not five. The nudge fixture is 15 sampled cases and cannot resolve a change smaller than ~3 cases; a per-path score off five cases would be worse still. More cases per path is the point of this fixture.",
|
||||||
|
"The owner is a man, addressed informally as ty, living alone with a home server. Every utterance is written the way he actually talks to her.",
|
||||||
|
"chat-formality-bait and chat-about-me exist to provoke the two persona breaks the nudge eval caught: the formal vy/vas plural, and talking about him in the third person.",
|
||||||
|
"want_any fragments are stems so Russian declension does not defeat the on-topic check. They are lowercased before comparison.",
|
||||||
|
"want_any is a plain substring test, so a fragment that is too short passes by accident: \"ты\" matches inside \"работы\", \"нет\" inside \"интернет\". Keep every fragment to three or more letters of a real stem.",
|
||||||
|
"Notes are written as the store would have them: short, first person, no punctuation discipline."
|
||||||
|
],
|
||||||
|
"cases": [
|
||||||
|
{
|
||||||
|
"id": "chat-how-are-you",
|
||||||
|
"path": "chat",
|
||||||
|
"utterance": "привет, как дела?",
|
||||||
|
"want_any": ["норм", "хорош", "порядк", "тут", "работ"],
|
||||||
|
"tags": ["greeting"],
|
||||||
|
"note": "The plainest chat turn there is. If the persona breaks anywhere it breaks here first."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "chat-formality-bait",
|
||||||
|
"path": "chat",
|
||||||
|
"utterance": "не могли бы вы подсказать, чем вы сейчас занимаетесь?",
|
||||||
|
"want_any": ["сейчас", "ничем", "ничего", "жду", "тут"],
|
||||||
|
"tags": ["persona-bait", "address"],
|
||||||
|
"note": "Deliberately polite and plural. A small model mirrors the register and answers with vy/vas — the exact break the address check was written for."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "chat-about-me",
|
||||||
|
"path": "chat",
|
||||||
|
"utterance": "расскажи обо мне",
|
||||||
|
"want_any": ["теб"],
|
||||||
|
"tags": ["persona-bait", "third-person"],
|
||||||
|
"note": "Baits the third person: she should say 'ты живёшь один', not 'он живёт один', as if reporting to somebody else."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "chat-bored-evening",
|
||||||
|
"path": "chat",
|
||||||
|
"utterance": "скучно что-то вечером, посоветуй чем заняться",
|
||||||
|
"want_any": ["можеш", "попробу", "почита", "прогул", "фильм", "серв"],
|
||||||
|
"tags": ["open-ended"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "chat-followup-server",
|
||||||
|
"path": "chat",
|
||||||
|
"utterance": "а стоит его вообще перезагружать?",
|
||||||
|
"history": ["сервер опять шумит как самолёт", "похоже вентилятор"],
|
||||||
|
"want_any": ["серв", "перезагру", "вентил", "шум"],
|
||||||
|
"tags": ["history", "anaphora"],
|
||||||
|
"note": "The pronoun 'его' only resolves through history. Also the one case where 'он' about the server is legitimate."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "chat-tired",
|
||||||
|
"path": "chat",
|
||||||
|
"utterance": "устал я сегодня, весь день за компом",
|
||||||
|
"want_any": ["отдохн", "устал", "перерыв", "спат", "день"],
|
||||||
|
"tags": ["tone"],
|
||||||
|
"note": "Invites the fake-concern and emotional-support drift; the reply should stay plain."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "chat-thanks",
|
||||||
|
"path": "chat",
|
||||||
|
"utterance": "спасибо, выручила",
|
||||||
|
"want_any": ["пожалуйст", "не за что", "рада", "обращ"],
|
||||||
|
"tags": ["persona", "feminine"],
|
||||||
|
"note": "Feminine self-reference is unavoidable in an answer to thanks: 'рада', not 'рад'."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "chat-what-can-you-do",
|
||||||
|
"path": "chat",
|
||||||
|
"utterance": "что ты вообще умеешь?",
|
||||||
|
"want_any": ["напомн", "замет", "запис", "могу", "умею"],
|
||||||
|
"tags": ["self-description", "feminine"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "chat-joke",
|
||||||
|
"path": "chat",
|
||||||
|
"utterance": "расскажи что-нибудь смешное",
|
||||||
|
"want_any": ["анекдот", "шутк", "смешн", "истори"],
|
||||||
|
"tags": ["open-ended"],
|
||||||
|
"note": "Longest free-form generation in the chat set — the most likely place for a truncated reply."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "query-router-password",
|
||||||
|
"path": "query",
|
||||||
|
"utterance": "что я записывал про пароль от роутера?",
|
||||||
|
"notes": ["пароль от роутера admin/xxK9tp — на наклейке снизу", "роутер висит в коридоре"],
|
||||||
|
"want_any": ["парол", "роутер", "наклейк"],
|
||||||
|
"tags": ["notes", "recall"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "query-bedtime-yesterday",
|
||||||
|
"path": "query",
|
||||||
|
"utterance": "напомни, во сколько я вчера лёг?",
|
||||||
|
"notes": ["лёг спать в 02:40", "сегодня встал в 9"],
|
||||||
|
"want_any": ["02:40", "2:40", "полтрет", "ноч"],
|
||||||
|
"tags": ["notes", "time"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "query-doctor-name",
|
||||||
|
"path": "query",
|
||||||
|
"utterance": "как звали того стоматолога, которого мне советовали?",
|
||||||
|
"notes": ["стоматолог Игорь Валерьевич, клиника на Ленина, советовал Дима"],
|
||||||
|
"want_any": ["игор", "валерьев", "стоматолог"],
|
||||||
|
"tags": ["notes", "recall"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "query-disk-plan",
|
||||||
|
"path": "query",
|
||||||
|
"utterance": "я что-то планировал с диском на сервере, что именно?",
|
||||||
|
"notes": ["купить второй hdd на 4тб под бэкапы", "перенести медиатеку с системного диска"],
|
||||||
|
"want_any": ["hdd", "бэкап", "диск", "4тб", "медиатек"],
|
||||||
|
"tags": ["notes", "homeserver"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "query-notes-do-not-answer",
|
||||||
|
"path": "query",
|
||||||
|
"utterance": "сколько я заплатил за домен?",
|
||||||
|
"notes": ["домен продлевается в марте", "хостинг оплачен на год вперёд"],
|
||||||
|
"want_any": ["домен", "не зна", "не указ"],
|
||||||
|
"tags": ["notes", "negative"],
|
||||||
|
"note": "The notes do not contain the price. The prompt tells her to say so; a made-up number is the failure being watched for."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "query-single-note",
|
||||||
|
"path": "query",
|
||||||
|
"utterance": "где лежит запасной ключ?",
|
||||||
|
"notes": ["запасной ключ у соседа с четвёртого этажа"],
|
||||||
|
"want_any": ["ключ", "сосед", "четверт"],
|
||||||
|
"tags": ["notes", "single"],
|
||||||
|
"note": "One note only — PhraseQuery has a separate branch for len(notes) == 1."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "query-polite-form",
|
||||||
|
"path": "query",
|
||||||
|
"utterance": "подскажите, пожалуйста, что у меня записано по машине?",
|
||||||
|
"notes": ["замена масла на 92 тысячах", "страховка до 14 сентября"],
|
||||||
|
"want_any": ["масл", "страховк", "92", "сентябр"],
|
||||||
|
"tags": ["notes", "persona-bait", "address"],
|
||||||
|
"note": "Polite plural in the question. The answer must still be ty."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "query-shopping",
|
||||||
|
"path": "query",
|
||||||
|
"utterance": "что мне надо было купить?",
|
||||||
|
"notes": ["купить кофе и фильтры", "закончилась паста"],
|
||||||
|
"want_any": ["кофе", "фильтр", "паст"],
|
||||||
|
"tags": ["notes", "list"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "query-wifi-guest",
|
||||||
|
"path": "query",
|
||||||
|
"utterance": "я записывал гостевой вайфай?",
|
||||||
|
"notes": ["гостевая сеть maven-guest, пароль 12345678 меняю раз в месяц"],
|
||||||
|
"want_any": ["guest", "гостев", "12345678", "парол"],
|
||||||
|
"tags": ["notes", "recall"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "know-sky-blue",
|
||||||
|
"path": "knowledge",
|
||||||
|
"utterance": "почему небо синее?",
|
||||||
|
"want_any": ["све", "рассеи", "атмосфер", "син", "волн"],
|
||||||
|
"tags": ["general"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "know-boil-egg",
|
||||||
|
"path": "knowledge",
|
||||||
|
"utterance": "сколько варить яйцо вкрутую?",
|
||||||
|
"want_any": ["минут", "8", "9", "10", "варит"],
|
||||||
|
"tags": ["general", "practical"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "know-ssd-vs-hdd",
|
||||||
|
"path": "knowledge",
|
||||||
|
"utterance": "чем ssd отличается от hdd?",
|
||||||
|
"want_any": ["ssd", "hdd", "быстр", "диск", "механич"],
|
||||||
|
"tags": ["general", "tech"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "know-cat-purr",
|
||||||
|
"path": "knowledge",
|
||||||
|
"utterance": "почему кошки мурчат?",
|
||||||
|
"want_any": ["кош", "мурч", "вибра", "успока"],
|
||||||
|
"tags": ["general"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "know-hiccups",
|
||||||
|
"path": "knowledge",
|
||||||
|
"utterance": "как быстро избавиться от икоты?",
|
||||||
|
"want_any": ["икот", "дыха", "вод", "задерж"],
|
||||||
|
"tags": ["general", "practical"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "know-polite-form",
|
||||||
|
"path": "knowledge",
|
||||||
|
"utterance": "не могли бы вы объяснить, что такое vpn?",
|
||||||
|
"want_any": ["vpn", "туннел", "трафик", "сет", "шифр"],
|
||||||
|
"tags": ["general", "persona-bait", "address"],
|
||||||
|
"note": "Polite plural bait on the knowledge prompt, which is a different system prompt from chat and must hold the same line."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "know-dont-know",
|
||||||
|
"path": "knowledge",
|
||||||
|
"utterance": "как зовут моего соседа снизу?",
|
||||||
|
"want_any": ["не зна", "не мог"],
|
||||||
|
"tags": ["general", "negative"],
|
||||||
|
"note": "Unanswerable without notes. Admitting it beats inventing a name; watching for the invention."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "know-water-per-day",
|
||||||
|
"path": "knowledge",
|
||||||
|
"utterance": "сколько воды в день надо пить?",
|
||||||
|
"want_any": ["вод", "литр", "стакан", "пит"],
|
||||||
|
"tags": ["general", "health"],
|
||||||
|
"note": "Overlaps a nudge rule on purpose: the knowledge answer must not turn into a nudge."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "know-thunder-delay",
|
||||||
|
"path": "knowledge",
|
||||||
|
"utterance": "почему гром слышно позже молнии?",
|
||||||
|
"want_any": ["звук", "све", "быстр", "гром", "молни"],
|
||||||
|
"tags": ["general"]
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
@@ -0,0 +1,58 @@
|
|||||||
|
package eval
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"math/rand"
|
||||||
|
"testing"
|
||||||
|
|
||||||
|
"github.com/kami/maven/internal/phraser"
|
||||||
|
)
|
||||||
|
|
||||||
|
// TestTemplateNudges scores the hand-written Russian templates on the same
|
||||||
|
// fixture the model is scored on. No model, no network — it runs in milliseconds.
|
||||||
|
//
|
||||||
|
// The bar is every case, not most of them: the templates are hand-written, so a
|
||||||
|
// failure is a bug in one line of Russian, not model variance.
|
||||||
|
func TestTemplateNudges(t *testing.T) {
|
||||||
|
f, err := Load()
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("Load: %v", err)
|
||||||
|
}
|
||||||
|
// Fixed seed: the score must not depend on which variant came up.
|
||||||
|
nt, err := phraser.NewNudgeTemplates(rand.NewSource(20260731))
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("NewNudgeTemplates: %v", err)
|
||||||
|
}
|
||||||
|
rep, err := Score(context.Background(), "ru templates", nt, f)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("Score: %v", err)
|
||||||
|
}
|
||||||
|
t.Log("\n" + rep.String())
|
||||||
|
t.Log("\n" + rep.Messages())
|
||||||
|
if rep.Passed != rep.Total {
|
||||||
|
t.Errorf("templates scored %d/%d, want every case:\n%s",
|
||||||
|
rep.Passed, rep.Total, rep.Failures())
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestTemplateNudgesEverySeed — one seed passing could be luck. Every variant of
|
||||||
|
// every rule has to pass every check, so sweep seeds until each has been used.
|
||||||
|
func TestTemplateNudgesEverySeed(t *testing.T) {
|
||||||
|
f, err := Load()
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("Load: %v", err)
|
||||||
|
}
|
||||||
|
for seed := int64(0); seed < 60; seed++ {
|
||||||
|
nt, err := phraser.NewNudgeTemplates(rand.NewSource(seed))
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("NewNudgeTemplates: %v", err)
|
||||||
|
}
|
||||||
|
rep, err := Score(context.Background(), "ru templates", nt, f)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("Score: %v", err)
|
||||||
|
}
|
||||||
|
if rep.Passed != rep.Total {
|
||||||
|
t.Errorf("seed %d: %d/%d\n%s", seed, rep.Passed, rep.Total, rep.Failures())
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,129 @@
|
|||||||
|
package phraser
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"encoding/json"
|
||||||
|
"net/http"
|
||||||
|
"net/http/httptest"
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
|
||||||
|
"github.com/kami/maven/internal/loop"
|
||||||
|
)
|
||||||
|
|
||||||
|
// grammarSpy stands in for llama-server: it records the grammar field of every
|
||||||
|
// request and always answers with a contract-shaped reply.
|
||||||
|
type grammarSpy struct {
|
||||||
|
srv *httptest.Server
|
||||||
|
grammars []string
|
||||||
|
}
|
||||||
|
|
||||||
|
func newGrammarSpy(t *testing.T) *grammarSpy {
|
||||||
|
t.Helper()
|
||||||
|
s := &grammarSpy{}
|
||||||
|
s.srv = httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||||
|
var req chatReq
|
||||||
|
if err := json.NewDecoder(r.Body).Decode(&req); err != nil {
|
||||||
|
t.Errorf("spy: decode request: %v", err)
|
||||||
|
}
|
||||||
|
s.grammars = append(s.grammars, req.Grammar)
|
||||||
|
w.Header().Set("Content-Type", "application/json")
|
||||||
|
w.Write([]byte(`{"choices":[{"message":{"content":"{\"response\": \"ага\", \"mood\": \"neutral\"}"}}]}`))
|
||||||
|
}))
|
||||||
|
t.Cleanup(s.srv.Close)
|
||||||
|
return s
|
||||||
|
}
|
||||||
|
|
||||||
|
// callAllPhrasingPaths hits every path that expects the JSON contract.
|
||||||
|
// LLMNudges must be set on the phraser under test: nudges come from templates
|
||||||
|
// by default and never reach the model at all.
|
||||||
|
func callAllPhrasingPaths(t *testing.T, p *LLMPhraser) {
|
||||||
|
t.Helper()
|
||||||
|
ctx := context.Background()
|
||||||
|
if _, err := p.PhraseNudge(ctx, loop.Candidate{Rule: loop.WaterRule(), Severity: loop.Sev1}); err != nil {
|
||||||
|
t.Fatalf("PhraseNudge: %v", err)
|
||||||
|
}
|
||||||
|
if _, err := p.PhraseChat(ctx, "привет", nil); err != nil {
|
||||||
|
t.Fatalf("PhraseChat: %v", err)
|
||||||
|
}
|
||||||
|
// Both branches: no notes (general knowledge) and with notes (grounded).
|
||||||
|
if _, err := p.PhraseQuery(ctx, "сколько воды я выпил", nil); err != nil {
|
||||||
|
t.Fatalf("PhraseQuery (no notes): %v", err)
|
||||||
|
}
|
||||||
|
if _, err := p.PhraseQuery(ctx, "сколько воды я выпил", []string{"два литра"}); err != nil {
|
||||||
|
t.Fatalf("PhraseQuery (notes): %v", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestGrammarIsAttachedToEveryPhrasingRequest(t *testing.T) {
|
||||||
|
if strings.TrimSpace(responseGrammar) == "" {
|
||||||
|
t.Fatal("responseGrammar is empty")
|
||||||
|
}
|
||||||
|
spy := newGrammarSpy(t)
|
||||||
|
p := NewLLMPhraserAt(spy.srv.URL, Config{LLMNudges: true})
|
||||||
|
|
||||||
|
callAllPhrasingPaths(t, p)
|
||||||
|
|
||||||
|
if len(spy.grammars) != 4 {
|
||||||
|
t.Fatalf("expected 4 requests, got %d", len(spy.grammars))
|
||||||
|
}
|
||||||
|
for i, g := range spy.grammars {
|
||||||
|
if g != responseGrammar {
|
||||||
|
t.Errorf("request %d carries grammar %q, want responseGrammar", i, g)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestNoGrammarConfigDisablesIt(t *testing.T) {
|
||||||
|
spy := newGrammarSpy(t)
|
||||||
|
p := NewLLMPhraserAt(spy.srv.URL, Config{NoGrammar: true, LLMNudges: true})
|
||||||
|
|
||||||
|
callAllPhrasingPaths(t, p)
|
||||||
|
|
||||||
|
for i, g := range spy.grammars {
|
||||||
|
if g != "" {
|
||||||
|
t.Errorf("request %d still carries a grammar with NoGrammar set: %q", i, g)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The grammar's string rule must accept any codepoint, not just ASCII. Replies
|
||||||
|
// are Russian: an ASCII-only class would constrain the model into empty replies.
|
||||||
|
func TestGrammarStringRuleIsNotASCIIOnly(t *testing.T) {
|
||||||
|
if !strings.Contains(responseGrammar, `([^"\\] | "\\" ["\\/bfnrt])`) {
|
||||||
|
t.Error("string rule is not the any-codepoint-except-quote-and-backslash class; Cyrillic replies would be impossible")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// What the grammar describes must survive the parser that reads it back — a
|
||||||
|
// Russian body with an escaped quote inside, hand-built to test the contract.
|
||||||
|
func TestGrammarShapedJSONParses(t *testing.T) {
|
||||||
|
raw := `{"response": "он сказал \"привет\" и ушёл.\nвот так.", "mood": "confused"}`
|
||||||
|
text, mood, err := parseResponseMood(raw)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("grammar-shaped JSON did not parse: %v", err)
|
||||||
|
}
|
||||||
|
if want := "он сказал \"привет\" и ушёл.\nвот так."; text != want {
|
||||||
|
t.Errorf("response = %q, want %q", text, want)
|
||||||
|
}
|
||||||
|
if mood != "confused" {
|
||||||
|
t.Errorf("mood = %q, want confused", mood)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Every mood the grammar permits is one the contract knows, and all five are there.
|
||||||
|
func TestGrammarMoodEnumMatchesTheContract(t *testing.T) {
|
||||||
|
for _, m := range []string{"neutral", "happy", "thinking", "tired", "confused"} {
|
||||||
|
if !strings.Contains(responseGrammar, `"\"`+m+`\""`) {
|
||||||
|
t.Errorf("mood %q missing from the grammar", m)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
// No sixth mood: the enum line lists exactly five alternatives.
|
||||||
|
for _, line := range strings.Split(responseGrammar, "\n") {
|
||||||
|
if strings.HasPrefix(line, "mood") {
|
||||||
|
if n := strings.Count(line, "|") + 1; n != 5 {
|
||||||
|
t.Errorf("mood rule lists %d alternatives, want 5: %s", n, line)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
+309
-52
@@ -18,6 +18,7 @@ import (
|
|||||||
"github.com/kami/maven/internal/delivery"
|
"github.com/kami/maven/internal/delivery"
|
||||||
"github.com/kami/maven/internal/dialogue"
|
"github.com/kami/maven/internal/dialogue"
|
||||||
"github.com/kami/maven/internal/loop"
|
"github.com/kami/maven/internal/loop"
|
||||||
|
"github.com/kami/maven/internal/persona"
|
||||||
"github.com/kami/maven/internal/router"
|
"github.com/kami/maven/internal/router"
|
||||||
)
|
)
|
||||||
|
|
||||||
@@ -30,6 +31,10 @@ type LLMPhraser struct {
|
|||||||
cmd *exec.Cmd
|
cmd *exec.Cmd
|
||||||
cancel context.CancelFunc
|
cancel context.CancelFunc
|
||||||
wg sync.WaitGroup
|
wg sync.WaitGroup
|
||||||
|
|
||||||
|
// tmpl — the hand-written Russian nudges. Default path for nudges; see
|
||||||
|
// Config.LLMNudges. nil only if the template file failed to load.
|
||||||
|
tmpl *NudgeTemplates
|
||||||
}
|
}
|
||||||
|
|
||||||
type Config struct {
|
type Config struct {
|
||||||
@@ -39,7 +44,31 @@ type Config struct {
|
|||||||
NGpuLayers int
|
NGpuLayers int
|
||||||
NCtx int
|
NCtx int
|
||||||
Timeout time.Duration
|
Timeout time.Duration
|
||||||
Persona string // optional prompt prefix tuning maven's character
|
|
||||||
|
// ContextBlock renders the shared context block (who he is, how to
|
||||||
|
// address him, the time) fresh for each turn. See internal/persona.
|
||||||
|
// nil ⇒ no block, the prompts stand alone.
|
||||||
|
ContextBlock func() string
|
||||||
|
|
||||||
|
// LLMNudges puts the model back in charge of nudge wording.
|
||||||
|
//
|
||||||
|
// Off by default, and that is a deliberate deprecation of LLM-phrased
|
||||||
|
// nudges: hand-written templates (nudges_ru_v1.json) word every nudge now.
|
||||||
|
// A nudge has nothing to be creative about, and measured over many runs the
|
||||||
|
// 0.8B broke the persona (formal "вы", plural imperatives, masculine
|
||||||
|
// self-reference) and invented facts and units. Templates score 15/15 on the
|
||||||
|
// nudge fixture, the model 11-13/15.
|
||||||
|
//
|
||||||
|
// The LLM path is kept, not deleted: flip this on to get it back. Chat,
|
||||||
|
// query and reminder phrasing are untouched and still go through the model.
|
||||||
|
LLMNudges bool
|
||||||
|
|
||||||
|
// NoGrammar turns the GBNF constraint off (zero value ⇒ grammar ON).
|
||||||
|
// The escape hatch exists because the target resident model — the
|
||||||
|
// locally CPT'd Qwen3-1.7B — does not exist yet: if its chat template
|
||||||
|
// ever fights the grammar, the fix should be a config flip on the
|
||||||
|
// deploy box, not a code change and a rebuild.
|
||||||
|
NoGrammar bool
|
||||||
}
|
}
|
||||||
|
|
||||||
func DefaultConfig(modelPath string) Config {
|
func DefaultConfig(modelPath string) Config {
|
||||||
@@ -59,6 +88,7 @@ func NewLLMPhraser(ctx context.Context, cfg Config) (*LLMPhraser, error) {
|
|||||||
cfg: cfg,
|
cfg: cfg,
|
||||||
client: &http.Client{Timeout: cfg.Timeout},
|
client: &http.Client{Timeout: cfg.Timeout},
|
||||||
cancel: cancel,
|
cancel: cancel,
|
||||||
|
tmpl: loadNudgeTemplates(),
|
||||||
}
|
}
|
||||||
if err := p.start(ctx); err != nil {
|
if err := p.start(ctx); err != nil {
|
||||||
cancel()
|
cancel()
|
||||||
@@ -80,9 +110,22 @@ func NewLLMPhraserAt(baseURL string, cfg Config) *LLMPhraser {
|
|||||||
client: &http.Client{Timeout: cfg.Timeout},
|
client: &http.Client{Timeout: cfg.Timeout},
|
||||||
port: strings.TrimSuffix(baseURL, "/"),
|
port: strings.TrimSuffix(baseURL, "/"),
|
||||||
cancel: func() {},
|
cancel: func() {},
|
||||||
|
tmpl: loadNudgeTemplates(),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// loadNudgeTemplates loads the Russian nudge templates. A broken template file
|
||||||
|
// must not stop the daemon booting, so a failure logs and leaves the LLM path
|
||||||
|
// in charge of nudges.
|
||||||
|
func loadNudgeTemplates() *NudgeTemplates {
|
||||||
|
nt, err := NewNudgeTemplates(nil)
|
||||||
|
if err != nil {
|
||||||
|
log.Printf("phraser: nudge templates unavailable, using the model: %v", err)
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
return nt
|
||||||
|
}
|
||||||
|
|
||||||
func (p *LLMPhraser) start(ctx context.Context) error {
|
func (p *LLMPhraser) start(ctx context.Context) error {
|
||||||
args := []string{
|
args := []string{
|
||||||
"-m", p.cfg.ModelPath,
|
"-m", p.cfg.ModelPath,
|
||||||
@@ -173,18 +216,30 @@ func (p *LLMPhraser) Close() error {
|
|||||||
}
|
}
|
||||||
|
|
||||||
func (p *LLMPhraser) PhraseNudge(ctx context.Context, c loop.Candidate) (delivery.PhrasedNudge, error) {
|
func (p *LLMPhraser) PhraseNudge(ctx context.Context, c loop.Candidate) (delivery.PhrasedNudge, error) {
|
||||||
|
// Templates first — see Config.LLMNudges for why this is the default.
|
||||||
|
if !p.cfg.LLMNudges && p.tmpl != nil {
|
||||||
|
return p.tmpl.PhraseNudge(ctx, c)
|
||||||
|
}
|
||||||
prompt := buildNudgePrompt(c)
|
prompt := buildNudgePrompt(c)
|
||||||
resp, err := p.chat(ctx, prompt)
|
resp, err := p.chat(ctx, prompt)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return delivery.PhrasedNudge{}, err
|
return delivery.PhrasedNudge{}, err
|
||||||
}
|
}
|
||||||
body, mood := parseResponseMood(resp)
|
body, mood, perr := parseResponseMood(resp)
|
||||||
|
if perr != nil {
|
||||||
|
// Truncated JSON. Not a nudge — use the plain Russian fallback.
|
||||||
|
log.Printf("phraser: PhraseNudge: %v", perr)
|
||||||
|
body, mood = "", ""
|
||||||
|
}
|
||||||
if body == "" {
|
if body == "" {
|
||||||
// fallback: try old body/summary format
|
// fallback: try old body/summary format
|
||||||
body, _ = parsePhrase(resp)
|
body, _ = parsePhrase(resp)
|
||||||
}
|
}
|
||||||
if body == "" {
|
if body == "" {
|
||||||
body = fmt.Sprintf("%s — %s", c.Rule.Name, sevLabel(c.Severity))
|
// The model said nothing usable. Say it in Russian anyway — this text
|
||||||
|
// goes straight to a Russian piper voice, so the old "water — care"
|
||||||
|
// fallback was unspeakable.
|
||||||
|
body = fallbackNudge(c)
|
||||||
}
|
}
|
||||||
if mood == "" {
|
if mood == "" {
|
||||||
mood = "neutral"
|
mood = "neutral"
|
||||||
@@ -199,13 +254,18 @@ func (p *LLMPhraser) PhraseQuery(ctx context.Context, utterance string, notes []
|
|||||||
if len(notes) == 0 {
|
if len(notes) == 0 {
|
||||||
// General knowledge — no notes to ground the answer. The system
|
// General knowledge — no notes to ground the answer. The system
|
||||||
// prompt is the single tested source in router.KnowledgePrompt.
|
// prompt is the single tested source in router.KnowledgePrompt.
|
||||||
sys := router.KnowledgePrompt()
|
sys := persona.Prepend(p.cfg.ContextBlock, router.KnowledgePrompt())
|
||||||
prompt := fmt.Sprintf("Пользователь спрашивает: \"%s\".", utterance)
|
prompt := fmt.Sprintf("Пользователь спрашивает: \"%s\".", utterance)
|
||||||
resp, err := p.chatWithSystem(ctx, sys, prompt, 256)
|
resp, err := p.chatWithSystem(ctx, sys, prompt, 768)
|
||||||
if err != nil || resp == "" {
|
if err != nil || resp == "" {
|
||||||
return "не знаю.", nil
|
return "не знаю.", nil
|
||||||
}
|
}
|
||||||
if text, _ := parseResponseMood(resp); text != "" {
|
text, _, perr := parseResponseMood(resp)
|
||||||
|
if perr != nil {
|
||||||
|
log.Printf("phraser: PhraseQuery: %v", perr)
|
||||||
|
return "не знаю.", nil
|
||||||
|
}
|
||||||
|
if text != "" {
|
||||||
return text, nil
|
return text, nil
|
||||||
}
|
}
|
||||||
return resp, nil
|
return resp, nil
|
||||||
@@ -215,17 +275,22 @@ func (p *LLMPhraser) PhraseQuery(ctx context.Context, utterance string, notes []
|
|||||||
}
|
}
|
||||||
sys := p.querySystemPrompt()
|
sys := p.querySystemPrompt()
|
||||||
prompt := fmt.Sprintf(
|
prompt := fmt.Sprintf(
|
||||||
`The user asks: "%s". Your notes matching the query contain: "%s". Answer them naturally and briefly. If the notes don't answer the question, say so.`,
|
`Он спрашивает: "%s". В твоих заметках по этому вопросу написано: "%s". Ответь ему коротко и своими словами. Если в заметках ответа нет — так и скажи.`,
|
||||||
utterance, strings.Join(notes, `"; "`),
|
utterance, strings.Join(notes, `"; "`),
|
||||||
)
|
)
|
||||||
resp, err := p.chatWithSystem(ctx, sys, prompt, 256)
|
resp, err := p.chatWithSystem(ctx, sys, prompt, 768)
|
||||||
if err != nil {
|
text, _, perr := parseResponseMood(resp)
|
||||||
|
if err != nil || perr != nil {
|
||||||
|
// Read the notes out rather than ship a broken fragment.
|
||||||
|
if perr != nil {
|
||||||
|
log.Printf("phraser: PhraseQuery: %v", perr)
|
||||||
|
}
|
||||||
if len(notes) == 1 {
|
if len(notes) == 1 {
|
||||||
return "вот что я нашла: " + notes[0], nil
|
return "вот что я нашла: " + notes[0], nil
|
||||||
}
|
}
|
||||||
return "вот что я нашла: " + strings.Join(notes, "; "), nil
|
return "вот что я нашла: " + strings.Join(notes, "; "), nil
|
||||||
}
|
}
|
||||||
if text, _ := parseResponseMood(resp); text != "" {
|
if text != "" {
|
||||||
return text, nil
|
return text, nil
|
||||||
}
|
}
|
||||||
return resp, nil
|
return resp, nil
|
||||||
@@ -235,7 +300,7 @@ func (p *LLMPhraser) PhraseQuery(ctx context.Context, utterance string, notes []
|
|||||||
// message array from dialogue history + the current user utterance. Falls back
|
// message array from dialogue history + the current user utterance. Falls back
|
||||||
// to a simple greeting on any LLM error — better to say something than nothing.
|
// to a simple greeting on any LLM error — better to say something than nothing.
|
||||||
func (p *LLMPhraser) PhraseChat(ctx context.Context, utterance string, history []dialogue.Turn) (string, error) {
|
func (p *LLMPhraser) PhraseChat(ctx context.Context, utterance string, history []dialogue.Turn) (string, error) {
|
||||||
sys := chatSystemPrompt(p.cfg.Persona)
|
sys := chatSystemPrompt(p.cfg.ContextBlock)
|
||||||
msgs := []chatMsg{
|
msgs := []chatMsg{
|
||||||
{Role: "system", Content: sys},
|
{Role: "system", Content: sys},
|
||||||
}
|
}
|
||||||
@@ -248,12 +313,17 @@ func (p *LLMPhraser) PhraseChat(ctx context.Context, utterance string, history [
|
|||||||
combined += utterance
|
combined += utterance
|
||||||
msgs = append(msgs, chatMsg{Role: "user", Content: strings.TrimSpace(combined)})
|
msgs = append(msgs, chatMsg{Role: "user", Content: strings.TrimSpace(combined)})
|
||||||
|
|
||||||
resp, err := p.chatWithMessages(ctx, msgs, 512)
|
resp, err := p.chatWithMessages(ctx, msgs, 768)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
log.Printf("phraser: PhraseChat: %v", err)
|
log.Printf("phraser: PhraseChat: %v", err)
|
||||||
return "поговорили.", nil
|
return "поговорили.", nil
|
||||||
}
|
}
|
||||||
if text, _ := parseResponseMood(resp); text != "" {
|
text, _, perr := parseResponseMood(resp)
|
||||||
|
if perr != nil {
|
||||||
|
log.Printf("phraser: PhraseChat: %v", perr)
|
||||||
|
return "поговорили.", nil
|
||||||
|
}
|
||||||
|
if text != "" {
|
||||||
return text, nil
|
return text, nil
|
||||||
}
|
}
|
||||||
// fallback: plain text without JSON
|
// fallback: plain text without JSON
|
||||||
@@ -264,17 +334,16 @@ func (p *LLMPhraser) PhraseChat(ctx context.Context, utterance string, history [
|
|||||||
}
|
}
|
||||||
|
|
||||||
// chatSystemPrompt returns the system prompt for conversational chat.
|
// chatSystemPrompt returns the system prompt for conversational chat.
|
||||||
// Prepends the configured persona when set.
|
// Prepends the shared context block when the phraser has one.
|
||||||
func chatSystemPrompt(persona string) string {
|
func chatSystemPrompt(block func() string) string {
|
||||||
base := `You are maven, a self-hosted personal assistant. You're talking with your owner.
|
// No self-introduction here: the persona block prepended one line above
|
||||||
Keep replies brief (1-3 sentences) and natural. You're helpful, curious, and a little warm.
|
// already says who she is, same as router.KnowledgePrompt.
|
||||||
Respond in the user's language (Russian or English, matching their last message).
|
base := `Ты разговариваешь с хозяином. О себе говоришь в женском роде ("я подумала", "я рада"). Он мужчина: обращайся к нему на "ты", в мужском роде ("ты сказал", "ты забыл"). Никогда не "вы"/"ваш" и никогда "он"/"его" — ты говоришь ему, а не о нём.
|
||||||
Never roleplay emotions you don't have, but stay friendly.
|
|
||||||
Respond ONLY with valid JSON: {"response": "...", "mood": "neutral"}. "response" is your reply text; "mood" reflects your tone (neutral/happy/thinking/tired/confused).`
|
Отвечай по-русски, коротко: одна-три фразы, живым языком. Ты доброжелательная, тебе интересно, но чувства не изображай.
|
||||||
if persona != "" {
|
|
||||||
base = persona + "\n\n" + base
|
Отвечай ТОЛЬКО одним объектом JSON: {"response": "...", "mood": "neutral"}. В "response" — твой ответ. В "mood" — ровно одно из: neutral, happy, thinking, tired, confused.`
|
||||||
}
|
return persona.Prepend(block, base)
|
||||||
return base
|
|
||||||
}
|
}
|
||||||
|
|
||||||
// chatWithMessages sends a full message array (system + history + current) to
|
// chatWithMessages sends a full message array (system + history + current) to
|
||||||
@@ -285,6 +354,7 @@ func (p *LLMPhraser) chatWithMessages(ctx context.Context, msgs []chatMsg, maxTo
|
|||||||
Messages: msgs,
|
Messages: msgs,
|
||||||
Temperature: 0.7,
|
Temperature: 0.7,
|
||||||
MaxTokens: maxTokens,
|
MaxTokens: maxTokens,
|
||||||
|
Grammar: p.grammar(),
|
||||||
}
|
}
|
||||||
body, err := json.Marshal(req)
|
body, err := json.Marshal(req)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
@@ -336,7 +406,12 @@ func (p *LLMPhraser) PhraseReminder(ctx context.Context, d loop.ReminderDecision
|
|||||||
if err != nil {
|
if err != nil {
|
||||||
return delivery.PhrasedReminder{}, err
|
return delivery.PhrasedReminder{}, err
|
||||||
}
|
}
|
||||||
body, mood := parseResponseMood(resp)
|
body, mood, perr := parseResponseMood(resp)
|
||||||
|
if perr != nil {
|
||||||
|
// Truncated JSON. Fall through to the reminder's own text.
|
||||||
|
log.Printf("phraser: PhraseReminder: %v", perr)
|
||||||
|
body, mood = "", ""
|
||||||
|
}
|
||||||
if body == "" {
|
if body == "" {
|
||||||
// fallback: try old body/summary format
|
// fallback: try old body/summary format
|
||||||
body, _ = parsePhrase(resp)
|
body, _ = parsePhrase(resp)
|
||||||
@@ -364,6 +439,43 @@ type chatReq struct {
|
|||||||
Messages []chatMsg `json:"messages"`
|
Messages []chatMsg `json:"messages"`
|
||||||
Temperature float64 `json:"temperature"`
|
Temperature float64 `json:"temperature"`
|
||||||
MaxTokens int `json:"max_tokens"`
|
MaxTokens int `json:"max_tokens"`
|
||||||
|
// Grammar is llama-server's `grammar` field (GBNF). Same wiring as
|
||||||
|
// internal/llm.Req.Grammar. Empty ⇒ unconstrained sampling.
|
||||||
|
Grammar string `json:"grammar,omitempty"`
|
||||||
|
}
|
||||||
|
|
||||||
|
// responseGrammar — GBNF constraining the model to the documented phrasing
|
||||||
|
// contract and nothing else: {"response": "<text>", "mood": "<enum>"}.
|
||||||
|
//
|
||||||
|
// Without it a 0.8B answers roughly one chat turn in three with open reasoning
|
||||||
|
// as plain text ("Thinking Process:" …), which no tag-stripper can remove and
|
||||||
|
// which eats the token budget before the JSON closes. Modelled on
|
||||||
|
// routeGrammar in internal/router/llmrouter.go so the two read alike.
|
||||||
|
//
|
||||||
|
// text accepts ANY codepoint except the two JSON must escape — the replies are
|
||||||
|
// Russian, so an ASCII-only rule would make every reply empty. The escape rule
|
||||||
|
// is what lets the model close a string it opened with a quote inside. Length
|
||||||
|
// is bounded so a repetition loop truncates the field, not the JSON object.
|
||||||
|
//
|
||||||
|
// That bound was 400 and 400 was too tight. Measured against Qwen3.5-0.8B: on
|
||||||
|
// "почему гром слышно позже молнии?" the reply came back exactly 400 characters
|
||||||
|
// long, cut mid-word ("Нужно записать и,"), at every token cap from 256 to 2048.
|
||||||
|
// So the token cap was never what stopped it — this rule was. 1000 characters is
|
||||||
|
// roughly six Russian sentences, still short enough to stop a repetition loop.
|
||||||
|
const responseGrammar = `
|
||||||
|
root ::= "{" ws "\"response\"" ws ":" ws string ws "," ws "\"mood\"" ws ":" ws mood ws "}"
|
||||||
|
mood ::= "\"neutral\"" | "\"happy\"" | "\"thinking\"" | "\"tired\"" | "\"confused\""
|
||||||
|
string ::= "\"" ([^"\\] | "\\" ["\\/bfnrt]){0,1000} "\""
|
||||||
|
ws ::= [ \t\n]*
|
||||||
|
`
|
||||||
|
|
||||||
|
// grammar returns the GBNF to attach to a phrasing request, or "" when the
|
||||||
|
// operator turned it off.
|
||||||
|
func (p *LLMPhraser) grammar() string {
|
||||||
|
if p.cfg.NoGrammar {
|
||||||
|
return ""
|
||||||
|
}
|
||||||
|
return responseGrammar
|
||||||
}
|
}
|
||||||
|
|
||||||
type chatResp struct {
|
type chatResp struct {
|
||||||
@@ -388,6 +500,7 @@ func (p *LLMPhraser) chatWithSystem(ctx context.Context, system, user string, ma
|
|||||||
},
|
},
|
||||||
Temperature: 0.7,
|
Temperature: 0.7,
|
||||||
MaxTokens: maxTokens,
|
MaxTokens: maxTokens,
|
||||||
|
Grammar: p.grammar(),
|
||||||
}
|
}
|
||||||
body, err := json.Marshal(req)
|
body, err := json.Marshal(req)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
@@ -428,40 +541,166 @@ func (p *LLMPhraser) chatWithSystem(ctx context.Context, system, user string, ma
|
|||||||
return stripThink(content), nil
|
return stripThink(content), nil
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// nudgeSystem — the phrasing contract for nudges.
|
||||||
|
//
|
||||||
|
// Written as filled-in examples, not as a schema with "..." in it. A 0.8B
|
||||||
|
// copies whatever sits in the response slot, so a literal placeholder there
|
||||||
|
// teaches it to answer with the placeholder. Measured: 7/15 nudges came back
|
||||||
|
// as "..." before this. See PHRASING-EVAL-31-07-2026.md.
|
||||||
|
//
|
||||||
|
// Russian only, feminine self-reference, second person masculine (the owner is
|
||||||
|
// a man). She talks TO him, informally, singular — never "вы", never "он".
|
||||||
|
// One short sentence — the nudge is spoken aloud.
|
||||||
|
//
|
||||||
|
// What the ban on обращения forbids is pet names ("дорогой", "милый"), not his
|
||||||
|
// name: "Ками, ноутбук на трёх процентах" is exactly how she talks, and the
|
||||||
|
// unqualified word read as forbidding that too. Hence "ласковые обращения".
|
||||||
|
//
|
||||||
|
// The examples also never claim a physical act. She has no hands and no smart
|
||||||
|
// plug — she can tell him the battery is at three percent, she cannot put the
|
||||||
|
// laptop on charge. An example that says she did teaches the model to invent
|
||||||
|
// actions Maven never took, which is worse than a missing nudge.
|
||||||
|
const nudgeSystem = `Ты — Maven, домашняя ассистентка. О себе говоришь в женском роде ("я проверила", "я записала"). Владелец — мужчина, обращайся к нему в мужском роде ("ты пил", "ты забыл").
|
||||||
|
Говоришь с ним на "ты", в единственном числе ("выпей", "встань"). Никогда не "вы"/"вас"/"ваш" и никогда "он"/"его" — ты говоришь ему, а не о нём.
|
||||||
|
|
||||||
|
Пиши ОДНО короткое напоминание по-русски: не больше 120 символов и не больше 16 слов. Только по делу.
|
||||||
|
|
||||||
|
Запрещено: ласковые обращения ("дорогой", "милый"), эмодзи, извинения ("прости", "извини"), вопросы о самочувствии, похвала, больше одного восклицательного знака, английские слова кроме имён сервисов.
|
||||||
|
|
||||||
|
Отвечай ТОЛЬКО одним объектом JSON с полями "response" и "mood".
|
||||||
|
"response" — сам текст напоминания.
|
||||||
|
"mood" — ровно одно из: neutral, happy, thinking, tired, confused.
|
||||||
|
|
||||||
|
Так выглядит правильный ответ по форме. Темы здесь посторонние — их в запросе не будет:
|
||||||
|
{"response": "Стиральная машина закончила. Развесь бельё.", "mood": "neutral"}
|
||||||
|
{"response": "Ками, ноутбук на трёх процентах. Поставь его на зарядку.", "mood": "confused"}
|
||||||
|
|
||||||
|
Это примеры ФОРМЫ, а не темы. Пиши только про ту ситуацию, которую тебе дали в запросе. Не копируй примеры и никогда не пиши "..." в поле response.`
|
||||||
|
|
||||||
func (p *LLMPhraser) systemPrompt() string {
|
func (p *LLMPhraser) systemPrompt() string {
|
||||||
base := `You are maven, a self-hosted personal assistant. Generate brief, natural nudge messages in the user's language (Russian or English). Respond ONLY with valid JSON: {"response": "full voice message", "mood": "neutral"}. "response" is what the user hears; "mood" reflects maven's tone (neutral/happy/thinking/tired/confused).`
|
return persona.Prepend(p.cfg.ContextBlock, nudgeSystem)
|
||||||
if p.cfg.Persona != "" {
|
|
||||||
base = p.cfg.Persona + "\n\n" + base
|
|
||||||
}
|
|
||||||
return base
|
|
||||||
}
|
}
|
||||||
|
|
||||||
// querySystemPrompt returns the system prompt for PhraseQuery (notes + general
|
// querySystemPrompt returns the system prompt for PhraseQuery (notes + general
|
||||||
// knowledge). Prepends the configured persona when set.
|
// knowledge). Prepends the configured persona when set.
|
||||||
func (p *LLMPhraser) querySystemPrompt() string {
|
func (p *LLMPhraser) querySystemPrompt() string {
|
||||||
base := "You are maven, a self-hosted personal assistant answering from your notes. Answer briefly and naturally in Russian starting with \"вот что я нашла: \". Respond ONLY with valid JSON: {\"response\": \"...\", \"mood\": \"neutral\"}."
|
// No self-introduction here: the persona block prepended one line above
|
||||||
if p.cfg.Persona != "" {
|
// already says who she is, same as router.KnowledgePrompt.
|
||||||
base = p.cfg.Persona + "\n\n" + base
|
base := "Ты отвечаешь ему по своим заметкам. Отвечай по-русски, коротко и своими словами, начинай с \"вот что я нашла: \". О себе — в женском роде (\"нашла\", \"записала\"). Он мужчина, обращайся к нему на \"ты\". Respond ONLY with valid JSON: {\"response\": \"...\", \"mood\": \"neutral\"}."
|
||||||
|
return persona.Prepend(p.cfg.ContextBlock, base)
|
||||||
|
}
|
||||||
|
|
||||||
|
// ruleTopics — Russian gloss for each built-in rule name. The rule names are
|
||||||
|
// English identifiers; a 0.8B asked to nudge about "netdata_critical" writes
|
||||||
|
// about nothing. The daemon knows what its own rules mean, so it says so.
|
||||||
|
var ruleTopics = map[string]string{
|
||||||
|
"water": "он давно не пил воду",
|
||||||
|
"meal": "он давно не ел",
|
||||||
|
"break": "он давно без перерыва, пора встать и размяться",
|
||||||
|
"service_down": "сервис не отвечает, лежит",
|
||||||
|
"netdata_critical": "критический алярм в netdata, проблема с диском или местом",
|
||||||
|
}
|
||||||
|
|
||||||
|
// ruleKeywords — the word the message must contain. The 0.8B drifts to
|
||||||
|
// whatever topic it saw last unless the required word is named outright.
|
||||||
|
var ruleKeywords = map[string]string{
|
||||||
|
"water": "воду",
|
||||||
|
"meal": "поешь",
|
||||||
|
"break": "перерыв",
|
||||||
|
"service_down": "сервис",
|
||||||
|
"netdata_critical": "диск",
|
||||||
|
}
|
||||||
|
|
||||||
|
// ruleTopic turns a rule name into a Russian description of the situation.
|
||||||
|
// "routine:зарядка" and "morning:утро" carry their own Russian suffix.
|
||||||
|
func ruleTopic(rule string) string {
|
||||||
|
if t, ok := ruleTopics[rule]; ok {
|
||||||
|
return t
|
||||||
}
|
}
|
||||||
return base
|
if i := strings.IndexByte(rule, ':'); i > 0 && i+1 < len(rule) {
|
||||||
|
switch rule[:i] {
|
||||||
|
case "morning":
|
||||||
|
return "утро, пора начать день: " + rule[i+1:]
|
||||||
|
default:
|
||||||
|
return "пора сделать по распорядку: " + rule[i+1:]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return rule
|
||||||
|
}
|
||||||
|
|
||||||
|
// ruleKeyword — the word the nudge must contain, or "" when the rule name's
|
||||||
|
// own Russian suffix already is that word.
|
||||||
|
func ruleKeyword(rule string) string {
|
||||||
|
if k, ok := ruleKeywords[rule]; ok {
|
||||||
|
return k
|
||||||
|
}
|
||||||
|
if i := strings.IndexByte(rule, ':'); i > 0 && i+1 < len(rule) {
|
||||||
|
return rule[i+1:]
|
||||||
|
}
|
||||||
|
return ""
|
||||||
|
}
|
||||||
|
|
||||||
|
// ruDur — duration in Russian. humanDur is English and its output was landing
|
||||||
|
// verbatim in the message.
|
||||||
|
func ruDur(d time.Duration) string {
|
||||||
|
if d < 0 {
|
||||||
|
d = 0
|
||||||
|
}
|
||||||
|
h, m := int(d.Hours()), int(d.Minutes())%60
|
||||||
|
switch {
|
||||||
|
case h >= 2:
|
||||||
|
return fmt.Sprintf("%d ч", h)
|
||||||
|
case h == 1 && m >= 30:
|
||||||
|
return "полтора часа"
|
||||||
|
case h == 1:
|
||||||
|
return "час"
|
||||||
|
default:
|
||||||
|
return fmt.Sprintf("%d мин", m)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// fallbackNudge — plain Russian for when the model returns nothing parseable.
|
||||||
|
var fallbackNudges = map[string]string{
|
||||||
|
"water": "Ты давно не пил воду.",
|
||||||
|
"meal": "Ты давно не ел, поешь.",
|
||||||
|
"break": "Пора сделать перерыв.",
|
||||||
|
"service_down": "Сервис не отвечает.",
|
||||||
|
"netdata_critical": "Критический алярм: проверь диск.",
|
||||||
|
}
|
||||||
|
|
||||||
|
func fallbackNudge(c loop.Candidate) string {
|
||||||
|
if s, ok := fallbackNudges[c.Rule.Name]; ok {
|
||||||
|
return s
|
||||||
|
}
|
||||||
|
if kw := ruleKeyword(c.Rule.Name); kw != "" {
|
||||||
|
return "Напоминаю: " + kw + "."
|
||||||
|
}
|
||||||
|
return "Напоминаю о деле."
|
||||||
}
|
}
|
||||||
|
|
||||||
func buildNudgePrompt(c loop.Candidate) string {
|
func buildNudgePrompt(c loop.Candidate) string {
|
||||||
var ctxParts []string
|
var ctxParts []string
|
||||||
ctxParts = append(ctxParts, fmt.Sprintf("Rule: %s", c.Rule.Name))
|
ctxParts = append(ctxParts, "Ситуация: "+ruleTopic(c.Rule.Name))
|
||||||
ctxParts = append(ctxParts, fmt.Sprintf("Severity: %s", sevLabel(c.Severity)))
|
if f, ok := c.State.Facts[c.Rule.Name]; ok && f.Key != "" && f.Key != c.Rule.Name {
|
||||||
|
ctxParts = append(ctxParts, "Что именно: "+f.Key)
|
||||||
|
}
|
||||||
if d, ok := c.State.Since(c.Rule.Name); ok {
|
if d, ok := c.State.Since(c.Rule.Name); ok {
|
||||||
ctxParts = append(ctxParts, fmt.Sprintf("Duration since last event: %s", humanDur(d)))
|
ctxParts = append(ctxParts, "Прошло: "+ruDur(d))
|
||||||
|
}
|
||||||
|
switch sevLabel(c.Severity) {
|
||||||
|
case "alarm":
|
||||||
|
ctxParts = append(ctxParts, "Срочно, скажи прямо.")
|
||||||
|
case "ops":
|
||||||
|
ctxParts = append(ctxParts, "Это про сервер, не про здоровье.")
|
||||||
|
}
|
||||||
|
tail := "Напиши напоминание про эту ситуацию. Одно предложение, по-русски, в JSON."
|
||||||
|
if kw := ruleKeyword(c.Rule.Name); kw != "" {
|
||||||
|
// Last line on purpose: a 0.8B weights the end of the prompt hardest,
|
||||||
|
// and without the required word it drifts back to the examples.
|
||||||
|
tail += " Ответ ДОЛЖЕН содержать слово «" + kw + "»."
|
||||||
}
|
}
|
||||||
|
|
||||||
return fmt.Sprintf(
|
return strings.Join(ctxParts, "\n") + "\n\n" + tail
|
||||||
`Generate a nudge message. Context:
|
|
||||||
%s
|
|
||||||
|
|
||||||
Respond as JSON: {"response": "...", "mood": "..."}`,
|
|
||||||
strings.Join(ctxParts, "\n"),
|
|
||||||
)
|
|
||||||
}
|
}
|
||||||
|
|
||||||
type responseMood struct {
|
type responseMood struct {
|
||||||
@@ -469,21 +708,39 @@ type responseMood struct {
|
|||||||
Mood string `json:"mood"`
|
Mood string `json:"mood"`
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// errBrokenJSON — the model started a JSON object and never finished it.
|
||||||
|
// That is a failed generation, not a reply. Callers must use their fallback.
|
||||||
|
var errBrokenJSON = fmt.Errorf("phraser: model output starts as JSON but does not parse")
|
||||||
|
|
||||||
// parseResponseMood extracts {"response","mood"} from LLM output, tolerant
|
// parseResponseMood extracts {"response","mood"} from LLM output, tolerant
|
||||||
// of thinking tokens and extra text before/after the JSON block. Returns
|
// of thinking tokens and extra text before/after the JSON block.
|
||||||
// ("", "") when no valid JSON is found.
|
//
|
||||||
func parseResponseMood(raw string) (response, mood string) {
|
// Three outcomes:
|
||||||
|
// - parsed fine → the fields, nil error.
|
||||||
|
// - output never looked like JSON → ("", "", nil). The caller may ship it
|
||||||
|
// as-is; small models sometimes answer in bare prose and that is fine.
|
||||||
|
// - output starts with "{" but does not parse → errBrokenJSON. The grammar
|
||||||
|
// guarantees a valid *prefix*, so a generation that hits the token cap
|
||||||
|
// mid-object comes back as a fragment like `{` or `{\n "`. Shipping that
|
||||||
|
// as a reply is the bug this error exists to stop.
|
||||||
|
func parseResponseMood(raw string) (response, mood string, err error) {
|
||||||
cleaned := strings.TrimSpace(raw)
|
cleaned := strings.TrimSpace(raw)
|
||||||
start := strings.Index(cleaned, "{")
|
start := strings.Index(cleaned, "{")
|
||||||
end := strings.LastIndex(cleaned, "}")
|
end := strings.LastIndex(cleaned, "}")
|
||||||
if start < 0 || end < 0 || end <= start {
|
if start < 0 || end < 0 || end <= start {
|
||||||
return "", ""
|
if strings.HasPrefix(cleaned, "{") {
|
||||||
|
return "", "", errBrokenJSON
|
||||||
|
}
|
||||||
|
return "", "", nil
|
||||||
}
|
}
|
||||||
var parsed responseMood
|
var parsed responseMood
|
||||||
if err := json.Unmarshal([]byte(cleaned[start:end+1]), &parsed); err != nil {
|
if e := json.Unmarshal([]byte(cleaned[start:end+1]), &parsed); e != nil {
|
||||||
return "", ""
|
if strings.HasPrefix(cleaned, "{") {
|
||||||
|
return "", "", errBrokenJSON
|
||||||
|
}
|
||||||
|
return "", "", nil
|
||||||
}
|
}
|
||||||
return parsed.Response, parsed.Mood
|
return parsed.Response, parsed.Mood, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
func parsePhrase(raw string) (body, summary string) {
|
func parsePhrase(raw string) (body, summary string) {
|
||||||
|
|||||||
@@ -0,0 +1,261 @@
|
|||||||
|
package phraser
|
||||||
|
|
||||||
|
// Hand-written Russian nudges instead of generated ones.
|
||||||
|
//
|
||||||
|
// Why: on a nudge there is nothing to be creative about. Measured over many
|
||||||
|
// runs, Qwen3.5-0.8B breaks the persona (formal "вы", plural imperatives,
|
||||||
|
// masculine self-reference) and invents facts and units — it once told him to
|
||||||
|
// boil an egg for "90-95 секунд". A nudge is five words of known content, so
|
||||||
|
// wording it with a model buys nothing and risks the persona every time.
|
||||||
|
//
|
||||||
|
// The wording lives in nudges_ru_v1.json so it can be edited without touching
|
||||||
|
// Go. This file only picks one and fills in the values.
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
_ "embed"
|
||||||
|
"encoding/json"
|
||||||
|
"fmt"
|
||||||
|
"math/rand"
|
||||||
|
"regexp"
|
||||||
|
"strings"
|
||||||
|
"sync"
|
||||||
|
"time"
|
||||||
|
"unicode"
|
||||||
|
|
||||||
|
"github.com/kami/maven/internal/delivery"
|
||||||
|
"github.com/kami/maven/internal/loop"
|
||||||
|
)
|
||||||
|
|
||||||
|
//go:embed nudges_ru_v1.json
|
||||||
|
var nudgeTemplateJSON []byte
|
||||||
|
|
||||||
|
// NudgeTemplateSchemaVersion — the version this code understands.
|
||||||
|
const NudgeTemplateSchemaVersion = 1
|
||||||
|
|
||||||
|
type nudgeRuleSet struct {
|
||||||
|
Mood string `json:"mood"`
|
||||||
|
Variants []string `json:"variants"`
|
||||||
|
}
|
||||||
|
|
||||||
|
type nudgeTemplateFile struct {
|
||||||
|
SchemaVersion int `json:"schema_version"`
|
||||||
|
Name string `json:"name"`
|
||||||
|
Notes []string `json:"notes"`
|
||||||
|
Rules map[string]nudgeRuleSet `json:"rules"`
|
||||||
|
}
|
||||||
|
|
||||||
|
// NudgeTemplates picks a hand-written Russian nudge for a candidate.
|
||||||
|
//
|
||||||
|
// Safe for concurrent use. Random, but never the same variant twice in a row
|
||||||
|
// for the same rule — being nagged with identical words is what makes a nudge
|
||||||
|
// easy to tune out.
|
||||||
|
type NudgeTemplates struct {
|
||||||
|
mu sync.Mutex
|
||||||
|
rnd *rand.Rand
|
||||||
|
last map[string]string // rule family -> the text used last time
|
||||||
|
file nudgeTemplateFile
|
||||||
|
}
|
||||||
|
|
||||||
|
// NewNudgeTemplates loads the embedded template file. Pass a source to make the
|
||||||
|
// picking reproducible in tests; nil means seed from the clock.
|
||||||
|
func NewNudgeTemplates(src rand.Source) (*NudgeTemplates, error) {
|
||||||
|
var f nudgeTemplateFile
|
||||||
|
if err := json.Unmarshal(nudgeTemplateJSON, &f); err != nil {
|
||||||
|
return nil, fmt.Errorf("nudge templates: parse: %w", err)
|
||||||
|
}
|
||||||
|
if f.SchemaVersion != NudgeTemplateSchemaVersion {
|
||||||
|
return nil, fmt.Errorf("nudge templates: schema_version %d, want %d",
|
||||||
|
f.SchemaVersion, NudgeTemplateSchemaVersion)
|
||||||
|
}
|
||||||
|
if len(f.Rules) == 0 {
|
||||||
|
return nil, fmt.Errorf("nudge templates: no rules")
|
||||||
|
}
|
||||||
|
if src == nil {
|
||||||
|
src = rand.NewSource(time.Now().UnixNano())
|
||||||
|
}
|
||||||
|
return &NudgeTemplates{
|
||||||
|
rnd: rand.New(src),
|
||||||
|
last: map[string]string{},
|
||||||
|
file: f,
|
||||||
|
}, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// PhraseNudge implements the nudge half of the Phraser interface, so the
|
||||||
|
// templates can be scored by the same harness as the model.
|
||||||
|
func (t *NudgeTemplates) PhraseNudge(_ context.Context, c loop.Candidate) (delivery.PhrasedNudge, error) {
|
||||||
|
body, mood := t.Nudge(c)
|
||||||
|
return delivery.PhrasedNudge{Candidate: c, Body: body, Summary: body, Mood: mood}, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// Nudge returns the text and the mood for one candidate. Never fails: if no
|
||||||
|
// template fits it uses the plain per-rule fallback.
|
||||||
|
func (t *NudgeTemplates) Nudge(c loop.Candidate) (body, mood string) {
|
||||||
|
rule := c.Rule.Name
|
||||||
|
family := t.family(rule)
|
||||||
|
set, ok := t.file.Rules[family]
|
||||||
|
if !ok {
|
||||||
|
return fallbackNudge(c), "neutral"
|
||||||
|
}
|
||||||
|
vals := nudgeValues(c)
|
||||||
|
|
||||||
|
// Only variants whose placeholders all have a value.
|
||||||
|
usable := make([]string, 0, len(set.Variants))
|
||||||
|
for _, v := range set.Variants {
|
||||||
|
if text, ok := fillTemplate(v, vals); ok {
|
||||||
|
usable = append(usable, text)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if len(usable) == 0 {
|
||||||
|
return fallbackNudge(c), "neutral"
|
||||||
|
}
|
||||||
|
|
||||||
|
mood = set.Mood
|
||||||
|
if mood == "" {
|
||||||
|
mood = "neutral"
|
||||||
|
}
|
||||||
|
return t.pick(family, usable), mood
|
||||||
|
}
|
||||||
|
|
||||||
|
// pick chooses at random, skipping whatever this rule said last time.
|
||||||
|
func (t *NudgeTemplates) pick(family string, usable []string) string {
|
||||||
|
t.mu.Lock()
|
||||||
|
defer t.mu.Unlock()
|
||||||
|
|
||||||
|
choices := usable
|
||||||
|
if len(usable) > 1 {
|
||||||
|
choices = make([]string, 0, len(usable))
|
||||||
|
for _, v := range usable {
|
||||||
|
if v != t.last[family] {
|
||||||
|
choices = append(choices, v)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if len(choices) == 0 { // every variant equals the last one
|
||||||
|
choices = usable
|
||||||
|
}
|
||||||
|
}
|
||||||
|
got := choices[t.rnd.Intn(len(choices))]
|
||||||
|
t.last[family] = got
|
||||||
|
return got
|
||||||
|
}
|
||||||
|
|
||||||
|
// family maps a rule name to a block in the template file: an exact match
|
||||||
|
// first, then the prefix of "routine:зарядка" / "morning:утро", then "default".
|
||||||
|
func (t *NudgeTemplates) family(rule string) string {
|
||||||
|
if _, ok := t.file.Rules[rule]; ok {
|
||||||
|
return rule
|
||||||
|
}
|
||||||
|
if i := strings.IndexByte(rule, ':'); i > 0 {
|
||||||
|
if _, ok := t.file.Rules[rule[:i]]; ok {
|
||||||
|
return rule[:i]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return "default"
|
||||||
|
}
|
||||||
|
|
||||||
|
// placeholderRE — the {name} slots a template may use.
|
||||||
|
var placeholderRE = regexp.MustCompile(`\{([a-z]+)\}`)
|
||||||
|
|
||||||
|
// nudgeValues collects what this candidate can fill in. A key missing here
|
||||||
|
// means every template needing it is skipped, so nothing half-filled is ever
|
||||||
|
// spoken.
|
||||||
|
func nudgeValues(c loop.Candidate) map[string]string {
|
||||||
|
vals := map[string]string{}
|
||||||
|
rule := c.Rule.Name
|
||||||
|
|
||||||
|
// {since} — only at hour scale. Below an hour the phrase would be minutes,
|
||||||
|
// and none of the templates read well with "сорок минут".
|
||||||
|
if d, ok := c.State.Since(rule); ok && d >= time.Hour {
|
||||||
|
if s := ruSinceWords(d); s != "" {
|
||||||
|
vals["since"] = s
|
||||||
|
}
|
||||||
|
}
|
||||||
|
// {service} — the aggregate fact's key carries the service name.
|
||||||
|
if f, ok := c.State.Fact(rule); ok && f.Key != "" && f.Key != rule {
|
||||||
|
vals["service"] = f.Key
|
||||||
|
}
|
||||||
|
// {what} — the Russian suffix of "routine:таблетки" / "morning:утро".
|
||||||
|
if i := strings.IndexByte(rule, ':'); i > 0 && i+1 < len(rule) {
|
||||||
|
vals["what"] = rule[i+1:]
|
||||||
|
}
|
||||||
|
return vals
|
||||||
|
}
|
||||||
|
|
||||||
|
// fillTemplate substitutes the placeholders. Returns false when a value is
|
||||||
|
// missing, so a raw "{since}" can never reach the text-to-speech voice.
|
||||||
|
func fillTemplate(tmpl string, vals map[string]string) (string, bool) {
|
||||||
|
missing := false
|
||||||
|
out := placeholderRE.ReplaceAllStringFunc(tmpl, func(m string) string {
|
||||||
|
name := m[1 : len(m)-1]
|
||||||
|
v, ok := vals[name]
|
||||||
|
if !ok || v == "" {
|
||||||
|
missing = true
|
||||||
|
return m
|
||||||
|
}
|
||||||
|
return v
|
||||||
|
})
|
||||||
|
if missing || strings.ContainsAny(out, "{}%") {
|
||||||
|
return "", false
|
||||||
|
}
|
||||||
|
return capitalizeFirst(out), true
|
||||||
|
}
|
||||||
|
|
||||||
|
// capitalizeFirst — a placeholder can start the sentence, and "полтора часа без
|
||||||
|
// перерыва" should be spoken as a sentence, not a fragment.
|
||||||
|
func capitalizeFirst(s string) string {
|
||||||
|
for i, r := range s {
|
||||||
|
return string(unicode.ToUpper(r)) + s[i+len(string(r)):]
|
||||||
|
}
|
||||||
|
return s
|
||||||
|
}
|
||||||
|
|
||||||
|
// hourWords — hours spelled out. "3 ч" is fine on a screen and wrong in a
|
||||||
|
// Russian voice, so the number goes out as words.
|
||||||
|
var hourWords = []string{
|
||||||
|
"ноль", "один", "два", "три", "четыре", "пять", "шесть", "семь", "восемь",
|
||||||
|
"девять", "десять", "одиннадцать", "двенадцать", "тринадцать",
|
||||||
|
"четырнадцать", "пятнадцать", "шестнадцать", "семнадцать", "восемнадцать",
|
||||||
|
"девятнадцать", "двадцать", "двадцать один", "двадцать два", "двадцать три",
|
||||||
|
}
|
||||||
|
|
||||||
|
// hourPlural — час / часа / часов by Russian counting rules.
|
||||||
|
func hourPlural(h int) string {
|
||||||
|
if h%100 >= 11 && h%100 <= 14 {
|
||||||
|
return "часов"
|
||||||
|
}
|
||||||
|
switch h % 10 {
|
||||||
|
case 1:
|
||||||
|
return "час"
|
||||||
|
case 2, 3, 4:
|
||||||
|
return "часа"
|
||||||
|
default:
|
||||||
|
return "часов"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// ruSinceWords — "полтора часа", "два с половиной часа", "семь часов".
|
||||||
|
// Empty string means "do not say it" (under an hour, or over a day).
|
||||||
|
func ruSinceWords(d time.Duration) string {
|
||||||
|
if d < time.Hour {
|
||||||
|
return ""
|
||||||
|
}
|
||||||
|
h := int(d.Hours())
|
||||||
|
m := int(d.Minutes()) % 60
|
||||||
|
if m >= 45 {
|
||||||
|
h++
|
||||||
|
m = 0
|
||||||
|
}
|
||||||
|
if h >= len(hourWords) {
|
||||||
|
return "больше суток"
|
||||||
|
}
|
||||||
|
if h == 1 {
|
||||||
|
if m >= 15 {
|
||||||
|
return "полтора часа"
|
||||||
|
}
|
||||||
|
return "час"
|
||||||
|
}
|
||||||
|
if m >= 15 {
|
||||||
|
return hourWords[h] + " с половиной часа"
|
||||||
|
}
|
||||||
|
return hourWords[h] + " " + hourPlural(h)
|
||||||
|
}
|
||||||
@@ -0,0 +1,202 @@
|
|||||||
|
package phraser
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"math/rand"
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"github.com/kami/maven/internal/loop"
|
||||||
|
"github.com/kami/maven/internal/store"
|
||||||
|
)
|
||||||
|
|
||||||
|
// cand builds a candidate the way a tick would.
|
||||||
|
func cand(rule string, sinceMin int, factKey string) loop.Candidate {
|
||||||
|
now := time.Date(2026, 7, 31, 21, 40, 0, 0, time.UTC)
|
||||||
|
st := loop.State{Now: now, Facts: map[string]store.Fact{}}
|
||||||
|
if sinceMin > 0 || factKey != "" {
|
||||||
|
key := rule
|
||||||
|
if factKey != "" {
|
||||||
|
key = factKey
|
||||||
|
}
|
||||||
|
st.Facts[rule] = store.Fact{Key: key, Ts: now.Add(-time.Duration(sinceMin) * time.Minute)}
|
||||||
|
}
|
||||||
|
return loop.Candidate{Rule: loop.Rule{Name: rule, Severity: loop.Sev1}, Severity: loop.Sev1, State: st}
|
||||||
|
}
|
||||||
|
|
||||||
|
func newTestTemplates(t *testing.T, seed int64) *NudgeTemplates {
|
||||||
|
t.Helper()
|
||||||
|
nt, err := NewNudgeTemplates(rand.NewSource(seed))
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("NewNudgeTemplates: %v", err)
|
||||||
|
}
|
||||||
|
return nt
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestNudgeTemplatesLoad(t *testing.T) {
|
||||||
|
nt := newTestTemplates(t, 1)
|
||||||
|
for _, rule := range []string{"water", "meal", "break", "service_down", "netdata_critical", "routine", "morning", "default"} {
|
||||||
|
set, ok := nt.file.Rules[rule]
|
||||||
|
if !ok {
|
||||||
|
t.Errorf("no templates for %q", rule)
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if len(set.Variants) < 5 {
|
||||||
|
t.Errorf("%s: only %d variants", rule, len(set.Variants))
|
||||||
|
}
|
||||||
|
// Every rule needs one variant that needs no value, or a candidate
|
||||||
|
// without context has nothing to say. routine and morning are exempt:
|
||||||
|
// they always carry a name and must always say it.
|
||||||
|
plain := 0
|
||||||
|
seen := map[string]bool{}
|
||||||
|
for _, v := range set.Variants {
|
||||||
|
if !placeholderRE.MatchString(v) {
|
||||||
|
plain++
|
||||||
|
}
|
||||||
|
if seen[v] {
|
||||||
|
t.Errorf("%s: duplicate variant %q", rule, v)
|
||||||
|
}
|
||||||
|
seen[v] = true
|
||||||
|
}
|
||||||
|
if plain == 0 && rule != "routine" && rule != "morning" {
|
||||||
|
t.Errorf("%s: every variant needs a placeholder value", rule)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The whole point of the picker: never the same words twice in a row.
|
||||||
|
func TestNudgeNoImmediateRepeat(t *testing.T) {
|
||||||
|
nt := newTestTemplates(t, 7)
|
||||||
|
prev := ""
|
||||||
|
for i := 0; i < 200; i++ {
|
||||||
|
body, _ := nt.Nudge(cand("water", 200, ""))
|
||||||
|
if body == prev {
|
||||||
|
t.Fatalf("repeat at %d: %q", i, body)
|
||||||
|
}
|
||||||
|
prev = body
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Same seed, same sequence — otherwise the fixture score would drift run to run.
|
||||||
|
func TestNudgeDeterministicWithSeed(t *testing.T) {
|
||||||
|
var runs [2][]string
|
||||||
|
for r := range runs {
|
||||||
|
nt := newTestTemplates(t, 42)
|
||||||
|
for i := 0; i < 20; i++ {
|
||||||
|
body, _ := nt.Nudge(cand("break", 100, ""))
|
||||||
|
runs[r] = append(runs[r], body)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
for i := range runs[0] {
|
||||||
|
if runs[0][i] != runs[1][i] {
|
||||||
|
t.Fatalf("run %d differs: %q vs %q", i, runs[0][i], runs[1][i])
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// A variant is only used when its value exists, and nothing half-filled ships.
|
||||||
|
func TestNudgeNoLeftoverPlaceholders(t *testing.T) {
|
||||||
|
nt := newTestTemplates(t, 3)
|
||||||
|
cases := []loop.Candidate{
|
||||||
|
cand("water", 0, ""), // no duration
|
||||||
|
cand("water", 30, ""), // under an hour
|
||||||
|
cand("water", 200, ""), // hours
|
||||||
|
cand("service_down", 3, "vaultwarden"),
|
||||||
|
cand("service_down", 3, ""), // no service name
|
||||||
|
cand("routine:таблетки", 0, ""),
|
||||||
|
cand("morning:утро", 0, ""),
|
||||||
|
cand("unknown_rule", 0, ""),
|
||||||
|
}
|
||||||
|
for _, c := range cases {
|
||||||
|
for i := 0; i < 40; i++ {
|
||||||
|
body, mood := nt.Nudge(c)
|
||||||
|
if body == "" {
|
||||||
|
t.Fatalf("%s: empty body", c.Rule.Name)
|
||||||
|
}
|
||||||
|
if strings.ContainsAny(body, "{}%") {
|
||||||
|
t.Fatalf("%s: unfilled template %q", c.Rule.Name, body)
|
||||||
|
}
|
||||||
|
if mood != "neutral" {
|
||||||
|
t.Fatalf("%s: mood %q", c.Rule.Name, mood)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The routine name must actually land in the text.
|
||||||
|
func TestNudgeSubstitutesWhat(t *testing.T) {
|
||||||
|
nt := newTestTemplates(t, 11)
|
||||||
|
for i := 0; i < 40; i++ {
|
||||||
|
body, _ := nt.Nudge(cand("routine:таблетки", 0, ""))
|
||||||
|
if !strings.Contains(strings.ToLower(body), "таблетки") {
|
||||||
|
t.Fatalf("routine text lost the name: %q", body)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestRuSinceWords(t *testing.T) {
|
||||||
|
cases := []struct {
|
||||||
|
min int
|
||||||
|
want string
|
||||||
|
}{
|
||||||
|
{30, ""},
|
||||||
|
{60, "час"},
|
||||||
|
{95, "полтора часа"},
|
||||||
|
{150, "два с половиной часа"},
|
||||||
|
{190, "три часа"},
|
||||||
|
{240, "четыре часа"},
|
||||||
|
{430, "семь часов"},
|
||||||
|
{660, "одиннадцать часов"},
|
||||||
|
{60 * 30, "больше суток"},
|
||||||
|
}
|
||||||
|
for _, c := range cases {
|
||||||
|
got := ruSinceWords(time.Duration(c.min) * time.Minute)
|
||||||
|
if got != c.want {
|
||||||
|
t.Errorf("%d min: got %q want %q", c.min, got, c.want)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Templates are the default: a nudge must not reach the model at all.
|
||||||
|
func TestLLMPhraserUsesTemplatesByDefault(t *testing.T) {
|
||||||
|
spy := newGrammarSpy(t)
|
||||||
|
p := NewLLMPhraserAt(spy.srv.URL, Config{})
|
||||||
|
pn, err := p.PhraseNudge(context.Background(), cand("water", 200, ""))
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("PhraseNudge: %v", err)
|
||||||
|
}
|
||||||
|
if len(spy.grammars) != 0 {
|
||||||
|
t.Errorf("nudge hit the model %d times, want 0", len(spy.grammars))
|
||||||
|
}
|
||||||
|
if !strings.Contains(strings.ToLower(pn.Body), "вод") {
|
||||||
|
t.Errorf("nudge is not the water template: %q", pn.Body)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// ...and the flag brings the model back.
|
||||||
|
func TestLLMNudgesFlagRestoresTheModel(t *testing.T) {
|
||||||
|
spy := newGrammarSpy(t)
|
||||||
|
p := NewLLMPhraserAt(spy.srv.URL, Config{LLMNudges: true})
|
||||||
|
pn, err := p.PhraseNudge(context.Background(), cand("water", 200, ""))
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("PhraseNudge: %v", err)
|
||||||
|
}
|
||||||
|
if len(spy.grammars) != 1 {
|
||||||
|
t.Fatalf("nudge hit the model %d times, want 1", len(spy.grammars))
|
||||||
|
}
|
||||||
|
if pn.Body != "ага" {
|
||||||
|
t.Errorf("body = %q, want the model's reply", pn.Body)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestNudgeTemplatesPhraseNudge(t *testing.T) {
|
||||||
|
nt := newTestTemplates(t, 5)
|
||||||
|
pn, err := nt.PhraseNudge(context.Background(), cand("water", 200, ""))
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("PhraseNudge: %v", err)
|
||||||
|
}
|
||||||
|
if pn.Body == "" || pn.Summary != pn.Body || pn.Mood != "neutral" {
|
||||||
|
t.Fatalf("bad nudge: %+v", pn)
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,129 @@
|
|||||||
|
{
|
||||||
|
"schema_version": 1,
|
||||||
|
"name": "russian nudge templates v1",
|
||||||
|
"notes": [
|
||||||
|
"Hand-written Russian nudges. Edit the wording here, no Go changes needed.",
|
||||||
|
"Rules: she is feminine about herself, he is a man addressed as ты. Never вы/вас/ваш, never plural imperatives (выпейте), never он/его about him.",
|
||||||
|
"One short sentence. No questions, no emoji, no pet names, no emotional support.",
|
||||||
|
"Placeholders: {since} how long it has been (only used when it is at least an hour), {service} the service name, {what} the routine name. A variant whose placeholder has no value is skipped, so every rule needs at least one variant with no placeholder. The exception is routine and morning: those only exist for rules like routine:таблетки that always carry a name, and a routine nudge that drops the name is useless.",
|
||||||
|
"mood must be one of: neutral, happy, thinking, tired, confused."
|
||||||
|
],
|
||||||
|
"rules": {
|
||||||
|
"water": {
|
||||||
|
"mood": "neutral",
|
||||||
|
"variants": [
|
||||||
|
"Ты не пил воду {since} — выпей стакан.",
|
||||||
|
"Пора выпить воды.",
|
||||||
|
"Стакан воды не помешает.",
|
||||||
|
"Воду ты не пил уже {since}.",
|
||||||
|
"Напоминаю про воду.",
|
||||||
|
"Сходи за водой, дела подождут.",
|
||||||
|
"Сделай глоток воды, пока помнишь.",
|
||||||
|
"Между делом выпей воды.",
|
||||||
|
"Вода — простое дело: выпей стакан.",
|
||||||
|
"Отвлекись на стакан воды."
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"meal": {
|
||||||
|
"mood": "neutral",
|
||||||
|
"variants": [
|
||||||
|
"Ты не ел {since} — поешь.",
|
||||||
|
"Пора поесть, сделай перекус.",
|
||||||
|
"Еда важнее ещё одного часа за столом.",
|
||||||
|
"Без еды уже {since}, поешь.",
|
||||||
|
"Напоминаю про еду — поешь.",
|
||||||
|
"Возьми перерыв на обед.",
|
||||||
|
"Сделай себе перекус, это пять минут.",
|
||||||
|
"Поешь, потом вернёшься к работе.",
|
||||||
|
"Поешь нормально, а не на ходу.",
|
||||||
|
"Еды не было {since} — разогрей что-нибудь."
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"break": {
|
||||||
|
"mood": "neutral",
|
||||||
|
"variants": [
|
||||||
|
"Ты за столом {since} — встань и разомнись.",
|
||||||
|
"Пора сделать перерыв.",
|
||||||
|
"Встань на пять минут.",
|
||||||
|
"{since} без перерыва — отойди от экрана.",
|
||||||
|
"Напоминаю про перерыв.",
|
||||||
|
"Разомни спину, потом продолжишь.",
|
||||||
|
"Короткая пауза не сорвёт дела.",
|
||||||
|
"Отойди от компьютера на минуту.",
|
||||||
|
"Сидишь без перерыва {since}.",
|
||||||
|
"Встань, пройдись, вернись."
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"service_down": {
|
||||||
|
"mood": "neutral",
|
||||||
|
"variants": [
|
||||||
|
"Сервис {service} не отвечает.",
|
||||||
|
"{service} упал — сервис не отвечает.",
|
||||||
|
"{service} не отвечает, сервис нужно поднимать.",
|
||||||
|
"Сервис {service} недоступен.",
|
||||||
|
"Проверь {service}: сервис не отвечает.",
|
||||||
|
"Сервис перестал отвечать.",
|
||||||
|
"Сервис {service} лежит, нужно смотреть.",
|
||||||
|
"{service} не отвечает уже {since}.",
|
||||||
|
"Мониторинг сообщает: {service} лежит.",
|
||||||
|
"Сервис {service} не отвечает, посмотри логи."
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"netdata_critical": {
|
||||||
|
"mood": "neutral",
|
||||||
|
"variants": [
|
||||||
|
"Netdata: критический алярм, проверь диск.",
|
||||||
|
"Критический алярм в netdata — посмотри диск.",
|
||||||
|
"Netdata поднял тревогу по диску.",
|
||||||
|
"Проверь диск: netdata ругается.",
|
||||||
|
"Алярм от netdata, критический.",
|
||||||
|
"Netdata: критический уровень, дело в диске.",
|
||||||
|
"Диск требует внимания — критический алярм в netdata.",
|
||||||
|
"Критический алярм: проверь место на диске.",
|
||||||
|
"Netdata сообщает о критической проблеме с диском.",
|
||||||
|
"Открой netdata: там критический алярм по диску."
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"routine": {
|
||||||
|
"mood": "neutral",
|
||||||
|
"variants": [
|
||||||
|
"По распорядку: {what}.",
|
||||||
|
"Пора — {what}.",
|
||||||
|
"Напоминаю: {what}.",
|
||||||
|
"В списке на сейчас: {what}.",
|
||||||
|
"{what} — сейчас самое время.",
|
||||||
|
"Не пропусти: {what}.",
|
||||||
|
"{what}: пора сделать.",
|
||||||
|
"Сейчас по плану {what}.",
|
||||||
|
"Твой распорядок: {what}.",
|
||||||
|
"{what} — по распорядку сейчас."
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"morning": {
|
||||||
|
"mood": "neutral",
|
||||||
|
"variants": [
|
||||||
|
"{what} — пора начать день.",
|
||||||
|
"{what}: пройди утренний список.",
|
||||||
|
"Начни {what} со списка.",
|
||||||
|
"{what}. Осталось пройти чеклист.",
|
||||||
|
"Утренний список ещё не пройден: {what}.",
|
||||||
|
"{what}: первый пункт списка за тобой.",
|
||||||
|
"{what} идёт, а список стоит.",
|
||||||
|
"{what}: не забудь про утренние дела.",
|
||||||
|
"По утреннему чеклисту ещё есть дела: {what}.",
|
||||||
|
"{what} — утренний список дел ещё ждёт."
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"default": {
|
||||||
|
"mood": "neutral",
|
||||||
|
"variants": [
|
||||||
|
"Напоминаю: есть дело.",
|
||||||
|
"Пора вернуться к отложенному делу.",
|
||||||
|
"Одно дело ждёт тебя.",
|
||||||
|
"Напоминаю про дело из списка.",
|
||||||
|
"В списке осталось дело.",
|
||||||
|
"Дело всё ещё не сделано."
|
||||||
|
]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -2,6 +2,7 @@ package router
|
|||||||
|
|
||||||
import (
|
import (
|
||||||
"context"
|
"context"
|
||||||
|
"fmt"
|
||||||
"math"
|
"math"
|
||||||
"unicode"
|
"unicode"
|
||||||
)
|
)
|
||||||
@@ -19,6 +20,24 @@ type Embedder interface {
|
|||||||
Close() error
|
Close() error
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// IdentifiedEmbedder — an embedder that can name itself. The name goes into
|
||||||
|
// the DB next to the vectors it wrote, so a later model swap is caught instead
|
||||||
|
// of silently returning nonsense scores (Vikunja #378).
|
||||||
|
type IdentifiedEmbedder interface {
|
||||||
|
Embedder
|
||||||
|
ID() string
|
||||||
|
}
|
||||||
|
|
||||||
|
// EmbedderID is the stable string stored alongside the vectors. It comes from
|
||||||
|
// the embedder itself — nobody hand-types a model name twice — and changes
|
||||||
|
// whenever the model or its dimension changes.
|
||||||
|
func EmbedderID(e Embedder) string {
|
||||||
|
if i, ok := e.(IdentifiedEmbedder); ok {
|
||||||
|
return i.ID()
|
||||||
|
}
|
||||||
|
return fmt.Sprintf("unknown@%d", e.Dim())
|
||||||
|
}
|
||||||
|
|
||||||
// AsymmetricEmbedder — an embedder that wants to know whether a text is a
|
// AsymmetricEmbedder — an embedder that wants to know whether a text is a
|
||||||
// search query or a stored passage. Recall is asymmetric: a short question
|
// search query or a stored passage. Recall is asymmetric: a short question
|
||||||
// goes in, a longer note comes out. The e5 family is trained for exactly that
|
// goes in, a longer note comes out. The e5 family is trained for exactly that
|
||||||
@@ -70,6 +89,10 @@ func NewHashEmbedder(dim int) *HashEmbedder {
|
|||||||
|
|
||||||
func (h *HashEmbedder) Dim() int { return h.dim }
|
func (h *HashEmbedder) Dim() int { return h.dim }
|
||||||
|
|
||||||
|
// ID names this embedder for the DB marker. The dimension is part of it
|
||||||
|
// because a HashEmbedder of another width is a different vector space.
|
||||||
|
func (h *HashEmbedder) ID() string { return fmt.Sprintf("hash@%d", h.dim) }
|
||||||
|
|
||||||
func (h *HashEmbedder) Close() error { return nil }
|
func (h *HashEmbedder) Close() error { return nil }
|
||||||
|
|
||||||
func (h *HashEmbedder) Embed(_ context.Context, text string) ([]float32, error) {
|
func (h *HashEmbedder) Embed(_ context.Context, text string) ([]float32, error) {
|
||||||
|
|||||||
@@ -0,0 +1,24 @@
|
|||||||
|
package router
|
||||||
|
|
||||||
|
import "testing"
|
||||||
|
|
||||||
|
func TestEmbedderIDFromModelPath(t *testing.T) {
|
||||||
|
got := modelIDFromPath("/opt/maven/models/embedder/multilingual-e5-small.onnx")
|
||||||
|
if got != "multilingual-e5-small@384" {
|
||||||
|
t.Fatalf("modelIDFromPath = %q", got)
|
||||||
|
}
|
||||||
|
// A different model file must produce a different id, even at 384 dim.
|
||||||
|
old := modelIDFromPath("/opt/maven/models/embedder/paraphrase-multilingual-MiniLM-L12-v2.onnx")
|
||||||
|
if old == got {
|
||||||
|
t.Fatal("two different models share one id")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestEmbedderIDIncludesDim(t *testing.T) {
|
||||||
|
if id := EmbedderID(NewHashEmbedder(1024)); id != "hash@1024" {
|
||||||
|
t.Fatalf("EmbedderID = %q", id)
|
||||||
|
}
|
||||||
|
if EmbedderID(NewHashEmbedder(1024)) == EmbedderID(NewHashEmbedder(384)) {
|
||||||
|
t.Fatal("dimension not part of the id")
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -1,11 +1,8 @@
|
|||||||
package eval
|
package eval
|
||||||
|
|
||||||
import (
|
import (
|
||||||
"bytes"
|
|
||||||
"context"
|
"context"
|
||||||
"encoding/json"
|
|
||||||
"fmt"
|
"fmt"
|
||||||
"net/http"
|
|
||||||
"os"
|
"os"
|
||||||
"strings"
|
"strings"
|
||||||
"testing"
|
"testing"
|
||||||
@@ -29,13 +26,20 @@ import (
|
|||||||
// a bake-off across checkpoints (#278, #250) produces tables you can tell
|
// a bake-off across checkpoints (#278, #250) produces tables you can tell
|
||||||
// apart. Point the variable at one server at a time.
|
// apart. Point the variable at one server at a time.
|
||||||
//
|
//
|
||||||
// Three configurations, because "the LLM router" is ambiguous and the three
|
// Two configurations, because "the LLM router" is ambiguous and the two numbers
|
||||||
// numbers answer different questions:
|
// answer different questions:
|
||||||
//
|
//
|
||||||
// llm-only — the model alone. Measures the prompt + grammar contract.
|
// llm-only — the model alone. Measures the prompt + grammar contract.
|
||||||
// cascade+llm — what #320 would actually ship: stage-0 grammar, then the
|
// cascade+llm — what #320 would actually ship: stage-0 grammar, then the
|
||||||
// model, then the classifier as the failure floor.
|
// model, then the classifier as the failure floor.
|
||||||
// llm-no-thinking — diagnostic only, not a shippable path (see below).
|
//
|
||||||
|
// There used to be a third, "thinking off", which looked 6 points better. It is
|
||||||
|
// gone: it was measured with a hand-rolled HTTP client that quietly dropped
|
||||||
|
// repeat_penalty, so the gap was the missing penalty and not the thinking mode.
|
||||||
|
// Re-measured with everything else held equal, thinking off scores exactly the
|
||||||
|
// same, case for case — and a direct probe shows this llama-server build ignores
|
||||||
|
// enable_thinking / reasoning_budget for this model anyway, so there was nothing
|
||||||
|
// to turn off. Full write-up in ROUTING-EVAL-31-07-2026.md (Vikunja #376).
|
||||||
func TestLLMRouterBaseline(t *testing.T) {
|
func TestLLMRouterBaseline(t *testing.T) {
|
||||||
base := os.Getenv("MAVEN_LLM_URL")
|
base := os.Getenv("MAVEN_LLM_URL")
|
||||||
if base == "" {
|
if base == "" {
|
||||||
@@ -57,12 +61,12 @@ func TestLLMRouterBaseline(t *testing.T) {
|
|||||||
}
|
}
|
||||||
|
|
||||||
ctx := context.Background()
|
ctx := context.Background()
|
||||||
model, err := ModelID(ctx, base)
|
model, err := llm.ModelID(ctx, base)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
// Not fatal: an unlabelled score is still a score. But say so loudly,
|
// Not fatal: an unlabelled score is still a score. But say so loudly,
|
||||||
// because an unlabelled row in a bake-off table is worthless.
|
// because an unlabelled row in a bake-off table is worthless.
|
||||||
t.Logf("could not read model id from %s: %v — reports will say %q", base, err, "unknown-model")
|
t.Logf("could not read model id from %s: %v — reports will say %q", base, err, llm.UnknownModel)
|
||||||
model = "unknown-model"
|
model = llm.UnknownModel
|
||||||
}
|
}
|
||||||
t.Logf("scoring model %s at %s", model, base)
|
t.Logf("scoring model %s at %s", model, base)
|
||||||
lr := router.NewLLMRouter(client)
|
lr := router.NewLLMRouter(client)
|
||||||
@@ -96,104 +100,18 @@ func TestLLMRouterBaseline(t *testing.T) {
|
|||||||
}
|
}
|
||||||
t.Log("\n" + repCascade.String() + repCascade.Failures())
|
t.Log("\n" + repCascade.String() + repCascade.Failures())
|
||||||
|
|
||||||
// llm-no-thinking: same prompt and grammar with the chat template's
|
|
||||||
// thinking mode off. Qwen3.5's template defaults thinking=1, so under a
|
|
||||||
// grammar the constrained JSON lands in reasoning_content with content
|
|
||||||
// empty — llm.Client's ReasoningContent fallback is what makes the router
|
|
||||||
// work at all today, by accident rather than design.
|
|
||||||
//
|
|
||||||
// MEASURED 2026-07-31: this variant scores identically to as-deployed
|
|
||||||
// (18/76, 48.7% intent-only, 2 errors, same p50). Thinking mode is a
|
|
||||||
// non-issue under a grammar — llama.cpp constrains the same token stream
|
|
||||||
// either way. Kept so the question stays answered instead of being
|
|
||||||
// re-asked, and so internal/llm does NOT grow a chat_template_kwargs field
|
|
||||||
// for a problem that does not exist.
|
|
||||||
repNoThink, err := Score(ctx, "llm-only ("+model+", thinking off) [diagnostic]",
|
|
||||||
RouterFunc(func(ctx context.Context, u string, now time.Time) (router.Decision, error) {
|
|
||||||
d, ok, err := router.NewLLMRouter(&noThinkCompleter{base: base, http: &http.Client{Timeout: 60 * time.Second}}).Route(ctx, u, now)
|
|
||||||
if err != nil {
|
|
||||||
return d, err
|
|
||||||
}
|
|
||||||
if !ok {
|
|
||||||
return d, fmt.Errorf("llm router declined without an error")
|
|
||||||
}
|
|
||||||
return d, nil
|
|
||||||
}), f)
|
|
||||||
if err != nil {
|
|
||||||
t.Fatalf("Score no-thinking: %v", err)
|
|
||||||
}
|
|
||||||
t.Log("\n" + repNoThink.String() + repNoThink.Failures())
|
|
||||||
|
|
||||||
// Reports rather than asserts — the numbers are inputs to the #320
|
// Reports rather than asserts — the numbers are inputs to the #320
|
||||||
// decision, and an assertion here would be this test inventing the bar.
|
// decision, and an assertion here would be this test inventing the bar.
|
||||||
// The one thing worth failing on is a harness fault: if every single case
|
// The one thing worth failing on is a harness fault: if every single case
|
||||||
// errors, the run measured infrastructure, not routing, and the report
|
// errors, the run measured infrastructure, not routing, and the report
|
||||||
// must not be mistaken for a score.
|
// must not be mistaken for a score.
|
||||||
for _, rep := range []Report{repLLM, repCascade, repNoThink} {
|
for _, rep := range []Report{repLLM, repCascade} {
|
||||||
if rep.Errors == rep.Total {
|
if rep.Errors == rep.Total {
|
||||||
t.Errorf("%s: all %d cases errored — harness fault, not a measurement", rep.Name, rep.Total)
|
t.Errorf("%s: all %d cases errored — harness fault, not a measurement", rep.Name, rep.Total)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// noThinkCompleter — llm.Client with chat_template_kwargs.enable_thinking
|
|
||||||
// false. A test-local copy rather than a change to internal/llm: whether the
|
|
||||||
// daemon should send it is the open question, and answering it here by adding
|
|
||||||
// the field would prejudge #320.
|
|
||||||
type noThinkCompleter struct {
|
|
||||||
base string
|
|
||||||
http *http.Client
|
|
||||||
}
|
|
||||||
|
|
||||||
func (c *noThinkCompleter) Complete(ctx context.Context, r llm.Req) (string, error) {
|
|
||||||
payload := map[string]any{
|
|
||||||
"messages": []map[string]string{
|
|
||||||
{"role": "system", "content": r.System},
|
|
||||||
{"role": "user", "content": r.User},
|
|
||||||
},
|
|
||||||
"max_tokens": r.MaxTokens,
|
|
||||||
"temperature": 0,
|
|
||||||
"grammar": r.Grammar,
|
|
||||||
"chat_template_kwargs": map[string]any{"enable_thinking": false},
|
|
||||||
}
|
|
||||||
b, err := json.Marshal(payload)
|
|
||||||
if err != nil {
|
|
||||||
return "", err
|
|
||||||
}
|
|
||||||
req, err := http.NewRequestWithContext(ctx, "POST", c.base+"/v1/chat/completions", bytes.NewReader(b))
|
|
||||||
if err != nil {
|
|
||||||
return "", err
|
|
||||||
}
|
|
||||||
req.Header.Set("Content-Type", "application/json")
|
|
||||||
resp, err := c.http.Do(req)
|
|
||||||
if err != nil {
|
|
||||||
return "", err
|
|
||||||
}
|
|
||||||
defer resp.Body.Close()
|
|
||||||
if resp.StatusCode != 200 {
|
|
||||||
return "", fmt.Errorf("status %d", resp.StatusCode)
|
|
||||||
}
|
|
||||||
var out struct {
|
|
||||||
Choices []struct {
|
|
||||||
Message struct {
|
|
||||||
Content string `json:"content"`
|
|
||||||
ReasoningContent string `json:"reasoning_content"`
|
|
||||||
} `json:"message"`
|
|
||||||
} `json:"choices"`
|
|
||||||
}
|
|
||||||
if err := json.NewDecoder(resp.Body).Decode(&out); err != nil {
|
|
||||||
return "", err
|
|
||||||
}
|
|
||||||
if len(out.Choices) == 0 {
|
|
||||||
return "", fmt.Errorf("no choices")
|
|
||||||
}
|
|
||||||
m := out.Choices[0].Message
|
|
||||||
if m.Content != "" {
|
|
||||||
return m.Content, nil
|
|
||||||
}
|
|
||||||
return m.ReasoningContent, nil
|
|
||||||
}
|
|
||||||
|
|
||||||
func ping(ctx context.Context, c *llm.Client) error {
|
func ping(ctx context.Context, c *llm.Client) error {
|
||||||
ctx, cancel := context.WithTimeout(ctx, 90*time.Second)
|
ctx, cancel := context.WithTimeout(ctx, 90*time.Second)
|
||||||
defer cancel()
|
defer cancel()
|
||||||
|
|||||||
@@ -22,6 +22,7 @@
|
|||||||
{ "id": "ru-query-011", "utterance": "почему сервер тормозит", "lang": "ru", "intent": "query", "tags": ["homelab", "hard"], "note": "diagnostic question, not a chat opener" },
|
{ "id": "ru-query-011", "utterance": "почему сервер тормозит", "lang": "ru", "intent": "query", "tags": ["homelab", "hard"], "note": "diagnostic question, not a chat opener" },
|
||||||
{ "id": "ru-query-012", "utterance": "какие заметки я оставил про полив", "lang": "ru", "intent": "query", "tags": ["recall"] },
|
{ "id": "ru-query-012", "utterance": "какие заметки я оставил про полив", "lang": "ru", "intent": "query", "tags": ["recall"] },
|
||||||
{ "id": "ru-query-013", "utterance": "во сколько у меня встреча", "lang": "ru", "intent": "query", "tags": ["calendar"] },
|
{ "id": "ru-query-013", "utterance": "во сколько у меня встреча", "lang": "ru", "intent": "query", "tags": ["calendar"] },
|
||||||
|
{ "id": "ru-query-019", "utterance": "что у меня стоит в календаре на послезавтра", "lang": "ru", "intent": "query", "tags": ["calendar", "hard"], "note": "agenda, not the clock: the daemon answers this from CalendarEvents inside the query branch, so the clock/date system rule must not swallow it" },
|
||||||
{ "id": "ru-query-014", "utterance": "я успеваю до дедлайна", "lang": "ru", "intent": "query", "tags": ["hard", "no-question-word"] },
|
{ "id": "ru-query-014", "utterance": "я успеваю до дедлайна", "lang": "ru", "intent": "query", "tags": ["hard", "no-question-word"] },
|
||||||
{ "id": "ru-query-015", "utterance": "сколько я прошёл шагов", "lang": "ru", "intent": "query", "tags": ["aggregate"] },
|
{ "id": "ru-query-015", "utterance": "сколько я прошёл шагов", "lang": "ru", "intent": "query", "tags": ["aggregate"] },
|
||||||
{ "id": "ru-query-016", "utterance": "покажи давление за неделю", "lang": "ru", "intent": "query", "tags": ["hard", "imperative"], "note": "imperative form but a read — must not route to act" },
|
{ "id": "ru-query-016", "utterance": "покажи давление за неделю", "lang": "ru", "intent": "query", "tags": ["hard", "imperative"], "note": "imperative form but a read — must not route to act" },
|
||||||
|
|||||||
@@ -3,5 +3,8 @@ package router
|
|||||||
// KnowledgePrompt returns the system prompt for general knowledge questions
|
// KnowledgePrompt returns the system prompt for general knowledge questions
|
||||||
// that the phraser uses when no notes match the query.
|
// that the phraser uses when no notes match the query.
|
||||||
func KnowledgePrompt() string {
|
func KnowledgePrompt() string {
|
||||||
return `Ты — Мавена, персональный ассистент. Ответь кратко из своих знаний. Если не знаешь — скажи "не знаю". Не выдумывай. Respond ONLY with valid JSON: {"response": "...", "mood": "neutral"}.`
|
// No self-introduction here: the shared persona block already says who she
|
||||||
|
// is, and this line used to disagree with it — a different name ("Мавена")
|
||||||
|
// and a masculine noun ("ассистент") in front of a feminine persona.
|
||||||
|
return `Ответь кратко из своих знаний. Если не знаешь — скажи "не знаю". Не выдумывай. Respond ONLY with valid JSON: {"response": "...", "mood": "neutral"}.`
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -11,10 +11,18 @@ func TestKnowledgePrompt(t *testing.T) {
|
|||||||
t.Fatal("KnowledgePrompt returned empty string")
|
t.Fatal("KnowledgePrompt returned empty string")
|
||||||
}
|
}
|
||||||
// Must contain key instructions
|
// Must contain key instructions
|
||||||
checks := []string{"Мавена", "не знаю", "не выдумывай"}
|
checks := []string{"не знаю", "не выдумывай"}
|
||||||
for _, c := range checks {
|
for _, c := range checks {
|
||||||
if !strings.Contains(strings.ToLower(prompt), strings.ToLower(c)) {
|
if !strings.Contains(strings.ToLower(prompt), strings.ToLower(c)) {
|
||||||
t.Errorf("KnowledgePrompt should mention %q", c)
|
t.Errorf("KnowledgePrompt should mention %q", c)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
// Who she is comes from the shared persona block now. This prompt used to
|
||||||
|
// say it too, with a different name and a masculine noun, which is the
|
||||||
|
// drift the block exists to stop.
|
||||||
|
for _, w := range []string{"Мавена", "ассистент"} {
|
||||||
|
if strings.Contains(prompt, w) {
|
||||||
|
t.Errorf("KnowledgePrompt should not introduce her (%q) — the persona block does", w)
|
||||||
|
}
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -44,10 +44,21 @@ ws ::= [ \t\n]*
|
|||||||
// Changed again 31-07-2026: added the "unknown" escape hatch so the model can
|
// Changed again 31-07-2026: added the "unknown" escape hatch so the model can
|
||||||
// admit it cannot route (Vikunja #359).
|
// admit it cannot route (Vikunja #359).
|
||||||
//
|
//
|
||||||
|
// Changed again 31-07-2026: added the clock/calendar rule (Vikunja #374). The
|
||||||
|
// prompt never said which side "который час" or "какое число завтра" belong on,
|
||||||
|
// so the model guessed — `system→query ×4` in every eval run. The rule sits
|
||||||
|
// above the question test on purpose: these utterances all carry a question
|
||||||
|
// word, so a later rule would never be reached. The boundary is what the
|
||||||
|
// daemon can actually answer: only replySystem in cmd/mavend/voice.go owns the
|
||||||
|
// clock and the calendar formatter, while the agenda ("что у меня завтра") is
|
||||||
|
// answered inside the query branch, so that side stays query.
|
||||||
|
//
|
||||||
// The training workspace keeps its own copy of this prompt for relabelling, and
|
// The training workspace keeps its own copy of this prompt for relabelling, and
|
||||||
// `llm/check_prompt_parity.py` there compares the two. That copy is in another
|
// `llm/check_prompt_parity.py` there compares the two. That copy is in another
|
||||||
// repo and was not touched, so parity will fail until it gets the same edits —
|
// repo and was not touched, so parity will fail until it gets the same edits —
|
||||||
// both the rule reorder and the "unknown" wording (Vikunja #362).
|
// both the rule reorder and the "unknown" wording (Vikunja #362) — and now the
|
||||||
|
// clock/calendar rule too. The training workspace is not checked out on this
|
||||||
|
// box at all, so it could not be updated here; #362 still covers the catch-up.
|
||||||
const routeSystem = `Классифицируй ровно одно сообщение пользователя. Верни ОДИН JSON-массив действий.
|
const routeSystem = `Классифицируй ровно одно сообщение пользователя. Верни ОДИН JSON-массив действий.
|
||||||
|
|
||||||
Ровно одно намерение: fact, reminder, note, query, act, chat, system.
|
Ровно одно намерение: fact, reminder, note, query, act, chat, system.
|
||||||
@@ -56,19 +67,21 @@ const routeSystem = `Классифицируй ровно одно сообще
|
|||||||
Классифицируй по цели пользователя. Порядок решения:
|
Классифицируй по цели пользователя. Порядок решения:
|
||||||
1. Хочет напоминание в будущем → reminder
|
1. Хочет напоминание в будущем → reminder
|
||||||
2. Явно просит сохранить информацию → note
|
2. Явно просит сохранить информацию → note
|
||||||
3. Задаёт вопрос: есть вопросительное слово (сколько, что, какой, когда, где, кто, почему, как) или знак «?» → query
|
3. Спрашивает только «который час» / «какое число» / «какой день недели» — сами часы или календарная дата, без своих данных → system
|
||||||
4. Хочет получить информацию, в том числе о своих же данных → query
|
4. Задаёт вопрос: есть вопросительное слово (сколько, что, какой, когда, где, кто, почему, как) или знак «?» → query
|
||||||
5. Утверждает: сообщает или обновляет текущее состояние/событие → fact
|
5. Хочет получить информацию, в том числе о своих же данных → query
|
||||||
6. Просит выполнить работу → act
|
6. Утверждает: сообщает или обновляет текущее состояние/событие → fact
|
||||||
7. Про ассистента, настройки или память → system
|
7. Просит выполнить работу → act
|
||||||
8. Реплика — обрывок или указание на неназванное («это», «то», «потом»), и без него непонятно, что именно нужно сделать → unknown
|
8. Про ассистента, настройки или память → system
|
||||||
9. Иначе → chat
|
9. Реплика — обрывок или указание на неназванное («это», «то», «потом»), и без него непонятно, что именно нужно сделать → unknown
|
||||||
|
10. Иначе → chat
|
||||||
|
|
||||||
Различия:
|
Различия:
|
||||||
- note — сохранить информацию, без напоминания. text = суть.
|
- note — сохранить информацию, без напоминания. text = суть.
|
||||||
- reminder — уведомить позже. text = что напомнить.
|
- reminder — уведомить позже. text = что напомнить.
|
||||||
- fact — неявное обновление: пользователь сообщает, что что-то в мире изменилось (текущее/изменённое состояние, случившееся событие). key/value.
|
- fact — неявное обновление: пользователь сообщает, что что-то в мире изменилось (текущее/изменённое состояние, случившееся событие). key/value.
|
||||||
- unknown — редкий случай. Ставь его, только если в самой реплике нет ни предмета, ни действия. Короткая, простая или незнакомая тема — это не причина для unknown: приветствие и болтовня — это chat, вопрос на любую тему — это query, просьба сделать что-то названное — это act.
|
- unknown — редкий случай. Ставь его, только если в самой реплике нет ни предмета, ни действия. Короткая, простая или незнакомая тема — это не причина для unknown: приветствие и болтовня — это chat, вопрос на любую тему — это query, просьба сделать что-то названное — это act.
|
||||||
|
- system против query — часы и календарная дата сами по себе (сколько времени, какое число, какой день недели — можно и про завтра, и про другой город) — это system. А что записано в календаре или в памяти («что у меня завтра», «какие есть напоминания») — это query. Если в реплике есть просьба (напомни, запиши, сделай), то названное время — просто деталь просьбы, и это не system.
|
||||||
- query против fact — решает форма реплики, а не тема. Вопрос о состоянии — это query, даже если названо то же самое, что бывает в fact. Только утверждение — это fact.
|
- query против fact — решает форма реплики, а не тема. Вопрос о состоянии — это query, даже если названо то же самое, что бывает в fact. Только утверждение — это fact.
|
||||||
|
|
||||||
Примеры:
|
Примеры:
|
||||||
@@ -82,6 +95,8 @@ const routeSystem = `Классифицируй ровно одно сообще
|
|||||||
"что такое docker?" → {"intent":"query","text":"что такое docker"}
|
"что такое docker?" → {"intent":"query","text":"что такое docker"}
|
||||||
"напиши письмо" → {"intent":"act","verb":"написать письмо"}
|
"напиши письмо" → {"intent":"act","verb":"написать письмо"}
|
||||||
"очисти память" → {"intent":"system"}
|
"очисти память" → {"intent":"system"}
|
||||||
|
"который час?" → {"intent":"system"}
|
||||||
|
"какое число завтра?" → {"intent":"system"}
|
||||||
"привет" → {"intent":"chat","text":"привет"}
|
"привет" → {"intent":"chat","text":"привет"}
|
||||||
"сделай это" → {"intent":"unknown"}
|
"сделай это" → {"intent":"unknown"}
|
||||||
"ну это" → {"intent":"unknown"}
|
"ну это" → {"intent":"unknown"}
|
||||||
|
|||||||
@@ -33,6 +33,7 @@ const (
|
|||||||
type onnxEmbedder struct {
|
type onnxEmbedder struct {
|
||||||
tokenizer *unigramTokenizer
|
tokenizer *unigramTokenizer
|
||||||
session *ort.DynamicSession[int64, float32]
|
session *ort.DynamicSession[int64, float32]
|
||||||
|
id string
|
||||||
}
|
}
|
||||||
|
|
||||||
func NewONNXEmbedder(modelPath, tokenizerPath, libPath string) (*onnxEmbedder, error) {
|
func NewONNXEmbedder(modelPath, tokenizerPath, libPath string) (*onnxEmbedder, error) {
|
||||||
@@ -58,11 +59,31 @@ func NewONNXEmbedder(modelPath, tokenizerPath, libPath string) (*onnxEmbedder, e
|
|||||||
return &onnxEmbedder{
|
return &onnxEmbedder{
|
||||||
tokenizer: tok,
|
tokenizer: tok,
|
||||||
session: session,
|
session: session,
|
||||||
|
id: modelIDFromPath(modelPath),
|
||||||
}, nil
|
}, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
func (e *onnxEmbedder) Dim() int { return embedDim }
|
func (e *onnxEmbedder) Dim() int { return embedDim }
|
||||||
|
|
||||||
|
// ID names the loaded model for the DB marker (Vikunja #378): the model file's
|
||||||
|
// own name plus the dimension, so pointing the config at another model changes
|
||||||
|
// the string on its own.
|
||||||
|
func (e *onnxEmbedder) ID() string { return e.id }
|
||||||
|
|
||||||
|
// modelIDFromPath turns /opt/.../multilingual-e5-small.onnx into
|
||||||
|
// "multilingual-e5-small@384".
|
||||||
|
func modelIDFromPath(modelPath string) string {
|
||||||
|
name := modelPath
|
||||||
|
if i := strings.LastIndexAny(name, "/\\"); i >= 0 {
|
||||||
|
name = name[i+1:]
|
||||||
|
}
|
||||||
|
name = strings.TrimSuffix(name, ".onnx")
|
||||||
|
if name == "" {
|
||||||
|
name = "onnx"
|
||||||
|
}
|
||||||
|
return fmt.Sprintf("%s@%d", name, embedDim)
|
||||||
|
}
|
||||||
|
|
||||||
// Embed treats the text as a query. The classifier compares one short
|
// Embed treats the text as a query. The classifier compares one short
|
||||||
// utterance to another short seed phrase, so both sides get the same prefix
|
// utterance to another short seed phrase, so both sides get the same prefix
|
||||||
// and the comparison stays fair. The recall path must call EmbedQuery and
|
// and the comparison stays fair. The recall path must call EmbedQuery and
|
||||||
|
|||||||
@@ -449,16 +449,30 @@ func (AnaphoraResolver) Resolve(text string) (ref string, ok bool) {
|
|||||||
return "", false
|
return "", false
|
||||||
}
|
}
|
||||||
|
|
||||||
// ParseCalendarDate detects RU calendar date words in text and returns the
|
// ParseCalendarDate detects RU/EN calendar day words in text and returns
|
||||||
// resolved time (midnight UTC+0 for "сегодня"/"today", next day for "завтра"/"tomorrow").
|
// midnight of that day in now's own time zone. Handles "сегодня", "завтра",
|
||||||
// Returns zero time + false if no match.
|
// "послезавтра", "вчера" (and the English words). Returns zero time + false
|
||||||
|
// if no match.
|
||||||
|
//
|
||||||
|
// "послезавтра" is checked before "завтра" because it contains it.
|
||||||
func ParseCalendarDate(text string, now time.Time) (time.Time, bool) {
|
func ParseCalendarDate(text string, now time.Time) (time.Time, bool) {
|
||||||
lower := strings.ToLower(text)
|
lower := strings.ToLower(text)
|
||||||
if strings.Contains(lower, "сегодня") || strings.Contains(lower, "today") {
|
switch {
|
||||||
return now.Truncate(24 * time.Hour), true
|
case strings.Contains(lower, "сегодня") || strings.Contains(lower, "today"):
|
||||||
}
|
return midnight(now, 0), true
|
||||||
if strings.Contains(lower, "завтра") || strings.Contains(lower, "tomorrow") {
|
case strings.Contains(lower, "послезавтра") || strings.Contains(lower, "day after tomorrow"):
|
||||||
return now.Truncate(24 * time.Hour).Add(24 * time.Hour), true
|
return midnight(now, 2), true
|
||||||
|
case strings.Contains(lower, "завтра") || strings.Contains(lower, "tomorrow"):
|
||||||
|
return midnight(now, 1), true
|
||||||
|
case strings.Contains(lower, "вчера") || strings.Contains(lower, "yesterday"):
|
||||||
|
return midnight(now, -1), true
|
||||||
}
|
}
|
||||||
return time.Time{}, false
|
return time.Time{}, false
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// midnight returns the start of the day that is `days` away from now, in
|
||||||
|
// now's time zone (now.Truncate(24h) would cut on a UTC boundary instead).
|
||||||
|
func midnight(now time.Time, days int) time.Time {
|
||||||
|
y, m, d := now.AddDate(0, 0, days).Date()
|
||||||
|
return time.Date(y, m, d, 0, 0, 0, 0, now.Location())
|
||||||
|
}
|
||||||
|
|||||||
@@ -45,6 +45,9 @@ func TestParseCalendarDate(t *testing.T) {
|
|||||||
{"расписание на завтра", time.Date(2026, 7, 7, 0, 0, 0, 0, time.UTC), true},
|
{"расписание на завтра", time.Date(2026, 7, 7, 0, 0, 0, 0, time.UTC), true},
|
||||||
{"what's today", time.Date(2026, 7, 6, 0, 0, 0, 0, time.UTC), true},
|
{"what's today", time.Date(2026, 7, 6, 0, 0, 0, 0, time.UTC), true},
|
||||||
{"tomorrow plans", time.Date(2026, 7, 7, 0, 0, 0, 0, time.UTC), true},
|
{"tomorrow plans", time.Date(2026, 7, 7, 0, 0, 0, 0, time.UTC), true},
|
||||||
|
{"какое число послезавтра", time.Date(2026, 7, 8, 0, 0, 0, 0, time.UTC), true},
|
||||||
|
{"что было вчера", time.Date(2026, 7, 5, 0, 0, 0, 0, time.UTC), true},
|
||||||
|
{"yesterday plans", time.Date(2026, 7, 5, 0, 0, 0, 0, time.UTC), true},
|
||||||
{"какая погода", time.Time{}, false},
|
{"какая погода", time.Time{}, false},
|
||||||
{"сколько времени", time.Time{}, false},
|
{"сколько времени", time.Time{}, false},
|
||||||
{"", time.Time{}, false},
|
{"", time.Time{}, false},
|
||||||
|
|||||||
@@ -0,0 +1,161 @@
|
|||||||
|
package store
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"encoding/json"
|
||||||
|
"fmt"
|
||||||
|
"time"
|
||||||
|
)
|
||||||
|
|
||||||
|
// EmbedFunc embeds one piece of stored text. The caller passes
|
||||||
|
// router.EmbedPassage — the STORE side of the query/passage asymmetry, which is
|
||||||
|
// the side every vector in the DB was written with. (Passing the query side
|
||||||
|
// would put the stored vectors in the wrong half of the space and quietly halve
|
||||||
|
// recall.) A func instead of an interface keeps this package free of any
|
||||||
|
// dependency on internal/router.
|
||||||
|
type EmbedFunc func(ctx context.Context, text string) ([]float32, error)
|
||||||
|
|
||||||
|
// BackfillResult is what the re-embed run did, for logging.
|
||||||
|
type BackfillResult struct {
|
||||||
|
Skipped bool // marker already matched — nothing to do
|
||||||
|
Notes int // rows rewritten in the notes table
|
||||||
|
Facts int // fact rows rewritten in memory_vectors
|
||||||
|
MemNotes int // note rows rewritten in memory_vectors
|
||||||
|
NoText int // memory_vectors rows with no text in their meta, left alone
|
||||||
|
Took time.Duration
|
||||||
|
}
|
||||||
|
|
||||||
|
// ReembedAll rewrites every stored vector with the currently configured
|
||||||
|
// embedder and then records that embedder as the one that owns the DB.
|
||||||
|
//
|
||||||
|
// Both places a vector lives are rewritten in the same pass: the `notes` table
|
||||||
|
// `embedding` column and the `memory_vectors` rows (notes AND facts). Doing
|
||||||
|
// only one would leave the two indexes disagreeing, which is worse than leaving
|
||||||
|
// both stale.
|
||||||
|
//
|
||||||
|
// Safe to re-run: if the marker already names the current embedder there is
|
||||||
|
// nothing to fix, so it returns immediately with Skipped set.
|
||||||
|
//
|
||||||
|
// Crash safety: everything — every vector and the marker — happens inside one
|
||||||
|
// transaction. If anything fails or the process dies partway, the transaction
|
||||||
|
// rolls back: no vectors changed and no marker written, so the next run does
|
||||||
|
// the whole job again. The marker is never set unless the full rewrite
|
||||||
|
// committed.
|
||||||
|
func (s *Store) ReembedAll(ctx context.Context, currentID string, embed EmbedFunc) (BackfillResult, error) {
|
||||||
|
start := time.Now()
|
||||||
|
var res BackfillResult
|
||||||
|
|
||||||
|
stored, err := s.Meta(ctx, metaKeyEmbedderID)
|
||||||
|
if err != nil {
|
||||||
|
return res, err
|
||||||
|
}
|
||||||
|
if stored == currentID {
|
||||||
|
res.Skipped = true
|
||||||
|
res.Took = time.Since(start)
|
||||||
|
return res, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
tx, err := s.db.BeginTx(ctx, nil)
|
||||||
|
if err != nil {
|
||||||
|
return res, fmt.Errorf("reembed: begin: %w", err)
|
||||||
|
}
|
||||||
|
defer tx.Rollback() // no-op once committed
|
||||||
|
|
||||||
|
// ----- notes table -----
|
||||||
|
type noteRow struct {
|
||||||
|
id int64
|
||||||
|
text string
|
||||||
|
}
|
||||||
|
var notes []noteRow
|
||||||
|
rows, err := tx.QueryContext(ctx, `SELECT id, text FROM notes WHERE text != ''`)
|
||||||
|
if err != nil {
|
||||||
|
return res, fmt.Errorf("reembed: read notes: %w", err)
|
||||||
|
}
|
||||||
|
for rows.Next() {
|
||||||
|
var n noteRow
|
||||||
|
if err := rows.Scan(&n.id, &n.text); err != nil {
|
||||||
|
rows.Close()
|
||||||
|
return res, fmt.Errorf("reembed: note row: %w", err)
|
||||||
|
}
|
||||||
|
notes = append(notes, n)
|
||||||
|
}
|
||||||
|
rows.Close()
|
||||||
|
if err := rows.Err(); err != nil {
|
||||||
|
return res, fmt.Errorf("reembed: notes: %w", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
for _, n := range notes {
|
||||||
|
vec, err := embed(ctx, n.text)
|
||||||
|
if err != nil {
|
||||||
|
return res, fmt.Errorf("reembed: embed note %d: %w", n.id, err)
|
||||||
|
}
|
||||||
|
if _, err := tx.ExecContext(ctx,
|
||||||
|
`UPDATE notes SET embedding = ? WHERE id = ?`, floatsToBlob(vec), n.id); err != nil {
|
||||||
|
return res, fmt.Errorf("reembed: write note %d: %w", n.id, err)
|
||||||
|
}
|
||||||
|
res.Notes++
|
||||||
|
}
|
||||||
|
|
||||||
|
// ----- memory_vectors (the unified index: notes AND facts) -----
|
||||||
|
// The text to re-embed is the one carried in the row's meta blob, which is
|
||||||
|
// exactly the text that was embedded when the row was written.
|
||||||
|
type vecRow struct {
|
||||||
|
id, text, kind string
|
||||||
|
}
|
||||||
|
var vecs []vecRow
|
||||||
|
rows, err = tx.QueryContext(ctx, `SELECT id, meta FROM memory_vectors`)
|
||||||
|
if err != nil {
|
||||||
|
return res, fmt.Errorf("reembed: read memory vectors: %w", err)
|
||||||
|
}
|
||||||
|
for rows.Next() {
|
||||||
|
var id, metaJSON string
|
||||||
|
if err := rows.Scan(&id, &metaJSON); err != nil {
|
||||||
|
rows.Close()
|
||||||
|
return res, fmt.Errorf("reembed: memory row: %w", err)
|
||||||
|
}
|
||||||
|
meta := map[string]string{}
|
||||||
|
if err := json.Unmarshal([]byte(metaJSON), &meta); err != nil {
|
||||||
|
rows.Close()
|
||||||
|
return res, fmt.Errorf("reembed: meta for %q: %w", id, err)
|
||||||
|
}
|
||||||
|
if meta["text"] == "" {
|
||||||
|
res.NoText++
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
vecs = append(vecs, vecRow{id: id, text: meta["text"], kind: meta["type"]})
|
||||||
|
}
|
||||||
|
rows.Close()
|
||||||
|
if err := rows.Err(); err != nil {
|
||||||
|
return res, fmt.Errorf("reembed: memory vectors: %w", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
for _, v := range vecs {
|
||||||
|
vec, err := embed(ctx, v.text)
|
||||||
|
if err != nil {
|
||||||
|
return res, fmt.Errorf("reembed: embed %q: %w", v.id, err)
|
||||||
|
}
|
||||||
|
if _, err := tx.ExecContext(ctx,
|
||||||
|
`UPDATE memory_vectors SET vec = ? WHERE id = ?`, encodeVec(vec), v.id); err != nil {
|
||||||
|
return res, fmt.Errorf("reembed: write %q: %w", v.id, err)
|
||||||
|
}
|
||||||
|
if v.kind == "fact" {
|
||||||
|
res.Facts++
|
||||||
|
} else {
|
||||||
|
res.MemNotes++
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Same transaction as the rewrite, on purpose: the marker can only exist if
|
||||||
|
// every vector above was written.
|
||||||
|
if _, err := tx.ExecContext(ctx,
|
||||||
|
`INSERT INTO meta (key, value) VALUES (?,?)
|
||||||
|
ON CONFLICT(key) DO UPDATE SET value = excluded.value`,
|
||||||
|
metaKeyEmbedderID, currentID); err != nil {
|
||||||
|
return res, fmt.Errorf("reembed: write marker: %w", err)
|
||||||
|
}
|
||||||
|
if err := tx.Commit(); err != nil {
|
||||||
|
return res, fmt.Errorf("reembed: commit: %w", err)
|
||||||
|
}
|
||||||
|
res.Took = time.Since(start)
|
||||||
|
return res, nil
|
||||||
|
}
|
||||||
@@ -0,0 +1,171 @@
|
|||||||
|
package store
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"errors"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
)
|
||||||
|
|
||||||
|
// markerVec is a recognisable vector: nothing in these tests writes it except
|
||||||
|
// the backfill, so finding it proves the row really was rewritten.
|
||||||
|
var markerVec = []float32{9, 9, 9}
|
||||||
|
|
||||||
|
func newEmbedder(calls *int) EmbedFunc {
|
||||||
|
return func(_ context.Context, _ string) ([]float32, error) {
|
||||||
|
*calls++
|
||||||
|
return markerVec, nil
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// seedOldVectors puts one note (notes table + unified index) and one fact
|
||||||
|
// (unified index only) in the DB, both carrying obviously-old vectors.
|
||||||
|
func seedOldVectors(t *testing.T, s *Store) {
|
||||||
|
t.Helper()
|
||||||
|
ctx := context.Background()
|
||||||
|
old := []float32{0.1, 0.2, 0.3}
|
||||||
|
id, err := s.WriteNote(ctx, time.Now(), "молоко в холодильнике", old, "voice")
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("WriteNote: %v", err)
|
||||||
|
}
|
||||||
|
mem := s.VectorMemory()
|
||||||
|
if err := mem.Insert(ctx, "note:1", old, map[string]string{
|
||||||
|
"type": "note", "text": "молоко в холодильнике",
|
||||||
|
}); err != nil {
|
||||||
|
t.Fatalf("Insert note vector: %v", err)
|
||||||
|
}
|
||||||
|
if err := mem.Insert(ctx, "fact:water:1", old, map[string]string{
|
||||||
|
"type": "fact", "text": "я пил воду",
|
||||||
|
}); err != nil {
|
||||||
|
t.Fatalf("Insert fact vector: %v", err)
|
||||||
|
}
|
||||||
|
_ = id
|
||||||
|
}
|
||||||
|
|
||||||
|
func noteVec(t *testing.T, s *Store) []float32 {
|
||||||
|
t.Helper()
|
||||||
|
var blob []byte
|
||||||
|
if err := s.db.QueryRow(`SELECT embedding FROM notes LIMIT 1`).Scan(&blob); err != nil {
|
||||||
|
t.Fatalf("read note embedding: %v", err)
|
||||||
|
}
|
||||||
|
return blobToFloats(blob)
|
||||||
|
}
|
||||||
|
|
||||||
|
func memVec(t *testing.T, s *Store, id string) []float32 {
|
||||||
|
t.Helper()
|
||||||
|
var blob []byte
|
||||||
|
if err := s.db.QueryRow(`SELECT vec FROM memory_vectors WHERE id = ?`, id).Scan(&blob); err != nil {
|
||||||
|
t.Fatalf("read memory vector %s: %v", id, err)
|
||||||
|
}
|
||||||
|
return decodeVec(blob)
|
||||||
|
}
|
||||||
|
|
||||||
|
func sameVec(a, b []float32) bool {
|
||||||
|
if len(a) != len(b) {
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
for i := range a {
|
||||||
|
if a[i] != b[i] {
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
|
||||||
|
// The deployed case: old vectors everywhere, no marker. Every vector in both
|
||||||
|
// places must be rewritten and the marker recorded.
|
||||||
|
func TestReembedAllRewritesEveryVector(t *testing.T) {
|
||||||
|
s := newTestStore(t)
|
||||||
|
ctx := context.Background()
|
||||||
|
seedOldVectors(t, s)
|
||||||
|
|
||||||
|
calls := 0
|
||||||
|
res, err := s.ReembedAll(ctx, "multilingual-e5-small@384", newEmbedder(&calls))
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("ReembedAll: %v", err)
|
||||||
|
}
|
||||||
|
if res.Skipped {
|
||||||
|
t.Fatal("first run should not skip")
|
||||||
|
}
|
||||||
|
if res.Notes != 1 || res.MemNotes != 1 || res.Facts != 1 {
|
||||||
|
t.Fatalf("counts: notes=%d memNotes=%d facts=%d", res.Notes, res.MemNotes, res.Facts)
|
||||||
|
}
|
||||||
|
if calls != 3 {
|
||||||
|
t.Fatalf("embedder called %d times, want 3", calls)
|
||||||
|
}
|
||||||
|
if !sameVec(noteVec(t, s), markerVec) {
|
||||||
|
t.Fatalf("notes table not rewritten: %v", noteVec(t, s))
|
||||||
|
}
|
||||||
|
if !sameVec(memVec(t, s, "note:1"), markerVec) {
|
||||||
|
t.Fatal("unified index note row not rewritten")
|
||||||
|
}
|
||||||
|
if !sameVec(memVec(t, s, "fact:water:1"), markerVec) {
|
||||||
|
t.Fatal("unified index fact row not rewritten")
|
||||||
|
}
|
||||||
|
got, err := s.Meta(ctx, metaKeyEmbedderID)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("Meta: %v", err)
|
||||||
|
}
|
||||||
|
if got != "multilingual-e5-small@384" {
|
||||||
|
t.Fatalf("marker = %q", got)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Re-running must do nothing at all — not a second pass over the same rows.
|
||||||
|
func TestReembedAllSecondRunIsNoop(t *testing.T) {
|
||||||
|
s := newTestStore(t)
|
||||||
|
ctx := context.Background()
|
||||||
|
seedOldVectors(t, s)
|
||||||
|
|
||||||
|
calls := 0
|
||||||
|
if _, err := s.ReembedAll(ctx, "e5@384", newEmbedder(&calls)); err != nil {
|
||||||
|
t.Fatalf("first run: %v", err)
|
||||||
|
}
|
||||||
|
first := calls
|
||||||
|
|
||||||
|
res, err := s.ReembedAll(ctx, "e5@384", newEmbedder(&calls))
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("second run: %v", err)
|
||||||
|
}
|
||||||
|
if !res.Skipped {
|
||||||
|
t.Fatal("second run should report Skipped")
|
||||||
|
}
|
||||||
|
if calls != first {
|
||||||
|
t.Fatalf("second run embedded %d more rows, want 0", calls-first)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// A failure partway must leave the DB exactly as it was: no marker, and the old
|
||||||
|
// vectors still in place (one transaction, rolled back).
|
||||||
|
func TestReembedAllPartialFailureLeavesMarkerUnset(t *testing.T) {
|
||||||
|
s := newTestStore(t)
|
||||||
|
ctx := context.Background()
|
||||||
|
seedOldVectors(t, s)
|
||||||
|
before := noteVec(t, s)
|
||||||
|
|
||||||
|
calls := 0
|
||||||
|
boom := func(_ context.Context, _ string) ([]float32, error) {
|
||||||
|
calls++
|
||||||
|
if calls == 2 {
|
||||||
|
return nil, errors.New("onnx blew up")
|
||||||
|
}
|
||||||
|
return markerVec, nil
|
||||||
|
}
|
||||||
|
if _, err := s.ReembedAll(ctx, "e5@384", boom); err == nil {
|
||||||
|
t.Fatal("expected an error")
|
||||||
|
}
|
||||||
|
got, err := s.Meta(ctx, metaKeyEmbedderID)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("Meta: %v", err)
|
||||||
|
}
|
||||||
|
if got != "" {
|
||||||
|
t.Fatalf("marker was set to %q after a failed run", got)
|
||||||
|
}
|
||||||
|
if !sameVec(noteVec(t, s), before) {
|
||||||
|
t.Fatal("a failed run left a partially rewritten notes table")
|
||||||
|
}
|
||||||
|
// And the mismatch warning must still fire, so the user knows to re-run.
|
||||||
|
if _, mismatch, err := s.CheckEmbedder(ctx, "e5@384"); err != nil || !mismatch {
|
||||||
|
t.Fatalf("CheckEmbedder after failed backfill: mismatch=%v err=%v", mismatch, err)
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -13,11 +13,15 @@ import (
|
|||||||
// unknown = a pending row found stale at startup: the process that started it
|
// unknown = a pending row found stale at startup: the process that started it
|
||||||
// is gone, and the send may or may not have reached the external channel.
|
// is gone, and the send may or may not have reached the external channel.
|
||||||
// Never auto-resolved into sent or failed — that would be guessing.
|
// Never auto-resolved into sent or failed — that would be guessing.
|
||||||
|
// dropped = the routing table deliberately suppressed this one (a care nudge
|
||||||
|
// while you're away). Nothing was sent and nothing went wrong; the row exists
|
||||||
|
// so "she dropped it" and "the rule never fired" don't look the same later.
|
||||||
const (
|
const (
|
||||||
DeliveryPending = "pending"
|
DeliveryPending = "pending"
|
||||||
DeliverySent = "sent"
|
DeliverySent = "sent"
|
||||||
DeliveryFailed = "failed"
|
DeliveryFailed = "failed"
|
||||||
DeliveryUnknown = "unknown"
|
DeliveryUnknown = "unknown"
|
||||||
|
DeliveryDropped = "dropped"
|
||||||
)
|
)
|
||||||
|
|
||||||
// BeginDeliveryAttempt durably records intent to send BEFORE the external
|
// BeginDeliveryAttempt durably records intent to send BEFORE the external
|
||||||
@@ -43,10 +47,11 @@ func (s *Store) BeginDeliveryAttempt(ctx context.Context, kind, rule string, rem
|
|||||||
}
|
}
|
||||||
|
|
||||||
// CompleteDeliveryAttempt records the sink's outcome for a prior
|
// CompleteDeliveryAttempt records the sink's outcome for a prior
|
||||||
// BeginDeliveryAttempt. status is "sent" or "failed" — never "pending" or
|
// BeginDeliveryAttempt. status is "sent", "failed" or "dropped" — never
|
||||||
// "unknown" (those are set only by Begin and reconciliation respectively).
|
// "pending" or "unknown" (those are set only by Begin and reconciliation
|
||||||
|
// respectively).
|
||||||
func (s *Store) CompleteDeliveryAttempt(ctx context.Context, id int64, status string, now time.Time) error {
|
func (s *Store) CompleteDeliveryAttempt(ctx context.Context, id int64, status string, now time.Time) error {
|
||||||
if status != DeliverySent && status != DeliveryFailed {
|
if status != DeliverySent && status != DeliveryFailed && status != DeliveryDropped {
|
||||||
return fmt.Errorf("store: invalid delivery completion status %q", status)
|
return fmt.Errorf("store: invalid delivery completion status %q", status)
|
||||||
}
|
}
|
||||||
_, err := s.db.ExecContext(ctx,
|
_, err := s.db.ExecContext(ctx,
|
||||||
|
|||||||
@@ -0,0 +1,34 @@
|
|||||||
|
package store
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
)
|
||||||
|
|
||||||
|
// TestDroppedDeliveryAttemptRoundTrips — Vikunja #370. A suppressed nudge is
|
||||||
|
// recorded as 'dropped'. The status column has a CHECK constraint, so this
|
||||||
|
// only works if migration #12 widened it; a fake outbox in a unit test would
|
||||||
|
// not catch that.
|
||||||
|
func TestDroppedDeliveryAttemptRoundTrips(t *testing.T) {
|
||||||
|
s := newTestStore(t)
|
||||||
|
ctx := context.Background()
|
||||||
|
now := time.Now()
|
||||||
|
|
||||||
|
id, err := s.BeginDeliveryAttempt(ctx, "nudge", "water", 0, "drop", "abc123", now)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("BeginDeliveryAttempt: %v", err)
|
||||||
|
}
|
||||||
|
if err := s.CompleteDeliveryAttempt(ctx, id, DeliveryDropped, now); err != nil {
|
||||||
|
t.Fatalf("CompleteDeliveryAttempt: %v", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
var status string
|
||||||
|
err = s.db.QueryRowContext(ctx, `SELECT status FROM delivery_attempts WHERE id = ?`, id).Scan(&status)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("read back: %v", err)
|
||||||
|
}
|
||||||
|
if status != DeliveryDropped {
|
||||||
|
t.Fatalf("status: want %q, got %q", DeliveryDropped, status)
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,74 @@
|
|||||||
|
package store
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"fmt"
|
||||||
|
"time"
|
||||||
|
)
|
||||||
|
|
||||||
|
// DialogueSessionRow — one saved follow-up session. Data is the session
|
||||||
|
// encoded by the dialogue package; the store does not look inside it.
|
||||||
|
type DialogueSessionRow struct {
|
||||||
|
ID string
|
||||||
|
Data []byte
|
||||||
|
Ts time.Time
|
||||||
|
TTL time.Duration
|
||||||
|
Expires time.Time
|
||||||
|
}
|
||||||
|
|
||||||
|
// SaveDialogueSession — write (or replace) the session for one dialogue id.
|
||||||
|
// One row per id: a newer turn overwrites the older state.
|
||||||
|
func (s *Store) SaveDialogueSession(ctx context.Context, id string, data []byte, ts time.Time, ttl time.Duration) error {
|
||||||
|
expires := ts.Add(ttl)
|
||||||
|
_, err := s.db.ExecContext(ctx, `
|
||||||
|
INSERT INTO dialogue_sessions (id, data, ts, ttl_ms, expires_ts) VALUES (?, ?, ?, ?, ?)
|
||||||
|
ON CONFLICT(id) DO UPDATE SET data = excluded.data,
|
||||||
|
ts = excluded.ts,
|
||||||
|
ttl_ms = excluded.ttl_ms,
|
||||||
|
expires_ts = excluded.expires_ts`,
|
||||||
|
id, data, ts.UnixMilli(), ttl.Milliseconds(), expires.UnixMilli())
|
||||||
|
if err != nil {
|
||||||
|
return fmt.Errorf("save dialogue session: %w", err)
|
||||||
|
}
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// DeleteDialogueSession — drop one session (ended, or expired).
|
||||||
|
func (s *Store) DeleteDialogueSession(ctx context.Context, id string) error {
|
||||||
|
if _, err := s.db.ExecContext(ctx, `DELETE FROM dialogue_sessions WHERE id = ?`, id); err != nil {
|
||||||
|
return fmt.Errorf("delete dialogue session: %w", err)
|
||||||
|
}
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// LoadDialogueSessions — return the sessions still alive at `now` and delete
|
||||||
|
// the ones that already ran out. An expired session is dead: it never comes
|
||||||
|
// back after a restart.
|
||||||
|
func (s *Store) LoadDialogueSessions(ctx context.Context, now time.Time) ([]DialogueSessionRow, error) {
|
||||||
|
if _, err := s.db.ExecContext(ctx,
|
||||||
|
`DELETE FROM dialogue_sessions WHERE expires_ts <= ?`, now.UnixMilli()); err != nil {
|
||||||
|
return nil, fmt.Errorf("prune dialogue sessions: %w", err)
|
||||||
|
}
|
||||||
|
rows, err := s.db.QueryContext(ctx,
|
||||||
|
`SELECT id, data, ts, ttl_ms, expires_ts FROM dialogue_sessions ORDER BY id`)
|
||||||
|
if err != nil {
|
||||||
|
return nil, fmt.Errorf("load dialogue sessions: %w", err)
|
||||||
|
}
|
||||||
|
defer rows.Close()
|
||||||
|
var out []DialogueSessionRow
|
||||||
|
for rows.Next() {
|
||||||
|
var r DialogueSessionRow
|
||||||
|
var tsMilli, ttlMilli, expMilli int64
|
||||||
|
if err := rows.Scan(&r.ID, &r.Data, &tsMilli, &ttlMilli, &expMilli); err != nil {
|
||||||
|
return nil, fmt.Errorf("scan dialogue session: %w", err)
|
||||||
|
}
|
||||||
|
r.Ts = time.UnixMilli(tsMilli).UTC()
|
||||||
|
r.TTL = time.Duration(ttlMilli) * time.Millisecond
|
||||||
|
r.Expires = time.UnixMilli(expMilli).UTC()
|
||||||
|
out = append(out, r)
|
||||||
|
}
|
||||||
|
if err := rows.Err(); err != nil {
|
||||||
|
return nil, fmt.Errorf("load dialogue sessions: %w", err)
|
||||||
|
}
|
||||||
|
return out, nil
|
||||||
|
}
|
||||||
@@ -0,0 +1,97 @@
|
|||||||
|
package store
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"database/sql"
|
||||||
|
"errors"
|
||||||
|
"fmt"
|
||||||
|
)
|
||||||
|
|
||||||
|
// metaKeyEmbedderID names the embedder that wrote the stored vectors.
|
||||||
|
//
|
||||||
|
// Why one value for the whole DB and not a column on every vector row: the
|
||||||
|
// vectors are only ever rewritten all at once (one backfill re-embeds every
|
||||||
|
// note and fact together), so a per-row marker would hold the same string in
|
||||||
|
// every row and cost a column on two tables for nothing.
|
||||||
|
const metaKeyEmbedderID = "embedder_id"
|
||||||
|
|
||||||
|
// Meta reads a single value from the meta table. Missing key ⇒ empty string.
|
||||||
|
func (s *Store) Meta(ctx context.Context, key string) (string, error) {
|
||||||
|
var v string
|
||||||
|
err := s.db.QueryRowContext(ctx, `SELECT value FROM meta WHERE key = ?`, key).Scan(&v)
|
||||||
|
if errors.Is(err, sql.ErrNoRows) {
|
||||||
|
return "", nil
|
||||||
|
}
|
||||||
|
if err != nil {
|
||||||
|
return "", fmt.Errorf("read meta %s: %w", key, err)
|
||||||
|
}
|
||||||
|
return v, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// SetMeta writes (or overwrites) a single meta value.
|
||||||
|
func (s *Store) SetMeta(ctx context.Context, key, value string) error {
|
||||||
|
_, err := s.db.ExecContext(ctx,
|
||||||
|
`INSERT INTO meta (key, value) VALUES (?,?)
|
||||||
|
ON CONFLICT(key) DO UPDATE SET value = excluded.value`, key, value)
|
||||||
|
if err != nil {
|
||||||
|
return fmt.Errorf("write meta %s: %w", key, err)
|
||||||
|
}
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// EmbedderUnknown is the stored id reported for a DB that already holds
|
||||||
|
// vectors but never recorded who wrote them.
|
||||||
|
const EmbedderUnknown = "unknown (written before this marker existed)"
|
||||||
|
|
||||||
|
// CheckEmbedder compares the embedder now configured against the one that
|
||||||
|
// wrote the stored vectors. Returns the stored id and whether it differs.
|
||||||
|
//
|
||||||
|
// Vectors from two different models live in different spaces, so cosine
|
||||||
|
// between them is noise rather than a low score — and both of our models are
|
||||||
|
// 384-dimensional, so nothing else catches it.
|
||||||
|
//
|
||||||
|
// Three cases, and the middle one is the one that actually matters:
|
||||||
|
//
|
||||||
|
// - marker present ⇒ compare the two ids.
|
||||||
|
// - marker absent but vectors already stored ⇒ this is a DB from before the
|
||||||
|
// marker, so we cannot know who wrote them. Report a mismatch. This is the
|
||||||
|
// real case on the deployed box: those vectors came from the old embedder,
|
||||||
|
// and claiming them for the current one would hide the exact problem the
|
||||||
|
// marker was added to catch.
|
||||||
|
// - marker absent and no vectors ⇒ fresh DB, claim it, nothing to fix.
|
||||||
|
//
|
||||||
|
// On a mismatch the fix is ReembedAll (backfill.go), run explicitly with
|
||||||
|
// `mavend -reembed`. Nothing is re-embedded here: that work is minutes of CPU
|
||||||
|
// on the laptop and must not stall a normal start.
|
||||||
|
func (s *Store) CheckEmbedder(ctx context.Context, currentID string) (stored string, mismatch bool, err error) {
|
||||||
|
stored, err = s.Meta(ctx, metaKeyEmbedderID)
|
||||||
|
if err != nil {
|
||||||
|
return "", false, err
|
||||||
|
}
|
||||||
|
if stored != "" {
|
||||||
|
return stored, stored != currentID, nil
|
||||||
|
}
|
||||||
|
n, err := s.countVectors(ctx)
|
||||||
|
if err != nil {
|
||||||
|
return "", false, err
|
||||||
|
}
|
||||||
|
if n > 0 {
|
||||||
|
return EmbedderUnknown, true, nil
|
||||||
|
}
|
||||||
|
return currentID, false, s.SetMeta(ctx, metaKeyEmbedderID, currentID)
|
||||||
|
}
|
||||||
|
|
||||||
|
// countVectors — how many stored rows carry an embedding. Used only to tell a
|
||||||
|
// fresh DB apart from one that predates the marker.
|
||||||
|
func (s *Store) countVectors(ctx context.Context) (int, error) {
|
||||||
|
var notes, vecs int
|
||||||
|
if err := s.db.QueryRowContext(ctx,
|
||||||
|
`SELECT count(*) FROM notes WHERE embedding IS NOT NULL`).Scan(¬es); err != nil {
|
||||||
|
return 0, fmt.Errorf("count note vectors: %w", err)
|
||||||
|
}
|
||||||
|
if err := s.db.QueryRowContext(ctx,
|
||||||
|
`SELECT count(*) FROM memory_vectors`).Scan(&vecs); err != nil {
|
||||||
|
return 0, fmt.Errorf("count memory vectors: %w", err)
|
||||||
|
}
|
||||||
|
return notes + vecs, nil
|
||||||
|
}
|
||||||
@@ -0,0 +1,100 @@
|
|||||||
|
package store
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
)
|
||||||
|
|
||||||
|
// A fresh DB has no marker yet, so the current embedder is recorded and
|
||||||
|
// nothing is flagged.
|
||||||
|
func TestCheckEmbedderFreshDBRecords(t *testing.T) {
|
||||||
|
s := newTestStore(t)
|
||||||
|
ctx := context.Background()
|
||||||
|
|
||||||
|
stored, mismatch, err := s.CheckEmbedder(ctx, "multilingual-e5-small@384")
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("CheckEmbedder: %v", err)
|
||||||
|
}
|
||||||
|
if mismatch {
|
||||||
|
t.Fatal("fresh DB reported a mismatch")
|
||||||
|
}
|
||||||
|
if stored != "multilingual-e5-small@384" {
|
||||||
|
t.Fatalf("stored = %q", stored)
|
||||||
|
}
|
||||||
|
got, err := s.Meta(ctx, metaKeyEmbedderID)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("Meta: %v", err)
|
||||||
|
}
|
||||||
|
if got != "multilingual-e5-small@384" {
|
||||||
|
t.Fatalf("marker not persisted, got %q", got)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The deployed box: notes were written by the old embedder, before the marker
|
||||||
|
// existed. Claiming them for the current one would hide exactly the problem
|
||||||
|
// the marker is for, so an unmarked DB that already holds vectors is a
|
||||||
|
// mismatch.
|
||||||
|
func TestCheckEmbedderUnmarkedDBWithVectorsIsMismatch(t *testing.T) {
|
||||||
|
s := newTestStore(t)
|
||||||
|
ctx := context.Background()
|
||||||
|
|
||||||
|
if _, err := s.WriteNote(ctx, time.Now(), "молоко в холодильнике", []float32{0.1, 0.2}, "voice"); err != nil {
|
||||||
|
t.Fatalf("WriteNote: %v", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
stored, mismatch, err := s.CheckEmbedder(ctx, "multilingual-e5-small@384")
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("CheckEmbedder: %v", err)
|
||||||
|
}
|
||||||
|
if !mismatch {
|
||||||
|
t.Fatal("an unmarked DB with stored vectors should report a mismatch")
|
||||||
|
}
|
||||||
|
if stored != EmbedderUnknown {
|
||||||
|
t.Fatalf("stored = %q, want %q", stored, EmbedderUnknown)
|
||||||
|
}
|
||||||
|
// It must NOT claim the DB — that would silence the warning on restart.
|
||||||
|
got, err := s.Meta(ctx, metaKeyEmbedderID)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("Meta: %v", err)
|
||||||
|
}
|
||||||
|
if got != "" {
|
||||||
|
t.Fatalf("marker written despite unknown provenance: %q", got)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Both models are 384-dim, so this is the only thing that catches the swap.
|
||||||
|
func TestCheckEmbedderDifferentModelMismatch(t *testing.T) {
|
||||||
|
s := newTestStore(t)
|
||||||
|
ctx := context.Background()
|
||||||
|
|
||||||
|
if err := s.SetMeta(ctx, metaKeyEmbedderID, "paraphrase-multilingual-MiniLM-L12-v2@384"); err != nil {
|
||||||
|
t.Fatalf("SetMeta: %v", err)
|
||||||
|
}
|
||||||
|
stored, mismatch, err := s.CheckEmbedder(ctx, "multilingual-e5-small@384")
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("CheckEmbedder: %v", err)
|
||||||
|
}
|
||||||
|
if !mismatch {
|
||||||
|
t.Fatal("different embedder not detected")
|
||||||
|
}
|
||||||
|
if stored != "paraphrase-multilingual-MiniLM-L12-v2@384" {
|
||||||
|
t.Fatalf("stored = %q", stored)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The same embedder must never raise a false alarm, including on re-check.
|
||||||
|
func TestCheckEmbedderSameModelNoAlarm(t *testing.T) {
|
||||||
|
s := newTestStore(t)
|
||||||
|
ctx := context.Background()
|
||||||
|
|
||||||
|
for i := 0; i < 2; i++ {
|
||||||
|
_, mismatch, err := s.CheckEmbedder(ctx, "multilingual-e5-small@384")
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("CheckEmbedder: %v", err)
|
||||||
|
}
|
||||||
|
if mismatch {
|
||||||
|
t.Fatalf("false alarm on pass %d", i)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -74,6 +74,44 @@ ALTER TABLE reminders ADD COLUMN next_fire_ts INTEGER;`, // #2
|
|||||||
`CREATE INDEX IF NOT EXISTS idx_nudges_snoozed ON nudges (outcome_ts) WHERE outcome = 'snoozed';`, // #8 — SnoozedUntil runs every tick; keep it off a full scan (Vikunja #364)
|
`CREATE INDEX IF NOT EXISTS idx_nudges_snoozed ON nudges (outcome_ts) WHERE outcome = 'snoozed';`, // #8 — SnoozedUntil runs every tick; keep it off a full scan (Vikunja #364)
|
||||||
`ALTER TABLE proposed_routines ADD COLUMN accepted_ts INTEGER;
|
`ALTER TABLE proposed_routines ADD COLUMN accepted_ts INTEGER;
|
||||||
ALTER TABLE proposed_routines ADD COLUMN last_fired_ts INTEGER;`, // #9 — accepted routines keep firing (Vikunja #366): the tick loop needs to know when a routine was accepted and when it last nudged
|
ALTER TABLE proposed_routines ADD COLUMN last_fired_ts INTEGER;`, // #9 — accepted routines keep firing (Vikunja #366): the tick loop needs to know when a routine was accepted and when it last nudged
|
||||||
|
|
||||||
|
`CREATE TABLE IF NOT EXISTS dialogue_sessions (
|
||||||
|
id TEXT PRIMARY KEY,
|
||||||
|
data BLOB NOT NULL,
|
||||||
|
ts INTEGER NOT NULL,
|
||||||
|
ttl_ms INTEGER NOT NULL,
|
||||||
|
expires_ts INTEGER NOT NULL
|
||||||
|
);
|
||||||
|
CREATE INDEX IF NOT EXISTS idx_dialogue_sessions_expires ON dialogue_sessions (expires_ts);`, // #10 — the follow-up session survives a restart (Vikunja #363); small, TTL-pruned table, not a history log
|
||||||
|
|
||||||
|
`CREATE TABLE IF NOT EXISTS meta (
|
||||||
|
key TEXT PRIMARY KEY,
|
||||||
|
value TEXT NOT NULL
|
||||||
|
);`, // #11 — small key/value table for facts about the DB itself; first key is embedder_id (Vikunja #378)
|
||||||
|
|
||||||
|
// #12 — a suppressed nudge gets a 'dropped' row (Vikunja #370). sqlite
|
||||||
|
// can't widen a CHECK constraint in place, so the table is rebuilt; the
|
||||||
|
// index goes with the old table and is recreated. The columns are listed
|
||||||
|
// out rather than `SELECT *` — copying by position would silently shuffle
|
||||||
|
// every row if the old table's column order ever differed from this one.
|
||||||
|
`CREATE TABLE delivery_attempts_v12 (
|
||||||
|
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||||
|
kind TEXT NOT NULL CHECK (kind IN ('nudge','reminder')),
|
||||||
|
rule TEXT NOT NULL DEFAULT '',
|
||||||
|
reminder_id INTEGER NOT NULL DEFAULT 0,
|
||||||
|
channel TEXT NOT NULL,
|
||||||
|
body_hash TEXT NOT NULL,
|
||||||
|
status TEXT NOT NULL DEFAULT 'pending' CHECK (status IN ('pending','sent','failed','unknown','dropped')),
|
||||||
|
created_ts INTEGER NOT NULL,
|
||||||
|
completed_ts INTEGER
|
||||||
|
);
|
||||||
|
INSERT INTO delivery_attempts_v12
|
||||||
|
(id, kind, rule, reminder_id, channel, body_hash, status, created_ts, completed_ts)
|
||||||
|
SELECT id, kind, rule, reminder_id, channel, body_hash, status, created_ts, completed_ts
|
||||||
|
FROM delivery_attempts;
|
||||||
|
DROP TABLE delivery_attempts;
|
||||||
|
ALTER TABLE delivery_attempts_v12 RENAME TO delivery_attempts;
|
||||||
|
CREATE INDEX IF NOT EXISTS idx_delivery_attempts_status ON delivery_attempts (status);`,
|
||||||
}
|
}
|
||||||
|
|
||||||
// migrate applies every migration with a number greater than the DB's current
|
// migrate applies every migration with a number greater than the DB's current
|
||||||
|
|||||||
+79
-4
@@ -1,8 +1,10 @@
|
|||||||
#!/usr/bin/env bash
|
#!/usr/bin/env bash
|
||||||
# Unified script to stop all Maven services.
|
# Unified script to stop all Maven services.
|
||||||
# Usage: ./kill-maven.sh
|
# Usage: ./kill-maven.sh
|
||||||
# - Graceful SIGTERM is attempted first.
|
# - Docker deploy: `docker compose stop` (see why below).
|
||||||
# - If any process lingers, force with SIGKILL.
|
# - Bare-metal / dev run: graceful SIGTERM first, SIGKILL if anything lingers.
|
||||||
|
# Exits non-zero if it cannot confirm everything is stopped. It must never say
|
||||||
|
# "stopped" unless it checked.
|
||||||
|
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
|
|
||||||
@@ -24,6 +26,73 @@ else
|
|||||||
LLM='llama-server.*\.gguf'
|
LLM='llama-server.*\.gguf'
|
||||||
fi
|
fi
|
||||||
|
|
||||||
|
COMPOSE_FILE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/docker-compose.yml"
|
||||||
|
|
||||||
|
# --- containerised deploy ------------------------------------------------
|
||||||
|
# docker-compose.yml does not set `pid: host`, so each container has its own
|
||||||
|
# PID namespace: pkill on the host sees nothing inside them. This script used
|
||||||
|
# to print "all stopped" while every daemon was still happily running. Stop the
|
||||||
|
# containers through compose instead — that actually reaches them.
|
||||||
|
#
|
||||||
|
# running_containers prints the ids of the project's running containers, or
|
||||||
|
# nothing. Empty output plus a non-zero return means "could not ask docker",
|
||||||
|
# which is different from "nothing is running" and is handled below.
|
||||||
|
running_containers() {
|
||||||
|
docker compose -f "$COMPOSE_FILE" ps -q --status running 2>/dev/null
|
||||||
|
}
|
||||||
|
|
||||||
|
DOCKER_OK=0
|
||||||
|
CONTAINERS=""
|
||||||
|
if command -v docker >/dev/null 2>&1 && [ -f "$COMPOSE_FILE" ]; then
|
||||||
|
if CONTAINERS="$(running_containers)"; then
|
||||||
|
DOCKER_OK=1
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
|
||||||
|
if [ "$DOCKER_OK" = 1 ] && [ -n "$CONTAINERS" ]; then
|
||||||
|
echo "--- Maven is running in containers: stopping via docker compose ---"
|
||||||
|
if ! docker compose -f "$COMPOSE_FILE" stop; then
|
||||||
|
echo "ERROR: 'docker compose stop' failed. Containers may still be running." >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
echo "--- Verifying containers are gone ---"
|
||||||
|
LEFT="$(running_containers || true)"
|
||||||
|
if [ -n "$LEFT" ]; then
|
||||||
|
echo "ERROR: containers still running after stop:" >&2
|
||||||
|
docker compose -f "$COMPOSE_FILE" ps >&2 || true
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
echo "All containers stopped."
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
|
|
||||||
|
# --- bare-metal / dev run -----------------------------------------------
|
||||||
|
# pgrep -f matches whole command lines, so a shell that merely mentions
|
||||||
|
# "mavend" (this script's own parent, for one) shows up. Drop ourselves and our
|
||||||
|
# parent, otherwise the SIGKILL sweep can take out the terminal you ran this in.
|
||||||
|
host_pids() {
|
||||||
|
pgrep -f "$PAT|$LLM" | grep -v -e "^$$\$" -e "^$PPID\$" | paste -sd, - || true
|
||||||
|
}
|
||||||
|
HOST_PIDS=$(host_pids)
|
||||||
|
|
||||||
|
if [ -z "$HOST_PIDS" ]; then
|
||||||
|
# Nothing on the host. Whether that means "already down" depends on whether
|
||||||
|
# we managed to ask docker, and the two must not read the same.
|
||||||
|
if [ "$DOCKER_OK" = 1 ]; then
|
||||||
|
# Docker answered and named no running containers, and there is nothing
|
||||||
|
# on the host either. That is a real answer: Maven is already stopped.
|
||||||
|
echo "Nothing to stop: no Maven processes and no running containers."
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
|
# We could not ask docker, so Maven may be alive in a container we cannot
|
||||||
|
# see. Saying "stopped" here is the exact false success this script had.
|
||||||
|
echo "ERROR: no Maven processes on this host, and docker could not be asked." >&2
|
||||||
|
echo " If this is the container deploy it may still be running:" >&2
|
||||||
|
echo " docker compose -f $COMPOSE_FILE stop" >&2
|
||||||
|
echo " Nothing was stopped. Check by hand before assuming Maven is down." >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
echo "--- Sending graceful SIGTERM to Maven services ---"
|
echo "--- Sending graceful SIGTERM to Maven services ---"
|
||||||
pkill -TERM -f "$PAT" || true
|
pkill -TERM -f "$PAT" || true
|
||||||
# mavend's Pdeathsig SIGKILLs its llama-server on exit, but sweep strays too
|
# mavend's Pdeathsig SIGKILLs its llama-server on exit, but sweep strays too
|
||||||
@@ -32,12 +101,18 @@ pkill -TERM -f "$LLM" || true
|
|||||||
|
|
||||||
echo "--- Verifying processes are gone ---"
|
echo "--- Verifying processes are gone ---"
|
||||||
sleep 1
|
sleep 1
|
||||||
PIDS=$(pgrep -d ',' -f "$PAT|$LLM") || PIDS=""
|
PIDS=$(host_pids)
|
||||||
if [ -n "$PIDS" ]; then
|
if [ -n "$PIDS" ]; then
|
||||||
echo "Warning: some processes still alive. PIDs: $PIDS"
|
echo "Warning: some processes still alive. PIDs: $PIDS"
|
||||||
echo "--- Force killing with SIGKILL ---"
|
echo "--- Force killing with SIGKILL ---"
|
||||||
echo "$PIDS" | tr ',' '\n' | xargs -r kill -9
|
echo "$PIDS" | tr ',' '\n' | xargs -r kill -9
|
||||||
|
sleep 1
|
||||||
|
LEFT=$(host_pids)
|
||||||
|
if [ -n "$LEFT" ]; then
|
||||||
|
echo "ERROR: still alive after SIGKILL. PIDs: $LEFT" >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
echo "Done (SIGKILL)."
|
echo "Done (SIGKILL)."
|
||||||
else
|
else
|
||||||
echo "All services gracefully stopped."
|
echo "All services gracefully stopped."
|
||||||
fi
|
fi
|
||||||
|
|||||||
Reference in New Issue
Block a user