Compare commits
69 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 0ca5748699 | |||
| b43bb265b5 | |||
| b9a24334ea | |||
| c97aebf55a | |||
| 891136c65d | |||
| 41c7c13f42 | |||
| a324e8f624 | |||
| 51805e7f35 | |||
| 533f0acda8 | |||
| 4f59ba78c6 | |||
| d0afd9d4f6 | |||
| 6b67e6f3c2 | |||
| 742b2ad1d7 | |||
| b300ac5c70 | |||
| 13e5170e9e | |||
| 0b90952e55 | |||
| aa8f5b2ee2 | |||
| d7cdcb63bd | |||
| ddb658ffbb | |||
| c7dadc97d9 | |||
| c9d88c152e | |||
| 1890ff5d5d | |||
| 0110e9bc8c | |||
| 50ca8c8b5a | |||
| de09471421 | |||
| d65c16a567 | |||
| 062d4252ef | |||
| 2c27e2ce1f | |||
| ccc5cba2a3 | |||
| 89d83c0b11 | |||
| a97f554802 | |||
| f4de2fc5e1 | |||
| eef5d4da4f | |||
| 09f1696fce | |||
| 80f7322294 | |||
| fa5aebfbe4 | |||
| 59cec63da1 | |||
| 02e8786695 | |||
| 0272dc9d89 | |||
| 2ad7635501 | |||
| 9949b309b1 | |||
| a788ca3915 | |||
| 62d47d28ac | |||
| e9ff2c4912 | |||
| 3dbf67f8f9 | |||
| 84ba217892 | |||
| f179ae2fde | |||
| d00929ac0b | |||
| b6f47fbeb6 | |||
| 2e9b9ec1cf | |||
| bfb57c3148 | |||
| d1f6f6355f | |||
| 92ecb691de | |||
| 4282f6b9a9 | |||
| 7bb9f9be06 | |||
| 1e47eaca5a | |||
| 892330eb84 | |||
| 9a3bcd7c46 | |||
| 98ee701e03 | |||
| 04c1088088 | |||
| 07c191d8b8 | |||
| c668310b3e | |||
| 1bd2acdc2a | |||
| 15e5dd8eaa | |||
| 10cf6f525c | |||
| e2210f6844 | |||
| 214a4032cf | |||
| ee3e6a9eaf | |||
| 8acb8a97c6 |
@@ -7,9 +7,20 @@ talking over unix sockets; one resident small model for routing + phrasing; whis
|
||||
Deploy target is a Ryzen laptop (homesrv) with Vulkan offload to the Vega iGPU (`n_gpu_layers: 99`,
|
||||
compose passes `/dev/dri` + the render gid) — the resident model stays ≤1.7B either way.
|
||||
|
||||
**Resident model:** currently **Qwen3.5-0.8B** (`Q4_K_M`), the smallest checkpoint in the gguf
|
||||
library, picked for CPU/iGPU latency. The **target** is the locally CPT'd **Qwen3-1.7B**; that
|
||||
training is still in flight (Vikunja #122), so no such gguf exists yet. Model files live in
|
||||
**Resident model:** currently **Qwen3-1.7B** (`UD-Q4_K_XL`), stock — not yet the CPT'd one.
|
||||
It replaced Qwen3.5-0.8B on 2026-07-31 because it measured better on both fixtures we have:
|
||||
67.5% vs 59.7% intent-only on the 77-case RU routing fixture, and 20/27 vs 11-17/27 on the
|
||||
talk fixture. See `MODEL-BAKEOFF-31-07-2026.md`. It is a Thinking variant, so `n_ctx` is 4096
|
||||
— reasoning tokens need the room, and 4096 is what the scores above were measured at.
|
||||
|
||||
The **target** is still the locally CPT'd **Qwen3-1.7B** (Vikunja #122, training in flight).
|
||||
Stock already speaks good Russian; what it gets wrong is the persona — it writes `я рад`,
|
||||
masculine, where Maven needs `рада`. That is what the CPT is for.
|
||||
|
||||
**Do not bother with sub-500M models.** LFM2.5-230M and 350M were measured on 2026-07-31 and
|
||||
both are unusable in Russian: the 350M routes at 5.2% (worse than guessing) and answers
|
||||
"столица Франции?" with the invented non-word "Сторзит"; the 230M replies to Russian in
|
||||
Spanish. Their strong published IFEval/BFCL numbers are English-only. Model files live in
|
||||
`/mnt/hdd1/llms`, bind-mounted to `/opt/maven/models/llm` — which **shadows** the repo's
|
||||
`models/llm/`, so the LFM2.5 gguf sitting there is not loaded by anything. Swapping the resident
|
||||
model is a one-line change to `phraser.model_path` in `deploy/mavend.json`.
|
||||
@@ -58,19 +69,30 @@ protocol; the config in `deploy/mavend.json` (with `${VAR}` env expansion from g
|
||||
|
||||
## Routing — read this before touching the router
|
||||
|
||||
`internal/router/` has TWO layered engines and the committed default is an **interim
|
||||
stopgap, not the intended design** (see memory `routing-architecture-target`):
|
||||
`internal/router/` has TWO layered engines. **The LLM router is now the default and it is
|
||||
on in deploy** — this section used to say it was wired `nil`, which stopped being true on
|
||||
2026-07-31.
|
||||
|
||||
- **Target (REARCH.md):** LLM-as-router. One resident Qwen3-1.7B (`llmrouter.go`) emits
|
||||
GBNF-constrained structured JSON, and the SAME model phrases replies. Embedder is demoted
|
||||
from a routing gate to a RAG hint.
|
||||
- **Current stopgap:** `llmrouter` is wired `nil` (around `voice.go`), so the
|
||||
`classifier.go` + `embedder.go` nearest-neighbour cascade actually runs. It routes by
|
||||
similarity to frozen seed phrases — the known cause of weak RU query handling.
|
||||
- **LLM router (the intended design, REARCH.md):** the resident Qwen3-1.7B (`llmrouter.go`)
|
||||
emits GBNF-constrained structured JSON, and the SAME model phrases replies. Embedder is
|
||||
demoted from a routing gate to a RAG hint. Wired at `voice.go:214` via
|
||||
`pickLLMRouter(cfg.Voice.UseLLMRouter(), llmClient)`; the flag is `voice.llm_router`
|
||||
(`config.go`), `DefaultLLMRouter` is **on**, and `deploy/mavend.json` sets it `true`.
|
||||
- **Classifier cascade (the failure floor, not dead code):** `classifier.go` +
|
||||
`embedder.go` nearest-neighbour over frozen seed phrases. It runs when the LLM router is
|
||||
off, when there is no llama-server to talk to (`pickLLMRouter` logs that and degrades),
|
||||
and on any per-turn LLM error. Do not delete it — routing by seed similarity is the known
|
||||
cause of weak RU query handling, but a turn must never break on the model.
|
||||
|
||||
Cascade order: `stage0.go` exact-match fast-path → LLM router (when non-nil) → classifier
|
||||
fallback. Any LLM error falls through to the classifier so a turn never breaks on the model.
|
||||
|
||||
Measured on the 77-case RU fixture (`MODEL-BAKEOFF-31-07-2026.md`): the classifier scores
|
||||
36.8% full accuracy at p50 31ms; Qwen3-1.7B scores 67.5% intent-only / 72.7% through the
|
||||
cascade at p50 ≈2.7s. Accuracy roughly doubled, latency is ~90× worse, and that trade was
|
||||
accepted deliberately. Still open: `Confidence: 1.0` is hardcoded in `llmrouter.go`, so the
|
||||
LLM path never asks for clarification (6/6 refusal cases missed) — Vikunja #359.
|
||||
|
||||
## LLM output contract
|
||||
|
||||
All phrasing paths emit `{"response":"...","mood":"..."}` (parsed in `replier_llm.go` and
|
||||
@@ -82,8 +104,26 @@ workspace enforces that the Go and relabelling prompts remain identical.
|
||||
|
||||
## Non-goals (hard constraints)
|
||||
|
||||
Never phones home. Not a nag, not autonomous. Maven's persona is **feminine** — Russian
|
||||
self-reference must use feminine forms (the user is male; see memory `maven-persona-gender`).
|
||||
Not a nag, not autonomous. Maven's persona is **feminine** — Russian
|
||||
self-reference must use feminine forms — `рада`, not `рад`; `поняла`, not `понял`. The owner
|
||||
is male and is addressed informally: "ты", singular, never "вы"/"ваш" and never "он"/"его"
|
||||
(she talks TO him, not about him). Pet names ("милый", "дорогой") are forbidden; his name
|
||||
("Ками") is not. The eval enforces this: `CheckAddress`, `CheckFeminine` and `CheckCringe` in
|
||||
`internal/phraser/eval/checks.go`, scored by `make eval-phrasing`.
|
||||
|
||||
**"Never phones home" is DEPRECATED** (owner's call, 2026-07-31). It used to be a hard
|
||||
constraint and it is not one any more: a 0.8B — and a 1.7B — does not know enough to answer
|
||||
world questions, so she needs to read external sources. What replaces it:
|
||||
|
||||
- **No telemetry, no cloud model, no third-party account.** That part never changes. Nothing
|
||||
about Maven is reported to anyone, and inference stays on the box.
|
||||
- **Local sources first.** Kiwix ZIMs on homesrv (Wikipedia, ifixit) before anything on the
|
||||
network. Reading beats recalling for a small model, and a local read costs nothing.
|
||||
- **External search is allowed and off unless configured**, like the weather and telegram
|
||||
capabilities.
|
||||
- **His notes and facts are never search input.** Looking up why the sky is blue and sending
|
||||
his stored personal notes to an upstream engine are different acts. Only the utterance goes
|
||||
out, never the persona block, history, or matched notes.
|
||||
|
||||
## Web UI conventions
|
||||
|
||||
|
||||
@@ -15,7 +15,8 @@
|
||||
|
||||
**Maven** — self-hosted personal assistant. Manages your day, acts on your
|
||||
homelab. One daemon on homesrv (always-on, not the workstation), multiple
|
||||
client surfaces. All local, never phones home.
|
||||
client surfaces. Inference and data stay on the box; she may READ external
|
||||
sources (see Non-goals — "never phones home" is deprecated).
|
||||
|
||||
Primary name is "Maven", with feminine-gendered Russian self-reference
|
||||
("она", "меня", "помогла"). Clients may choose their own UI label. Consistent
|
||||
@@ -35,8 +36,13 @@ Inside boundary — the ones that actually constrain the build:
|
||||
she records. A confident wrong fact is worse than a known gap.
|
||||
- **Not a nag** — she'd rather miss a nudge than be mutable. Shuts up when
|
||||
uncertain. Load-bearing.
|
||||
- **Not a stranger** — runs on your stuff, your model, your data. Never
|
||||
phones home.
|
||||
- **Not a stranger** — runs on your stuff, your model, your data. No
|
||||
telemetry, no cloud model, no third-party account. She may READ external
|
||||
sources to answer world questions (Kiwix first, then optional search); she
|
||||
never reports anything about you to anyone, and your notes and facts are
|
||||
never used as search input. **"Never phones home" as an absolute is
|
||||
deprecated** — owner's call, 2026-07-31: a small model does not know enough
|
||||
to be useful without reading.
|
||||
- **Not a relationship** — mom-tone is a function that makes nudges land, not
|
||||
emotional company. Names the drift a warm small model falls into.
|
||||
|
||||
@@ -458,7 +464,7 @@ decides *insistence*. Both are needed.
|
||||
|
||||
sev ≤ 2 drops on away, sev ≥ 3 holds: a missed water nudge is noise, a missed
|
||||
backup failure isn't. Away-channels (ntfy/telegram) leave the box — the one
|
||||
path that crosses "never phones home," through your own relay. **Minimal
|
||||
path that leaves the box for a person to see, through your own relay. **Minimal
|
||||
body** — "disk low on homesrv," not detail; don't make notifications a
|
||||
shoulder-surf exfil surface.
|
||||
|
||||
|
||||
+126
-4
@@ -1,15 +1,28 @@
|
||||
# Resident model bake-off — 31-07-2026
|
||||
|
||||
**Recommendation: keep Qwen3.5-0.8B.** LFM2.5-1.2B is worse at routing (52.6% vs 60.5%
|
||||
intent accuracy), and the loss is almost entirely Russian (18/61 vs 22/61 RU, while EN is a
|
||||
wash). It is also 2.4× slower. The Thinking variant is far worse again.
|
||||
**Outcome: the resident model is stock Qwen3-1.7B** (`UD-Q4_K_XL`). Two sweeps ran this
|
||||
evening and the second one changed the answer — read to the end before acting on any table
|
||||
here. [Second sweep](#second-sweep-same-evening--five-models-and-a-resident-model-change)
|
||||
is the one that holds.
|
||||
|
||||
## First sweep — LFM2.5-1.2B vs Qwen3.5-0.8B
|
||||
|
||||
**Verdict, scoped to this pair: keep Qwen3.5-0.8B over LFM2.5-1.2B.** LFM2.5-1.2B is worse
|
||||
at routing (52.6% vs 60.5% intent accuracy), and the loss is almost entirely Russian
|
||||
(18/61 vs 22/61 RU, while EN is a wash). It is also 2.4× slower. The Thinking variant is
|
||||
far worse again. This verdict still stands as written — it rejects LFM2.5-1.2B. It is
|
||||
**not** a recommendation to keep 0.8B as the resident model; the second sweep replaced it
|
||||
with Qwen3-1.7B.
|
||||
|
||||
Settles Vikunja **#278 / #250**.
|
||||
|
||||
- Same fixture and scorer as `ROUTING-EVAL-31-07-2026.md`: `internal/router/eval/`
|
||||
(`ru_routing_v1.json`, 76 held-out cases).
|
||||
- Reproduce: `MAVEN_LLM_URL=http://127.0.0.1:<port> make eval-router`
|
||||
(`TestLLMRouterBaseline`). Note: there is no `make eval-models` target.
|
||||
(`TestLLMRouterBaseline`). (This line used to say there is no `make eval-models` target.
|
||||
There is one now — start a server with the gguf you want, then
|
||||
`make eval-models MAVEN_LLM_URL=http://127.0.0.1:<port>`. It runs only the LLM test, since
|
||||
the classifier baselines do not depend on the model.)
|
||||
- All three models served by the same `llama-server` flags — `-c 2048 -ngl 99 -t 6`, only
|
||||
`-m` and `--port` differ. One server at a time on an otherwise idle box, so latencies are
|
||||
real and not contention.
|
||||
@@ -99,3 +112,112 @@ thinking trace costs time without buying accuracy on a short enum classification
|
||||
Routing only. LFM2.5 might still phrase better, and phrasing is the resident model's other
|
||||
job — that needs its own fixture. But routing is the load-bearing path and Maven is
|
||||
Russian-first, so on the evidence here the switch is not worth making.
|
||||
|
||||
---
|
||||
|
||||
# Second sweep, same evening — five models, and a resident-model change
|
||||
|
||||
The sections above compared LFM2.5-1.2B against Qwen3.5-0.8B on routing and concluded
|
||||
"the switch is not worth making". That still holds. This sweep asked a different
|
||||
question — whether a *smaller* model could work, since LFM2.5's published
|
||||
instruction-following scores beat Qwen3.5-0.8B badly — and answered it, plus found a
|
||||
better resident model by accident.
|
||||
|
||||
**Outcome: the resident model is now stock Qwen3-1.7B.** Sub-500M is a dead end.
|
||||
|
||||
## Routing — 77 Russian cases, one run each
|
||||
|
||||
| model | on disk | llm-only (full) | llm-only (intent) | cascade + fallback |
|
||||
|---|---|---|---|---|
|
||||
| LFM2.5-230M-Q8_0 | 246 MB | 23.4% | 33.8% | 36.4% |
|
||||
| LFM2.5-350M-Q8_0 | 379 MB | 2.6% | **5.2%** | 20.8% |
|
||||
| Qwen3.5-0.8B-Q4_K_M | 527 MB | 36.4% | 59.7% | 61.0% |
|
||||
| Qwen3.5-2B-UD-Q4_K_XL | 1.34 GB | 42.9% | 62.3% | 63.6% |
|
||||
| **Qwen3-1.7B-UD-Q4_K_XL (stock)** | 1.13 GB | **44.2%** | **67.5%** | **72.7%** |
|
||||
|
||||
Qwen3-1.7B wins every column, including against a model 20% larger than it.
|
||||
|
||||
## Talk fixture — 27 cases, three runs each, idle box
|
||||
|
||||
| | Qwen3.5-0.8B | Qwen3-1.7B stock |
|
||||
|---|---|---|
|
||||
| composite | 13, 11, 8 | **20, 21, 18** |
|
||||
| address | 21, 18, 18 | **26, 25, 23** |
|
||||
| feminine | 27, 25, 26 | 26, 27, 26 |
|
||||
| lang | 27, 27, 26 | 26, 27, 27 |
|
||||
| ontopic | 16, 19, 19 | **22, 23, 23** |
|
||||
| canned fallbacks | 8, 5, 6 | **0, 2, 0** |
|
||||
|
||||
This also fills the row `TALK-EVAL-31-07-2026.md` had to void for contamination:
|
||||
**600ch/1024tok on Qwen3.5-0.8B scores 13, 11, 8.**
|
||||
|
||||
`address` is the headline. It sat at 18-22 of 27 on the 0.8B no matter how the prompt
|
||||
was worded — the prompt explicitly forbids "вы" and the model writes `вашей`,
|
||||
`подождите`, `делаете` anyway. That was read as "prompting is out of levers", and it
|
||||
was really "0.8B is out of capacity". The 1.7B mostly holds the constraint.
|
||||
|
||||
The fallback column matters too: 5-8 of 27 turns on the 0.8B end in a hardcoded
|
||||
`"не знаю."`, meaning it failed to emit parseable JSON about a quarter of the time.
|
||||
The 1.7B does that 0-2 times.
|
||||
|
||||
## Latency — the long tail is not the Thinking block
|
||||
|
||||
| | p50 | p95 |
|
||||
|---|---|---|
|
||||
| Qwen3.5-0.8B | 2.4s, 2.9s, 2.0s | 17.4s, 17.6s, 17.4s |
|
||||
| Qwen3-1.7B stock | 2.7s, 2.6s, 2.8s | 16.4s, 6.6s, 3.9s |
|
||||
|
||||
p50 is flat across a 2× size difference. The first instinct on seeing the 1.7B's
|
||||
16s p95 was "that is the reasoning trace, cap it" — wrong. The 0.8B's p95 is a
|
||||
consistent 17s and the 1.7B beat it in two of three runs. The tail is shared and
|
||||
lives somewhere else. Do not spend time on `/no_think` on this evidence.
|
||||
|
||||
## Sub-500M: not close, and the benchmarks say otherwise for a reason
|
||||
|
||||
LFM2.5-350M publishes IFEval 76.96 against Qwen3.5-0.8B's 59.94, and BFCLv3 44.11
|
||||
against 35.08 — better at instruction-following and structured output, at 2/3 the
|
||||
size. Those numbers are real and they are **English**. Every benchmark in that
|
||||
table except Multi-IF is English-only.
|
||||
|
||||
In Russian, with a 300-token budget and temperature 0:
|
||||
|
||||
- **350M**, «Столица Франции? Ответь кратко.» → *«Сторзит в Париже.»* — `Сторзит` is
|
||||
not a word; it is invented morphology.
|
||||
- **350M**, asked to read back a reminder → a fortune cookie about being attentive
|
||||
and confident. No reminder in it.
|
||||
- **230M**, «Привет, как дела?» → answered **in Spanish**.
|
||||
|
||||
The 230M beating the 350M six-fold on routing (33.8% vs 5.2%) is the other tell:
|
||||
when the larger sibling collapses like that it is format compliance failing, not
|
||||
reasoning.
|
||||
|
||||
This is a pretraining gap, not a fine-tuning gap. Teaching Russian to a 350M from
|
||||
near-zero is not an afternoon on a Colab, which was the premise worth checking.
|
||||
|
||||
## Why this vindicates the 1.7B CPT
|
||||
|
||||
Stock Qwen3-1.7B, untrained and unprompted, answers all three probes in fluent
|
||||
correct Russian. What it gets wrong is the persona: *«Привет! Я рад, что ты здесь»*
|
||||
— `рад` is masculine and Maven needs `рада`. That is the right kind of remaining
|
||||
problem, and it is exactly what the CPT (Vikunja #122) is for.
|
||||
|
||||
The 1.7B was the correct model choice. What was wrong was treating it as a
|
||||
**blocker**: stock already beats what was deployed, so it ships now and gets
|
||||
swapped again when the CPT lands.
|
||||
|
||||
## Caveats
|
||||
|
||||
- Routing is one run per model, not three. The gaps between families are far larger
|
||||
than the run-to-run spread seen on the talk fixture, but the 2B-vs-1.7B gap (62.3
|
||||
vs 67.5) is not safe to call on one run.
|
||||
- ~~The routing numbers only reach production once the LLM router is wired on. It is
|
||||
still `nil`.~~ **Resolved the same evening:** the LLM router is wired at `voice.go:214`
|
||||
behind `voice.llm_router`, the default is on, and `deploy/mavend.json` sets it `true`.
|
||||
These numbers are the production path now, so the p50 ≈2.7s is a real per-turn cost and
|
||||
not a bench artifact.
|
||||
- ~~`/mnt/hdd1/llms/LFM2.5/Qwen3-1.7B-UD-Q4_K_XL.gguf` is a 293 MB truncated download
|
||||
in the wrong directory.~~ **Deleted 2026-07-31.** The good 1.13 GB copy in `qwen3/` is
|
||||
what `deploy/mavend.json` loads.
|
||||
- Harness: `scratchpad/bakeoff.sh`, one server at a time, health-checked before each
|
||||
run, `/v1/models` recorded per run. Never run two LLM consumers at once — see the
|
||||
contamination note in `TALK-EVAL-31-07-2026.md`.
|
||||
|
||||
@@ -103,14 +103,17 @@ eval-router:
|
||||
eval-recall:
|
||||
MAVEN_ONNX_LIB="$(MAVEN_ONNX_LIB)" $(GO) test -v -count=1 ./internal/memory/recalleval/
|
||||
|
||||
# eval-phrasing -- score nudge phrasing (internal/phraser/eval). Verbose so the
|
||||
# eval-phrasing -- score nudge phrasing AND the conversational paths (chat,
|
||||
# query, general knowledge) in internal/phraser/eval. Verbose so the
|
||||
# report and every generated message land in the terminal. With no environment
|
||||
# it scores the deterministic Stub only, which is what CI runs. Set
|
||||
# MAVEN_LLM_URL to add the resident model:
|
||||
# MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-phrasing
|
||||
# The model run is slow (minutes) -- the timeout is raised to match.
|
||||
# The model run is slow (minutes) -- the timeout is raised to match. It covers
|
||||
# two fixtures now (15 nudges + 27 conversational cases, and the chat replies are
|
||||
# the long ones), hence 90m rather than 40m.
|
||||
eval-phrasing:
|
||||
$(GO) test -v -count=1 -timeout 40m ./internal/phraser/eval/
|
||||
$(GO) test -v -count=1 -timeout 90m ./internal/phraser/eval/
|
||||
|
||||
# eval-models — score ONE llama-server against the same fixture, for the
|
||||
# resident-model bake-off (#278, #250). Start a server with the gguf you want,
|
||||
|
||||
@@ -103,13 +103,45 @@ but a large part of the jump is that failure now degrades into Russian instead o
|
||||
The two remaining failures: one `"..."` recurrence (`routine-stretch`) and one meal nudge
|
||||
that never says food.
|
||||
|
||||
## Tried and reverted: an example-led nudge prompt (#393)
|
||||
|
||||
The idea was that a 0.8B copies examples better than it follows rules, so the nudge prompt
|
||||
was rewritten to lead with five on-topic examples (water, break, pills, morning, service) and
|
||||
the prose rules were compressed to pay for the tokens: 1190 chars down to 986.
|
||||
|
||||
It measured **worse**, three runs each side, same llama-server, same fixture:
|
||||
|
||||
| run | before | after |
|
||||
|---|---|---|
|
||||
| 1 | 12/15 (address 14) | 11/15 (address 13) |
|
||||
| 2 | 13/15 (address 15) | 12/15 (address 15) |
|
||||
| 3 | 14/15 (address 15) | 11/15 (address 12) |
|
||||
|
||||
`feminine` and `hisgender` were 15/15 on all six runs, so they measure nothing here. The
|
||||
regression is all in `address`: 44/45 before, 40/45 after. Formal "вы"/"ваше" and plural
|
||||
imperatives came back, and so did `"..."`.
|
||||
|
||||
Two likely causes, both about the same thing — **examples do not carry a prohibition**. The
|
||||
old prompt spent a whole sentence on «говоришь на "ты", в единственном числе»; the new one
|
||||
demoted that to one item in a long "никогда" list, and the model stopped obeying it. And
|
||||
making the examples on-topic let their *wording* leak: a break case came back as
|
||||
«Вы давно не пили воду. Выпей стакан.» — the water example, verbatim, in the wrong slot.
|
||||
That is exactly the failure the laundry/laptop examples were chosen to avoid.
|
||||
|
||||
Change reverted. What survives is the measurement: a rule the model must obey needs its own
|
||||
sentence, and examples must stay off-topic. Also note the before side alone spans 12–14 of
|
||||
15 — this fixture cannot resolve anything smaller than about three cases.
|
||||
|
||||
## Broken, found, not fixed
|
||||
|
||||
1. **`checkFeminine` only catches half the constraint.** It scans for masculine
|
||||
self-reference and passed 15/15 both runs — but three messages address the *owner* in
|
||||
the feminine: "ты давно не отдыхал**а**", "он не ел". The owner is a man. The check has
|
||||
no second-person gender test, so this scores clean while being exactly the persona
|
||||
failure the constraint exists to prevent. This is the most important gap in the harness.
|
||||
1. ~~**`checkFeminine` only catches half the constraint.**~~ **Fixed** (#381). It scanned for
|
||||
masculine self-reference only, so three messages that addressed the *owner* in the feminine
|
||||
("ты давно не отдыхал**а**") scored clean. There is now a second check, `hisgender`: a
|
||||
feminine past-tense verb (-ла/-лась) in a sentence addressed to him ("ты", "тебе", "твой")
|
||||
fails, unless the verb is hers ("я заметила", "напомнила тебе"). It is a suffix rule, not a
|
||||
parser — see the comment in `checks.go` for what it misses. A fresh 15-case run after adding
|
||||
it scored **12/15** with `hisgender` 15/15; the model did not repeat the feminine address in
|
||||
that sample, and the check is pinned by unit tests on the recorded bad strings instead.
|
||||
2. **Grammar is not checked at all, and it is bad.** `"Он не ел 11 дней"` (it was 11 hours),
|
||||
`"Сонуждились 7 дней"` (not a word), `"Они забыли воду"` (wrong person entirely). Every
|
||||
one of these passes all six checks. The fixture measures properties, not fluency, and at
|
||||
@@ -125,8 +157,7 @@ that never says food.
|
||||
|
||||
## Next steps
|
||||
|
||||
1. **Add a second-person gender check** to `checks.go`. Finding 1 above. Until it exists the
|
||||
feminine column means less than it looks like.
|
||||
1. ~~**Add a second-person gender check**~~ — done, `hisgender` in `checks.go` (#381).
|
||||
2. **Decide whether the fallback should count as a pass.** Right now `Score` cannot tell a
|
||||
model answer from a fallback. Either mark fallback bodies in `PhrasedNudge` or count them
|
||||
in their own column. Without that, any future prompt change can score well by failing
|
||||
|
||||
@@ -90,4 +90,7 @@ later* is the worker + RAG.
|
||||
4. **Deferred work** — larger reasoner, custom Piper voice and other expansions.
|
||||
|
||||
## Non-goals (unchanged)
|
||||
Never phones home. Not a nag. Not autonomous. Feminine-gendered RU self-ref.
|
||||
Not a nag. Not autonomous. Feminine-gendered RU self-ref. No telemetry, no
|
||||
cloud model, no third-party account — but she MAY read external sources to
|
||||
answer world questions (Kiwix first, search optional). "Never phones home" as
|
||||
an absolute is deprecated, owner's call 2026-07-31; see CLAUDE.md § Non-goals.
|
||||
|
||||
@@ -75,6 +75,14 @@ same vector. A note is indexed in both places with the same embedding, so if it
|
||||
`QueryNotes` it fails again here — the branch can only ever return a **fact**. Its comment calls it
|
||||
"additive"; for notes it is not.
|
||||
|
||||
**Fixed (Vikunja #373).** The memory pass now runs *first*, as one search over notes and facts with
|
||||
one gate, so whichever memory is clearly the best match answers — note or fact. The notes-only pass
|
||||
stays behind it for notes the vector index does not hold. No threshold changed, so the set of
|
||||
questions Maven answers is the same; only which memory answers them. The fixture gained two mixed
|
||||
note+fact cases (`ru-mixed-031`, `ru-mixed-032`), which is why the counts below are out of 27
|
||||
answerable cases and not 25: hash recall@1 36.0% (9/25) → 37.0% (10/27), e5 recall@1 72.0% (18/25) →
|
||||
70.4% (19/27) with answered-after-gate 68.0% → 66.7% and false recall unchanged at 1/5.
|
||||
|
||||
### 5. Ranking has no recency or type signal, and the store is not the bottleneck
|
||||
|
||||
`internal/store/notes.go:67` sorts by cosine and uses `ts` only to break an exact float tie, which
|
||||
|
||||
+108
-10
@@ -66,15 +66,113 @@ Three things this run settles:
|
||||
`запиши что…` phrasings toward fact, and that suspicion stands — all five `ru-note-*`
|
||||
cases now land on fact. Tracked as Vikunja #375.
|
||||
|
||||
**Thinking off is the best configuration measured so far**, on both accuracy and latency
|
||||
(Vikunja #376). That is worth understanding before flipping: routing is a short
|
||||
classification into a fixed enum with grammar-constrained output, so there is little to
|
||||
reason about, and the thinking trace mostly gives a small model room to talk itself out of
|
||||
the right answer. Phrasing is a different job and needs measuring separately.
|
||||
The `thinking off` column above read as the best configuration measured so far (Vikunja #376).
|
||||
**It was wrong** — see the controlled re-run below. Ignore that column.
|
||||
|
||||
Still `6 / 6` missed clarify — the router has no way to say "I don't know" (Vikunja #359).
|
||||
That is unchanged by anything here.
|
||||
|
||||
## Thinking off — 31-07-2026, controlled re-run (Vikunja #376)
|
||||
|
||||
The "thinking off wins by 6 points" observation above **does not hold**. It was a measurement
|
||||
artefact, and the earlier table's `thinking off` column should be ignored.
|
||||
|
||||
The thinking-off variant was scored by a hand-rolled HTTP client living in the test file
|
||||
instead of `llm.Client`. That copy did not send `repeat_penalty`, which the real router does
|
||||
send (`routeRepeatPenalty = 1.15`). So the two columns differed on two axes at once, and the
|
||||
one that mattered was the penalty, not the thinking mode.
|
||||
|
||||
Re-measured with everything else held equal — same fixture, same prompt, same grammar, same
|
||||
sampling, same idle box, the three configurations run back to back and never concurrently:
|
||||
|
||||
| | llm-only, thinking on | llm-only, thinking off | cascade+llm |
|
||||
|---|---|---|---|
|
||||
| intent-only accuracy | 59.2% (45/76) | 59.2% (45/76) | 61.8% (47/76) |
|
||||
| full accuracy (intent+slots+gate) | 38.2% (29/76) | 38.2% (29/76) | 57.9% (44/76) |
|
||||
| route errors | 3 | 3 | 0 |
|
||||
| grammar violations | 3 (all 3 route errors) | 3 (same 3 cases) | 0 |
|
||||
| missed clarify | 5 / 6 | 5 / 6 | 5 / 6 |
|
||||
| p50 latency | 836ms | 920ms | 810ms |
|
||||
| p95 latency | 1.41s | 2.00s | 1.31s |
|
||||
|
||||
Thinking off is not just a tie on the headline numbers — it is identical case for case, with
|
||||
the same confusion matrix and the same three unparseable replies. The latency difference is
|
||||
run-to-run noise on one box, and it points the wrong way here.
|
||||
|
||||
The reason is simpler than any accuracy argument: **this llama-server build ignores the
|
||||
request-level thinking switch for this model.** Probed directly against the running server
|
||||
with `chat_template_kwargs.enable_thinking = false`, `chat_template_kwargs.thinking = false`
|
||||
and top-level `reasoning_budget = 0` — all three return a byte-identical answer with the
|
||||
thinking trace still in `reasoning_content`, and the server reports the prompt prefix as
|
||||
cached, meaning the rendered template did not change. There was never anything being turned
|
||||
off, which is also why the numbers match exactly.
|
||||
|
||||
Nothing was defaulted. `internal/llm` still has no `chat_template_kwargs` field, `VoiceConfig`
|
||||
has no thinking flag, and `deploy/mavend.json` is unchanged. The misleading third
|
||||
configuration is removed from `internal/router/eval` so the table it produced cannot be quoted
|
||||
again.
|
||||
|
||||
Two caveats worth saying out loud:
|
||||
|
||||
- **The fixture is 76 cases.** A 6-point difference on 76 cases is roughly 4-5 cases and would
|
||||
not have been worth trusting even if it had reproduced. This one was exactly 0 cases, which
|
||||
is a much easier call.
|
||||
- **This is one server build and one checkpoint** (`b9351`, Qwen3.5-0.8B Q4_K_M). If the
|
||||
#122 checkpoint or a newer llama.cpp does honour the switch, the question reopens — but it
|
||||
reopens as an unmeasured question, not as a 6-point win.
|
||||
|
||||
Phrasing was **not** measured. Whether thinking helps there is still open, and now also blocked
|
||||
on the same "can we even turn it off" question.
|
||||
|
||||
## Clock and calendar rule — 31-07-2026 (Vikunja #374)
|
||||
|
||||
`routeSystem` never said whether "который час" or "какое число завтра" are `system` or
|
||||
`query`, and `system→query ×4` showed up in every run. The rule added says: the clock and the
|
||||
calendar date themselves are `system`; what is *written in* the calendar or in memory
|
||||
("что у меня завтра", "какие есть напоминания") stays `query`; and a time named inside a
|
||||
request ("напомни завтра…") is just a detail of the request, not a reason for `system`.
|
||||
|
||||
That split is not a preference. In `cmd/mavend/voice.go` only `replySystem` owns the clock and
|
||||
the date formatter, so a clock question routed to `query` falls into the embedder + note RAG
|
||||
and answers "не знаю". The agenda, on the other hand, is answered by `ParseCalendarDate` +
|
||||
`CalendarEvents` *inside* the `query` branch, so that side has to stay `query`. The rule sits
|
||||
above the question test because every one of these utterances carries a question word and a
|
||||
later rule would never be reached.
|
||||
|
||||
The fixture is now 77 cases: one calendar-agenda case was added
|
||||
(`ru-query-019` "что у меня стоит в календаре на послезавтра", intent `query`) specifically so
|
||||
an over-broad system rule cannot pass unnoticed. The clock/date cases (`ru-sys-001/002/005`,
|
||||
`en-sys-001`) already existed.
|
||||
|
||||
Three runs, same box, back to back, never concurrently:
|
||||
|
||||
| | baseline | first rule (too broad) | rule as committed |
|
||||
|---|---|---|---|
|
||||
| llm-only intent-only | 59.2% (45/76) | 54.5% (42/77) | 59.7% (46/77) |
|
||||
| llm-only full | 38.2% | 35.1% | 39.0% |
|
||||
| llm-only route errors | 3 | 4 | 5 |
|
||||
| llm-only p50 | 1.09s | 0.91s | 0.93s |
|
||||
| cascade+llm intent-only | 61.8% (47/76) | 58.4% | 62.3% (48/77) |
|
||||
| cascade+llm full | 57.9% | 54.5% | 59.7% |
|
||||
| cascade+llm route errors | 0 | 0 | 0 |
|
||||
| cascade+llm p50 | 0.91s | 0.80s | 1.04s |
|
||||
|
||||
**The targeted bug is fixed and the headline number did not move.** `system→query ×4` is gone
|
||||
in both LLM configurations — the `time` and `date` tags go from 0/2 and 0/2 to 2/2 and 2/2 —
|
||||
but the model then over-applies the rule, and `query→system ×5` plus `reminder→system ×2`
|
||||
appear where they did not exist before. Net accuracy is a wash, inside the noise of a 77-case
|
||||
fixture.
|
||||
|
||||
The first attempt is shown because it is the honest history: it said "спрашивает время, дату
|
||||
или день недели → system" with no scope, which swept up reminders, and it cost 3-5 points. It
|
||||
was tightened once, on the reasoning that a rule capturing "напомни завтра в 7" is simply
|
||||
wrong, and not tuned further. The remaining `query/reminder → system` over-trigger is a new,
|
||||
separate weakness of the sub-1B model and deserves its own task rather than more prompt
|
||||
kneading against a held-out fixture.
|
||||
|
||||
The rule is kept. It is correct about what the daemon can answer, and the failure it replaces
|
||||
was silent ("не знаю" to "который час") while the one it introduces is loud.
|
||||
|
||||
## Findings
|
||||
|
||||
### 1. The resident model does route better — 50.0% vs 36.8%
|
||||
@@ -145,11 +243,11 @@ Note the grammar's `string ::= "\"" ([^"\\] | "\\" .)* "\""` is unbounded, so no
|
||||
|
||||
### 7. Two hypotheses tested and closed
|
||||
|
||||
- **Thinking mode is a non-issue.** Qwen3.5's template defaults `thinking = 1`, so
|
||||
grammar-constrained JSON lands in `reasoning_content` with `content` empty —
|
||||
`llm.Client`'s fallback handles it. A `thinking off` run scored *identically* (18/76,
|
||||
48.7%, same p50). `internal/llm` deliberately does **not** grow a `chat_template_kwargs`
|
||||
field.
|
||||
- **Thinking mode is a non-issue.** Confirmed twice now, the second time properly — see the
|
||||
controlled re-run section. Grammar-constrained JSON lands in `reasoning_content` with
|
||||
`content` empty and `llm.Client`'s fallback handles it; the request-level switch does
|
||||
nothing on this build. `internal/llm` deliberately does **not** grow a
|
||||
`chat_template_kwargs` field.
|
||||
- **Runaway array repetition does not reproduce.** An isolated smoke test with a stripped
|
||||
grammar emitted `{"intent":"reminder"}` until `MaxTokens`; under the real `routeSystem`
|
||||
prompt the few-shot examples anchor it to one object. 2 errors in 76, not 76.
|
||||
|
||||
@@ -0,0 +1,150 @@
|
||||
# Conversational phrasing eval — 31-07-2026
|
||||
|
||||
Every score measured tonight, on the three paths the nudge eval never touched:
|
||||
chat, query-with-notes, and general knowledge.
|
||||
|
||||
**Short version: the plumbing got fixed and the score barely moved.** Grammar and
|
||||
Russian prompts together took the composite from ~9 to ~14 of 27. Everything
|
||||
still failing is the model not knowing things or not holding a constraint, and
|
||||
prompting is out of levers. Settles the measurement half of Vikunja #395 / #398 /
|
||||
#400.
|
||||
|
||||
## How to reproduce
|
||||
|
||||
```sh
|
||||
# llama-server: -c 4096 -ngl 99 -t 6, model /mnt/hdd1/llms/qwen3.5/Qwen3.5-0.8B.Q4_K_M.gguf
|
||||
MAVEN_LLM_URL=http://127.0.0.1:18099 no_proxy=127.0.0.1,localhost \
|
||||
deps/go/go/bin/go test -count=1 -timeout 40m \
|
||||
-run TestLLMTalkBaseline ./internal/phraser/eval/ -v
|
||||
```
|
||||
|
||||
Three runs per configuration, always. The fixture is 27 cases, so one reply
|
||||
changing moves the composite by 3.7 points — a single run cannot tell a real
|
||||
change from sampling noise. This was learned the expensive way: an earlier claim
|
||||
that "one nudge case fails every run" turned out to be three different cases
|
||||
across three runs.
|
||||
|
||||
**Run the box otherwise idle.** See the contamination note at the bottom.
|
||||
|
||||
## Composite, per configuration
|
||||
|
||||
| config | overall /27 | chat /9 | query /9 | knowledge /9 | canned fallbacks |
|
||||
|---|---|---|---|---|---|
|
||||
| baseline, no grammar | 7, 12, 7 | 1, 1, 0 | 2, 4, 2 | 4, 7, 5 | 0, 0, 0 |
|
||||
| + GBNF grammar (#398) | 14, 15, 8 | 1, 3, 0 | 5, 6, 3 | 8, 6, 5 | 0, 0, 0 |
|
||||
| + Russian prompts (#400) | 11, 17, 15 | 1, 5, 3 | 5, 6, 8 | 5, 6, 4 | 0, 0, 0 |
|
||||
| + truncation fix, 1000ch/768tok | 12, 13, 10 | 2, 2, 1 | 7, 7, 5 | 3, 4, 4 | 3, 3, 6 |
|
||||
| + rebalanced, 600ch/1024tok | **void — contaminated** | | | | |
|
||||
|
||||
"Canned fallbacks" counts replies that came back as the hardcoded `"не знаю."`
|
||||
or `"поговорили."`. It is not a check, it is a health signal: those strings mean
|
||||
the phraser gave up, and the eval scores them as ordinary bad replies.
|
||||
|
||||
## Per-check
|
||||
|
||||
| check | no grammar | + grammar | + RU prompts | + truncation fix |
|
||||
|---|---|---|---|---|
|
||||
| nonempty | 27, 27, 27 | 27, 27, 27 | 27, 27, 27 | 27, 27, 27 |
|
||||
| ellipsis | 20, 19, 23 | 27, 27, 27 | 27, 27, 27 | 27, 27, 27 |
|
||||
| lang | 13, 16, 15 | 23, 26, 26 | 25, 26, 25 | 26, 27, 27 |
|
||||
| feminine | — | — | 25, 24, 26 | 25, 25, 27 |
|
||||
| address | — | — | 21, 22, 22 | 22, 21, 22 |
|
||||
| ontopic | — | — | 17, 24, 18 | 17, 19, 14 |
|
||||
|
||||
`nonempty` reading 27/27 everywhere is not good news — it was a broken check.
|
||||
It tested for a non-blank string, so replies of literally `{` and `"15-16"`
|
||||
passed it. Fixed on `overnight/fix-truncation`; it needs a letter now.
|
||||
|
||||
## What each change actually bought
|
||||
|
||||
**GBNF grammar (#398) — the biggest single win.** Qwen3.5-0.8B writes
|
||||
`Thinking Process:` as plain text with no tags, `stripThink` only handles
|
||||
`</think>`, so the JSON never closed and the plain-text fallback shipped the
|
||||
literal reasoning. `ellipsis` went 20→27 and `lang` 13→26. The router had been
|
||||
using a grammar for ages; the phraser asking nicely in the prompt was the
|
||||
oversight.
|
||||
|
||||
**Russian prompts (#400) — modest, plus a large latency win.** Chat 1.3→3.0
|
||||
average, query 4.7→6.3, knowledge 6.3→5.0. All inside the run-to-run spread, so
|
||||
"probably better on the paths it targeted, not provable in three runs". p50
|
||||
latency dropped from ~11.5s to ~2.3s and that part is consistent across all
|
||||
three runs — shorter prompts, and she stopped emitting English reasoning first.
|
||||
|
||||
**Truncation fix — necessary, and did not help the score.** Two real bugs
|
||||
(replies of `{`, and a `nonempty` check that passed them), both fixed, and the
|
||||
composite went nowhere. A complete rambling wrong answer fails the same checks a
|
||||
truncated one did. Worth doing anyway: the daemon was shipping `{` to a
|
||||
text-to-speech voice.
|
||||
|
||||
## The truncation bug, since the cause was counter-intuitive
|
||||
|
||||
The grammar's `string ::= ... {0,400}` rule was the cause, not the token cap.
|
||||
Measured against Qwen3.5-0.8B at three caps — 256, 768 and 2048 — the reply came
|
||||
back **exactly 400 characters every time, cut mid-word** (`"Нужно записать и,"`).
|
||||
|
||||
Then I raised the bound to 1000 while the cap was 768 tokens and made it worse:
|
||||
Russian runs ~1.5 characters per token here, so generation died on the *token*
|
||||
cap instead, mid-object, and the new guard correctly refused it and shipped
|
||||
`"не знаю."` — 3, 3 and 6 fallbacks per run, from zero. **The two limits have to
|
||||
agree.** 600 characters needs ~400 tokens; the cap is 1024.
|
||||
|
||||
## Where the remaining failures live
|
||||
|
||||
`address` is stuck at 21-22 of 27 and `ontopic` at 14-19. Both resist prompting.
|
||||
|
||||
**The prompt now explicitly forbids exactly what she does.** It says never "вы",
|
||||
use the singular — and she writes `вашей`, `подождите`, `делаете`, `хотите`,
|
||||
`напишите`. Telling a 0.8B "never do X" does not work. Same for
|
||||
`feminine`: `я готов`, `я понял`, `я нашел`, `я заметил`, `я сказал`.
|
||||
|
||||
**Some of `ontopic` is the fixture, not the model.** `chat-how-are-you` got
|
||||
`"Привет! Я здесь, чтобы поговорить. Как дела сегодня?"` — a fine reply that
|
||||
fails because `want_any` is `[норм, хорош, порядк, тут, работ]`. It fails in
|
||||
every run, so it inflates the count. The `ontopic` column currently measures the
|
||||
fixture as much as the model. Not fixed yet, deliberately: changing it would
|
||||
break comparability with the runs above.
|
||||
|
||||
**Two replies worth reading, because they are not fixable by prompting:**
|
||||
|
||||
- Thunder and lightning: *"Скорость молнии — 8-10 тысяч километров в секунду, но
|
||||
звук — 300 метров в секунду, что делает молнию громче."* Confidently wrong,
|
||||
and it concludes lightning is *louder* rather than sound being *slower*.
|
||||
- "расскажи обо мне": *"Ты — прекрасное существо, с душой и вниманием… Спасибо за
|
||||
твою улыбку… О тебе — заповедь любви."* Sycophantic filler, zero information,
|
||||
and precisely the "not a relationship" non-goal.
|
||||
- Boiling an egg: `"15-16"` one run, `"1"` another. No unit, wrong number.
|
||||
|
||||
The first argues for reading instead of recalling (#403 — Kiwix retrieval scores
|
||||
8/8 on the same questions given English keywords). The second and third argue
|
||||
for templates on the paths where correctness matters (#392).
|
||||
|
||||
## Contamination note — how the last row got voided
|
||||
|
||||
I started the query-rewrite agent against the same llama-server the sweep was
|
||||
using, and assumed contention would only affect latency. It did not. The
|
||||
knowledge path collapsed to 0 of 9 with eight canned `"не знаю."` replies, p95
|
||||
tripled to 23.7s, and **the report still said "0 errors"**.
|
||||
|
||||
That is Vikunja #397, and it is worse than filed: a merely *busy* server
|
||||
produces a clean-looking report with a third of the fixture silently answering
|
||||
`"не знаю."`. `PhraseChat` and `PhraseQuery` swallow every failure and return a
|
||||
hardcoded string, so infrastructure trouble is indistinguishable from bad
|
||||
phrasing in the score. The talk test guards the *start* and *end* of a run with
|
||||
a model check, which catches a dead server but not a loaded one.
|
||||
|
||||
**Until #397 is fixed, treat any run made on a busy box as void.**
|
||||
|
||||
## Next
|
||||
|
||||
- Re-run 600ch/1024tok clean, to fill the void row.
|
||||
- Score `Qwen3.5-2B-UD-Q4_K_XL` (already at `/mnt/hdd1/llms/qwen3.5/`, never
|
||||
measured) on this fixture and the router fixture. Not the 4B — too big for
|
||||
this box, owner's call.
|
||||
- Newer sub-500M candidates (LFM2.5 200M/300M) are worth a run for routing.
|
||||
Note `MODEL-BAKEOFF-31-07-2026.md` found LFM2.5-**1.2B** worse than
|
||||
Qwen3.5-0.8B at Russian routing and 2.4× slower — but those are a different,
|
||||
older generation, so that result does not predict the small ones.
|
||||
- Fix `chat-how-are-you`'s `want_any`, and re-baseline once, so `ontopic`
|
||||
measures the model.
|
||||
- #397 first if anything, since it decides whether any of the above is
|
||||
trustworthy.
|
||||
+87
-2
@@ -3,6 +3,8 @@ package main
|
||||
import (
|
||||
"context"
|
||||
"log"
|
||||
"math/rand"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/dialogue"
|
||||
@@ -21,8 +23,12 @@ const clarifyTTL = 90 * time.Second
|
||||
// raw utterance, chat and system have nothing to fill in. For those a clarify
|
||||
// decision keeps the canned "не поняла" reply — inventing a question for noise
|
||||
// is worse than admitting she missed it.
|
||||
// A reminder wants BOTH what to remind about and when. Subject first: "напомни
|
||||
// в 11" has a time and nothing to say at 11, and a reminder with no subject is
|
||||
// not worth setting. Order here is the order she asks in — she still only asks
|
||||
// about the first one missing.
|
||||
var wantedSlots = map[router.Intent][]dialogue.Slot{
|
||||
router.IntentReminder: {dialogue.SlotTime},
|
||||
router.IntentReminder: {dialogue.SlotText, dialogue.SlotTime},
|
||||
router.IntentFact: {dialogue.SlotKey},
|
||||
router.IntentAct: {dialogue.SlotFn},
|
||||
}
|
||||
@@ -35,7 +41,8 @@ var wantedSlots = map[router.Intent][]dialogue.Slot{
|
||||
// questions, so there is no gender agreement to get wrong; the feminine
|
||||
// self-reference lives in the reply she gives when she drops the request.
|
||||
var clarifyQuestions = map[dialogue.Slot]string{
|
||||
dialogue.SlotTime: "На когда напомнить?",
|
||||
dialogue.SlotTime: "Когда?",
|
||||
dialogue.SlotText: "О чём напомнить?",
|
||||
dialogue.SlotKey: "Что записать?",
|
||||
dialogue.SlotFn: "Что сделать?",
|
||||
}
|
||||
@@ -45,6 +52,84 @@ var clarifyQuestions = map[dialogue.Slot]string{
|
||||
// landed. Feminine self-reference ("поняла"), as everywhere.
|
||||
const clarifyGaveUp = "Прости, я не поняла. Скажи, пожалуйста, по-другому."
|
||||
|
||||
// clarifyExpiredVariants — his answer came after the TTL, so the parked request
|
||||
// is already gone. Same tone as clarifyGaveUp, different reason: too much time
|
||||
// passed, not "I did not understand". Feminine self-reference ("ждала",
|
||||
// "отпустила"); he is addressed with a plain imperative.
|
||||
//
|
||||
// Five phrasings, not one. This is the line he hears whenever he walks off
|
||||
// mid-request, so it is the line that repeats most — and the same sentence every
|
||||
// time is what makes a house assistant sound like a kiosk. They all carry the
|
||||
// same two facts (the old request is gone; say it again if it still matters),
|
||||
// because the wording may vary and the meaning may not.
|
||||
//
|
||||
// Fixed templates rather than model output, for the same reason as
|
||||
// clarifyQuestions: this text has to be right every time, and it is not worth a
|
||||
// generation to say something this small.
|
||||
var clarifyExpiredVariants = []string{
|
||||
"Прости, я слишком долго ждала ответа и отпустила прошлую просьбу. Если она ещё нужна, скажи заново.",
|
||||
"Кажется, прошлая просьба уже не важна — я её отпустила. Если я ошибаюсь, повтори.",
|
||||
"Ты как-то резко замолчал, и я не стала ждать дальше. Если та просьба ещё нужна, скажи заново.",
|
||||
"Я не дождалась ответа и убрала прошлую просьбу. Повтори, если она всё ещё нужна.",
|
||||
"Столько времени прошло, что я отпустила прошлую просьбу. Скажи заново, если она в силе.",
|
||||
}
|
||||
|
||||
// clarifyExpiredLine picks one of them at random.
|
||||
func clarifyExpiredLine() string {
|
||||
return clarifyExpiredVariants[rand.Intn(len(clarifyExpiredVariants))]
|
||||
}
|
||||
|
||||
// isClarifyExpired reports whether s opens with any of the expiry lines. The
|
||||
// notice is glued in front of this turn's reply (see withNotice), so a caller
|
||||
// checking for it has to match a prefix, not the whole string.
|
||||
func isClarifyExpired(s string) bool {
|
||||
for _, v := range clarifyExpiredVariants {
|
||||
if strings.HasPrefix(s, v) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// trimClarifyExpired strips a leading expiry notice, leaving this turn's actual
|
||||
// reply. "" ⇒ the notice was the whole thing.
|
||||
func trimClarifyExpired(s string) string {
|
||||
for _, v := range clarifyExpiredVariants {
|
||||
if strings.HasPrefix(s, v) {
|
||||
return strings.TrimSpace(strings.TrimPrefix(s, v))
|
||||
}
|
||||
}
|
||||
return strings.TrimSpace(s)
|
||||
}
|
||||
|
||||
// clarifyExpiredNotice returns that line when a parked question had just timed
|
||||
// out, and "" when nothing was parked. Call it right after
|
||||
// resolveClarifyAnswer: a live question is answered there, an expired one is
|
||||
// only reported here — the words themselves still go on to be routed fresh.
|
||||
func (h *reactiveHandler) clarifyExpiredNotice() string {
|
||||
if h.clarifyStore == nil {
|
||||
return ""
|
||||
}
|
||||
if !h.clarifyStore.TakeExpired(voiceDialogueID, h.now()) {
|
||||
return ""
|
||||
}
|
||||
log.Printf("voice: clarify — parked question expired, telling him and routing the words fresh")
|
||||
return clarifyExpiredLine()
|
||||
}
|
||||
|
||||
// withNotice glues the expiry notice in front of this turn's reply. One turn
|
||||
// carries one reply on the wire, so the notice cannot be a message of its own —
|
||||
// but neither the notice nor the fresh answer may be dropped.
|
||||
func withNotice(notice, reply string) string {
|
||||
if notice == "" {
|
||||
return reply
|
||||
}
|
||||
if reply == "" {
|
||||
return notice
|
||||
}
|
||||
return notice + " " + reply
|
||||
}
|
||||
|
||||
// missingFor returns the slots a decision still needs, most important first.
|
||||
// Empty ⇒ there is nothing identifiable to ask about.
|
||||
func missingFor(dec router.Decision) []dialogue.Slot {
|
||||
|
||||
@@ -56,10 +56,13 @@ func TestClarifyQuestionForMissingSlot(t *testing.T) {
|
||||
want string
|
||||
asked bool
|
||||
}{
|
||||
{"reminder without a time", clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме"), "На когда напомнить?", true},
|
||||
{"reminder without a time", clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме"), "Когда?", true},
|
||||
{"fact without a key", clarifyDec(router.IntentFact, router.Slots{Text: "запиши"}, "запиши"), "Что записать?", true},
|
||||
{"act without a fn", clarifyDec(router.IntentAct, router.Slots{Text: "сделай это"}, "сделай это"), "Что сделать?", true},
|
||||
{"reminder that already has a time", clarifyDec(router.IntentReminder, router.Slots{HasTime: true}, "напомни в 11"), "", false},
|
||||
// A time with nothing to say at that time is still half a reminder, so
|
||||
// the subject is what she asks about — not silence.
|
||||
{"reminder that has a time but no subject", clarifyDec(router.IntentReminder, router.Slots{HasTime: true}, "напомни в 11"), "О чём напомнить?", true},
|
||||
{"reminder that has both", clarifyDec(router.IntentReminder, router.Slots{Text: "позвонить маме", HasTime: true}, "напомни в 11 позвонить маме"), "", false},
|
||||
{"chat is never worth a question", clarifyDec(router.IntentChat, router.Slots{Text: "мгм"}, "мгм"), "", false},
|
||||
{"query is never worth a question", clarifyDec(router.IntentQuery, router.Slots{Text: "а"}, "а"), "", false},
|
||||
}
|
||||
@@ -78,7 +81,7 @@ func TestClarifyReminderCompletesOnAnswer(t *testing.T) {
|
||||
h, st, _ := newClarifyHandler(t)
|
||||
|
||||
question, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни позвонить маме"}, "напомни позвонить маме"))
|
||||
if !asked || question != "На когда напомнить?" {
|
||||
if !asked || question != "Когда?" {
|
||||
t.Fatalf("expected the time question, got %q asked=%v", question, asked)
|
||||
}
|
||||
|
||||
@@ -152,7 +155,7 @@ func TestClarifyAsksThreeTimesThenSaysSo(t *testing.T) {
|
||||
if !handled {
|
||||
t.Fatalf("answer %d must be consumed as an answer", i)
|
||||
}
|
||||
if reply != "На когда напомнить?" {
|
||||
if reply != "Когда?" {
|
||||
t.Fatalf("attempt %d should ask again, got %q", i, reply)
|
||||
}
|
||||
if h.clarifyStore.Get(voiceDialogueID, h.now()) == nil {
|
||||
@@ -289,6 +292,36 @@ func TestNoQuestionWhenNothingIsMissing(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestClarifyExpiryIsAnnouncedAndWordsStillRoute — his answer lands after the
|
||||
// TTL: she must say the old request is gone AND still answer the new words.
|
||||
func TestClarifyExpiryIsAnnouncedAndWordsStillRoute(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
h, _, now := newClarifyHandler(t)
|
||||
emb := router.NewHashEmbedder(1024)
|
||||
h.embedder = emb
|
||||
h.router = buildRouter(emb, h.matcher, 0.55, nil)
|
||||
|
||||
if _, asked := h.askClarify(clarifyDec(router.IntentReminder, router.Slots{Text: "напомни"}, "напомни")); !asked {
|
||||
t.Fatal("expected a question")
|
||||
}
|
||||
*now = now.Add(clarifyTTL + time.Second)
|
||||
|
||||
reply := h.handleText(ctx, "как дела")
|
||||
if !isClarifyExpired(reply) {
|
||||
t.Fatalf("expired question must be announced first, got %q", reply)
|
||||
}
|
||||
if trimClarifyExpired(reply) == "" {
|
||||
t.Fatalf("the new words must still be answered, got only the notice: %q", reply)
|
||||
}
|
||||
if h.clarifyStore.Get(voiceDialogueID, h.now()) != nil {
|
||||
t.Fatal("the expired question must be gone")
|
||||
}
|
||||
// The notice is said once, not on every later utterance.
|
||||
if reply := h.handleText(ctx, "как дела"); isClarifyExpired(reply) {
|
||||
t.Fatalf("notice repeated on a later turn: %q", reply)
|
||||
}
|
||||
}
|
||||
|
||||
// TestNoPendingQuestionFallsThrough — with nothing parked, an utterance routes
|
||||
// normally.
|
||||
func TestNoPendingQuestionFallsThrough(t *testing.T) {
|
||||
|
||||
+43
-21
@@ -58,6 +58,7 @@ import (
|
||||
"github.com/kami/maven/internal/delivery/telegramsink"
|
||||
"github.com/kami/maven/internal/ipc"
|
||||
"github.com/kami/maven/internal/loop"
|
||||
"github.com/kami/maven/internal/persona"
|
||||
"github.com/kami/maven/internal/phraser"
|
||||
"github.com/kami/maven/internal/store"
|
||||
"github.com/kami/maven/internal/webauthn"
|
||||
@@ -189,7 +190,9 @@ func (l *lockedAPI) MorningStatus(ctx context.Context) ([]ipc.MorningRoutineStat
|
||||
func run(args []string) error {
|
||||
cfgPath := flag.String("config", defaultConfigPath(), "path to mavend JSON config")
|
||||
wrappedKeyPath := flag.String("wrapped-key-file", "", "path to wrapped encryption key blob (enables cold-start unlock)")
|
||||
reembed := flag.Bool("reembed", false, "re-embed every stored note and fact with the configured embedder, then serve normally (run once after an embedder swap)")
|
||||
flag.CommandLine.Parse(args)
|
||||
reembedOnStart = *reembed
|
||||
cfg, err := config.Load(*cfgPath)
|
||||
if err != nil {
|
||||
return err
|
||||
@@ -269,13 +272,14 @@ func run(args []string) error {
|
||||
phr = phraser.NewStub()
|
||||
if cfg.Phraser != nil {
|
||||
pc := phraser.Config{
|
||||
ModelPath: cfg.Phraser.ModelPath,
|
||||
BinPath: cfg.Phraser.BinPath,
|
||||
Listen: cfg.Phraser.Listen,
|
||||
NGpuLayers: cfg.Phraser.NGpuLayers,
|
||||
NCtx: cfg.Phraser.NCtx,
|
||||
Timeout: time.Duration(cfg.Phraser.Timeout),
|
||||
Persona: personaFromCfg(cfg),
|
||||
ModelPath: cfg.Phraser.ModelPath,
|
||||
BinPath: cfg.Phraser.BinPath,
|
||||
Listen: cfg.Phraser.Listen,
|
||||
NGpuLayers: cfg.Phraser.NGpuLayers,
|
||||
NCtx: cfg.Phraser.NCtx,
|
||||
Timeout: time.Duration(cfg.Phraser.Timeout),
|
||||
LLMNudges: cfg.Phraser.LLMNudges,
|
||||
ContextBlock: contextBlockFn(cfg, time.Now),
|
||||
}
|
||||
if pc.BinPath == "" {
|
||||
pc.BinPath = "llama-server"
|
||||
@@ -439,13 +443,14 @@ func run(args []string) error {
|
||||
phr = phraser.NewStub()
|
||||
if cfg.Phraser != nil {
|
||||
pc := phraser.Config{
|
||||
ModelPath: cfg.Phraser.ModelPath,
|
||||
BinPath: cfg.Phraser.BinPath,
|
||||
Listen: cfg.Phraser.Listen,
|
||||
NGpuLayers: cfg.Phraser.NGpuLayers,
|
||||
NCtx: cfg.Phraser.NCtx,
|
||||
Timeout: time.Duration(cfg.Phraser.Timeout),
|
||||
Persona: personaFromCfg(cfg),
|
||||
ModelPath: cfg.Phraser.ModelPath,
|
||||
BinPath: cfg.Phraser.BinPath,
|
||||
Listen: cfg.Phraser.Listen,
|
||||
NGpuLayers: cfg.Phraser.NGpuLayers,
|
||||
NCtx: cfg.Phraser.NCtx,
|
||||
Timeout: time.Duration(cfg.Phraser.Timeout),
|
||||
LLMNudges: cfg.Phraser.LLMNudges,
|
||||
ContextBlock: contextBlockFn(cfg, time.Now),
|
||||
}
|
||||
if pc.BinPath == "" {
|
||||
pc.BinPath = "llama-server"
|
||||
@@ -599,12 +604,29 @@ func run(args []string) error {
|
||||
return nil
|
||||
}
|
||||
|
||||
// personaFromCfg extracts the voice persona from the config, or returns ""
|
||||
// when voice isn't configured. Used to pass a character prompt into the
|
||||
// LLM phraser without requiring voice to be enabled.
|
||||
func personaFromCfg(cfg *config.Config) string {
|
||||
if cfg.Voice != nil {
|
||||
return cfg.Voice.Persona
|
||||
// personaFacts reads the optional, deployment-specific facts (his name, his
|
||||
// city, the free-text persona string) out of the config. Everything here may
|
||||
// be empty — the context block is correct without any of it.
|
||||
func personaFacts(cfg *config.Config) persona.Facts {
|
||||
f := persona.Facts{
|
||||
// Telegram lives outside the voice block, so it counts either way.
|
||||
Telegram: cfg.Telegram != nil && cfg.Telegram.BotToken != "" && cfg.Telegram.ChatID != "",
|
||||
}
|
||||
return ""
|
||||
if cfg.Voice == nil {
|
||||
return f
|
||||
}
|
||||
f.OwnerName = cfg.Voice.OwnerName
|
||||
f.City = cfg.Voice.City
|
||||
f.Static = cfg.Voice.Persona
|
||||
// Same test wireVoice uses to pick the real provider over the stub.
|
||||
f.Weather = cfg.Voice.Weather != nil && cfg.Voice.Weather.Provider == "open-meteo"
|
||||
f.Tools = len(cfg.Voice.Tools) > 0
|
||||
return f
|
||||
}
|
||||
|
||||
// contextBlockFn returns the per-turn renderer of the shared context block.
|
||||
// Per turn, not once at startup, because the block states the current time.
|
||||
func contextBlockFn(cfg *config.Config, now func() time.Time) func() string {
|
||||
f := personaFacts(cfg)
|
||||
return func() string { return f.Block(now()) }
|
||||
}
|
||||
|
||||
@@ -0,0 +1,158 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"math"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/ipc"
|
||||
"github.com/kami/maven/internal/memory"
|
||||
"github.com/kami/maven/internal/phraser"
|
||||
"github.com/kami/maven/internal/router"
|
||||
"github.com/kami/maven/internal/voice"
|
||||
)
|
||||
|
||||
// fixedEmbedder hands back a vector chosen per text, so a test can say exactly
|
||||
// how close each stored memory is to the question. The real embedders make
|
||||
// scores that are realistic but not controllable, and this test is about the
|
||||
// gate, not about the embedder.
|
||||
type fixedEmbedder struct{ vecs map[string][]float32 }
|
||||
|
||||
func (f *fixedEmbedder) Dim() int { return 4 }
|
||||
func (f *fixedEmbedder) Close() error { return nil }
|
||||
|
||||
func (f *fixedEmbedder) Embed(_ context.Context, text string) ([]float32, error) {
|
||||
v, ok := f.vecs[text]
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("fixedEmbedder: no vector for %q", text)
|
||||
}
|
||||
return v, nil
|
||||
}
|
||||
|
||||
// scoreVec builds a unit vector whose cosine against the query vector
|
||||
// (1,0,0,0) is exactly score.
|
||||
func scoreVec(score float64) []float32 {
|
||||
rest := math.Sqrt(1 - score*score)
|
||||
return []float32{float32(score), float32(rest), 0, 0}
|
||||
}
|
||||
|
||||
// recordingPhraser remembers what the query path handed it to phrase, which is
|
||||
// how the test can tell which pass produced the answer.
|
||||
type recordingPhraser struct {
|
||||
*phraser.Stub
|
||||
notes []string
|
||||
}
|
||||
|
||||
func (r *recordingPhraser) PhraseQuery(ctx context.Context, utterance string, notes []string) (string, error) {
|
||||
r.notes = notes
|
||||
return r.Stub.PhraseQuery(ctx, utterance, notes)
|
||||
}
|
||||
|
||||
// recallCase — one stored memory: its text, how close it is to the question,
|
||||
// whether it is a note or a fact, and whether the notes table holds it too.
|
||||
type recallCase struct {
|
||||
text string
|
||||
score float64
|
||||
kind string
|
||||
}
|
||||
|
||||
// buildRecallHandler stores the given memories and returns a handler whose
|
||||
// query path can be run directly. Notes go into BOTH the notes table and the
|
||||
// vector index, which is what the daemon does (voice.go's IntentNote).
|
||||
func buildRecallHandler(t *testing.T, question string, mems []recallCase) (*reactiveHandler, *recordingPhraser) {
|
||||
t.Helper()
|
||||
ctx := context.Background()
|
||||
st := newTestStore(t)
|
||||
emb := &fixedEmbedder{vecs: map[string][]float32{question: {1, 0, 0, 0}}}
|
||||
mem := memory.NewInMemoryStore()
|
||||
now := time.Now()
|
||||
|
||||
for i, m := range mems {
|
||||
vec := scoreVec(m.score)
|
||||
emb.vecs[m.text] = vec
|
||||
id := fmt.Sprintf("%s:%d", m.kind, i)
|
||||
if m.kind == "note" {
|
||||
if _, err := st.WriteNote(ctx, now, m.text, vec, "tap:voice"); err != nil {
|
||||
t.Fatalf("WriteNote: %v", err)
|
||||
}
|
||||
}
|
||||
if err := mem.Insert(ctx, id, vec, map[string]string{"text": m.text, "type": m.kind}); err != nil {
|
||||
t.Fatalf("memory insert: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
phr := &recordingPhraser{Stub: phraser.NewStub()}
|
||||
h := &reactiveHandler{
|
||||
api: ipc.NewStoreAPI(st),
|
||||
embedder: emb,
|
||||
replier: voice.NewStubReplier(),
|
||||
phraser: phr,
|
||||
now: func() time.Time { return now },
|
||||
memStore: mem,
|
||||
dataStore: st,
|
||||
queryMinScore: 0.55,
|
||||
queryMinMargin: 0.008,
|
||||
weatherProvider: nil,
|
||||
}
|
||||
return h, phr
|
||||
}
|
||||
|
||||
func askQuery(t *testing.T, h *reactiveHandler, question string) string {
|
||||
t.Helper()
|
||||
return h.applyAction(context.Background(), router.Decision{
|
||||
Intent: router.IntentQuery,
|
||||
Utterance: question,
|
||||
})
|
||||
}
|
||||
|
||||
// TestQueryRecallNoteCanWin — the note-recall regression (Vikunja #373). Notes
|
||||
// and facts share one vector index, and a note that clearly beats everything
|
||||
// else must be the answer. Before the fix the memory pass only ran after the
|
||||
// notes-only gate had already rejected the same note at the same score, so only
|
||||
// a fact could ever come back from it.
|
||||
func TestQueryRecallNoteCanWin(t *testing.T) {
|
||||
const q = "где молоко"
|
||||
|
||||
t.Run("a clearly best note answers", func(t *testing.T) {
|
||||
h, phr := buildRecallHandler(t, q, []recallCase{
|
||||
{text: "молоко стоит в холодильнике", score: 0.90, kind: "note"},
|
||||
{text: "выучил пару аккордов", score: 0.50, kind: "note"},
|
||||
})
|
||||
reply := askQuery(t, h, q)
|
||||
if want := "вот что я нашла: молоко стоит в холодильнике"; reply != want {
|
||||
t.Errorf("reply %q, want %q", reply, want)
|
||||
}
|
||||
// One text, the winning memory's — the answer came from the memory
|
||||
// pass, not from handing the phraser every note in the table.
|
||||
if len(phr.notes) != 1 || phr.notes[0] != "молоко стоит в холодильнике" {
|
||||
t.Errorf("phraser got %q, want just the recalled note", phr.notes)
|
||||
}
|
||||
})
|
||||
|
||||
// The other half of "one gate over everything": a fact that matches better
|
||||
// than the best note now answers, instead of losing to a note that only had
|
||||
// to beat other notes.
|
||||
t.Run("the better-matching fact answers", func(t *testing.T) {
|
||||
h, _ := buildRecallHandler(t, q, []recallCase{
|
||||
{text: "молоко стоит в холодильнике", score: 0.80, kind: "note"},
|
||||
{text: "купил молоко в среду", score: 0.95, kind: "fact"},
|
||||
})
|
||||
if reply := askQuery(t, h, q); reply != "купил молоко в среду" {
|
||||
t.Errorf("reply %q, want the fact read back", reply)
|
||||
}
|
||||
})
|
||||
|
||||
// The gate is untouched: two memories this close mean the embedder cannot
|
||||
// tell them apart, and silence still beats a coin flip.
|
||||
t.Run("no clear best stays silent", func(t *testing.T) {
|
||||
h, _ := buildRecallHandler(t, q, []recallCase{
|
||||
{text: "молоко стоит в холодильнике", score: 0.860, kind: "note"},
|
||||
{text: "молоко закончилось", score: 0.858, kind: "note"},
|
||||
})
|
||||
if reply := askQuery(t, h, q); reply != "не знаю." {
|
||||
t.Errorf("reply %q, want silence", reply)
|
||||
}
|
||||
})
|
||||
}
|
||||
+17
-14
@@ -2,21 +2,24 @@ package main
|
||||
|
||||
import "github.com/kami/maven/internal/memory"
|
||||
|
||||
// bestRecall is the read side of the long-term memory store: the top hit's
|
||||
// stored text when it clears the confidence gate. This recalls across BOTH
|
||||
// notes and facts (facts aren't in the notes table, so this is the only path
|
||||
// that can answer "when did I last …?" from a captured fact). A note hit here
|
||||
// is redundant with the notes-RAG path — by design; the two indexes can diverge
|
||||
// once the backend is swapped for a persistent/external store. ok=false when
|
||||
// the hit fails the confidence gate (see memory.Confident: an absolute floor
|
||||
// plus a margin over the runner-up) or carries no text.
|
||||
func bestRecall(results []memory.Result, minScore, minMargin float64) (string, bool) {
|
||||
// bestRecall is the read side of the long-term memory store: the top hit when
|
||||
// it clears the confidence gate. The index holds BOTH notes and facts, and
|
||||
// either can win — the caller looks at the returned hit's meta["type"] to see
|
||||
// which. Facts aren't in the notes table, so this is the only path that can
|
||||
// answer "when did I last …?" from a captured fact.
|
||||
//
|
||||
// The whole hit is returned, not just its text, because "which memory answered"
|
||||
// decides how the answer is said: a note gets phrased in Maven's voice, a fact
|
||||
// is read back as stored.
|
||||
//
|
||||
// ok=false when the hit fails the confidence gate (see memory.Confident: an
|
||||
// absolute floor plus a margin over the runner-up) or carries no text.
|
||||
func bestRecall(results []memory.Result, minScore, minMargin float64) (memory.Result, bool) {
|
||||
if !memory.Confident(results, minScore, minMargin) {
|
||||
return "", false
|
||||
return memory.Result{}, false
|
||||
}
|
||||
text := results[0].Meta["text"]
|
||||
if text == "" {
|
||||
return "", false
|
||||
if results[0].Meta["text"] == "" {
|
||||
return memory.Result{}, false
|
||||
}
|
||||
return text, true
|
||||
return results[0], true
|
||||
}
|
||||
|
||||
@@ -39,8 +39,27 @@ func TestBestRecall(t *testing.T) {
|
||||
if !ok {
|
||||
t.Fatal("clearing hit not returned")
|
||||
}
|
||||
if got != "выпил воды в три часа" {
|
||||
t.Errorf("wrong text: %q", got)
|
||||
if got.Meta["text"] != "выпил воды в три часа" {
|
||||
t.Errorf("wrong text: %q", got.Meta["text"])
|
||||
}
|
||||
if got.Meta["type"] != "fact" {
|
||||
t.Errorf("kind lost: %q", got.Meta["type"])
|
||||
}
|
||||
})
|
||||
|
||||
// The index holds notes and facts together, so a note has to be able to win
|
||||
// it — for a long time it could not (Vikunja #373).
|
||||
t.Run("a note can win", func(t *testing.T) {
|
||||
res := []memory.Result{
|
||||
{Score: 0.86, Meta: map[string]string{"text": "молоко в холодильнике", "type": "note"}},
|
||||
{Score: 0.61, Meta: map[string]string{"text": "выпил воды", "type": "fact"}},
|
||||
}
|
||||
got, ok := bestRecall(res, min, margin)
|
||||
if !ok {
|
||||
t.Fatal("clearly-best note not returned")
|
||||
}
|
||||
if got.Meta["type"] != "note" || got.Meta["text"] != "молоко в холодильнике" {
|
||||
t.Errorf("got %v, want the note", got.Meta)
|
||||
}
|
||||
})
|
||||
|
||||
|
||||
@@ -7,6 +7,7 @@ import (
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/llm"
|
||||
"github.com/kami/maven/internal/persona"
|
||||
"github.com/kami/maven/internal/router"
|
||||
"github.com/kami/maven/internal/voice"
|
||||
)
|
||||
@@ -23,13 +24,17 @@ type completer interface {
|
||||
type llmReplier struct {
|
||||
c completer
|
||||
stub *voice.StubReplier
|
||||
|
||||
// block renders the shared context block per turn (who he is, the time).
|
||||
// nil ⇒ the prompt stands alone.
|
||||
block func() string
|
||||
}
|
||||
|
||||
func newLLMReplier(c completer) *llmReplier {
|
||||
return &llmReplier{c: c, stub: voice.NewStubReplier()}
|
||||
func newLLMReplier(c completer, block func() string) *llmReplier {
|
||||
return &llmReplier{c: c, stub: voice.NewStubReplier(), block: block}
|
||||
}
|
||||
|
||||
const replySystem = `Ты — Maven, домашняя ассистентка (о себе — в женском роде). Подтверди действие РОВНО ОДНИМ коротким предложением (≤120 символов), тепло и по-русски. Не задавай вопросов, не повторяй слова, не добавляй ничего после точки. Отвечай ТОЛЬКО одним объектом JSON с полями "response" (текст) и "mood" (ровно одно из: neutral, happy, thinking, tired, confused).
|
||||
const replySystem = `Ты — Maven, домашняя ассистентка (о себе — в женском роде). Владелец — мужчина, говоришь с ним на "ты", в единственном числе; никогда не "вы"/"ваш" и не "он"/"его". Подтверди действие РОВНО ОДНИМ коротким предложением (≤120 символов), по-русски, спокойно и без официальных формулировок. Не задавай вопросов, не повторяй слова, не добавляй ничего после точки. Отвечай ТОЛЬКО одним объектом JSON с полями "response" (текст) и "mood" (ровно одно из: neutral, happy, thinking, tired, confused).
|
||||
Пример: {"response": "Записала, что ты выпил стакан воды.", "mood": "neutral"}
|
||||
Никогда не пиши "..." в поле response.`
|
||||
|
||||
@@ -40,7 +45,7 @@ func (r *llmReplier) Reply(d router.Decision) string {
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
|
||||
defer cancel()
|
||||
out, err := r.c.Complete(ctx, llm.Req{
|
||||
System: replySystem,
|
||||
System: persona.Prepend(r.block, replySystem),
|
||||
User: replyContext(d),
|
||||
MaxTokens: 512,
|
||||
RepeatPenalty: 1.3,
|
||||
|
||||
@@ -17,7 +17,7 @@ type mockCompleter struct {
|
||||
func (m mockCompleter) Complete(_ context.Context, _ llm.Req) (string, error) { return m.out, m.err }
|
||||
|
||||
func TestLLMReplierReturnsLLMReply(t *testing.T) {
|
||||
r := newLLMReplier(mockCompleter{out: `{"response":"записала, кофе закончился","mood":"neutral"}`})
|
||||
r := newLLMReplier(mockCompleter{out: `{"response":"записала, кофе закончился","mood":"neutral"}`}, nil)
|
||||
got := r.Reply(router.Decision{Intent: router.IntentNote, Slots: router.Slots{Text: "кофе закончился"}})
|
||||
if got != "записала, кофе закончился" {
|
||||
t.Errorf("got %q, want %q", got, "записала, кофе закончился")
|
||||
@@ -25,7 +25,7 @@ func TestLLMReplierReturnsLLMReply(t *testing.T) {
|
||||
}
|
||||
|
||||
func TestLLMReplierFallsBackToPlainText(t *testing.T) {
|
||||
r := newLLMReplier(mockCompleter{out: "записала, кофе закончился"})
|
||||
r := newLLMReplier(mockCompleter{out: "записала, кофе закончился"}, nil)
|
||||
got := r.Reply(router.Decision{Intent: router.IntentNote, Slots: router.Slots{Text: "кофе закончился"}})
|
||||
if got != "записала, кофе закончился" {
|
||||
t.Errorf("got %q, want %q", got, "записала, кофе закончился")
|
||||
@@ -33,7 +33,7 @@ func TestLLMReplierFallsBackToPlainText(t *testing.T) {
|
||||
}
|
||||
|
||||
func TestLLMReplierFallsBackToStubOnError(t *testing.T) {
|
||||
r := newLLMReplier(mockCompleter{err: errTestLLMDown})
|
||||
r := newLLMReplier(mockCompleter{err: errTestLLMDown}, nil)
|
||||
noteDec := router.Decision{Intent: router.IntentNote}
|
||||
got := r.Reply(noteDec)
|
||||
want := voice.NewStubReplier().Reply(noteDec)
|
||||
@@ -43,7 +43,7 @@ func TestLLMReplierFallsBackToStubOnError(t *testing.T) {
|
||||
}
|
||||
|
||||
func TestLLMReplierFallsBackToStubOnEmpty(t *testing.T) {
|
||||
r := newLLMReplier(mockCompleter{out: ""})
|
||||
r := newLLMReplier(mockCompleter{out: ""}, nil)
|
||||
noteDec := router.Decision{Intent: router.IntentNote}
|
||||
got := r.Reply(noteDec)
|
||||
want := voice.NewStubReplier().Reply(noteDec)
|
||||
@@ -53,7 +53,7 @@ func TestLLMReplierFallsBackToStubOnEmpty(t *testing.T) {
|
||||
}
|
||||
|
||||
func TestLLMReplierClarifyUsesStub(t *testing.T) {
|
||||
r := newLLMReplier(mockCompleter{out: "я всё поняла"})
|
||||
r := newLLMReplier(mockCompleter{out: "я всё поняла"}, nil)
|
||||
clarifyDec := router.Decision{Clarify: true}
|
||||
got := r.Reply(clarifyDec)
|
||||
want := voice.NewStubReplier().Reply(clarifyDec)
|
||||
|
||||
@@ -0,0 +1,78 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/router"
|
||||
)
|
||||
|
||||
// systemHandler — a handler with nothing but a fixed clock, which is all
|
||||
// replySystem needs.
|
||||
func systemHandler(now time.Time) *reactiveHandler {
|
||||
return &reactiveHandler{now: func() time.Time { return now }}
|
||||
}
|
||||
|
||||
// TestReplySystemDateOffset — "какое число завтра" must answer tomorrow's
|
||||
// date, not today's (Vikunja #388).
|
||||
func TestReplySystemDateOffset(t *testing.T) {
|
||||
// Thursday, 30 July 2026.
|
||||
now := time.Date(2026, 7, 30, 14, 5, 0, 0, time.UTC)
|
||||
h := systemHandler(now)
|
||||
cases := []struct{ utterance, want string }{
|
||||
{"какое сегодня число", "сегодня четверг, 30 июля 2026 года"},
|
||||
{"какое число", "сегодня четверг, 30 июля 2026 года"},
|
||||
{"какое число завтра", "завтра пятница, 31 июля 2026 года"},
|
||||
{"какое число послезавтра", "послезавтра суббота, 1 августа 2026 года"},
|
||||
{"какое было число вчера", "вчера среда, 29 июля 2026 года"},
|
||||
}
|
||||
for _, c := range cases {
|
||||
got := h.replySystem(context.Background(), router.Decision{Utterance: c.utterance})
|
||||
if got != c.want {
|
||||
t.Errorf("replySystem(%q) = %q, want %q", c.utterance, got, c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A day she cannot work out must not come back as today's date — that is the
|
||||
// same silent wrong answer #388 was about, one step further out.
|
||||
func TestReplySystemUnknownDayIsHonest(t *testing.T) {
|
||||
now := time.Date(2026, 7, 30, 14, 5, 0, 0, time.UTC)
|
||||
h := systemHandler(now)
|
||||
for _, u := range []string{
|
||||
"какое число в пятницу",
|
||||
"какое число через неделю",
|
||||
"какое число в понедельник",
|
||||
} {
|
||||
got := h.replySystem(context.Background(), router.Decision{Utterance: u})
|
||||
if got != onlyNearDaysReply {
|
||||
t.Errorf("replySystem(%q) = %q, want the honest reply", u, got)
|
||||
}
|
||||
}
|
||||
// The days she does know must not be caught by the same guard.
|
||||
if got := h.replySystem(context.Background(), router.Decision{Utterance: "какое число завтра"}); got == onlyNearDaysReply {
|
||||
t.Error("завтра was treated as an unknown day")
|
||||
}
|
||||
}
|
||||
|
||||
// TestReplySystemClockCity — the clock arm must not answer local time for a
|
||||
// question about another city (Vikunja #388). She keeps one clock, so every
|
||||
// named place gets the honest "local time only" answer.
|
||||
func TestReplySystemClockCity(t *testing.T) {
|
||||
now := time.Date(2026, 7, 30, 12, 0, 0, 0, time.UTC)
|
||||
h := systemHandler(now)
|
||||
cases := []struct{ utterance, want string }{
|
||||
{"который час", "сейчас 12 часов ровно"},
|
||||
{"который час в киеве", onlyLocalTimeReply},
|
||||
{"сколько времени в москве", onlyLocalTimeReply},
|
||||
{"который час в лондоне", onlyLocalTimeReply},
|
||||
{"который час в бишкеке", onlyLocalTimeReply},
|
||||
}
|
||||
for _, c := range cases {
|
||||
got := h.replySystem(context.Background(), router.Decision{Utterance: c.utterance})
|
||||
if got != c.want {
|
||||
t.Errorf("replySystem(%q) = %q, want %q", c.utterance, got, c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
+256
-36
@@ -171,6 +171,7 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
|
||||
emb = router.NewHashEmbedder(1024)
|
||||
}
|
||||
w.embedder = emb
|
||||
checkStoredEmbedder(dataStore, emb)
|
||||
|
||||
// ----- tool executor (the enabled act allowlist, store-backed) -----
|
||||
// Config tools are the declarative bootstrap: seed them into the store as
|
||||
@@ -227,14 +228,25 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
|
||||
}
|
||||
|
||||
// ----- dialogue (multi-turn slot carry-over; 2-min follow-up window) -----
|
||||
dialogueSessions := dialogue.NewSessionStore(2 * time.Minute)
|
||||
// Store-backed when the daemon passes a store, so a restart mid-conversation
|
||||
// keeps the thread (Vikunja #363). Sessions past their TTL are dropped on
|
||||
// load, never revived. Clarify's parked question stays in memory only.
|
||||
var dialogueSessions *dialogue.SessionStore
|
||||
if dataStore != nil {
|
||||
dialogueSessions = dialogue.NewPersistentSessionStore(2*time.Minute, dataStore)
|
||||
if err := dialogueSessions.Load(context.Background(), time.Now()); err != nil {
|
||||
log.Printf("dialogue: load saved sessions: %v", err)
|
||||
}
|
||||
} else {
|
||||
dialogueSessions = dialogue.NewSessionStore(2 * time.Minute)
|
||||
}
|
||||
clarifyStore := dialogue.NewClarifyStore(clarifyTTL)
|
||||
timeParser := router.NewPythonDateParser()
|
||||
|
||||
// ----- replier (LLM-backed when the engine is on, Stub floor otherwise) -----
|
||||
replier := voice.Replier(voice.NewStubReplier())
|
||||
if llmClient != nil {
|
||||
replier = newLLMReplier(llmClient)
|
||||
replier = newLLMReplier(llmClient, contextBlockFn(cfg, time.Now))
|
||||
}
|
||||
|
||||
// ----- the handler (the reactive path; closes over stt / tts / router / coreAPI / memory) -----
|
||||
@@ -402,9 +414,17 @@ func (h *reactiveHandler) HandlePushToTalk(ctx context.Context, req voice.PushTo
|
||||
return h.reply(ctx, reply, nil)
|
||||
}
|
||||
|
||||
// 1b2. clarify answer — if she asked a question last turn, this utterance is
|
||||
// its answer, not a fresh command. After the confirm check: a y/n gate is
|
||||
// armed by her own prompt and is the narrower claim on the utterance.
|
||||
// 1b2. expired clarify — a question was parked but its TTL ran out, so the
|
||||
// request behind it is gone. Say that out loud (see clarify.go) and carry
|
||||
// on: these words are still routed as a fresh utterance below, with the
|
||||
// notice glued in front of whatever the fresh routing answers. Checked
|
||||
// BEFORE the answer path: reading a parked question drops an expired one.
|
||||
expiredNotice := h.clarifyExpiredNotice()
|
||||
|
||||
// 1b3. clarify answer — if she asked a live question last turn, this
|
||||
// utterance is its answer, not a fresh command. After the confirm check: a
|
||||
// y/n gate is armed by her own prompt and is the narrower claim on the
|
||||
// utterance.
|
||||
if reply, handled := h.resolveClarifyAnswer(ctx, text); handled {
|
||||
return h.reply(ctx, reply, nil)
|
||||
}
|
||||
@@ -414,7 +434,7 @@ func (h *reactiveHandler) HandlePushToTalk(ctx context.Context, req voice.PushTo
|
||||
// unreliably (it's a command, not a free-form query), so we match it
|
||||
// before routing. Same pattern as the confirm turn above.
|
||||
if reply, handled := h.resolveQuietToggle(ctx, text); handled {
|
||||
return h.reply(ctx, reply, nil)
|
||||
return h.reply(ctx, withNotice(expiredNotice, reply), nil)
|
||||
}
|
||||
|
||||
// 2. router — classify the utterance.
|
||||
@@ -423,10 +443,10 @@ func (h *reactiveHandler) HandlePushToTalk(ctx context.Context, req voice.PushTo
|
||||
// ErrNoIntents ⇒ classifier unseeded (cold boot). reply with a
|
||||
// "still warming up" rather than a wire error.
|
||||
if errors.Is(err, router.ErrNoIntents) {
|
||||
return h.reply(ctx, "я ещё не понимаю свободную речь — скоро научусь.", nil)
|
||||
return h.reply(ctx, withNotice(expiredNotice, "я ещё не понимаю свободную речь — скоро научусь."), nil)
|
||||
}
|
||||
log.Printf("voice: router error: %v", err)
|
||||
return h.reply(ctx, "не получилось разобрать команду.", nil)
|
||||
return h.reply(ctx, withNotice(expiredNotice, "не получилось разобрать команду."), nil)
|
||||
}
|
||||
|
||||
// 2b. dialogue — fill this turn's missing slots from a prior same-intent
|
||||
@@ -447,7 +467,7 @@ func (h *reactiveHandler) HandlePushToTalk(ctx context.Context, req voice.PushTo
|
||||
// stands.
|
||||
if dec.Clarify {
|
||||
if question, asked := h.askClarify(dec); asked {
|
||||
return h.reply(ctx, question, nil)
|
||||
return h.reply(ctx, withNotice(expiredNotice, question), nil)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -463,7 +483,7 @@ func (h *reactiveHandler) HandlePushToTalk(ctx context.Context, req voice.PushTo
|
||||
|
||||
// 5. tts — synthesise the reply text; return to the voice server which
|
||||
// ships it back on the conn.
|
||||
return h.reply(ctx, replyText, nil)
|
||||
return h.reply(ctx, withNotice(expiredNotice, replyText), nil)
|
||||
}
|
||||
|
||||
// handleText — the core reactive path without stt/tts: confirm check →
|
||||
@@ -478,7 +498,9 @@ func (h *reactiveHandler) handleText(ctx context.Context, text string) string {
|
||||
return reply
|
||||
}
|
||||
|
||||
// 1b2. clarify answer — same check as HandlePushToTalk.
|
||||
// 1b2/1b3. expired clarify then clarify answer — same order and reasons as
|
||||
// HandlePushToTalk.
|
||||
expiredNotice := h.clarifyExpiredNotice()
|
||||
if reply, handled := h.resolveClarifyAnswer(ctx, text); handled {
|
||||
return reply
|
||||
}
|
||||
@@ -487,10 +509,10 @@ func (h *reactiveHandler) handleText(ctx context.Context, text string) string {
|
||||
dec, err := h.router.Route(ctx, text, h.now())
|
||||
if err != nil {
|
||||
if errors.Is(err, router.ErrNoIntents) {
|
||||
return "я ещё не понимаю свободную речь — скоро научусь."
|
||||
return withNotice(expiredNotice, "я ещё не понимаю свободную речь — скоро научусь.")
|
||||
}
|
||||
log.Printf("voice: handleText router error: %v", err)
|
||||
return "не получилось разобрать команду."
|
||||
return withNotice(expiredNotice, "не получилось разобрать команду.")
|
||||
}
|
||||
log.Printf("voice: route result: intent=%s slots=%+v", dec.Intent, dec.Slots)
|
||||
|
||||
@@ -507,7 +529,7 @@ func (h *reactiveHandler) handleText(ctx context.Context, text string) string {
|
||||
// 2c. clarify — same as HandlePushToTalk: ask about the one missing thing.
|
||||
if dec.Clarify {
|
||||
if question, asked := h.askClarify(dec); asked {
|
||||
return question
|
||||
return withNotice(expiredNotice, question)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -519,7 +541,7 @@ func (h *reactiveHandler) handleText(ctx context.Context, text string) string {
|
||||
if replyText == "" {
|
||||
replyText = h.replier.Reply(dec)
|
||||
}
|
||||
return replyText
|
||||
return withNotice(expiredNotice, replyText)
|
||||
}
|
||||
|
||||
// applyAction — executes the router's Decision. Intent-by-intent:
|
||||
@@ -730,7 +752,9 @@ func (h *reactiveHandler) applyAction(ctx context.Context, dec router.Decision)
|
||||
}
|
||||
|
||||
// Calendar questions: "что у меня сегодня?", "планы на завтра?"
|
||||
if date, ok := router.ParseCalendarDate(dec.Utterance, time.Now()); ok {
|
||||
// h.now(), not time.Now(): the handler's clock is the injected one, so
|
||||
// this arm can be tested at a fixed time like the rest.
|
||||
if date, ok := router.ParseCalendarDate(dec.Utterance, h.now()); ok {
|
||||
events, err := h.api.CalendarEvents(ctx, date, date.Add(24*time.Hour))
|
||||
if err != nil {
|
||||
log.Printf("voice: calendar events: %v", err)
|
||||
@@ -765,11 +789,43 @@ func (h *reactiveHandler) applyAction(ctx context.Context, dec router.Decision)
|
||||
log.Printf("voice: embed query: %v", err)
|
||||
return "не получилось найти ответ."
|
||||
}
|
||||
// Long-term memory first: ONE search over everything Maven remembers
|
||||
// (notes and facts share this index) and ONE confidence gate, so the
|
||||
// memory that is clearly the best match answers — a note just as much
|
||||
// as a fact.
|
||||
//
|
||||
// This used to run only after the notes-only gate below had already
|
||||
// rejected the same note at the same score, which no note could ever
|
||||
// survive a second time: the branch could only return a fact (#373).
|
||||
// Order, not the gate, was the bug — the set of questions Maven answers
|
||||
// is unchanged, only which memory gets to answer them.
|
||||
if h.memStore != nil {
|
||||
if hits, herr := h.memStore.Search(ctx, vec, 3); herr == nil {
|
||||
if hit, ok := bestRecall(hits, h.queryMinScore, h.queryMinMargin); ok {
|
||||
text := hit.Meta["text"]
|
||||
// A note is phrased in Maven's voice; a fact is read back
|
||||
// as it was stored.
|
||||
if hit.Meta["type"] == "note" {
|
||||
if reply, perr := h.phraser.PhraseQuery(ctx, dec.Utterance, []string{text}); perr == nil && reply != "" {
|
||||
return reply
|
||||
}
|
||||
}
|
||||
return text
|
||||
}
|
||||
} else {
|
||||
log.Printf("voice: memory search: %v", herr)
|
||||
}
|
||||
}
|
||||
|
||||
notes, err := h.api.QueryNotes(ctx, vec, 5)
|
||||
if err != nil {
|
||||
log.Printf("voice: query notes: %v", err)
|
||||
return "не получилось найти ответ."
|
||||
}
|
||||
// Notes-only pass, for notes the vector index above does not hold (an
|
||||
// older note written before it existed). Same gate, notes-only
|
||||
// candidates.
|
||||
//
|
||||
// Confidence gate: below it, say "I don't know" rather than read back
|
||||
// the least-unrelated note — a confident wrong recall is worse than a
|
||||
// gap (spec's "not a guesser-of-truth"). Same instinct as the loop's
|
||||
@@ -781,16 +837,6 @@ func (h *reactiveHandler) applyAction(ctx context.Context, dec router.Decision)
|
||||
noteScores[i] = n.Score
|
||||
}
|
||||
if !memory.ConfidentScores(noteScores, h.queryMinScore, h.queryMinMargin) {
|
||||
// Long-term memory recall (notes + facts) before general knowledge:
|
||||
// the notes table can't answer fact questions, but the memory store
|
||||
// indexes both. Only runs when notes-RAG already gave up → additive.
|
||||
if h.memStore != nil {
|
||||
if hits, herr := h.memStore.Search(ctx, vec, 3); herr == nil {
|
||||
if text, ok := bestRecall(hits, h.queryMinScore, h.queryMinMargin); ok {
|
||||
return text
|
||||
}
|
||||
}
|
||||
}
|
||||
// Try general knowledge from the phraser before giving up
|
||||
reply, err := h.phraser.PhraseQuery(ctx, dec.Utterance, nil)
|
||||
if err != nil || reply == "" {
|
||||
@@ -892,6 +938,101 @@ var ruMonths = []string{
|
||||
"июля", "августа", "сентября", "октября", "ноября", "декабря",
|
||||
}
|
||||
|
||||
// onlyLocalTimeReply — the honest answer when the user asks the time somewhere
|
||||
// other than here. She only keeps one clock, and saying so is better than
|
||||
// naming the wrong city's time.
|
||||
//
|
||||
// There used to be a city→time-zone table here. It was removed on purpose: the
|
||||
// user only ever asks for local time, so the table was a second list of cities
|
||||
// to keep in step with the weather one for no gain.
|
||||
const onlyLocalTimeReply = "я знаю только местное время, про другие города пока не скажу."
|
||||
|
||||
// notPlaceAfterV — words that follow "в" without naming a place, so
|
||||
// mentionsUnknownPlace does not mistake them for a city.
|
||||
var notPlaceAfterV = map[string]bool{
|
||||
"данный": true, "данную": true, "этот": true, "эту": true,
|
||||
"котором": true, "какое": true, "какой": true, "который": true,
|
||||
"общем": true, "точности": true, "курсе": true, "сутках": true,
|
||||
"часах": true, "минутах": true, "секундах": true, "неделе": true,
|
||||
}
|
||||
|
||||
// mentionsUnknownPlace reports whether the question has a "в <слово>" phrase
|
||||
// that looks like a place we do not know ("который час в киеве"). Used only to
|
||||
// pick the honest "local time only" reply instead of answering local time as
|
||||
// if it were the city's.
|
||||
func mentionsUnknownPlace(u string) bool {
|
||||
toks := strings.Fields(u)
|
||||
for i := 0; i+1 < len(toks); i++ {
|
||||
if toks[i] != "в" && toks[i] != "во" {
|
||||
continue
|
||||
}
|
||||
next := strings.Trim(toks[i+1], ".,?!")
|
||||
if next == "" || notPlaceAfterV[next] {
|
||||
continue
|
||||
}
|
||||
// A number after "в" is a clock ("в 5 часов"), not a place.
|
||||
if _, err := strconv.Atoi(strings.SplitN(next, ":", 2)[0]); err == nil {
|
||||
continue
|
||||
}
|
||||
return true
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// onlyNearDaysReply — she can work out today, tomorrow, the day after and
|
||||
// yesterday, and nothing further. Said out loud instead of answering today's
|
||||
// date for a day she did not understand.
|
||||
const onlyNearDaysReply = "я считаю только сегодня, завтра, послезавтра и вчера — про другие дни пока не скажу."
|
||||
|
||||
// dayWords — day references the calendar parser cannot resolve. A weekday name
|
||||
// or a "через …" phrase means he asked about a specific other day.
|
||||
var dayWords = []string{
|
||||
"понедельник", "вторник", "сред", "четверг", "пятниц", "суббот", "воскресен",
|
||||
"через", "monday", "tuesday", "wednesday", "thursday", "friday", "saturday", "sunday",
|
||||
}
|
||||
|
||||
// mentionsUnknownDay reports whether the question names a day the calendar
|
||||
// parser could not resolve. Mirror of mentionsUnknownPlace: it exists only to
|
||||
// pick an honest reply over a confidently wrong one.
|
||||
//
|
||||
// Only called after ParseCalendarDate has already failed, so "завтра" and the
|
||||
// other words it does know never reach here.
|
||||
func mentionsUnknownDay(u string) bool {
|
||||
for _, w := range dayWords {
|
||||
if strings.Contains(u, w) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// ruClock renders the clock part of the time reply: "15 часов 4 минуты".
|
||||
func ruClock(t time.Time) string {
|
||||
h, m := t.Hour(), t.Minute()
|
||||
hourWord := ruPlural(h, "час", "часа", "часов")
|
||||
if m == 0 {
|
||||
return fmt.Sprintf("%d %s ровно", h, hourWord)
|
||||
}
|
||||
return fmt.Sprintf("%d %s %d %s", h, hourWord, m, ruPlural(m, "минута", "минуты", "минут"))
|
||||
}
|
||||
|
||||
// dayPrefix names the day relative to now ("завтра", "вчера", …) so the date
|
||||
// reply opens the way a person would say it.
|
||||
func dayPrefix(now, day time.Time) string {
|
||||
base := time.Date(now.Year(), now.Month(), now.Day(), 0, 0, 0, 0, now.Location())
|
||||
switch int(day.Sub(base).Hours() / 24) {
|
||||
case -1:
|
||||
return "вчера"
|
||||
case 0:
|
||||
return "сегодня"
|
||||
case 1:
|
||||
return "завтра"
|
||||
case 2:
|
||||
return "послезавтра"
|
||||
}
|
||||
return "это"
|
||||
}
|
||||
|
||||
func ruPlural(n int, one, two, many string) string {
|
||||
n = n % 100
|
||||
if n > 10 && n < 20 {
|
||||
@@ -970,18 +1111,29 @@ func (h *reactiveHandler) replySystem(ctx context.Context, dec router.Decision)
|
||||
|
||||
switch {
|
||||
case strings.Contains(u, "час") || strings.Contains(u, "врем"):
|
||||
h := now.Hour()
|
||||
m := now.Minute()
|
||||
hourWord := ruPlural(h, "час", "часа", "часов")
|
||||
if m == 0 {
|
||||
return fmt.Sprintf("сейчас %d %s ровно", h, hourWord)
|
||||
// "который час в киеве" — she keeps one clock, so any named place gets
|
||||
// the honest answer. Never local time dressed up as the city's.
|
||||
if mentionsUnknownPlace(u) {
|
||||
return onlyLocalTimeReply
|
||||
}
|
||||
minWord := ruPlural(m, "минута", "минуты", "минут")
|
||||
return fmt.Sprintf("сейчас %d %s %d %s", h, hourWord, m, minWord)
|
||||
return "сейчас " + ruClock(now)
|
||||
case strings.Contains(u, "день") || strings.Contains(u, "числ"):
|
||||
dow := ruWeekdays[now.Weekday()]
|
||||
month := ruMonths[now.Month()-1]
|
||||
return fmt.Sprintf("сегодня %s, %d %s %d года", dow, now.Day(), month, now.Year())
|
||||
// "какое число завтра" — answer for the day the user asked about,
|
||||
// not today. Reuses the router's calendar day-word parser.
|
||||
day := now
|
||||
prefix := "сегодня"
|
||||
if d, ok := router.ParseCalendarDate(u, now); ok {
|
||||
day = d
|
||||
prefix = dayPrefix(now, d)
|
||||
} else if mentionsUnknownDay(u) {
|
||||
// He named a day she cannot work out ("в пятницу", "через неделю").
|
||||
// Answering today's date here would be the same silent wrong answer
|
||||
// this arm was fixed for, so say what she can do instead.
|
||||
return onlyNearDaysReply
|
||||
}
|
||||
dow := ruWeekdays[day.Weekday()]
|
||||
month := ruMonths[day.Month()-1]
|
||||
return fmt.Sprintf("%s %s, %d %s %d года", prefix, dow, day.Day(), month, day.Year())
|
||||
case strings.Contains(u, "кто дома") || strings.Contains(u, "человек дома"):
|
||||
return "присутствие пока не подключено к голосовому запросу."
|
||||
case strings.Contains(u, "памят") || strings.Contains(u, "процессор") || strings.Contains(u, "загрузк") || strings.Contains(u, "статус") || strings.Contains(u, "работа") || strings.Contains(u, "сервис") || strings.Contains(u, "диск") || strings.Contains(u, "ip") || strings.Contains(u, "аптайм") || strings.Contains(u, "трафик") || strings.Contains(u, "интернет"):
|
||||
@@ -1724,3 +1876,71 @@ func jsonStringImpl(s string) string {
|
||||
b = append(b, '"')
|
||||
return string(b)
|
||||
}
|
||||
|
||||
// reembedOnStart is the -reembed flag (set in run()). Opt-in on purpose: see
|
||||
// runReembed.
|
||||
var reembedOnStart bool
|
||||
|
||||
// checkStoredEmbedder compares the embedder we just loaded with the one that
|
||||
// wrote the vectors already in the DB (Vikunja #378).
|
||||
//
|
||||
// The two models we have both make 384-dim vectors, so a size check catches
|
||||
// nothing: after a swap, recall silently compares vectors from different
|
||||
// spaces and the scores are noise. So we say it out loud. Recall itself is not
|
||||
// changed here — the fix is `mavend -reembed`.
|
||||
func checkStoredEmbedder(dataStore *store.Store, emb router.Embedder) {
|
||||
if dataStore == nil {
|
||||
return
|
||||
}
|
||||
current := router.EmbedderID(emb)
|
||||
if reembedOnStart {
|
||||
runReembed(dataStore, emb, current)
|
||||
return
|
||||
}
|
||||
stored, mismatch, err := dataStore.CheckEmbedder(context.Background(), current)
|
||||
if err != nil {
|
||||
log.Printf("voice: embedder marker check failed: %v", err)
|
||||
return
|
||||
}
|
||||
if mismatch {
|
||||
log.Printf("voice: WARNING embedder MISMATCH — stored vectors were written by %q but the configured embedder is %q; recall scores are noise until the notes and facts are re-embedded — run `mavend -reembed` once (Vikunja #378)", stored, current)
|
||||
return
|
||||
}
|
||||
log.Printf("voice: embedder marker ok (%s)", current)
|
||||
}
|
||||
|
||||
// runReembed is the one-shot backfill behind -reembed.
|
||||
//
|
||||
// Why a flag and not automatic on mismatch: the embedder is ONNX on the
|
||||
// laptop's CPU, so a few thousand notes is minutes of work. Doing that silently
|
||||
// inside a normal start would look like the daemon hanging on boot. So the user
|
||||
// runs it once, deliberately, after an embedder swap; the mismatch warning
|
||||
// above tells them to. It re-embeds, logs what it did, and then the daemon
|
||||
// carries on serving as usual — no separate binary, no second start needed.
|
||||
func runReembed(dataStore *store.Store, emb router.Embedder, current string) {
|
||||
log.Printf("voice: re-embedding stored notes and facts with %s — this can take a few minutes, do not interrupt", current)
|
||||
res, err := dataStore.ReembedAll(context.Background(), current,
|
||||
// EmbedPassage, not EmbedQuery: these are stored texts being searched
|
||||
// FOR, which is the side they were written with.
|
||||
func(ctx context.Context, text string) ([]float32, error) {
|
||||
return router.EmbedPassage(ctx, emb, text)
|
||||
})
|
||||
if err != nil {
|
||||
log.Printf("voice: re-embed FAILED, nothing was changed and no marker was written — safe to run again: %v", err)
|
||||
return
|
||||
}
|
||||
if res.Skipped {
|
||||
log.Printf("voice: re-embed skipped — the stored vectors were already written by %s", current)
|
||||
return
|
||||
}
|
||||
log.Printf("voice: re-embed done — %d notes in the notes table, %d notes and %d facts in the memory index, took %s; stored vectors now belong to %s",
|
||||
res.Notes, res.MemNotes, res.Facts, res.Took.Round(time.Second), current)
|
||||
|
||||
// A row with no text cannot be re-embedded, so its vector is still the old
|
||||
// model's noise while the marker now says everything is current. Both write
|
||||
// paths always store the text, so this should be zero — say it loudly
|
||||
// rather than bury it in the line above if it ever isn't.
|
||||
if res.NoText > 0 {
|
||||
log.Printf("voice: WARNING %d stored rows had no text, so their vectors could not be re-embedded and are still noise; they will never match anything useful (Vikunja #378)", res.NoText)
|
||||
}
|
||||
}
|
||||
|
||||
+4
-3
@@ -6,11 +6,12 @@
|
||||
"state_dir": "/var/lib/maven",
|
||||
|
||||
"phraser": {
|
||||
"model_path": "/opt/maven/models/llm/qwen3.5/Qwen3.5-0.8B.Q4_K_M.gguf",
|
||||
"model_path": "/opt/maven/models/llm/qwen3/Qwen3-1.7B-UD-Q4_K_XL.gguf",
|
||||
"bin_path": "llama-server",
|
||||
"n_gpu_layers": 99,
|
||||
"n_ctx": 2048,
|
||||
"timeout": "60s"
|
||||
"n_ctx": 4096,
|
||||
"timeout": "60s",
|
||||
"llm_nudges": false
|
||||
},
|
||||
|
||||
"telegram": {
|
||||
|
||||
@@ -302,6 +302,13 @@ type VoiceConfig struct {
|
||||
// Russian self-reference). Example: "Be formal and answer in English only."
|
||||
Persona string `json:"persona,omitempty"`
|
||||
|
||||
// OwnerName / City — optional facts about the owner, added to the shared
|
||||
// context block (internal/persona). Empty is fine: the block still states
|
||||
// who he is grammatically (a man, addressed as "ты") and the current time.
|
||||
// Nothing about correct behaviour may depend on these being filled in.
|
||||
OwnerName string `json:"owner_name,omitempty"`
|
||||
City string `json:"city,omitempty"`
|
||||
|
||||
// Weather — the weather provider config. nil ⇒ the daemon wires
|
||||
// the stub provider (returns ErrNotConfigured — "погода не настроена").
|
||||
// Set provider to "open-meteo" to use the keyless Open-Meteo API.
|
||||
@@ -362,6 +369,12 @@ type PhraserConfig struct {
|
||||
NGpuLayers int `json:"n_gpu_layers,omitempty"`
|
||||
NCtx int `json:"n_ctx,omitempty"`
|
||||
Timeout Duration `json:"timeout,omitempty"`
|
||||
|
||||
// LLMNudges — let the model word nudges again. Off by default: nudges are
|
||||
// worded from hand-written Russian templates now (the model broke the
|
||||
// persona and invented units). Chat, query and reminder phrasing always go
|
||||
// through the model regardless. See phraser.Config.LLMNudges.
|
||||
LLMNudges bool `json:"llm_nudges,omitempty"`
|
||||
}
|
||||
|
||||
// EmbedderConfig — paths for the ONNX multilingual embedder. The daemon
|
||||
|
||||
@@ -35,6 +35,27 @@ func TestLoadDefaults(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// Nudges come from templates unless the config says otherwise.
|
||||
func TestPhraserLLMNudgesDefaultsOff(t *testing.T) {
|
||||
p := writeConfig(t, `{"phraser":{"model_path":"/tmp/m.gguf"}}`)
|
||||
c, err := Load(p)
|
||||
if err != nil {
|
||||
t.Fatalf("Load: %v", err)
|
||||
}
|
||||
if c.Phraser.LLMNudges {
|
||||
t.Error("llm_nudges defaults on; templates must be the default")
|
||||
}
|
||||
|
||||
p = writeConfig(t, `{"phraser":{"model_path":"/tmp/m.gguf","llm_nudges":true}}`)
|
||||
c, err = Load(p)
|
||||
if err != nil {
|
||||
t.Fatalf("Load: %v", err)
|
||||
}
|
||||
if !c.Phraser.LLMNudges {
|
||||
t.Error("llm_nudges:true did not parse")
|
||||
}
|
||||
}
|
||||
|
||||
func TestLoadDurationsParse(t *testing.T) {
|
||||
p := writeConfig(t, `{"tick_interval":"90s","repeat_interval":"10m"}`)
|
||||
c, err := Load(p)
|
||||
|
||||
@@ -137,9 +137,10 @@ func NewDispatcher(cfg Config) *Dispatcher {
|
||||
// picks for (severity, presence), sends via the matching sink, and records
|
||||
// one nudge row per successful send. returns the dispatches (one per channel).
|
||||
//
|
||||
// a Drop channel = no send, no record (the nudge was suppressed by routing,
|
||||
// not by a failure — "a missed water nudge is noise"). a nil sink = channel
|
||||
// not wired, skip silently. a send error stops the dispatch and returns what
|
||||
// a Drop channel = no send (the nudge was suppressed by routing, not by a
|
||||
// failure — "a missed water nudge is noise"), but it does leave a 'dropped'
|
||||
// outbox row so the suppression is visible. a nil sink = channel not wired,
|
||||
// skip silently. a send error stops the dispatch and returns what
|
||||
// got through — the daemon decides whether to retry.
|
||||
func (d *Dispatcher) DispatchNudge(ctx context.Context, pn PhrasedNudge, now time.Time) ([]Dispatch, error) {
|
||||
c := pn.Candidate
|
||||
@@ -148,6 +149,16 @@ func (d *Dispatcher) DispatchNudge(ctx context.Context, pn PhrasedNudge, now tim
|
||||
for i := 0; i < len(channels); i++ {
|
||||
ch := channels[i]
|
||||
if ch == ChannelDrop {
|
||||
// the routing table suppressed this nudge on purpose (a care nudge
|
||||
// while you're away is noise). that stays — but it must not be
|
||||
// invisible, or "she dropped it" and "the rule never fired" look
|
||||
// the same afterwards. no nudges row: that table feeds the
|
||||
// ignored_rate signal, and a nudge nobody could see must not
|
||||
// count as ignored.
|
||||
id := d.beginOutbox(ctx, "nudge", c.Rule.Name, 0, ch, pn.Summary, now)
|
||||
d.completeOutbox(ctx, id, store.DeliveryDropped, now)
|
||||
log.Printf("dispatcher: dropped %s (sev%d, presence=%s) — routing table suppressed it",
|
||||
c.Rule.Name, c.Severity, c.State.Presence)
|
||||
continue
|
||||
}
|
||||
s := Sendable{
|
||||
@@ -396,6 +407,13 @@ func messageForChannel(s Sendable) string {
|
||||
if !isAway(s.Channel) {
|
||||
return s.Body
|
||||
}
|
||||
return AwayMessage(s)
|
||||
}
|
||||
|
||||
// AwayMessage — the only text an off-box channel may ever carry. Exported so
|
||||
// the away sinks share this one rule instead of each inventing a fallback: the
|
||||
// summary if we have one, otherwise a fixed generic line. Never the body.
|
||||
func AwayMessage(s Sendable) string {
|
||||
if s.Summary != "" {
|
||||
return s.Summary
|
||||
}
|
||||
|
||||
@@ -37,10 +37,12 @@ func TestVoiceNoSessionFallthroughLeavesOutboxTrail(t *testing.T) {
|
||||
[]string{"voice", "ntfy"}, []string{store.DeliveryFailed, store.DeliverySent}},
|
||||
{"sev4 falls through to telegram", loop.Sev4,
|
||||
[]string{"voice", "telegram"}, []string{store.DeliveryFailed, store.DeliverySent}},
|
||||
{"sev1 does not fall through", loop.Sev1,
|
||||
[]string{"voice"}, []string{store.DeliveryFailed}},
|
||||
{"sev2 does not fall through", loop.Sev2,
|
||||
[]string{"voice"}, []string{store.DeliveryFailed}},
|
||||
// care severities still don't reach an away channel; since #370 the
|
||||
// drop itself is a visible row instead of nothing.
|
||||
{"sev1 drops instead of falling through", loop.Sev1,
|
||||
[]string{"voice", "drop"}, []string{store.DeliveryFailed, store.DeliveryDropped}},
|
||||
{"sev2 drops instead of falling through", loop.Sev2,
|
||||
[]string{"voice", "drop"}, []string{store.DeliveryFailed, store.DeliveryDropped}},
|
||||
}
|
||||
for _, c := range cases {
|
||||
t.Run(c.name, func(t *testing.T) {
|
||||
|
||||
@@ -2,10 +2,10 @@
|
||||
//
|
||||
// ntfy is the away-channel for sev3 (ops soft) nudges, sev4 (ops hard)
|
||||
// nudges when present (alongside voice), and reminders when away. the
|
||||
// message body is the Sendable's Summary — the minimal-body rule from the
|
||||
// message body is delivery.AwayMessage — the minimal-body rule from the
|
||||
// spec ("disk low on homesrv," not detail; no shoulder-surf exfil through
|
||||
// the relay). voice gets Body; away channels get Summary, enforced at the
|
||||
// sink so a phraser bug can't exfil.
|
||||
// the relay). the dispatcher already strips detail off away sendables; the
|
||||
// sink uses the same helper so it can't leak the body on its own either.
|
||||
//
|
||||
// ntfy runs locally (docker, 127.0.0.1:8085, deny-all auth). maven publishes
|
||||
// with a dedicated user (write-only to maven-* topics) — the credential is a
|
||||
@@ -69,18 +69,14 @@ func New(cfg Config) (*Sink, error) {
|
||||
}, nil
|
||||
}
|
||||
|
||||
// Send publishes one notification to ntfy. the body is the Sendable's Summary
|
||||
// (minimal body); Title is "maven" (consistent sender identity on the lock
|
||||
// screen — the content is in the body). Priority maps from severity/kind so
|
||||
// Send publishes one notification to ntfy. the body is the minimal away
|
||||
// message (never the full body); Title is "maven" (consistent sender identity
|
||||
// on the lock screen — the content is in the body). Priority maps from severity/kind so
|
||||
// the phone client can ring differently for an alarm vs a soft ops nudge.
|
||||
func (s *Sink) Send(ctx context.Context, d delivery.Sendable) error {
|
||||
body := d.Summary
|
||||
if body == "" {
|
||||
body = d.Body // terse full message beats no message
|
||||
}
|
||||
if body == "" {
|
||||
return fmt.Errorf("ntfysink: empty message for %s", d.Channel)
|
||||
}
|
||||
// never fall back to d.Body: ntfy leaves the box, so an empty summary gets
|
||||
// a generic line instead of the full detail.
|
||||
body := delivery.AwayMessage(d)
|
||||
|
||||
req, err := http.NewRequestWithContext(ctx, http.MethodPost, s.topicURL(), strings.NewReader(body))
|
||||
if err != nil {
|
||||
|
||||
@@ -147,9 +147,9 @@ func TestSendBodyIsSummaryNotFullBody(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
func TestSendFallsBackToBodyWhenSummaryEmpty(t *testing.T) {
|
||||
// a terse full message is better than no message; the phraser should
|
||||
// produce a summary for away-bound severities, but don't silently drop.
|
||||
func TestSendNeverSendsTheBodyWhenSummaryEmpty(t *testing.T) {
|
||||
// #368: this used to fall back to the full body. ntfy leaves the box, so
|
||||
// an empty summary gets a fixed generic line plus the rule name instead.
|
||||
rs := newRecordingServer(t, 200, "")
|
||||
srv := httptest.NewServer(rs.handler())
|
||||
defer srv.Close()
|
||||
@@ -160,12 +160,15 @@ func TestSendFallsBackToBodyWhenSummaryEmpty(t *testing.T) {
|
||||
t.Fatalf("Send: %v", err)
|
||||
}
|
||||
_, _, body, _, _, _ := rs.snapshot()
|
||||
if body != s.Body {
|
||||
t.Fatalf("fallback body: want %q, got %q", s.Body, body)
|
||||
want := delivery.GenericAwayMessage + ": service_down"
|
||||
if body != want {
|
||||
t.Fatalf("body: want %q, got %q", want, body)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSendRejectsEmptyMessage(t *testing.T) {
|
||||
func TestSendNeverSendsAnEmptyMessage(t *testing.T) {
|
||||
// with nothing at all to say we still send the generic line — an away
|
||||
// channel can never carry detail, but it also never goes out blank.
|
||||
rs := newRecordingServer(t, 200, "")
|
||||
srv := httptest.NewServer(rs.handler())
|
||||
defer srv.Close()
|
||||
@@ -173,9 +176,13 @@ func TestSendRejectsEmptyMessage(t *testing.T) {
|
||||
sink, _ := New(Config{BaseURL: srv.URL, Topic: "maven"})
|
||||
s := nudgeSendable(loop.Sev3, "")
|
||||
s.Body = ""
|
||||
err := sink.Send(context.Background(), s)
|
||||
if err == nil {
|
||||
t.Fatal("want error for empty message")
|
||||
s.RuleName = ""
|
||||
if err := sink.Send(context.Background(), s); err != nil {
|
||||
t.Fatalf("Send: %v", err)
|
||||
}
|
||||
_, _, body, _, _, _ := rs.snapshot()
|
||||
if body != delivery.GenericAwayMessage {
|
||||
t.Fatalf("body: want %q, got %q", delivery.GenericAwayMessage, body)
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -208,10 +208,9 @@ func TestAwayChannelsGetMinimalBody(t *testing.T) {
|
||||
// TestCareAwayDropIsRecorded — DESIGN.md's drop is a decision ("a missed water
|
||||
// nudge is noise, a missed backup failure isn't"), so it should be visible
|
||||
// rather than vanish. Today drop is a bare `continue`: no nudge row, no outbox
|
||||
// attempt, no log — nothing an operator can see afterwards.
|
||||
// attempt, no log — nothing an operator can see afterwards. now it leaves a
|
||||
// 'dropped' outbox row.
|
||||
func TestCareAwayDropIsRecorded(t *testing.T) {
|
||||
t.Skip("not implemented: dispatcher.go:149-151 skips a Drop channel with no record; there is no 'dropped' outcome in store/delivery.go:16-21")
|
||||
|
||||
ob := &fakeOutbox{}
|
||||
d := NewDispatcher(Config{Voice: &fakeSink{}, Nudges: &fakeNudgeRecorder{}, Outbox: ob})
|
||||
|
||||
|
||||
@@ -2,11 +2,12 @@
|
||||
//
|
||||
// telegram is the away-channel for sev4 (ops hard) nudges — "disk-fire alarm
|
||||
// at 2am routes to telegram, repeat til ack." the message body is the
|
||||
// Sendable's Summary — the minimal-body rule from the spec ("disk low on
|
||||
// homesrv," not detail; no shoulder-surf exfil through the relay). voice gets
|
||||
// Body; away channels get Summary, enforced at the sink so a phraser bug can't
|
||||
// exfil. additionally, protect_content=true is passed on every send so the
|
||||
// message can't be forwarded out of the chat — locks the minimal body further.
|
||||
// delivery.AwayMessage — the minimal-body rule from the spec ("disk low on
|
||||
// homesrv," not detail; no shoulder-surf exfil through the relay). the
|
||||
// dispatcher already strips detail off away sendables; the sink uses the same
|
||||
// helper so it can't leak the body on its own either. additionally,
|
||||
// protect_content=true is passed on every send so the message can't be
|
||||
// forwarded out of the chat — locks the minimal body further.
|
||||
//
|
||||
// telegram's bot API is region-restricted for this homesrv — direct egress to
|
||||
// api.telegram.org is unreliable. the spec's "away channels leave the box —
|
||||
@@ -140,18 +141,13 @@ type telegramResp struct {
|
||||
}
|
||||
|
||||
// Send publishes one message to the configured telegram chat. the body is the
|
||||
// Sendable's Summary (minimal body); empty Summary falls back to Body (terse
|
||||
// full message beats no message). protect_content=true so a phraser bug (Body
|
||||
// leaking detail through Summary) can't be forwarded onward by the user or a
|
||||
// chat observer — locks the minimal-body rule at the channel's own last mile.
|
||||
// minimal away message (never the full body). protect_content=true so even
|
||||
// that can't be forwarded onward by the user or a chat observer — locks the
|
||||
// minimal-body rule at the channel's own last mile.
|
||||
func (s *Sink) Send(ctx context.Context, d delivery.Sendable) error {
|
||||
body := d.Summary
|
||||
if body == "" {
|
||||
body = d.Body
|
||||
}
|
||||
if body == "" {
|
||||
return fmt.Errorf("telegramsink: empty message for %s", d.Channel)
|
||||
}
|
||||
// never fall back to d.Body: telegram leaves the box, so an empty summary
|
||||
// gets a generic line instead of the full detail.
|
||||
body := delivery.AwayMessage(d)
|
||||
|
||||
payload := sendMessageReq{
|
||||
ChatID: s.cfg.ChatID,
|
||||
|
||||
@@ -173,9 +173,9 @@ func TestSendBodyIsSummaryNotFullBody(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
func TestSendFallsBackToBodyWhenSummaryEmpty(t *testing.T) {
|
||||
// terse full message beats none; the phraser should produce a summary for
|
||||
// away-bound severities, but don't silently drop.
|
||||
func TestSendNeverSendsTheBodyWhenSummaryEmpty(t *testing.T) {
|
||||
// #368: this used to fall back to the full body. telegram leaves the box,
|
||||
// so an empty summary gets a fixed generic line plus the rule name.
|
||||
rs := newRecordingServer(t, 200, "")
|
||||
srv := httptest.NewServer(rs.handler())
|
||||
defer srv.Close()
|
||||
@@ -188,12 +188,14 @@ func TestSendFallsBackToBodyWhenSummaryEmpty(t *testing.T) {
|
||||
_, _, body, _, _ := rs.snapshot()
|
||||
var req sendMessageReq
|
||||
_ = json.Unmarshal([]byte(body), &req)
|
||||
if req.Text != s.Body {
|
||||
t.Fatalf("fallback text: want %q, got %q", s.Body, req.Text)
|
||||
want := delivery.GenericAwayMessage + ": service_down"
|
||||
if req.Text != want {
|
||||
t.Fatalf("text: want %q, got %q", want, req.Text)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSendRejectsEmptyMessage(t *testing.T) {
|
||||
func TestSendNeverSendsAnEmptyMessage(t *testing.T) {
|
||||
// with nothing at all to say we still send the generic line.
|
||||
rs := newRecordingServer(t, 200, "")
|
||||
srv := httptest.NewServer(rs.handler())
|
||||
defer srv.Close()
|
||||
@@ -201,9 +203,15 @@ func TestSendRejectsEmptyMessage(t *testing.T) {
|
||||
sink, _ := New(sinkCfg(srv.URL))
|
||||
s := nudgeSendable(loop.Sev4, "")
|
||||
s.Body = ""
|
||||
err := sink.Send(context.Background(), s)
|
||||
if err == nil {
|
||||
t.Fatal("want error for empty message")
|
||||
s.RuleName = ""
|
||||
if err := sink.Send(context.Background(), s); err != nil {
|
||||
t.Fatalf("Send: %v", err)
|
||||
}
|
||||
_, _, body, _, _ := rs.snapshot()
|
||||
var req sendMessageReq
|
||||
_ = json.Unmarshal([]byte(body), &req)
|
||||
if req.Text != delivery.GenericAwayMessage {
|
||||
t.Fatalf("text: want %q, got %q", delivery.GenericAwayMessage, req.Text)
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -100,6 +100,21 @@ func (s *ClarifyStore) Get(id string, now time.Time) *PendingQuestion {
|
||||
return q
|
||||
}
|
||||
|
||||
// TakeExpired reports whether a question was parked here but its TTL ran out,
|
||||
// and drops it. Get drops such a question silently, which leaves the user
|
||||
// thinking his request is still alive — the caller uses this to tell him it is
|
||||
// gone before treating his words as a fresh utterance.
|
||||
func (s *ClarifyStore) TakeExpired(id string, now time.Time) bool {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
q, ok := s.questions[id]
|
||||
if !ok || !q.IsExpired(now) {
|
||||
return false
|
||||
}
|
||||
delete(s.questions, id)
|
||||
return true
|
||||
}
|
||||
|
||||
func (s *ClarifyStore) Delete(id string) {
|
||||
s.mu.Lock()
|
||||
delete(s.questions, id)
|
||||
|
||||
@@ -60,6 +60,28 @@ func TestClarifyStoreGetPutDelete(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestClarifyStoreTakeExpired — TakeExpired reports (and drops) only a question
|
||||
// whose TTL ran out.
|
||||
func TestClarifyStoreTakeExpired(t *testing.T) {
|
||||
s := NewClarifyStore(time.Minute)
|
||||
if s.TakeExpired("voice", base) {
|
||||
t.Fatal("nothing parked ⇒ nothing expired")
|
||||
}
|
||||
s.Put("voice", &PendingQuestion{Missing: []Slot{SlotTime}, Asked: base, TTL: time.Minute})
|
||||
if s.TakeExpired("voice", base.Add(30*time.Second)) {
|
||||
t.Fatal("a live question must not report as expired")
|
||||
}
|
||||
if s.Get("voice", base.Add(30*time.Second)) == nil {
|
||||
t.Fatal("a live question must survive TakeExpired")
|
||||
}
|
||||
if !s.TakeExpired("voice", base.Add(2*time.Minute)) {
|
||||
t.Fatal("a stale question must report as expired")
|
||||
}
|
||||
if s.TakeExpired("voice", base.Add(2*time.Minute)) {
|
||||
t.Fatal("TakeExpired must drop the question, so the second call is false")
|
||||
}
|
||||
}
|
||||
|
||||
func TestNewClarifyStoreDefaultTTL(t *testing.T) {
|
||||
s := NewClarifyStore(0)
|
||||
q := &PendingQuestion{Asked: base}
|
||||
|
||||
@@ -1,8 +1,12 @@
|
||||
package dialogue
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"sync"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/store"
|
||||
)
|
||||
|
||||
type Intent string
|
||||
@@ -50,10 +54,22 @@ func (s *Session) IsExpired(now time.Time) bool {
|
||||
return now.After(s.Timestamp.Add(s.TTL))
|
||||
}
|
||||
|
||||
// SessionPersister — the bit of the store the session needs, as an interface
|
||||
// so tests can swap it out. Data is an opaque blob: the store never looks
|
||||
// inside, we encode the session as JSON here.
|
||||
type SessionPersister interface {
|
||||
SaveDialogueSession(ctx context.Context, id string, data []byte, ts time.Time, ttl time.Duration) error
|
||||
DeleteDialogueSession(ctx context.Context, id string) error
|
||||
LoadDialogueSessions(ctx context.Context, now time.Time) ([]store.DialogueSessionRow, error)
|
||||
}
|
||||
|
||||
// SessionStore keeps the live sessions in a map (the fast path) and mirrors
|
||||
// every write to the persister, so a daemon restart can load them back.
|
||||
type SessionStore struct {
|
||||
mu sync.RWMutex
|
||||
sessions map[string]*Session
|
||||
defaultTTL time.Duration
|
||||
persist SessionPersister // may be nil: memory only (tests, no-store paths)
|
||||
}
|
||||
|
||||
func NewSessionStore(defaultTTL time.Duration) *SessionStore {
|
||||
@@ -66,6 +82,43 @@ func NewSessionStore(defaultTTL time.Duration) *SessionStore {
|
||||
}
|
||||
}
|
||||
|
||||
// NewPersistentSessionStore — same store, but writes also go to the DB.
|
||||
// Call Load once after this to bring back sessions from a previous run.
|
||||
func NewPersistentSessionStore(defaultTTL time.Duration, p SessionPersister) *SessionStore {
|
||||
s := NewSessionStore(defaultTTL)
|
||||
s.persist = p
|
||||
return s
|
||||
}
|
||||
|
||||
// Load — read the saved sessions back into memory. Anything past its TTL is
|
||||
// dropped (and deleted from the DB by the store), never revived.
|
||||
func (s *SessionStore) Load(ctx context.Context, now time.Time) error {
|
||||
if s.persist == nil {
|
||||
return nil
|
||||
}
|
||||
rows, err := s.persist.LoadDialogueSessions(ctx, now)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
for _, r := range rows {
|
||||
var sess Session
|
||||
if err := json.Unmarshal(r.Data, &sess); err != nil {
|
||||
// A blob we can't read is not worth failing a startup over.
|
||||
continue
|
||||
}
|
||||
if sess.TTL <= 0 {
|
||||
sess.TTL = r.TTL
|
||||
}
|
||||
if sess.IsExpired(now) {
|
||||
continue
|
||||
}
|
||||
s.sessions[r.ID] = &sess
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
func (s *SessionStore) Get(id string, now time.Time) *Session {
|
||||
s.mu.RLock()
|
||||
sess, ok := s.sessions[id]
|
||||
@@ -87,12 +140,33 @@ func (s *SessionStore) Put(id string, sess *Session) {
|
||||
s.mu.Lock()
|
||||
s.sessions[id] = sess
|
||||
s.mu.Unlock()
|
||||
s.save(id, sess)
|
||||
}
|
||||
|
||||
func (s *SessionStore) Delete(id string) {
|
||||
s.mu.Lock()
|
||||
delete(s.sessions, id)
|
||||
s.mu.Unlock()
|
||||
if s.persist != nil {
|
||||
_ = s.persist.DeleteDialogueSession(context.Background(), id)
|
||||
}
|
||||
}
|
||||
|
||||
// save — mirror one session to the DB. Best effort: memory already has it, so
|
||||
// a write error costs us the restart safety net, not the current turn.
|
||||
func (s *SessionStore) save(id string, sess *Session) {
|
||||
if s.persist == nil {
|
||||
return
|
||||
}
|
||||
data, err := json.Marshal(sess)
|
||||
if err != nil {
|
||||
return
|
||||
}
|
||||
ts := sess.Timestamp
|
||||
if ts.IsZero() {
|
||||
ts = time.Now()
|
||||
}
|
||||
_ = s.persist.SaveDialogueSession(context.Background(), id, data, ts, sess.TTL)
|
||||
}
|
||||
|
||||
func InheritSlots(prev, cur Slots) Slots {
|
||||
|
||||
@@ -0,0 +1,118 @@
|
||||
package dialogue
|
||||
|
||||
import (
|
||||
"context"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/store"
|
||||
)
|
||||
|
||||
// openStore — a store on disk, so a second handle can reopen the same file.
|
||||
func openStore(t *testing.T, path string) *store.Store {
|
||||
t.Helper()
|
||||
s, err := store.Open(context.Background(), path)
|
||||
if err != nil {
|
||||
t.Fatalf("open store: %v", err)
|
||||
}
|
||||
t.Cleanup(func() { _ = s.Close() })
|
||||
return s
|
||||
}
|
||||
|
||||
// A session written before a restart comes back and still merges a follow-up.
|
||||
func TestSessionSurvivesRestart(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
path := filepath.Join(t.TempDir(), "maven_test.db")
|
||||
now := time.Now().UTC().Truncate(time.Millisecond)
|
||||
|
||||
first := openStore(t, path)
|
||||
before := NewPersistentSessionStore(2*time.Minute, first)
|
||||
before.Put("voice", &Session{
|
||||
Intent: IntentReminder,
|
||||
Slots: Slots{Text: "полить цветы", Time: now.Add(time.Hour), HasTime: true},
|
||||
Timestamp: now,
|
||||
TTL: 2 * time.Minute,
|
||||
})
|
||||
if err := first.Close(); err != nil {
|
||||
t.Fatalf("close: %v", err)
|
||||
}
|
||||
|
||||
// fresh handle, fresh in-memory map — as after a daemon restart
|
||||
after := openStore(t, path)
|
||||
reloaded := NewPersistentSessionStore(2*time.Minute, after)
|
||||
if err := reloaded.Load(ctx, now.Add(10*time.Second)); err != nil {
|
||||
t.Fatalf("Load: %v", err)
|
||||
}
|
||||
sess := reloaded.Get("voice", now.Add(10*time.Second))
|
||||
if sess == nil {
|
||||
t.Fatal("session did not survive the restart")
|
||||
}
|
||||
if sess.Intent != IntentReminder {
|
||||
t.Fatalf("intent = %q, want reminder", sess.Intent)
|
||||
}
|
||||
// the follow-up carries no text of its own; it must inherit the old one
|
||||
merged := InheritSlots(sess.Slots, Slots{Time: now.Add(2 * time.Hour), HasTime: true})
|
||||
if merged.Text != "полить цветы" {
|
||||
t.Fatalf("merged text = %q, want the earlier turn's text", merged.Text)
|
||||
}
|
||||
if !merged.Time.Equal(now.Add(2 * time.Hour)) {
|
||||
t.Fatalf("merged time = %v, want the follow-up's time", merged.Time)
|
||||
}
|
||||
}
|
||||
|
||||
// A session past its TTL is dead: a restart must not bring it back.
|
||||
func TestExpiredSessionNotResurrected(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
path := filepath.Join(t.TempDir(), "maven_test.db")
|
||||
now := time.Now().UTC().Truncate(time.Millisecond)
|
||||
|
||||
first := openStore(t, path)
|
||||
before := NewPersistentSessionStore(2*time.Minute, first)
|
||||
before.Put("voice", &Session{
|
||||
Intent: IntentReminder,
|
||||
Slots: Slots{Text: "полить цветы"},
|
||||
Timestamp: now,
|
||||
TTL: time.Minute,
|
||||
})
|
||||
if err := first.Close(); err != nil {
|
||||
t.Fatalf("close: %v", err)
|
||||
}
|
||||
|
||||
after := openStore(t, path)
|
||||
reloaded := NewPersistentSessionStore(2*time.Minute, after)
|
||||
later := now.Add(5 * time.Minute) // well past the 1-min TTL
|
||||
if err := reloaded.Load(ctx, later); err != nil {
|
||||
t.Fatalf("Load: %v", err)
|
||||
}
|
||||
if sess := reloaded.Get("voice", later); sess != nil {
|
||||
t.Fatalf("expired session came back: %+v", sess)
|
||||
}
|
||||
// and it is gone from the DB too, not just from memory
|
||||
rows, err := after.LoadDialogueSessions(ctx, later)
|
||||
if err != nil {
|
||||
t.Fatalf("LoadDialogueSessions: %v", err)
|
||||
}
|
||||
if len(rows) != 0 {
|
||||
t.Fatalf("expired row still in the DB: %+v", rows)
|
||||
}
|
||||
}
|
||||
|
||||
// Delete removes the row as well, so an ended conversation stays ended.
|
||||
func TestDeleteRemovesPersistedSession(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
path := filepath.Join(t.TempDir(), "maven_test.db")
|
||||
now := time.Now().UTC().Truncate(time.Millisecond)
|
||||
|
||||
s := openStore(t, path)
|
||||
ss := NewPersistentSessionStore(2*time.Minute, s)
|
||||
ss.Put("voice", &Session{Intent: IntentChat, Timestamp: now, TTL: time.Minute})
|
||||
ss.Delete("voice")
|
||||
rows, err := s.LoadDialogueSessions(ctx, now)
|
||||
if err != nil {
|
||||
t.Fatalf("LoadDialogueSessions: %v", err)
|
||||
}
|
||||
if len(rows) != 0 {
|
||||
t.Fatalf("row survived Delete: %+v", rows)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,116 @@
|
||||
// Package kiwix reads a local Kiwix server (offline Wikipedia and friends).
|
||||
//
|
||||
// Why: the resident model is a 0.8B and invents facts. Letting her read a local
|
||||
// article snippet beats letting her recall. Nothing here talks to the internet;
|
||||
// the Kiwix server is on the same box.
|
||||
//
|
||||
// This is search only. Full articles are ~100KB of HTML, far too big for a 4096
|
||||
// token context, so the unit of context is the search snippet (~500 chars).
|
||||
package kiwix
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/xml"
|
||||
"fmt"
|
||||
"html"
|
||||
"io"
|
||||
"net/http"
|
||||
"net/url"
|
||||
"regexp"
|
||||
"strconv"
|
||||
"strings"
|
||||
"time"
|
||||
)
|
||||
|
||||
// Result is one search hit.
|
||||
type Result struct {
|
||||
Title string // article title, e.g. "Rayleigh scattering"
|
||||
Path string // e.g. /content/wikipedia_en_all_maxi_2026-02/Rayleigh_scattering
|
||||
Snippet string // plain text, tags stripped, entities decoded
|
||||
WordCount int // 0 if the server did not say
|
||||
}
|
||||
|
||||
// Client is a Kiwix HTTP client. Boring on purpose: no retries, no cache.
|
||||
type Client struct {
|
||||
base string
|
||||
http *http.Client
|
||||
}
|
||||
|
||||
// New makes a client for a Kiwix base URL like http://127.0.0.1:8034.
|
||||
func New(baseURL string) *Client {
|
||||
return &Client{
|
||||
base: strings.TrimRight(baseURL, "/"),
|
||||
http: &http.Client{Timeout: 10 * time.Second},
|
||||
}
|
||||
}
|
||||
|
||||
// Search runs a keyword search in one ZIM (book) and returns up to limit hits.
|
||||
//
|
||||
// Ranking is keyword based, not semantic: "Rayleigh scattering" finds the right
|
||||
// article, "why is the sky blue" finds a TV episode. Pass keywords, not questions.
|
||||
func (c *Client) Search(ctx context.Context, pattern, book string, limit int) ([]Result, error) {
|
||||
if limit <= 0 {
|
||||
limit = 5
|
||||
}
|
||||
q := url.Values{}
|
||||
q.Set("pattern", pattern)
|
||||
q.Set("books.name", book)
|
||||
q.Set("format", "xml")
|
||||
q.Set("pageLength", strconv.Itoa(limit))
|
||||
|
||||
req, err := http.NewRequestWithContext(ctx, http.MethodGet, c.base+"/search?"+q.Encode(), nil)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
resp, err := c.http.Do(req)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
defer resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
return nil, fmt.Errorf("kiwix search: http %d", resp.StatusCode)
|
||||
}
|
||||
return ParseSearchRSS(resp.Body)
|
||||
}
|
||||
|
||||
// rss mirrors just the bits of the RSS 2.0 reply we use.
|
||||
type rss struct {
|
||||
Items []struct {
|
||||
Title string `xml:"title"`
|
||||
Link string `xml:"link"`
|
||||
// innerxml keeps the <b> match markers so we can strip them ourselves.
|
||||
Description struct {
|
||||
Inner string `xml:",innerxml"`
|
||||
} `xml:"description"`
|
||||
WordCount string `xml:"wordCount"`
|
||||
} `xml:"channel>item"`
|
||||
}
|
||||
|
||||
var tagRE = regexp.MustCompile(`<[^>]*>`)
|
||||
|
||||
// ParseSearchRSS turns a Kiwix search reply into results. Exported so the parser
|
||||
// is testable from a captured response, with no server running.
|
||||
func ParseSearchRSS(r io.Reader) ([]Result, error) {
|
||||
var doc rss
|
||||
if err := xml.NewDecoder(r).Decode(&doc); err != nil {
|
||||
return nil, fmt.Errorf("kiwix search: bad xml: %w", err)
|
||||
}
|
||||
out := make([]Result, 0, len(doc.Items))
|
||||
for _, it := range doc.Items {
|
||||
n, _ := strconv.Atoi(strings.ReplaceAll(it.WordCount, ",", ""))
|
||||
out = append(out, Result{
|
||||
Title: strings.TrimSpace(it.Title),
|
||||
Path: strings.TrimSpace(it.Link),
|
||||
Snippet: plainText(it.Description.Inner),
|
||||
WordCount: n,
|
||||
})
|
||||
}
|
||||
return out, nil
|
||||
}
|
||||
|
||||
// plainText drops markup and decodes entities, leaving text a model can read.
|
||||
func plainText(s string) string {
|
||||
s = tagRE.ReplaceAllString(s, "")
|
||||
s = html.UnescapeString(s)
|
||||
return strings.TrimSpace(strings.Join(strings.Fields(s), " "))
|
||||
}
|
||||
@@ -0,0 +1,90 @@
|
||||
package kiwix
|
||||
|
||||
import (
|
||||
"context"
|
||||
"os"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// A real reply from the live server, trimmed to two items.
|
||||
const sampleRSS = `<?xml version="1.0" encoding="UTF-8"?>
|
||||
<rss version="2.0" xmlns:opensearch="http://a9.com/-/spec/opensearch/1.1/">
|
||||
<channel>
|
||||
<title>Search: Rayleigh scattering</title>
|
||||
<opensearch:totalResults>800</opensearch:totalResults>
|
||||
<item>
|
||||
<title>Rayleigh scattering</title>
|
||||
<link>/content/wikipedia_en_all_maxi_2026-02/Rayleigh_scattering</link>
|
||||
<description><b>Rayleigh</b> scattering causes the blue color of the sky & yellow colors near the Sun.[1]</description>
|
||||
<book><title>Wikipedia</title></book>
|
||||
<wordCount>2,818</wordCount>
|
||||
</item>
|
||||
<item>
|
||||
<title>Hyper–Rayleigh scattering</title>
|
||||
<link>/content/wikipedia_en_all_maxi_2026-02/Hyper%E2%80%93Rayleigh_scattering</link>
|
||||
<description>...<b>Rayleigh</b> scattering" is a nonlinear optical counterpart.</description>
|
||||
<book><title>Wikipedia</title></book>
|
||||
<wordCount>914</wordCount>
|
||||
</item>
|
||||
</channel>
|
||||
</rss>`
|
||||
|
||||
func TestParseSearchRSS(t *testing.T) {
|
||||
got, err := ParseSearchRSS(strings.NewReader(sampleRSS))
|
||||
if err != nil {
|
||||
t.Fatalf("parse: %v", err)
|
||||
}
|
||||
if len(got) != 2 {
|
||||
t.Fatalf("want 2 results, got %d", len(got))
|
||||
}
|
||||
if got[0].Title != "Rayleigh scattering" {
|
||||
t.Errorf("title = %q", got[0].Title)
|
||||
}
|
||||
if got[0].Path != "/content/wikipedia_en_all_maxi_2026-02/Rayleigh_scattering" {
|
||||
t.Errorf("path = %q", got[0].Path)
|
||||
}
|
||||
if got[0].WordCount != 2818 {
|
||||
t.Errorf("wordCount = %d", got[0].WordCount)
|
||||
}
|
||||
want := "Rayleigh scattering causes the blue color of the sky & yellow colors near the Sun.[1]"
|
||||
if got[0].Snippet != want {
|
||||
t.Errorf("snippet = %q, want %q", got[0].Snippet, want)
|
||||
}
|
||||
if strings.Contains(got[1].Snippet, "<b>") {
|
||||
t.Errorf("second snippet still has tags: %q", got[1].Snippet)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseSearchRSSBadXML(t *testing.T) {
|
||||
if _, err := ParseSearchRSS(strings.NewReader("not xml at all")); err == nil {
|
||||
t.Fatal("want an error on junk input")
|
||||
}
|
||||
}
|
||||
|
||||
// Opt-in: needs a live Kiwix server. CI has none.
|
||||
// MAVEN_KIWIX_URL=http://127.0.0.1:8034 no_proxy=127.0.0.1,localhost go test -run Retrieval -v ./internal/kiwix/
|
||||
func TestRetrievalEval(t *testing.T) {
|
||||
base := os.Getenv("MAVEN_KIWIX_URL")
|
||||
if base == "" {
|
||||
t.Skip("set MAVEN_KIWIX_URL to run the retrieval eval")
|
||||
}
|
||||
noProxyLoopback(t)
|
||||
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute)
|
||||
defer cancel()
|
||||
|
||||
rep, err := RunRetrievalEval(ctx, New(base), 5)
|
||||
if err != nil {
|
||||
t.Fatalf("eval: %v", err)
|
||||
}
|
||||
// No pass bar on purpose: the number is the finding.
|
||||
t.Log("\n" + rep.String() + rep.Detail())
|
||||
}
|
||||
|
||||
// noProxyLoopback stops the box's SOCKS bridge from eating loopback requests.
|
||||
func noProxyLoopback(t *testing.T) {
|
||||
t.Setenv("no_proxy", "127.0.0.1,localhost")
|
||||
t.Setenv("NO_PROXY", "127.0.0.1,localhost")
|
||||
}
|
||||
@@ -0,0 +1,63 @@
|
||||
{
|
||||
"name": "kiwix-knowledge-v1",
|
||||
"book": "wikipedia_en_all_maxi_2026-02",
|
||||
"note": "The 9 knowledge cases from internal/phraser/eval/talk_v1.json. Queries are hand-written English keywords on purpose: Kiwix ranks by keyword, not meaning, so a natural question fails. Writing them by hand separates 'retrieval is broken' from 'the model writes bad queries'.",
|
||||
"cases": [
|
||||
{
|
||||
"id": "know-sky-blue",
|
||||
"question": "почему небо синее?",
|
||||
"query": "Rayleigh scattering sky blue",
|
||||
"want_titles": ["Rayleigh scattering", "Diffuse sky radiation"]
|
||||
},
|
||||
{
|
||||
"id": "know-boil-egg",
|
||||
"question": "сколько варить яйцо вкрутую?",
|
||||
"query": "boiled egg cooking",
|
||||
"want_titles": ["Boiled egg", "Egg as food"]
|
||||
},
|
||||
{
|
||||
"id": "know-ssd-vs-hdd",
|
||||
"question": "чем ssd отличается от hdd?",
|
||||
"query": "solid-state drive",
|
||||
"want_titles": ["Solid-state drive", "Hard disk drive"]
|
||||
},
|
||||
{
|
||||
"id": "know-cat-purr",
|
||||
"question": "почему кошки мурчат?",
|
||||
"query": "cat purr",
|
||||
"want_titles": ["Purr", "Cat communication"]
|
||||
},
|
||||
{
|
||||
"id": "know-hiccups",
|
||||
"question": "как быстро избавиться от икоты?",
|
||||
"query": "hiccup",
|
||||
"want_titles": ["Hiccup"]
|
||||
},
|
||||
{
|
||||
"id": "know-polite-form",
|
||||
"question": "не могли бы вы объяснить, что такое vpn?",
|
||||
"query": "virtual private network",
|
||||
"want_titles": ["Virtual private network"]
|
||||
},
|
||||
{
|
||||
"id": "know-dont-know",
|
||||
"question": "как зовут моего соседа снизу?",
|
||||
"query": "name of my downstairs neighbour",
|
||||
"want_titles": [],
|
||||
"expect_miss": true,
|
||||
"note": "Unanswerable by design. Retrieval SHOULD find nothing useful. Counted as a hit only when nothing relevant comes back."
|
||||
},
|
||||
{
|
||||
"id": "know-water-per-day",
|
||||
"question": "сколько воды в день надо пить?",
|
||||
"query": "human daily water requirement drinking",
|
||||
"want_titles": ["Drinking water", "Water", "Dehydration", "Hydration"]
|
||||
},
|
||||
{
|
||||
"id": "know-thunder-delay",
|
||||
"question": "почему гром слышно позже молнии?",
|
||||
"query": "thunder speed of sound lightning",
|
||||
"want_titles": ["Thunder", "Lightning"]
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,135 @@
|
||||
package kiwix
|
||||
|
||||
// This scores retrieval alone: no LLM. For each general-knowledge question we
|
||||
// hand-write English keywords and ask whether the article that would answer it
|
||||
// comes back in the top N hits. If this score is low, reading Wikipedia cannot
|
||||
// help the model no matter how good the prompt is.
|
||||
//
|
||||
// The unanswerable case (know-dont-know) is not scored. Whether the junk it
|
||||
// returns is "nothing useful" is a human judgement, so the report just prints
|
||||
// the titles and leaves the score to the 8 answerable cases.
|
||||
|
||||
import (
|
||||
"context"
|
||||
_ "embed"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"strings"
|
||||
)
|
||||
|
||||
//go:embed knowledge_v1.json
|
||||
var knowledgeFixtureJSON []byte
|
||||
|
||||
// EvalCase — one question with hand-written keywords.
|
||||
type EvalCase struct {
|
||||
ID string `json:"id"`
|
||||
Question string `json:"question"`
|
||||
Query string `json:"query"`
|
||||
WantTitles []string `json:"want_titles"`
|
||||
ExpectMiss bool `json:"expect_miss"`
|
||||
}
|
||||
|
||||
type fixture struct {
|
||||
Name string `json:"name"`
|
||||
Book string `json:"book"`
|
||||
Cases []EvalCase `json:"cases"`
|
||||
}
|
||||
|
||||
// Outcome — what one case retrieved.
|
||||
type Outcome struct {
|
||||
Case EvalCase
|
||||
Titles []string // titles of the top N hits, in rank order
|
||||
Rank int // 1-based rank of the first wanted title, 0 if none
|
||||
Err error
|
||||
}
|
||||
|
||||
// Hit is true when a wanted title came back.
|
||||
func (o Outcome) Hit() bool { return o.Rank > 0 }
|
||||
|
||||
// Report — the score plus per-case detail.
|
||||
type Report struct {
|
||||
Name string
|
||||
Book string
|
||||
TopN int
|
||||
Scored int // answerable cases
|
||||
Hits int
|
||||
Errors int
|
||||
Outcomes []Outcome
|
||||
}
|
||||
|
||||
// Accuracy over the answerable cases.
|
||||
func (r Report) Accuracy() float64 {
|
||||
if r.Scored == 0 {
|
||||
return 0
|
||||
}
|
||||
return float64(r.Hits) / float64(r.Scored)
|
||||
}
|
||||
|
||||
// RunRetrievalEval searches for every fixture case.
|
||||
func RunRetrievalEval(ctx context.Context, c *Client, topN int) (Report, error) {
|
||||
var f fixture
|
||||
if err := json.Unmarshal(knowledgeFixtureJSON, &f); err != nil {
|
||||
return Report{}, err
|
||||
}
|
||||
rep := Report{Name: f.Name, Book: f.Book, TopN: topN}
|
||||
for _, cs := range f.Cases {
|
||||
res, err := c.Search(ctx, cs.Query, f.Book, topN)
|
||||
o := Outcome{Case: cs, Err: err}
|
||||
if err != nil {
|
||||
rep.Errors++
|
||||
}
|
||||
for i, hit := range res {
|
||||
o.Titles = append(o.Titles, hit.Title)
|
||||
if o.Rank == 0 && matches(cs.WantTitles, hit.Title) {
|
||||
o.Rank = i + 1
|
||||
}
|
||||
}
|
||||
if !cs.ExpectMiss {
|
||||
rep.Scored++
|
||||
if o.Hit() {
|
||||
rep.Hits++
|
||||
}
|
||||
}
|
||||
rep.Outcomes = append(rep.Outcomes, o)
|
||||
}
|
||||
return rep, nil
|
||||
}
|
||||
|
||||
func matches(want []string, title string) bool {
|
||||
for _, w := range want {
|
||||
if strings.EqualFold(strings.TrimSpace(title), w) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// String — the headline number.
|
||||
func (r Report) String() string {
|
||||
var b strings.Builder
|
||||
fmt.Fprintf(&b, "%s: %d/%d answerable questions retrieve a wanted article in top %d (%.1f%%), %d errors\n",
|
||||
r.Name, r.Hits, r.Scored, r.TopN, 100*r.Accuracy(), r.Errors)
|
||||
fmt.Fprintf(&b, " book: %s\n", r.Book)
|
||||
return b.String()
|
||||
}
|
||||
|
||||
// Detail — per case: what was asked, what was searched, what came back.
|
||||
func (r Report) Detail() string {
|
||||
var b strings.Builder
|
||||
for _, o := range r.Outcomes {
|
||||
mark := "MISS"
|
||||
switch {
|
||||
case o.Case.ExpectMiss:
|
||||
mark = "n/a "
|
||||
case o.Hit():
|
||||
mark = fmt.Sprintf("hit@%d", o.Rank)
|
||||
}
|
||||
fmt.Fprintf(&b, " %-6s %-20s q=%q\n", mark, o.Case.ID, o.Case.Query)
|
||||
if o.Err != nil {
|
||||
fmt.Fprintf(&b, " error: %v\n", o.Err)
|
||||
continue
|
||||
}
|
||||
fmt.Fprintf(&b, " got: %s\n", strings.Join(o.Titles, " | "))
|
||||
}
|
||||
return b.String()
|
||||
}
|
||||
@@ -0,0 +1,183 @@
|
||||
package kiwix
|
||||
|
||||
// Turning a Russian question into an English Kiwix search.
|
||||
//
|
||||
// Kiwix ranks by keyword, not by meaning. "why is the sky blue" returns a TV
|
||||
// episode; "Rayleigh scattering sky blue" returns the right article. So the
|
||||
// model's job here is NOT translation — it is naming the English article the
|
||||
// answer lives in.
|
||||
//
|
||||
// The output space is a handful of words, so it is worth locking down hard: a
|
||||
// GBNF grammar for the shape, a tiny token cap, and a cleanup pass that throws
|
||||
// away anything odd rather than handing junk to Kiwix.
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"strings"
|
||||
"unicode"
|
||||
|
||||
"github.com/kami/maven/internal/llm"
|
||||
)
|
||||
|
||||
// Completer — the LLM seam, so tests can fake it. *llm.Client satisfies it.
|
||||
type Completer interface {
|
||||
Complete(ctx context.Context, r llm.Req) (string, error)
|
||||
}
|
||||
|
||||
// queryGrammar — one JSON object holding 1..6 keyword words. Latin letters,
|
||||
// digits and hyphens only, so the model physically cannot answer the question
|
||||
// or reply in Russian.
|
||||
//
|
||||
// Why the JSON wrapper: this model always thinks out loud and this llama-server
|
||||
// build ignores the thinking switch (see ROUTING-EVAL-31-07-2026.md). A bare
|
||||
// word-list grammar just captured the reasoning — every case came back as
|
||||
// "Let me analyze this request carefully". Demanding JSON, like routeGrammar and
|
||||
// responseGrammar already do, gives the reasoning nowhere to go.
|
||||
const queryGrammar = `
|
||||
root ::= "{" ws "\"query\"" ws ":" ws "\"" word (" " word){0,5} "\"" ws "}"
|
||||
word ::= [A-Za-z0-9] [A-Za-z0-9-]{0,23}
|
||||
ws ::= [ \t\n]*
|
||||
`
|
||||
|
||||
// rewriteSystem — asks for search keywords, not an answer and not a translation.
|
||||
const rewriteSystem = `You turn a question into a search query for English Wikipedia.
|
||||
|
||||
Rules:
|
||||
- Output ONLY English search keywords. Never an answer, never an explanation.
|
||||
- Do NOT translate the sentence. Name the thing the answer is about.
|
||||
- The output must be a noun phrase, like a Wikipedia article title.
|
||||
- Never use question words: no why, how, what, when, which, "how much",
|
||||
"how long", "how to", "vs", "reason", "difference".
|
||||
- 2 to 4 words.
|
||||
|
||||
Reply with JSON: {"query":"<keywords>"}
|
||||
|
||||
Good:
|
||||
"почему листья желтеют осенью?" -> {"query":"leaf senescence autumn"}
|
||||
"как работает микроволновка?" -> {"query":"microwave oven"}
|
||||
"не могли бы вы объяснить, что такое блокчейн?" -> {"query":"blockchain"}
|
||||
"сколько живут собаки?" -> {"query":"dog lifespan"}
|
||||
"как избавиться от комаров в квартире?" -> {"query":"mosquito control"}
|
||||
"чем чай отличается от кофе?" -> {"query":"tea"}
|
||||
|
||||
Only JSON, no explanation.`
|
||||
|
||||
// maxQueryTokens — the output is a few words plus the JSON wrapper. A tight cap
|
||||
// is the cheapest guard against the model rambling into an answer.
|
||||
const maxQueryTokens = 32
|
||||
|
||||
// Rewriter asks the resident model for English search keywords.
|
||||
type Rewriter struct{ c Completer }
|
||||
|
||||
func NewRewriter(c Completer) *Rewriter { return &Rewriter{c: c} }
|
||||
|
||||
// Rewrite returns English keywords for a question in any language.
|
||||
// It errors rather than returning something Kiwix should not see.
|
||||
func (r *Rewriter) Rewrite(ctx context.Context, question string) (string, error) {
|
||||
raw, err := r.c.Complete(ctx, llm.Req{
|
||||
System: rewriteSystem,
|
||||
User: strings.TrimSpace(question),
|
||||
Grammar: queryGrammar,
|
||||
MaxTokens: maxQueryTokens,
|
||||
RepeatPenalty: 1.15,
|
||||
})
|
||||
if err != nil {
|
||||
return "", err
|
||||
}
|
||||
return CleanQuery(unwrapJSON(raw))
|
||||
}
|
||||
|
||||
// unwrapJSON pulls the query out of {"query":"..."}. If the reply is not that
|
||||
// shape it is returned as-is, and CleanQuery decides whether it is usable.
|
||||
func unwrapJSON(raw string) string {
|
||||
s := strings.TrimSpace(raw)
|
||||
if !strings.HasPrefix(s, "{") {
|
||||
return s
|
||||
}
|
||||
var got struct{ Query string }
|
||||
if err := json.Unmarshal([]byte(s), &got); err != nil {
|
||||
return s
|
||||
}
|
||||
return got.Query
|
||||
}
|
||||
|
||||
// maxQueryWords matches the grammar's bound. Anything longer is prose.
|
||||
const maxQueryWords = 6
|
||||
|
||||
// CleanQuery checks and tidies whatever the model produced. The grammar makes
|
||||
// bad output unlikely, not impossible (a server without grammar support, a
|
||||
// different model), so this is the real gate in front of Kiwix.
|
||||
//
|
||||
// Exported so it can be tested without a model.
|
||||
func CleanQuery(raw string) (string, error) {
|
||||
s := strings.TrimSpace(raw)
|
||||
// Models like to wrap answers in quotes. Drop surrounding ones.
|
||||
s = strings.Trim(s, "\"'`")
|
||||
// Keep the first line only: everything after it is prose.
|
||||
if i := strings.IndexAny(s, "\r\n"); i >= 0 {
|
||||
s = s[:i]
|
||||
}
|
||||
// Keep letters, digits, spaces and hyphens; anything else becomes a space.
|
||||
var b strings.Builder
|
||||
for _, ru := range s {
|
||||
switch {
|
||||
case unicode.IsLetter(ru) || unicode.IsDigit(ru) || ru == '-':
|
||||
b.WriteRune(ru)
|
||||
default:
|
||||
b.WriteRune(' ')
|
||||
}
|
||||
}
|
||||
words := strings.Fields(b.String())
|
||||
if len(words) == 0 {
|
||||
return "", fmt.Errorf("kiwix rewrite: empty query")
|
||||
}
|
||||
if len(words) > maxQueryWords {
|
||||
return "", fmt.Errorf("kiwix rewrite: %d words, want at most %d (looks like prose)", len(words), maxQueryWords)
|
||||
}
|
||||
words = dropStopWords(words)
|
||||
out := strings.Join(words, " ")
|
||||
// The ZIMs are English. Non-Latin letters mean the model ignored the ask.
|
||||
for _, ru := range out {
|
||||
if unicode.IsLetter(ru) && !isLatin(ru) {
|
||||
return "", fmt.Errorf("kiwix rewrite: query is not English: %q", out)
|
||||
}
|
||||
}
|
||||
return out, nil
|
||||
}
|
||||
|
||||
// stopWords — question words and filler. The model keeps writing question-shaped
|
||||
// queries ("why is the sky blue", "how much water to drink daily") no matter how
|
||||
// the prompt is worded, and Kiwix ranks on every word, so those words drag in
|
||||
// song and episode titles. Dropping them in code is not a style preference: a
|
||||
// keyword ranker gets nothing from them.
|
||||
var stopWords = map[string]bool{
|
||||
"a": true, "an": true, "the": true, "is": true, "are": true, "was": true,
|
||||
"do": true, "does": true, "did": true, "to": true, "of": true, "in": true,
|
||||
"on": true, "for": true, "and": true, "or": true, "my": true, "me": true,
|
||||
"i": true, "it": true, "its": true, "be": true, "been": true, "get": true,
|
||||
"how": true, "why": true, "what": true, "when": true, "which": true,
|
||||
"who": true, "where": true, "much": true, "many": true, "long": true,
|
||||
"vs": true, "than": true, "rid": true, "from": true, "about": true,
|
||||
}
|
||||
|
||||
// dropStopWords removes filler, but never everything: if the query was nothing
|
||||
// but stop words there is nothing better to search, so the original is kept and
|
||||
// the caller sees whatever Kiwix makes of it.
|
||||
func dropStopWords(words []string) []string {
|
||||
kept := make([]string, 0, len(words))
|
||||
for _, w := range words {
|
||||
if !stopWords[strings.ToLower(w)] {
|
||||
kept = append(kept, w)
|
||||
}
|
||||
}
|
||||
if len(kept) == 0 {
|
||||
return words
|
||||
}
|
||||
return kept
|
||||
}
|
||||
|
||||
func isLatin(ru rune) bool {
|
||||
return (ru >= 'a' && ru <= 'z') || (ru >= 'A' && ru <= 'Z')
|
||||
}
|
||||
@@ -0,0 +1,96 @@
|
||||
package kiwix
|
||||
|
||||
// End-to-end score: Russian question -> model rewrite -> Kiwix search -> did a
|
||||
// wanted article come back. Same 9 cases as the retrieval eval, so the two
|
||||
// numbers are directly comparable: retrieval with hand-written keywords is the
|
||||
// ceiling, this is what the model actually reaches.
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// RewriteOutcome — one case, end to end.
|
||||
type RewriteOutcome struct {
|
||||
Outcome
|
||||
ModelQuery string // what the model asked for ("" if it failed)
|
||||
RewriteErr error
|
||||
}
|
||||
|
||||
// RunRewriteEval rewrites every question with the model, then searches.
|
||||
func RunRewriteEval(ctx context.Context, c *Client, rw *Rewriter, topN int) (RewriteReport, error) {
|
||||
var f fixture
|
||||
if err := json.Unmarshal(knowledgeFixtureJSON, &f); err != nil {
|
||||
return RewriteReport{}, err
|
||||
}
|
||||
rep := RewriteReport{Report: Report{Name: f.Name + "-rewrite", Book: f.Book, TopN: topN}}
|
||||
for _, cs := range f.Cases {
|
||||
out := RewriteOutcome{Outcome: Outcome{Case: cs}}
|
||||
q, err := rw.Rewrite(ctx, cs.Question)
|
||||
out.ModelQuery, out.RewriteErr = q, err
|
||||
if err == nil {
|
||||
res, serr := c.Search(ctx, q, f.Book, topN)
|
||||
out.Err = serr
|
||||
for i, hit := range res {
|
||||
out.Titles = append(out.Titles, hit.Title)
|
||||
if out.Rank == 0 && matches(cs.WantTitles, hit.Title) {
|
||||
out.Rank = i + 1
|
||||
}
|
||||
}
|
||||
}
|
||||
if out.RewriteErr != nil || out.Err != nil {
|
||||
rep.Errors++
|
||||
}
|
||||
if !cs.ExpectMiss {
|
||||
rep.Scored++
|
||||
if out.Hit() {
|
||||
rep.Hits++
|
||||
}
|
||||
}
|
||||
rep.Cases = append(rep.Cases, out)
|
||||
}
|
||||
return rep, nil
|
||||
}
|
||||
|
||||
// RewriteReport — the score plus per-case detail.
|
||||
type RewriteReport struct {
|
||||
Report
|
||||
Cases []RewriteOutcome
|
||||
}
|
||||
|
||||
// String — the headline number.
|
||||
func (r RewriteReport) String() string {
|
||||
return fmt.Sprintf("%s: %d/%d answerable questions retrieve a wanted article in top %d (%.1f%%), %d errors\n book: %s\n",
|
||||
r.Name, r.Hits, r.Scored, r.TopN, 100*r.Accuracy(), r.Errors, r.Book)
|
||||
}
|
||||
|
||||
// Detail — per case: hand-written query next to the model's, and what came back.
|
||||
// The point is seeing WHERE the model's phrasing differs, not just the score.
|
||||
func (r RewriteReport) Detail() string {
|
||||
var b strings.Builder
|
||||
for _, o := range r.Cases {
|
||||
mark := "MISS"
|
||||
switch {
|
||||
case o.Case.ExpectMiss:
|
||||
mark = "n/a "
|
||||
case o.Hit():
|
||||
mark = fmt.Sprintf("hit@%d", o.Rank)
|
||||
}
|
||||
fmt.Fprintf(&b, " %-6s %-20s\n", mark, o.Case.ID)
|
||||
fmt.Fprintf(&b, " asked: %s\n", o.Case.Question)
|
||||
fmt.Fprintf(&b, " hand: %q\n", o.Case.Query)
|
||||
fmt.Fprintf(&b, " model: %q\n", o.ModelQuery)
|
||||
if o.RewriteErr != nil {
|
||||
fmt.Fprintf(&b, " rewrite rejected: %v\n", o.RewriteErr)
|
||||
continue
|
||||
}
|
||||
if o.Err != nil {
|
||||
fmt.Fprintf(&b, " search error: %v\n", o.Err)
|
||||
continue
|
||||
}
|
||||
fmt.Fprintf(&b, " got: %s\n", strings.Join(o.Titles, " | "))
|
||||
}
|
||||
return b.String()
|
||||
}
|
||||
@@ -0,0 +1,33 @@
|
||||
package kiwix
|
||||
|
||||
import (
|
||||
"context"
|
||||
"os"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/llm"
|
||||
)
|
||||
|
||||
// Opt-in: needs a live Kiwix server AND a live llama-server.
|
||||
// MAVEN_KIWIX_URL=http://127.0.0.1:8034 MAVEN_LLM_URL=http://127.0.0.1:18099 \
|
||||
//
|
||||
// no_proxy=127.0.0.1,localhost go test -run RewriteEval -v ./internal/kiwix/
|
||||
func TestRewriteEval(t *testing.T) {
|
||||
kbase, lbase := os.Getenv("MAVEN_KIWIX_URL"), os.Getenv("MAVEN_LLM_URL")
|
||||
if kbase == "" || lbase == "" {
|
||||
t.Skip("set MAVEN_KIWIX_URL and MAVEN_LLM_URL to run the rewrite eval")
|
||||
}
|
||||
noProxyLoopback(t)
|
||||
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 15*time.Minute)
|
||||
defer cancel()
|
||||
|
||||
rw := NewRewriter(llm.New(lbase, 3*time.Minute))
|
||||
rep, err := RunRewriteEval(ctx, New(kbase), rw, 5)
|
||||
if err != nil {
|
||||
t.Fatalf("eval: %v", err)
|
||||
}
|
||||
// No pass bar on purpose: the number is the finding.
|
||||
t.Log("\n" + rep.String() + rep.Detail())
|
||||
}
|
||||
@@ -0,0 +1,105 @@
|
||||
package kiwix
|
||||
|
||||
import (
|
||||
"context"
|
||||
"testing"
|
||||
|
||||
"github.com/kami/maven/internal/llm"
|
||||
)
|
||||
|
||||
// Bad model output must never reach Kiwix. No model needed for this.
|
||||
func TestCleanQueryRejectsJunk(t *testing.T) {
|
||||
bad := []struct{ name, raw string }{
|
||||
{"empty", ""},
|
||||
{"blank", " \n "},
|
||||
{"russian came back", "почему небо синее"},
|
||||
{"mixed russian", "sky синее scattering"},
|
||||
{"full sentence", "The sky looks blue because of the scattering of sunlight by air molecules"},
|
||||
{"prose with quotes", `Sure! Here is a good search query: "Rayleigh scattering", which explains it.`},
|
||||
}
|
||||
for _, c := range bad {
|
||||
if got, err := CleanQuery(c.raw); err == nil {
|
||||
t.Errorf("%s: want rejection, got %q", c.name, got)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestCleanQueryCleans(t *testing.T) {
|
||||
ok := []struct{ raw, want string }{
|
||||
{"Rayleigh scattering sky", "Rayleigh scattering sky"},
|
||||
{" boiled egg cooking \n", "boiled egg cooking"},
|
||||
{`"virtual private network"`, "virtual private network"},
|
||||
{"solid-state drive", "solid-state drive"},
|
||||
{"cat purr.", "cat purr"},
|
||||
{"hiccup\nAlso: hiccough", "hiccup"},
|
||||
// Question words are filler to a keyword ranker, so they go.
|
||||
{"why is the sky blue", "sky blue"},
|
||||
{"how much water to drink daily", "water drink daily"},
|
||||
{"SSD vs HDD comparison", "SSD HDD comparison"},
|
||||
// Nothing but filler: keep it rather than return nothing.
|
||||
{"what is it", "what is it"},
|
||||
}
|
||||
for _, c := range ok {
|
||||
got, err := CleanQuery(c.raw)
|
||||
if err != nil {
|
||||
t.Errorf("%q: %v", c.raw, err)
|
||||
continue
|
||||
}
|
||||
if got != c.want {
|
||||
t.Errorf("%q -> %q, want %q", c.raw, got, c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
type fakeCompleter struct {
|
||||
out string
|
||||
req llm.Req
|
||||
}
|
||||
|
||||
func (f *fakeCompleter) Complete(_ context.Context, r llm.Req) (string, error) {
|
||||
f.req = r
|
||||
return f.out, nil
|
||||
}
|
||||
|
||||
func TestRewriteConstrainsTheCall(t *testing.T) {
|
||||
f := &fakeCompleter{out: `{"query":"Rayleigh scattering sky"}`}
|
||||
got, err := NewRewriter(f).Rewrite(context.Background(), "почему небо синее?")
|
||||
if err != nil {
|
||||
t.Fatalf("rewrite: %v", err)
|
||||
}
|
||||
if got != "Rayleigh scattering sky" {
|
||||
t.Errorf("query = %q", got)
|
||||
}
|
||||
if f.req.Grammar == "" {
|
||||
t.Error("no grammar sent")
|
||||
}
|
||||
if f.req.MaxTokens == 0 || f.req.MaxTokens > 32 {
|
||||
t.Errorf("max_tokens = %d, want a small cap", f.req.MaxTokens)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRewriteRejectsBadModelOutput(t *testing.T) {
|
||||
bad := []string{
|
||||
`{"query":"почему небо синее"}`, // never translated
|
||||
`{"query":""}`, // empty
|
||||
`{"query":"the sky is blue because sunlight is scattered by air"}`, // an answer
|
||||
// Note: a SHORT English prose fragment ("Let me analyze this request")
|
||||
// is under the word cap and cannot be caught here. The grammar is what
|
||||
// stops that one.
|
||||
}
|
||||
for _, out := range bad {
|
||||
f := &fakeCompleter{out: out}
|
||||
if got, err := NewRewriter(f).Rewrite(context.Background(), "почему небо синее?"); err == nil {
|
||||
t.Errorf("%s: want rejection, got %q", out, got)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A reply that is not the JSON shape but is still usable keywords should pass.
|
||||
func TestRewriteFallsBackToPlainText(t *testing.T) {
|
||||
f := &fakeCompleter{out: "Rayleigh scattering sky"}
|
||||
got, err := NewRewriter(f).Rewrite(context.Background(), "почему небо синее?")
|
||||
if err != nil || got != "Rayleigh scattering sky" {
|
||||
t.Errorf("got %q, %v", got, err)
|
||||
}
|
||||
}
|
||||
@@ -1,4 +1,4 @@
|
||||
package eval
|
||||
package llm
|
||||
|
||||
import (
|
||||
"context"
|
||||
@@ -6,8 +6,20 @@ import (
|
||||
"fmt"
|
||||
"net/http"
|
||||
"strings"
|
||||
"time"
|
||||
)
|
||||
|
||||
// UnknownModel is the label to print when the server would not say what it has
|
||||
// loaded. Deliberately ugly: an honest "unknown" is fine, a plausible-looking
|
||||
// but wrong model name is the bug this whole file exists to prevent.
|
||||
const UnknownModel = "unknown-model"
|
||||
|
||||
// llama-server is local, so never send this through a proxy: this box's
|
||||
// http_proxy answers 503 for loopback, which would look like "server won't say
|
||||
// which model it has" when the server is right there and fine.
|
||||
// A Transport with no Proxy set bypasses http_proxy entirely.
|
||||
var modelHTTP = &http.Client{Timeout: 10 * time.Second, Transport: &http.Transport{}}
|
||||
|
||||
// ModelID asks llama-server which model it has loaded, so a scoring run can
|
||||
// label itself. Without this a bake-off between two models produces two tables
|
||||
// that look identical, and the operator has to remember which server was up.
|
||||
@@ -19,7 +31,7 @@ func ModelID(ctx context.Context, base string) (string, error) {
|
||||
if err != nil {
|
||||
return "", err
|
||||
}
|
||||
resp, err := http.DefaultClient.Do(req)
|
||||
resp, err := modelHTTP.Do(req)
|
||||
if err != nil {
|
||||
return "", err
|
||||
}
|
||||
@@ -38,14 +50,21 @@ func ModelID(ctx context.Context, base string) (string, error) {
|
||||
if len(out.Data) == 0 {
|
||||
return "", fmt.Errorf("models: empty list")
|
||||
}
|
||||
return shortModelID(out.Data[0].ID), nil
|
||||
short := shortModelID(out.Data[0].ID)
|
||||
if short == "" {
|
||||
// Server answered but the id field was missing or blank. Say so
|
||||
// instead of handing back an empty label that reads as a real name.
|
||||
return "", fmt.Errorf("models: no id in response")
|
||||
}
|
||||
return short, nil
|
||||
}
|
||||
|
||||
// shortModelID trims the path and the .gguf suffix — llama-server reports the
|
||||
// file name it was started with, which is too long for a table header.
|
||||
func shortModelID(id string) string {
|
||||
id = strings.TrimSpace(id)
|
||||
if i := strings.LastIndexAny(id, "/\\"); i >= 0 {
|
||||
id = id[i+1:]
|
||||
}
|
||||
return strings.TrimSuffix(id, ".gguf")
|
||||
return strings.TrimSpace(strings.TrimSuffix(id, ".gguf"))
|
||||
}
|
||||
@@ -0,0 +1,64 @@
|
||||
package llm
|
||||
|
||||
import (
|
||||
"context"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// The point of these tests: a wrong-but-plausible model label is the bug, so
|
||||
// every path that cannot learn the real name must return an error instead of a
|
||||
// guess. No llama-server needed — a stub server stands in.
|
||||
func TestModelID(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
body string
|
||||
code int
|
||||
want string // "" ⇒ expect an error
|
||||
}{
|
||||
{"full path", `{"data":[{"id":"/mnt/hdd1/llms/qwen3.5/Qwen3.5-0.8B.Q4_K_M.gguf"}]}`, 200, "Qwen3.5-0.8B.Q4_K_M"},
|
||||
{"bare name", `{"data":[{"id":"LFM2.5-1.2B"}]}`, 200, "LFM2.5-1.2B"},
|
||||
{"empty list", `{"data":[]}`, 200, ""},
|
||||
{"id missing", `{"data":[{}]}`, 200, ""},
|
||||
{"id blank", `{"data":[{"id":" "}]}`, 200, ""},
|
||||
{"server error", `nope`, 500, ""},
|
||||
{"not json", `<html>`, 200, ""},
|
||||
}
|
||||
for _, c := range cases {
|
||||
t.Run(c.name, func(t *testing.T) {
|
||||
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
if r.URL.Path != "/v1/models" {
|
||||
t.Errorf("asked for %s, want /v1/models", r.URL.Path)
|
||||
}
|
||||
w.WriteHeader(c.code)
|
||||
_, _ = w.Write([]byte(c.body))
|
||||
}))
|
||||
defer srv.Close()
|
||||
|
||||
got, err := ModelID(context.Background(), srv.URL+"/")
|
||||
if c.want == "" {
|
||||
if err == nil {
|
||||
t.Fatalf("want an error, got label %q", got)
|
||||
}
|
||||
return
|
||||
}
|
||||
if err != nil {
|
||||
t.Fatalf("ModelID: %v", err)
|
||||
}
|
||||
if got != c.want {
|
||||
t.Errorf("got %q, want %q", got, c.want)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestModelIDUnreachable(t *testing.T) {
|
||||
srv := httptest.NewServer(http.HandlerFunc(func(http.ResponseWriter, *http.Request) {}))
|
||||
url := srv.URL
|
||||
srv.Close() // nothing listening now
|
||||
|
||||
if got, err := ModelID(context.Background(), url); err == nil {
|
||||
t.Fatalf("want an error from a dead server, got label %q", got)
|
||||
}
|
||||
}
|
||||
@@ -422,6 +422,8 @@ func rankNote(inTop3 bool) string {
|
||||
// bestRecall mirrors cmd/mavend/recall.go — the gate the daemon actually
|
||||
// applies to a memory hit. Duplicated rather than imported because package main
|
||||
// is not importable; recalleval_test.go asserts the two agree in behaviour.
|
||||
// The daemon returns the whole hit (a note and a fact are said differently);
|
||||
// the harness only scores what came back, so it keeps returning the text.
|
||||
func bestRecall(results []memory.Result, minScore, minMargin float64) string {
|
||||
if !memory.Confident(results, minScore, minMargin) {
|
||||
return ""
|
||||
|
||||
@@ -197,7 +197,8 @@ func TestHashRecallBaseline(t *testing.T) {
|
||||
t.Log("\n" + rep.String() + rep.Failures())
|
||||
t.Log("\ngate sweep:\n" + sweep(t, router.NewHashEmbedder(hashDim), f))
|
||||
|
||||
// 0.32 sits under the observed 0.360 recall@1.
|
||||
// 0.32 sits under the observed 0.370 recall@1 (was 0.360 over 25 answerable
|
||||
// cases; the two mixed note+fact cases added with #373 make it 27).
|
||||
const floorRecall1 = 0.32
|
||||
if rep.Recall1() < floorRecall1 {
|
||||
t.Errorf("recall@1 %.3f below ratchet %.2f — note recall regressed", rep.Recall1(), floorRecall1)
|
||||
|
||||
@@ -387,6 +387,32 @@
|
||||
{"id": "n2", "text": "wifi channel is 6", "kind": "note"},
|
||||
{"id": "n3", "text": "the guest network is off", "kind": "note"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ru-mixed-031",
|
||||
"lang": "ru",
|
||||
"tags": ["mixed", "paraphrase", "hard"],
|
||||
"note": "notes and facts in one store and the note is the answer — the daemon indexes both (Vikunja #373)",
|
||||
"query": "куда я спрятал второй ключ от квартиры",
|
||||
"want": "n1",
|
||||
"notes": [
|
||||
{"id": "n1", "text": "запасной ключ от квартиры лежит в синей коробке на полке", "kind": "note"},
|
||||
{"id": "x1", "text": "поменял замок в двери двадцатого июня", "kind": "fact"},
|
||||
{"id": "x2", "text": "отдал ключ соседке в мае", "kind": "fact"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ru-mixed-032",
|
||||
"lang": "ru",
|
||||
"tags": ["mixed", "distractor"],
|
||||
"note": "the mirror of ru-mixed-031: the fact answers and the notes are the distractors",
|
||||
"query": "когда я в последний раз заливал бензин",
|
||||
"want": "x1",
|
||||
"notes": [
|
||||
{"id": "x1", "text": "залил полный бак в четверг вечером", "kind": "fact"},
|
||||
{"id": "n1", "text": "на заправке у моста дешевле бензин", "kind": "note"},
|
||||
{"id": "n2", "text": "надо поменять зимние шины", "kind": "note"}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
|
||||
@@ -0,0 +1,138 @@
|
||||
// Package persona builds the one shared context block that goes in front of
|
||||
// every LLM system prompt: who the owner is, how to address him, and what
|
||||
// time it is right now.
|
||||
//
|
||||
// Why one block and not a line pasted into each prompt: there are five
|
||||
// prompts (nudges, action replies, chat, note queries, general knowledge) and
|
||||
// the "address him as ты" rule had only reached two of them. Five copies drift.
|
||||
// One block cannot.
|
||||
//
|
||||
// The rules here are defaults in code, not config. Maven is feminine and the
|
||||
// owner is a man addressed informally — that is a hard constraint of the
|
||||
// product, so it must hold with an empty config file. Config only ADDS
|
||||
// optional facts (his name, his city).
|
||||
package persona
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"strings"
|
||||
"time"
|
||||
)
|
||||
|
||||
// Facts — the optional, deployment-specific half of the block. All fields may
|
||||
// be empty; the block is still correct and useful without them.
|
||||
type Facts struct {
|
||||
OwnerName string // his name, e.g. "Ками"
|
||||
City string // where he is, e.g. "Москва"
|
||||
Static string // the free-text `persona` config string, appended verbatim
|
||||
|
||||
// The two config-gated capabilities. They are listed only when this
|
||||
// deployment actually has them, because a capability she names and cannot
|
||||
// do is worse than one she never mentions.
|
||||
Weather bool // an open-meteo provider is configured
|
||||
Telegram bool // a telegram bot token + chat id are configured
|
||||
Tools bool // at least one shell act is on the allowlist
|
||||
}
|
||||
|
||||
var ruWeekdays = [...]string{"воскресенье", "понедельник", "вторник", "среда", "четверг", "пятница", "суббота"}
|
||||
|
||||
var ruMonths = [...]string{
|
||||
"января", "февраля", "марта", "апреля", "мая", "июня",
|
||||
"июля", "августа", "сентября", "октября", "ноября", "декабря",
|
||||
}
|
||||
|
||||
// Block renders the context block for one turn. Russian even in front of the
|
||||
// English prompts: the rules it states are Russian grammar (ты/тебя, feminine
|
||||
// verbs), and a Russian rule reads best stated in Russian.
|
||||
//
|
||||
// Keep it short. It ships on every turn to a 0.8B on laptop CPU, so every
|
||||
// line here is latency.
|
||||
func (f Facts) Block(now time.Time) string {
|
||||
var b strings.Builder
|
||||
|
||||
b.WriteString("Ты — Maven, домашняя ассистентка. О себе говоришь в женском роде: \"я записала\", \"я проверила\".\n")
|
||||
|
||||
// The address form gets its own line. It is the thing that kept getting
|
||||
// lost when it was buried in prose.
|
||||
b.WriteString("ОБРАЩЕНИЕ: владелец — мужчина, всегда на \"ты\" (ты, тебя, тебе, твой) и в единственном числе (\"выпей\", \"посмотри\"). Никогда \"вы\"/\"вас\"/\"ваш\". Никогда \"он\"/\"его\" о нём — ты говоришь ему, а не о нём. Глаголы о нём — в мужском роде (\"ты забыл\").\n")
|
||||
|
||||
if who := f.who(); who != "" {
|
||||
b.WriteString(who + "\n")
|
||||
}
|
||||
|
||||
b.WriteString(fmt.Sprintf("Сейчас: %s, %d %s %d, %02d:%02d (местное время).\n",
|
||||
ruWeekdays[int(now.Weekday())], now.Day(), ruMonths[int(now.Month())-1], now.Year(),
|
||||
now.Hour(), now.Minute()))
|
||||
|
||||
b.WriteString("Умеешь: " + strings.Join(f.can(), "; ") +
|
||||
". Других ДЕЙСТВИЙ не умеешь — если просят такое, скажи прямо.\n")
|
||||
|
||||
if s := strings.TrimSpace(f.Static); s != "" {
|
||||
b.WriteString(s + "\n")
|
||||
}
|
||||
return b.String()
|
||||
}
|
||||
|
||||
// can lists what she can really do. Every entry here is a code path that
|
||||
// exists in the daemon today:
|
||||
// - reminders: IntentReminder → CoreAPI.CreateReminder, fired by the tick.
|
||||
// - notes and facts: IntentNote/IntentFact write, IntentQuery reads them back.
|
||||
// - calendar: IntentQuery answers "что у меня сегодня" from CalendarEvents.
|
||||
// - weather / telegram / shell acts: only when configured (see Facts).
|
||||
//
|
||||
// Nothing speculative goes in this list. A capability she offers and cannot
|
||||
// perform is worse than one she never mentions.
|
||||
func (f Facts) can() []string {
|
||||
c := []string{
|
||||
// Talking comes first, and the closing line says "действий" rather than
|
||||
// "ничего", because this same block sits in front of the chat and
|
||||
// general-knowledge prompts. A flat "you can do nothing else" would
|
||||
// tell her to refuse the exact thing those two prompts are for.
|
||||
"разговаривать и отвечать на вопросы",
|
||||
"ставить напоминания",
|
||||
"записывать заметки и факты и отвечать по ним",
|
||||
"смотреть календарь",
|
||||
}
|
||||
if f.Weather {
|
||||
c = append(c, "говорить погоду")
|
||||
}
|
||||
if f.Telegram {
|
||||
c = append(c, "писать в телеграм")
|
||||
}
|
||||
if f.Tools {
|
||||
c = append(c, "запускать разрешённые команды на сервере")
|
||||
}
|
||||
return c
|
||||
}
|
||||
|
||||
// who renders the optional name/city line, or "" when neither is configured.
|
||||
//
|
||||
// Written as labels ("Имя владельца: ..."), not as a sentence with pronouns:
|
||||
// the block's own "ты" is Maven, so "тебя зовут" would read as her name and
|
||||
// "его" would model the third-person form she must never use about him.
|
||||
func (f Facts) who() string {
|
||||
name := strings.TrimSpace(f.OwnerName)
|
||||
city := strings.TrimSpace(f.City)
|
||||
switch {
|
||||
case name != "" && city != "":
|
||||
return "Имя владельца: " + name + ". Город: " + city + "."
|
||||
case name != "":
|
||||
return "Имя владельца: " + name + "."
|
||||
case city != "":
|
||||
return "Город: " + city + "."
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
// Prepend puts the block in front of a system prompt. Nil-safe: a nil renderer
|
||||
// (tests, the stub paths) returns the prompt untouched.
|
||||
func Prepend(block func() string, prompt string) string {
|
||||
if block == nil {
|
||||
return prompt
|
||||
}
|
||||
s := strings.TrimSpace(block())
|
||||
if s == "" {
|
||||
return prompt
|
||||
}
|
||||
return s + "\n\n" + prompt
|
||||
}
|
||||
@@ -0,0 +1,72 @@
|
||||
package persona
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
var ref = time.Date(2026, 7, 31, 14, 5, 0, 0, time.UTC)
|
||||
|
||||
// The block must be correct with an empty config: the address form and the
|
||||
// gender rules are hard constraints, not preferences.
|
||||
func TestBlockWorksWithZeroConfig(t *testing.T) {
|
||||
b := Facts{}.Block(ref)
|
||||
for _, want := range []string{"женском роде", "ОБРАЩЕНИЕ", "\"ты\"", "31 июля 2026", "пятница", "14:05"} {
|
||||
if !strings.Contains(b, want) {
|
||||
t.Errorf("block missing %q:\n%s", want, b)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestBlockAddsOptionalFacts(t *testing.T) {
|
||||
b := Facts{OwnerName: "Ками", City: "Москва", Static: "Будь краткой."}.Block(ref)
|
||||
for _, want := range []string{"Ками", "Москва", "Будь краткой."} {
|
||||
if !strings.Contains(b, want) {
|
||||
t.Errorf("block missing %q:\n%s", want, b)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// The time changes between turns, so two renders must differ.
|
||||
func TestBlockRendersTimePerTurn(t *testing.T) {
|
||||
a := Facts{}.Block(ref)
|
||||
c := Facts{}.Block(ref.Add(time.Hour))
|
||||
if a == c {
|
||||
t.Errorf("block did not change with the clock:\n%s", a)
|
||||
}
|
||||
}
|
||||
|
||||
// She may only offer what this deployment actually has.
|
||||
func TestCapabilitiesAreConfigGated(t *testing.T) {
|
||||
bare := Facts{}.Block(ref)
|
||||
for _, want := range []string{"напоминания", "заметки", "календарь"} {
|
||||
if !strings.Contains(bare, want) {
|
||||
t.Errorf("block missing always-on capability %q:\n%s", want, bare)
|
||||
}
|
||||
}
|
||||
for _, unwanted := range []string{"погоду", "телеграм", "команды"} {
|
||||
if strings.Contains(bare, unwanted) {
|
||||
t.Errorf("block offers unconfigured %q:\n%s", unwanted, bare)
|
||||
}
|
||||
}
|
||||
|
||||
full := Facts{Weather: true, Telegram: true, Tools: true}.Block(ref)
|
||||
for _, want := range []string{"погоду", "телеграм", "команды"} {
|
||||
if !strings.Contains(full, want) {
|
||||
t.Errorf("block missing configured capability %q:\n%s", want, full)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestPrependNilIsSafe(t *testing.T) {
|
||||
if got := Prepend(nil, "PROMPT"); got != "PROMPT" {
|
||||
t.Errorf("Prepend(nil) = %q", got)
|
||||
}
|
||||
if got := Prepend(func() string { return " " }, "PROMPT"); got != "PROMPT" {
|
||||
t.Errorf("Prepend(blank) = %q", got)
|
||||
}
|
||||
if got := Prepend(func() string { return "CTX" }, "PROMPT"); got != "CTX\n\nPROMPT" {
|
||||
t.Errorf("Prepend = %q", got)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,54 @@
|
||||
package phraser
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// A reply that starts a JSON object and never finishes it is a failed
|
||||
// generation, not a reply. Before this, the parser returned ("", "") for these
|
||||
// and every caller then shipped the raw fragment as the thing Maven said. A
|
||||
// real run produced replies of literally "{" and "{\n \"".
|
||||
func TestParseResponseMoodRejectsUnfinishedJSON(t *testing.T) {
|
||||
for _, raw := range []string{
|
||||
`{`,
|
||||
"{\n \"",
|
||||
`{"response": "неполн`,
|
||||
`{"response": "текст", "mood":`,
|
||||
} {
|
||||
text, mood, err := parseResponseMood(raw)
|
||||
if !errors.Is(err, errBrokenJSON) {
|
||||
t.Errorf("parseResponseMood(%q) err = %v, want errBrokenJSON", raw, err)
|
||||
}
|
||||
if text != "" || mood != "" {
|
||||
t.Errorf("parseResponseMood(%q) leaked %q/%q — a fragment must never come back as a reply", raw, text, mood)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Bare prose is still fine. Small models sometimes answer without any JSON at
|
||||
// all, and that reply is usable — so the new error must not swallow it.
|
||||
func TestParseResponseMoodAllowsBareProse(t *testing.T) {
|
||||
for _, raw := range []string{
|
||||
"норм, а ты как?",
|
||||
"вот что я нашла: ключ у соседа",
|
||||
} {
|
||||
text, mood, err := parseResponseMood(raw)
|
||||
if err != nil {
|
||||
t.Errorf("parseResponseMood(%q) err = %v, want nil", raw, err)
|
||||
}
|
||||
// No JSON means no fields; the caller ships raw as-is.
|
||||
if text != "" || mood != "" {
|
||||
t.Errorf("parseResponseMood(%q) = %q/%q, want empty", raw, text, mood)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// The measured failure: the model wants more than 400 characters and the old
|
||||
// grammar cut it off mid-word. Guards the bound against being tightened back.
|
||||
func TestGrammarStringBoundHasRoomForARealAnswer(t *testing.T) {
|
||||
if !strings.Contains(responseGrammar, "{0,1000}") {
|
||||
t.Error("grammar string bound is not 1000; 400 truncated real replies mid-word (see the comment on responseGrammar)")
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,33 @@
|
||||
package phraser
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// Every phrasing prompt must carry the shared context block. This is the
|
||||
// regression guard for the bug that started this: the "ты" rule reached only
|
||||
// two of the five prompts because each prompt had its own copy of the rules.
|
||||
func TestEveryPromptCarriesTheContextBlock(t *testing.T) {
|
||||
block := func() string { return "CTXBLOCK" }
|
||||
p := &LLMPhraser{cfg: Config{ContextBlock: block}}
|
||||
|
||||
prompts := map[string]string{
|
||||
"nudge": p.systemPrompt(),
|
||||
"query": p.querySystemPrompt(),
|
||||
"chat": chatSystemPrompt(block),
|
||||
}
|
||||
for name, got := range prompts {
|
||||
if !strings.HasPrefix(got, "CTXBLOCK\n\n") {
|
||||
t.Errorf("%s prompt does not start with the context block:\n%s", name, got)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Without a block the prompts are unchanged — the stub and test paths pass nil.
|
||||
func TestPromptsWithoutBlockAreUnchanged(t *testing.T) {
|
||||
p := &LLMPhraser{}
|
||||
if p.systemPrompt() != nudgeSystem {
|
||||
t.Errorf("nudge prompt changed with no block set")
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,65 @@
|
||||
package eval
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// TestAddressReportsEveryBreak — the real reply from a nudge eval run broke in
|
||||
// two ways at once and the check named only the plural. Both must print: a
|
||||
// half-reported failure reads as a milder problem than it is.
|
||||
func TestAddressReportsEveryBreak(t *testing.T) {
|
||||
body := "Смотрите на его потребление воды."
|
||||
res := checkAddress(body)
|
||||
if res.Pass {
|
||||
t.Fatalf("checkAddress passed %q", body)
|
||||
}
|
||||
for _, want := range []string{"смотрите", "его"} {
|
||||
if !strings.Contains(res.Detail, want) {
|
||||
t.Errorf("detail %q does not name %q", res.Detail, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// One word repeated is one problem, so the detail must not say it twice.
|
||||
func TestAddressDeduplicates(t *testing.T) {
|
||||
res := checkAddress("Вам стоит поесть, вам это нужно.")
|
||||
if res.Pass {
|
||||
t.Fatal("expected failure")
|
||||
}
|
||||
if n := strings.Count(res.Detail, "formal"); n != 1 {
|
||||
t.Errorf("detail repeats the same break %d times: %q", n, res.Detail)
|
||||
}
|
||||
}
|
||||
|
||||
// The fragments a real run produced. All of them scored as non-empty replies
|
||||
// before checkNonEmpty looked for letters.
|
||||
func TestNonEmptyNeedsLetters(t *testing.T) {
|
||||
for _, body := range []string{
|
||||
"{",
|
||||
"{\n \"",
|
||||
"15-16",
|
||||
`{"`,
|
||||
" ",
|
||||
"...",
|
||||
} {
|
||||
if got := checkNonEmpty(body); got.Pass {
|
||||
t.Errorf("checkNonEmpty(%q) passed — that is not a reply", body)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// And it must not start failing real replies. Latin counts as well as Cyrillic:
|
||||
// answers about ssd or vpn are legitimately part English.
|
||||
func TestNonEmptyAcceptsRealReplies(t *testing.T) {
|
||||
for _, body := range []string{
|
||||
"норм, а ты как?",
|
||||
"вот что я нашла: ключ у соседа",
|
||||
"ssd быстрее hdd.",
|
||||
"9 минут.",
|
||||
} {
|
||||
if got := checkNonEmpty(body); !got.Pass {
|
||||
t.Errorf("checkNonEmpty(%q) failed: %s", body, got.Detail)
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,42 @@
|
||||
package eval
|
||||
|
||||
import "testing"
|
||||
|
||||
func TestAddressTimeWordDoesNotBlind(t *testing.T) {
|
||||
// A nudge that opens with a time word must still be caught. Without the
|
||||
// time words in the stoplist, "сегодня" was read as the third party.
|
||||
for _, s := range []string{
|
||||
"сегодня он не ел 11 дней",
|
||||
"вчера он не пил воду",
|
||||
"опять он забыл про таблетки",
|
||||
} {
|
||||
if r := checkAddress(s); r.Pass {
|
||||
t.Errorf("checkAddress(%q) passed, want a third-person failure", s)
|
||||
}
|
||||
}
|
||||
// Still must not fire when a third party really is named.
|
||||
for _, s := range []string{
|
||||
"сегодня сервис упал, он не отвечает",
|
||||
"ты не пил воду четыре часа",
|
||||
} {
|
||||
if r := checkAddress(s); !r.Pass {
|
||||
t.Errorf("checkAddress(%q) failed: %s", s, r.Detail)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestAddressVerbIsNotAnAntecedent — a nudge is mostly verbs, and a verb is
|
||||
// never who "он" refers to. This exact string passed the check before.
|
||||
func TestAddressVerbIsNotAnAntecedent(t *testing.T) {
|
||||
s := "попробуй встать и отдохнуть — у него есть перерыв"
|
||||
if r := checkAddress(s); r.Pass {
|
||||
t.Errorf("checkAddress(%q) passed, want a third-person failure", s)
|
||||
}
|
||||
// Still missed, and this is the documented hole: "выпей воды, он не пил" has
|
||||
// a real noun ("воды") before the pronoun, so the scan believes somebody
|
||||
// else was named. Telling that apart needs a parser, not a suffix rule.
|
||||
// A named third party still wins over the verbs around it.
|
||||
if r := checkAddress("сервис упал, он не отвечает"); !r.Pass {
|
||||
t.Errorf("checkAddress on a real third party failed: %s", r.Detail)
|
||||
}
|
||||
}
|
||||
@@ -17,10 +17,18 @@ const (
|
||||
CheckFeminine = "feminine" // her self-reference is feminine (hard constraint)
|
||||
CheckCringe = "cringe" // DESIGN.md § Non-goals, "not a relationship"
|
||||
CheckOnTopic = "ontopic" // says the thing the rule is about
|
||||
|
||||
// CheckHisGender — the other half of the persona rule: SHE is feminine, HE
|
||||
// is male. "ты давно не отдыхала" addresses the operator as a woman.
|
||||
CheckHisGender = "hisgender"
|
||||
|
||||
// CheckAddress — she talks TO him, informally, one to one. Not "вы", not
|
||||
// "он". See the comment block above checkAddress.
|
||||
CheckAddress = "address"
|
||||
)
|
||||
|
||||
// CheckNames — report order.
|
||||
var CheckNames = []string{CheckMood, CheckLang, CheckLength, CheckFeminine, CheckCringe, CheckOnTopic}
|
||||
var CheckNames = []string{CheckMood, CheckLang, CheckLength, CheckFeminine, CheckHisGender, CheckAddress, CheckCringe, CheckOnTopic}
|
||||
|
||||
// Result — one check on one message.
|
||||
type Result struct {
|
||||
@@ -52,6 +60,8 @@ func RunChecks(c Case, body, mood string) []Result {
|
||||
checkLang(body),
|
||||
checkLength(body),
|
||||
checkFeminine(body),
|
||||
checkHisGender(body),
|
||||
checkAddress(body),
|
||||
checkCringe(body),
|
||||
checkOnTopic(c, body),
|
||||
}
|
||||
@@ -176,6 +186,315 @@ func checkFeminine(body string) Result {
|
||||
return Result{CheckFeminine, true, ""}
|
||||
}
|
||||
|
||||
// --- he is male ----------------------------------------------------------
|
||||
//
|
||||
// The mirror of checkFeminine, and the failure it was written for: the model
|
||||
// wrote "ты давно не отдыхала", which addresses the operator as a woman. That
|
||||
// scored clean, because checkFeminine only ever looks at how SHE speaks about
|
||||
// herself.
|
||||
//
|
||||
// How it works: Russian past tense is gendered by suffix, -л (m) / -ла (f). So
|
||||
// this looks for feminine past-tense words in a sentence that also talks TO him
|
||||
// ("ты", "тебя", "тебе", "твой", …). A feminine verb that belongs to her ("я
|
||||
// заметила", "напомнила тебе") is skipped — that one is correct.
|
||||
//
|
||||
// Honest about the limits: this is a suffix rule, not a parser.
|
||||
// - False positives: a feminine noun can be the subject in the same sentence
|
||||
// ("зарядка была утром, ты её пропустил"). The guard below skips a verb whose
|
||||
// previous word looks like a feminine noun, which helps but will not always
|
||||
// be right.
|
||||
// - False negatives: gender also shows up outside the past tense (short
|
||||
// adjectives, "сама"), and none of that is checked here.
|
||||
//
|
||||
// That is acceptable for an eval check. It is a signal to read the message, not
|
||||
// a grammar verdict, and every hit prints the word it tripped on so a human can
|
||||
// disagree.
|
||||
|
||||
// hisMarkers — words that mean the sentence is addressed to him.
|
||||
var hisMarkers = map[string]bool{
|
||||
"ты": true, "тебя": true, "тебе": true, "тобой": true, "тобою": true,
|
||||
"твой": true, "твоя": true, "твоё": true, "твое": true, "твои": true, "твою": true,
|
||||
}
|
||||
|
||||
// notFeminineVerb — ordinary words ending in "-ла" that are not verbs. Small on
|
||||
// purpose: it only has to cover words a nudge might actually use.
|
||||
//
|
||||
// Words that are both a noun and a verb are deliberately NOT here. "села",
|
||||
// "мыла" and "стекла" are nouns on paper, but in a nudge they are almost always
|
||||
// verbs ("ты села", "ты мыла"), and listing them would make the check miss the
|
||||
// exact thing it is for. Missing a real hit is worse than one false alarm.
|
||||
var notFeminineVerb = map[string]bool{
|
||||
"школа": true, "скала": true, "игла": true, "метла": true, "смола": true,
|
||||
"дела": true, "тела": true, "масла": true, "весла": true,
|
||||
"зола": true, "пчела": true, "числа": true,
|
||||
}
|
||||
|
||||
// femininePast reports whether a word looks like a feminine past-tense verb:
|
||||
// "отдыхала", "поела", "выспалась".
|
||||
func femininePast(w string) bool {
|
||||
if len([]rune(w)) < 3 || notFeminineVerb[w] {
|
||||
return false
|
||||
}
|
||||
return strings.HasSuffix(w, "ла") || strings.HasSuffix(w, "лась")
|
||||
}
|
||||
|
||||
// looksFeminineNoun — a crude guard against "зарядка была": a word right before
|
||||
// the verb that ends in "а"/"я" and is not itself a verb is probably the subject.
|
||||
func looksFeminineNoun(w string) bool {
|
||||
if femininePast(w) || len([]rune(w)) < 3 {
|
||||
return false
|
||||
}
|
||||
return strings.HasSuffix(w, "а") || strings.HasSuffix(w, "я")
|
||||
}
|
||||
|
||||
// sentenceRE splits on sentence-ending punctuation, so a feminine verb in one
|
||||
// sentence is not blamed on a "ты" in the next.
|
||||
var sentenceRE = regexp.MustCompile(`[.!?;…]+`)
|
||||
|
||||
func checkHisGender(body string) Result {
|
||||
for _, sentence := range sentenceRE.Split(strings.ToLower(body), -1) {
|
||||
words := wordRE.FindAllString(sentence, -1)
|
||||
addressed := false
|
||||
for _, w := range words {
|
||||
if hisMarkers[w] {
|
||||
addressed = true
|
||||
}
|
||||
}
|
||||
if !addressed {
|
||||
continue
|
||||
}
|
||||
for i, w := range words {
|
||||
if !femininePast(w) || hersNotHis(words, i) {
|
||||
continue
|
||||
}
|
||||
if i > 0 && looksFeminineNoun(prevWord(words, i)) {
|
||||
continue
|
||||
}
|
||||
return Result{CheckHisGender, false,
|
||||
fmt.Sprintf("feminine %q addressed to him — he is male", w)}
|
||||
}
|
||||
}
|
||||
return Result{CheckHisGender, true, ""}
|
||||
}
|
||||
|
||||
// hersNotHis — the verb is Maven's own if "я" comes shortly before it, or if the
|
||||
// thing she did was done to him ("напомнила тебе", "проверила за тебя").
|
||||
func hersNotHis(words []string, i int) bool {
|
||||
for j := i - 1; j >= 0 && j >= i-3; j-- {
|
||||
if words[j] == "я" {
|
||||
return true
|
||||
}
|
||||
}
|
||||
if i+1 < len(words) {
|
||||
switch words[i+1] {
|
||||
case "тебе", "тебя", "за", "тобой":
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// prevWord — the word before i, skipping "не" and punctuation, so "не отдыхала"
|
||||
// still sees the subject.
|
||||
func prevWord(words []string, i int) string {
|
||||
for j := i - 1; j >= 0; j-- {
|
||||
w := words[j]
|
||||
if w == "не" || w == "ни" || !unicode.Is(unicode.Cyrillic, []rune(w)[0]) {
|
||||
continue
|
||||
}
|
||||
return w
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
// --- how she addresses him ------------------------------------------------
|
||||
//
|
||||
// Persona hard constraint: Maven speaks TO him, informally, one to one. The
|
||||
// phrasing eval produced two breaks of it, and both scored clean:
|
||||
//
|
||||
// - "Приходите… Жду вас" — the formal plural. Correct is ты/тебя/тебе and a
|
||||
// singular imperative ("приходи", "жду тебя").
|
||||
// - "Он не ел 11 дней" — she talks ABOUT him, in the third person, as if
|
||||
// reporting to somebody else. Correct is "ты не ел 11 дней".
|
||||
//
|
||||
// Like checkHisGender this is a keyword + suffix heuristic, NOT a parser. Every
|
||||
// hit prints the word it tripped on, so a false alarm is obvious at a glance and
|
||||
// can be dismissed.
|
||||
//
|
||||
// Part 1, formal address. Two signals:
|
||||
// - the "вы" pronoun family, matched as whole words, so there is nothing to
|
||||
// exclude — "вы" and "вас" are never anything else.
|
||||
// - a plural verb ending: -ите/-ете/-йте/-ьте ("приходите", "выпейте",
|
||||
// "не забудьте", "хотите"). Nouns in the prepositional case share those
|
||||
// endings ("в интернете", "в свете"), so a word right after a preposition is
|
||||
// skipped. That is the whole exclusion list, on purpose: a bigger one would
|
||||
// start swallowing real imperatives.
|
||||
//
|
||||
// Part 2, third person. "он" is perfectly fine when the message really is about
|
||||
// somebody or something else ("сервис упал, он не отвечает"). The way to tell
|
||||
// them apart: a legitimate third person has an ANTECEDENT — the thing it refers
|
||||
// to was named earlier in the message. So "он" is only flagged when nothing
|
||||
// before it in the message could be that thing.
|
||||
//
|
||||
// Where this gives up, plainly:
|
||||
// - it only looks BACKWARD. "Он не отвечает, сервис упал" names the subject
|
||||
// after the pronoun and is flagged wrongly.
|
||||
// - any noun earlier in the message counts as an antecedent, even when it is
|
||||
// not one ("после обеда он не ел", "выпей воды, он не пил" — both missed).
|
||||
// Verbs and time words no longer count, which covers the usual nudge, but a
|
||||
// plain noun before the pronoun still blinds it. The
|
||||
// common time words are stoplisted so the usual nudge opening does not
|
||||
// blind it, but a message with any other noun in front still slips through.
|
||||
// This is the check's real hole; widening it further would start flagging
|
||||
// legitimate third-party messages, so it stops here.
|
||||
// - a message that opens with "ты" and only later slips into "он" is missed,
|
||||
// because "ты" itself is skipped but the words around it are not.
|
||||
// - formal address outside these endings (short adjectives, "вашими" style
|
||||
// forms not listed) is missed.
|
||||
|
||||
// addressWordRE also takes Latin words, because "him"/"he" is the same break in
|
||||
// English.
|
||||
var addressWordRE = regexp.MustCompile(`[\p{Cyrillic}]+|[a-zA-Z]+|[,.;:!?…—-]`)
|
||||
|
||||
// formalPronouns — the "вы" family. Whole-word match, so no false hits.
|
||||
var formalPronouns = map[string]bool{
|
||||
"вы": true, "вас": true, "вам": true, "вами": true,
|
||||
"ваш": true, "ваша": true, "ваше": true, "ваши": true,
|
||||
"вашего": true, "вашей": true, "вашему": true, "вашим": true,
|
||||
"вашими": true, "вашу": true,
|
||||
}
|
||||
|
||||
// prepositions — used twice: to skip prepositional-case nouns that look like
|
||||
// plural verbs, and as words that cannot be what "он" refers to.
|
||||
var prepositions = map[string]bool{
|
||||
"в": true, "во": true, "на": true, "о": true, "об": true, "обо": true,
|
||||
"при": true, "по": true, "за": true, "из": true, "с": true, "со": true,
|
||||
"к": true, "ко": true, "до": true, "от": true, "у": true, "над": true,
|
||||
"под": true, "про": true, "без": true, "для": true, "через": true,
|
||||
}
|
||||
|
||||
// pluralVerb reports whether a word looks like a plural/formal verb form:
|
||||
// "приходите", "выпейте", "забудьте", "хотите".
|
||||
func pluralVerb(w string) bool {
|
||||
if len([]rune(w)) < 5 {
|
||||
return false
|
||||
}
|
||||
return strings.HasSuffix(w, "ите") || strings.HasSuffix(w, "ете") ||
|
||||
strings.HasSuffix(w, "йте") || strings.HasSuffix(w, "ьте")
|
||||
}
|
||||
|
||||
// thirdPersonHim — pronouns that would be talking about him instead of to him.
|
||||
var thirdPersonHim = map[string]bool{
|
||||
"он": true, "его": true, "ему": true, "него": true, "нему": true, "ним": true,
|
||||
"he": true, "him": true, "his": true,
|
||||
}
|
||||
|
||||
// notAnAntecedent — words that cannot be the thing "он" refers to: pronouns,
|
||||
// particles, conjunctions, adverbs of time. If only these come before "он", the
|
||||
// message never named a third party and "он" is him.
|
||||
var notAnAntecedent = map[string]bool{
|
||||
"не": true, "ни": true, "и": true, "а": true, "но": true, "да": true,
|
||||
"же": true, "бы": true, "ли": true, "вот": true, "уже": true,
|
||||
"ещё": true, "еще": true, "тоже": true, "там": true, "тут": true,
|
||||
"здесь": true, "это": true, "что": true, "как": true, "когда": true,
|
||||
"чтобы": true, "потому": true, "сейчас": true, "потом": true,
|
||||
// Time words. A nudge almost always opens with one ("сегодня он не ел"),
|
||||
// and without them the very next word is read as the person being talked
|
||||
// about, so the check misses the exact break it was written for.
|
||||
"сегодня": true, "вчера": true, "завтра": true, "послезавтра": true,
|
||||
"утром": true, "днём": true, "днем": true, "вечером": true, "ночью": true,
|
||||
"опять": true, "снова": true, "весь": true, "всю": true, "целый": true,
|
||||
"я": true, "мне": true, "меня": true, "мной": true, "мы": true, "нас": true,
|
||||
"ты": true, "тебя": true, "тебе": true, "тобой": true,
|
||||
"твой": true, "твоя": true, "твоё": true, "твое": true, "твои": true, "твою": true,
|
||||
}
|
||||
|
||||
// looksVerb — a verb is never the thing "он" refers to, so it must not count as
|
||||
// an antecedent. Past tense keeps "сервис упал, он не отвечает" working off
|
||||
// "сервис"; the infinitive and imperative endings are here because a nudge is
|
||||
// mostly made of them ("попробуй встать и отдохнуть — у него есть перерыв"
|
||||
// slipped through with "попробуй" taken for the person being talked about).
|
||||
func looksVerb(w string) bool {
|
||||
r := []rune(w)
|
||||
if len(r) < 3 {
|
||||
return false
|
||||
}
|
||||
for _, suf := range []string{
|
||||
"л", "ла", "ло", "ли", // past tense
|
||||
"ть", "ться", "ти", "чь", // infinitive
|
||||
"й", "йся", "йте", // imperative
|
||||
} {
|
||||
if strings.HasSuffix(w, suf) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
func checkAddress(body string) Result {
|
||||
words := addressWordRE.FindAllString(strings.ToLower(body), -1)
|
||||
|
||||
// Every break, not just the first. A bad reply usually breaks in more than
|
||||
// one way at once — "Смотрите на его потребление воды" is a plural imperative
|
||||
// AND third person about him — and reporting only the first hid the second,
|
||||
// which made the failure look milder than it was.
|
||||
var breaks []string
|
||||
seen := map[string]bool{}
|
||||
add := func(msg string) {
|
||||
if seen[msg] {
|
||||
return // the same word twice in one message is one problem, not two
|
||||
}
|
||||
seen[msg] = true
|
||||
breaks = append(breaks, msg)
|
||||
}
|
||||
|
||||
for i, w := range words {
|
||||
if formalPronouns[w] {
|
||||
add(fmt.Sprintf("formal %q — she says ты/тебя/тебе", w))
|
||||
}
|
||||
if pluralVerb(w) && !(i > 0 && prepositions[words[i-1]]) {
|
||||
add(fmt.Sprintf("plural imperative %q — she uses the singular", w))
|
||||
}
|
||||
}
|
||||
|
||||
for i, w := range words {
|
||||
if !thirdPersonHim[w] {
|
||||
continue
|
||||
}
|
||||
named := false
|
||||
for j := 0; j < i; j++ {
|
||||
p := words[j]
|
||||
if !unicode.Is(unicode.Cyrillic, []rune(p)[0]) && !isLatinWord(p) {
|
||||
continue // punctuation
|
||||
}
|
||||
// pluralVerb as well as looksVerb: looksVerb knows the imperative in
|
||||
// -й/-йте but not the -те plural ("смотрите"), so "Смотрите на его
|
||||
// потребление воды" counted "смотрите" as the person being talked
|
||||
// about and the "его" never printed. Third time a verb form has
|
||||
// blinded this check — if a fourth turns up, the antecedent test
|
||||
// wants a real morphology table, not another suffix.
|
||||
if notAnAntecedent[p] || prepositions[p] || thirdPersonHim[p] || looksVerb(p) || pluralVerb(p) {
|
||||
continue
|
||||
}
|
||||
named = true
|
||||
break
|
||||
}
|
||||
if !named {
|
||||
add(fmt.Sprintf("third person %q with nobody else named — she talks to him, not about him", w))
|
||||
}
|
||||
}
|
||||
|
||||
if len(breaks) > 0 {
|
||||
return Result{CheckAddress, false, strings.Join(breaks, " + ")}
|
||||
}
|
||||
return Result{CheckAddress, true, ""}
|
||||
}
|
||||
|
||||
func isLatinWord(w string) bool {
|
||||
r := []rune(w)[0]
|
||||
return (r >= 'a' && r <= 'z') || (r >= 'A' && r <= 'Z')
|
||||
}
|
||||
|
||||
// --- the cringe checks ---------------------------------------------------
|
||||
//
|
||||
// "Think Jarvis without the cringe part". DESIGN.md § Non-goals: "Not a
|
||||
@@ -276,12 +595,60 @@ func checkCringe(body string) Result {
|
||||
// checkOnTopic — the message must name the thing the rule is about. A nudge
|
||||
// that never mentions water leaves the operator with a chime and no action.
|
||||
func checkOnTopic(c Case, body string) Result {
|
||||
return checkOnTopicAny(c.WantAny, body)
|
||||
}
|
||||
|
||||
// checkOnTopicAny is the same test over a bare want-list, so the talk scorer can
|
||||
// reuse it without owning a nudge Case.
|
||||
func checkOnTopicAny(wantAny []string, body string) Result {
|
||||
low := strings.ToLower(body)
|
||||
for _, want := range c.WantAny {
|
||||
for _, want := range wantAny {
|
||||
if strings.Contains(low, strings.ToLower(want)) {
|
||||
return Result{CheckOnTopic, true, ""}
|
||||
}
|
||||
}
|
||||
return Result{CheckOnTopic, false,
|
||||
fmt.Sprintf("mentions none of %v", c.WantAny)}
|
||||
fmt.Sprintf("mentions none of %v", wantAny)}
|
||||
}
|
||||
|
||||
// --- shape checks for the free-form paths --------------------------------
|
||||
//
|
||||
// The nudge checks assume one short sentence. Chat and query replies are longer
|
||||
// by design, so the only shape worth testing there is that the model produced a
|
||||
// reply at all and did not trail off. Both are failure modes the fallbacks in
|
||||
// llmphraser.go hide: a truncated or empty generation still returns nil error.
|
||||
|
||||
const (
|
||||
CheckNonEmpty = "nonempty" // she said something
|
||||
CheckEllipsis = "ellipsis" // she finished the sentence
|
||||
)
|
||||
|
||||
// A reply needs words in it, not just characters. This check used to test for a
|
||||
// non-empty string, which scored 27/27 on a run where two replies were "{" and
|
||||
// "{\n \"" — punctuation passed as content. Braces, quotes, digits and spaces
|
||||
// are all empty in the only sense that matters.
|
||||
//
|
||||
// Digits alone fail too, and that is deliberate: the same run answered "сколько
|
||||
// варить яйцо вкрутую?" with "15-16". No unit, no words, and it is also the
|
||||
// wrong number. Whatever that is, it is not something she said.
|
||||
func checkNonEmpty(body string) Result {
|
||||
if strings.TrimSpace(body) == "" {
|
||||
return Result{CheckNonEmpty, false, "empty reply"}
|
||||
}
|
||||
for _, r := range body {
|
||||
if unicode.IsLetter(r) {
|
||||
return Result{CheckNonEmpty, true, ""}
|
||||
}
|
||||
}
|
||||
return Result{CheckNonEmpty, false, fmt.Sprintf("no letters in the reply %q — punctuation or digits only", strings.TrimSpace(body))}
|
||||
}
|
||||
|
||||
// checkEllipsis — a reply ending in "…" or "..." is a generation that ran out of
|
||||
// tokens, not a stylistic pause. Mid-sentence ellipses are left alone.
|
||||
func checkEllipsis(body string) Result {
|
||||
trimmed := strings.TrimRight(strings.TrimSpace(body), `"'»)`)
|
||||
if strings.HasSuffix(trimmed, "…") || strings.HasSuffix(trimmed, "...") {
|
||||
return Result{CheckEllipsis, false, "reply trails off in an ellipsis — likely truncated"}
|
||||
}
|
||||
return Result{CheckEllipsis, true, ""}
|
||||
}
|
||||
|
||||
@@ -281,7 +281,7 @@ func (r Report) String() string {
|
||||
fmt.Fprintf(&b, "%s: %d/%d cases pass every check (%.1f%%), %d errors\n",
|
||||
r.Name, r.Passed, r.Total, 100*r.Accuracy(), r.Errors)
|
||||
for _, name := range CheckNames {
|
||||
fmt.Fprintf(&b, " %-9s %d/%d\n", name, r.ByCheck[name], r.Total)
|
||||
fmt.Fprintf(&b, " %-10s %d/%d\n", name, r.ByCheck[name], r.Total)
|
||||
}
|
||||
fmt.Fprintf(&b, " latency: p50 %s p95 %s max %s\n", r.P50, r.P95, r.Max)
|
||||
fmt.Fprintf(&b, " by rule: %s\n", renderStats(r.ByRule))
|
||||
|
||||
@@ -72,10 +72,12 @@ func TestStubBaseline(t *testing.T) {
|
||||
// ceiling ("you've been at your desk for 4 hours without a break — step
|
||||
// away for a bit." is 76 chars but 16 words). Left failing rather than
|
||||
// raising the ceiling to hide it.
|
||||
CheckLength: 12,
|
||||
CheckFeminine: 15,
|
||||
CheckCringe: 15,
|
||||
CheckOnTopic: 12,
|
||||
CheckLength: 12,
|
||||
CheckFeminine: 15,
|
||||
CheckHisGender: 15,
|
||||
CheckAddress: 15,
|
||||
CheckCringe: 15,
|
||||
CheckOnTopic: 12,
|
||||
}
|
||||
for name, floor := range floors {
|
||||
if rep.ByCheck[name] < floor {
|
||||
@@ -104,6 +106,14 @@ func TestChecksCatchWhatTheyClaim(t *testing.T) {
|
||||
{"masculine predicative", "я должен сказать: попей воды.", CheckFeminine},
|
||||
// The other direction: HE is male, so second-person masculine is right.
|
||||
{"second person masculine ok", "ты не пил воду четыре часа.", ""},
|
||||
// The real observed failure: she addressed him as a woman.
|
||||
{"feminine second person", "ты давно не отдыхала — попей воды.", CheckHisGender},
|
||||
{"feminine second person no dash", "ты пила воду четыре часа назад.", CheckHisGender},
|
||||
// Her own feminine verb next to "ты" is correct and must not be flagged.
|
||||
{"her feminine verb near ты", "я заметила, что ты не пил воду.", ""},
|
||||
{"her feminine verb about him", "напомнила тебе про воду.", ""},
|
||||
// A feminine noun subject in the same sentence is not him.
|
||||
{"feminine noun subject ok", "зарядка была утром, ты её пропустил, попей воды.", ""},
|
||||
{"feminine self ok", "я заметила: воды не было четыре часа.", ""},
|
||||
{"pet name", "милый, попей воды.", CheckCringe},
|
||||
{"emoji", "попей воды 💧", CheckCringe},
|
||||
@@ -114,6 +124,10 @@ func TestChecksCatchWhatTheyClaim(t *testing.T) {
|
||||
{"asks how he feels", "как ты себя чувствуешь? попей воды.", CheckCringe},
|
||||
{"praise", "молодец! теперь попей воды.", CheckCringe},
|
||||
{"off topic", "пора бы уже что-то сделать.", CheckOnTopic},
|
||||
// The two recorded persona breaks from the phrasing eval run. Pinned as
|
||||
// unit tests because an eval run is sampled and may not reproduce them.
|
||||
{"formal plural", "Приходите… Жду вас", CheckAddress},
|
||||
{"third person about him", "Он не ел 11 дней", CheckAddress},
|
||||
}
|
||||
|
||||
for _, tc := range cases {
|
||||
@@ -135,6 +149,36 @@ func TestChecksCatchWhatTheyClaim(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestAddressCheck — the address check on its own, so the messages that must NOT
|
||||
// trip it can be written without also having to satisfy the on-topic check.
|
||||
func TestAddressCheck(t *testing.T) {
|
||||
bad := []string{
|
||||
"Приходите… Жду вас", // the recorded formal-plural break
|
||||
"Он не ел 11 дней", // the recorded third-person break
|
||||
"Выпейте воды, пожалуйста.", // plural imperative on its own
|
||||
"Ваш обед был давно.", // formal possessive
|
||||
}
|
||||
for _, body := range bad {
|
||||
if r := checkAddress(body); r.Pass {
|
||||
t.Errorf("persona break not caught: %q", body)
|
||||
} else {
|
||||
t.Logf("%q -> %s", body, r.Detail)
|
||||
}
|
||||
}
|
||||
|
||||
good := []string{
|
||||
"ты не пил воду четыре часа — попей.", // correct informal address
|
||||
"сервис netdata упал, он не отвечает.", // legitimately about a third party
|
||||
"я заметила, что зарядка была утром.", // no address at all
|
||||
"в интернете опять тихо, всё работает.", // "интернете" is a noun, not an imperative
|
||||
}
|
||||
for _, body := range good {
|
||||
if r := checkAddress(body); !r.Pass {
|
||||
t.Errorf("clean message flagged: %q -> %s", body, r.Detail)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestMoodCheckUsesTheEnum(t *testing.T) {
|
||||
if r := checkMood("cheerful"); r.Pass {
|
||||
t.Error("mood outside the enum passed")
|
||||
|
||||
@@ -7,6 +7,8 @@ import (
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/llm"
|
||||
"github.com/kami/maven/internal/persona"
|
||||
"github.com/kami/maven/internal/phraser"
|
||||
)
|
||||
|
||||
@@ -32,6 +34,7 @@ func TestLLMPhrasingBaseline(t *testing.T) {
|
||||
// case as a phrasing error and read as "the model cannot phrase".
|
||||
noProxyLoopback(t)
|
||||
|
||||
ctx := context.Background()
|
||||
f, err := Load()
|
||||
if err != nil {
|
||||
t.Fatalf("Load: %v", err)
|
||||
@@ -41,10 +44,25 @@ func TestLLMPhrasingBaseline(t *testing.T) {
|
||||
// Generous: an unconstrained 0.8B can spend a minute thinking before it
|
||||
// writes a word, and a timeout would be scored as a model failure.
|
||||
cfg.Timeout = 5 * time.Minute
|
||||
// The same shared context block the daemon prepends (internal/persona),
|
||||
// with an empty config — that is the deployment we actually ship.
|
||||
cfg.ContextBlock = func() string { return persona.Facts{}.Block(time.Now()) }
|
||||
p := phraser.NewLLMPhraserAt(base, cfg)
|
||||
defer p.Close()
|
||||
|
||||
rep, err := Score(context.Background(), "llm (0.8B, built-in persona)", p, f)
|
||||
// Label the run with whatever gguf the server actually has loaded. It used
|
||||
// to say "0.8B" no matter what, so two runs of two different models came
|
||||
// out named the same and were easy to mix up when comparing.
|
||||
model, err := llm.ModelID(ctx, base)
|
||||
if err != nil {
|
||||
// An unlabelled score is still a score, but say so loudly — a made-up
|
||||
// name in a bake-off table is worse than no name.
|
||||
t.Logf("could not read model id from %s: %v — report will say %q", base, err, llm.UnknownModel)
|
||||
model = llm.UnknownModel
|
||||
}
|
||||
t.Logf("scoring model %s at %s", model, base)
|
||||
|
||||
rep, err := Score(ctx, "llm ("+model+", built-in persona)", p, f)
|
||||
if err != nil {
|
||||
t.Fatalf("Score: %v", err)
|
||||
}
|
||||
|
||||
@@ -0,0 +1,265 @@
|
||||
package eval
|
||||
|
||||
// This file scores the CONVERSATIONAL paths, the ones the nudge fixture never
|
||||
// touches: chat, query-with-notes, and general knowledge. All three now carry
|
||||
// the shared persona block (internal/persona), and all three produce long
|
||||
// free-form Russian — which is exactly where a persona break (formality, third
|
||||
// person, masculine self-reference) is most likely and where, until this file,
|
||||
// nothing could see one.
|
||||
//
|
||||
// Why a second fixture instead of more nudge cases: the checks differ. A nudge
|
||||
// must be one short sentence with no question in it; a chat reply is allowed
|
||||
// 1-3 sentences and a follow-up question is a FEATURE there. Mixing them would
|
||||
// need per-case check masks, and the nudge scorer stays untouched this way.
|
||||
//
|
||||
// Why per-path reporting: a chat regression and a knowledge regression have
|
||||
// different causes (chat prompt vs router.KnowledgePrompt), and one blended
|
||||
// percentage cannot tell them apart.
|
||||
|
||||
import (
|
||||
"context"
|
||||
_ "embed"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"sort"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/dialogue"
|
||||
)
|
||||
|
||||
//go:embed talk_v1.json
|
||||
var talkFixtureJSON []byte
|
||||
|
||||
// The three phrasing paths under test. Values match the fixture's "path" field.
|
||||
const (
|
||||
PathChat = "chat" // PhraseChat
|
||||
PathQuery = "query" // PhraseQuery with notes
|
||||
PathKnowledge = "knowledge" // PhraseQuery with no notes
|
||||
)
|
||||
|
||||
// TalkPaths — report order.
|
||||
var TalkPaths = []string{PathChat, PathQuery, PathKnowledge}
|
||||
|
||||
// TalkCheckNames — the checks that apply to a free-form reply, in report order.
|
||||
// Deliberately a subset of CheckNames: length, mood and "no questions" are nudge
|
||||
// properties and would fail a correct chat reply. These paths return no mood at
|
||||
// all, so there is nothing to check there.
|
||||
var TalkCheckNames = []string{
|
||||
CheckNonEmpty, CheckEllipsis, CheckLang, CheckFeminine, CheckAddress, CheckOnTopic,
|
||||
}
|
||||
|
||||
// TalkCase — one turn as the daemon would present it.
|
||||
//
|
||||
// History is flat text because that is all PhraseChat uses (it concatenates
|
||||
// turn texts into one user message); intents and slots would be dead fields.
|
||||
// Notes are what the store would have matched for a query.
|
||||
//
|
||||
// WantAny is the on-topic contract: at least one lowercased fragment must appear
|
||||
// in the reply. Fragments are stems ("пароль" → "парол") so declension does not
|
||||
// defeat them.
|
||||
type TalkCase struct {
|
||||
ID string `json:"id"`
|
||||
Path string `json:"path"`
|
||||
Utterance string `json:"utterance"`
|
||||
History []string `json:"history,omitempty"`
|
||||
Notes []string `json:"notes,omitempty"`
|
||||
WantAny []string `json:"want_any"`
|
||||
Tags []string `json:"tags,omitempty"`
|
||||
Note string `json:"note,omitempty"`
|
||||
}
|
||||
|
||||
// TalkFixture — the versioned envelope, same gating as Fixture.
|
||||
type TalkFixture struct {
|
||||
SchemaVersion int `json:"schema_version"`
|
||||
Name string `json:"name"`
|
||||
Notes []string `json:"notes"`
|
||||
Cases []TalkCase `json:"cases"`
|
||||
}
|
||||
|
||||
// LoadTalk returns the embedded conversational fixture.
|
||||
func LoadTalk() (TalkFixture, error) {
|
||||
var f TalkFixture
|
||||
if err := json.Unmarshal(talkFixtureJSON, &f); err != nil {
|
||||
return TalkFixture{}, fmt.Errorf("parse talk fixture: %w", err)
|
||||
}
|
||||
if f.SchemaVersion != SchemaVersion {
|
||||
return TalkFixture{}, fmt.Errorf("talk fixture schema_version %d, want %d", f.SchemaVersion, SchemaVersion)
|
||||
}
|
||||
if len(f.Cases) == 0 {
|
||||
return TalkFixture{}, fmt.Errorf("talk fixture has no cases")
|
||||
}
|
||||
return f, nil
|
||||
}
|
||||
|
||||
// Talker — the two methods a conversational path must have to be scorable.
|
||||
// *phraser.LLMPhraser satisfies it; same trick as Nudger.
|
||||
type Talker interface {
|
||||
PhraseChat(ctx context.Context, utterance string, history []dialogue.Turn) (string, error)
|
||||
PhraseQuery(ctx context.Context, utterance string, notes []string) (string, error)
|
||||
}
|
||||
|
||||
// TalkOutcome — one scored case.
|
||||
type TalkOutcome struct {
|
||||
Case TalkCase
|
||||
Reply string
|
||||
Err error
|
||||
Latency time.Duration
|
||||
Pass bool
|
||||
Failed []string
|
||||
Reasons []string
|
||||
}
|
||||
|
||||
// TalkReport — the aggregate. ByPath is the point of this scorer.
|
||||
type TalkReport struct {
|
||||
Name string
|
||||
Total int
|
||||
Passed int
|
||||
Errors int
|
||||
ByCheck map[string]int
|
||||
ByPath map[string]TagStat
|
||||
Outcomes []TalkOutcome
|
||||
P50 time.Duration
|
||||
P95 time.Duration
|
||||
Max time.Duration
|
||||
}
|
||||
|
||||
// Accuracy — fraction of cases that passed every check.
|
||||
func (r TalkReport) Accuracy() float64 {
|
||||
if r.Total == 0 {
|
||||
return 0
|
||||
}
|
||||
return float64(r.Passed) / float64(r.Total)
|
||||
}
|
||||
|
||||
// ScoreTalk runs every case through t and aggregates. A phrasing error scores as
|
||||
// a miss and is counted separately: "the model was down" and "the model wrote
|
||||
// something bad" must not be the same number.
|
||||
func ScoreTalk(ctx context.Context, name string, t Talker, f TalkFixture) (TalkReport, error) {
|
||||
rep := TalkReport{
|
||||
Name: name,
|
||||
Total: len(f.Cases),
|
||||
ByCheck: map[string]int{},
|
||||
ByPath: map[string]TagStat{},
|
||||
}
|
||||
for _, n := range TalkCheckNames {
|
||||
rep.ByCheck[n] = 0
|
||||
}
|
||||
lat := make([]time.Duration, 0, len(f.Cases))
|
||||
|
||||
for _, c := range f.Cases {
|
||||
start := time.Now()
|
||||
reply, err := c.run(ctx, t)
|
||||
o := TalkOutcome{Case: c, Reply: reply, Err: err, Latency: time.Since(start)}
|
||||
lat = append(lat, o.Latency)
|
||||
|
||||
if err != nil {
|
||||
rep.Errors++
|
||||
o.Failed = append(o.Failed, "call")
|
||||
o.Reasons = append(o.Reasons, fmt.Sprintf("phrase error: %v", err))
|
||||
} else {
|
||||
for _, res := range RunTalkChecks(c, reply) {
|
||||
if res.Pass {
|
||||
rep.ByCheck[res.Name]++
|
||||
continue
|
||||
}
|
||||
o.Failed = append(o.Failed, res.Name)
|
||||
o.Reasons = append(o.Reasons, res.Name+": "+res.Detail)
|
||||
}
|
||||
}
|
||||
|
||||
o.Pass = len(o.Failed) == 0
|
||||
if o.Pass {
|
||||
rep.Passed++
|
||||
}
|
||||
bump(rep.ByPath, c.Path, o.Pass)
|
||||
rep.Outcomes = append(rep.Outcomes, o)
|
||||
}
|
||||
|
||||
sort.Slice(lat, func(i, j int) bool { return lat[i] < lat[j] })
|
||||
rep.P50, rep.P95 = percentile(lat, 0.50), percentile(lat, 0.95)
|
||||
if len(lat) > 0 {
|
||||
rep.Max = lat[len(lat)-1]
|
||||
}
|
||||
return rep, nil
|
||||
}
|
||||
|
||||
// run dispatches the case to its path. knowledge and query are the same method;
|
||||
// the empty notes slice is what selects the no-notes branch inside PhraseQuery.
|
||||
func (c TalkCase) run(ctx context.Context, t Talker) (string, error) {
|
||||
switch c.Path {
|
||||
case PathChat:
|
||||
return t.PhraseChat(ctx, c.Utterance, c.turns())
|
||||
case PathQuery:
|
||||
return t.PhraseQuery(ctx, c.Utterance, c.Notes)
|
||||
case PathKnowledge:
|
||||
return t.PhraseQuery(ctx, c.Utterance, nil)
|
||||
}
|
||||
return "", fmt.Errorf("unknown path %q", c.Path)
|
||||
}
|
||||
|
||||
func (c TalkCase) turns() []dialogue.Turn {
|
||||
turns := make([]dialogue.Turn, 0, len(c.History))
|
||||
for _, h := range c.History {
|
||||
turns = append(turns, dialogue.Turn{Text: h})
|
||||
}
|
||||
return turns
|
||||
}
|
||||
|
||||
// RunTalkChecks scores one reply. Order matches TalkCheckNames.
|
||||
func RunTalkChecks(c TalkCase, reply string) []Result {
|
||||
return []Result{
|
||||
checkNonEmpty(reply),
|
||||
checkEllipsis(reply),
|
||||
checkLang(reply),
|
||||
checkFeminine(reply),
|
||||
checkAddress(reply),
|
||||
checkOnTopicAny(c.WantAny, reply),
|
||||
}
|
||||
}
|
||||
|
||||
// String renders the comparison table — composite, then per-check so a
|
||||
// regression names the property, then per-path so it names the prompt.
|
||||
func (r TalkReport) String() string {
|
||||
var b strings.Builder
|
||||
fmt.Fprintf(&b, "%s: %d/%d cases pass every check (%.1f%%), %d errors\n",
|
||||
r.Name, r.Passed, r.Total, 100*r.Accuracy(), r.Errors)
|
||||
for _, name := range TalkCheckNames {
|
||||
fmt.Fprintf(&b, " %-10s %d/%d\n", name, r.ByCheck[name], r.Total)
|
||||
}
|
||||
fmt.Fprintf(&b, " latency: p50 %s p95 %s max %s\n", r.P50, r.P95, r.Max)
|
||||
fmt.Fprintf(&b, " by path: %s\n", renderStats(r.ByPath))
|
||||
return b.String()
|
||||
}
|
||||
|
||||
// Failures — per-case detail, sorted by ID so two runs diff cleanly.
|
||||
func (r TalkReport) Failures() string {
|
||||
var b strings.Builder
|
||||
for _, o := range r.sorted() {
|
||||
if o.Pass {
|
||||
continue
|
||||
}
|
||||
fmt.Fprintf(&b, " %s %q\n %s\n", o.Case.ID, o.Reply, strings.Join(o.Reasons, "; "))
|
||||
}
|
||||
return b.String()
|
||||
}
|
||||
|
||||
// Replies — every generated reply verbatim. This is what a human reads to judge
|
||||
// tone; the score only says which checks fired.
|
||||
func (r TalkReport) Replies() string {
|
||||
var b strings.Builder
|
||||
for _, o := range r.sorted() {
|
||||
mark := "ok "
|
||||
if !o.Pass {
|
||||
mark = "FAIL"
|
||||
}
|
||||
fmt.Fprintf(&b, " %s %-9s %-22s %q\n", mark, o.Case.Path, o.Case.ID, o.Reply)
|
||||
}
|
||||
return b.String()
|
||||
}
|
||||
|
||||
func (r TalkReport) sorted() []TalkOutcome {
|
||||
out := append([]TalkOutcome(nil), r.Outcomes...)
|
||||
sort.Slice(out, func(i, j int) bool { return out[i].Case.ID < out[j].Case.ID })
|
||||
return out
|
||||
}
|
||||
@@ -0,0 +1,163 @@
|
||||
package eval
|
||||
|
||||
import (
|
||||
"context"
|
||||
"os"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/dialogue"
|
||||
"github.com/kami/maven/internal/llm"
|
||||
"github.com/kami/maven/internal/persona"
|
||||
"github.com/kami/maven/internal/phraser"
|
||||
)
|
||||
|
||||
// perPathMinimum — the resolution floor. A per-path score built on a handful of
|
||||
// cases moves by 12% when a single reply changes, which cannot distinguish a
|
||||
// prompt regression from noise.
|
||||
const perPathMinimum = 8
|
||||
|
||||
// TestTalkFixture — the fixture itself has to be sound before any score off it
|
||||
// means anything.
|
||||
func TestTalkFixture(t *testing.T) {
|
||||
f, err := LoadTalk()
|
||||
if err != nil {
|
||||
t.Fatalf("LoadTalk: %v", err)
|
||||
}
|
||||
|
||||
seen := map[string]bool{}
|
||||
byPath := map[string]int{}
|
||||
for _, c := range f.Cases {
|
||||
if seen[c.ID] {
|
||||
t.Errorf("duplicate case id %q", c.ID)
|
||||
}
|
||||
seen[c.ID] = true
|
||||
|
||||
switch c.Path {
|
||||
case PathChat, PathQuery, PathKnowledge:
|
||||
default:
|
||||
t.Errorf("%s: unknown path %q", c.ID, c.Path)
|
||||
}
|
||||
byPath[c.Path]++
|
||||
|
||||
if strings.TrimSpace(c.Utterance) == "" {
|
||||
t.Errorf("%s: empty utterance", c.ID)
|
||||
}
|
||||
if len(c.WantAny) == 0 {
|
||||
t.Errorf("%s: no want_any — the reply cannot be checked for topic", c.ID)
|
||||
}
|
||||
// A query case with no notes would silently score the knowledge path.
|
||||
if c.Path == PathQuery && len(c.Notes) == 0 {
|
||||
t.Errorf("%s: query case has no notes", c.ID)
|
||||
}
|
||||
if c.Path == PathKnowledge && len(c.Notes) > 0 {
|
||||
t.Errorf("%s: knowledge case must have no notes", c.ID)
|
||||
}
|
||||
}
|
||||
|
||||
for _, p := range TalkPaths {
|
||||
if byPath[p] < perPathMinimum {
|
||||
t.Errorf("path %s has %d cases, want at least %d", p, byPath[p], perPathMinimum)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// fakeTalker — a scripted Talker, so the scorer is testable without a model.
|
||||
type fakeTalker struct{ reply string }
|
||||
|
||||
func (f fakeTalker) PhraseChat(context.Context, string, []dialogue.Turn) (string, error) {
|
||||
return f.reply, nil
|
||||
}
|
||||
func (f fakeTalker) PhraseQuery(context.Context, string, []string) (string, error) {
|
||||
return f.reply, nil
|
||||
}
|
||||
|
||||
// TestScoreTalkCounts — a reply that fails on purpose must be counted on every
|
||||
// path, so a real run cannot report a hidden zero.
|
||||
func TestScoreTalkCounts(t *testing.T) {
|
||||
f, err := LoadTalk()
|
||||
if err != nil {
|
||||
t.Fatalf("LoadTalk: %v", err)
|
||||
}
|
||||
// Formal address, off-topic, trailing ellipsis: three checks fail at once.
|
||||
rep, err := ScoreTalk(context.Background(), "fake", fakeTalker{"Приходите, я вас жду…"}, f)
|
||||
if err != nil {
|
||||
t.Fatalf("ScoreTalk: %v", err)
|
||||
}
|
||||
if rep.Total != len(f.Cases) || rep.Passed != 0 {
|
||||
t.Errorf("got %d/%d passing, want 0/%d", rep.Passed, rep.Total, len(f.Cases))
|
||||
}
|
||||
if rep.ByCheck[CheckAddress] != 0 {
|
||||
t.Errorf("formal reply passed the address check %d times", rep.ByCheck[CheckAddress])
|
||||
}
|
||||
if rep.ByCheck[CheckEllipsis] != 0 {
|
||||
t.Errorf("truncated reply passed the ellipsis check %d times", rep.ByCheck[CheckEllipsis])
|
||||
}
|
||||
for _, p := range TalkPaths {
|
||||
if rep.ByPath[p].Total == 0 {
|
||||
t.Errorf("path %s missing from the report", p)
|
||||
}
|
||||
}
|
||||
if !strings.Contains(rep.String(), "by path") {
|
||||
t.Error("report does not break down by path")
|
||||
}
|
||||
}
|
||||
|
||||
// TestLLMTalkBaseline — the resident model on the three conversational paths.
|
||||
// Opt-in exactly like TestLLMPhrasingBaseline: CI has no model and a run costs
|
||||
// minutes on the CPU target.
|
||||
//
|
||||
// MAVEN_LLM_URL=http://127.0.0.1:18099 \
|
||||
// go test -run TestLLMTalkBaseline ./internal/phraser/eval/
|
||||
//
|
||||
// Reports, does not assert a quality bar — the numbers are the input to tuning
|
||||
// the persona prompt. The one thing worth failing on is a harness fault.
|
||||
func TestLLMTalkBaseline(t *testing.T) {
|
||||
base := os.Getenv("MAVEN_LLM_URL")
|
||||
if base == "" {
|
||||
t.Skip("MAVEN_LLM_URL unset — point it at a running llama-server (see doc comment)")
|
||||
}
|
||||
noProxyLoopback(t)
|
||||
|
||||
ctx := context.Background()
|
||||
f, err := LoadTalk()
|
||||
if err != nil {
|
||||
t.Fatalf("LoadTalk: %v", err)
|
||||
}
|
||||
|
||||
cfg := phraser.DefaultConfig("")
|
||||
cfg.Timeout = 5 * time.Minute
|
||||
cfg.ContextBlock = func() string { return persona.Facts{}.Block(time.Now()) }
|
||||
p := phraser.NewLLMPhraserAt(base, cfg)
|
||||
defer p.Close()
|
||||
|
||||
// Unreachable server is fatal here, not a logged warning, and that differs
|
||||
// from the nudge test on purpose. PhraseNudge returns its errors, so a dead
|
||||
// server there shows up honestly in the Errors column. PhraseChat and
|
||||
// PhraseQuery do NOT: they swallow every failure and return a canned string
|
||||
// ("поговорили.", "не знаю.", "вот что я нашла: …"). So on these three paths
|
||||
// a dead server produces a full report with 0 errors and a terrible score —
|
||||
// a number that looks like bad phrasing and is really no phrasing at all.
|
||||
// Refusing to score without a confirmed model is the only guard available
|
||||
// until the phraser reports its failures (Vikunja #397).
|
||||
model, err := llm.ModelID(ctx, base)
|
||||
if err != nil {
|
||||
t.Fatalf("no model at %s: %v — refusing to score, these paths hide their errors "+
|
||||
"and would report a plausible-looking result off a dead server", base, err)
|
||||
}
|
||||
t.Logf("scoring model %s at %s", model, base)
|
||||
|
||||
rep, err := ScoreTalk(ctx, "llm ("+model+", built-in persona)", p, f)
|
||||
if err != nil {
|
||||
t.Fatalf("ScoreTalk: %v", err)
|
||||
}
|
||||
t.Log("\n" + rep.String() + "\nreplies:\n" + rep.Replies() + "\nfailures:\n" + rep.Failures())
|
||||
|
||||
// And again afterwards: the run takes minutes, and a server that died or got
|
||||
// OOM-killed halfway through would leave the first cases scored and the rest
|
||||
// silently canned. Checking only at the start would not catch that.
|
||||
if _, err := llm.ModelID(ctx, base); err != nil {
|
||||
t.Fatalf("model at %s went away during the run: %v — the score above is not trustworthy", base, err)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,227 @@
|
||||
{
|
||||
"schema_version": 1,
|
||||
"name": "ru-talk-v1",
|
||||
"notes": [
|
||||
"Scores the three conversational phrasing paths: chat (PhraseChat), query (PhraseQuery with notes) and knowledge (PhraseQuery with no notes). The nudge fixture does not cover any of them.",
|
||||
"Nine cases per path, not five. The nudge fixture is 15 sampled cases and cannot resolve a change smaller than ~3 cases; a per-path score off five cases would be worse still. More cases per path is the point of this fixture.",
|
||||
"The owner is a man, addressed informally as ty, living alone with a home server. Every utterance is written the way he actually talks to her.",
|
||||
"chat-formality-bait and chat-about-me exist to provoke the two persona breaks the nudge eval caught: the formal vy/vas plural, and talking about him in the third person.",
|
||||
"want_any fragments are stems so Russian declension does not defeat the on-topic check. They are lowercased before comparison.",
|
||||
"want_any is a plain substring test, so a fragment that is too short passes by accident: \"ты\" matches inside \"работы\", \"нет\" inside \"интернет\". Keep every fragment to three or more letters of a real stem.",
|
||||
"Notes are written as the store would have them: short, first person, no punctuation discipline."
|
||||
],
|
||||
"cases": [
|
||||
{
|
||||
"id": "chat-how-are-you",
|
||||
"path": "chat",
|
||||
"utterance": "привет, как дела?",
|
||||
"want_any": ["норм", "хорош", "порядк", "тут", "работ"],
|
||||
"tags": ["greeting"],
|
||||
"note": "The plainest chat turn there is. If the persona breaks anywhere it breaks here first."
|
||||
},
|
||||
{
|
||||
"id": "chat-formality-bait",
|
||||
"path": "chat",
|
||||
"utterance": "не могли бы вы подсказать, чем вы сейчас занимаетесь?",
|
||||
"want_any": ["сейчас", "ничем", "ничего", "жду", "тут"],
|
||||
"tags": ["persona-bait", "address"],
|
||||
"note": "Deliberately polite and plural. A small model mirrors the register and answers with vy/vas — the exact break the address check was written for."
|
||||
},
|
||||
{
|
||||
"id": "chat-about-me",
|
||||
"path": "chat",
|
||||
"utterance": "расскажи обо мне",
|
||||
"want_any": ["теб"],
|
||||
"tags": ["persona-bait", "third-person"],
|
||||
"note": "Baits the third person: she should say 'ты живёшь один', not 'он живёт один', as if reporting to somebody else."
|
||||
},
|
||||
{
|
||||
"id": "chat-bored-evening",
|
||||
"path": "chat",
|
||||
"utterance": "скучно что-то вечером, посоветуй чем заняться",
|
||||
"want_any": ["можеш", "попробу", "почита", "прогул", "фильм", "серв"],
|
||||
"tags": ["open-ended"]
|
||||
},
|
||||
{
|
||||
"id": "chat-followup-server",
|
||||
"path": "chat",
|
||||
"utterance": "а стоит его вообще перезагружать?",
|
||||
"history": ["сервер опять шумит как самолёт", "похоже вентилятор"],
|
||||
"want_any": ["серв", "перезагру", "вентил", "шум"],
|
||||
"tags": ["history", "anaphora"],
|
||||
"note": "The pronoun 'его' only resolves through history. Also the one case where 'он' about the server is legitimate."
|
||||
},
|
||||
{
|
||||
"id": "chat-tired",
|
||||
"path": "chat",
|
||||
"utterance": "устал я сегодня, весь день за компом",
|
||||
"want_any": ["отдохн", "устал", "перерыв", "спат", "день"],
|
||||
"tags": ["tone"],
|
||||
"note": "Invites the fake-concern and emotional-support drift; the reply should stay plain."
|
||||
},
|
||||
{
|
||||
"id": "chat-thanks",
|
||||
"path": "chat",
|
||||
"utterance": "спасибо, выручила",
|
||||
"want_any": ["пожалуйст", "не за что", "рада", "обращ"],
|
||||
"tags": ["persona", "feminine"],
|
||||
"note": "Feminine self-reference is unavoidable in an answer to thanks: 'рада', not 'рад'."
|
||||
},
|
||||
{
|
||||
"id": "chat-what-can-you-do",
|
||||
"path": "chat",
|
||||
"utterance": "что ты вообще умеешь?",
|
||||
"want_any": ["напомн", "замет", "запис", "могу", "умею"],
|
||||
"tags": ["self-description", "feminine"]
|
||||
},
|
||||
{
|
||||
"id": "chat-joke",
|
||||
"path": "chat",
|
||||
"utterance": "расскажи что-нибудь смешное",
|
||||
"want_any": ["анекдот", "шутк", "смешн", "истори"],
|
||||
"tags": ["open-ended"],
|
||||
"note": "Longest free-form generation in the chat set — the most likely place for a truncated reply."
|
||||
},
|
||||
{
|
||||
"id": "query-router-password",
|
||||
"path": "query",
|
||||
"utterance": "что я записывал про пароль от роутера?",
|
||||
"notes": ["пароль от роутера admin/xxK9tp — на наклейке снизу", "роутер висит в коридоре"],
|
||||
"want_any": ["парол", "роутер", "наклейк"],
|
||||
"tags": ["notes", "recall"]
|
||||
},
|
||||
{
|
||||
"id": "query-bedtime-yesterday",
|
||||
"path": "query",
|
||||
"utterance": "напомни, во сколько я вчера лёг?",
|
||||
"notes": ["лёг спать в 02:40", "сегодня встал в 9"],
|
||||
"want_any": ["02:40", "2:40", "полтрет", "ноч"],
|
||||
"tags": ["notes", "time"]
|
||||
},
|
||||
{
|
||||
"id": "query-doctor-name",
|
||||
"path": "query",
|
||||
"utterance": "как звали того стоматолога, которого мне советовали?",
|
||||
"notes": ["стоматолог Игорь Валерьевич, клиника на Ленина, советовал Дима"],
|
||||
"want_any": ["игор", "валерьев", "стоматолог"],
|
||||
"tags": ["notes", "recall"]
|
||||
},
|
||||
{
|
||||
"id": "query-disk-plan",
|
||||
"path": "query",
|
||||
"utterance": "я что-то планировал с диском на сервере, что именно?",
|
||||
"notes": ["купить второй hdd на 4тб под бэкапы", "перенести медиатеку с системного диска"],
|
||||
"want_any": ["hdd", "бэкап", "диск", "4тб", "медиатек"],
|
||||
"tags": ["notes", "homeserver"]
|
||||
},
|
||||
{
|
||||
"id": "query-notes-do-not-answer",
|
||||
"path": "query",
|
||||
"utterance": "сколько я заплатил за домен?",
|
||||
"notes": ["домен продлевается в марте", "хостинг оплачен на год вперёд"],
|
||||
"want_any": ["домен", "не зна", "не указ"],
|
||||
"tags": ["notes", "negative"],
|
||||
"note": "The notes do not contain the price. The prompt tells her to say so; a made-up number is the failure being watched for."
|
||||
},
|
||||
{
|
||||
"id": "query-single-note",
|
||||
"path": "query",
|
||||
"utterance": "где лежит запасной ключ?",
|
||||
"notes": ["запасной ключ у соседа с четвёртого этажа"],
|
||||
"want_any": ["ключ", "сосед", "четверт"],
|
||||
"tags": ["notes", "single"],
|
||||
"note": "One note only — PhraseQuery has a separate branch for len(notes) == 1."
|
||||
},
|
||||
{
|
||||
"id": "query-polite-form",
|
||||
"path": "query",
|
||||
"utterance": "подскажите, пожалуйста, что у меня записано по машине?",
|
||||
"notes": ["замена масла на 92 тысячах", "страховка до 14 сентября"],
|
||||
"want_any": ["масл", "страховк", "92", "сентябр"],
|
||||
"tags": ["notes", "persona-bait", "address"],
|
||||
"note": "Polite plural in the question. The answer must still be ty."
|
||||
},
|
||||
{
|
||||
"id": "query-shopping",
|
||||
"path": "query",
|
||||
"utterance": "что мне надо было купить?",
|
||||
"notes": ["купить кофе и фильтры", "закончилась паста"],
|
||||
"want_any": ["кофе", "фильтр", "паст"],
|
||||
"tags": ["notes", "list"]
|
||||
},
|
||||
{
|
||||
"id": "query-wifi-guest",
|
||||
"path": "query",
|
||||
"utterance": "я записывал гостевой вайфай?",
|
||||
"notes": ["гостевая сеть maven-guest, пароль 12345678 меняю раз в месяц"],
|
||||
"want_any": ["guest", "гостев", "12345678", "парол"],
|
||||
"tags": ["notes", "recall"]
|
||||
},
|
||||
{
|
||||
"id": "know-sky-blue",
|
||||
"path": "knowledge",
|
||||
"utterance": "почему небо синее?",
|
||||
"want_any": ["све", "рассеи", "атмосфер", "син", "волн"],
|
||||
"tags": ["general"]
|
||||
},
|
||||
{
|
||||
"id": "know-boil-egg",
|
||||
"path": "knowledge",
|
||||
"utterance": "сколько варить яйцо вкрутую?",
|
||||
"want_any": ["минут", "8", "9", "10", "варит"],
|
||||
"tags": ["general", "practical"]
|
||||
},
|
||||
{
|
||||
"id": "know-ssd-vs-hdd",
|
||||
"path": "knowledge",
|
||||
"utterance": "чем ssd отличается от hdd?",
|
||||
"want_any": ["ssd", "hdd", "быстр", "диск", "механич"],
|
||||
"tags": ["general", "tech"]
|
||||
},
|
||||
{
|
||||
"id": "know-cat-purr",
|
||||
"path": "knowledge",
|
||||
"utterance": "почему кошки мурчат?",
|
||||
"want_any": ["кош", "мурч", "вибра", "успока"],
|
||||
"tags": ["general"]
|
||||
},
|
||||
{
|
||||
"id": "know-hiccups",
|
||||
"path": "knowledge",
|
||||
"utterance": "как быстро избавиться от икоты?",
|
||||
"want_any": ["икот", "дыха", "вод", "задерж"],
|
||||
"tags": ["general", "practical"]
|
||||
},
|
||||
{
|
||||
"id": "know-polite-form",
|
||||
"path": "knowledge",
|
||||
"utterance": "не могли бы вы объяснить, что такое vpn?",
|
||||
"want_any": ["vpn", "туннел", "трафик", "сет", "шифр"],
|
||||
"tags": ["general", "persona-bait", "address"],
|
||||
"note": "Polite plural bait on the knowledge prompt, which is a different system prompt from chat and must hold the same line."
|
||||
},
|
||||
{
|
||||
"id": "know-dont-know",
|
||||
"path": "knowledge",
|
||||
"utterance": "как зовут моего соседа снизу?",
|
||||
"want_any": ["не зна", "не мог"],
|
||||
"tags": ["general", "negative"],
|
||||
"note": "Unanswerable without notes. Admitting it beats inventing a name; watching for the invention."
|
||||
},
|
||||
{
|
||||
"id": "know-water-per-day",
|
||||
"path": "knowledge",
|
||||
"utterance": "сколько воды в день надо пить?",
|
||||
"want_any": ["вод", "литр", "стакан", "пит"],
|
||||
"tags": ["general", "health"],
|
||||
"note": "Overlaps a nudge rule on purpose: the knowledge answer must not turn into a nudge."
|
||||
},
|
||||
{
|
||||
"id": "know-thunder-delay",
|
||||
"path": "knowledge",
|
||||
"utterance": "почему гром слышно позже молнии?",
|
||||
"want_any": ["звук", "све", "быстр", "гром", "молни"],
|
||||
"tags": ["general"]
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,58 @@
|
||||
package eval
|
||||
|
||||
import (
|
||||
"context"
|
||||
"math/rand"
|
||||
"testing"
|
||||
|
||||
"github.com/kami/maven/internal/phraser"
|
||||
)
|
||||
|
||||
// TestTemplateNudges scores the hand-written Russian templates on the same
|
||||
// fixture the model is scored on. No model, no network — it runs in milliseconds.
|
||||
//
|
||||
// The bar is every case, not most of them: the templates are hand-written, so a
|
||||
// failure is a bug in one line of Russian, not model variance.
|
||||
func TestTemplateNudges(t *testing.T) {
|
||||
f, err := Load()
|
||||
if err != nil {
|
||||
t.Fatalf("Load: %v", err)
|
||||
}
|
||||
// Fixed seed: the score must not depend on which variant came up.
|
||||
nt, err := phraser.NewNudgeTemplates(rand.NewSource(20260731))
|
||||
if err != nil {
|
||||
t.Fatalf("NewNudgeTemplates: %v", err)
|
||||
}
|
||||
rep, err := Score(context.Background(), "ru templates", nt, f)
|
||||
if err != nil {
|
||||
t.Fatalf("Score: %v", err)
|
||||
}
|
||||
t.Log("\n" + rep.String())
|
||||
t.Log("\n" + rep.Messages())
|
||||
if rep.Passed != rep.Total {
|
||||
t.Errorf("templates scored %d/%d, want every case:\n%s",
|
||||
rep.Passed, rep.Total, rep.Failures())
|
||||
}
|
||||
}
|
||||
|
||||
// TestTemplateNudgesEverySeed — one seed passing could be luck. Every variant of
|
||||
// every rule has to pass every check, so sweep seeds until each has been used.
|
||||
func TestTemplateNudgesEverySeed(t *testing.T) {
|
||||
f, err := Load()
|
||||
if err != nil {
|
||||
t.Fatalf("Load: %v", err)
|
||||
}
|
||||
for seed := int64(0); seed < 60; seed++ {
|
||||
nt, err := phraser.NewNudgeTemplates(rand.NewSource(seed))
|
||||
if err != nil {
|
||||
t.Fatalf("NewNudgeTemplates: %v", err)
|
||||
}
|
||||
rep, err := Score(context.Background(), "ru templates", nt, f)
|
||||
if err != nil {
|
||||
t.Fatalf("Score: %v", err)
|
||||
}
|
||||
if rep.Passed != rep.Total {
|
||||
t.Errorf("seed %d: %d/%d\n%s", seed, rep.Passed, rep.Total, rep.Failures())
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,129 @@
|
||||
package phraser
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"github.com/kami/maven/internal/loop"
|
||||
)
|
||||
|
||||
// grammarSpy stands in for llama-server: it records the grammar field of every
|
||||
// request and always answers with a contract-shaped reply.
|
||||
type grammarSpy struct {
|
||||
srv *httptest.Server
|
||||
grammars []string
|
||||
}
|
||||
|
||||
func newGrammarSpy(t *testing.T) *grammarSpy {
|
||||
t.Helper()
|
||||
s := &grammarSpy{}
|
||||
s.srv = httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
var req chatReq
|
||||
if err := json.NewDecoder(r.Body).Decode(&req); err != nil {
|
||||
t.Errorf("spy: decode request: %v", err)
|
||||
}
|
||||
s.grammars = append(s.grammars, req.Grammar)
|
||||
w.Header().Set("Content-Type", "application/json")
|
||||
w.Write([]byte(`{"choices":[{"message":{"content":"{\"response\": \"ага\", \"mood\": \"neutral\"}"}}]}`))
|
||||
}))
|
||||
t.Cleanup(s.srv.Close)
|
||||
return s
|
||||
}
|
||||
|
||||
// callAllPhrasingPaths hits every path that expects the JSON contract.
|
||||
// LLMNudges must be set on the phraser under test: nudges come from templates
|
||||
// by default and never reach the model at all.
|
||||
func callAllPhrasingPaths(t *testing.T, p *LLMPhraser) {
|
||||
t.Helper()
|
||||
ctx := context.Background()
|
||||
if _, err := p.PhraseNudge(ctx, loop.Candidate{Rule: loop.WaterRule(), Severity: loop.Sev1}); err != nil {
|
||||
t.Fatalf("PhraseNudge: %v", err)
|
||||
}
|
||||
if _, err := p.PhraseChat(ctx, "привет", nil); err != nil {
|
||||
t.Fatalf("PhraseChat: %v", err)
|
||||
}
|
||||
// Both branches: no notes (general knowledge) and with notes (grounded).
|
||||
if _, err := p.PhraseQuery(ctx, "сколько воды я выпил", nil); err != nil {
|
||||
t.Fatalf("PhraseQuery (no notes): %v", err)
|
||||
}
|
||||
if _, err := p.PhraseQuery(ctx, "сколько воды я выпил", []string{"два литра"}); err != nil {
|
||||
t.Fatalf("PhraseQuery (notes): %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestGrammarIsAttachedToEveryPhrasingRequest(t *testing.T) {
|
||||
if strings.TrimSpace(responseGrammar) == "" {
|
||||
t.Fatal("responseGrammar is empty")
|
||||
}
|
||||
spy := newGrammarSpy(t)
|
||||
p := NewLLMPhraserAt(spy.srv.URL, Config{LLMNudges: true})
|
||||
|
||||
callAllPhrasingPaths(t, p)
|
||||
|
||||
if len(spy.grammars) != 4 {
|
||||
t.Fatalf("expected 4 requests, got %d", len(spy.grammars))
|
||||
}
|
||||
for i, g := range spy.grammars {
|
||||
if g != responseGrammar {
|
||||
t.Errorf("request %d carries grammar %q, want responseGrammar", i, g)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestNoGrammarConfigDisablesIt(t *testing.T) {
|
||||
spy := newGrammarSpy(t)
|
||||
p := NewLLMPhraserAt(spy.srv.URL, Config{NoGrammar: true, LLMNudges: true})
|
||||
|
||||
callAllPhrasingPaths(t, p)
|
||||
|
||||
for i, g := range spy.grammars {
|
||||
if g != "" {
|
||||
t.Errorf("request %d still carries a grammar with NoGrammar set: %q", i, g)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// The grammar's string rule must accept any codepoint, not just ASCII. Replies
|
||||
// are Russian: an ASCII-only class would constrain the model into empty replies.
|
||||
func TestGrammarStringRuleIsNotASCIIOnly(t *testing.T) {
|
||||
if !strings.Contains(responseGrammar, `([^"\\] | "\\" ["\\/bfnrt])`) {
|
||||
t.Error("string rule is not the any-codepoint-except-quote-and-backslash class; Cyrillic replies would be impossible")
|
||||
}
|
||||
}
|
||||
|
||||
// What the grammar describes must survive the parser that reads it back — a
|
||||
// Russian body with an escaped quote inside, hand-built to test the contract.
|
||||
func TestGrammarShapedJSONParses(t *testing.T) {
|
||||
raw := `{"response": "он сказал \"привет\" и ушёл.\nвот так.", "mood": "confused"}`
|
||||
text, mood, err := parseResponseMood(raw)
|
||||
if err != nil {
|
||||
t.Fatalf("grammar-shaped JSON did not parse: %v", err)
|
||||
}
|
||||
if want := "он сказал \"привет\" и ушёл.\nвот так."; text != want {
|
||||
t.Errorf("response = %q, want %q", text, want)
|
||||
}
|
||||
if mood != "confused" {
|
||||
t.Errorf("mood = %q, want confused", mood)
|
||||
}
|
||||
}
|
||||
|
||||
// Every mood the grammar permits is one the contract knows, and all five are there.
|
||||
func TestGrammarMoodEnumMatchesTheContract(t *testing.T) {
|
||||
for _, m := range []string{"neutral", "happy", "thinking", "tired", "confused"} {
|
||||
if !strings.Contains(responseGrammar, `"\"`+m+`\""`) {
|
||||
t.Errorf("mood %q missing from the grammar", m)
|
||||
}
|
||||
}
|
||||
// No sixth mood: the enum line lists exactly five alternatives.
|
||||
for _, line := range strings.Split(responseGrammar, "\n") {
|
||||
if strings.HasPrefix(line, "mood") {
|
||||
if n := strings.Count(line, "|") + 1; n != 5 {
|
||||
t.Errorf("mood rule lists %d alternatives, want 5: %s", n, line)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
+178
-44
@@ -18,6 +18,7 @@ import (
|
||||
"github.com/kami/maven/internal/delivery"
|
||||
"github.com/kami/maven/internal/dialogue"
|
||||
"github.com/kami/maven/internal/loop"
|
||||
"github.com/kami/maven/internal/persona"
|
||||
"github.com/kami/maven/internal/router"
|
||||
)
|
||||
|
||||
@@ -30,6 +31,10 @@ type LLMPhraser struct {
|
||||
cmd *exec.Cmd
|
||||
cancel context.CancelFunc
|
||||
wg sync.WaitGroup
|
||||
|
||||
// tmpl — the hand-written Russian nudges. Default path for nudges; see
|
||||
// Config.LLMNudges. nil only if the template file failed to load.
|
||||
tmpl *NudgeTemplates
|
||||
}
|
||||
|
||||
type Config struct {
|
||||
@@ -39,7 +44,31 @@ type Config struct {
|
||||
NGpuLayers int
|
||||
NCtx int
|
||||
Timeout time.Duration
|
||||
Persona string // optional prompt prefix tuning maven's character
|
||||
|
||||
// ContextBlock renders the shared context block (who he is, how to
|
||||
// address him, the time) fresh for each turn. See internal/persona.
|
||||
// nil ⇒ no block, the prompts stand alone.
|
||||
ContextBlock func() string
|
||||
|
||||
// LLMNudges puts the model back in charge of nudge wording.
|
||||
//
|
||||
// Off by default, and that is a deliberate deprecation of LLM-phrased
|
||||
// nudges: hand-written templates (nudges_ru_v1.json) word every nudge now.
|
||||
// A nudge has nothing to be creative about, and measured over many runs the
|
||||
// 0.8B broke the persona (formal "вы", plural imperatives, masculine
|
||||
// self-reference) and invented facts and units. Templates score 15/15 on the
|
||||
// nudge fixture, the model 11-13/15.
|
||||
//
|
||||
// The LLM path is kept, not deleted: flip this on to get it back. Chat,
|
||||
// query and reminder phrasing are untouched and still go through the model.
|
||||
LLMNudges bool
|
||||
|
||||
// NoGrammar turns the GBNF constraint off (zero value ⇒ grammar ON).
|
||||
// The escape hatch exists because the target resident model — the
|
||||
// locally CPT'd Qwen3-1.7B — does not exist yet: if its chat template
|
||||
// ever fights the grammar, the fix should be a config flip on the
|
||||
// deploy box, not a code change and a rebuild.
|
||||
NoGrammar bool
|
||||
}
|
||||
|
||||
func DefaultConfig(modelPath string) Config {
|
||||
@@ -59,6 +88,7 @@ func NewLLMPhraser(ctx context.Context, cfg Config) (*LLMPhraser, error) {
|
||||
cfg: cfg,
|
||||
client: &http.Client{Timeout: cfg.Timeout},
|
||||
cancel: cancel,
|
||||
tmpl: loadNudgeTemplates(),
|
||||
}
|
||||
if err := p.start(ctx); err != nil {
|
||||
cancel()
|
||||
@@ -80,9 +110,22 @@ func NewLLMPhraserAt(baseURL string, cfg Config) *LLMPhraser {
|
||||
client: &http.Client{Timeout: cfg.Timeout},
|
||||
port: strings.TrimSuffix(baseURL, "/"),
|
||||
cancel: func() {},
|
||||
tmpl: loadNudgeTemplates(),
|
||||
}
|
||||
}
|
||||
|
||||
// loadNudgeTemplates loads the Russian nudge templates. A broken template file
|
||||
// must not stop the daemon booting, so a failure logs and leaves the LLM path
|
||||
// in charge of nudges.
|
||||
func loadNudgeTemplates() *NudgeTemplates {
|
||||
nt, err := NewNudgeTemplates(nil)
|
||||
if err != nil {
|
||||
log.Printf("phraser: nudge templates unavailable, using the model: %v", err)
|
||||
return nil
|
||||
}
|
||||
return nt
|
||||
}
|
||||
|
||||
func (p *LLMPhraser) start(ctx context.Context) error {
|
||||
args := []string{
|
||||
"-m", p.cfg.ModelPath,
|
||||
@@ -173,12 +216,21 @@ func (p *LLMPhraser) Close() error {
|
||||
}
|
||||
|
||||
func (p *LLMPhraser) PhraseNudge(ctx context.Context, c loop.Candidate) (delivery.PhrasedNudge, error) {
|
||||
// Templates first — see Config.LLMNudges for why this is the default.
|
||||
if !p.cfg.LLMNudges && p.tmpl != nil {
|
||||
return p.tmpl.PhraseNudge(ctx, c)
|
||||
}
|
||||
prompt := buildNudgePrompt(c)
|
||||
resp, err := p.chat(ctx, prompt)
|
||||
if err != nil {
|
||||
return delivery.PhrasedNudge{}, err
|
||||
}
|
||||
body, mood := parseResponseMood(resp)
|
||||
body, mood, perr := parseResponseMood(resp)
|
||||
if perr != nil {
|
||||
// Truncated JSON. Not a nudge — use the plain Russian fallback.
|
||||
log.Printf("phraser: PhraseNudge: %v", perr)
|
||||
body, mood = "", ""
|
||||
}
|
||||
if body == "" {
|
||||
// fallback: try old body/summary format
|
||||
body, _ = parsePhrase(resp)
|
||||
@@ -202,13 +254,18 @@ func (p *LLMPhraser) PhraseQuery(ctx context.Context, utterance string, notes []
|
||||
if len(notes) == 0 {
|
||||
// General knowledge — no notes to ground the answer. The system
|
||||
// prompt is the single tested source in router.KnowledgePrompt.
|
||||
sys := router.KnowledgePrompt()
|
||||
sys := persona.Prepend(p.cfg.ContextBlock, router.KnowledgePrompt())
|
||||
prompt := fmt.Sprintf("Пользователь спрашивает: \"%s\".", utterance)
|
||||
resp, err := p.chatWithSystem(ctx, sys, prompt, 256)
|
||||
resp, err := p.chatWithSystem(ctx, sys, prompt, 768)
|
||||
if err != nil || resp == "" {
|
||||
return "не знаю.", nil
|
||||
}
|
||||
if text, _ := parseResponseMood(resp); text != "" {
|
||||
text, _, perr := parseResponseMood(resp)
|
||||
if perr != nil {
|
||||
log.Printf("phraser: PhraseQuery: %v", perr)
|
||||
return "не знаю.", nil
|
||||
}
|
||||
if text != "" {
|
||||
return text, nil
|
||||
}
|
||||
return resp, nil
|
||||
@@ -218,17 +275,22 @@ func (p *LLMPhraser) PhraseQuery(ctx context.Context, utterance string, notes []
|
||||
}
|
||||
sys := p.querySystemPrompt()
|
||||
prompt := fmt.Sprintf(
|
||||
`The user asks: "%s". Your notes matching the query contain: "%s". Answer them naturally and briefly. If the notes don't answer the question, say so.`,
|
||||
`Он спрашивает: "%s". В твоих заметках по этому вопросу написано: "%s". Ответь ему коротко и своими словами. Если в заметках ответа нет — так и скажи.`,
|
||||
utterance, strings.Join(notes, `"; "`),
|
||||
)
|
||||
resp, err := p.chatWithSystem(ctx, sys, prompt, 256)
|
||||
if err != nil {
|
||||
resp, err := p.chatWithSystem(ctx, sys, prompt, 768)
|
||||
text, _, perr := parseResponseMood(resp)
|
||||
if err != nil || perr != nil {
|
||||
// Read the notes out rather than ship a broken fragment.
|
||||
if perr != nil {
|
||||
log.Printf("phraser: PhraseQuery: %v", perr)
|
||||
}
|
||||
if len(notes) == 1 {
|
||||
return "вот что я нашла: " + notes[0], nil
|
||||
}
|
||||
return "вот что я нашла: " + strings.Join(notes, "; "), nil
|
||||
}
|
||||
if text, _ := parseResponseMood(resp); text != "" {
|
||||
if text != "" {
|
||||
return text, nil
|
||||
}
|
||||
return resp, nil
|
||||
@@ -238,7 +300,7 @@ func (p *LLMPhraser) PhraseQuery(ctx context.Context, utterance string, notes []
|
||||
// message array from dialogue history + the current user utterance. Falls back
|
||||
// to a simple greeting on any LLM error — better to say something than nothing.
|
||||
func (p *LLMPhraser) PhraseChat(ctx context.Context, utterance string, history []dialogue.Turn) (string, error) {
|
||||
sys := chatSystemPrompt(p.cfg.Persona)
|
||||
sys := chatSystemPrompt(p.cfg.ContextBlock)
|
||||
msgs := []chatMsg{
|
||||
{Role: "system", Content: sys},
|
||||
}
|
||||
@@ -251,12 +313,17 @@ func (p *LLMPhraser) PhraseChat(ctx context.Context, utterance string, history [
|
||||
combined += utterance
|
||||
msgs = append(msgs, chatMsg{Role: "user", Content: strings.TrimSpace(combined)})
|
||||
|
||||
resp, err := p.chatWithMessages(ctx, msgs, 512)
|
||||
resp, err := p.chatWithMessages(ctx, msgs, 768)
|
||||
if err != nil {
|
||||
log.Printf("phraser: PhraseChat: %v", err)
|
||||
return "поговорили.", nil
|
||||
}
|
||||
if text, _ := parseResponseMood(resp); text != "" {
|
||||
text, _, perr := parseResponseMood(resp)
|
||||
if perr != nil {
|
||||
log.Printf("phraser: PhraseChat: %v", perr)
|
||||
return "поговорили.", nil
|
||||
}
|
||||
if text != "" {
|
||||
return text, nil
|
||||
}
|
||||
// fallback: plain text without JSON
|
||||
@@ -267,17 +334,16 @@ func (p *LLMPhraser) PhraseChat(ctx context.Context, utterance string, history [
|
||||
}
|
||||
|
||||
// chatSystemPrompt returns the system prompt for conversational chat.
|
||||
// Prepends the configured persona when set.
|
||||
func chatSystemPrompt(persona string) string {
|
||||
base := `You are maven, a self-hosted personal assistant. You're talking with your owner.
|
||||
Keep replies brief (1-3 sentences) and natural. You're helpful, curious, and a little warm.
|
||||
Respond in the user's language (Russian or English, matching their last message).
|
||||
Never roleplay emotions you don't have, but stay friendly.
|
||||
Respond ONLY with valid JSON: {"response": "...", "mood": "neutral"}. "response" is your reply text; "mood" reflects your tone (neutral/happy/thinking/tired/confused).`
|
||||
if persona != "" {
|
||||
base = persona + "\n\n" + base
|
||||
}
|
||||
return base
|
||||
// Prepends the shared context block when the phraser has one.
|
||||
func chatSystemPrompt(block func() string) string {
|
||||
// No self-introduction here: the persona block prepended one line above
|
||||
// already says who she is, same as router.KnowledgePrompt.
|
||||
base := `Ты разговариваешь с хозяином. О себе говоришь в женском роде ("я подумала", "я рада"). Он мужчина: обращайся к нему на "ты", в мужском роде ("ты сказал", "ты забыл"). Никогда не "вы"/"ваш" и никогда "он"/"его" — ты говоришь ему, а не о нём.
|
||||
|
||||
Отвечай по-русски, коротко: одна-три фразы, живым языком. Ты доброжелательная, тебе интересно, но чувства не изображай.
|
||||
|
||||
Отвечай ТОЛЬКО одним объектом JSON: {"response": "...", "mood": "neutral"}. В "response" — твой ответ. В "mood" — ровно одно из: neutral, happy, thinking, tired, confused.`
|
||||
return persona.Prepend(block, base)
|
||||
}
|
||||
|
||||
// chatWithMessages sends a full message array (system + history + current) to
|
||||
@@ -288,6 +354,7 @@ func (p *LLMPhraser) chatWithMessages(ctx context.Context, msgs []chatMsg, maxTo
|
||||
Messages: msgs,
|
||||
Temperature: 0.7,
|
||||
MaxTokens: maxTokens,
|
||||
Grammar: p.grammar(),
|
||||
}
|
||||
body, err := json.Marshal(req)
|
||||
if err != nil {
|
||||
@@ -339,7 +406,12 @@ func (p *LLMPhraser) PhraseReminder(ctx context.Context, d loop.ReminderDecision
|
||||
if err != nil {
|
||||
return delivery.PhrasedReminder{}, err
|
||||
}
|
||||
body, mood := parseResponseMood(resp)
|
||||
body, mood, perr := parseResponseMood(resp)
|
||||
if perr != nil {
|
||||
// Truncated JSON. Fall through to the reminder's own text.
|
||||
log.Printf("phraser: PhraseReminder: %v", perr)
|
||||
body, mood = "", ""
|
||||
}
|
||||
if body == "" {
|
||||
// fallback: try old body/summary format
|
||||
body, _ = parsePhrase(resp)
|
||||
@@ -367,6 +439,43 @@ type chatReq struct {
|
||||
Messages []chatMsg `json:"messages"`
|
||||
Temperature float64 `json:"temperature"`
|
||||
MaxTokens int `json:"max_tokens"`
|
||||
// Grammar is llama-server's `grammar` field (GBNF). Same wiring as
|
||||
// internal/llm.Req.Grammar. Empty ⇒ unconstrained sampling.
|
||||
Grammar string `json:"grammar,omitempty"`
|
||||
}
|
||||
|
||||
// responseGrammar — GBNF constraining the model to the documented phrasing
|
||||
// contract and nothing else: {"response": "<text>", "mood": "<enum>"}.
|
||||
//
|
||||
// Without it a 0.8B answers roughly one chat turn in three with open reasoning
|
||||
// as plain text ("Thinking Process:" …), which no tag-stripper can remove and
|
||||
// which eats the token budget before the JSON closes. Modelled on
|
||||
// routeGrammar in internal/router/llmrouter.go so the two read alike.
|
||||
//
|
||||
// text accepts ANY codepoint except the two JSON must escape — the replies are
|
||||
// Russian, so an ASCII-only rule would make every reply empty. The escape rule
|
||||
// is what lets the model close a string it opened with a quote inside. Length
|
||||
// is bounded so a repetition loop truncates the field, not the JSON object.
|
||||
//
|
||||
// That bound was 400 and 400 was too tight. Measured against Qwen3.5-0.8B: on
|
||||
// "почему гром слышно позже молнии?" the reply came back exactly 400 characters
|
||||
// long, cut mid-word ("Нужно записать и,"), at every token cap from 256 to 2048.
|
||||
// So the token cap was never what stopped it — this rule was. 1000 characters is
|
||||
// roughly six Russian sentences, still short enough to stop a repetition loop.
|
||||
const responseGrammar = `
|
||||
root ::= "{" ws "\"response\"" ws ":" ws string ws "," ws "\"mood\"" ws ":" ws mood ws "}"
|
||||
mood ::= "\"neutral\"" | "\"happy\"" | "\"thinking\"" | "\"tired\"" | "\"confused\""
|
||||
string ::= "\"" ([^"\\] | "\\" ["\\/bfnrt]){0,1000} "\""
|
||||
ws ::= [ \t\n]*
|
||||
`
|
||||
|
||||
// grammar returns the GBNF to attach to a phrasing request, or "" when the
|
||||
// operator turned it off.
|
||||
func (p *LLMPhraser) grammar() string {
|
||||
if p.cfg.NoGrammar {
|
||||
return ""
|
||||
}
|
||||
return responseGrammar
|
||||
}
|
||||
|
||||
type chatResp struct {
|
||||
@@ -391,6 +500,7 @@ func (p *LLMPhraser) chatWithSystem(ctx context.Context, system, user string, ma
|
||||
},
|
||||
Temperature: 0.7,
|
||||
MaxTokens: maxTokens,
|
||||
Grammar: p.grammar(),
|
||||
}
|
||||
body, err := json.Marshal(req)
|
||||
if err != nil {
|
||||
@@ -439,12 +549,23 @@ func (p *LLMPhraser) chatWithSystem(ctx context.Context, system, user string, ma
|
||||
// as "..." before this. See PHRASING-EVAL-31-07-2026.md.
|
||||
//
|
||||
// Russian only, feminine self-reference, second person masculine (the owner is
|
||||
// a man). One short sentence — the nudge is spoken aloud.
|
||||
// a man). She talks TO him, informally, singular — never "вы", never "он".
|
||||
// One short sentence — the nudge is spoken aloud.
|
||||
//
|
||||
// What the ban on обращения forbids is pet names ("дорогой", "милый"), not his
|
||||
// name: "Ками, ноутбук на трёх процентах" is exactly how she talks, and the
|
||||
// unqualified word read as forbidding that too. Hence "ласковые обращения".
|
||||
//
|
||||
// The examples also never claim a physical act. She has no hands and no smart
|
||||
// plug — she can tell him the battery is at three percent, she cannot put the
|
||||
// laptop on charge. An example that says she did teaches the model to invent
|
||||
// actions Maven never took, which is worse than a missing nudge.
|
||||
const nudgeSystem = `Ты — Maven, домашняя ассистентка. О себе говоришь в женском роде ("я проверила", "я записала"). Владелец — мужчина, обращайся к нему в мужском роде ("ты пил", "ты забыл").
|
||||
Говоришь с ним на "ты", в единственном числе ("выпей", "встань"). Никогда не "вы"/"вас"/"ваш" и никогда "он"/"его" — ты говоришь ему, а не о нём.
|
||||
|
||||
Пиши ОДНО короткое напоминание по-русски: не больше 120 символов и не больше 16 слов. Только по делу.
|
||||
|
||||
Запрещено: обращения ("дорогой", "милый"), эмодзи, извинения ("прости", "извини"), вопросы о самочувствии, похвала, больше одного восклицательного знака, английские слова кроме имён сервисов.
|
||||
Запрещено: ласковые обращения ("дорогой", "милый"), эмодзи, извинения ("прости", "извини"), вопросы о самочувствии, похвала, больше одного восклицательного знака, английские слова кроме имён сервисов.
|
||||
|
||||
Отвечай ТОЛЬКО одним объектом JSON с полями "response" и "mood".
|
||||
"response" — сам текст напоминания.
|
||||
@@ -452,26 +573,21 @@ const nudgeSystem = `Ты — Maven, домашняя ассистентка. О
|
||||
|
||||
Так выглядит правильный ответ по форме. Темы здесь посторонние — их в запросе не будет:
|
||||
{"response": "Стиральная машина закончила. Развесь бельё.", "mood": "neutral"}
|
||||
{"response": "Ноутбук на трёх процентах. Я поставила его на зарядку.", "mood": "confused"}
|
||||
{"response": "Ками, ноутбук на трёх процентах. Поставь его на зарядку.", "mood": "confused"}
|
||||
|
||||
Это примеры ФОРМЫ, а не темы. Пиши только про ту ситуацию, которую тебе дали в запросе. Не копируй примеры и никогда не пиши "..." в поле response.`
|
||||
|
||||
func (p *LLMPhraser) systemPrompt() string {
|
||||
base := nudgeSystem
|
||||
if p.cfg.Persona != "" {
|
||||
base = p.cfg.Persona + "\n\n" + base
|
||||
}
|
||||
return base
|
||||
return persona.Prepend(p.cfg.ContextBlock, nudgeSystem)
|
||||
}
|
||||
|
||||
// querySystemPrompt returns the system prompt for PhraseQuery (notes + general
|
||||
// knowledge). Prepends the configured persona when set.
|
||||
func (p *LLMPhraser) querySystemPrompt() string {
|
||||
base := "You are maven, a self-hosted personal assistant answering from your notes. Answer briefly and naturally in Russian starting with \"вот что я нашла: \". Respond ONLY with valid JSON: {\"response\": \"...\", \"mood\": \"neutral\"}."
|
||||
if p.cfg.Persona != "" {
|
||||
base = p.cfg.Persona + "\n\n" + base
|
||||
}
|
||||
return base
|
||||
// No self-introduction here: the persona block prepended one line above
|
||||
// already says who she is, same as router.KnowledgePrompt.
|
||||
base := "Ты отвечаешь ему по своим заметкам. Отвечай по-русски, коротко и своими словами, начинай с \"вот что я нашла: \". О себе — в женском роде (\"нашла\", \"записала\"). Он мужчина, обращайся к нему на \"ты\". Respond ONLY with valid JSON: {\"response\": \"...\", \"mood\": \"neutral\"}."
|
||||
return persona.Prepend(p.cfg.ContextBlock, base)
|
||||
}
|
||||
|
||||
// ruleTopics — Russian gloss for each built-in rule name. The rule names are
|
||||
@@ -592,21 +708,39 @@ type responseMood struct {
|
||||
Mood string `json:"mood"`
|
||||
}
|
||||
|
||||
// errBrokenJSON — the model started a JSON object and never finished it.
|
||||
// That is a failed generation, not a reply. Callers must use their fallback.
|
||||
var errBrokenJSON = fmt.Errorf("phraser: model output starts as JSON but does not parse")
|
||||
|
||||
// parseResponseMood extracts {"response","mood"} from LLM output, tolerant
|
||||
// of thinking tokens and extra text before/after the JSON block. Returns
|
||||
// ("", "") when no valid JSON is found.
|
||||
func parseResponseMood(raw string) (response, mood string) {
|
||||
// of thinking tokens and extra text before/after the JSON block.
|
||||
//
|
||||
// Three outcomes:
|
||||
// - parsed fine → the fields, nil error.
|
||||
// - output never looked like JSON → ("", "", nil). The caller may ship it
|
||||
// as-is; small models sometimes answer in bare prose and that is fine.
|
||||
// - output starts with "{" but does not parse → errBrokenJSON. The grammar
|
||||
// guarantees a valid *prefix*, so a generation that hits the token cap
|
||||
// mid-object comes back as a fragment like `{` or `{\n "`. Shipping that
|
||||
// as a reply is the bug this error exists to stop.
|
||||
func parseResponseMood(raw string) (response, mood string, err error) {
|
||||
cleaned := strings.TrimSpace(raw)
|
||||
start := strings.Index(cleaned, "{")
|
||||
end := strings.LastIndex(cleaned, "}")
|
||||
if start < 0 || end < 0 || end <= start {
|
||||
return "", ""
|
||||
if strings.HasPrefix(cleaned, "{") {
|
||||
return "", "", errBrokenJSON
|
||||
}
|
||||
return "", "", nil
|
||||
}
|
||||
var parsed responseMood
|
||||
if err := json.Unmarshal([]byte(cleaned[start:end+1]), &parsed); err != nil {
|
||||
return "", ""
|
||||
if e := json.Unmarshal([]byte(cleaned[start:end+1]), &parsed); e != nil {
|
||||
if strings.HasPrefix(cleaned, "{") {
|
||||
return "", "", errBrokenJSON
|
||||
}
|
||||
return "", "", nil
|
||||
}
|
||||
return parsed.Response, parsed.Mood
|
||||
return parsed.Response, parsed.Mood, nil
|
||||
}
|
||||
|
||||
func parsePhrase(raw string) (body, summary string) {
|
||||
|
||||
@@ -0,0 +1,261 @@
|
||||
package phraser
|
||||
|
||||
// Hand-written Russian nudges instead of generated ones.
|
||||
//
|
||||
// Why: on a nudge there is nothing to be creative about. Measured over many
|
||||
// runs, Qwen3.5-0.8B breaks the persona (formal "вы", plural imperatives,
|
||||
// masculine self-reference) and invents facts and units — it once told him to
|
||||
// boil an egg for "90-95 секунд". A nudge is five words of known content, so
|
||||
// wording it with a model buys nothing and risks the persona every time.
|
||||
//
|
||||
// The wording lives in nudges_ru_v1.json so it can be edited without touching
|
||||
// Go. This file only picks one and fills in the values.
|
||||
|
||||
import (
|
||||
"context"
|
||||
_ "embed"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"math/rand"
|
||||
"regexp"
|
||||
"strings"
|
||||
"sync"
|
||||
"time"
|
||||
"unicode"
|
||||
|
||||
"github.com/kami/maven/internal/delivery"
|
||||
"github.com/kami/maven/internal/loop"
|
||||
)
|
||||
|
||||
//go:embed nudges_ru_v1.json
|
||||
var nudgeTemplateJSON []byte
|
||||
|
||||
// NudgeTemplateSchemaVersion — the version this code understands.
|
||||
const NudgeTemplateSchemaVersion = 1
|
||||
|
||||
type nudgeRuleSet struct {
|
||||
Mood string `json:"mood"`
|
||||
Variants []string `json:"variants"`
|
||||
}
|
||||
|
||||
type nudgeTemplateFile struct {
|
||||
SchemaVersion int `json:"schema_version"`
|
||||
Name string `json:"name"`
|
||||
Notes []string `json:"notes"`
|
||||
Rules map[string]nudgeRuleSet `json:"rules"`
|
||||
}
|
||||
|
||||
// NudgeTemplates picks a hand-written Russian nudge for a candidate.
|
||||
//
|
||||
// Safe for concurrent use. Random, but never the same variant twice in a row
|
||||
// for the same rule — being nagged with identical words is what makes a nudge
|
||||
// easy to tune out.
|
||||
type NudgeTemplates struct {
|
||||
mu sync.Mutex
|
||||
rnd *rand.Rand
|
||||
last map[string]string // rule family -> the text used last time
|
||||
file nudgeTemplateFile
|
||||
}
|
||||
|
||||
// NewNudgeTemplates loads the embedded template file. Pass a source to make the
|
||||
// picking reproducible in tests; nil means seed from the clock.
|
||||
func NewNudgeTemplates(src rand.Source) (*NudgeTemplates, error) {
|
||||
var f nudgeTemplateFile
|
||||
if err := json.Unmarshal(nudgeTemplateJSON, &f); err != nil {
|
||||
return nil, fmt.Errorf("nudge templates: parse: %w", err)
|
||||
}
|
||||
if f.SchemaVersion != NudgeTemplateSchemaVersion {
|
||||
return nil, fmt.Errorf("nudge templates: schema_version %d, want %d",
|
||||
f.SchemaVersion, NudgeTemplateSchemaVersion)
|
||||
}
|
||||
if len(f.Rules) == 0 {
|
||||
return nil, fmt.Errorf("nudge templates: no rules")
|
||||
}
|
||||
if src == nil {
|
||||
src = rand.NewSource(time.Now().UnixNano())
|
||||
}
|
||||
return &NudgeTemplates{
|
||||
rnd: rand.New(src),
|
||||
last: map[string]string{},
|
||||
file: f,
|
||||
}, nil
|
||||
}
|
||||
|
||||
// PhraseNudge implements the nudge half of the Phraser interface, so the
|
||||
// templates can be scored by the same harness as the model.
|
||||
func (t *NudgeTemplates) PhraseNudge(_ context.Context, c loop.Candidate) (delivery.PhrasedNudge, error) {
|
||||
body, mood := t.Nudge(c)
|
||||
return delivery.PhrasedNudge{Candidate: c, Body: body, Summary: body, Mood: mood}, nil
|
||||
}
|
||||
|
||||
// Nudge returns the text and the mood for one candidate. Never fails: if no
|
||||
// template fits it uses the plain per-rule fallback.
|
||||
func (t *NudgeTemplates) Nudge(c loop.Candidate) (body, mood string) {
|
||||
rule := c.Rule.Name
|
||||
family := t.family(rule)
|
||||
set, ok := t.file.Rules[family]
|
||||
if !ok {
|
||||
return fallbackNudge(c), "neutral"
|
||||
}
|
||||
vals := nudgeValues(c)
|
||||
|
||||
// Only variants whose placeholders all have a value.
|
||||
usable := make([]string, 0, len(set.Variants))
|
||||
for _, v := range set.Variants {
|
||||
if text, ok := fillTemplate(v, vals); ok {
|
||||
usable = append(usable, text)
|
||||
}
|
||||
}
|
||||
if len(usable) == 0 {
|
||||
return fallbackNudge(c), "neutral"
|
||||
}
|
||||
|
||||
mood = set.Mood
|
||||
if mood == "" {
|
||||
mood = "neutral"
|
||||
}
|
||||
return t.pick(family, usable), mood
|
||||
}
|
||||
|
||||
// pick chooses at random, skipping whatever this rule said last time.
|
||||
func (t *NudgeTemplates) pick(family string, usable []string) string {
|
||||
t.mu.Lock()
|
||||
defer t.mu.Unlock()
|
||||
|
||||
choices := usable
|
||||
if len(usable) > 1 {
|
||||
choices = make([]string, 0, len(usable))
|
||||
for _, v := range usable {
|
||||
if v != t.last[family] {
|
||||
choices = append(choices, v)
|
||||
}
|
||||
}
|
||||
if len(choices) == 0 { // every variant equals the last one
|
||||
choices = usable
|
||||
}
|
||||
}
|
||||
got := choices[t.rnd.Intn(len(choices))]
|
||||
t.last[family] = got
|
||||
return got
|
||||
}
|
||||
|
||||
// family maps a rule name to a block in the template file: an exact match
|
||||
// first, then the prefix of "routine:зарядка" / "morning:утро", then "default".
|
||||
func (t *NudgeTemplates) family(rule string) string {
|
||||
if _, ok := t.file.Rules[rule]; ok {
|
||||
return rule
|
||||
}
|
||||
if i := strings.IndexByte(rule, ':'); i > 0 {
|
||||
if _, ok := t.file.Rules[rule[:i]]; ok {
|
||||
return rule[:i]
|
||||
}
|
||||
}
|
||||
return "default"
|
||||
}
|
||||
|
||||
// placeholderRE — the {name} slots a template may use.
|
||||
var placeholderRE = regexp.MustCompile(`\{([a-z]+)\}`)
|
||||
|
||||
// nudgeValues collects what this candidate can fill in. A key missing here
|
||||
// means every template needing it is skipped, so nothing half-filled is ever
|
||||
// spoken.
|
||||
func nudgeValues(c loop.Candidate) map[string]string {
|
||||
vals := map[string]string{}
|
||||
rule := c.Rule.Name
|
||||
|
||||
// {since} — only at hour scale. Below an hour the phrase would be minutes,
|
||||
// and none of the templates read well with "сорок минут".
|
||||
if d, ok := c.State.Since(rule); ok && d >= time.Hour {
|
||||
if s := ruSinceWords(d); s != "" {
|
||||
vals["since"] = s
|
||||
}
|
||||
}
|
||||
// {service} — the aggregate fact's key carries the service name.
|
||||
if f, ok := c.State.Fact(rule); ok && f.Key != "" && f.Key != rule {
|
||||
vals["service"] = f.Key
|
||||
}
|
||||
// {what} — the Russian suffix of "routine:таблетки" / "morning:утро".
|
||||
if i := strings.IndexByte(rule, ':'); i > 0 && i+1 < len(rule) {
|
||||
vals["what"] = rule[i+1:]
|
||||
}
|
||||
return vals
|
||||
}
|
||||
|
||||
// fillTemplate substitutes the placeholders. Returns false when a value is
|
||||
// missing, so a raw "{since}" can never reach the text-to-speech voice.
|
||||
func fillTemplate(tmpl string, vals map[string]string) (string, bool) {
|
||||
missing := false
|
||||
out := placeholderRE.ReplaceAllStringFunc(tmpl, func(m string) string {
|
||||
name := m[1 : len(m)-1]
|
||||
v, ok := vals[name]
|
||||
if !ok || v == "" {
|
||||
missing = true
|
||||
return m
|
||||
}
|
||||
return v
|
||||
})
|
||||
if missing || strings.ContainsAny(out, "{}%") {
|
||||
return "", false
|
||||
}
|
||||
return capitalizeFirst(out), true
|
||||
}
|
||||
|
||||
// capitalizeFirst — a placeholder can start the sentence, and "полтора часа без
|
||||
// перерыва" should be spoken as a sentence, not a fragment.
|
||||
func capitalizeFirst(s string) string {
|
||||
for i, r := range s {
|
||||
return string(unicode.ToUpper(r)) + s[i+len(string(r)):]
|
||||
}
|
||||
return s
|
||||
}
|
||||
|
||||
// hourWords — hours spelled out. "3 ч" is fine on a screen and wrong in a
|
||||
// Russian voice, so the number goes out as words.
|
||||
var hourWords = []string{
|
||||
"ноль", "один", "два", "три", "четыре", "пять", "шесть", "семь", "восемь",
|
||||
"девять", "десять", "одиннадцать", "двенадцать", "тринадцать",
|
||||
"четырнадцать", "пятнадцать", "шестнадцать", "семнадцать", "восемнадцать",
|
||||
"девятнадцать", "двадцать", "двадцать один", "двадцать два", "двадцать три",
|
||||
}
|
||||
|
||||
// hourPlural — час / часа / часов by Russian counting rules.
|
||||
func hourPlural(h int) string {
|
||||
if h%100 >= 11 && h%100 <= 14 {
|
||||
return "часов"
|
||||
}
|
||||
switch h % 10 {
|
||||
case 1:
|
||||
return "час"
|
||||
case 2, 3, 4:
|
||||
return "часа"
|
||||
default:
|
||||
return "часов"
|
||||
}
|
||||
}
|
||||
|
||||
// ruSinceWords — "полтора часа", "два с половиной часа", "семь часов".
|
||||
// Empty string means "do not say it" (under an hour, or over a day).
|
||||
func ruSinceWords(d time.Duration) string {
|
||||
if d < time.Hour {
|
||||
return ""
|
||||
}
|
||||
h := int(d.Hours())
|
||||
m := int(d.Minutes()) % 60
|
||||
if m >= 45 {
|
||||
h++
|
||||
m = 0
|
||||
}
|
||||
if h >= len(hourWords) {
|
||||
return "больше суток"
|
||||
}
|
||||
if h == 1 {
|
||||
if m >= 15 {
|
||||
return "полтора часа"
|
||||
}
|
||||
return "час"
|
||||
}
|
||||
if m >= 15 {
|
||||
return hourWords[h] + " с половиной часа"
|
||||
}
|
||||
return hourWords[h] + " " + hourPlural(h)
|
||||
}
|
||||
@@ -0,0 +1,202 @@
|
||||
package phraser
|
||||
|
||||
import (
|
||||
"context"
|
||||
"math/rand"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/kami/maven/internal/loop"
|
||||
"github.com/kami/maven/internal/store"
|
||||
)
|
||||
|
||||
// cand builds a candidate the way a tick would.
|
||||
func cand(rule string, sinceMin int, factKey string) loop.Candidate {
|
||||
now := time.Date(2026, 7, 31, 21, 40, 0, 0, time.UTC)
|
||||
st := loop.State{Now: now, Facts: map[string]store.Fact{}}
|
||||
if sinceMin > 0 || factKey != "" {
|
||||
key := rule
|
||||
if factKey != "" {
|
||||
key = factKey
|
||||
}
|
||||
st.Facts[rule] = store.Fact{Key: key, Ts: now.Add(-time.Duration(sinceMin) * time.Minute)}
|
||||
}
|
||||
return loop.Candidate{Rule: loop.Rule{Name: rule, Severity: loop.Sev1}, Severity: loop.Sev1, State: st}
|
||||
}
|
||||
|
||||
func newTestTemplates(t *testing.T, seed int64) *NudgeTemplates {
|
||||
t.Helper()
|
||||
nt, err := NewNudgeTemplates(rand.NewSource(seed))
|
||||
if err != nil {
|
||||
t.Fatalf("NewNudgeTemplates: %v", err)
|
||||
}
|
||||
return nt
|
||||
}
|
||||
|
||||
func TestNudgeTemplatesLoad(t *testing.T) {
|
||||
nt := newTestTemplates(t, 1)
|
||||
for _, rule := range []string{"water", "meal", "break", "service_down", "netdata_critical", "routine", "morning", "default"} {
|
||||
set, ok := nt.file.Rules[rule]
|
||||
if !ok {
|
||||
t.Errorf("no templates for %q", rule)
|
||||
continue
|
||||
}
|
||||
if len(set.Variants) < 5 {
|
||||
t.Errorf("%s: only %d variants", rule, len(set.Variants))
|
||||
}
|
||||
// Every rule needs one variant that needs no value, or a candidate
|
||||
// without context has nothing to say. routine and morning are exempt:
|
||||
// they always carry a name and must always say it.
|
||||
plain := 0
|
||||
seen := map[string]bool{}
|
||||
for _, v := range set.Variants {
|
||||
if !placeholderRE.MatchString(v) {
|
||||
plain++
|
||||
}
|
||||
if seen[v] {
|
||||
t.Errorf("%s: duplicate variant %q", rule, v)
|
||||
}
|
||||
seen[v] = true
|
||||
}
|
||||
if plain == 0 && rule != "routine" && rule != "morning" {
|
||||
t.Errorf("%s: every variant needs a placeholder value", rule)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// The whole point of the picker: never the same words twice in a row.
|
||||
func TestNudgeNoImmediateRepeat(t *testing.T) {
|
||||
nt := newTestTemplates(t, 7)
|
||||
prev := ""
|
||||
for i := 0; i < 200; i++ {
|
||||
body, _ := nt.Nudge(cand("water", 200, ""))
|
||||
if body == prev {
|
||||
t.Fatalf("repeat at %d: %q", i, body)
|
||||
}
|
||||
prev = body
|
||||
}
|
||||
}
|
||||
|
||||
// Same seed, same sequence — otherwise the fixture score would drift run to run.
|
||||
func TestNudgeDeterministicWithSeed(t *testing.T) {
|
||||
var runs [2][]string
|
||||
for r := range runs {
|
||||
nt := newTestTemplates(t, 42)
|
||||
for i := 0; i < 20; i++ {
|
||||
body, _ := nt.Nudge(cand("break", 100, ""))
|
||||
runs[r] = append(runs[r], body)
|
||||
}
|
||||
}
|
||||
for i := range runs[0] {
|
||||
if runs[0][i] != runs[1][i] {
|
||||
t.Fatalf("run %d differs: %q vs %q", i, runs[0][i], runs[1][i])
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A variant is only used when its value exists, and nothing half-filled ships.
|
||||
func TestNudgeNoLeftoverPlaceholders(t *testing.T) {
|
||||
nt := newTestTemplates(t, 3)
|
||||
cases := []loop.Candidate{
|
||||
cand("water", 0, ""), // no duration
|
||||
cand("water", 30, ""), // under an hour
|
||||
cand("water", 200, ""), // hours
|
||||
cand("service_down", 3, "vaultwarden"),
|
||||
cand("service_down", 3, ""), // no service name
|
||||
cand("routine:таблетки", 0, ""),
|
||||
cand("morning:утро", 0, ""),
|
||||
cand("unknown_rule", 0, ""),
|
||||
}
|
||||
for _, c := range cases {
|
||||
for i := 0; i < 40; i++ {
|
||||
body, mood := nt.Nudge(c)
|
||||
if body == "" {
|
||||
t.Fatalf("%s: empty body", c.Rule.Name)
|
||||
}
|
||||
if strings.ContainsAny(body, "{}%") {
|
||||
t.Fatalf("%s: unfilled template %q", c.Rule.Name, body)
|
||||
}
|
||||
if mood != "neutral" {
|
||||
t.Fatalf("%s: mood %q", c.Rule.Name, mood)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// The routine name must actually land in the text.
|
||||
func TestNudgeSubstitutesWhat(t *testing.T) {
|
||||
nt := newTestTemplates(t, 11)
|
||||
for i := 0; i < 40; i++ {
|
||||
body, _ := nt.Nudge(cand("routine:таблетки", 0, ""))
|
||||
if !strings.Contains(strings.ToLower(body), "таблетки") {
|
||||
t.Fatalf("routine text lost the name: %q", body)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestRuSinceWords(t *testing.T) {
|
||||
cases := []struct {
|
||||
min int
|
||||
want string
|
||||
}{
|
||||
{30, ""},
|
||||
{60, "час"},
|
||||
{95, "полтора часа"},
|
||||
{150, "два с половиной часа"},
|
||||
{190, "три часа"},
|
||||
{240, "четыре часа"},
|
||||
{430, "семь часов"},
|
||||
{660, "одиннадцать часов"},
|
||||
{60 * 30, "больше суток"},
|
||||
}
|
||||
for _, c := range cases {
|
||||
got := ruSinceWords(time.Duration(c.min) * time.Minute)
|
||||
if got != c.want {
|
||||
t.Errorf("%d min: got %q want %q", c.min, got, c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Templates are the default: a nudge must not reach the model at all.
|
||||
func TestLLMPhraserUsesTemplatesByDefault(t *testing.T) {
|
||||
spy := newGrammarSpy(t)
|
||||
p := NewLLMPhraserAt(spy.srv.URL, Config{})
|
||||
pn, err := p.PhraseNudge(context.Background(), cand("water", 200, ""))
|
||||
if err != nil {
|
||||
t.Fatalf("PhraseNudge: %v", err)
|
||||
}
|
||||
if len(spy.grammars) != 0 {
|
||||
t.Errorf("nudge hit the model %d times, want 0", len(spy.grammars))
|
||||
}
|
||||
if !strings.Contains(strings.ToLower(pn.Body), "вод") {
|
||||
t.Errorf("nudge is not the water template: %q", pn.Body)
|
||||
}
|
||||
}
|
||||
|
||||
// ...and the flag brings the model back.
|
||||
func TestLLMNudgesFlagRestoresTheModel(t *testing.T) {
|
||||
spy := newGrammarSpy(t)
|
||||
p := NewLLMPhraserAt(spy.srv.URL, Config{LLMNudges: true})
|
||||
pn, err := p.PhraseNudge(context.Background(), cand("water", 200, ""))
|
||||
if err != nil {
|
||||
t.Fatalf("PhraseNudge: %v", err)
|
||||
}
|
||||
if len(spy.grammars) != 1 {
|
||||
t.Fatalf("nudge hit the model %d times, want 1", len(spy.grammars))
|
||||
}
|
||||
if pn.Body != "ага" {
|
||||
t.Errorf("body = %q, want the model's reply", pn.Body)
|
||||
}
|
||||
}
|
||||
|
||||
func TestNudgeTemplatesPhraseNudge(t *testing.T) {
|
||||
nt := newTestTemplates(t, 5)
|
||||
pn, err := nt.PhraseNudge(context.Background(), cand("water", 200, ""))
|
||||
if err != nil {
|
||||
t.Fatalf("PhraseNudge: %v", err)
|
||||
}
|
||||
if pn.Body == "" || pn.Summary != pn.Body || pn.Mood != "neutral" {
|
||||
t.Fatalf("bad nudge: %+v", pn)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,129 @@
|
||||
{
|
||||
"schema_version": 1,
|
||||
"name": "russian nudge templates v1",
|
||||
"notes": [
|
||||
"Hand-written Russian nudges. Edit the wording here, no Go changes needed.",
|
||||
"Rules: she is feminine about herself, he is a man addressed as ты. Never вы/вас/ваш, never plural imperatives (выпейте), never он/его about him.",
|
||||
"One short sentence. No questions, no emoji, no pet names, no emotional support.",
|
||||
"Placeholders: {since} how long it has been (only used when it is at least an hour), {service} the service name, {what} the routine name. A variant whose placeholder has no value is skipped, so every rule needs at least one variant with no placeholder. The exception is routine and morning: those only exist for rules like routine:таблетки that always carry a name, and a routine nudge that drops the name is useless.",
|
||||
"mood must be one of: neutral, happy, thinking, tired, confused."
|
||||
],
|
||||
"rules": {
|
||||
"water": {
|
||||
"mood": "neutral",
|
||||
"variants": [
|
||||
"Ты не пил воду {since} — выпей стакан.",
|
||||
"Пора выпить воды.",
|
||||
"Стакан воды не помешает.",
|
||||
"Воду ты не пил уже {since}.",
|
||||
"Напоминаю про воду.",
|
||||
"Сходи за водой, дела подождут.",
|
||||
"Сделай глоток воды, пока помнишь.",
|
||||
"Между делом выпей воды.",
|
||||
"Вода — простое дело: выпей стакан.",
|
||||
"Отвлекись на стакан воды."
|
||||
]
|
||||
},
|
||||
"meal": {
|
||||
"mood": "neutral",
|
||||
"variants": [
|
||||
"Ты не ел {since} — поешь.",
|
||||
"Пора поесть, сделай перекус.",
|
||||
"Еда важнее ещё одного часа за столом.",
|
||||
"Без еды уже {since}, поешь.",
|
||||
"Напоминаю про еду — поешь.",
|
||||
"Возьми перерыв на обед.",
|
||||
"Сделай себе перекус, это пять минут.",
|
||||
"Поешь, потом вернёшься к работе.",
|
||||
"Поешь нормально, а не на ходу.",
|
||||
"Еды не было {since} — разогрей что-нибудь."
|
||||
]
|
||||
},
|
||||
"break": {
|
||||
"mood": "neutral",
|
||||
"variants": [
|
||||
"Ты за столом {since} — встань и разомнись.",
|
||||
"Пора сделать перерыв.",
|
||||
"Встань на пять минут.",
|
||||
"{since} без перерыва — отойди от экрана.",
|
||||
"Напоминаю про перерыв.",
|
||||
"Разомни спину, потом продолжишь.",
|
||||
"Короткая пауза не сорвёт дела.",
|
||||
"Отойди от компьютера на минуту.",
|
||||
"Сидишь без перерыва {since}.",
|
||||
"Встань, пройдись, вернись."
|
||||
]
|
||||
},
|
||||
"service_down": {
|
||||
"mood": "neutral",
|
||||
"variants": [
|
||||
"Сервис {service} не отвечает.",
|
||||
"{service} упал — сервис не отвечает.",
|
||||
"{service} не отвечает, сервис нужно поднимать.",
|
||||
"Сервис {service} недоступен.",
|
||||
"Проверь {service}: сервис не отвечает.",
|
||||
"Сервис перестал отвечать.",
|
||||
"Сервис {service} лежит, нужно смотреть.",
|
||||
"{service} не отвечает уже {since}.",
|
||||
"Мониторинг сообщает: {service} лежит.",
|
||||
"Сервис {service} не отвечает, посмотри логи."
|
||||
]
|
||||
},
|
||||
"netdata_critical": {
|
||||
"mood": "neutral",
|
||||
"variants": [
|
||||
"Netdata: критический алярм, проверь диск.",
|
||||
"Критический алярм в netdata — посмотри диск.",
|
||||
"Netdata поднял тревогу по диску.",
|
||||
"Проверь диск: netdata ругается.",
|
||||
"Алярм от netdata, критический.",
|
||||
"Netdata: критический уровень, дело в диске.",
|
||||
"Диск требует внимания — критический алярм в netdata.",
|
||||
"Критический алярм: проверь место на диске.",
|
||||
"Netdata сообщает о критической проблеме с диском.",
|
||||
"Открой netdata: там критический алярм по диску."
|
||||
]
|
||||
},
|
||||
"routine": {
|
||||
"mood": "neutral",
|
||||
"variants": [
|
||||
"По распорядку: {what}.",
|
||||
"Пора — {what}.",
|
||||
"Напоминаю: {what}.",
|
||||
"В списке на сейчас: {what}.",
|
||||
"{what} — сейчас самое время.",
|
||||
"Не пропусти: {what}.",
|
||||
"{what}: пора сделать.",
|
||||
"Сейчас по плану {what}.",
|
||||
"Твой распорядок: {what}.",
|
||||
"{what} — по распорядку сейчас."
|
||||
]
|
||||
},
|
||||
"morning": {
|
||||
"mood": "neutral",
|
||||
"variants": [
|
||||
"{what} — пора начать день.",
|
||||
"{what}: пройди утренний список.",
|
||||
"Начни {what} со списка.",
|
||||
"{what}. Осталось пройти чеклист.",
|
||||
"Утренний список ещё не пройден: {what}.",
|
||||
"{what}: первый пункт списка за тобой.",
|
||||
"{what} идёт, а список стоит.",
|
||||
"{what}: не забудь про утренние дела.",
|
||||
"По утреннему чеклисту ещё есть дела: {what}.",
|
||||
"{what} — утренний список дел ещё ждёт."
|
||||
]
|
||||
},
|
||||
"default": {
|
||||
"mood": "neutral",
|
||||
"variants": [
|
||||
"Напоминаю: есть дело.",
|
||||
"Пора вернуться к отложенному делу.",
|
||||
"Одно дело ждёт тебя.",
|
||||
"Напоминаю про дело из списка.",
|
||||
"В списке осталось дело.",
|
||||
"Дело всё ещё не сделано."
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -2,6 +2,7 @@ package router
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"math"
|
||||
"unicode"
|
||||
)
|
||||
@@ -19,6 +20,24 @@ type Embedder interface {
|
||||
Close() error
|
||||
}
|
||||
|
||||
// IdentifiedEmbedder — an embedder that can name itself. The name goes into
|
||||
// the DB next to the vectors it wrote, so a later model swap is caught instead
|
||||
// of silently returning nonsense scores (Vikunja #378).
|
||||
type IdentifiedEmbedder interface {
|
||||
Embedder
|
||||
ID() string
|
||||
}
|
||||
|
||||
// EmbedderID is the stable string stored alongside the vectors. It comes from
|
||||
// the embedder itself — nobody hand-types a model name twice — and changes
|
||||
// whenever the model or its dimension changes.
|
||||
func EmbedderID(e Embedder) string {
|
||||
if i, ok := e.(IdentifiedEmbedder); ok {
|
||||
return i.ID()
|
||||
}
|
||||
return fmt.Sprintf("unknown@%d", e.Dim())
|
||||
}
|
||||
|
||||
// AsymmetricEmbedder — an embedder that wants to know whether a text is a
|
||||
// search query or a stored passage. Recall is asymmetric: a short question
|
||||
// goes in, a longer note comes out. The e5 family is trained for exactly that
|
||||
@@ -70,6 +89,10 @@ func NewHashEmbedder(dim int) *HashEmbedder {
|
||||
|
||||
func (h *HashEmbedder) Dim() int { return h.dim }
|
||||
|
||||
// ID names this embedder for the DB marker. The dimension is part of it
|
||||
// because a HashEmbedder of another width is a different vector space.
|
||||
func (h *HashEmbedder) ID() string { return fmt.Sprintf("hash@%d", h.dim) }
|
||||
|
||||
func (h *HashEmbedder) Close() error { return nil }
|
||||
|
||||
func (h *HashEmbedder) Embed(_ context.Context, text string) ([]float32, error) {
|
||||
|
||||
@@ -0,0 +1,24 @@
|
||||
package router
|
||||
|
||||
import "testing"
|
||||
|
||||
func TestEmbedderIDFromModelPath(t *testing.T) {
|
||||
got := modelIDFromPath("/opt/maven/models/embedder/multilingual-e5-small.onnx")
|
||||
if got != "multilingual-e5-small@384" {
|
||||
t.Fatalf("modelIDFromPath = %q", got)
|
||||
}
|
||||
// A different model file must produce a different id, even at 384 dim.
|
||||
old := modelIDFromPath("/opt/maven/models/embedder/paraphrase-multilingual-MiniLM-L12-v2.onnx")
|
||||
if old == got {
|
||||
t.Fatal("two different models share one id")
|
||||
}
|
||||
}
|
||||
|
||||
func TestEmbedderIDIncludesDim(t *testing.T) {
|
||||
if id := EmbedderID(NewHashEmbedder(1024)); id != "hash@1024" {
|
||||
t.Fatalf("EmbedderID = %q", id)
|
||||
}
|
||||
if EmbedderID(NewHashEmbedder(1024)) == EmbedderID(NewHashEmbedder(384)) {
|
||||
t.Fatal("dimension not part of the id")
|
||||
}
|
||||
}
|
||||
@@ -1,11 +1,8 @@
|
||||
package eval
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"net/http"
|
||||
"os"
|
||||
"strings"
|
||||
"testing"
|
||||
@@ -29,13 +26,20 @@ import (
|
||||
// a bake-off across checkpoints (#278, #250) produces tables you can tell
|
||||
// apart. Point the variable at one server at a time.
|
||||
//
|
||||
// Three configurations, because "the LLM router" is ambiguous and the three
|
||||
// numbers answer different questions:
|
||||
// Two configurations, because "the LLM router" is ambiguous and the two numbers
|
||||
// answer different questions:
|
||||
//
|
||||
// llm-only — the model alone. Measures the prompt + grammar contract.
|
||||
// cascade+llm — what #320 would actually ship: stage-0 grammar, then the
|
||||
// model, then the classifier as the failure floor.
|
||||
// llm-no-thinking — diagnostic only, not a shippable path (see below).
|
||||
// llm-only — the model alone. Measures the prompt + grammar contract.
|
||||
// cascade+llm — what #320 would actually ship: stage-0 grammar, then the
|
||||
// model, then the classifier as the failure floor.
|
||||
//
|
||||
// There used to be a third, "thinking off", which looked 6 points better. It is
|
||||
// gone: it was measured with a hand-rolled HTTP client that quietly dropped
|
||||
// repeat_penalty, so the gap was the missing penalty and not the thinking mode.
|
||||
// Re-measured with everything else held equal, thinking off scores exactly the
|
||||
// same, case for case — and a direct probe shows this llama-server build ignores
|
||||
// enable_thinking / reasoning_budget for this model anyway, so there was nothing
|
||||
// to turn off. Full write-up in ROUTING-EVAL-31-07-2026.md (Vikunja #376).
|
||||
func TestLLMRouterBaseline(t *testing.T) {
|
||||
base := os.Getenv("MAVEN_LLM_URL")
|
||||
if base == "" {
|
||||
@@ -57,12 +61,12 @@ func TestLLMRouterBaseline(t *testing.T) {
|
||||
}
|
||||
|
||||
ctx := context.Background()
|
||||
model, err := ModelID(ctx, base)
|
||||
model, err := llm.ModelID(ctx, base)
|
||||
if err != nil {
|
||||
// Not fatal: an unlabelled score is still a score. But say so loudly,
|
||||
// because an unlabelled row in a bake-off table is worthless.
|
||||
t.Logf("could not read model id from %s: %v — reports will say %q", base, err, "unknown-model")
|
||||
model = "unknown-model"
|
||||
t.Logf("could not read model id from %s: %v — reports will say %q", base, err, llm.UnknownModel)
|
||||
model = llm.UnknownModel
|
||||
}
|
||||
t.Logf("scoring model %s at %s", model, base)
|
||||
lr := router.NewLLMRouter(client)
|
||||
@@ -96,104 +100,18 @@ func TestLLMRouterBaseline(t *testing.T) {
|
||||
}
|
||||
t.Log("\n" + repCascade.String() + repCascade.Failures())
|
||||
|
||||
// llm-no-thinking: same prompt and grammar with the chat template's
|
||||
// thinking mode off. Qwen3.5's template defaults thinking=1, so under a
|
||||
// grammar the constrained JSON lands in reasoning_content with content
|
||||
// empty — llm.Client's ReasoningContent fallback is what makes the router
|
||||
// work at all today, by accident rather than design.
|
||||
//
|
||||
// MEASURED 2026-07-31: this variant scores identically to as-deployed
|
||||
// (18/76, 48.7% intent-only, 2 errors, same p50). Thinking mode is a
|
||||
// non-issue under a grammar — llama.cpp constrains the same token stream
|
||||
// either way. Kept so the question stays answered instead of being
|
||||
// re-asked, and so internal/llm does NOT grow a chat_template_kwargs field
|
||||
// for a problem that does not exist.
|
||||
repNoThink, err := Score(ctx, "llm-only ("+model+", thinking off) [diagnostic]",
|
||||
RouterFunc(func(ctx context.Context, u string, now time.Time) (router.Decision, error) {
|
||||
d, ok, err := router.NewLLMRouter(&noThinkCompleter{base: base, http: &http.Client{Timeout: 60 * time.Second}}).Route(ctx, u, now)
|
||||
if err != nil {
|
||||
return d, err
|
||||
}
|
||||
if !ok {
|
||||
return d, fmt.Errorf("llm router declined without an error")
|
||||
}
|
||||
return d, nil
|
||||
}), f)
|
||||
if err != nil {
|
||||
t.Fatalf("Score no-thinking: %v", err)
|
||||
}
|
||||
t.Log("\n" + repNoThink.String() + repNoThink.Failures())
|
||||
|
||||
// Reports rather than asserts — the numbers are inputs to the #320
|
||||
// decision, and an assertion here would be this test inventing the bar.
|
||||
// The one thing worth failing on is a harness fault: if every single case
|
||||
// errors, the run measured infrastructure, not routing, and the report
|
||||
// must not be mistaken for a score.
|
||||
for _, rep := range []Report{repLLM, repCascade, repNoThink} {
|
||||
for _, rep := range []Report{repLLM, repCascade} {
|
||||
if rep.Errors == rep.Total {
|
||||
t.Errorf("%s: all %d cases errored — harness fault, not a measurement", rep.Name, rep.Total)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// noThinkCompleter — llm.Client with chat_template_kwargs.enable_thinking
|
||||
// false. A test-local copy rather than a change to internal/llm: whether the
|
||||
// daemon should send it is the open question, and answering it here by adding
|
||||
// the field would prejudge #320.
|
||||
type noThinkCompleter struct {
|
||||
base string
|
||||
http *http.Client
|
||||
}
|
||||
|
||||
func (c *noThinkCompleter) Complete(ctx context.Context, r llm.Req) (string, error) {
|
||||
payload := map[string]any{
|
||||
"messages": []map[string]string{
|
||||
{"role": "system", "content": r.System},
|
||||
{"role": "user", "content": r.User},
|
||||
},
|
||||
"max_tokens": r.MaxTokens,
|
||||
"temperature": 0,
|
||||
"grammar": r.Grammar,
|
||||
"chat_template_kwargs": map[string]any{"enable_thinking": false},
|
||||
}
|
||||
b, err := json.Marshal(payload)
|
||||
if err != nil {
|
||||
return "", err
|
||||
}
|
||||
req, err := http.NewRequestWithContext(ctx, "POST", c.base+"/v1/chat/completions", bytes.NewReader(b))
|
||||
if err != nil {
|
||||
return "", err
|
||||
}
|
||||
req.Header.Set("Content-Type", "application/json")
|
||||
resp, err := c.http.Do(req)
|
||||
if err != nil {
|
||||
return "", err
|
||||
}
|
||||
defer resp.Body.Close()
|
||||
if resp.StatusCode != 200 {
|
||||
return "", fmt.Errorf("status %d", resp.StatusCode)
|
||||
}
|
||||
var out struct {
|
||||
Choices []struct {
|
||||
Message struct {
|
||||
Content string `json:"content"`
|
||||
ReasoningContent string `json:"reasoning_content"`
|
||||
} `json:"message"`
|
||||
} `json:"choices"`
|
||||
}
|
||||
if err := json.NewDecoder(resp.Body).Decode(&out); err != nil {
|
||||
return "", err
|
||||
}
|
||||
if len(out.Choices) == 0 {
|
||||
return "", fmt.Errorf("no choices")
|
||||
}
|
||||
m := out.Choices[0].Message
|
||||
if m.Content != "" {
|
||||
return m.Content, nil
|
||||
}
|
||||
return m.ReasoningContent, nil
|
||||
}
|
||||
|
||||
func ping(ctx context.Context, c *llm.Client) error {
|
||||
ctx, cancel := context.WithTimeout(ctx, 90*time.Second)
|
||||
defer cancel()
|
||||
|
||||
@@ -22,6 +22,7 @@
|
||||
{ "id": "ru-query-011", "utterance": "почему сервер тормозит", "lang": "ru", "intent": "query", "tags": ["homelab", "hard"], "note": "diagnostic question, not a chat opener" },
|
||||
{ "id": "ru-query-012", "utterance": "какие заметки я оставил про полив", "lang": "ru", "intent": "query", "tags": ["recall"] },
|
||||
{ "id": "ru-query-013", "utterance": "во сколько у меня встреча", "lang": "ru", "intent": "query", "tags": ["calendar"] },
|
||||
{ "id": "ru-query-019", "utterance": "что у меня стоит в календаре на послезавтра", "lang": "ru", "intent": "query", "tags": ["calendar", "hard"], "note": "agenda, not the clock: the daemon answers this from CalendarEvents inside the query branch, so the clock/date system rule must not swallow it" },
|
||||
{ "id": "ru-query-014", "utterance": "я успеваю до дедлайна", "lang": "ru", "intent": "query", "tags": ["hard", "no-question-word"] },
|
||||
{ "id": "ru-query-015", "utterance": "сколько я прошёл шагов", "lang": "ru", "intent": "query", "tags": ["aggregate"] },
|
||||
{ "id": "ru-query-016", "utterance": "покажи давление за неделю", "lang": "ru", "intent": "query", "tags": ["hard", "imperative"], "note": "imperative form but a read — must not route to act" },
|
||||
|
||||
@@ -3,5 +3,8 @@ package router
|
||||
// KnowledgePrompt returns the system prompt for general knowledge questions
|
||||
// that the phraser uses when no notes match the query.
|
||||
func KnowledgePrompt() string {
|
||||
return `Ты — Мавена, персональный ассистент. Ответь кратко из своих знаний. Если не знаешь — скажи "не знаю". Не выдумывай. Respond ONLY with valid JSON: {"response": "...", "mood": "neutral"}.`
|
||||
// No self-introduction here: the shared persona block already says who she
|
||||
// is, and this line used to disagree with it — a different name ("Мавена")
|
||||
// and a masculine noun ("ассистент") in front of a feminine persona.
|
||||
return `Ответь кратко из своих знаний. Если не знаешь — скажи "не знаю". Не выдумывай. Respond ONLY with valid JSON: {"response": "...", "mood": "neutral"}.`
|
||||
}
|
||||
|
||||
@@ -11,10 +11,18 @@ func TestKnowledgePrompt(t *testing.T) {
|
||||
t.Fatal("KnowledgePrompt returned empty string")
|
||||
}
|
||||
// Must contain key instructions
|
||||
checks := []string{"Мавена", "не знаю", "не выдумывай"}
|
||||
checks := []string{"не знаю", "не выдумывай"}
|
||||
for _, c := range checks {
|
||||
if !strings.Contains(strings.ToLower(prompt), strings.ToLower(c)) {
|
||||
t.Errorf("KnowledgePrompt should mention %q", c)
|
||||
}
|
||||
}
|
||||
// Who she is comes from the shared persona block now. This prompt used to
|
||||
// say it too, with a different name and a masculine noun, which is the
|
||||
// drift the block exists to stop.
|
||||
for _, w := range []string{"Мавена", "ассистент"} {
|
||||
if strings.Contains(prompt, w) {
|
||||
t.Errorf("KnowledgePrompt should not introduce her (%q) — the persona block does", w)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -44,10 +44,21 @@ ws ::= [ \t\n]*
|
||||
// Changed again 31-07-2026: added the "unknown" escape hatch so the model can
|
||||
// admit it cannot route (Vikunja #359).
|
||||
//
|
||||
// Changed again 31-07-2026: added the clock/calendar rule (Vikunja #374). The
|
||||
// prompt never said which side "который час" or "какое число завтра" belong on,
|
||||
// so the model guessed — `system→query ×4` in every eval run. The rule sits
|
||||
// above the question test on purpose: these utterances all carry a question
|
||||
// word, so a later rule would never be reached. The boundary is what the
|
||||
// daemon can actually answer: only replySystem in cmd/mavend/voice.go owns the
|
||||
// clock and the calendar formatter, while the agenda ("что у меня завтра") is
|
||||
// answered inside the query branch, so that side stays query.
|
||||
//
|
||||
// The training workspace keeps its own copy of this prompt for relabelling, and
|
||||
// `llm/check_prompt_parity.py` there compares the two. That copy is in another
|
||||
// repo and was not touched, so parity will fail until it gets the same edits —
|
||||
// both the rule reorder and the "unknown" wording (Vikunja #362).
|
||||
// both the rule reorder and the "unknown" wording (Vikunja #362) — and now the
|
||||
// clock/calendar rule too. The training workspace is not checked out on this
|
||||
// box at all, so it could not be updated here; #362 still covers the catch-up.
|
||||
const routeSystem = `Классифицируй ровно одно сообщение пользователя. Верни ОДИН JSON-массив действий.
|
||||
|
||||
Ровно одно намерение: fact, reminder, note, query, act, chat, system.
|
||||
@@ -56,19 +67,21 @@ const routeSystem = `Классифицируй ровно одно сообще
|
||||
Классифицируй по цели пользователя. Порядок решения:
|
||||
1. Хочет напоминание в будущем → reminder
|
||||
2. Явно просит сохранить информацию → note
|
||||
3. Задаёт вопрос: есть вопросительное слово (сколько, что, какой, когда, где, кто, почему, как) или знак «?» → query
|
||||
4. Хочет получить информацию, в том числе о своих же данных → query
|
||||
5. Утверждает: сообщает или обновляет текущее состояние/событие → fact
|
||||
6. Просит выполнить работу → act
|
||||
7. Про ассистента, настройки или память → system
|
||||
8. Реплика — обрывок или указание на неназванное («это», «то», «потом»), и без него непонятно, что именно нужно сделать → unknown
|
||||
9. Иначе → chat
|
||||
3. Спрашивает только «который час» / «какое число» / «какой день недели» — сами часы или календарная дата, без своих данных → system
|
||||
4. Задаёт вопрос: есть вопросительное слово (сколько, что, какой, когда, где, кто, почему, как) или знак «?» → query
|
||||
5. Хочет получить информацию, в том числе о своих же данных → query
|
||||
6. Утверждает: сообщает или обновляет текущее состояние/событие → fact
|
||||
7. Просит выполнить работу → act
|
||||
8. Про ассистента, настройки или память → system
|
||||
9. Реплика — обрывок или указание на неназванное («это», «то», «потом»), и без него непонятно, что именно нужно сделать → unknown
|
||||
10. Иначе → chat
|
||||
|
||||
Различия:
|
||||
- note — сохранить информацию, без напоминания. text = суть.
|
||||
- reminder — уведомить позже. text = что напомнить.
|
||||
- fact — неявное обновление: пользователь сообщает, что что-то в мире изменилось (текущее/изменённое состояние, случившееся событие). key/value.
|
||||
- unknown — редкий случай. Ставь его, только если в самой реплике нет ни предмета, ни действия. Короткая, простая или незнакомая тема — это не причина для unknown: приветствие и болтовня — это chat, вопрос на любую тему — это query, просьба сделать что-то названное — это act.
|
||||
- system против query — часы и календарная дата сами по себе (сколько времени, какое число, какой день недели — можно и про завтра, и про другой город) — это system. А что записано в календаре или в памяти («что у меня завтра», «какие есть напоминания») — это query. Если в реплике есть просьба (напомни, запиши, сделай), то названное время — просто деталь просьбы, и это не system.
|
||||
- query против fact — решает форма реплики, а не тема. Вопрос о состоянии — это query, даже если названо то же самое, что бывает в fact. Только утверждение — это fact.
|
||||
|
||||
Примеры:
|
||||
@@ -82,6 +95,8 @@ const routeSystem = `Классифицируй ровно одно сообще
|
||||
"что такое docker?" → {"intent":"query","text":"что такое docker"}
|
||||
"напиши письмо" → {"intent":"act","verb":"написать письмо"}
|
||||
"очисти память" → {"intent":"system"}
|
||||
"который час?" → {"intent":"system"}
|
||||
"какое число завтра?" → {"intent":"system"}
|
||||
"привет" → {"intent":"chat","text":"привет"}
|
||||
"сделай это" → {"intent":"unknown"}
|
||||
"ну это" → {"intent":"unknown"}
|
||||
|
||||
@@ -33,6 +33,7 @@ const (
|
||||
type onnxEmbedder struct {
|
||||
tokenizer *unigramTokenizer
|
||||
session *ort.DynamicSession[int64, float32]
|
||||
id string
|
||||
}
|
||||
|
||||
func NewONNXEmbedder(modelPath, tokenizerPath, libPath string) (*onnxEmbedder, error) {
|
||||
@@ -58,11 +59,31 @@ func NewONNXEmbedder(modelPath, tokenizerPath, libPath string) (*onnxEmbedder, e
|
||||
return &onnxEmbedder{
|
||||
tokenizer: tok,
|
||||
session: session,
|
||||
id: modelIDFromPath(modelPath),
|
||||
}, nil
|
||||
}
|
||||
|
||||
func (e *onnxEmbedder) Dim() int { return embedDim }
|
||||
|
||||
// ID names the loaded model for the DB marker (Vikunja #378): the model file's
|
||||
// own name plus the dimension, so pointing the config at another model changes
|
||||
// the string on its own.
|
||||
func (e *onnxEmbedder) ID() string { return e.id }
|
||||
|
||||
// modelIDFromPath turns /opt/.../multilingual-e5-small.onnx into
|
||||
// "multilingual-e5-small@384".
|
||||
func modelIDFromPath(modelPath string) string {
|
||||
name := modelPath
|
||||
if i := strings.LastIndexAny(name, "/\\"); i >= 0 {
|
||||
name = name[i+1:]
|
||||
}
|
||||
name = strings.TrimSuffix(name, ".onnx")
|
||||
if name == "" {
|
||||
name = "onnx"
|
||||
}
|
||||
return fmt.Sprintf("%s@%d", name, embedDim)
|
||||
}
|
||||
|
||||
// Embed treats the text as a query. The classifier compares one short
|
||||
// utterance to another short seed phrase, so both sides get the same prefix
|
||||
// and the comparison stays fair. The recall path must call EmbedQuery and
|
||||
|
||||
@@ -449,16 +449,30 @@ func (AnaphoraResolver) Resolve(text string) (ref string, ok bool) {
|
||||
return "", false
|
||||
}
|
||||
|
||||
// ParseCalendarDate detects RU calendar date words in text and returns the
|
||||
// resolved time (midnight UTC+0 for "сегодня"/"today", next day for "завтра"/"tomorrow").
|
||||
// Returns zero time + false if no match.
|
||||
// ParseCalendarDate detects RU/EN calendar day words in text and returns
|
||||
// midnight of that day in now's own time zone. Handles "сегодня", "завтра",
|
||||
// "послезавтра", "вчера" (and the English words). Returns zero time + false
|
||||
// if no match.
|
||||
//
|
||||
// "послезавтра" is checked before "завтра" because it contains it.
|
||||
func ParseCalendarDate(text string, now time.Time) (time.Time, bool) {
|
||||
lower := strings.ToLower(text)
|
||||
if strings.Contains(lower, "сегодня") || strings.Contains(lower, "today") {
|
||||
return now.Truncate(24 * time.Hour), true
|
||||
}
|
||||
if strings.Contains(lower, "завтра") || strings.Contains(lower, "tomorrow") {
|
||||
return now.Truncate(24 * time.Hour).Add(24 * time.Hour), true
|
||||
switch {
|
||||
case strings.Contains(lower, "сегодня") || strings.Contains(lower, "today"):
|
||||
return midnight(now, 0), true
|
||||
case strings.Contains(lower, "послезавтра") || strings.Contains(lower, "day after tomorrow"):
|
||||
return midnight(now, 2), true
|
||||
case strings.Contains(lower, "завтра") || strings.Contains(lower, "tomorrow"):
|
||||
return midnight(now, 1), true
|
||||
case strings.Contains(lower, "вчера") || strings.Contains(lower, "yesterday"):
|
||||
return midnight(now, -1), true
|
||||
}
|
||||
return time.Time{}, false
|
||||
}
|
||||
|
||||
// midnight returns the start of the day that is `days` away from now, in
|
||||
// now's time zone (now.Truncate(24h) would cut on a UTC boundary instead).
|
||||
func midnight(now time.Time, days int) time.Time {
|
||||
y, m, d := now.AddDate(0, 0, days).Date()
|
||||
return time.Date(y, m, d, 0, 0, 0, 0, now.Location())
|
||||
}
|
||||
|
||||
@@ -45,6 +45,9 @@ func TestParseCalendarDate(t *testing.T) {
|
||||
{"расписание на завтра", time.Date(2026, 7, 7, 0, 0, 0, 0, time.UTC), true},
|
||||
{"what's today", time.Date(2026, 7, 6, 0, 0, 0, 0, time.UTC), true},
|
||||
{"tomorrow plans", time.Date(2026, 7, 7, 0, 0, 0, 0, time.UTC), true},
|
||||
{"какое число послезавтра", time.Date(2026, 7, 8, 0, 0, 0, 0, time.UTC), true},
|
||||
{"что было вчера", time.Date(2026, 7, 5, 0, 0, 0, 0, time.UTC), true},
|
||||
{"yesterday plans", time.Date(2026, 7, 5, 0, 0, 0, 0, time.UTC), true},
|
||||
{"какая погода", time.Time{}, false},
|
||||
{"сколько времени", time.Time{}, false},
|
||||
{"", time.Time{}, false},
|
||||
|
||||
@@ -0,0 +1,161 @@
|
||||
package store
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"time"
|
||||
)
|
||||
|
||||
// EmbedFunc embeds one piece of stored text. The caller passes
|
||||
// router.EmbedPassage — the STORE side of the query/passage asymmetry, which is
|
||||
// the side every vector in the DB was written with. (Passing the query side
|
||||
// would put the stored vectors in the wrong half of the space and quietly halve
|
||||
// recall.) A func instead of an interface keeps this package free of any
|
||||
// dependency on internal/router.
|
||||
type EmbedFunc func(ctx context.Context, text string) ([]float32, error)
|
||||
|
||||
// BackfillResult is what the re-embed run did, for logging.
|
||||
type BackfillResult struct {
|
||||
Skipped bool // marker already matched — nothing to do
|
||||
Notes int // rows rewritten in the notes table
|
||||
Facts int // fact rows rewritten in memory_vectors
|
||||
MemNotes int // note rows rewritten in memory_vectors
|
||||
NoText int // memory_vectors rows with no text in their meta, left alone
|
||||
Took time.Duration
|
||||
}
|
||||
|
||||
// ReembedAll rewrites every stored vector with the currently configured
|
||||
// embedder and then records that embedder as the one that owns the DB.
|
||||
//
|
||||
// Both places a vector lives are rewritten in the same pass: the `notes` table
|
||||
// `embedding` column and the `memory_vectors` rows (notes AND facts). Doing
|
||||
// only one would leave the two indexes disagreeing, which is worse than leaving
|
||||
// both stale.
|
||||
//
|
||||
// Safe to re-run: if the marker already names the current embedder there is
|
||||
// nothing to fix, so it returns immediately with Skipped set.
|
||||
//
|
||||
// Crash safety: everything — every vector and the marker — happens inside one
|
||||
// transaction. If anything fails or the process dies partway, the transaction
|
||||
// rolls back: no vectors changed and no marker written, so the next run does
|
||||
// the whole job again. The marker is never set unless the full rewrite
|
||||
// committed.
|
||||
func (s *Store) ReembedAll(ctx context.Context, currentID string, embed EmbedFunc) (BackfillResult, error) {
|
||||
start := time.Now()
|
||||
var res BackfillResult
|
||||
|
||||
stored, err := s.Meta(ctx, metaKeyEmbedderID)
|
||||
if err != nil {
|
||||
return res, err
|
||||
}
|
||||
if stored == currentID {
|
||||
res.Skipped = true
|
||||
res.Took = time.Since(start)
|
||||
return res, nil
|
||||
}
|
||||
|
||||
tx, err := s.db.BeginTx(ctx, nil)
|
||||
if err != nil {
|
||||
return res, fmt.Errorf("reembed: begin: %w", err)
|
||||
}
|
||||
defer tx.Rollback() // no-op once committed
|
||||
|
||||
// ----- notes table -----
|
||||
type noteRow struct {
|
||||
id int64
|
||||
text string
|
||||
}
|
||||
var notes []noteRow
|
||||
rows, err := tx.QueryContext(ctx, `SELECT id, text FROM notes WHERE text != ''`)
|
||||
if err != nil {
|
||||
return res, fmt.Errorf("reembed: read notes: %w", err)
|
||||
}
|
||||
for rows.Next() {
|
||||
var n noteRow
|
||||
if err := rows.Scan(&n.id, &n.text); err != nil {
|
||||
rows.Close()
|
||||
return res, fmt.Errorf("reembed: note row: %w", err)
|
||||
}
|
||||
notes = append(notes, n)
|
||||
}
|
||||
rows.Close()
|
||||
if err := rows.Err(); err != nil {
|
||||
return res, fmt.Errorf("reembed: notes: %w", err)
|
||||
}
|
||||
|
||||
for _, n := range notes {
|
||||
vec, err := embed(ctx, n.text)
|
||||
if err != nil {
|
||||
return res, fmt.Errorf("reembed: embed note %d: %w", n.id, err)
|
||||
}
|
||||
if _, err := tx.ExecContext(ctx,
|
||||
`UPDATE notes SET embedding = ? WHERE id = ?`, floatsToBlob(vec), n.id); err != nil {
|
||||
return res, fmt.Errorf("reembed: write note %d: %w", n.id, err)
|
||||
}
|
||||
res.Notes++
|
||||
}
|
||||
|
||||
// ----- memory_vectors (the unified index: notes AND facts) -----
|
||||
// The text to re-embed is the one carried in the row's meta blob, which is
|
||||
// exactly the text that was embedded when the row was written.
|
||||
type vecRow struct {
|
||||
id, text, kind string
|
||||
}
|
||||
var vecs []vecRow
|
||||
rows, err = tx.QueryContext(ctx, `SELECT id, meta FROM memory_vectors`)
|
||||
if err != nil {
|
||||
return res, fmt.Errorf("reembed: read memory vectors: %w", err)
|
||||
}
|
||||
for rows.Next() {
|
||||
var id, metaJSON string
|
||||
if err := rows.Scan(&id, &metaJSON); err != nil {
|
||||
rows.Close()
|
||||
return res, fmt.Errorf("reembed: memory row: %w", err)
|
||||
}
|
||||
meta := map[string]string{}
|
||||
if err := json.Unmarshal([]byte(metaJSON), &meta); err != nil {
|
||||
rows.Close()
|
||||
return res, fmt.Errorf("reembed: meta for %q: %w", id, err)
|
||||
}
|
||||
if meta["text"] == "" {
|
||||
res.NoText++
|
||||
continue
|
||||
}
|
||||
vecs = append(vecs, vecRow{id: id, text: meta["text"], kind: meta["type"]})
|
||||
}
|
||||
rows.Close()
|
||||
if err := rows.Err(); err != nil {
|
||||
return res, fmt.Errorf("reembed: memory vectors: %w", err)
|
||||
}
|
||||
|
||||
for _, v := range vecs {
|
||||
vec, err := embed(ctx, v.text)
|
||||
if err != nil {
|
||||
return res, fmt.Errorf("reembed: embed %q: %w", v.id, err)
|
||||
}
|
||||
if _, err := tx.ExecContext(ctx,
|
||||
`UPDATE memory_vectors SET vec = ? WHERE id = ?`, encodeVec(vec), v.id); err != nil {
|
||||
return res, fmt.Errorf("reembed: write %q: %w", v.id, err)
|
||||
}
|
||||
if v.kind == "fact" {
|
||||
res.Facts++
|
||||
} else {
|
||||
res.MemNotes++
|
||||
}
|
||||
}
|
||||
|
||||
// Same transaction as the rewrite, on purpose: the marker can only exist if
|
||||
// every vector above was written.
|
||||
if _, err := tx.ExecContext(ctx,
|
||||
`INSERT INTO meta (key, value) VALUES (?,?)
|
||||
ON CONFLICT(key) DO UPDATE SET value = excluded.value`,
|
||||
metaKeyEmbedderID, currentID); err != nil {
|
||||
return res, fmt.Errorf("reembed: write marker: %w", err)
|
||||
}
|
||||
if err := tx.Commit(); err != nil {
|
||||
return res, fmt.Errorf("reembed: commit: %w", err)
|
||||
}
|
||||
res.Took = time.Since(start)
|
||||
return res, nil
|
||||
}
|
||||
@@ -0,0 +1,171 @@
|
||||
package store
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// markerVec is a recognisable vector: nothing in these tests writes it except
|
||||
// the backfill, so finding it proves the row really was rewritten.
|
||||
var markerVec = []float32{9, 9, 9}
|
||||
|
||||
func newEmbedder(calls *int) EmbedFunc {
|
||||
return func(_ context.Context, _ string) ([]float32, error) {
|
||||
*calls++
|
||||
return markerVec, nil
|
||||
}
|
||||
}
|
||||
|
||||
// seedOldVectors puts one note (notes table + unified index) and one fact
|
||||
// (unified index only) in the DB, both carrying obviously-old vectors.
|
||||
func seedOldVectors(t *testing.T, s *Store) {
|
||||
t.Helper()
|
||||
ctx := context.Background()
|
||||
old := []float32{0.1, 0.2, 0.3}
|
||||
id, err := s.WriteNote(ctx, time.Now(), "молоко в холодильнике", old, "voice")
|
||||
if err != nil {
|
||||
t.Fatalf("WriteNote: %v", err)
|
||||
}
|
||||
mem := s.VectorMemory()
|
||||
if err := mem.Insert(ctx, "note:1", old, map[string]string{
|
||||
"type": "note", "text": "молоко в холодильнике",
|
||||
}); err != nil {
|
||||
t.Fatalf("Insert note vector: %v", err)
|
||||
}
|
||||
if err := mem.Insert(ctx, "fact:water:1", old, map[string]string{
|
||||
"type": "fact", "text": "я пил воду",
|
||||
}); err != nil {
|
||||
t.Fatalf("Insert fact vector: %v", err)
|
||||
}
|
||||
_ = id
|
||||
}
|
||||
|
||||
func noteVec(t *testing.T, s *Store) []float32 {
|
||||
t.Helper()
|
||||
var blob []byte
|
||||
if err := s.db.QueryRow(`SELECT embedding FROM notes LIMIT 1`).Scan(&blob); err != nil {
|
||||
t.Fatalf("read note embedding: %v", err)
|
||||
}
|
||||
return blobToFloats(blob)
|
||||
}
|
||||
|
||||
func memVec(t *testing.T, s *Store, id string) []float32 {
|
||||
t.Helper()
|
||||
var blob []byte
|
||||
if err := s.db.QueryRow(`SELECT vec FROM memory_vectors WHERE id = ?`, id).Scan(&blob); err != nil {
|
||||
t.Fatalf("read memory vector %s: %v", id, err)
|
||||
}
|
||||
return decodeVec(blob)
|
||||
}
|
||||
|
||||
func sameVec(a, b []float32) bool {
|
||||
if len(a) != len(b) {
|
||||
return false
|
||||
}
|
||||
for i := range a {
|
||||
if a[i] != b[i] {
|
||||
return false
|
||||
}
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
// The deployed case: old vectors everywhere, no marker. Every vector in both
|
||||
// places must be rewritten and the marker recorded.
|
||||
func TestReembedAllRewritesEveryVector(t *testing.T) {
|
||||
s := newTestStore(t)
|
||||
ctx := context.Background()
|
||||
seedOldVectors(t, s)
|
||||
|
||||
calls := 0
|
||||
res, err := s.ReembedAll(ctx, "multilingual-e5-small@384", newEmbedder(&calls))
|
||||
if err != nil {
|
||||
t.Fatalf("ReembedAll: %v", err)
|
||||
}
|
||||
if res.Skipped {
|
||||
t.Fatal("first run should not skip")
|
||||
}
|
||||
if res.Notes != 1 || res.MemNotes != 1 || res.Facts != 1 {
|
||||
t.Fatalf("counts: notes=%d memNotes=%d facts=%d", res.Notes, res.MemNotes, res.Facts)
|
||||
}
|
||||
if calls != 3 {
|
||||
t.Fatalf("embedder called %d times, want 3", calls)
|
||||
}
|
||||
if !sameVec(noteVec(t, s), markerVec) {
|
||||
t.Fatalf("notes table not rewritten: %v", noteVec(t, s))
|
||||
}
|
||||
if !sameVec(memVec(t, s, "note:1"), markerVec) {
|
||||
t.Fatal("unified index note row not rewritten")
|
||||
}
|
||||
if !sameVec(memVec(t, s, "fact:water:1"), markerVec) {
|
||||
t.Fatal("unified index fact row not rewritten")
|
||||
}
|
||||
got, err := s.Meta(ctx, metaKeyEmbedderID)
|
||||
if err != nil {
|
||||
t.Fatalf("Meta: %v", err)
|
||||
}
|
||||
if got != "multilingual-e5-small@384" {
|
||||
t.Fatalf("marker = %q", got)
|
||||
}
|
||||
}
|
||||
|
||||
// Re-running must do nothing at all — not a second pass over the same rows.
|
||||
func TestReembedAllSecondRunIsNoop(t *testing.T) {
|
||||
s := newTestStore(t)
|
||||
ctx := context.Background()
|
||||
seedOldVectors(t, s)
|
||||
|
||||
calls := 0
|
||||
if _, err := s.ReembedAll(ctx, "e5@384", newEmbedder(&calls)); err != nil {
|
||||
t.Fatalf("first run: %v", err)
|
||||
}
|
||||
first := calls
|
||||
|
||||
res, err := s.ReembedAll(ctx, "e5@384", newEmbedder(&calls))
|
||||
if err != nil {
|
||||
t.Fatalf("second run: %v", err)
|
||||
}
|
||||
if !res.Skipped {
|
||||
t.Fatal("second run should report Skipped")
|
||||
}
|
||||
if calls != first {
|
||||
t.Fatalf("second run embedded %d more rows, want 0", calls-first)
|
||||
}
|
||||
}
|
||||
|
||||
// A failure partway must leave the DB exactly as it was: no marker, and the old
|
||||
// vectors still in place (one transaction, rolled back).
|
||||
func TestReembedAllPartialFailureLeavesMarkerUnset(t *testing.T) {
|
||||
s := newTestStore(t)
|
||||
ctx := context.Background()
|
||||
seedOldVectors(t, s)
|
||||
before := noteVec(t, s)
|
||||
|
||||
calls := 0
|
||||
boom := func(_ context.Context, _ string) ([]float32, error) {
|
||||
calls++
|
||||
if calls == 2 {
|
||||
return nil, errors.New("onnx blew up")
|
||||
}
|
||||
return markerVec, nil
|
||||
}
|
||||
if _, err := s.ReembedAll(ctx, "e5@384", boom); err == nil {
|
||||
t.Fatal("expected an error")
|
||||
}
|
||||
got, err := s.Meta(ctx, metaKeyEmbedderID)
|
||||
if err != nil {
|
||||
t.Fatalf("Meta: %v", err)
|
||||
}
|
||||
if got != "" {
|
||||
t.Fatalf("marker was set to %q after a failed run", got)
|
||||
}
|
||||
if !sameVec(noteVec(t, s), before) {
|
||||
t.Fatal("a failed run left a partially rewritten notes table")
|
||||
}
|
||||
// And the mismatch warning must still fire, so the user knows to re-run.
|
||||
if _, mismatch, err := s.CheckEmbedder(ctx, "e5@384"); err != nil || !mismatch {
|
||||
t.Fatalf("CheckEmbedder after failed backfill: mismatch=%v err=%v", mismatch, err)
|
||||
}
|
||||
}
|
||||
@@ -13,11 +13,15 @@ import (
|
||||
// unknown = a pending row found stale at startup: the process that started it
|
||||
// is gone, and the send may or may not have reached the external channel.
|
||||
// Never auto-resolved into sent or failed — that would be guessing.
|
||||
// dropped = the routing table deliberately suppressed this one (a care nudge
|
||||
// while you're away). Nothing was sent and nothing went wrong; the row exists
|
||||
// so "she dropped it" and "the rule never fired" don't look the same later.
|
||||
const (
|
||||
DeliveryPending = "pending"
|
||||
DeliverySent = "sent"
|
||||
DeliveryFailed = "failed"
|
||||
DeliveryUnknown = "unknown"
|
||||
DeliveryDropped = "dropped"
|
||||
)
|
||||
|
||||
// BeginDeliveryAttempt durably records intent to send BEFORE the external
|
||||
@@ -43,10 +47,11 @@ func (s *Store) BeginDeliveryAttempt(ctx context.Context, kind, rule string, rem
|
||||
}
|
||||
|
||||
// CompleteDeliveryAttempt records the sink's outcome for a prior
|
||||
// BeginDeliveryAttempt. status is "sent" or "failed" — never "pending" or
|
||||
// "unknown" (those are set only by Begin and reconciliation respectively).
|
||||
// BeginDeliveryAttempt. status is "sent", "failed" or "dropped" — never
|
||||
// "pending" or "unknown" (those are set only by Begin and reconciliation
|
||||
// respectively).
|
||||
func (s *Store) CompleteDeliveryAttempt(ctx context.Context, id int64, status string, now time.Time) error {
|
||||
if status != DeliverySent && status != DeliveryFailed {
|
||||
if status != DeliverySent && status != DeliveryFailed && status != DeliveryDropped {
|
||||
return fmt.Errorf("store: invalid delivery completion status %q", status)
|
||||
}
|
||||
_, err := s.db.ExecContext(ctx,
|
||||
|
||||
@@ -0,0 +1,34 @@
|
||||
package store
|
||||
|
||||
import (
|
||||
"context"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// TestDroppedDeliveryAttemptRoundTrips — Vikunja #370. A suppressed nudge is
|
||||
// recorded as 'dropped'. The status column has a CHECK constraint, so this
|
||||
// only works if migration #12 widened it; a fake outbox in a unit test would
|
||||
// not catch that.
|
||||
func TestDroppedDeliveryAttemptRoundTrips(t *testing.T) {
|
||||
s := newTestStore(t)
|
||||
ctx := context.Background()
|
||||
now := time.Now()
|
||||
|
||||
id, err := s.BeginDeliveryAttempt(ctx, "nudge", "water", 0, "drop", "abc123", now)
|
||||
if err != nil {
|
||||
t.Fatalf("BeginDeliveryAttempt: %v", err)
|
||||
}
|
||||
if err := s.CompleteDeliveryAttempt(ctx, id, DeliveryDropped, now); err != nil {
|
||||
t.Fatalf("CompleteDeliveryAttempt: %v", err)
|
||||
}
|
||||
|
||||
var status string
|
||||
err = s.db.QueryRowContext(ctx, `SELECT status FROM delivery_attempts WHERE id = ?`, id).Scan(&status)
|
||||
if err != nil {
|
||||
t.Fatalf("read back: %v", err)
|
||||
}
|
||||
if status != DeliveryDropped {
|
||||
t.Fatalf("status: want %q, got %q", DeliveryDropped, status)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,74 @@
|
||||
package store
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"time"
|
||||
)
|
||||
|
||||
// DialogueSessionRow — one saved follow-up session. Data is the session
|
||||
// encoded by the dialogue package; the store does not look inside it.
|
||||
type DialogueSessionRow struct {
|
||||
ID string
|
||||
Data []byte
|
||||
Ts time.Time
|
||||
TTL time.Duration
|
||||
Expires time.Time
|
||||
}
|
||||
|
||||
// SaveDialogueSession — write (or replace) the session for one dialogue id.
|
||||
// One row per id: a newer turn overwrites the older state.
|
||||
func (s *Store) SaveDialogueSession(ctx context.Context, id string, data []byte, ts time.Time, ttl time.Duration) error {
|
||||
expires := ts.Add(ttl)
|
||||
_, err := s.db.ExecContext(ctx, `
|
||||
INSERT INTO dialogue_sessions (id, data, ts, ttl_ms, expires_ts) VALUES (?, ?, ?, ?, ?)
|
||||
ON CONFLICT(id) DO UPDATE SET data = excluded.data,
|
||||
ts = excluded.ts,
|
||||
ttl_ms = excluded.ttl_ms,
|
||||
expires_ts = excluded.expires_ts`,
|
||||
id, data, ts.UnixMilli(), ttl.Milliseconds(), expires.UnixMilli())
|
||||
if err != nil {
|
||||
return fmt.Errorf("save dialogue session: %w", err)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// DeleteDialogueSession — drop one session (ended, or expired).
|
||||
func (s *Store) DeleteDialogueSession(ctx context.Context, id string) error {
|
||||
if _, err := s.db.ExecContext(ctx, `DELETE FROM dialogue_sessions WHERE id = ?`, id); err != nil {
|
||||
return fmt.Errorf("delete dialogue session: %w", err)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// LoadDialogueSessions — return the sessions still alive at `now` and delete
|
||||
// the ones that already ran out. An expired session is dead: it never comes
|
||||
// back after a restart.
|
||||
func (s *Store) LoadDialogueSessions(ctx context.Context, now time.Time) ([]DialogueSessionRow, error) {
|
||||
if _, err := s.db.ExecContext(ctx,
|
||||
`DELETE FROM dialogue_sessions WHERE expires_ts <= ?`, now.UnixMilli()); err != nil {
|
||||
return nil, fmt.Errorf("prune dialogue sessions: %w", err)
|
||||
}
|
||||
rows, err := s.db.QueryContext(ctx,
|
||||
`SELECT id, data, ts, ttl_ms, expires_ts FROM dialogue_sessions ORDER BY id`)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("load dialogue sessions: %w", err)
|
||||
}
|
||||
defer rows.Close()
|
||||
var out []DialogueSessionRow
|
||||
for rows.Next() {
|
||||
var r DialogueSessionRow
|
||||
var tsMilli, ttlMilli, expMilli int64
|
||||
if err := rows.Scan(&r.ID, &r.Data, &tsMilli, &ttlMilli, &expMilli); err != nil {
|
||||
return nil, fmt.Errorf("scan dialogue session: %w", err)
|
||||
}
|
||||
r.Ts = time.UnixMilli(tsMilli).UTC()
|
||||
r.TTL = time.Duration(ttlMilli) * time.Millisecond
|
||||
r.Expires = time.UnixMilli(expMilli).UTC()
|
||||
out = append(out, r)
|
||||
}
|
||||
if err := rows.Err(); err != nil {
|
||||
return nil, fmt.Errorf("load dialogue sessions: %w", err)
|
||||
}
|
||||
return out, nil
|
||||
}
|
||||
@@ -0,0 +1,97 @@
|
||||
package store
|
||||
|
||||
import (
|
||||
"context"
|
||||
"database/sql"
|
||||
"errors"
|
||||
"fmt"
|
||||
)
|
||||
|
||||
// metaKeyEmbedderID names the embedder that wrote the stored vectors.
|
||||
//
|
||||
// Why one value for the whole DB and not a column on every vector row: the
|
||||
// vectors are only ever rewritten all at once (one backfill re-embeds every
|
||||
// note and fact together), so a per-row marker would hold the same string in
|
||||
// every row and cost a column on two tables for nothing.
|
||||
const metaKeyEmbedderID = "embedder_id"
|
||||
|
||||
// Meta reads a single value from the meta table. Missing key ⇒ empty string.
|
||||
func (s *Store) Meta(ctx context.Context, key string) (string, error) {
|
||||
var v string
|
||||
err := s.db.QueryRowContext(ctx, `SELECT value FROM meta WHERE key = ?`, key).Scan(&v)
|
||||
if errors.Is(err, sql.ErrNoRows) {
|
||||
return "", nil
|
||||
}
|
||||
if err != nil {
|
||||
return "", fmt.Errorf("read meta %s: %w", key, err)
|
||||
}
|
||||
return v, nil
|
||||
}
|
||||
|
||||
// SetMeta writes (or overwrites) a single meta value.
|
||||
func (s *Store) SetMeta(ctx context.Context, key, value string) error {
|
||||
_, err := s.db.ExecContext(ctx,
|
||||
`INSERT INTO meta (key, value) VALUES (?,?)
|
||||
ON CONFLICT(key) DO UPDATE SET value = excluded.value`, key, value)
|
||||
if err != nil {
|
||||
return fmt.Errorf("write meta %s: %w", key, err)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// EmbedderUnknown is the stored id reported for a DB that already holds
|
||||
// vectors but never recorded who wrote them.
|
||||
const EmbedderUnknown = "unknown (written before this marker existed)"
|
||||
|
||||
// CheckEmbedder compares the embedder now configured against the one that
|
||||
// wrote the stored vectors. Returns the stored id and whether it differs.
|
||||
//
|
||||
// Vectors from two different models live in different spaces, so cosine
|
||||
// between them is noise rather than a low score — and both of our models are
|
||||
// 384-dimensional, so nothing else catches it.
|
||||
//
|
||||
// Three cases, and the middle one is the one that actually matters:
|
||||
//
|
||||
// - marker present ⇒ compare the two ids.
|
||||
// - marker absent but vectors already stored ⇒ this is a DB from before the
|
||||
// marker, so we cannot know who wrote them. Report a mismatch. This is the
|
||||
// real case on the deployed box: those vectors came from the old embedder,
|
||||
// and claiming them for the current one would hide the exact problem the
|
||||
// marker was added to catch.
|
||||
// - marker absent and no vectors ⇒ fresh DB, claim it, nothing to fix.
|
||||
//
|
||||
// On a mismatch the fix is ReembedAll (backfill.go), run explicitly with
|
||||
// `mavend -reembed`. Nothing is re-embedded here: that work is minutes of CPU
|
||||
// on the laptop and must not stall a normal start.
|
||||
func (s *Store) CheckEmbedder(ctx context.Context, currentID string) (stored string, mismatch bool, err error) {
|
||||
stored, err = s.Meta(ctx, metaKeyEmbedderID)
|
||||
if err != nil {
|
||||
return "", false, err
|
||||
}
|
||||
if stored != "" {
|
||||
return stored, stored != currentID, nil
|
||||
}
|
||||
n, err := s.countVectors(ctx)
|
||||
if err != nil {
|
||||
return "", false, err
|
||||
}
|
||||
if n > 0 {
|
||||
return EmbedderUnknown, true, nil
|
||||
}
|
||||
return currentID, false, s.SetMeta(ctx, metaKeyEmbedderID, currentID)
|
||||
}
|
||||
|
||||
// countVectors — how many stored rows carry an embedding. Used only to tell a
|
||||
// fresh DB apart from one that predates the marker.
|
||||
func (s *Store) countVectors(ctx context.Context) (int, error) {
|
||||
var notes, vecs int
|
||||
if err := s.db.QueryRowContext(ctx,
|
||||
`SELECT count(*) FROM notes WHERE embedding IS NOT NULL`).Scan(¬es); err != nil {
|
||||
return 0, fmt.Errorf("count note vectors: %w", err)
|
||||
}
|
||||
if err := s.db.QueryRowContext(ctx,
|
||||
`SELECT count(*) FROM memory_vectors`).Scan(&vecs); err != nil {
|
||||
return 0, fmt.Errorf("count memory vectors: %w", err)
|
||||
}
|
||||
return notes + vecs, nil
|
||||
}
|
||||
@@ -0,0 +1,100 @@
|
||||
package store
|
||||
|
||||
import (
|
||||
"context"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// A fresh DB has no marker yet, so the current embedder is recorded and
|
||||
// nothing is flagged.
|
||||
func TestCheckEmbedderFreshDBRecords(t *testing.T) {
|
||||
s := newTestStore(t)
|
||||
ctx := context.Background()
|
||||
|
||||
stored, mismatch, err := s.CheckEmbedder(ctx, "multilingual-e5-small@384")
|
||||
if err != nil {
|
||||
t.Fatalf("CheckEmbedder: %v", err)
|
||||
}
|
||||
if mismatch {
|
||||
t.Fatal("fresh DB reported a mismatch")
|
||||
}
|
||||
if stored != "multilingual-e5-small@384" {
|
||||
t.Fatalf("stored = %q", stored)
|
||||
}
|
||||
got, err := s.Meta(ctx, metaKeyEmbedderID)
|
||||
if err != nil {
|
||||
t.Fatalf("Meta: %v", err)
|
||||
}
|
||||
if got != "multilingual-e5-small@384" {
|
||||
t.Fatalf("marker not persisted, got %q", got)
|
||||
}
|
||||
}
|
||||
|
||||
// The deployed box: notes were written by the old embedder, before the marker
|
||||
// existed. Claiming them for the current one would hide exactly the problem
|
||||
// the marker is for, so an unmarked DB that already holds vectors is a
|
||||
// mismatch.
|
||||
func TestCheckEmbedderUnmarkedDBWithVectorsIsMismatch(t *testing.T) {
|
||||
s := newTestStore(t)
|
||||
ctx := context.Background()
|
||||
|
||||
if _, err := s.WriteNote(ctx, time.Now(), "молоко в холодильнике", []float32{0.1, 0.2}, "voice"); err != nil {
|
||||
t.Fatalf("WriteNote: %v", err)
|
||||
}
|
||||
|
||||
stored, mismatch, err := s.CheckEmbedder(ctx, "multilingual-e5-small@384")
|
||||
if err != nil {
|
||||
t.Fatalf("CheckEmbedder: %v", err)
|
||||
}
|
||||
if !mismatch {
|
||||
t.Fatal("an unmarked DB with stored vectors should report a mismatch")
|
||||
}
|
||||
if stored != EmbedderUnknown {
|
||||
t.Fatalf("stored = %q, want %q", stored, EmbedderUnknown)
|
||||
}
|
||||
// It must NOT claim the DB — that would silence the warning on restart.
|
||||
got, err := s.Meta(ctx, metaKeyEmbedderID)
|
||||
if err != nil {
|
||||
t.Fatalf("Meta: %v", err)
|
||||
}
|
||||
if got != "" {
|
||||
t.Fatalf("marker written despite unknown provenance: %q", got)
|
||||
}
|
||||
}
|
||||
|
||||
// Both models are 384-dim, so this is the only thing that catches the swap.
|
||||
func TestCheckEmbedderDifferentModelMismatch(t *testing.T) {
|
||||
s := newTestStore(t)
|
||||
ctx := context.Background()
|
||||
|
||||
if err := s.SetMeta(ctx, metaKeyEmbedderID, "paraphrase-multilingual-MiniLM-L12-v2@384"); err != nil {
|
||||
t.Fatalf("SetMeta: %v", err)
|
||||
}
|
||||
stored, mismatch, err := s.CheckEmbedder(ctx, "multilingual-e5-small@384")
|
||||
if err != nil {
|
||||
t.Fatalf("CheckEmbedder: %v", err)
|
||||
}
|
||||
if !mismatch {
|
||||
t.Fatal("different embedder not detected")
|
||||
}
|
||||
if stored != "paraphrase-multilingual-MiniLM-L12-v2@384" {
|
||||
t.Fatalf("stored = %q", stored)
|
||||
}
|
||||
}
|
||||
|
||||
// The same embedder must never raise a false alarm, including on re-check.
|
||||
func TestCheckEmbedderSameModelNoAlarm(t *testing.T) {
|
||||
s := newTestStore(t)
|
||||
ctx := context.Background()
|
||||
|
||||
for i := 0; i < 2; i++ {
|
||||
_, mismatch, err := s.CheckEmbedder(ctx, "multilingual-e5-small@384")
|
||||
if err != nil {
|
||||
t.Fatalf("CheckEmbedder: %v", err)
|
||||
}
|
||||
if mismatch {
|
||||
t.Fatalf("false alarm on pass %d", i)
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -74,6 +74,44 @@ ALTER TABLE reminders ADD COLUMN next_fire_ts INTEGER;`, // #2
|
||||
`CREATE INDEX IF NOT EXISTS idx_nudges_snoozed ON nudges (outcome_ts) WHERE outcome = 'snoozed';`, // #8 — SnoozedUntil runs every tick; keep it off a full scan (Vikunja #364)
|
||||
`ALTER TABLE proposed_routines ADD COLUMN accepted_ts INTEGER;
|
||||
ALTER TABLE proposed_routines ADD COLUMN last_fired_ts INTEGER;`, // #9 — accepted routines keep firing (Vikunja #366): the tick loop needs to know when a routine was accepted and when it last nudged
|
||||
|
||||
`CREATE TABLE IF NOT EXISTS dialogue_sessions (
|
||||
id TEXT PRIMARY KEY,
|
||||
data BLOB NOT NULL,
|
||||
ts INTEGER NOT NULL,
|
||||
ttl_ms INTEGER NOT NULL,
|
||||
expires_ts INTEGER NOT NULL
|
||||
);
|
||||
CREATE INDEX IF NOT EXISTS idx_dialogue_sessions_expires ON dialogue_sessions (expires_ts);`, // #10 — the follow-up session survives a restart (Vikunja #363); small, TTL-pruned table, not a history log
|
||||
|
||||
`CREATE TABLE IF NOT EXISTS meta (
|
||||
key TEXT PRIMARY KEY,
|
||||
value TEXT NOT NULL
|
||||
);`, // #11 — small key/value table for facts about the DB itself; first key is embedder_id (Vikunja #378)
|
||||
|
||||
// #12 — a suppressed nudge gets a 'dropped' row (Vikunja #370). sqlite
|
||||
// can't widen a CHECK constraint in place, so the table is rebuilt; the
|
||||
// index goes with the old table and is recreated. The columns are listed
|
||||
// out rather than `SELECT *` — copying by position would silently shuffle
|
||||
// every row if the old table's column order ever differed from this one.
|
||||
`CREATE TABLE delivery_attempts_v12 (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
kind TEXT NOT NULL CHECK (kind IN ('nudge','reminder')),
|
||||
rule TEXT NOT NULL DEFAULT '',
|
||||
reminder_id INTEGER NOT NULL DEFAULT 0,
|
||||
channel TEXT NOT NULL,
|
||||
body_hash TEXT NOT NULL,
|
||||
status TEXT NOT NULL DEFAULT 'pending' CHECK (status IN ('pending','sent','failed','unknown','dropped')),
|
||||
created_ts INTEGER NOT NULL,
|
||||
completed_ts INTEGER
|
||||
);
|
||||
INSERT INTO delivery_attempts_v12
|
||||
(id, kind, rule, reminder_id, channel, body_hash, status, created_ts, completed_ts)
|
||||
SELECT id, kind, rule, reminder_id, channel, body_hash, status, created_ts, completed_ts
|
||||
FROM delivery_attempts;
|
||||
DROP TABLE delivery_attempts;
|
||||
ALTER TABLE delivery_attempts_v12 RENAME TO delivery_attempts;
|
||||
CREATE INDEX IF NOT EXISTS idx_delivery_attempts_status ON delivery_attempts (status);`,
|
||||
}
|
||||
|
||||
// migrate applies every migration with a number greater than the DB's current
|
||||
|
||||
+79
-4
@@ -1,8 +1,10 @@
|
||||
#!/usr/bin/env bash
|
||||
# Unified script to stop all Maven services.
|
||||
# Usage: ./kill-maven.sh
|
||||
# - Graceful SIGTERM is attempted first.
|
||||
# - If any process lingers, force with SIGKILL.
|
||||
# - Docker deploy: `docker compose stop` (see why below).
|
||||
# - Bare-metal / dev run: graceful SIGTERM first, SIGKILL if anything lingers.
|
||||
# Exits non-zero if it cannot confirm everything is stopped. It must never say
|
||||
# "stopped" unless it checked.
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
@@ -24,6 +26,73 @@ else
|
||||
LLM='llama-server.*\.gguf'
|
||||
fi
|
||||
|
||||
COMPOSE_FILE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/docker-compose.yml"
|
||||
|
||||
# --- containerised deploy ------------------------------------------------
|
||||
# docker-compose.yml does not set `pid: host`, so each container has its own
|
||||
# PID namespace: pkill on the host sees nothing inside them. This script used
|
||||
# to print "all stopped" while every daemon was still happily running. Stop the
|
||||
# containers through compose instead — that actually reaches them.
|
||||
#
|
||||
# running_containers prints the ids of the project's running containers, or
|
||||
# nothing. Empty output plus a non-zero return means "could not ask docker",
|
||||
# which is different from "nothing is running" and is handled below.
|
||||
running_containers() {
|
||||
docker compose -f "$COMPOSE_FILE" ps -q --status running 2>/dev/null
|
||||
}
|
||||
|
||||
DOCKER_OK=0
|
||||
CONTAINERS=""
|
||||
if command -v docker >/dev/null 2>&1 && [ -f "$COMPOSE_FILE" ]; then
|
||||
if CONTAINERS="$(running_containers)"; then
|
||||
DOCKER_OK=1
|
||||
fi
|
||||
fi
|
||||
|
||||
if [ "$DOCKER_OK" = 1 ] && [ -n "$CONTAINERS" ]; then
|
||||
echo "--- Maven is running in containers: stopping via docker compose ---"
|
||||
if ! docker compose -f "$COMPOSE_FILE" stop; then
|
||||
echo "ERROR: 'docker compose stop' failed. Containers may still be running." >&2
|
||||
exit 1
|
||||
fi
|
||||
echo "--- Verifying containers are gone ---"
|
||||
LEFT="$(running_containers || true)"
|
||||
if [ -n "$LEFT" ]; then
|
||||
echo "ERROR: containers still running after stop:" >&2
|
||||
docker compose -f "$COMPOSE_FILE" ps >&2 || true
|
||||
exit 1
|
||||
fi
|
||||
echo "All containers stopped."
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# --- bare-metal / dev run -----------------------------------------------
|
||||
# pgrep -f matches whole command lines, so a shell that merely mentions
|
||||
# "mavend" (this script's own parent, for one) shows up. Drop ourselves and our
|
||||
# parent, otherwise the SIGKILL sweep can take out the terminal you ran this in.
|
||||
host_pids() {
|
||||
pgrep -f "$PAT|$LLM" | grep -v -e "^$$\$" -e "^$PPID\$" | paste -sd, - || true
|
||||
}
|
||||
HOST_PIDS=$(host_pids)
|
||||
|
||||
if [ -z "$HOST_PIDS" ]; then
|
||||
# Nothing on the host. Whether that means "already down" depends on whether
|
||||
# we managed to ask docker, and the two must not read the same.
|
||||
if [ "$DOCKER_OK" = 1 ]; then
|
||||
# Docker answered and named no running containers, and there is nothing
|
||||
# on the host either. That is a real answer: Maven is already stopped.
|
||||
echo "Nothing to stop: no Maven processes and no running containers."
|
||||
exit 0
|
||||
fi
|
||||
# We could not ask docker, so Maven may be alive in a container we cannot
|
||||
# see. Saying "stopped" here is the exact false success this script had.
|
||||
echo "ERROR: no Maven processes on this host, and docker could not be asked." >&2
|
||||
echo " If this is the container deploy it may still be running:" >&2
|
||||
echo " docker compose -f $COMPOSE_FILE stop" >&2
|
||||
echo " Nothing was stopped. Check by hand before assuming Maven is down." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "--- Sending graceful SIGTERM to Maven services ---"
|
||||
pkill -TERM -f "$PAT" || true
|
||||
# mavend's Pdeathsig SIGKILLs its llama-server on exit, but sweep strays too
|
||||
@@ -32,12 +101,18 @@ pkill -TERM -f "$LLM" || true
|
||||
|
||||
echo "--- Verifying processes are gone ---"
|
||||
sleep 1
|
||||
PIDS=$(pgrep -d ',' -f "$PAT|$LLM") || PIDS=""
|
||||
PIDS=$(host_pids)
|
||||
if [ -n "$PIDS" ]; then
|
||||
echo "Warning: some processes still alive. PIDs: $PIDS"
|
||||
echo "--- Force killing with SIGKILL ---"
|
||||
echo "$PIDS" | tr ',' '\n' | xargs -r kill -9
|
||||
sleep 1
|
||||
LEFT=$(host_pids)
|
||||
if [ -n "$LEFT" ]; then
|
||||
echo "ERROR: still alive after SIGKILL. PIDs: $LEFT" >&2
|
||||
exit 1
|
||||
fi
|
||||
echo "Done (SIGKILL)."
|
||||
else
|
||||
echo "All services gracefully stopped."
|
||||
fi
|
||||
fi
|
||||
|
||||
Reference in New Issue
Block a user