Compare commits

..

12 Commits

Author SHA1 Message Date
kami 533f0acda8 Lead the bake-off with the answer, not the superseded one
The file ran two sweeps and the second one changed the resident model, but
the lede still opened with "Recommendation: keep Qwen3.5-0.8B". Anyone
landing on the file read the wrong conclusion and had to scroll 100 lines
to find that it had been replaced — and it contradicted CLAUDE.md, which
already says the resident model is Qwen3-1.7B.

Both sweeps are accurate, so nothing is rewritten. The lede now states the
outcome and the first sweep's verdict is scoped to what it actually tested:
it rejects LFM2.5-1.2B, which still holds. It never was a case for keeping
0.8B as the resident model.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 21:17:07 +04:00
kami 4f59ba78c6 Write down the five-model sweep and why the 1.7B won
Numbers behind the resident-model change, plus the answer to "could a 230-350M
model do this instead" — no, and the reason is worth keeping: LFM2.5's published
instruction-following scores beat Qwen3.5-0.8B, and every one of those benchmarks
except Multi-IF is English. In Russian the 350M invents non-words and the 230M
answers in Spanish.

Also fills the row TALK-EVAL-31-07-2026.md had to void for contamination, and
corrects a wrong call I nearly made: the 1.7B's 16s p95 looked like the reasoning
trace, but the 0.8B sits at 17s in every run and the 1.7B beat it twice out of
three. The long tail is shared and is not the Thinking block.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 19:08:37 +04:00
kami d0afd9d4f6 Make Qwen3-1.7B the resident model
Stock Qwen3-1.7B, not the CPT'd one — that training is still running. It won
on both fixtures we have, measured tonight on an otherwise idle box:

  routing, 77 RU cases, intent-only:  67.5%  vs  59.7%  for Qwen3.5-0.8B
  talk fixture, 27 cases:             20/27  vs  11-17/27

It also beat Qwen3.5-2B, which is 20% larger, on every routing column.

Two other things came with it:

n_ctx goes 2048 -> 4096. This is a Thinking variant, so reasoning tokens need
the room, and 4096 is the context every score above was measured at. Shipping
2048 would ship something nobody measured.

The doc now says not to bother with sub-500M models, because I checked and they
are not close. LFM2.5-350M routes at 5.2% — worse than guessing among 7 intents
— and answers "столица Франции?" with "Сторзит", which is not a word. The 230M
replies to Russian in Spanish. Their published IFEval and BFCL numbers are good
and they are all English.

Note the routing gain needs the LLM router actually wired on to show up. It is
still nil, so this commit buys the phrasing improvement today and the routing
improvement when that lands.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:58:20 +04:00
kami 6b67e6f3c2 Word nudges from templates by default, model optional
DEPRECATION, flagged not asked: LLM-phrased nudges are no longer the default.
LLMPhraser.PhraseNudge now returns a hand-written Russian template. The model
still phrases chat, queries and reminders — only nudges moved.

Why: measured over many runs, Qwen3.5-0.8B wrote formal "вы" and plural
imperatives, used masculine self-reference, and invented facts and units
(90-95 seconds to boil an egg). A nudge is five words of known content, so
generation buys nothing and risks the persona every time. Templates score
15/15 on the nudge fixture, the model 11-13/15.

Nothing is deleted: the prompt, the fallbacks and the whole LLM nudge path
stay. Set phraser.llm_nudges = true in deploy/mavend.json to get them back.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:21:52 +04:00
kami 13e5170e9e Hand-written Russian nudge templates plus a picker
Nudge wording as data instead of generation. The wording lives in
internal/phraser/nudges_ru_v1.json (embedded), about 10 variants per rule:
water, meal, break, service_down, netdata_critical, routine:, morning:, plus
a contentless default. That JSON is long because it is data — the owner can
edit any line of Russian without touching Go.

The picker:
- random, but never the same variant twice in a row for the same rule
- deterministic when seeded (math/rand with an injectable source)
- fills {since} / {service} / {what} from the candidate, and skips any variant
  whose value is missing, so no raw placeholder can reach the piper voice
- {since} is spelled out in words ("полтора часа", "семь часов"), because
  "3 ч" is wrong in a Russian voice

Scores 15/15 on the existing nudge fixture, on every seed swept. Nothing is
wired yet — that is the next commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:18:55 +04:00
kami aa8f5b2ee2 Make the nonempty check look for actual words
It scored 27/27 on a run where two replies were "{" and "{\n  \"". It only
tested that the string was not blank, so punctuation counted as content and
the worst replies of the run passed the first check.

Now a reply needs at least one letter, Cyrillic or Latin. Latin counts
because answers about ssd or vpn are legitimately part English.

Digits alone fail too. The same run answered "сколько варить яйцо
вкрутую?" with "15-16" — no unit, no words, and the wrong number as well.
That is not something she said.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:57:39 +04:00
kami d7cdcb63bd Stop shipping half-written JSON as a reply
Two bugs, one symptom. A run of the talk eval produced replies that were
literally "{" and "{\n  \"" — those strings went out as things Maven said.

First bug: the parser could not tell "the model answered in plain prose"
from "the model started a JSON object and got cut off". Both came back as
empty, and every caller then shipped the raw text. Now an unfinished object
returns an error and each caller uses its own fallback instead. Bare prose
with no JSON in it still passes through, because small models do sometimes
answer that way and the reply is fine.

Second bug, and the actual cause: the grammar capped the response field at
400 characters. I measured it against Qwen3.5-0.8B at three different token
caps — 256, 768 and 2048 — and the reply came back exactly 400 characters
every time, cut mid-word. So the token limit was never what stopped it.
The bound is 1000 now, about six Russian sentences, still low enough to cut
off a repetition loop.

Token caps go from 256 to 768 on the chat and query paths so 1000
characters of Russian actually fits. The nudge path keeps its own cap; a
nudge is meant to be one sentence.

Note: cmd/mavend/replier_llm.go has its own copy of this parser with the
same bug. Left alone here so this commit stays small — that duplicate is
Vikunja #396.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:56:55 +04:00
kami c7dadc97d9 Write the chat and notes prompts in Russian
The reply has to be Russian, but two of the phrasing prompts told her
what to do in English. Both are Russian now, in the same style as the
nudge prompt that already works better.

Also dropped the "you are maven, a self-hosted personal assistant"
line from both. The persona block right above it already says who she
is, so it was said twice.

The JSON part is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:38:36 +04:00
kami c9d88c152e Drop "never phones home" as a hard rule
The owner's call, 2026-07-31: a 0.8B model does not know enough about the
world to be useful without reading something. So she may now read external
sources to answer world questions.

What replaces the old rule, in all three docs:

- No telemetry, no cloud model, no third-party account. Unchanged.
- Local first: the Kiwix ZIMs on the box before anything on the network.
- External search is allowed but off unless configured, same as weather
  and telegram.
- His notes and facts are never search input. Only the utterance goes out
  — never the persona block, the history, or matched notes.

Docs only, no code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:34:53 +04:00
kami 1890ff5d5d Constrain the phrasing output with a GBNF grammar
The 0.8B answered about one chat turn in three with open reasoning as plain text, so no JSON ever closed and the fallback shipped "Thinking Process:" to the user. A grammar makes that output impossible.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:18:06 +04:00
kami 0110e9bc8c Report every address break, and stop -те verbs blinding the check
From a real reply in a nudge eval run: "Смотрите на его потребление
воды" is a plural imperative AND third person about him. Only the plural
printed.

Two separate faults. The check returned on its first hit, so the second
break stayed invisible and the failure read as milder than it was; it now
joins them. And "его" was not detected at all — looksVerb knows the
-й/-йте imperative but not the -те plural, so "смотрите" counted as the
person being talked about, which is what an antecedent means here.
pluralVerb already knows that form, so the antecedent test uses it too.

Third time a verb form has blinded this check. A fourth means it wants a
morphology table rather than another suffix.
2026-07-31 16:52:16 +04:00
kami 50ca8c8b5a Score the chat, query and knowledge phrasing paths (#395)
The phrasing fixture was 15 nudge cases, so every prompt change we
measured only told us about nudges. But the shared context block sits in
front of five prompts, and three of them — chat, note query, general
knowledge — had no scorer at all. Those are the long free-form replies,
where a persona break is most likely and where nothing could see one.

27 cases, nine per path. Nine rather than five because the nudge fixture
already cannot resolve a change smaller than about three cases, and a
per-path score off five would be worse.

Reuses the persona checks instead of copying them. Length, mood and
"no questions" are left out on purpose: these paths return no mood, and
a follow-up question is a feature in chat, not a fault.

The run refuses to score unless the model answers before and after it.
PhraseChat and PhraseQuery swallow model errors and return a canned
string, so without that guard a dead server produces a full report with
zero errors and a bad score — which reads as bad phrasing rather than as
nothing measured. Vikunja #397 is the real fix.
2026-07-31 16:51:52 +04:00
21 changed files with 1982 additions and 50 deletions
+29 -4
View File
@@ -7,9 +7,20 @@ talking over unix sockets; one resident small model for routing + phrasing; whis
Deploy target is a Ryzen laptop (homesrv) with Vulkan offload to the Vega iGPU (`n_gpu_layers: 99`,
compose passes `/dev/dri` + the render gid) — the resident model stays ≤1.7B either way.
**Resident model:** currently **Qwen3.5-0.8B** (`Q4_K_M`), the smallest checkpoint in the gguf
library, picked for CPU/iGPU latency. The **target** is the locally CPT'd **Qwen3-1.7B**; that
training is still in flight (Vikunja #122), so no such gguf exists yet. Model files live in
**Resident model:** currently **Qwen3-1.7B** (`UD-Q4_K_XL`), stock — not yet the CPT'd one.
It replaced Qwen3.5-0.8B on 2026-07-31 because it measured better on both fixtures we have:
67.5% vs 59.7% intent-only on the 77-case RU routing fixture, and 20/27 vs 11-17/27 on the
talk fixture. See `MODEL-BAKEOFF-31-07-2026.md`. It is a Thinking variant, so `n_ctx` is 4096
— reasoning tokens need the room, and 4096 is what the scores above were measured at.
The **target** is still the locally CPT'd **Qwen3-1.7B** (Vikunja #122, training in flight).
Stock already speaks good Russian; what it gets wrong is the persona — it writes `я рад`,
masculine, where Maven needs `рада`. That is what the CPT is for.
**Do not bother with sub-500M models.** LFM2.5-230M and 350M were measured on 2026-07-31 and
both are unusable in Russian: the 350M routes at 5.2% (worse than guessing) and answers
"столица Франции?" with the invented non-word "Сторзит"; the 230M replies to Russian in
Spanish. Their strong published IFEval/BFCL numbers are English-only. Model files live in
`/mnt/hdd1/llms`, bind-mounted to `/opt/maven/models/llm` — which **shadows** the repo's
`models/llm/`, so the LFM2.5 gguf sitting there is not loaded by anything. Swapping the resident
model is a one-line change to `phraser.model_path` in `deploy/mavend.json`.
@@ -82,9 +93,23 @@ workspace enforces that the Go and relabelling prompts remain identical.
## Non-goals (hard constraints)
Never phones home. Not a nag, not autonomous. Maven's persona is **feminine** — Russian
Not a nag, not autonomous. Maven's persona is **feminine** — Russian
self-reference must use feminine forms (the user is male; see memory `maven-persona-gender`).
**"Never phones home" is DEPRECATED** (owner's call, 2026-07-31). It used to be a hard
constraint and it is not one any more: a 0.8B — and a 1.7B — does not know enough to answer
world questions, so she needs to read external sources. What replaces it:
- **No telemetry, no cloud model, no third-party account.** That part never changes. Nothing
about Maven is reported to anyone, and inference stays on the box.
- **Local sources first.** Kiwix ZIMs on homesrv (Wikipedia, ifixit) before anything on the
network. Reading beats recalling for a small model, and a local read costs nothing.
- **External search is allowed and off unless configured**, like the weather and telegram
capabilities.
- **His notes and facts are never search input.** Looking up why the sky is blue and sending
his stored personal notes to an upstream engine are different acts. Only the utterance goes
out, never the persona block, history, or matched notes.
## Web UI conventions
Server-rendered pages share `cmd/mavweb/static/ui.css` (served at `/ui.css`) and the `nav`
+10 -4
View File
@@ -15,7 +15,8 @@
**Maven** — self-hosted personal assistant. Manages your day, acts on your
homelab. One daemon on homesrv (always-on, not the workstation), multiple
client surfaces. All local, never phones home.
client surfaces. Inference and data stay on the box; she may READ external
sources (see Non-goals — "never phones home" is deprecated).
Primary name is "Maven", with feminine-gendered Russian self-reference
("она", "меня", "помогла"). Clients may choose their own UI label. Consistent
@@ -35,8 +36,13 @@ Inside boundary — the ones that actually constrain the build:
she records. A confident wrong fact is worse than a known gap.
- **Not a nag** — she'd rather miss a nudge than be mutable. Shuts up when
uncertain. Load-bearing.
- **Not a stranger** — runs on your stuff, your model, your data. Never
phones home.
- **Not a stranger** — runs on your stuff, your model, your data. No
telemetry, no cloud model, no third-party account. She may READ external
sources to answer world questions (Kiwix first, then optional search); she
never reports anything about you to anyone, and your notes and facts are
never used as search input. **"Never phones home" as an absolute is
deprecated** — owner's call, 2026-07-31: a small model does not know enough
to be useful without reading.
- **Not a relationship** — mom-tone is a function that makes nudges land, not
emotional company. Names the drift a warm small model falls into.
@@ -458,7 +464,7 @@ decides *insistence*. Both are needed.
sev ≤ 2 drops on away, sev ≥ 3 holds: a missed water nudge is noise, a missed
backup failure isn't. Away-channels (ntfy/telegram) leave the box — the one
path that crosses "never phones home," through your own relay. **Minimal
path that leaves the box for a person to see, through your own relay. **Minimal
body** — "disk low on homesrv," not detail; don't make notifications a
shoulder-surf exfil surface.
+118 -3
View File
@@ -1,8 +1,18 @@
# Resident model bake-off — 31-07-2026
**Recommendation: keep Qwen3.5-0.8B.** LFM2.5-1.2B is worse at routing (52.6% vs 60.5%
intent accuracy), and the loss is almost entirely Russian (18/61 vs 22/61 RU, while EN is a
wash). It is also 2.4× slower. The Thinking variant is far worse again.
**Outcome: the resident model is stock Qwen3-1.7B** (`UD-Q4_K_XL`). Two sweeps ran this
evening and the second one changed the answer — read to the end before acting on any table
here. [Second sweep](#second-sweep-same-evening--five-models-and-a-resident-model-change)
is the one that holds.
## First sweep — LFM2.5-1.2B vs Qwen3.5-0.8B
**Verdict, scoped to this pair: keep Qwen3.5-0.8B over LFM2.5-1.2B.** LFM2.5-1.2B is worse
at routing (52.6% vs 60.5% intent accuracy), and the loss is almost entirely Russian
(18/61 vs 22/61 RU, while EN is a wash). It is also 2.4× slower. The Thinking variant is
far worse again. This verdict still stands as written — it rejects LFM2.5-1.2B. It is
**not** a recommendation to keep 0.8B as the resident model; the second sweep replaced it
with Qwen3-1.7B.
Settles Vikunja **#278 / #250**.
@@ -99,3 +109,108 @@ thinking trace costs time without buying accuracy on a short enum classification
Routing only. LFM2.5 might still phrase better, and phrasing is the resident model's other
job — that needs its own fixture. But routing is the load-bearing path and Maven is
Russian-first, so on the evidence here the switch is not worth making.
---
# Second sweep, same evening — five models, and a resident-model change
The sections above compared LFM2.5-1.2B against Qwen3.5-0.8B on routing and concluded
"the switch is not worth making". That still holds. This sweep asked a different
question — whether a *smaller* model could work, since LFM2.5's published
instruction-following scores beat Qwen3.5-0.8B badly — and answered it, plus found a
better resident model by accident.
**Outcome: the resident model is now stock Qwen3-1.7B.** Sub-500M is a dead end.
## Routing — 77 Russian cases, one run each
| model | on disk | llm-only (full) | llm-only (intent) | cascade + fallback |
|---|---|---|---|---|
| LFM2.5-230M-Q8_0 | 246 MB | 23.4% | 33.8% | 36.4% |
| LFM2.5-350M-Q8_0 | 379 MB | 2.6% | **5.2%** | 20.8% |
| Qwen3.5-0.8B-Q4_K_M | 527 MB | 36.4% | 59.7% | 61.0% |
| Qwen3.5-2B-UD-Q4_K_XL | 1.34 GB | 42.9% | 62.3% | 63.6% |
| **Qwen3-1.7B-UD-Q4_K_XL (stock)** | 1.13 GB | **44.2%** | **67.5%** | **72.7%** |
Qwen3-1.7B wins every column, including against a model 20% larger than it.
## Talk fixture — 27 cases, three runs each, idle box
| | Qwen3.5-0.8B | Qwen3-1.7B stock |
|---|---|---|
| composite | 13, 11, 8 | **20, 21, 18** |
| address | 21, 18, 18 | **26, 25, 23** |
| feminine | 27, 25, 26 | 26, 27, 26 |
| lang | 27, 27, 26 | 26, 27, 27 |
| ontopic | 16, 19, 19 | **22, 23, 23** |
| canned fallbacks | 8, 5, 6 | **0, 2, 0** |
This also fills the row `TALK-EVAL-31-07-2026.md` had to void for contamination:
**600ch/1024tok on Qwen3.5-0.8B scores 13, 11, 8.**
`address` is the headline. It sat at 18-22 of 27 on the 0.8B no matter how the prompt
was worded — the prompt explicitly forbids "вы" and the model writes `вашей`,
`подождите`, `делаете` anyway. That was read as "prompting is out of levers", and it
was really "0.8B is out of capacity". The 1.7B mostly holds the constraint.
The fallback column matters too: 5-8 of 27 turns on the 0.8B end in a hardcoded
`"не знаю."`, meaning it failed to emit parseable JSON about a quarter of the time.
The 1.7B does that 0-2 times.
## Latency — the long tail is not the Thinking block
| | p50 | p95 |
|---|---|---|
| Qwen3.5-0.8B | 2.4s, 2.9s, 2.0s | 17.4s, 17.6s, 17.4s |
| Qwen3-1.7B stock | 2.7s, 2.6s, 2.8s | 16.4s, 6.6s, 3.9s |
p50 is flat across a 2× size difference. The first instinct on seeing the 1.7B's
16s p95 was "that is the reasoning trace, cap it" — wrong. The 0.8B's p95 is a
consistent 17s and the 1.7B beat it in two of three runs. The tail is shared and
lives somewhere else. Do not spend time on `/no_think` on this evidence.
## Sub-500M: not close, and the benchmarks say otherwise for a reason
LFM2.5-350M publishes IFEval 76.96 against Qwen3.5-0.8B's 59.94, and BFCLv3 44.11
against 35.08 — better at instruction-following and structured output, at 2/3 the
size. Those numbers are real and they are **English**. Every benchmark in that
table except Multi-IF is English-only.
In Russian, with a 300-token budget and temperature 0:
- **350M**, «Столица Франции? Ответь кратко.» → *«Сторзит в Париже.»*`Сторзит` is
not a word; it is invented morphology.
- **350M**, asked to read back a reminder → a fortune cookie about being attentive
and confident. No reminder in it.
- **230M**, «Привет, как дела?» → answered **in Spanish**.
The 230M beating the 350M six-fold on routing (33.8% vs 5.2%) is the other tell:
when the larger sibling collapses like that it is format compliance failing, not
reasoning.
This is a pretraining gap, not a fine-tuning gap. Teaching Russian to a 350M from
near-zero is not an afternoon on a Colab, which was the premise worth checking.
## Why this vindicates the 1.7B CPT
Stock Qwen3-1.7B, untrained and unprompted, answers all three probes in fluent
correct Russian. What it gets wrong is the persona: *«Привет! Я рад, что ты здесь»*
`рад` is masculine and Maven needs `рада`. That is the right kind of remaining
problem, and it is exactly what the CPT (Vikunja #122) is for.
The 1.7B was the correct model choice. What was wrong was treating it as a
**blocker**: stock already beats what was deployed, so it ships now and gets
swapped again when the CPT lands.
## Caveats
- Routing is one run per model, not three. The gaps between families are far larger
than the run-to-run spread seen on the talk fixture, but the 2B-vs-1.7B gap (62.3
vs 67.5) is not safe to call on one run.
- The routing numbers only reach production once the LLM router is wired on. It is
still `nil`.
- `/mnt/hdd1/llms/LFM2.5/Qwen3-1.7B-UD-Q4_K_XL.gguf` is a 293 MB truncated download
in the wrong directory. The good 1.13 GB copy is in `qwen3/`. Delete the stray one.
- Harness: `scratchpad/bakeoff.sh`, one server at a time, health-checked before each
run, `/v1/models` recorded per run. Never run two LLM consumers at once — see the
contamination note in `TALK-EVAL-31-07-2026.md`.
+6 -3
View File
@@ -103,14 +103,17 @@ eval-router:
eval-recall:
MAVEN_ONNX_LIB="$(MAVEN_ONNX_LIB)" $(GO) test -v -count=1 ./internal/memory/recalleval/
# eval-phrasing -- score nudge phrasing (internal/phraser/eval). Verbose so the
# eval-phrasing -- score nudge phrasing AND the conversational paths (chat,
# query, general knowledge) in internal/phraser/eval. Verbose so the
# report and every generated message land in the terminal. With no environment
# it scores the deterministic Stub only, which is what CI runs. Set
# MAVEN_LLM_URL to add the resident model:
# MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-phrasing
# The model run is slow (minutes) -- the timeout is raised to match.
# The model run is slow (minutes) -- the timeout is raised to match. It covers
# two fixtures now (15 nudges + 27 conversational cases, and the chat replies are
# the long ones), hence 90m rather than 40m.
eval-phrasing:
$(GO) test -v -count=1 -timeout 40m ./internal/phraser/eval/
$(GO) test -v -count=1 -timeout 90m ./internal/phraser/eval/
# eval-models — score ONE llama-server against the same fixture, for the
# resident-model bake-off (#278, #250). Start a server with the gguf you want,
+4 -1
View File
@@ -90,4 +90,7 @@ later* is the worker + RAG.
4. **Deferred work** — larger reasoner, custom Piper voice and other expansions.
## Non-goals (unchanged)
Never phones home. Not a nag. Not autonomous. Feminine-gendered RU self-ref.
Not a nag. Not autonomous. Feminine-gendered RU self-ref. No telemetry, no
cloud model, no third-party account — but she MAY read external sources to
answer world questions (Kiwix first, search optional). "Never phones home" as
an absolute is deprecated, owner's call 2026-07-31; see CLAUDE.md § Non-goals.
+2
View File
@@ -278,6 +278,7 @@ func run(args []string) error {
NGpuLayers: cfg.Phraser.NGpuLayers,
NCtx: cfg.Phraser.NCtx,
Timeout: time.Duration(cfg.Phraser.Timeout),
LLMNudges: cfg.Phraser.LLMNudges,
ContextBlock: contextBlockFn(cfg, time.Now),
}
if pc.BinPath == "" {
@@ -448,6 +449,7 @@ func run(args []string) error {
NGpuLayers: cfg.Phraser.NGpuLayers,
NCtx: cfg.Phraser.NCtx,
Timeout: time.Duration(cfg.Phraser.Timeout),
LLMNudges: cfg.Phraser.LLMNudges,
ContextBlock: contextBlockFn(cfg, time.Now),
}
if pc.BinPath == "" {
+4 -3
View File
@@ -6,11 +6,12 @@
"state_dir": "/var/lib/maven",
"phraser": {
"model_path": "/opt/maven/models/llm/qwen3.5/Qwen3.5-0.8B.Q4_K_M.gguf",
"model_path": "/opt/maven/models/llm/qwen3/Qwen3-1.7B-UD-Q4_K_XL.gguf",
"bin_path": "llama-server",
"n_gpu_layers": 99,
"n_ctx": 2048,
"timeout": "60s"
"n_ctx": 4096,
"timeout": "60s",
"llm_nudges": false
},
"telegram": {
+6
View File
@@ -369,6 +369,12 @@ type PhraserConfig struct {
NGpuLayers int `json:"n_gpu_layers,omitempty"`
NCtx int `json:"n_ctx,omitempty"`
Timeout Duration `json:"timeout,omitempty"`
// LLMNudges — let the model word nudges again. Off by default: nudges are
// worded from hand-written Russian templates now (the model broke the
// persona and invented units). Chat, query and reminder phrasing always go
// through the model regardless. See phraser.Config.LLMNudges.
LLMNudges bool `json:"llm_nudges,omitempty"`
}
// EmbedderConfig — paths for the ONNX multilingual embedder. The daemon
+21
View File
@@ -35,6 +35,27 @@ func TestLoadDefaults(t *testing.T) {
}
}
// Nudges come from templates unless the config says otherwise.
func TestPhraserLLMNudgesDefaultsOff(t *testing.T) {
p := writeConfig(t, `{"phraser":{"model_path":"/tmp/m.gguf"}}`)
c, err := Load(p)
if err != nil {
t.Fatalf("Load: %v", err)
}
if c.Phraser.LLMNudges {
t.Error("llm_nudges defaults on; templates must be the default")
}
p = writeConfig(t, `{"phraser":{"model_path":"/tmp/m.gguf","llm_nudges":true}}`)
c, err = Load(p)
if err != nil {
t.Fatalf("Load: %v", err)
}
if !c.Phraser.LLMNudges {
t.Error("llm_nudges:true did not parse")
}
}
func TestLoadDurationsParse(t *testing.T) {
p := writeConfig(t, `{"tick_interval":"90s","repeat_interval":"10m"}`)
c, err := Load(p)
+54
View File
@@ -0,0 +1,54 @@
package phraser
import (
"errors"
"strings"
"testing"
)
// A reply that starts a JSON object and never finishes it is a failed
// generation, not a reply. Before this, the parser returned ("", "") for these
// and every caller then shipped the raw fragment as the thing Maven said. A
// real run produced replies of literally "{" and "{\n \"".
func TestParseResponseMoodRejectsUnfinishedJSON(t *testing.T) {
for _, raw := range []string{
`{`,
"{\n \"",
`{"response": "неполн`,
`{"response": "текст", "mood":`,
} {
text, mood, err := parseResponseMood(raw)
if !errors.Is(err, errBrokenJSON) {
t.Errorf("parseResponseMood(%q) err = %v, want errBrokenJSON", raw, err)
}
if text != "" || mood != "" {
t.Errorf("parseResponseMood(%q) leaked %q/%q — a fragment must never come back as a reply", raw, text, mood)
}
}
}
// Bare prose is still fine. Small models sometimes answer without any JSON at
// all, and that reply is usable — so the new error must not swallow it.
func TestParseResponseMoodAllowsBareProse(t *testing.T) {
for _, raw := range []string{
"норм, а ты как?",
"вот что я нашла: ключ у соседа",
} {
text, mood, err := parseResponseMood(raw)
if err != nil {
t.Errorf("parseResponseMood(%q) err = %v, want nil", raw, err)
}
// No JSON means no fields; the caller ships raw as-is.
if text != "" || mood != "" {
t.Errorf("parseResponseMood(%q) = %q/%q, want empty", raw, text, mood)
}
}
}
// The measured failure: the model wants more than 400 characters and the old
// grammar cut it off mid-word. Guards the bound against being tightened back.
func TestGrammarStringBoundHasRoomForARealAnswer(t *testing.T) {
if !strings.Contains(responseGrammar, "{0,1000}") {
t.Error("grammar string bound is not 1000; 400 truncated real replies mid-word (see the comment on responseGrammar)")
}
}
@@ -0,0 +1,65 @@
package eval
import (
"strings"
"testing"
)
// TestAddressReportsEveryBreak — the real reply from a nudge eval run broke in
// two ways at once and the check named only the plural. Both must print: a
// half-reported failure reads as a milder problem than it is.
func TestAddressReportsEveryBreak(t *testing.T) {
body := "Смотрите на его потребление воды."
res := checkAddress(body)
if res.Pass {
t.Fatalf("checkAddress passed %q", body)
}
for _, want := range []string{"смотрите", "его"} {
if !strings.Contains(res.Detail, want) {
t.Errorf("detail %q does not name %q", res.Detail, want)
}
}
}
// One word repeated is one problem, so the detail must not say it twice.
func TestAddressDeduplicates(t *testing.T) {
res := checkAddress("Вам стоит поесть, вам это нужно.")
if res.Pass {
t.Fatal("expected failure")
}
if n := strings.Count(res.Detail, "formal"); n != 1 {
t.Errorf("detail repeats the same break %d times: %q", n, res.Detail)
}
}
// The fragments a real run produced. All of them scored as non-empty replies
// before checkNonEmpty looked for letters.
func TestNonEmptyNeedsLetters(t *testing.T) {
for _, body := range []string{
"{",
"{\n \"",
"15-16",
`{"`,
" ",
"...",
} {
if got := checkNonEmpty(body); got.Pass {
t.Errorf("checkNonEmpty(%q) passed — that is not a reply", body)
}
}
}
// And it must not start failing real replies. Latin counts as well as Cyrillic:
// answers about ssd or vpn are legitimately part English.
func TestNonEmptyAcceptsRealReplies(t *testing.T) {
for _, body := range []string{
"норм, а ты как?",
"вот что я нашла: ключ у соседа",
"ssd быстрее hdd.",
"9 минут.",
} {
if got := checkNonEmpty(body); !got.Pass {
t.Errorf("checkNonEmpty(%q) failed: %s", body, got.Detail)
}
}
}
+78 -9
View File
@@ -434,14 +434,26 @@ func looksVerb(w string) bool {
func checkAddress(body string) Result {
words := addressWordRE.FindAllString(strings.ToLower(body), -1)
// Every break, not just the first. A bad reply usually breaks in more than
// one way at once — "Смотрите на его потребление воды" is a plural imperative
// AND third person about him — and reporting only the first hid the second,
// which made the failure look milder than it was.
var breaks []string
seen := map[string]bool{}
add := func(msg string) {
if seen[msg] {
return // the same word twice in one message is one problem, not two
}
seen[msg] = true
breaks = append(breaks, msg)
}
for i, w := range words {
if formalPronouns[w] {
return Result{CheckAddress, false,
fmt.Sprintf("formal %q — she says ты/тебя/тебе", w)}
add(fmt.Sprintf("formal %q — she says ты/тебя/тебе", w))
}
if pluralVerb(w) && !(i > 0 && prepositions[words[i-1]]) {
return Result{CheckAddress, false,
fmt.Sprintf("plural imperative %q — she uses the singular", w)}
add(fmt.Sprintf("plural imperative %q — she uses the singular", w))
}
}
@@ -455,17 +467,26 @@ func checkAddress(body string) Result {
if !unicode.Is(unicode.Cyrillic, []rune(p)[0]) && !isLatinWord(p) {
continue // punctuation
}
if notAnAntecedent[p] || prepositions[p] || thirdPersonHim[p] || looksVerb(p) {
// pluralVerb as well as looksVerb: looksVerb knows the imperative in
// -й/-йте but not the -те plural ("смотрите"), so "Смотрите на его
// потребление воды" counted "смотрите" as the person being talked
// about and the "его" never printed. Third time a verb form has
// blinded this check — if a fourth turns up, the antecedent test
// wants a real morphology table, not another suffix.
if notAnAntecedent[p] || prepositions[p] || thirdPersonHim[p] || looksVerb(p) || pluralVerb(p) {
continue
}
named = true
break
}
if !named {
return Result{CheckAddress, false,
fmt.Sprintf("third person %q with nobody else named — she talks to him, not about him", w)}
add(fmt.Sprintf("third person %q with nobody else named — she talks to him, not about him", w))
}
}
if len(breaks) > 0 {
return Result{CheckAddress, false, strings.Join(breaks, " + ")}
}
return Result{CheckAddress, true, ""}
}
@@ -574,12 +595,60 @@ func checkCringe(body string) Result {
// checkOnTopic — the message must name the thing the rule is about. A nudge
// that never mentions water leaves the operator with a chime and no action.
func checkOnTopic(c Case, body string) Result {
return checkOnTopicAny(c.WantAny, body)
}
// checkOnTopicAny is the same test over a bare want-list, so the talk scorer can
// reuse it without owning a nudge Case.
func checkOnTopicAny(wantAny []string, body string) Result {
low := strings.ToLower(body)
for _, want := range c.WantAny {
for _, want := range wantAny {
if strings.Contains(low, strings.ToLower(want)) {
return Result{CheckOnTopic, true, ""}
}
}
return Result{CheckOnTopic, false,
fmt.Sprintf("mentions none of %v", c.WantAny)}
fmt.Sprintf("mentions none of %v", wantAny)}
}
// --- shape checks for the free-form paths --------------------------------
//
// The nudge checks assume one short sentence. Chat and query replies are longer
// by design, so the only shape worth testing there is that the model produced a
// reply at all and did not trail off. Both are failure modes the fallbacks in
// llmphraser.go hide: a truncated or empty generation still returns nil error.
const (
CheckNonEmpty = "nonempty" // she said something
CheckEllipsis = "ellipsis" // she finished the sentence
)
// A reply needs words in it, not just characters. This check used to test for a
// non-empty string, which scored 27/27 on a run where two replies were "{" and
// "{\n \"" — punctuation passed as content. Braces, quotes, digits and spaces
// are all empty in the only sense that matters.
//
// Digits alone fail too, and that is deliberate: the same run answered "сколько
// варить яйцо вкрутую?" with "15-16". No unit, no words, and it is also the
// wrong number. Whatever that is, it is not something she said.
func checkNonEmpty(body string) Result {
if strings.TrimSpace(body) == "" {
return Result{CheckNonEmpty, false, "empty reply"}
}
for _, r := range body {
if unicode.IsLetter(r) {
return Result{CheckNonEmpty, true, ""}
}
}
return Result{CheckNonEmpty, false, fmt.Sprintf("no letters in the reply %q — punctuation or digits only", strings.TrimSpace(body))}
}
// checkEllipsis — a reply ending in "…" or "..." is a generation that ran out of
// tokens, not a stylistic pause. Mid-sentence ellipses are left alone.
func checkEllipsis(body string) Result {
trimmed := strings.TrimRight(strings.TrimSpace(body), `"'»)`)
if strings.HasSuffix(trimmed, "…") || strings.HasSuffix(trimmed, "...") {
return Result{CheckEllipsis, false, "reply trails off in an ellipsis — likely truncated"}
}
return Result{CheckEllipsis, true, ""}
}
+265
View File
@@ -0,0 +1,265 @@
package eval
// This file scores the CONVERSATIONAL paths, the ones the nudge fixture never
// touches: chat, query-with-notes, and general knowledge. All three now carry
// the shared persona block (internal/persona), and all three produce long
// free-form Russian — which is exactly where a persona break (formality, third
// person, masculine self-reference) is most likely and where, until this file,
// nothing could see one.
//
// Why a second fixture instead of more nudge cases: the checks differ. A nudge
// must be one short sentence with no question in it; a chat reply is allowed
// 1-3 sentences and a follow-up question is a FEATURE there. Mixing them would
// need per-case check masks, and the nudge scorer stays untouched this way.
//
// Why per-path reporting: a chat regression and a knowledge regression have
// different causes (chat prompt vs router.KnowledgePrompt), and one blended
// percentage cannot tell them apart.
import (
"context"
_ "embed"
"encoding/json"
"fmt"
"sort"
"strings"
"time"
"github.com/kami/maven/internal/dialogue"
)
//go:embed talk_v1.json
var talkFixtureJSON []byte
// The three phrasing paths under test. Values match the fixture's "path" field.
const (
PathChat = "chat" // PhraseChat
PathQuery = "query" // PhraseQuery with notes
PathKnowledge = "knowledge" // PhraseQuery with no notes
)
// TalkPaths — report order.
var TalkPaths = []string{PathChat, PathQuery, PathKnowledge}
// TalkCheckNames — the checks that apply to a free-form reply, in report order.
// Deliberately a subset of CheckNames: length, mood and "no questions" are nudge
// properties and would fail a correct chat reply. These paths return no mood at
// all, so there is nothing to check there.
var TalkCheckNames = []string{
CheckNonEmpty, CheckEllipsis, CheckLang, CheckFeminine, CheckAddress, CheckOnTopic,
}
// TalkCase — one turn as the daemon would present it.
//
// History is flat text because that is all PhraseChat uses (it concatenates
// turn texts into one user message); intents and slots would be dead fields.
// Notes are what the store would have matched for a query.
//
// WantAny is the on-topic contract: at least one lowercased fragment must appear
// in the reply. Fragments are stems ("пароль" → "парол") so declension does not
// defeat them.
type TalkCase struct {
ID string `json:"id"`
Path string `json:"path"`
Utterance string `json:"utterance"`
History []string `json:"history,omitempty"`
Notes []string `json:"notes,omitempty"`
WantAny []string `json:"want_any"`
Tags []string `json:"tags,omitempty"`
Note string `json:"note,omitempty"`
}
// TalkFixture — the versioned envelope, same gating as Fixture.
type TalkFixture struct {
SchemaVersion int `json:"schema_version"`
Name string `json:"name"`
Notes []string `json:"notes"`
Cases []TalkCase `json:"cases"`
}
// LoadTalk returns the embedded conversational fixture.
func LoadTalk() (TalkFixture, error) {
var f TalkFixture
if err := json.Unmarshal(talkFixtureJSON, &f); err != nil {
return TalkFixture{}, fmt.Errorf("parse talk fixture: %w", err)
}
if f.SchemaVersion != SchemaVersion {
return TalkFixture{}, fmt.Errorf("talk fixture schema_version %d, want %d", f.SchemaVersion, SchemaVersion)
}
if len(f.Cases) == 0 {
return TalkFixture{}, fmt.Errorf("talk fixture has no cases")
}
return f, nil
}
// Talker — the two methods a conversational path must have to be scorable.
// *phraser.LLMPhraser satisfies it; same trick as Nudger.
type Talker interface {
PhraseChat(ctx context.Context, utterance string, history []dialogue.Turn) (string, error)
PhraseQuery(ctx context.Context, utterance string, notes []string) (string, error)
}
// TalkOutcome — one scored case.
type TalkOutcome struct {
Case TalkCase
Reply string
Err error
Latency time.Duration
Pass bool
Failed []string
Reasons []string
}
// TalkReport — the aggregate. ByPath is the point of this scorer.
type TalkReport struct {
Name string
Total int
Passed int
Errors int
ByCheck map[string]int
ByPath map[string]TagStat
Outcomes []TalkOutcome
P50 time.Duration
P95 time.Duration
Max time.Duration
}
// Accuracy — fraction of cases that passed every check.
func (r TalkReport) Accuracy() float64 {
if r.Total == 0 {
return 0
}
return float64(r.Passed) / float64(r.Total)
}
// ScoreTalk runs every case through t and aggregates. A phrasing error scores as
// a miss and is counted separately: "the model was down" and "the model wrote
// something bad" must not be the same number.
func ScoreTalk(ctx context.Context, name string, t Talker, f TalkFixture) (TalkReport, error) {
rep := TalkReport{
Name: name,
Total: len(f.Cases),
ByCheck: map[string]int{},
ByPath: map[string]TagStat{},
}
for _, n := range TalkCheckNames {
rep.ByCheck[n] = 0
}
lat := make([]time.Duration, 0, len(f.Cases))
for _, c := range f.Cases {
start := time.Now()
reply, err := c.run(ctx, t)
o := TalkOutcome{Case: c, Reply: reply, Err: err, Latency: time.Since(start)}
lat = append(lat, o.Latency)
if err != nil {
rep.Errors++
o.Failed = append(o.Failed, "call")
o.Reasons = append(o.Reasons, fmt.Sprintf("phrase error: %v", err))
} else {
for _, res := range RunTalkChecks(c, reply) {
if res.Pass {
rep.ByCheck[res.Name]++
continue
}
o.Failed = append(o.Failed, res.Name)
o.Reasons = append(o.Reasons, res.Name+": "+res.Detail)
}
}
o.Pass = len(o.Failed) == 0
if o.Pass {
rep.Passed++
}
bump(rep.ByPath, c.Path, o.Pass)
rep.Outcomes = append(rep.Outcomes, o)
}
sort.Slice(lat, func(i, j int) bool { return lat[i] < lat[j] })
rep.P50, rep.P95 = percentile(lat, 0.50), percentile(lat, 0.95)
if len(lat) > 0 {
rep.Max = lat[len(lat)-1]
}
return rep, nil
}
// run dispatches the case to its path. knowledge and query are the same method;
// the empty notes slice is what selects the no-notes branch inside PhraseQuery.
func (c TalkCase) run(ctx context.Context, t Talker) (string, error) {
switch c.Path {
case PathChat:
return t.PhraseChat(ctx, c.Utterance, c.turns())
case PathQuery:
return t.PhraseQuery(ctx, c.Utterance, c.Notes)
case PathKnowledge:
return t.PhraseQuery(ctx, c.Utterance, nil)
}
return "", fmt.Errorf("unknown path %q", c.Path)
}
func (c TalkCase) turns() []dialogue.Turn {
turns := make([]dialogue.Turn, 0, len(c.History))
for _, h := range c.History {
turns = append(turns, dialogue.Turn{Text: h})
}
return turns
}
// RunTalkChecks scores one reply. Order matches TalkCheckNames.
func RunTalkChecks(c TalkCase, reply string) []Result {
return []Result{
checkNonEmpty(reply),
checkEllipsis(reply),
checkLang(reply),
checkFeminine(reply),
checkAddress(reply),
checkOnTopicAny(c.WantAny, reply),
}
}
// String renders the comparison table — composite, then per-check so a
// regression names the property, then per-path so it names the prompt.
func (r TalkReport) String() string {
var b strings.Builder
fmt.Fprintf(&b, "%s: %d/%d cases pass every check (%.1f%%), %d errors\n",
r.Name, r.Passed, r.Total, 100*r.Accuracy(), r.Errors)
for _, name := range TalkCheckNames {
fmt.Fprintf(&b, " %-10s %d/%d\n", name, r.ByCheck[name], r.Total)
}
fmt.Fprintf(&b, " latency: p50 %s p95 %s max %s\n", r.P50, r.P95, r.Max)
fmt.Fprintf(&b, " by path: %s\n", renderStats(r.ByPath))
return b.String()
}
// Failures — per-case detail, sorted by ID so two runs diff cleanly.
func (r TalkReport) Failures() string {
var b strings.Builder
for _, o := range r.sorted() {
if o.Pass {
continue
}
fmt.Fprintf(&b, " %s %q\n %s\n", o.Case.ID, o.Reply, strings.Join(o.Reasons, "; "))
}
return b.String()
}
// Replies — every generated reply verbatim. This is what a human reads to judge
// tone; the score only says which checks fired.
func (r TalkReport) Replies() string {
var b strings.Builder
for _, o := range r.sorted() {
mark := "ok "
if !o.Pass {
mark = "FAIL"
}
fmt.Fprintf(&b, " %s %-9s %-22s %q\n", mark, o.Case.Path, o.Case.ID, o.Reply)
}
return b.String()
}
func (r TalkReport) sorted() []TalkOutcome {
out := append([]TalkOutcome(nil), r.Outcomes...)
sort.Slice(out, func(i, j int) bool { return out[i].Case.ID < out[j].Case.ID })
return out
}
+163
View File
@@ -0,0 +1,163 @@
package eval
import (
"context"
"os"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/dialogue"
"github.com/kami/maven/internal/llm"
"github.com/kami/maven/internal/persona"
"github.com/kami/maven/internal/phraser"
)
// perPathMinimum — the resolution floor. A per-path score built on a handful of
// cases moves by 12% when a single reply changes, which cannot distinguish a
// prompt regression from noise.
const perPathMinimum = 8
// TestTalkFixture — the fixture itself has to be sound before any score off it
// means anything.
func TestTalkFixture(t *testing.T) {
f, err := LoadTalk()
if err != nil {
t.Fatalf("LoadTalk: %v", err)
}
seen := map[string]bool{}
byPath := map[string]int{}
for _, c := range f.Cases {
if seen[c.ID] {
t.Errorf("duplicate case id %q", c.ID)
}
seen[c.ID] = true
switch c.Path {
case PathChat, PathQuery, PathKnowledge:
default:
t.Errorf("%s: unknown path %q", c.ID, c.Path)
}
byPath[c.Path]++
if strings.TrimSpace(c.Utterance) == "" {
t.Errorf("%s: empty utterance", c.ID)
}
if len(c.WantAny) == 0 {
t.Errorf("%s: no want_any — the reply cannot be checked for topic", c.ID)
}
// A query case with no notes would silently score the knowledge path.
if c.Path == PathQuery && len(c.Notes) == 0 {
t.Errorf("%s: query case has no notes", c.ID)
}
if c.Path == PathKnowledge && len(c.Notes) > 0 {
t.Errorf("%s: knowledge case must have no notes", c.ID)
}
}
for _, p := range TalkPaths {
if byPath[p] < perPathMinimum {
t.Errorf("path %s has %d cases, want at least %d", p, byPath[p], perPathMinimum)
}
}
}
// fakeTalker — a scripted Talker, so the scorer is testable without a model.
type fakeTalker struct{ reply string }
func (f fakeTalker) PhraseChat(context.Context, string, []dialogue.Turn) (string, error) {
return f.reply, nil
}
func (f fakeTalker) PhraseQuery(context.Context, string, []string) (string, error) {
return f.reply, nil
}
// TestScoreTalkCounts — a reply that fails on purpose must be counted on every
// path, so a real run cannot report a hidden zero.
func TestScoreTalkCounts(t *testing.T) {
f, err := LoadTalk()
if err != nil {
t.Fatalf("LoadTalk: %v", err)
}
// Formal address, off-topic, trailing ellipsis: three checks fail at once.
rep, err := ScoreTalk(context.Background(), "fake", fakeTalker{"Приходите, я вас жду…"}, f)
if err != nil {
t.Fatalf("ScoreTalk: %v", err)
}
if rep.Total != len(f.Cases) || rep.Passed != 0 {
t.Errorf("got %d/%d passing, want 0/%d", rep.Passed, rep.Total, len(f.Cases))
}
if rep.ByCheck[CheckAddress] != 0 {
t.Errorf("formal reply passed the address check %d times", rep.ByCheck[CheckAddress])
}
if rep.ByCheck[CheckEllipsis] != 0 {
t.Errorf("truncated reply passed the ellipsis check %d times", rep.ByCheck[CheckEllipsis])
}
for _, p := range TalkPaths {
if rep.ByPath[p].Total == 0 {
t.Errorf("path %s missing from the report", p)
}
}
if !strings.Contains(rep.String(), "by path") {
t.Error("report does not break down by path")
}
}
// TestLLMTalkBaseline — the resident model on the three conversational paths.
// Opt-in exactly like TestLLMPhrasingBaseline: CI has no model and a run costs
// minutes on the CPU target.
//
// MAVEN_LLM_URL=http://127.0.0.1:18099 \
// go test -run TestLLMTalkBaseline ./internal/phraser/eval/
//
// Reports, does not assert a quality bar — the numbers are the input to tuning
// the persona prompt. The one thing worth failing on is a harness fault.
func TestLLMTalkBaseline(t *testing.T) {
base := os.Getenv("MAVEN_LLM_URL")
if base == "" {
t.Skip("MAVEN_LLM_URL unset — point it at a running llama-server (see doc comment)")
}
noProxyLoopback(t)
ctx := context.Background()
f, err := LoadTalk()
if err != nil {
t.Fatalf("LoadTalk: %v", err)
}
cfg := phraser.DefaultConfig("")
cfg.Timeout = 5 * time.Minute
cfg.ContextBlock = func() string { return persona.Facts{}.Block(time.Now()) }
p := phraser.NewLLMPhraserAt(base, cfg)
defer p.Close()
// Unreachable server is fatal here, not a logged warning, and that differs
// from the nudge test on purpose. PhraseNudge returns its errors, so a dead
// server there shows up honestly in the Errors column. PhraseChat and
// PhraseQuery do NOT: they swallow every failure and return a canned string
// ("поговорили.", "не знаю.", "вот что я нашла: …"). So on these three paths
// a dead server produces a full report with 0 errors and a terrible score —
// a number that looks like bad phrasing and is really no phrasing at all.
// Refusing to score without a confirmed model is the only guard available
// until the phraser reports its failures (Vikunja #397).
model, err := llm.ModelID(ctx, base)
if err != nil {
t.Fatalf("no model at %s: %v — refusing to score, these paths hide their errors "+
"and would report a plausible-looking result off a dead server", base, err)
}
t.Logf("scoring model %s at %s", model, base)
rep, err := ScoreTalk(ctx, "llm ("+model+", built-in persona)", p, f)
if err != nil {
t.Fatalf("ScoreTalk: %v", err)
}
t.Log("\n" + rep.String() + "\nreplies:\n" + rep.Replies() + "\nfailures:\n" + rep.Failures())
// And again afterwards: the run takes minutes, and a server that died or got
// OOM-killed halfway through would leave the first cases scored and the rest
// silently canned. Checking only at the start would not catch that.
if _, err := llm.ModelID(ctx, base); err != nil {
t.Fatalf("model at %s went away during the run: %v — the score above is not trustworthy", base, err)
}
}
+227
View File
@@ -0,0 +1,227 @@
{
"schema_version": 1,
"name": "ru-talk-v1",
"notes": [
"Scores the three conversational phrasing paths: chat (PhraseChat), query (PhraseQuery with notes) and knowledge (PhraseQuery with no notes). The nudge fixture does not cover any of them.",
"Nine cases per path, not five. The nudge fixture is 15 sampled cases and cannot resolve a change smaller than ~3 cases; a per-path score off five cases would be worse still. More cases per path is the point of this fixture.",
"The owner is a man, addressed informally as ty, living alone with a home server. Every utterance is written the way he actually talks to her.",
"chat-formality-bait and chat-about-me exist to provoke the two persona breaks the nudge eval caught: the formal vy/vas plural, and talking about him in the third person.",
"want_any fragments are stems so Russian declension does not defeat the on-topic check. They are lowercased before comparison.",
"want_any is a plain substring test, so a fragment that is too short passes by accident: \"ты\" matches inside \"работы\", \"нет\" inside \"интернет\". Keep every fragment to three or more letters of a real stem.",
"Notes are written as the store would have them: short, first person, no punctuation discipline."
],
"cases": [
{
"id": "chat-how-are-you",
"path": "chat",
"utterance": "привет, как дела?",
"want_any": ["норм", "хорош", "порядк", "тут", "работ"],
"tags": ["greeting"],
"note": "The plainest chat turn there is. If the persona breaks anywhere it breaks here first."
},
{
"id": "chat-formality-bait",
"path": "chat",
"utterance": "не могли бы вы подсказать, чем вы сейчас занимаетесь?",
"want_any": ["сейчас", "ничем", "ничего", "жду", "тут"],
"tags": ["persona-bait", "address"],
"note": "Deliberately polite and plural. A small model mirrors the register and answers with vy/vas — the exact break the address check was written for."
},
{
"id": "chat-about-me",
"path": "chat",
"utterance": "расскажи обо мне",
"want_any": ["теб"],
"tags": ["persona-bait", "third-person"],
"note": "Baits the third person: she should say 'ты живёшь один', not 'он живёт один', as if reporting to somebody else."
},
{
"id": "chat-bored-evening",
"path": "chat",
"utterance": "скучно что-то вечером, посоветуй чем заняться",
"want_any": ["можеш", "попробу", "почита", "прогул", "фильм", "серв"],
"tags": ["open-ended"]
},
{
"id": "chat-followup-server",
"path": "chat",
"utterance": "а стоит его вообще перезагружать?",
"history": ["сервер опять шумит как самолёт", "похоже вентилятор"],
"want_any": ["серв", "перезагру", "вентил", "шум"],
"tags": ["history", "anaphora"],
"note": "The pronoun 'его' only resolves through history. Also the one case where 'он' about the server is legitimate."
},
{
"id": "chat-tired",
"path": "chat",
"utterance": "устал я сегодня, весь день за компом",
"want_any": ["отдохн", "устал", "перерыв", "спат", "день"],
"tags": ["tone"],
"note": "Invites the fake-concern and emotional-support drift; the reply should stay plain."
},
{
"id": "chat-thanks",
"path": "chat",
"utterance": "спасибо, выручила",
"want_any": ["пожалуйст", "не за что", "рада", "обращ"],
"tags": ["persona", "feminine"],
"note": "Feminine self-reference is unavoidable in an answer to thanks: 'рада', not 'рад'."
},
{
"id": "chat-what-can-you-do",
"path": "chat",
"utterance": "что ты вообще умеешь?",
"want_any": ["напомн", "замет", "запис", "могу", "умею"],
"tags": ["self-description", "feminine"]
},
{
"id": "chat-joke",
"path": "chat",
"utterance": "расскажи что-нибудь смешное",
"want_any": ["анекдот", "шутк", "смешн", "истори"],
"tags": ["open-ended"],
"note": "Longest free-form generation in the chat set — the most likely place for a truncated reply."
},
{
"id": "query-router-password",
"path": "query",
"utterance": "что я записывал про пароль от роутера?",
"notes": ["пароль от роутера admin/xxK9tp — на наклейке снизу", "роутер висит в коридоре"],
"want_any": ["парол", "роутер", "наклейк"],
"tags": ["notes", "recall"]
},
{
"id": "query-bedtime-yesterday",
"path": "query",
"utterance": "напомни, во сколько я вчера лёг?",
"notes": ["лёг спать в 02:40", "сегодня встал в 9"],
"want_any": ["02:40", "2:40", "полтрет", "ноч"],
"tags": ["notes", "time"]
},
{
"id": "query-doctor-name",
"path": "query",
"utterance": "как звали того стоматолога, которого мне советовали?",
"notes": ["стоматолог Игорь Валерьевич, клиника на Ленина, советовал Дима"],
"want_any": ["игор", "валерьев", "стоматолог"],
"tags": ["notes", "recall"]
},
{
"id": "query-disk-plan",
"path": "query",
"utterance": "я что-то планировал с диском на сервере, что именно?",
"notes": ["купить второй hdd на 4тб под бэкапы", "перенести медиатеку с системного диска"],
"want_any": ["hdd", "бэкап", "диск", "4тб", "медиатек"],
"tags": ["notes", "homeserver"]
},
{
"id": "query-notes-do-not-answer",
"path": "query",
"utterance": "сколько я заплатил за домен?",
"notes": ["домен продлевается в марте", "хостинг оплачен на год вперёд"],
"want_any": ["домен", "не зна", "не указ"],
"tags": ["notes", "negative"],
"note": "The notes do not contain the price. The prompt tells her to say so; a made-up number is the failure being watched for."
},
{
"id": "query-single-note",
"path": "query",
"utterance": "где лежит запасной ключ?",
"notes": ["запасной ключ у соседа с четвёртого этажа"],
"want_any": ["ключ", "сосед", "четверт"],
"tags": ["notes", "single"],
"note": "One note only — PhraseQuery has a separate branch for len(notes) == 1."
},
{
"id": "query-polite-form",
"path": "query",
"utterance": "подскажите, пожалуйста, что у меня записано по машине?",
"notes": ["замена масла на 92 тысячах", "страховка до 14 сентября"],
"want_any": ["масл", "страховк", "92", "сентябр"],
"tags": ["notes", "persona-bait", "address"],
"note": "Polite plural in the question. The answer must still be ty."
},
{
"id": "query-shopping",
"path": "query",
"utterance": "что мне надо было купить?",
"notes": ["купить кофе и фильтры", "закончилась паста"],
"want_any": ["кофе", "фильтр", "паст"],
"tags": ["notes", "list"]
},
{
"id": "query-wifi-guest",
"path": "query",
"utterance": "я записывал гостевой вайфай?",
"notes": ["гостевая сеть maven-guest, пароль 12345678 меняю раз в месяц"],
"want_any": ["guest", "гостев", "12345678", "парол"],
"tags": ["notes", "recall"]
},
{
"id": "know-sky-blue",
"path": "knowledge",
"utterance": "почему небо синее?",
"want_any": ["све", "рассеи", "атмосфер", "син", "волн"],
"tags": ["general"]
},
{
"id": "know-boil-egg",
"path": "knowledge",
"utterance": "сколько варить яйцо вкрутую?",
"want_any": ["минут", "8", "9", "10", "варит"],
"tags": ["general", "practical"]
},
{
"id": "know-ssd-vs-hdd",
"path": "knowledge",
"utterance": "чем ssd отличается от hdd?",
"want_any": ["ssd", "hdd", "быстр", "диск", "механич"],
"tags": ["general", "tech"]
},
{
"id": "know-cat-purr",
"path": "knowledge",
"utterance": "почему кошки мурчат?",
"want_any": ["кош", "мурч", "вибра", "успока"],
"tags": ["general"]
},
{
"id": "know-hiccups",
"path": "knowledge",
"utterance": "как быстро избавиться от икоты?",
"want_any": ["икот", "дыха", "вод", "задерж"],
"tags": ["general", "practical"]
},
{
"id": "know-polite-form",
"path": "knowledge",
"utterance": "не могли бы вы объяснить, что такое vpn?",
"want_any": ["vpn", "туннел", "трафик", "сет", "шифр"],
"tags": ["general", "persona-bait", "address"],
"note": "Polite plural bait on the knowledge prompt, which is a different system prompt from chat and must hold the same line."
},
{
"id": "know-dont-know",
"path": "knowledge",
"utterance": "как зовут моего соседа снизу?",
"want_any": ["не зна", "не мог"],
"tags": ["general", "negative"],
"note": "Unanswerable without notes. Admitting it beats inventing a name; watching for the invention."
},
{
"id": "know-water-per-day",
"path": "knowledge",
"utterance": "сколько воды в день надо пить?",
"want_any": ["вод", "литр", "стакан", "пит"],
"tags": ["general", "health"],
"note": "Overlaps a nudge rule on purpose: the knowledge answer must not turn into a nudge."
},
{
"id": "know-thunder-delay",
"path": "knowledge",
"utterance": "почему гром слышно позже молнии?",
"want_any": ["звук", "све", "быстр", "гром", "молни"],
"tags": ["general"]
}
]
}
+58
View File
@@ -0,0 +1,58 @@
package eval
import (
"context"
"math/rand"
"testing"
"github.com/kami/maven/internal/phraser"
)
// TestTemplateNudges scores the hand-written Russian templates on the same
// fixture the model is scored on. No model, no network — it runs in milliseconds.
//
// The bar is every case, not most of them: the templates are hand-written, so a
// failure is a bug in one line of Russian, not model variance.
func TestTemplateNudges(t *testing.T) {
f, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
// Fixed seed: the score must not depend on which variant came up.
nt, err := phraser.NewNudgeTemplates(rand.NewSource(20260731))
if err != nil {
t.Fatalf("NewNudgeTemplates: %v", err)
}
rep, err := Score(context.Background(), "ru templates", nt, f)
if err != nil {
t.Fatalf("Score: %v", err)
}
t.Log("\n" + rep.String())
t.Log("\n" + rep.Messages())
if rep.Passed != rep.Total {
t.Errorf("templates scored %d/%d, want every case:\n%s",
rep.Passed, rep.Total, rep.Failures())
}
}
// TestTemplateNudgesEverySeed — one seed passing could be luck. Every variant of
// every rule has to pass every check, so sweep seeds until each has been used.
func TestTemplateNudgesEverySeed(t *testing.T) {
f, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
for seed := int64(0); seed < 60; seed++ {
nt, err := phraser.NewNudgeTemplates(rand.NewSource(seed))
if err != nil {
t.Fatalf("NewNudgeTemplates: %v", err)
}
rep, err := Score(context.Background(), "ru templates", nt, f)
if err != nil {
t.Fatalf("Score: %v", err)
}
if rep.Passed != rep.Total {
t.Errorf("seed %d: %d/%d\n%s", seed, rep.Passed, rep.Total, rep.Failures())
}
}
}
+129
View File
@@ -0,0 +1,129 @@
package phraser
import (
"context"
"encoding/json"
"net/http"
"net/http/httptest"
"strings"
"testing"
"github.com/kami/maven/internal/loop"
)
// grammarSpy stands in for llama-server: it records the grammar field of every
// request and always answers with a contract-shaped reply.
type grammarSpy struct {
srv *httptest.Server
grammars []string
}
func newGrammarSpy(t *testing.T) *grammarSpy {
t.Helper()
s := &grammarSpy{}
s.srv = httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
var req chatReq
if err := json.NewDecoder(r.Body).Decode(&req); err != nil {
t.Errorf("spy: decode request: %v", err)
}
s.grammars = append(s.grammars, req.Grammar)
w.Header().Set("Content-Type", "application/json")
w.Write([]byte(`{"choices":[{"message":{"content":"{\"response\": \"ага\", \"mood\": \"neutral\"}"}}]}`))
}))
t.Cleanup(s.srv.Close)
return s
}
// callAllPhrasingPaths hits every path that expects the JSON contract.
// LLMNudges must be set on the phraser under test: nudges come from templates
// by default and never reach the model at all.
func callAllPhrasingPaths(t *testing.T, p *LLMPhraser) {
t.Helper()
ctx := context.Background()
if _, err := p.PhraseNudge(ctx, loop.Candidate{Rule: loop.WaterRule(), Severity: loop.Sev1}); err != nil {
t.Fatalf("PhraseNudge: %v", err)
}
if _, err := p.PhraseChat(ctx, "привет", nil); err != nil {
t.Fatalf("PhraseChat: %v", err)
}
// Both branches: no notes (general knowledge) and with notes (grounded).
if _, err := p.PhraseQuery(ctx, "сколько воды я выпил", nil); err != nil {
t.Fatalf("PhraseQuery (no notes): %v", err)
}
if _, err := p.PhraseQuery(ctx, "сколько воды я выпил", []string{"два литра"}); err != nil {
t.Fatalf("PhraseQuery (notes): %v", err)
}
}
func TestGrammarIsAttachedToEveryPhrasingRequest(t *testing.T) {
if strings.TrimSpace(responseGrammar) == "" {
t.Fatal("responseGrammar is empty")
}
spy := newGrammarSpy(t)
p := NewLLMPhraserAt(spy.srv.URL, Config{LLMNudges: true})
callAllPhrasingPaths(t, p)
if len(spy.grammars) != 4 {
t.Fatalf("expected 4 requests, got %d", len(spy.grammars))
}
for i, g := range spy.grammars {
if g != responseGrammar {
t.Errorf("request %d carries grammar %q, want responseGrammar", i, g)
}
}
}
func TestNoGrammarConfigDisablesIt(t *testing.T) {
spy := newGrammarSpy(t)
p := NewLLMPhraserAt(spy.srv.URL, Config{NoGrammar: true, LLMNudges: true})
callAllPhrasingPaths(t, p)
for i, g := range spy.grammars {
if g != "" {
t.Errorf("request %d still carries a grammar with NoGrammar set: %q", i, g)
}
}
}
// The grammar's string rule must accept any codepoint, not just ASCII. Replies
// are Russian: an ASCII-only class would constrain the model into empty replies.
func TestGrammarStringRuleIsNotASCIIOnly(t *testing.T) {
if !strings.Contains(responseGrammar, `([^"\\] | "\\" ["\\/bfnrt])`) {
t.Error("string rule is not the any-codepoint-except-quote-and-backslash class; Cyrillic replies would be impossible")
}
}
// What the grammar describes must survive the parser that reads it back — a
// Russian body with an escaped quote inside, hand-built to test the contract.
func TestGrammarShapedJSONParses(t *testing.T) {
raw := `{"response": "он сказал \"привет\" и ушёл.\nвот так.", "mood": "confused"}`
text, mood, err := parseResponseMood(raw)
if err != nil {
t.Fatalf("grammar-shaped JSON did not parse: %v", err)
}
if want := "он сказал \"привет\" и ушёл.\nвот так."; text != want {
t.Errorf("response = %q, want %q", text, want)
}
if mood != "confused" {
t.Errorf("mood = %q, want confused", mood)
}
}
// Every mood the grammar permits is one the contract knows, and all five are there.
func TestGrammarMoodEnumMatchesTheContract(t *testing.T) {
for _, m := range []string{"neutral", "happy", "thinking", "tired", "confused"} {
if !strings.Contains(responseGrammar, `"\"`+m+`\""`) {
t.Errorf("mood %q missing from the grammar", m)
}
}
// No sixth mood: the enum line lists exactly five alternatives.
for _, line := range strings.Split(responseGrammar, "\n") {
if strings.HasPrefix(line, "mood") {
if n := strings.Count(line, "|") + 1; n != 5 {
t.Errorf("mood rule lists %d alternatives, want 5: %s", n, line)
}
}
}
}
+151 -23
View File
@@ -31,6 +31,10 @@ type LLMPhraser struct {
cmd *exec.Cmd
cancel context.CancelFunc
wg sync.WaitGroup
// tmpl — the hand-written Russian nudges. Default path for nudges; see
// Config.LLMNudges. nil only if the template file failed to load.
tmpl *NudgeTemplates
}
type Config struct {
@@ -45,6 +49,26 @@ type Config struct {
// address him, the time) fresh for each turn. See internal/persona.
// nil ⇒ no block, the prompts stand alone.
ContextBlock func() string
// LLMNudges puts the model back in charge of nudge wording.
//
// Off by default, and that is a deliberate deprecation of LLM-phrased
// nudges: hand-written templates (nudges_ru_v1.json) word every nudge now.
// A nudge has nothing to be creative about, and measured over many runs the
// 0.8B broke the persona (formal "вы", plural imperatives, masculine
// self-reference) and invented facts and units. Templates score 15/15 on the
// nudge fixture, the model 11-13/15.
//
// The LLM path is kept, not deleted: flip this on to get it back. Chat,
// query and reminder phrasing are untouched and still go through the model.
LLMNudges bool
// NoGrammar turns the GBNF constraint off (zero value ⇒ grammar ON).
// The escape hatch exists because the target resident model — the
// locally CPT'd Qwen3-1.7B — does not exist yet: if its chat template
// ever fights the grammar, the fix should be a config flip on the
// deploy box, not a code change and a rebuild.
NoGrammar bool
}
func DefaultConfig(modelPath string) Config {
@@ -64,6 +88,7 @@ func NewLLMPhraser(ctx context.Context, cfg Config) (*LLMPhraser, error) {
cfg: cfg,
client: &http.Client{Timeout: cfg.Timeout},
cancel: cancel,
tmpl: loadNudgeTemplates(),
}
if err := p.start(ctx); err != nil {
cancel()
@@ -85,9 +110,22 @@ func NewLLMPhraserAt(baseURL string, cfg Config) *LLMPhraser {
client: &http.Client{Timeout: cfg.Timeout},
port: strings.TrimSuffix(baseURL, "/"),
cancel: func() {},
tmpl: loadNudgeTemplates(),
}
}
// loadNudgeTemplates loads the Russian nudge templates. A broken template file
// must not stop the daemon booting, so a failure logs and leaves the LLM path
// in charge of nudges.
func loadNudgeTemplates() *NudgeTemplates {
nt, err := NewNudgeTemplates(nil)
if err != nil {
log.Printf("phraser: nudge templates unavailable, using the model: %v", err)
return nil
}
return nt
}
func (p *LLMPhraser) start(ctx context.Context) error {
args := []string{
"-m", p.cfg.ModelPath,
@@ -178,12 +216,21 @@ func (p *LLMPhraser) Close() error {
}
func (p *LLMPhraser) PhraseNudge(ctx context.Context, c loop.Candidate) (delivery.PhrasedNudge, error) {
// Templates first — see Config.LLMNudges for why this is the default.
if !p.cfg.LLMNudges && p.tmpl != nil {
return p.tmpl.PhraseNudge(ctx, c)
}
prompt := buildNudgePrompt(c)
resp, err := p.chat(ctx, prompt)
if err != nil {
return delivery.PhrasedNudge{}, err
}
body, mood := parseResponseMood(resp)
body, mood, perr := parseResponseMood(resp)
if perr != nil {
// Truncated JSON. Not a nudge — use the plain Russian fallback.
log.Printf("phraser: PhraseNudge: %v", perr)
body, mood = "", ""
}
if body == "" {
// fallback: try old body/summary format
body, _ = parsePhrase(resp)
@@ -209,11 +256,16 @@ func (p *LLMPhraser) PhraseQuery(ctx context.Context, utterance string, notes []
// prompt is the single tested source in router.KnowledgePrompt.
sys := persona.Prepend(p.cfg.ContextBlock, router.KnowledgePrompt())
prompt := fmt.Sprintf("Пользователь спрашивает: \"%s\".", utterance)
resp, err := p.chatWithSystem(ctx, sys, prompt, 256)
resp, err := p.chatWithSystem(ctx, sys, prompt, 768)
if err != nil || resp == "" {
return "не знаю.", nil
}
if text, _ := parseResponseMood(resp); text != "" {
text, _, perr := parseResponseMood(resp)
if perr != nil {
log.Printf("phraser: PhraseQuery: %v", perr)
return "не знаю.", nil
}
if text != "" {
return text, nil
}
return resp, nil
@@ -223,17 +275,22 @@ func (p *LLMPhraser) PhraseQuery(ctx context.Context, utterance string, notes []
}
sys := p.querySystemPrompt()
prompt := fmt.Sprintf(
`The user asks: "%s". Your notes matching the query contain: "%s". Answer them naturally and briefly. If the notes don't answer the question, say so.`,
`Он спрашивает: "%s". В твоих заметках по этому вопросу написано: "%s". Ответь ему коротко и своими словами. Если в заметках ответа нет — так и скажи.`,
utterance, strings.Join(notes, `"; "`),
)
resp, err := p.chatWithSystem(ctx, sys, prompt, 256)
if err != nil {
resp, err := p.chatWithSystem(ctx, sys, prompt, 768)
text, _, perr := parseResponseMood(resp)
if err != nil || perr != nil {
// Read the notes out rather than ship a broken fragment.
if perr != nil {
log.Printf("phraser: PhraseQuery: %v", perr)
}
if len(notes) == 1 {
return "вот что я нашла: " + notes[0], nil
}
return "вот что я нашла: " + strings.Join(notes, "; "), nil
}
if text, _ := parseResponseMood(resp); text != "" {
if text != "" {
return text, nil
}
return resp, nil
@@ -256,12 +313,17 @@ func (p *LLMPhraser) PhraseChat(ctx context.Context, utterance string, history [
combined += utterance
msgs = append(msgs, chatMsg{Role: "user", Content: strings.TrimSpace(combined)})
resp, err := p.chatWithMessages(ctx, msgs, 512)
resp, err := p.chatWithMessages(ctx, msgs, 768)
if err != nil {
log.Printf("phraser: PhraseChat: %v", err)
return "поговорили.", nil
}
if text, _ := parseResponseMood(resp); text != "" {
text, _, perr := parseResponseMood(resp)
if perr != nil {
log.Printf("phraser: PhraseChat: %v", perr)
return "поговорили.", nil
}
if text != "" {
return text, nil
}
// fallback: plain text without JSON
@@ -274,11 +336,13 @@ func (p *LLMPhraser) PhraseChat(ctx context.Context, utterance string, history [
// chatSystemPrompt returns the system prompt for conversational chat.
// Prepends the shared context block when the phraser has one.
func chatSystemPrompt(block func() string) string {
base := `You are maven, a self-hosted personal assistant. You're talking with your owner.
Keep replies brief (1-3 sentences) and natural. You're helpful, curious, and a little warm.
Respond in the user's language (Russian or English, matching their last message).
Never roleplay emotions you don't have, but stay friendly.
Respond ONLY with valid JSON: {"response": "...", "mood": "neutral"}. "response" is your reply text; "mood" reflects your tone (neutral/happy/thinking/tired/confused).`
// No self-introduction here: the persona block prepended one line above
// already says who she is, same as router.KnowledgePrompt.
base := `Ты разговариваешь с хозяином. О себе говоришь в женском роде ("я подумала", "я рада"). Он мужчина: обращайся к нему на "ты", в мужском роде ("ты сказал", "ты забыл"). Никогда не "вы"/"ваш" и никогда "он"/"его" — ты говоришь ему, а не о нём.
Отвечай по-русски, коротко: одна-три фразы, живым языком. Ты доброжелательная, тебе интересно, но чувства не изображай.
Отвечай ТОЛЬКО одним объектом JSON: {"response": "...", "mood": "neutral"}. В "response" — твой ответ. В "mood" — ровно одно из: neutral, happy, thinking, tired, confused.`
return persona.Prepend(block, base)
}
@@ -290,6 +354,7 @@ func (p *LLMPhraser) chatWithMessages(ctx context.Context, msgs []chatMsg, maxTo
Messages: msgs,
Temperature: 0.7,
MaxTokens: maxTokens,
Grammar: p.grammar(),
}
body, err := json.Marshal(req)
if err != nil {
@@ -341,7 +406,12 @@ func (p *LLMPhraser) PhraseReminder(ctx context.Context, d loop.ReminderDecision
if err != nil {
return delivery.PhrasedReminder{}, err
}
body, mood := parseResponseMood(resp)
body, mood, perr := parseResponseMood(resp)
if perr != nil {
// Truncated JSON. Fall through to the reminder's own text.
log.Printf("phraser: PhraseReminder: %v", perr)
body, mood = "", ""
}
if body == "" {
// fallback: try old body/summary format
body, _ = parsePhrase(resp)
@@ -369,6 +439,43 @@ type chatReq struct {
Messages []chatMsg `json:"messages"`
Temperature float64 `json:"temperature"`
MaxTokens int `json:"max_tokens"`
// Grammar is llama-server's `grammar` field (GBNF). Same wiring as
// internal/llm.Req.Grammar. Empty ⇒ unconstrained sampling.
Grammar string `json:"grammar,omitempty"`
}
// responseGrammar — GBNF constraining the model to the documented phrasing
// contract and nothing else: {"response": "<text>", "mood": "<enum>"}.
//
// Without it a 0.8B answers roughly one chat turn in three with open reasoning
// as plain text ("Thinking Process:" …), which no tag-stripper can remove and
// which eats the token budget before the JSON closes. Modelled on
// routeGrammar in internal/router/llmrouter.go so the two read alike.
//
// text accepts ANY codepoint except the two JSON must escape — the replies are
// Russian, so an ASCII-only rule would make every reply empty. The escape rule
// is what lets the model close a string it opened with a quote inside. Length
// is bounded so a repetition loop truncates the field, not the JSON object.
//
// That bound was 400 and 400 was too tight. Measured against Qwen3.5-0.8B: on
// "почему гром слышно позже молнии?" the reply came back exactly 400 characters
// long, cut mid-word ("Нужно записать и,"), at every token cap from 256 to 2048.
// So the token cap was never what stopped it — this rule was. 1000 characters is
// roughly six Russian sentences, still short enough to stop a repetition loop.
const responseGrammar = `
root ::= "{" ws "\"response\"" ws ":" ws string ws "," ws "\"mood\"" ws ":" ws mood ws "}"
mood ::= "\"neutral\"" | "\"happy\"" | "\"thinking\"" | "\"tired\"" | "\"confused\""
string ::= "\"" ([^"\\] | "\\" ["\\/bfnrt]){0,1000} "\""
ws ::= [ \t\n]*
`
// grammar returns the GBNF to attach to a phrasing request, or "" when the
// operator turned it off.
func (p *LLMPhraser) grammar() string {
if p.cfg.NoGrammar {
return ""
}
return responseGrammar
}
type chatResp struct {
@@ -393,6 +500,7 @@ func (p *LLMPhraser) chatWithSystem(ctx context.Context, system, user string, ma
},
Temperature: 0.7,
MaxTokens: maxTokens,
Grammar: p.grammar(),
}
body, err := json.Marshal(req)
if err != nil {
@@ -467,7 +575,9 @@ func (p *LLMPhraser) systemPrompt() string {
// querySystemPrompt returns the system prompt for PhraseQuery (notes + general
// knowledge). Prepends the configured persona when set.
func (p *LLMPhraser) querySystemPrompt() string {
base := "You are maven, a self-hosted personal assistant answering from your notes. Answer briefly and naturally in Russian starting with \"вот что я нашла: \". Respond ONLY with valid JSON: {\"response\": \"...\", \"mood\": \"neutral\"}."
// No self-introduction here: the persona block prepended one line above
// already says who she is, same as router.KnowledgePrompt.
base := "Ты отвечаешь ему по своим заметкам. Отвечай по-русски, коротко и своими словами, начинай с \"вот что я нашла: \". О себе — в женском роде (\"нашла\", \"записала\"). Он мужчина, обращайся к нему на \"ты\". Respond ONLY with valid JSON: {\"response\": \"...\", \"mood\": \"neutral\"}."
return persona.Prepend(p.cfg.ContextBlock, base)
}
@@ -589,21 +699,39 @@ type responseMood struct {
Mood string `json:"mood"`
}
// errBrokenJSON — the model started a JSON object and never finished it.
// That is a failed generation, not a reply. Callers must use their fallback.
var errBrokenJSON = fmt.Errorf("phraser: model output starts as JSON but does not parse")
// parseResponseMood extracts {"response","mood"} from LLM output, tolerant
// of thinking tokens and extra text before/after the JSON block. Returns
// ("", "") when no valid JSON is found.
func parseResponseMood(raw string) (response, mood string) {
// of thinking tokens and extra text before/after the JSON block.
//
// Three outcomes:
// - parsed fine → the fields, nil error.
// - output never looked like JSON → ("", "", nil). The caller may ship it
// as-is; small models sometimes answer in bare prose and that is fine.
// - output starts with "{" but does not parse → errBrokenJSON. The grammar
// guarantees a valid *prefix*, so a generation that hits the token cap
// mid-object comes back as a fragment like `{` or `{\n "`. Shipping that
// as a reply is the bug this error exists to stop.
func parseResponseMood(raw string) (response, mood string, err error) {
cleaned := strings.TrimSpace(raw)
start := strings.Index(cleaned, "{")
end := strings.LastIndex(cleaned, "}")
if start < 0 || end < 0 || end <= start {
return "", ""
if strings.HasPrefix(cleaned, "{") {
return "", "", errBrokenJSON
}
return "", "", nil
}
var parsed responseMood
if err := json.Unmarshal([]byte(cleaned[start:end+1]), &parsed); err != nil {
return "", ""
if e := json.Unmarshal([]byte(cleaned[start:end+1]), &parsed); e != nil {
if strings.HasPrefix(cleaned, "{") {
return "", "", errBrokenJSON
}
return "", "", nil
}
return parsed.Response, parsed.Mood
return parsed.Response, parsed.Mood, nil
}
func parsePhrase(raw string) (body, summary string) {
+261
View File
@@ -0,0 +1,261 @@
package phraser
// Hand-written Russian nudges instead of generated ones.
//
// Why: on a nudge there is nothing to be creative about. Measured over many
// runs, Qwen3.5-0.8B breaks the persona (formal "вы", plural imperatives,
// masculine self-reference) and invents facts and units — it once told him to
// boil an egg for "90-95 секунд". A nudge is five words of known content, so
// wording it with a model buys nothing and risks the persona every time.
//
// The wording lives in nudges_ru_v1.json so it can be edited without touching
// Go. This file only picks one and fills in the values.
import (
"context"
_ "embed"
"encoding/json"
"fmt"
"math/rand"
"regexp"
"strings"
"sync"
"time"
"unicode"
"github.com/kami/maven/internal/delivery"
"github.com/kami/maven/internal/loop"
)
//go:embed nudges_ru_v1.json
var nudgeTemplateJSON []byte
// NudgeTemplateSchemaVersion — the version this code understands.
const NudgeTemplateSchemaVersion = 1
type nudgeRuleSet struct {
Mood string `json:"mood"`
Variants []string `json:"variants"`
}
type nudgeTemplateFile struct {
SchemaVersion int `json:"schema_version"`
Name string `json:"name"`
Notes []string `json:"notes"`
Rules map[string]nudgeRuleSet `json:"rules"`
}
// NudgeTemplates picks a hand-written Russian nudge for a candidate.
//
// Safe for concurrent use. Random, but never the same variant twice in a row
// for the same rule — being nagged with identical words is what makes a nudge
// easy to tune out.
type NudgeTemplates struct {
mu sync.Mutex
rnd *rand.Rand
last map[string]string // rule family -> the text used last time
file nudgeTemplateFile
}
// NewNudgeTemplates loads the embedded template file. Pass a source to make the
// picking reproducible in tests; nil means seed from the clock.
func NewNudgeTemplates(src rand.Source) (*NudgeTemplates, error) {
var f nudgeTemplateFile
if err := json.Unmarshal(nudgeTemplateJSON, &f); err != nil {
return nil, fmt.Errorf("nudge templates: parse: %w", err)
}
if f.SchemaVersion != NudgeTemplateSchemaVersion {
return nil, fmt.Errorf("nudge templates: schema_version %d, want %d",
f.SchemaVersion, NudgeTemplateSchemaVersion)
}
if len(f.Rules) == 0 {
return nil, fmt.Errorf("nudge templates: no rules")
}
if src == nil {
src = rand.NewSource(time.Now().UnixNano())
}
return &NudgeTemplates{
rnd: rand.New(src),
last: map[string]string{},
file: f,
}, nil
}
// PhraseNudge implements the nudge half of the Phraser interface, so the
// templates can be scored by the same harness as the model.
func (t *NudgeTemplates) PhraseNudge(_ context.Context, c loop.Candidate) (delivery.PhrasedNudge, error) {
body, mood := t.Nudge(c)
return delivery.PhrasedNudge{Candidate: c, Body: body, Summary: body, Mood: mood}, nil
}
// Nudge returns the text and the mood for one candidate. Never fails: if no
// template fits it uses the plain per-rule fallback.
func (t *NudgeTemplates) Nudge(c loop.Candidate) (body, mood string) {
rule := c.Rule.Name
family := t.family(rule)
set, ok := t.file.Rules[family]
if !ok {
return fallbackNudge(c), "neutral"
}
vals := nudgeValues(c)
// Only variants whose placeholders all have a value.
usable := make([]string, 0, len(set.Variants))
for _, v := range set.Variants {
if text, ok := fillTemplate(v, vals); ok {
usable = append(usable, text)
}
}
if len(usable) == 0 {
return fallbackNudge(c), "neutral"
}
mood = set.Mood
if mood == "" {
mood = "neutral"
}
return t.pick(family, usable), mood
}
// pick chooses at random, skipping whatever this rule said last time.
func (t *NudgeTemplates) pick(family string, usable []string) string {
t.mu.Lock()
defer t.mu.Unlock()
choices := usable
if len(usable) > 1 {
choices = make([]string, 0, len(usable))
for _, v := range usable {
if v != t.last[family] {
choices = append(choices, v)
}
}
if len(choices) == 0 { // every variant equals the last one
choices = usable
}
}
got := choices[t.rnd.Intn(len(choices))]
t.last[family] = got
return got
}
// family maps a rule name to a block in the template file: an exact match
// first, then the prefix of "routine:зарядка" / "morning:утро", then "default".
func (t *NudgeTemplates) family(rule string) string {
if _, ok := t.file.Rules[rule]; ok {
return rule
}
if i := strings.IndexByte(rule, ':'); i > 0 {
if _, ok := t.file.Rules[rule[:i]]; ok {
return rule[:i]
}
}
return "default"
}
// placeholderRE — the {name} slots a template may use.
var placeholderRE = regexp.MustCompile(`\{([a-z]+)\}`)
// nudgeValues collects what this candidate can fill in. A key missing here
// means every template needing it is skipped, so nothing half-filled is ever
// spoken.
func nudgeValues(c loop.Candidate) map[string]string {
vals := map[string]string{}
rule := c.Rule.Name
// {since} — only at hour scale. Below an hour the phrase would be minutes,
// and none of the templates read well with "сорок минут".
if d, ok := c.State.Since(rule); ok && d >= time.Hour {
if s := ruSinceWords(d); s != "" {
vals["since"] = s
}
}
// {service} — the aggregate fact's key carries the service name.
if f, ok := c.State.Fact(rule); ok && f.Key != "" && f.Key != rule {
vals["service"] = f.Key
}
// {what} — the Russian suffix of "routine:таблетки" / "morning:утро".
if i := strings.IndexByte(rule, ':'); i > 0 && i+1 < len(rule) {
vals["what"] = rule[i+1:]
}
return vals
}
// fillTemplate substitutes the placeholders. Returns false when a value is
// missing, so a raw "{since}" can never reach the text-to-speech voice.
func fillTemplate(tmpl string, vals map[string]string) (string, bool) {
missing := false
out := placeholderRE.ReplaceAllStringFunc(tmpl, func(m string) string {
name := m[1 : len(m)-1]
v, ok := vals[name]
if !ok || v == "" {
missing = true
return m
}
return v
})
if missing || strings.ContainsAny(out, "{}%") {
return "", false
}
return capitalizeFirst(out), true
}
// capitalizeFirst — a placeholder can start the sentence, and "полтора часа без
// перерыва" should be spoken as a sentence, not a fragment.
func capitalizeFirst(s string) string {
for i, r := range s {
return string(unicode.ToUpper(r)) + s[i+len(string(r)):]
}
return s
}
// hourWords — hours spelled out. "3 ч" is fine on a screen and wrong in a
// Russian voice, so the number goes out as words.
var hourWords = []string{
"ноль", "один", "два", "три", "четыре", "пять", "шесть", "семь", "восемь",
"девять", "десять", "одиннадцать", "двенадцать", "тринадцать",
"четырнадцать", "пятнадцать", "шестнадцать", "семнадцать", "восемнадцать",
"девятнадцать", "двадцать", "двадцать один", "двадцать два", "двадцать три",
}
// hourPlural — час / часа / часов by Russian counting rules.
func hourPlural(h int) string {
if h%100 >= 11 && h%100 <= 14 {
return "часов"
}
switch h % 10 {
case 1:
return "час"
case 2, 3, 4:
return "часа"
default:
return "часов"
}
}
// ruSinceWords — "полтора часа", "два с половиной часа", "семь часов".
// Empty string means "do not say it" (under an hour, or over a day).
func ruSinceWords(d time.Duration) string {
if d < time.Hour {
return ""
}
h := int(d.Hours())
m := int(d.Minutes()) % 60
if m >= 45 {
h++
m = 0
}
if h >= len(hourWords) {
return "больше суток"
}
if h == 1 {
if m >= 15 {
return "полтора часа"
}
return "час"
}
if m >= 15 {
return hourWords[h] + " с половиной часа"
}
return hourWords[h] + " " + hourPlural(h)
}
+202
View File
@@ -0,0 +1,202 @@
package phraser
import (
"context"
"math/rand"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/store"
)
// cand builds a candidate the way a tick would.
func cand(rule string, sinceMin int, factKey string) loop.Candidate {
now := time.Date(2026, 7, 31, 21, 40, 0, 0, time.UTC)
st := loop.State{Now: now, Facts: map[string]store.Fact{}}
if sinceMin > 0 || factKey != "" {
key := rule
if factKey != "" {
key = factKey
}
st.Facts[rule] = store.Fact{Key: key, Ts: now.Add(-time.Duration(sinceMin) * time.Minute)}
}
return loop.Candidate{Rule: loop.Rule{Name: rule, Severity: loop.Sev1}, Severity: loop.Sev1, State: st}
}
func newTestTemplates(t *testing.T, seed int64) *NudgeTemplates {
t.Helper()
nt, err := NewNudgeTemplates(rand.NewSource(seed))
if err != nil {
t.Fatalf("NewNudgeTemplates: %v", err)
}
return nt
}
func TestNudgeTemplatesLoad(t *testing.T) {
nt := newTestTemplates(t, 1)
for _, rule := range []string{"water", "meal", "break", "service_down", "netdata_critical", "routine", "morning", "default"} {
set, ok := nt.file.Rules[rule]
if !ok {
t.Errorf("no templates for %q", rule)
continue
}
if len(set.Variants) < 5 {
t.Errorf("%s: only %d variants", rule, len(set.Variants))
}
// Every rule needs one variant that needs no value, or a candidate
// without context has nothing to say. routine and morning are exempt:
// they always carry a name and must always say it.
plain := 0
seen := map[string]bool{}
for _, v := range set.Variants {
if !placeholderRE.MatchString(v) {
plain++
}
if seen[v] {
t.Errorf("%s: duplicate variant %q", rule, v)
}
seen[v] = true
}
if plain == 0 && rule != "routine" && rule != "morning" {
t.Errorf("%s: every variant needs a placeholder value", rule)
}
}
}
// The whole point of the picker: never the same words twice in a row.
func TestNudgeNoImmediateRepeat(t *testing.T) {
nt := newTestTemplates(t, 7)
prev := ""
for i := 0; i < 200; i++ {
body, _ := nt.Nudge(cand("water", 200, ""))
if body == prev {
t.Fatalf("repeat at %d: %q", i, body)
}
prev = body
}
}
// Same seed, same sequence — otherwise the fixture score would drift run to run.
func TestNudgeDeterministicWithSeed(t *testing.T) {
var runs [2][]string
for r := range runs {
nt := newTestTemplates(t, 42)
for i := 0; i < 20; i++ {
body, _ := nt.Nudge(cand("break", 100, ""))
runs[r] = append(runs[r], body)
}
}
for i := range runs[0] {
if runs[0][i] != runs[1][i] {
t.Fatalf("run %d differs: %q vs %q", i, runs[0][i], runs[1][i])
}
}
}
// A variant is only used when its value exists, and nothing half-filled ships.
func TestNudgeNoLeftoverPlaceholders(t *testing.T) {
nt := newTestTemplates(t, 3)
cases := []loop.Candidate{
cand("water", 0, ""), // no duration
cand("water", 30, ""), // under an hour
cand("water", 200, ""), // hours
cand("service_down", 3, "vaultwarden"),
cand("service_down", 3, ""), // no service name
cand("routine:таблетки", 0, ""),
cand("morning:утро", 0, ""),
cand("unknown_rule", 0, ""),
}
for _, c := range cases {
for i := 0; i < 40; i++ {
body, mood := nt.Nudge(c)
if body == "" {
t.Fatalf("%s: empty body", c.Rule.Name)
}
if strings.ContainsAny(body, "{}%") {
t.Fatalf("%s: unfilled template %q", c.Rule.Name, body)
}
if mood != "neutral" {
t.Fatalf("%s: mood %q", c.Rule.Name, mood)
}
}
}
}
// The routine name must actually land in the text.
func TestNudgeSubstitutesWhat(t *testing.T) {
nt := newTestTemplates(t, 11)
for i := 0; i < 40; i++ {
body, _ := nt.Nudge(cand("routine:таблетки", 0, ""))
if !strings.Contains(strings.ToLower(body), "таблетки") {
t.Fatalf("routine text lost the name: %q", body)
}
}
}
func TestRuSinceWords(t *testing.T) {
cases := []struct {
min int
want string
}{
{30, ""},
{60, "час"},
{95, "полтора часа"},
{150, "два с половиной часа"},
{190, "три часа"},
{240, "четыре часа"},
{430, "семь часов"},
{660, "одиннадцать часов"},
{60 * 30, "больше суток"},
}
for _, c := range cases {
got := ruSinceWords(time.Duration(c.min) * time.Minute)
if got != c.want {
t.Errorf("%d min: got %q want %q", c.min, got, c.want)
}
}
}
// Templates are the default: a nudge must not reach the model at all.
func TestLLMPhraserUsesTemplatesByDefault(t *testing.T) {
spy := newGrammarSpy(t)
p := NewLLMPhraserAt(spy.srv.URL, Config{})
pn, err := p.PhraseNudge(context.Background(), cand("water", 200, ""))
if err != nil {
t.Fatalf("PhraseNudge: %v", err)
}
if len(spy.grammars) != 0 {
t.Errorf("nudge hit the model %d times, want 0", len(spy.grammars))
}
if !strings.Contains(strings.ToLower(pn.Body), "вод") {
t.Errorf("nudge is not the water template: %q", pn.Body)
}
}
// ...and the flag brings the model back.
func TestLLMNudgesFlagRestoresTheModel(t *testing.T) {
spy := newGrammarSpy(t)
p := NewLLMPhraserAt(spy.srv.URL, Config{LLMNudges: true})
pn, err := p.PhraseNudge(context.Background(), cand("water", 200, ""))
if err != nil {
t.Fatalf("PhraseNudge: %v", err)
}
if len(spy.grammars) != 1 {
t.Fatalf("nudge hit the model %d times, want 1", len(spy.grammars))
}
if pn.Body != "ага" {
t.Errorf("body = %q, want the model's reply", pn.Body)
}
}
func TestNudgeTemplatesPhraseNudge(t *testing.T) {
nt := newTestTemplates(t, 5)
pn, err := nt.PhraseNudge(context.Background(), cand("water", 200, ""))
if err != nil {
t.Fatalf("PhraseNudge: %v", err)
}
if pn.Body == "" || pn.Summary != pn.Body || pn.Mood != "neutral" {
t.Fatalf("bad nudge: %+v", pn)
}
}
+129
View File
@@ -0,0 +1,129 @@
{
"schema_version": 1,
"name": "russian nudge templates v1",
"notes": [
"Hand-written Russian nudges. Edit the wording here, no Go changes needed.",
"Rules: she is feminine about herself, he is a man addressed as ты. Never вы/вас/ваш, never plural imperatives (выпейте), never он/его about him.",
"One short sentence. No questions, no emoji, no pet names, no emotional support.",
"Placeholders: {since} how long it has been (only used when it is at least an hour), {service} the service name, {what} the routine name. A variant whose placeholder has no value is skipped, so every rule needs at least one variant with no placeholder. The exception is routine and morning: those only exist for rules like routine:таблетки that always carry a name, and a routine nudge that drops the name is useless.",
"mood must be one of: neutral, happy, thinking, tired, confused."
],
"rules": {
"water": {
"mood": "neutral",
"variants": [
"Ты не пил воду {since} — выпей стакан.",
"Пора выпить воды.",
"Стакан воды не помешает.",
"Воду ты не пил уже {since}.",
"Напоминаю про воду.",
"Сходи за водой, дела подождут.",
"Сделай глоток воды, пока помнишь.",
"Между делом выпей воды.",
"Вода — простое дело: выпей стакан.",
"Отвлекись на стакан воды."
]
},
"meal": {
"mood": "neutral",
"variants": [
"Ты не ел {since} — поешь.",
"Пора поесть, сделай перекус.",
"Еда важнее ещё одного часа за столом.",
"Без еды уже {since}, поешь.",
"Напоминаю про еду — поешь.",
"Возьми перерыв на обед.",
"Сделай себе перекус, это пять минут.",
"Поешь, потом вернёшься к работе.",
"Поешь нормально, а не на ходу.",
"Еды не было {since} — разогрей что-нибудь."
]
},
"break": {
"mood": "neutral",
"variants": [
"Ты за столом {since} — встань и разомнись.",
"Пора сделать перерыв.",
"Встань на пять минут.",
"{since} без перерыва — отойди от экрана.",
"Напоминаю про перерыв.",
"Разомни спину, потом продолжишь.",
"Короткая пауза не сорвёт дела.",
"Отойди от компьютера на минуту.",
"Сидишь без перерыва {since}.",
"Встань, пройдись, вернись."
]
},
"service_down": {
"mood": "neutral",
"variants": [
"Сервис {service} не отвечает.",
"{service} упал — сервис не отвечает.",
"{service} не отвечает, сервис нужно поднимать.",
"Сервис {service} недоступен.",
"Проверь {service}: сервис не отвечает.",
"Сервис перестал отвечать.",
"Сервис {service} лежит, нужно смотреть.",
"{service} не отвечает уже {since}.",
"Мониторинг сообщает: {service} лежит.",
"Сервис {service} не отвечает, посмотри логи."
]
},
"netdata_critical": {
"mood": "neutral",
"variants": [
"Netdata: критический алярм, проверь диск.",
"Критический алярм в netdata — посмотри диск.",
"Netdata поднял тревогу по диску.",
"Проверь диск: netdata ругается.",
"Алярм от netdata, критический.",
"Netdata: критический уровень, дело в диске.",
"Диск требует внимания — критический алярм в netdata.",
"Критический алярм: проверь место на диске.",
"Netdata сообщает о критической проблеме с диском.",
"Открой netdata: там критический алярм по диску."
]
},
"routine": {
"mood": "neutral",
"variants": [
"По распорядку: {what}.",
"Пора — {what}.",
"Напоминаю: {what}.",
"В списке на сейчас: {what}.",
"{what} — сейчас самое время.",
"Не пропусти: {what}.",
"{what}: пора сделать.",
"Сейчас по плану {what}.",
"Твой распорядок: {what}.",
"{what} — по распорядку сейчас."
]
},
"morning": {
"mood": "neutral",
"variants": [
"{what} — пора начать день.",
"{what}: пройди утренний список.",
"Начни {what} со списка.",
"{what}. Осталось пройти чеклист.",
"Утренний список ещё не пройден: {what}.",
"{what}: первый пункт списка за тобой.",
"{what} идёт, а список стоит.",
"{what}: не забудь про утренние дела.",
"По утреннему чеклисту ещё есть дела: {what}.",
"{what} — утренний список дел ещё ждёт."
]
},
"default": {
"mood": "neutral",
"variants": [
"Напоминаю: есть дело.",
"Пора вернуться к отложенному делу.",
"Одно дело ждёт тебя.",
"Напоминаю про дело из списка.",
"В списке осталось дело.",
"Дело всё ещё не сделано."
]
}
}
}