Compare commits

..

19 Commits

Author SHA1 Message Date
kami 93c1a41d4a Route with the resident model by default
The two things that made this unsafe are fixed: the router can now
refuse, and slot extraction runs on its decisions.

On the held-out fixture it gets 63.2% of intents right against the
classifier's 50.0%, with no route errors. It costs about a second a
turn instead of 30ms.

The flag is a pointer now, so leaving it out of the config means on
and only writing false turns it off.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:47:57 +04:00
kami bce5ed210c Merge branch 'worktree-agent-a76ce40c73601d90d' into overnight-jul31 2026-07-31 11:44:32 +04:00
kami c31f0d1001 Extract slots for LLM router decisions too
An LLM-routed reminder came back with no parsed time and an act with no
fn, because only the classifier path ran the extractor. Now the router
runs the same extraction after an LLM decision and fills only the empty
slots. No time in the utterance still means no time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:44:06 +04:00
kami 34521c30b8 Merge branch 'worktree-agent-ad5da57e47b822152' into overnight-jul31 2026-07-31 11:39:35 +04:00
kami 1d48755d12 Record the recall numbers after the embedder swap
recall@1 60% to 72%, answered 48% to 72%, latency 3x better. But false
recall went 1/5 to 5/5: e5 packs every score into a narrow high band, so
the 0.55 gate now admits everything. Left the gate alone as instructed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:38:18 +04:00
kami f6d5a2a7a4 Swap the embedder to multilingual-e5-small (Vikunja #371, #372)
The old model was a symmetric paraphrase model, so it scored "do these
look alike" instead of "does this note answer this question". Also fixes
the file mismatch: the Makefile, the deploy config and both evals now all
name the same quantized file, and the quantized one is what gets measured.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:38:18 +04:00
kami 751c2a705f Embed a question and a stored note differently (Vikunja #371)
Note recall is asymmetric: a short question goes in, a longer note comes
out. Adds EmbedQuery/EmbedPassage helpers and the e5 prefixes, and points
the note/fact write path at the passage side and the query path at the
query side. Reviewers: the three call sites in voice.go.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:38:07 +04:00
kami 5d5b0cfd49 Merge branch 'worktree-agent-a0ab2a9b7439296e3' into overnight-jul31 2026-07-31 11:37:53 +04:00
kami bd16ca69e5 Let the LLM router answer "unknown" when it cannot route
Chose an 8th enum value over a confidence number: the model already picks
one enum token, so it costs nothing in the grammar, while a score from a
0.8B model would be uncalibrated noise. A refusal returns "no decision"
with no error, which is the fall-through the caller already uses for a
bad parse, so the classifier and its clarify gate take the turn.

Reviewers: the prompt's counter-examples matter most — a small model will
over-use any easy escape hatch. The training workspace copy of the prompt
still needs the same edit (Vikunja #362).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:35:37 +04:00
kami 0ed386eca6 Re-measure the router on a quiet box and record the numbers
The earlier before/after was taken while another eval shared
llama-server. This run had the box to itself.

Intent accuracy 61.8% llm-only, 63.2% cascade, 67.1% with thinking off.
The prompt fix holds. note→fact shows up here too, so it is real.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 11:34:57 +04:00
kami 1c4eab2107 Merge commit '94eb92f' into overnight-jul31
# Conflicts:
#	Makefile
2026-07-31 10:11:56 +04:00
kami 0914e0a3d5 Merge commit '4ba9a6f' into overnight-jul31
# Conflicts:
#	Makefile
2026-07-31 10:11:29 +04:00
kami c860808528 Make make test actually gate on gofmt and vet
DESIGN.md has always said `make test` is "gofmt + vet + -race, no
exceptions". It only ever ran the tests, which is how nine files drifted
out of format without anyone noticing.

`test` now depends on `fmt-check` and `vet`. Checked that fmt-check does
fail when a file is unformatted, so the gate is real and not decorative.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 10:10:27 +04:00
kami f7442c3aea Run gofmt over the seven files that had drifted
Formatting only: import order, and statements that were packed onto one
line split out. `git diff -w` shows nothing but that.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 10:09:32 +04:00
kami 75b067ac51 Merge the accepted-routine fix and drop reminder_id from accept
Two merge fixes on top of the branch:

- migrations: keep both new steps, snooze stays #8, the routine columns
  become #9. Both agents had numbered theirs #8.
- accepting no longer takes a reminder id, on the web surface too. The
  web accept path had the same one-shot-reminder bug the voice path did,
  so both now just flip the status and let the tick loop schedule.

The test that asserted "accept creates a reminder and links it" asserted
the bug. It now asserts that accepting creates no reminder.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 10:07:37 +04:00
kami 424d1b3446 Fire accepted routines every interval, not once (Vikunja #366)
The tick loop now reads accepted routines from the store and nudges when
their interval has passed; accepting no longer builds a one-shot reminder.
Look at routine.DueAccepted for the schedule rule (no catch-up backlog) and
at fireAcceptedRoutines for the restraint gate — routines do not bypass it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:45:30 +04:00
kami 54dc43516b Add accepted-routine timestamps to the store (Vikunja #366)
Data layer only. Migration #8 adds accepted_ts and last_fired_ts to
proposed_routines, plus ListAcceptedRoutines and MarkRoutineFired so the
tick loop can own the schedule. Accepting no longer links a reminder id.
Look at the TODO(vikunja#366) in cmd/mavend/tick.go for the next commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:41:58 +04:00
kami 4ba9a6f422 Add a deterministic scorer for nudge phrasing (Vikunja #323)
Review internal/phraser/eval/checks.go -- it IS the measurement. Each check
names in a comment which DESIGN.md line it defends: length, feminine
self-reference (windowed around "я" so the operator's own masculine
second-person forms are not flagged), the cringe list (pet names, emoji,
"!!", fake concern, apology, emotional support, asking how he feels,
praise), on-topic, mood enum. No send/veto signal anywhere, per
DESIGN.md § "Rules decide, LLM phrases".
Fixture (158 lines) and tests (252) do not count toward the diff ceiling;
the scorer itself is still ~650. Splitting eval.go from checks.go would
give two commits neither of which measures anything.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:30:52 +04:00
kami 94eb92fb15 Label LLM eval runs with the model llama-server has loaded
The bake-off in #278/#250 needs two models' scores side by side, and the
report names only carried the config, so the rows were indistinguishable.
ModelID reads /v1/models instead of taking a string that goes stale.
New target: make eval-models MAVEN_LLM_URL=...

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:17:20 +04:00
47 changed files with 2139 additions and 186 deletions
+8 -5
View File
@@ -43,14 +43,17 @@ notes. Without it, the floor `HashEmbedder` is used — deterministic but weak
(Russian recall rarely clears the confidence gate, many commands fall to
"clarify").
**Download the embedder** (ONNX, ~90 MB):
**Download the embedder** (ONNX, ~120 MB):
```sh
make download-embedder
```
This fetches `paraphrase-multilingual-MiniLM-L12-v2` (384-dim, 12-layer,
supports 50+ languages including Russian) to `models/embedder/`.
This fetches `multilingual-e5-small` (384-dim, 12-layer, Russian and English)
to `models/embedder/multilingual-e5-small/`. It is an asymmetric retrieval
model: the code puts `query: ` in front of a question and `passage: ` in front
of a stored note, which is how e5 was trained. The quantized file is the one
that is downloaded, deployed and measured.
**Also need ONNX Runtime** (`libonnxruntime.so`):
@@ -64,8 +67,8 @@ sudo cp onnxruntime-linux-x64-1.15.1/lib/libonnxruntime.so* /usr/local/lib/
```json
"voice": {
"embedder": {
"model_path": "models/embedder/model_quantized.onnx",
"tokenizer_path": "models/embedder/tokenizer.json",
"model_path": "models/embedder/multilingual-e5-small/model_quantized.onnx",
"tokenizer_path": "models/embedder/multilingual-e5-small/tokenizer.json",
"lib_path": "/usr/local/lib/libonnxruntime.so"
}
}
+46 -5
View File
@@ -16,7 +16,7 @@ PIPER_BIN := $(shell pwd)/deps/piper/piper
PIPER_MODEL := $(shell pwd)/models/tts/ru_RU-irina-medium.onnx
PIPER_ESPEAK := $(shell pwd)/deps/piper/espeak-ng-data
.PHONY: all build build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav clean test run-stt run-tts run-web download-embedder deps-go eval-router eval-recall
.PHONY: all build build-stt build-tts build-daemon build-client build-waked build-web build-poll build-caldav clean test fmt-check vet run-stt run-tts run-web download-embedder deps-go eval-router eval-recall eval-phrasing eval-models
all: build
@@ -69,7 +69,20 @@ deps-go:
done
$(GO) version
test:
# fmt-check fails if any file needs gofmt. DESIGN.md has always said `make
# test` gates on gofmt and vet; it did not, so nine files quietly drifted.
# Run `gofmt -w` on whatever this prints.
fmt-check:
@bad=$$(gofmt -l internal cmd); \
if [ -n "$$bad" ]; then \
echo "these files need gofmt:"; echo "$$bad"; exit 1; \
fi
vet:
CGO_CFLAGS="$(CGO_CFLAGS)" CGO_LDFLAGS="$(CGO_LDFLAGS)" LD_LIBRARY_PATH="$(shell pwd)/deps/lib" \
$(GO) vet ./internal/... ./cmd/...
test: fmt-check vet
CGO_CFLAGS="$(CGO_CFLAGS)" CGO_LDFLAGS="$(CGO_LDFLAGS)" LD_LIBRARY_PATH="$(shell pwd)/deps/lib" \
$(GO) test -race -coverprofile=coverage.out ./internal/... ./cmd/...
@@ -90,6 +103,30 @@ eval-router:
eval-recall:
MAVEN_ONNX_LIB="$(MAVEN_ONNX_LIB)" $(GO) test -v -count=1 ./internal/memory/recalleval/
# eval-phrasing -- score nudge phrasing (internal/phraser/eval). Verbose so the
# report and every generated message land in the terminal. With no environment
# it scores the deterministic Stub only, which is what CI runs. Set
# MAVEN_LLM_URL to add the resident model:
# MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-phrasing
# The model run is slow (minutes) -- the timeout is raised to match.
eval-phrasing:
$(GO) test -v -count=1 -timeout 40m ./internal/phraser/eval/
# eval-models — score ONE llama-server against the same fixture, for the
# resident-model bake-off (#278, #250). Start a server with the gguf you want,
# then:
#
# make eval-models MAVEN_LLM_URL=http://127.0.0.1:18100
#
# The report names carry the model llama-server reports, so runs from two
# checkpoints stay apart. Only the LLM test runs — the classifier baselines do
# not depend on the model and take the ONNX runtime with them.
MAVEN_LLM_URL ?= http://127.0.0.1:18099
eval-models:
MAVEN_LLM_URL="$(MAVEN_LLM_URL)" $(GO) test -v -count=1 -timeout 60m \
-run TestLLMRouterBaseline ./internal/router/eval/
run-stt: build-stt
LD_LIBRARY_PATH="$(shell pwd)/deps/lib" \
./mavsttd -socket /tmp/maven/stt.sock -model $(WHISPER_MODEL)
@@ -115,9 +152,13 @@ deps-piper:
-o /tmp/piper.tar.gz
tar -xzf /tmp/piper.tar.gz -C deps/
EMBEDDER_DIR := $(shell pwd)/models/embedder
EMBEDDER_MODEL_URL := https://huggingface.co/Xenova/paraphrase-multilingual-MiniLM-L12-v2/resolve/main/onnx/model_quantized.onnx
EMBEDDER_TOKENIZER_URL := https://huggingface.co/Xenova/paraphrase-multilingual-MiniLM-L12-v2/resolve/main/tokenizer.json
# multilingual-e5-small: an asymmetric retrieval model. It is trained to match
# a short question against a longer passage, which is what note recall is.
# The quantized file is the one we download, deploy and measure — see
# RECALL-EVAL-31-07-2026.md.
EMBEDDER_DIR := $(shell pwd)/models/embedder/multilingual-e5-small
EMBEDDER_MODEL_URL := https://huggingface.co/Xenova/multilingual-e5-small/resolve/main/onnx/model_quantized.onnx
EMBEDDER_TOKENIZER_URL := https://huggingface.co/Xenova/multilingual-e5-small/resolve/main/tokenizer.json
download-embedder:
mkdir -p $(EMBEDDER_DIR)
+60
View File
@@ -83,6 +83,66 @@ sqlite-backed `store.MemoryStore` and `memory.InMemoryStore` identically — bot
(`internal/store/memory.go:64`) at ~150µs over 42 rows against a ~59ms query embed. An ANN index is
not the problem to solve.
## Re-measured after the embedder swap — 31-07-2026, later the same day
Changed: `models/embedder/` is now **multilingual-e5-small** (quantized, 118MB), with `query: ` in
front of a question and `passage: ` in front of a stored note (Vikunja #371). `deploy/mavend.json`
and `make download-embedder` now name the same file, and it is the quantized one — that is what the
column below measures (Vikunja #372). Everything else is unchanged: same fixture, same store, same
0.55 gate. The old column is the baseline and is left as it was.
| | recall+onnx, MiniLM (baseline) | recall+onnx, e5-small (new) |
|---|---|---|
| **recall@1** | 60.0% (15/25) | **72.0% (18/25)** |
| recall@3 | 80.0% (20/25) | 84.0% (21/25) |
| **answered after the 0.55 gate** | 48.0% (12/25) | **72.0% (18/25)** |
| wrong note on top / tie on top | 10 / 0 | 7 / 0 |
| ranked first, then silenced by the gate | 3 | 0 |
| **false recall** | 1/5 (20%) | **5/5 (100%)** |
| top-1 score when right, min / median | 0.559 / 0.678 | 0.791 / 0.857 |
| top-1 when it must stay silent, median / max | 0.470 / 0.567 | 0.815 / 0.835 |
| RU / EN / `hard` cases passed | 13/24 / 3/6 / 2/11 | 14/24 / 4/6 / 5/11 |
| latency p50 / p95 / max | 59ms / 148ms / 194ms | 18ms / 37ms / 49ms |
### What moved
Ranking got better and got faster. Half the previously-unwinnable `hard` cases now pass (2/11 →
5/11), the guitar note no longer beats the docker-logs note, and the gate stops silencing notes that
already ranked first. The quantized e5 is also ~3x quicker than the fp32 MiniLM it replaces.
### What got worse: the gate is now a no-op
e5 packs every cosine into a narrow high band. Right-note scores start at 0.791; must-stay-silent
scores reach 0.835. **The distributions still overlap, and now they overlap above the gate**, so
0.55 admits everything and false recall goes from 1/5 to 5/5. The sweep:
```
gate 0.500.70: answered 18/25 (72%) false recall 5/5
gate 0.80: answered 17/25 (68%) false recall 4/5
gate 0.90: answered 0/25 ( 0%) false recall 0/5
```
There is no value that keeps real recall and rejects made-up questions — same conclusion as before,
now with a wider band and no room at all. `query_min_score` was left at 0.55 as instructed. **The
recommendation is to leave it there and stop tuning it**: any number under ~0.79 is a no-op and
anything above starts cutting real recall long before it stops the false ones. The fix is a margin
gate (`top1 top2 > δ`), next-steps item 3, which is now the top item.
### The prefixes did not do the work
A control run with both prefixes set to the empty string scored the **same** recall@1 (72%), a
slightly better recall@3 (88%) and the same 5/5 false recall. So on this fixture the gain comes from
the model, not from the `query:` / `passage:` split. The prefixes are kept because they are how e5
was trained and the split is the right shape for the read path, but they are not worth defending on
this evidence — a bigger fixture may say otherwise.
### Stored vectors from the old model are now junk
Cosine between a MiniLM vector and an e5 vector means nothing. Every row already in `notes` and in
the vector memory table was written by the old model, so after this deploy they will score as noise
against a new query. A live database needs every note and fact re-embedded before recall works at
all. Filed as its own task.
## Next steps — ordered by value-to-risk; nothing here is a decision
1. **Swap the embedder to `multilingual-e5-small` with `query:`/`passage:` prefixes.** One config
+35
View File
@@ -40,6 +40,41 @@ model → classifier as failure floor.
Never compare a hash-embedder run to an ONNX one.
## Re-measured after the prompt fix
The table above is the **baseline at commit `46259b4`**, kept as-is. The prompt fix (query
tested before fact, plus `repeat_penalty` and a bounded grammar string) was then measured on
an otherwise idle box — no other eval sharing llama-server, so these latencies are real
rather than contention.
| | llm-only (0.8B) | cascade+llm (0.8B) | llm-only, thinking off |
|---|---|---|---|
| **intent-only accuracy** | 48.7% → **61.8%** | 50.0% → **63.2%** | **67.1%** |
| full accuracy (intent+slots+gate) | 23.7% → **38.2%** | 32.9% → **47.4%** | **42.1%** |
| route errors | 2 → **0** | 0 → 0 | **0** |
| p50 / p95 latency | **1.08s / 1.55s** | **1.04s / 1.53s** | **0.93s / 1.41s** |
Three things this run settles:
1. **The prompt fix holds.** An earlier contended run reported 60.5% / 36.8% for llm-only;
the quiet run gives 61.8% / 38.2%. Close enough to call the gain real, and the earlier
run's 4-5s latency figures were contention, not the model.
2. **`query→fact` fell from ×15 to ×7**, and both unparseable replies are gone. Zero route
errors in every LLM configuration.
3. **`note→fact ×4` is real, not noise.** It shows up in the quiet run too. The agent that
wrote the prompt fix suspected its own change might have caused it by pulling assertive
`запиши что…` phrasings toward fact, and that suspicion stands — all five `ru-note-*`
cases now land on fact. Tracked as Vikunja #375.
**Thinking off is the best configuration measured so far**, on both accuracy and latency
(Vikunja #376). That is worth understanding before flipping: routing is a short
classification into a fixed enum with grammar-constrained output, so there is little to
reason about, and the thinking trace mostly gives a small model room to talk itself out of
the right answer. Phrasing is a different job and needs measuring separately.
Still `6 / 6` missed clarify — the router has no way to say "I don't know" (Vikunja #359).
That is unchanged by anything here.
## Findings
### 1. The resident model does route better — 50.0% vs 36.8%
+12 -12
View File
@@ -57,10 +57,10 @@ import (
"github.com/kami/maven/internal/delivery/ntfysink"
"github.com/kami/maven/internal/delivery/telegramsink"
"github.com/kami/maven/internal/ipc"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/phraser"
"github.com/kami/maven/internal/store"
"github.com/kami/maven/internal/webauthn"
"github.com/kami/maven/internal/loop"
)
var errLocked = errors.New("mavend: daemon locked — complete passkey assertion first")
@@ -161,12 +161,12 @@ func (l *lockedAPI) EnableTool(ctx context.Context, name string, cmd []string, d
return errLocked
}
func (l *lockedAPI) DisableTool(ctx context.Context, name string) error { return errLocked }
func (l *lockedAPI) DeleteTool(ctx context.Context, name string) error { return errLocked }
func (l *lockedAPI) DeleteTool(ctx context.Context, name string) error { return errLocked }
func (l *lockedAPI) ListProposedRoutines(ctx context.Context) ([]ipc.ProposedRoutine, error) {
return nil, errLocked
}
func (l *lockedAPI) DismissProposedRoutine(ctx context.Context, id int64) error { return errLocked }
func (l *lockedAPI) AcceptProposedRoutine(ctx context.Context, id, remID int64) error {
func (l *lockedAPI) AcceptProposedRoutine(ctx context.Context, id int64) error {
return errLocked
}
func (l *lockedAPI) LookupTool(ctx context.Context, name string) (ipc.Tool, error) {
@@ -247,15 +247,15 @@ func run(args []string) error {
// ----- daemon components (only wired when unlocked) -----
// Pre-declare so the unlock path can wire them later.
var (
gatherer *loop.Gatherer
rules []loop.Rule
phr phraser.Phraser
voiceW *voiceWiring
dispatcher *delivery.Dispatcher
tl *tickLoop
coreAPI ipc.CoreAPI
eco *ecosystemWiring
factWorker *factEnrichmentWorker
gatherer *loop.Gatherer
rules []loop.Rule
phr phraser.Phraser
voiceW *voiceWiring
dispatcher *delivery.Dispatcher
tl *tickLoop
coreAPI ipc.CoreAPI
eco *ecosystemWiring
factWorker *factEnrichmentWorker
)
if !locked {
+4 -1
View File
@@ -9,7 +9,10 @@ import (
"github.com/kami/maven/internal/voice"
)
type mockCompleter struct{ out string; err error }
type mockCompleter struct {
out string
err error
}
func (m mockCompleter) Complete(_ context.Context, _ llm.Req) (string, error) { return m.out, m.err }
+62
View File
@@ -186,6 +186,10 @@ func (t *tickLoop) tick(ctx context.Context, now time.Time) {
// LLM-phrased — so a routine can't hallucinate. severity comes from config.
t.fireRoutines(ctx, now, state)
// accepted routines: patterns the user confirmed. read straight from the
// store each tick so the schedule survives a restart.
t.fireAcceptedRoutines(ctx, now, state)
// morning routines: daily checklists (medicine/water/pets/...), nagged at
// most once per day per routine, and only for items still unevidenced at
// nudge time. See internal/morning for the "why not four timers" rationale.
@@ -377,6 +381,64 @@ func (t *tickLoop) fireRoutines(ctx context.Context, now time.Time, state loop.S
}
}
// fireAcceptedRoutines nudges about the routines the user accepted, once per
// interval (Vikunja #366). Accepting used to create a single reminder, so a
// non-weekly routine fired once and went quiet forever; the schedule lives in
// the proposed_routines row now and the loop re-reads it every tick.
//
// A routine is a care-class nudge and goes through the restraint gate like any
// other: quiet hours, away presence and snooze all suppress it. Reminders bypass
// that gate; routines must not. A suppressed nudge is NOT marked fired, so it
// goes out on the next tick that the gate allows — one nudge, held, not dropped
// and not repeated.
//
// The body is literal text built from the detected action and object, not
// LLM-phrased, so a routine can't hallucinate. It nudges; it never acts.
func (t *tickLoop) fireAcceptedRoutines(ctx context.Context, now time.Time, state loop.State) {
rows, err := t.store.ListAcceptedRoutines(ctx)
if err != nil {
log.Printf("tick: list accepted routines: %v", err)
return
}
accepted := make([]routine.Accepted, 0, len(rows))
for _, r := range rows {
if r.AcceptedTs == nil {
continue // accepted before the schedule column existed — no clock to start from.
}
accepted = append(accepted, routine.Accepted{
ID: r.ID,
Name: r.Action + " " + r.Object,
IntervalDays: r.IntervalDays,
Accepted: *r.AcceptedTs,
LastFired: r.LastFiredTs,
})
}
for _, a := range routine.DueAccepted(accepted, now) {
rule := loop.Rule{Name: "routine:" + a.Name, Severity: loop.Sev1}
if !loop.Gate(state, rule) {
continue
}
body := "пора: " + a.Name
pn := delivery.PhrasedNudge{
Candidate: loop.Candidate{Rule: rule, Severity: rule.Severity, State: state},
Body: body,
Summary: body,
}
sent, err := t.dispatcher.DispatchNudge(ctx, pn, now)
if err != nil {
log.Printf("tick: dispatch accepted routine %d: %v", a.ID, err)
continue
}
if len(sent) == 0 {
continue // routing dropped it — leave it due.
}
if err := t.store.MarkRoutineFired(ctx, a.ID, now); err != nil {
log.Printf("tick: mark routine %d fired: %v", a.ID, err)
}
}
}
// fireMorningRoutines checks each configured checklist against today's facts
// and dispatches a nag listing exactly what's still missing, at most once per
// routine per calendar day. Fact reads happen here (not in loop.Gatherer)
+109
View File
@@ -96,6 +96,115 @@ func TestTickFiresRoutineWhenScheduleCrosses(t *testing.T) {
}
}
// TestTickFiresAcceptedRoutineEveryInterval — Vikunja #366. An accepted routine
// with a 3-day interval must nudge every 3 days, not once. It also must not
// replay the occurrences it slept through: after a 30-day gap it nudges once.
func TestTickFiresAcceptedRoutineEveryInterval(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
accepted := refNow()
id, err := st.CreateProposedRoutine(ctx, "полить", "цветы", 3.0, accepted)
if err != nil {
t.Fatalf("CreateProposedRoutine: %v", err)
}
if err := st.AcceptProposedRoutine(ctx, id, accepted); err != nil {
t.Fatalf("AcceptProposedRoutine: %v", err)
}
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
const rule = "routine:полить цветы"
// Same day as the accept: not due yet.
markPresent(t, st, ctx, accepted)
tl.tick(ctx, accepted.Add(time.Hour))
if n := countSends(sink, rule); n != 0 {
t.Fatalf("routine fired %d times before its first interval passed, want 0", n)
}
// Three days later: the first nudge.
first := accepted.Add(3 * 24 * time.Hour)
markPresent(t, st, ctx, first)
tl.tick(ctx, first)
if n := countSends(sink, rule); n != 1 {
t.Fatalf("first interval: sends = %d, want 1", n)
}
// Next day: still inside the interval, silent.
sink.sends = nil
markPresent(t, st, ctx, first.Add(24*time.Hour))
tl.tick(ctx, first.Add(24*time.Hour))
if n := countSends(sink, rule); n != 0 {
t.Fatalf("mid-interval: sends = %d, want 0", n)
}
// Three days after the first nudge: it fires again. This is the bug —
// a one-shot reminder would never come back.
second := first.Add(3 * 24 * time.Hour)
markPresent(t, st, ctx, second)
tl.tick(ctx, second)
if n := countSends(sink, rule); n != 1 {
t.Fatalf("second interval: sends = %d, want 1 (a routine repeats)", n)
}
// A long silence must not turn into a backlog of missed nudges.
sink.sends = nil
late := second.Add(30 * 24 * time.Hour)
markPresent(t, st, ctx, late)
tl.tick(ctx, late)
if n := countSends(sink, rule); n != 1 {
t.Fatalf("after a 30-day gap: sends = %d, want exactly 1 (no backlog)", n)
}
}
// TestTickAcceptedRoutineRespectsQuietHours — routines are not reminders: they
// do not inherit the reminder gate bypass. Away presence drops a care-class
// nudge, and the routine stays due so it nudges once the user is back.
func TestTickAcceptedRoutineRespectsGate(t *testing.T) {
st := newTestStore(t)
ctx := context.Background()
accepted := refNow()
id, err := st.CreateProposedRoutine(ctx, "полить", "цветы", 3.0, accepted)
if err != nil {
t.Fatalf("CreateProposedRoutine: %v", err)
}
if err := st.AcceptProposedRoutine(ctx, id, accepted); err != nil {
t.Fatalf("AcceptProposedRoutine: %v", err)
}
sink := &fakeSink{}
tl := newTestTickLoop(t, st, sink, nil)
const rule = "routine:полить цветы"
// No presence probes at all ⇒ away ⇒ the care gate blocks the nudge.
due := accepted.Add(3 * 24 * time.Hour)
tl.tick(ctx, due)
if n := countSends(sink, rule); n != 0 {
t.Fatalf("away: sends = %d, want 0 (routine must not bypass the gate)", n)
}
// Back at the desk a minute later: the nudge that was held now goes out.
back := due.Add(time.Minute)
markPresent(t, st, ctx, back)
tl.tick(ctx, back)
if n := countSends(sink, rule); n != 1 {
t.Fatalf("present again: sends = %d, want 1", n)
}
}
// countSends counts captured sends for one rule name.
func countSends(sink *fakeSink, rule string) int {
n := 0
for _, s := range sink.sends {
if s.RuleName == rule {
n++
}
}
return n
}
// refNow — fixed tick time so presence decay + since durations are deterministic.
func refNow() time.Time { return time.Date(2026, 6, 30, 12, 0, 0, 0, time.UTC) }
+13 -24
View File
@@ -206,11 +206,11 @@ func wireVoice(cfg *config.Config, coreAPI ipc.CoreAPI, phr phraser.Phraser, mem
if threshold <= 0 {
threshold = config.DefaultRouterThreshold
}
// Both routing paths are weak on held-out utterances — the classifier gets
// 36.8% of intents right, the resident model 50.0% and much slower. Off by
// default (see config.VoiceConfig.LLMRouter); the classifier always stays
// wired as the fallback, so a model error never breaks a turn.
rtr := buildRouter(emb, matcher, threshold, pickLLMRouter(cfg.Voice.LLMRouter, llmClient))
// The resident model routes by default: 63.2% of held-out intents right
// against the classifier's 50.0%, at about 1s a turn instead of 30ms (see
// config.VoiceConfig.LLMRouter). The classifier always stays wired as the
// fallback, so a model error never breaks a turn.
rtr := buildRouter(emb, matcher, threshold, pickLLMRouter(cfg.Voice.UseLLMRouter(), llmClient))
// ----- sessions registry (shared with voicesink) -----
sessions := voice.NewSessions()
@@ -559,7 +559,7 @@ func (h *reactiveHandler) applyAction(ctx context.Context, dec router.Decision)
// fail the fact write). Facts aren't in the notes table, so this is the
// only recall path for them — "когда я пил воду?" reads back from here.
if h.memStore != nil {
if vec, err := h.embedder.Embed(ctx, dec.Utterance); err != nil {
if vec, err := router.EmbedPassage(ctx, h.embedder, dec.Utterance); err != nil {
log.Printf("voice: embed fact for memory: %v", err)
} else if err := h.memStore.Insert(ctx, "fact:"+dec.Slots.Key+":"+strconv.FormatInt(now.Unix(), 10), vec, map[string]string{
"source": "voice",
@@ -673,7 +673,7 @@ func (h *reactiveHandler) applyAction(ctx context.Context, dec router.Decision)
// embed the note text with the same model the classifier uses, persist
// via CoreAPI (source=tap:voice). Semantic recall lives in `notes`, not
// facts — no predicate reads it (spec's two-memory split).
vec, err := h.embedder.Embed(ctx, dec.Utterance)
vec, err := router.EmbedPassage(ctx, h.embedder, dec.Utterance)
if err != nil {
log.Printf("voice: embed note: %v", err)
return "не получилось сохранить заметку."
@@ -750,7 +750,7 @@ func (h *reactiveHandler) applyAction(ctx context.Context, dec router.Decision)
return fmt.Sprintf("в %s сейчас %.0f градусов, %s.", w.Location, w.Temperature, w.Condition)
}
vec, err := h.embedder.Embed(ctx, dec.Utterance)
vec, err := router.EmbedQuery(ctx, h.embedder, dec.Utterance)
if err != nil {
log.Printf("voice: embed query: %v", err)
return "не получилось найти ответ."
@@ -1421,23 +1421,12 @@ func (h *reactiveHandler) resolveConfirm(ctx context.Context, text string) (stri
switch classifyConfirm(text) {
case confirmYes:
h.pendingRoutine = nil
// Create a recurring reminder at the detected interval.
// Weekly patterns get a cron expression; arbitrary intervals
// fire once and the detector re-proposes on the next cycle.
intervalDur := time.Duration(pr.interval * 24 * float64(time.Hour))
fire := h.now().Add(intervalDur)
cron := ""
if pr.interval >= 6.5 && pr.interval <= 7.5 {
cron = fmt.Sprintf("0 %d * * %d", fire.Hour(), int(fire.Weekday()))
}
payload := fmt.Sprintf(`{"text":"%s %s"}`, pr.action, pr.object)
remID, err := h.api.CreateReminder(ctx, fire, payload, cron)
if err != nil {
log.Printf("voice: create routine reminder: %v", err)
return "не получилось поставить напоминание.", true
}
if err := h.dataStore.AcceptProposedRoutine(ctx, pr.routineID, remID); err != nil {
// Only record the acceptance. The tick loop reads accepted
// routines and nudges on their own interval. Building a reminder
// here made a routine fire exactly once (Vikunja #366).
if err := h.dataStore.AcceptProposedRoutine(ctx, pr.routineID, h.now()); err != nil {
log.Printf("voice: accept proposed routine: %v", err)
return "не получилось запомнить рутину.", true
}
return "буду напоминать.", true
case confirmNo:
+4 -1
View File
@@ -91,7 +91,10 @@ func handleEcosystem(w http.ResponseWriter, r *http.Request, urls ecoURLs) {
var d ecoData
var wg sync.WaitGroup
wg.Add(3)
go func() { defer wg.Done(); d.Nexus.Err = getEco(ctx, urls.nexus, "/api/v1/entities?limit=50", &d.Nexus.Rows) }()
go func() {
defer wg.Done()
d.Nexus.Err = getEco(ctx, urls.nexus, "/api/v1/entities?limit=50", &d.Nexus.Rows)
}()
go func() {
defer wg.Done()
d.Praxis.Err = getEco(ctx, urls.praxis, "/api/v1/items?limit=50", &d.Praxis.Rows)
+11 -8
View File
@@ -928,7 +928,6 @@ type routineCore struct {
routines []ipc.ProposedRoutine
dismissed int64
acceptedID int64
acceptedRe int64
remCron string
}
@@ -941,8 +940,8 @@ func (c *routineCore) DismissProposedRoutine(_ context.Context, id int64) error
return nil
}
func (c *routineCore) AcceptProposedRoutine(_ context.Context, id, remID int64) error {
c.acceptedID, c.acceptedRe = id, remID
func (c *routineCore) AcceptProposedRoutine(_ context.Context, id int64) error {
c.acceptedID = id
return nil
}
@@ -990,18 +989,22 @@ func TestHandleRoutines_Accept_RequiresStepUp(t *testing.T) {
}
}
func TestHandleRoutines_Accept_CreatesReminderAndLinksIt(t *testing.T) {
// Accepting only flips the status. It used to also create a one-shot reminder,
// which is why a non-weekly routine fired once and then went quiet forever
// (Vikunja #366). The tick loop owns the schedule now, so a reminder here would
// be a second, competing schedule.
func TestHandleRoutines_Accept_FlipsStatusAndMakesNoReminder(t *testing.T) {
core := weeklyRoutineCore()
rr := httptest.NewRecorder()
handleRoutines(rr, postRoutine("accept", "3"), core, stepUpSession(), false)
if rr.Code != http.StatusOK {
t.Fatalf("status = %d, want 200; body=%s", rr.Code, rr.Body.String())
}
if core.acceptedID != 3 || core.acceptedRe != 77 {
t.Fatalf("accepted id=%d reminder=%d, want 3 and 77", core.acceptedID, core.acceptedRe)
if core.acceptedID != 3 {
t.Fatalf("accepted id = %d, want 3", core.acceptedID)
}
if core.remCron == "" {
t.Fatal("a weekly pattern should get a cron expression")
if core.remCron != "" {
t.Fatalf("accepting must not create a reminder, got cron %q", core.remCron)
}
}
+5 -14
View File
@@ -898,20 +898,11 @@ func acceptRoutine(ctx context.Context, core ipc.CoreAPI, id int64) error {
return errors.New("no such proposed routine")
}
fire := time.Now().Add(time.Duration(found.IntervalDays * 24 * float64(time.Hour)))
cron := ""
if found.IntervalDays >= 6.5 && found.IntervalDays <= 7.5 {
cron = fmt.Sprintf("0 %d * * %d", fire.Hour(), int(fire.Weekday()))
}
payload, err := json.Marshal(map[string]string{"text": found.Action + " " + found.Object})
if err != nil {
return err
}
remID, err := core.CreateReminder(ctx, fire, string(payload), cron)
if err != nil {
return err
}
return core.AcceptProposedRoutine(ctx, id, remID)
// No reminder is created here. Accepting only flips the status; the tick
// loop reads accepted routines and nudges on the interval (Vikunja #366).
// The old code made a one-shot reminder, so a non-weekly routine fired
// once and then went quiet forever.
return core.AcceptProposedRoutine(ctx, id)
}
func handleTrace(w http.ResponseWriter, r *http.Request, core ipc.CoreAPI) {
+3 -3
View File
@@ -36,11 +36,11 @@
"stt": { "socket": "/run/maven/stt.sock", "lang": "ru" },
"tts": { "socket": "/run/maven/tts.sock", "lang": "ru" },
"embedder": {
"model_path": "/opt/maven/models/embedder/model.onnx",
"tokenizer_path": "/opt/maven/models/embedder/tokenizer.json",
"model_path": "/opt/maven/models/embedder/multilingual-e5-small/model_quantized.onnx",
"tokenizer_path": "/opt/maven/models/embedder/multilingual-e5-small/tokenizer.json",
"lib_path": "/opt/maven/lib/libonnxruntime.so"
},
"llm_router": false,
"llm_router": true,
"tool_timeout": "30s",
"tools": [
{ "name": "status", "cmd": ["systemctl", "status"], "scope": "homelab", "destructive": false },
+1 -1
View File
@@ -446,7 +446,7 @@ func (r *recordingAPI) RevertFact(_ context.Context, _ string) (int64, error) {
func (r *recordingAPI) ListProposedRoutines(_ context.Context) ([]ipc.ProposedRoutine, error) {
return nil, nil
}
func (r *recordingAPI) AcceptProposedRoutine(_ context.Context, _, _ int64) error {
func (r *recordingAPI) AcceptProposedRoutine(_ context.Context, _ int64) error {
return nil
}
func (r *recordingAPI) DismissProposedRoutine(_ context.Context, _ int64) error {
+33 -11
View File
@@ -258,18 +258,25 @@ type VoiceConfig struct {
RouterThreshold float64 `json:"router_threshold,omitempty"`
// LLMRouter — route with the resident model instead of the embedding
// classifier. Measured on the held-out fixture (ROUTING-EVAL-31-07-2026.md)
// the model gets 50.0% of intents right against the classifier's 36.8%, but
// it costs about 800ms per turn instead of 30ms.
// classifier. On by default since Vikunja #320.
//
// TODO: the default stays false until two things land.
// 1. The LLM router cannot refuse. LLMRouter.Route hardcodes
// Confidence: 1.0, so the stage-3 clarify gate never fires and an
// unclear utterance becomes a confident wrong action (Vikunja #359).
// 2. Extractor.Extract never runs on an LLM decision, so acts arrive with
// no Fn and reminders with no Time.
// Turning this on today makes routing more accurate and less safe.
LLMRouter bool `json:"llm_router,omitempty"`
// Measured on the held-out fixture (ROUTING-EVAL-31-07-2026.md): 63.2% of
// intents right against the classifier's 50.0%, and no route errors. It
// costs about 1s per turn instead of 30ms.
//
// It is safe to leave on. The model can refuse — it answers "unknown" when
// it cannot route, and the turn drops to the classifier and its clarify
// gate. Any LLM error does the same, so a turn never breaks on the model.
// Slot extraction runs on LLM decisions too, so acts get their Fn and
// reminders their Time.
//
// Set it false to go back to the classifier, e.g. on a box with no
// llama-server or when 1s a turn is too slow.
//
// It is a pointer so that "missing from the file" and "explicitly false"
// are different things: missing means on, false means off. Read it with
// UseLLMRouter(), not directly.
LLMRouter *bool `json:"llm_router,omitempty"`
// QueryMinScore — the note-recall confidence gate. Top cosine below this
// ⇒ "I don't know" instead of a guess. Tuned for the ONNX embedder (0.55);
@@ -405,6 +412,8 @@ const (
DefaultRouterThreshold = 0.55
DefaultQueryMinScore = 0.55
DefaultToolTimeout = 30 * time.Second
// DefaultLLMRouter — route with the resident model unless told otherwise.
DefaultLLMRouter = true
DefaultFactEnrichmentInterval = 30 * time.Second
)
@@ -489,6 +498,10 @@ func (c *Config) applyDefaults() {
if c.Voice.ToolTimeout <= 0 {
c.Voice.ToolTimeout = Duration(DefaultToolTimeout)
}
if c.Voice.LLMRouter == nil {
on := DefaultLLMRouter
c.Voice.LLMRouter = &on
}
}
// routines: default severity to care-class (1) — the safe floor: a
@@ -507,6 +520,15 @@ func (c *Config) applyDefaults() {
}
}
// UseLLMRouter reports whether to route with the resident model. Unset means
// on; only an explicit false in the config turns it off.
func (v *VoiceConfig) UseLLMRouter() bool {
if v == nil || v.LLMRouter == nil {
return DefaultLLMRouter
}
return *v.LLMRouter
}
func (c *Config) validate() error {
if c.Phraser != nil {
if c.Phraser.ModelPath == "" {
+16 -4
View File
@@ -171,14 +171,26 @@ func TestWeatherConfigNilOK(t *testing.T) {
}
}
func TestLLMRouterDefaultsOff(t *testing.T) {
func TestLLMRouterDefaultsOn(t *testing.T) {
p := writeConfig(t, `{"voice":{"enabled":true,"bind":"127.0.0.1:9100"}}`)
c, err := Load(p)
if err != nil {
t.Fatalf("Load: %v", err)
}
if c.Voice.LLMRouter {
t.Error("voice.llm_router absent should mean false")
if !c.Voice.UseLLMRouter() {
t.Error("voice.llm_router absent should mean on")
}
}
// Missing and explicitly false must not mean the same thing.
func TestLLMRouterExplicitFalseTurnsItOff(t *testing.T) {
p := writeConfig(t, `{"voice":{"enabled":true,"bind":"127.0.0.1:9100","llm_router":false}}`)
c, err := Load(p)
if err != nil {
t.Fatalf("Load: %v", err)
}
if c.Voice.UseLLMRouter() {
t.Error("voice.llm_router false should turn it off")
}
}
@@ -188,7 +200,7 @@ func TestLLMRouterRead(t *testing.T) {
if err != nil {
t.Fatalf("Load: %v", err)
}
if !c.Voice.LLMRouter {
if !c.Voice.UseLLMRouter() {
t.Error("voice.llm_router true was not read")
}
}
+4 -5
View File
@@ -238,8 +238,7 @@ type dismissProposedRoutineReq struct {
}
type acceptProposedRoutineReq struct {
ID int64 `json:"id"`
ReminderID int64 `json:"reminder_id"`
ID int64 `json:"id"`
}
// CoreAPI — what core exposes to modules. One Go interface, satisfied by:
@@ -291,9 +290,9 @@ type CoreAPI interface {
ListProposedRoutines(ctx context.Context) ([]ProposedRoutine, error)
// DismissProposedRoutine flips a proposed routine to 'dismissed'.
DismissProposedRoutine(ctx context.Context, id int64) error
// AcceptProposedRoutine flips a proposed routine to 'accepted' and links
// the reminder that will fire it. The caller creates the reminder first.
AcceptProposedRoutine(ctx context.Context, id, reminderID int64) error
// AcceptProposedRoutine flips a proposed routine to 'accepted'. The tick
// loop takes the schedule from there — no reminder is created (Vikunja #366).
AcceptProposedRoutine(ctx context.Context, id int64) error
// TickTrace returns the most recent tick's rule trace. The daemon caches
// this after every tick; the store adapter returns an error (trace is not
+2 -2
View File
@@ -430,8 +430,8 @@ func (c *Client) DismissProposedRoutine(ctx context.Context, id int64) error {
return c.call(ctx, MethodDismissProposedRoutine, dismissProposedRoutineReq{ID: id}, nil)
}
func (c *Client) AcceptProposedRoutine(ctx context.Context, id, reminderID int64) error {
return c.call(ctx, MethodAcceptProposedRoutine, acceptProposedRoutineReq{ID: id, ReminderID: reminderID}, nil)
func (c *Client) AcceptProposedRoutine(ctx context.Context, id int64) error {
return c.call(ctx, MethodAcceptProposedRoutine, acceptProposedRoutineReq{ID: id}, nil)
}
func (c *Client) Chat(ctx context.Context, text string) (string, error) {
+1 -1
View File
@@ -478,7 +478,7 @@ func (a *chatTestAPI) DeleteTool(ctx context.Context, name string) error {
func (a *chatTestAPI) ListProposedRoutines(ctx context.Context) ([]ProposedRoutine, error) {
return nil, ErrUnknownMethod
}
func (a *chatTestAPI) AcceptProposedRoutine(ctx context.Context, id, remID int64) error {
func (a *chatTestAPI) AcceptProposedRoutine(ctx context.Context, id int64) error {
return nil
}
func (a *chatTestAPI) DismissProposedRoutine(ctx context.Context, id int64) error {
+3 -3
View File
@@ -253,8 +253,8 @@ func (a *storeAPI) DismissProposedRoutine(ctx context.Context, id int64) error {
return mapErr(a.s.DismissProposedRoutine(ctx, id))
}
func (a *storeAPI) AcceptProposedRoutine(ctx context.Context, id, reminderID int64) error {
return mapErr(a.s.AcceptProposedRoutine(ctx, id, reminderID))
func (a *storeAPI) AcceptProposedRoutine(ctx context.Context, id int64) error {
return mapErr(a.s.AcceptProposedRoutine(ctx, id, time.Now().UTC()))
}
func toTool(t store.Tool) Tool {
@@ -783,7 +783,7 @@ func (s *Server) dispatch(ctx context.Context, req Request) (json.RawMessage, er
if err := unmarshalParams(req.Params, &p); err != nil {
return nil, err
}
return marshalResult(nil), api.AcceptProposedRoutine(ctx, p.ID, p.ReminderID)
return marshalResult(nil), api.AcceptProposedRoutine(ctx, p.ID)
case MethodRevertFact:
var p struct {
+27 -27
View File
@@ -13,39 +13,39 @@ import (
type Method string
const (
MethodWriteFact Method = "write_fact"
MethodLatestFact Method = "latest_fact"
MethodLatestFactBySource Method = "latest_fact_by_source"
MethodSince Method = "since"
MethodPresence Method = "presence"
MethodCreateReminder Method = "create_reminder"
MethodMarkReminder Method = "mark_reminder"
MethodListReminders Method = "list_reminders"
MethodRecordNudge Method = "record_nudge"
MethodResolveNudge Method = "resolve_nudge"
MethodRecentOutcomes Method = "recent_outcomes"
MethodRecentFacts Method = "recent_facts"
MethodCalendarEvents Method = "calendar_events"
MethodRecentNudges Method = "recent_nudges"
MethodWriteNote Method = "write_note"
MethodQueryNotes Method = "query_notes"
MethodRecentNotes Method = "recent_notes"
MethodProposeTool Method = "propose_tool"
MethodEnableTool Method = "enable_tool"
MethodDisableTool Method = "disable_tool"
MethodAssertStepUp Method = "assert_stepup"
MethodStoreEncryptionKey Method = "store_encryption_key"
MethodUnlock Method = "unlock"
MethodLookupTool Method = "lookup_tool"
MethodWriteFact Method = "write_fact"
MethodLatestFact Method = "latest_fact"
MethodLatestFactBySource Method = "latest_fact_by_source"
MethodSince Method = "since"
MethodPresence Method = "presence"
MethodCreateReminder Method = "create_reminder"
MethodMarkReminder Method = "mark_reminder"
MethodListReminders Method = "list_reminders"
MethodRecordNudge Method = "record_nudge"
MethodResolveNudge Method = "resolve_nudge"
MethodRecentOutcomes Method = "recent_outcomes"
MethodRecentFacts Method = "recent_facts"
MethodCalendarEvents Method = "calendar_events"
MethodRecentNudges Method = "recent_nudges"
MethodWriteNote Method = "write_note"
MethodQueryNotes Method = "query_notes"
MethodRecentNotes Method = "recent_notes"
MethodProposeTool Method = "propose_tool"
MethodEnableTool Method = "enable_tool"
MethodDisableTool Method = "disable_tool"
MethodAssertStepUp Method = "assert_stepup"
MethodStoreEncryptionKey Method = "store_encryption_key"
MethodUnlock Method = "unlock"
MethodLookupTool Method = "lookup_tool"
MethodListTools Method = "list_tools"
MethodDeleteTool Method = "delete_tool"
MethodListProposedRoutines Method = "list_proposed_routines"
MethodDismissProposedRoutine Method = "dismiss_proposed_routine"
MethodAcceptProposedRoutine Method = "accept_proposed_routine"
MethodRevertFact Method = "revert_fact"
MethodTickTrace Method = "tick_trace"
MethodMorningStatus Method = "morning_status"
MethodChat Method = "chat"
MethodTickTrace Method = "tick_trace"
MethodMorningStatus Method = "morning_status"
MethodChat Method = "chat"
)
// Request — one frame from module to core. Params is the JSON-encoded argument
+1 -1
View File
@@ -19,7 +19,7 @@ func TestComplete(t *testing.T) {
t.Errorf("path = %q, want /v1/chat/completions", r.URL.Path)
}
var reqBody struct {
Messages []struct {
Messages []struct {
Role string `json:"role"`
Content string `json:"content"`
} `json:"messages"`
+27 -5
View File
@@ -117,18 +117,40 @@ type cachingEmbedder struct {
seen map[string][]float32
}
var _ router.AsymmetricEmbedder = (*cachingEmbedder)(nil)
func (c *cachingEmbedder) Dim() int { return c.inner.Dim() }
func (c *cachingEmbedder) Close() error { return nil } // the caller owns inner
func (c *cachingEmbedder) Embed(ctx context.Context, text string) ([]float32, error) {
if v, ok := c.seen[text]; ok {
return c.cached(ctx, "embed:"+text, func() ([]float32, error) {
return c.inner.Embed(ctx, text)
})
}
// The two sides of an asymmetric embedder give different vectors for the same
// string, so the cache key has to say which side asked.
func (c *cachingEmbedder) EmbedQuery(ctx context.Context, text string) ([]float32, error) {
return c.cached(ctx, "query:"+text, func() ([]float32, error) {
return router.EmbedQuery(ctx, c.inner, text)
})
}
func (c *cachingEmbedder) EmbedPassage(ctx context.Context, text string) ([]float32, error) {
return c.cached(ctx, "passage:"+text, func() ([]float32, error) {
return router.EmbedPassage(ctx, c.inner, text)
})
}
func (c *cachingEmbedder) cached(_ context.Context, key string, embed func() ([]float32, error)) ([]float32, error) {
if v, ok := c.seen[key]; ok {
return v, nil
}
v, err := c.inner.Embed(ctx, text)
v, err := embed()
if err != nil {
return nil, err
}
c.seen[text] = v
c.seen[key] = v
return v, nil
}
@@ -305,7 +327,7 @@ func scoreCase(ctx context.Context, emb router.Embedder, newStore NewStore, minS
all := append(append([]StoredNote(nil), c.Notes...), filler...)
for _, n := range all {
vec, err := emb.Embed(ctx, n.Text)
vec, err := router.EmbedPassage(ctx, emb, n.Text)
if err != nil {
return Outcome{}, fmt.Errorf("%s: embed note %s: %w", c.ID, n.ID, err)
}
@@ -317,7 +339,7 @@ func scoreCase(ctx context.Context, emb router.Embedder, newStore NewStore, minS
o := Outcome{Case: c}
start := time.Now()
qvec, err := emb.Embed(ctx, c.Query)
qvec, err := router.EmbedQuery(ctx, emb, c.Query)
if err != nil {
o.Latency = time.Since(start)
o.Err = err
@@ -246,8 +246,8 @@ func TestONNXRecall(t *testing.T) {
if lib == "" {
t.Skip("MAVEN_ONNX_LIB unset — see AGENTS.md § Embedder model for intent routing")
}
model := filepath.Join("../../..", "models/embedder/model.onnx")
tok := filepath.Join("../../..", "models/embedder/tokenizer.json")
model := filepath.Join("../../..", "models/embedder/multilingual-e5-small/model_quantized.onnx")
tok := filepath.Join("../../..", "models/embedder/multilingual-e5-small/tokenizer.json")
for _, p := range []string{lib, model, tok} {
if _, err := os.Stat(p); err != nil {
t.Skipf("missing %s: %v", p, err)
+287
View File
@@ -0,0 +1,287 @@
package eval
import (
"fmt"
"regexp"
"strings"
"unicode"
"unicode/utf8"
)
// The check names, in report order. Every check is a string or length test — no
// model grades another model here.
const (
CheckMood = "mood" // mood is in the documented enum
CheckLang = "lang" // the operator's language, not the prompt's
CheckLength = "length" // a nudge is one sentence, not a paragraph
CheckFeminine = "feminine" // her self-reference is feminine (hard constraint)
CheckCringe = "cringe" // DESIGN.md § Non-goals, "not a relationship"
CheckOnTopic = "ontopic" // says the thing the rule is about
)
// CheckNames — report order.
var CheckNames = []string{CheckMood, CheckLang, CheckLength, CheckFeminine, CheckCringe, CheckOnTopic}
// Result — one check on one message.
type Result struct {
Name string
Pass bool
Detail string
}
// Moods — the fixed enum from the LLM output contract. Not extended here; the
// contract lives in the daemon and the harness only reads it.
var Moods = map[string]bool{
"neutral": true, "happy": true, "thinking": true, "tired": true, "confused": true,
}
// Length ceilings. Justification: the nudge is spoken by piper at roughly 14
// characters per second, so 120 characters is about 8 seconds of speech. The
// operator has AuDHD — past one short sentence a nudge stops being a nudge and
// becomes something to tune out, which is exactly the "not a nag" failure. The
// word ceiling catches the same thing for languages that pack more per byte.
const (
MaxChars = 120
MaxWords = 16
)
// RunChecks scores one message. Order matches CheckNames.
func RunChecks(c Case, body, mood string) []Result {
return []Result{
checkMood(mood),
checkLang(body),
checkLength(body),
checkFeminine(body),
checkCringe(body),
checkOnTopic(c, body),
}
}
func checkMood(mood string) Result {
if Moods[mood] {
return Result{CheckMood, true, ""}
}
return Result{CheckMood, false, fmt.Sprintf("mood %q not in the enum", mood)}
}
// checkLang — the operator is Russian-speaking and the nudge is spoken aloud by
// a Russian piper voice. An English nudge is not a tone problem, it is an
// unusable one.
func checkLang(body string) Result {
cyr, lat := 0, 0
for _, r := range body {
switch {
case unicode.Is(unicode.Cyrillic, r):
cyr++
case r >= 'a' && r <= 'z', r >= 'A' && r <= 'Z':
lat++
}
}
if cyr > lat {
return Result{CheckLang, true, ""}
}
return Result{CheckLang, false, fmt.Sprintf("not Russian (%d cyrillic vs %d latin letters)", cyr, lat)}
}
func checkLength(body string) Result {
chars := utf8.RuneCountInString(body)
words := len(strings.Fields(body))
if chars <= MaxChars && words <= MaxWords {
return Result{CheckLength, true, ""}
}
return Result{CheckLength, false,
fmt.Sprintf("%d chars / %d words, ceiling %d / %d", chars, words, MaxChars, MaxWords)}
}
// --- feminine self-reference ---------------------------------------------
//
// The hard constraint (CLAUDE.md, DESIGN.md § Identity): Maven's Russian
// self-reference is feminine. The operator is male, so second-person forms
// addressed to him are MASCULINE and must not be flagged — "ты не пил воду" is
// correct, "я напомнил" is not. Both directions matter, which is why this is a
// windowed scan around "я" and not a bare search for masculine endings.
var wordRE = regexp.MustCompile(`[\p{Cyrillic}]+|[,.;:!?…—-]`)
// secondPerson — pronouns that end the self-reference window. Everything after
// one of these is about him, not about her.
var secondPerson = map[string]bool{
"ты": true, "тебе": true, "тебя": true, "тобой": true,
"вы": true, "вам": true, "вас": true,
"он": true, "она": true, "оно": true, "они": true,
}
// masculinePredicative — short adjectives with no verb ending to key off.
var masculinePredicative = map[string]bool{
"должен": true, "готов": true, "рад": true, "уверен": true,
"обязан": true, "сам": true, "занят": true, "прав": true,
}
// nounsEndingInL — the false positives of "ends in л ⇒ masculine past tense".
// Small on purpose: it only has to cover nouns a nudge might actually use.
var nounsEndingInL = map[string]bool{
"стол": true, "стул": true, "пол": true, "зал": true, "гол": true,
"узел": true, "отдел": true, "файл": true, "канал": true, "угол": true,
"футбол": true, "вокзал": true, "металл": true, "интервал": true,
"уровень": true, "мускул": true, "апрель": true, "июль": true, "рубль": true,
}
// masculinePast reports whether a word looks like a masculine past-tense verb.
// Russian past tense is gendered by suffix: -л (m), -ла (f). A 0.8B with weak
// Russian defaults to the masculine form, which is the exact drift being
// measured.
func masculinePast(w string) bool {
if len([]rune(w)) < 3 || nounsEndingInL[w] {
return false
}
return strings.HasSuffix(w, "л") || strings.HasSuffix(w, "лся")
}
func checkFeminine(body string) Result {
words := wordRE.FindAllString(strings.ToLower(body), -1)
for i, w := range words {
if w != "я" {
continue
}
// Scan the next few words. Stop at punctuation or at a second-person
// pronoun: past that point the sentence is about him and masculine is
// correct.
for j := i + 1; j < len(words) && j <= i+3; j++ {
nw := words[j]
if len(nw) == 1 && !unicode.Is(unicode.Cyrillic, []rune(nw)[0]) {
break
}
if secondPerson[nw] {
break
}
if masculinePast(nw) || masculinePredicative[nw] {
return Result{CheckFeminine, false,
fmt.Sprintf("masculine self-reference %q after \"я\"", nw)}
}
}
}
// Second pass: self-reference with the pronoun dropped — "напомнил тебе",
// "проверил за тебя". A masculine past-tense verb whose object is HIM can
// only be her speaking about herself.
for i, w := range words {
if !masculinePast(w) || i+1 >= len(words) {
continue
}
next := words[i+1]
if next == "тебе" || next == "тебя" || next == "за" {
return Result{CheckFeminine, false,
fmt.Sprintf("masculine self-reference %q before %q", w, next)}
}
}
return Result{CheckFeminine, true, ""}
}
// --- the cringe checks ---------------------------------------------------
//
// "Think Jarvis without the cringe part". DESIGN.md § Non-goals: "Not a
// relationship — mom-tone is a function that makes nudges land, not emotional
// company. Names the drift a warm small model falls into." Each pattern below
// is one shape of that drift. They are deliberately specific: a check that
// flags any warmth at all would make the nudges robotic, which is the other
// failure.
type cringePattern struct {
// what the pattern is defending against, shown in the failure detail.
why string
pat *regexp.Regexp
}
var cringePatterns = []cringePattern{
{
// Endearments. "Not a relationship" — a pet name reframes a nudge as
// intimacy, and the operator asked for mother-like, not girlfriend-like.
why: "pet name / endearment",
// No \b around the Russian alternatives: Go's RE2 \b is ASCII-only and
// never matches at a Cyrillic boundary, so anchoring them would make
// this check silently always pass.
pat: regexp.MustCompile(`(?i)(милый|дорогой|солнышко|солнце моё|солнце мое|зайчик|котик|сладкий|малыш|дружок|родной|любимый|\bhoney\b|\bsweetie\b|\bdarling\b|\bbuddy\b)`),
},
{
// Emoji. The nudge is spoken aloud; an emoji is either silence or a TTS
// artefact. Also the single loudest cringe signal in a small model.
why: "emoji",
pat: nil, // handled by hasEmoji, ranges don't fit a regexp cleanly
},
{
// Exclamation pileup. One "!" is emphasis; two is a cheerleader.
why: "more than one exclamation mark",
pat: regexp.MustCompile(`!.*!|!!`),
},
{
// Fake concern. She has no feelings to report, and reporting them makes
// the nudge about her instead of about the water.
why: "fake concern opener",
pat: regexp.MustCompile(`(?i)(я волну|я беспоко|беспокоюсь|переживаю|я забочусь|я тревож|мне тревожно|i'?m worried)`),
},
{
// Apologising. The rule decided she speaks. Apologising for a greenlit
// nudge undermines the one thing that makes nudges land.
why: "apology",
pat: regexp.MustCompile(`(?i)(извини|прости|сожалею|прошу прощения|не хочу мешать|не хочу отвлекать|sorry|apolog)`),
},
{
// Offering emotional support. The explicit "not emotional company" line.
why: "offer of emotional support",
pat: regexp.MustCompile(`(?i)(я рядом|я здесь для теб|ты не один|всё будет хорошо|все будет хорошо|не переживай|я с тобой|обнимаю|я поддерж|держись)`),
},
{
// Asking how he feels. Turns a one-way nudge into a conversation he now
// owes an answer to — the most reliable way to make him mute it.
why: "asking how he feels",
pat: regexp.MustCompile(`(?i)(как ты\s*[?.!]|как ты себя|как самочувств|как настроение|всё ли в порядке|все ли в порядке|ты в порядке|how are you)`),
},
{
// Praise for compliance. Rewards make the nudge a training exercise;
// "not a relationship" again, from the other side.
why: "praise / reward framing",
pat: regexp.MustCompile(`(?i)(молодец|умница|ты справ|гордюсь|горжусь|отличная работа|так держать|good job|proud of you)`),
},
}
// hasEmoji — the pictographic ranges plus the variation selector. Cyrillic and
// ordinary punctuation are far below all of these.
func hasEmoji(s string) bool {
for _, r := range s {
switch {
case r >= 0x1F000 && r <= 0x1FAFF,
r >= 0x2600 && r <= 0x27BF,
r >= 0x2B00 && r <= 0x2BFF,
r == 0xFE0F, r == 0x203C, r == 0x2049:
return true
}
}
return false
}
func checkCringe(body string) Result {
for _, c := range cringePatterns {
if c.pat == nil {
if hasEmoji(body) {
return Result{CheckCringe, false, c.why}
}
continue
}
if m := c.pat.FindString(body); m != "" {
return Result{CheckCringe, false, fmt.Sprintf("%s (%q)", c.why, m)}
}
}
return Result{CheckCringe, true, ""}
}
// checkOnTopic — the message must name the thing the rule is about. A nudge
// that never mentions water leaves the operator with a chime and no action.
func checkOnTopic(c Case, body string) Result {
low := strings.ToLower(body)
for _, want := range c.WantAny {
if strings.Contains(low, strings.ToLower(want)) {
return Result{CheckOnTopic, true, ""}
}
}
return Result{CheckOnTopic, false,
fmt.Sprintf("mentions none of %v", c.WantAny)}
}
+334
View File
@@ -0,0 +1,334 @@
// Package eval scores nudge phrasing — the sentences the operator actually
// hears. It is the phrasing counterpart to internal/router/eval.
//
// Why a separate package from phraser: the fixture must be scorable by BOTH
// phrasing paths (the deterministic Stub and the resident model) from outside
// the phraser package, and a _test.go file inside phraser cannot be imported.
// So the fixture is embedded here and the scorer takes a Nudger interface.
//
// Why deterministic checks and not model judgement: the resident model is a
// 0.8B. It cannot grade its own tone. Every check in checks.go is a string or
// length test that a human can read and disagree with. A score here is a claim
// about measurable properties, not about whether a sentence is good.
//
// DESIGN.md § "Rules decide, LLM phrases" is why there is no send/veto signal
// anywhere in this package: the rule already decided she speaks. The phraser
// only words it, so a nudge the model refuses to write is a failure, never a
// legitimate outcome.
package eval
import (
"context"
_ "embed"
"encoding/json"
"fmt"
"sort"
"strings"
"time"
"github.com/kami/maven/internal/delivery"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/store"
)
//go:embed nudges_v1.json
var fixtureJSON []byte
// SchemaVersion — the version this package understands.
const SchemaVersion = 1
// Case — one nudge situation, as a real tick would present it. The fields are
// the (rule, severity, context) input DESIGN.md names, flattened to JSON.
//
// WantAny is the on-topic contract: at least one of these lowercased fragments
// must appear in the message. A water nudge that never mentions water is a
// failure however charming it reads. Fragments are stems ("вод") so declension
// does not defeat the check, and they list both languages because the Stub is
// still English (see the writeup).
type Case struct {
ID string `json:"id"`
Rule string `json:"rule"`
Severity int `json:"severity"`
// SinceMinutes — age of the rule's own fact. 0 means "no such fact", which
// is the branch where the phraser has no duration to name.
SinceMinutes int `json:"since_minutes"`
// FactKey/FactValue/FactSource — the aggregate fact behind the ops rules.
// service_down phrasing reads the key for the service name.
FactKey string `json:"fact_key,omitempty"`
FactValue string `json:"fact_value,omitempty"`
FactSource string `json:"fact_source,omitempty"`
// QuietHours/CalendarBusy — the bad moments. The gate already let this
// nudge through (ops outranks quiet hours), so the phrasing still has to be
// short and plain rather than apologetic about the timing.
QuietHours bool `json:"quiet_hours,omitempty"`
CalendarBusy bool `json:"calendar_busy,omitempty"`
WantAny []string `json:"want_any"`
Tags []string `json:"tags,omitempty"`
Note string `json:"note,omitempty"`
}
// Fixture — the versioned envelope. SchemaVersion gates the loader so an older
// binary refuses a fixture it would misread instead of reporting a wrong score.
type Fixture struct {
SchemaVersion int `json:"schema_version"`
Name string `json:"name"`
ReferenceNow string `json:"reference_now"`
Notes []string `json:"notes"`
Cases []Case `json:"cases"`
}
// Load returns the embedded fixture.
func Load() (Fixture, error) {
var f Fixture
if err := json.Unmarshal(fixtureJSON, &f); err != nil {
return Fixture{}, fmt.Errorf("parse fixture: %w", err)
}
if f.SchemaVersion != SchemaVersion {
return Fixture{}, fmt.Errorf("fixture schema_version %d, want %d", f.SchemaVersion, SchemaVersion)
}
if len(f.Cases) == 0 {
return Fixture{}, fmt.Errorf("fixture has no cases")
}
return f, nil
}
// Now — the fixture's reference clock, so fact ages are reproducible.
func (f Fixture) Now() (time.Time, error) {
t, err := time.Parse(time.RFC3339, f.ReferenceNow)
if err != nil {
return time.Time{}, fmt.Errorf("parse reference_now %q: %w", f.ReferenceNow, err)
}
return t, nil
}
// Candidate rebuilds the loop.Candidate a tick would hand the phraser.
func (c Case) Candidate(now time.Time) loop.Candidate {
state := loop.State{
Now: now,
Facts: map[string]store.Fact{},
QuietHours: c.QuietHours,
CalendarBusy: c.CalendarBusy,
}
if c.SinceMinutes > 0 {
state.Facts[c.Rule] = store.Fact{
Key: c.Rule,
Ts: now.Add(-time.Duration(c.SinceMinutes) * time.Minute),
}
}
if c.FactKey != "" {
state.Facts[c.Rule] = store.Fact{
Key: c.FactKey,
Value: c.FactValue,
Source: c.FactSource,
Ts: now.Add(-time.Duration(c.SinceMinutes) * time.Minute),
}
}
sev := loop.Severity(c.Severity)
return loop.Candidate{
Rule: loop.Rule{Name: c.Rule, Severity: sev},
Severity: sev,
State: state,
}
}
// Nudger — the one thing a phrasing path must do to be scorable. Both
// *phraser.Stub and *phraser.LLMPhraser satisfy it.
type Nudger interface {
PhraseNudge(ctx context.Context, c loop.Candidate) (delivery.PhrasedNudge, error)
}
// Outcome — one scored case. Failed lists the check names that did not pass,
// Reasons the human-readable detail. Failed is empty exactly when Pass is true.
type Outcome struct {
Case Case
Body string
Mood string
Err error
Latency time.Duration
Pass bool
Failed []string
Reasons []string
}
// Report — the aggregate. ByCheck is the useful part: one composite percentage
// hides which property broke, and tuning a prompt needs to know.
type Report struct {
Name string
Total int
Passed int
Errors int
ByCheck map[string]int
ByRule map[string]TagStat
Outcomes []Outcome
P50 time.Duration
P95 time.Duration
Max time.Duration
}
// TagStat — passed/total for one slice of the fixture.
type TagStat struct{ Passed, Total int }
// Accuracy — fraction of cases that passed every check.
func (r Report) Accuracy() float64 {
if r.Total == 0 {
return 0
}
return float64(r.Passed) / float64(r.Total)
}
// Score runs every case through p and aggregates. A phrasing error scores as a
// miss and is counted in Errors — "the model was down" and "the model wrote
// something bad" are different numbers and a prompt change must not be able to
// hide behind the first one.
//
// Latency is wall-clock per PhraseNudge call. On the CPU/iGPU target a nudge
// the model takes a minute to word has already missed its moment.
func Score(ctx context.Context, name string, p Nudger, f Fixture) (Report, error) {
now, err := f.Now()
if err != nil {
return Report{}, err
}
rep := Report{
Name: name,
Total: len(f.Cases),
ByCheck: map[string]int{},
ByRule: map[string]TagStat{},
}
for _, name := range CheckNames {
rep.ByCheck[name] = 0
}
lat := make([]time.Duration, 0, len(f.Cases))
for _, c := range f.Cases {
start := time.Now()
pn, err := p.PhraseNudge(ctx, c.Candidate(now))
o := Outcome{Case: c, Body: pn.Body, Mood: pn.Mood, Err: err, Latency: time.Since(start)}
lat = append(lat, o.Latency)
if err != nil {
rep.Errors++
o.Failed = append(o.Failed, "call")
o.Reasons = append(o.Reasons, fmt.Sprintf("phrase error: %v", err))
} else {
for _, res := range RunChecks(c, pn.Body, pn.Mood) {
if res.Pass {
rep.ByCheck[res.Name]++
continue
}
o.Failed = append(o.Failed, res.Name)
o.Reasons = append(o.Reasons, res.Name+": "+res.Detail)
}
}
o.Pass = len(o.Failed) == 0
if o.Pass {
rep.Passed++
}
bump(rep.ByRule, ruleFamily(c.Rule), o.Pass)
rep.Outcomes = append(rep.Outcomes, o)
}
sort.Slice(lat, func(i, j int) bool { return lat[i] < lat[j] })
rep.P50, rep.P95 = percentile(lat, 0.50), percentile(lat, 0.95)
if len(lat) > 0 {
rep.Max = lat[len(lat)-1]
}
return rep, nil
}
// ruleFamily collapses "routine:зарядка" to "routine" so the per-rule table
// stays readable however many routines the operator configures.
func ruleFamily(rule string) string {
if i := strings.IndexByte(rule, ':'); i > 0 {
return rule[:i]
}
return rule
}
func bump(m map[string]TagStat, key string, pass bool) {
if key == "" {
return
}
s := m[key]
s.Total++
if pass {
s.Passed++
}
m[key] = s
}
// percentile — nearest-rank on a pre-sorted slice. No interpolation: with ~15
// samples an interpolated p95 invents a latency no call actually took.
func percentile(sorted []time.Duration, p float64) time.Duration {
if len(sorted) == 0 {
return 0
}
i := int(p * float64(len(sorted)))
if i >= len(sorted) {
i = len(sorted) - 1
}
return sorted[i]
}
// String renders the comparison table — composite score, then per-check so a
// regression names the property it broke, then latency.
func (r Report) String() string {
var b strings.Builder
fmt.Fprintf(&b, "%s: %d/%d cases pass every check (%.1f%%), %d errors\n",
r.Name, r.Passed, r.Total, 100*r.Accuracy(), r.Errors)
for _, name := range CheckNames {
fmt.Fprintf(&b, " %-9s %d/%d\n", name, r.ByCheck[name], r.Total)
}
fmt.Fprintf(&b, " latency: p50 %s p95 %s max %s\n", r.P50, r.P95, r.Max)
fmt.Fprintf(&b, " by rule: %s\n", renderStats(r.ByRule))
return b.String()
}
// Failures — per-case detail, sorted by ID so two runs diff cleanly.
func (r Report) Failures() string {
var b strings.Builder
for _, o := range r.sorted() {
if o.Pass {
continue
}
fmt.Fprintf(&b, " %s %q\n %s\n", o.Case.ID, o.Body, strings.Join(o.Reasons, "; "))
}
return b.String()
}
// Messages — every generated message verbatim, pass or fail. This is what a
// human reads to judge tone; the score only says which checks fired.
func (r Report) Messages() string {
var b strings.Builder
for _, o := range r.sorted() {
mark := "ok "
if !o.Pass {
mark = "FAIL"
}
fmt.Fprintf(&b, " %s %-22s [%s] %q\n", mark, o.Case.ID, o.Mood, o.Body)
}
return b.String()
}
func (r Report) sorted() []Outcome {
out := append([]Outcome(nil), r.Outcomes...)
sort.Slice(out, func(i, j int) bool { return out[i].Case.ID < out[j].Case.ID })
return out
}
func renderStats(m map[string]TagStat) string {
keys := make([]string, 0, len(m))
for k := range m {
keys = append(keys, k)
}
sort.Strings(keys)
parts := make([]string, 0, len(keys))
for _, k := range keys {
parts = append(parts, fmt.Sprintf("%s %d/%d", k, m[k].Passed, m[k].Total))
}
return strings.Join(parts, " ")
}
+179
View File
@@ -0,0 +1,179 @@
package eval
import (
"context"
"strings"
"testing"
"github.com/kami/maven/internal/loop"
"github.com/kami/maven/internal/phraser"
)
func TestLoadFixture(t *testing.T) {
f, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
if _, err := f.Now(); err != nil {
t.Fatalf("Now: %v", err)
}
seen := map[string]bool{}
for _, c := range f.Cases {
if seen[c.ID] {
t.Errorf("duplicate case id %q", c.ID)
}
seen[c.ID] = true
if c.Rule == "" || c.Severity < 1 || c.Severity > 4 {
t.Errorf("%s: rule %q severity %d", c.ID, c.Rule, c.Severity)
}
if len(c.WantAny) == 0 {
t.Errorf("%s: no want_any, the on-topic check would always pass", c.ID)
}
}
// Coverage floor: all five loop rules plus both minted families, or the
// fixture measures a subset and the score does not mean what it says.
for _, rule := range []string{"water", "meal", "break", "service_down", "netdata_critical", "routine", "morning"} {
found := false
for _, c := range f.Cases {
if ruleFamily(c.Rule) == rule {
found = true
}
}
if !found {
t.Errorf("no case for rule family %q", rule)
}
}
}
// TestStubBaseline is the CI ratchet: the deterministic Stub, no model, no
// network. The floor is low on purpose — the Stub is English template phrasing,
// so it fails `lang` on every case by construction. The point of the ratchet is
// that the checks keep running and the Stub does not get worse, not that the
// Stub is good.
func TestStubBaseline(t *testing.T) {
f, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
rep, err := Score(context.Background(), "stub (deterministic floor)", phraser.NewStub(), f)
if err != nil {
t.Fatalf("Score: %v", err)
}
t.Log("\n" + rep.String() + rep.Messages())
if rep.Errors != 0 {
t.Errorf("stub returned %d errors — the deterministic path must never fail", rep.Errors)
}
// Per-check ratchets rather than one composite: the Stub's composite is 0
// (it never passes `lang`), so a composite floor would catch nothing.
floors := map[string]int{
CheckMood: 15,
// 12, not 15: the Stub's `break` template genuinely runs past the
// ceiling ("you've been at your desk for 4 hours without a break — step
// away for a bit." is 76 chars but 16 words). Left failing rather than
// raising the ceiling to hide it.
CheckLength: 12,
CheckFeminine: 15,
CheckCringe: 15,
CheckOnTopic: 12,
}
for name, floor := range floors {
if rep.ByCheck[name] < floor {
t.Errorf("check %s: %d/%d, below ratchet %d — phrasing regressed",
name, rep.ByCheck[name], rep.Total, floor)
}
}
}
// TestChecksCatchWhatTheyClaim — the checks are the measurement, so they get
// their own tests. Without these, a bad regexp would silently make every
// phrasing run look clean.
func TestChecksCatchWhatTheyClaim(t *testing.T) {
water := Case{Rule: "water", WantAny: []string{"вод"}}
cases := []struct {
name string
body string
want string // the check that must fail, "" for a clean message
}{
{"clean", "уже четыре часа без воды — попей.", ""},
{"long", "уже четыре часа без воды, а это довольно много, и вообще пить надо регулярно, иначе будет плохо совсем", CheckLength},
{"english", "you haven't had water in 4 hours, drink something", CheckLang},
{"masculine self", "я напомнил про воду.", CheckFeminine},
{"masculine dropped pronoun", "напомнил тебе про воду.", CheckFeminine},
{"masculine predicative", "я должен сказать: попей воды.", CheckFeminine},
// The other direction: HE is male, so second-person masculine is right.
{"second person masculine ok", "ты не пил воду четыре часа.", ""},
{"feminine self ok", "я заметила: воды не было четыре часа.", ""},
{"pet name", "милый, попей воды.", CheckCringe},
{"emoji", "попей воды 💧", CheckCringe},
{"exclamations", "попей воды!!", CheckCringe},
{"fake concern", "я беспокоюсь: воды не было четыре часа.", CheckCringe},
{"apology", "извини, что отвлекаю — попей воды.", CheckCringe},
{"emotional support", "я рядом, ты не один. попей воды.", CheckCringe},
{"asks how he feels", "как ты себя чувствуешь? попей воды.", CheckCringe},
{"praise", "молодец! теперь попей воды.", CheckCringe},
{"off topic", "пора бы уже что-то сделать.", CheckOnTopic},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
var failed []string
for _, r := range RunChecks(water, tc.body, "neutral") {
if !r.Pass {
failed = append(failed, r.Name+"("+r.Detail+")")
}
}
joined := strings.Join(failed, " ")
switch {
case tc.want == "" && len(failed) > 0:
t.Errorf("clean message flagged: %s", joined)
case tc.want != "" && !strings.Contains(joined, tc.want+"("):
t.Errorf("want %s to fail, got %q", tc.want, joined)
}
})
}
}
func TestMoodCheckUsesTheEnum(t *testing.T) {
if r := checkMood("cheerful"); r.Pass {
t.Error("mood outside the enum passed")
}
if r := checkMood(""); r.Pass {
t.Error("empty mood passed")
}
for m := range Moods {
if r := checkMood(m); !r.Pass {
t.Errorf("enum mood %q failed", m)
}
}
}
// TestCandidateCarriesTheContext — the whole harness is worthless if the
// Candidate it builds does not carry the duration the prompt is supposed to
// name.
func TestCandidateCarriesTheContext(t *testing.T) {
f, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
now, _ := f.Now()
for _, c := range f.Cases {
cand := c.Candidate(now)
if cand.Rule.Name != c.Rule || cand.Severity != loop.Severity(c.Severity) {
t.Errorf("%s: candidate lost rule or severity", c.ID)
}
if c.SinceMinutes > 0 {
d, ok := cand.State.Since(c.Rule)
if !ok || int(d.Minutes()) != c.SinceMinutes {
t.Errorf("%s: since %v ok=%v, want %d minutes", c.ID, d, ok, c.SinceMinutes)
}
}
if c.FactKey != "" {
fact, ok := cand.State.Fact(c.Rule)
if !ok || fact.Key != c.FactKey {
t.Errorf("%s: fact key %q, want %q", c.ID, fact.Key, c.FactKey)
}
}
}
}
+73
View File
@@ -0,0 +1,73 @@
package eval
import (
"context"
"os"
"strings"
"testing"
"time"
"github.com/kami/maven/internal/phraser"
)
// TestLLMPhrasingBaseline — the resident model wording real nudges. Opt-in,
// same shape as internal/router/eval's MAVEN_LLM_URL gate, because CI has no
// model and a phrasing run costs minutes on the CPU target.
//
// llama-server -m /mnt/hdd1/llms/qwen3.5/Qwen3.5-0.8B.Q4_K_M.gguf \
// --host 127.0.0.1 --port 18099 -c 2048 -ngl 99
// MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-phrasing
//
// It reports and does not assert a quality bar. The numbers are the input to
// tuning the persona prompt; an assertion here would be the test inventing the
// bar rather than measuring against it. The one thing worth failing on is a
// harness fault — every case erroring means the run measured infrastructure.
func TestLLMPhrasingBaseline(t *testing.T) {
base := os.Getenv("MAVEN_LLM_URL")
if base == "" {
t.Skip("MAVEN_LLM_URL unset — point it at a running llama-server (see doc comment)")
}
// A local llama-server must not go through an HTTP proxy. This box proxies
// loopback through a SOCKS bridge that answers 503, which would score every
// case as a phrasing error and read as "the model cannot phrase".
noProxyLoopback(t)
f, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
cfg := phraser.DefaultConfig("")
// Generous: an unconstrained 0.8B can spend a minute thinking before it
// writes a word, and a timeout would be scored as a model failure.
cfg.Timeout = 5 * time.Minute
p := phraser.NewLLMPhraserAt(base, cfg)
defer p.Close()
rep, err := Score(context.Background(), "llm (0.8B, built-in persona)", p, f)
if err != nil {
t.Fatalf("Score: %v", err)
}
t.Log("\n" + rep.String() + "\nmessages:\n" + rep.Messages() + "\nfailures:\n" + rep.Failures())
if rep.Errors == rep.Total {
t.Errorf("all %d cases errored — harness fault, not a measurement", rep.Total)
}
}
// noProxyLoopback appends the loopback host to no_proxy before any request, so
// http.ProxyFromEnvironment (which caches the environment on first use) sees it.
func noProxyLoopback(t *testing.T) {
t.Helper()
for _, key := range []string{"no_proxy", "NO_PROXY"} {
cur := os.Getenv(key)
if strings.Contains(cur, "127.0.0.1") {
continue
}
if cur == "" {
t.Setenv(key, "127.0.0.1,localhost")
continue
}
t.Setenv(key, cur+",127.0.0.1,localhost")
}
}
+158
View File
@@ -0,0 +1,158 @@
{
"schema_version": 1,
"name": "nudge phrasing v1",
"reference_now": "2026-07-31T21:40:00+03:00",
"notes": [
"The five loop rules (water, meal, break, service_down, netdata_critical) plus one routine: and one morning: case, at the severities they actually ship with.",
"The bad-moment cases (quiet_hours, calendar_busy) are here because the gate already let them through — ops outranks quiet hours. The phrasing must stay short and plain, not apologise for the timing.",
"want_any lists stems, not whole words, so declension does not defeat the on-topic check. English stems are included because the deterministic Stub is still English.",
"There is no expected message. The checks measure properties, not similarity to a reference sentence — a fixed golden string would just freeze one arbitrary phrasing."
],
"cases": [
{
"id": "water-3h",
"rule": "water",
"severity": 1,
"since_minutes": 190,
"want_any": ["вод", "попей", "пить", "выпей", "напит", "water", "drink"],
"tags": ["care"],
"note": "The base case. Just over the 3h predicate."
},
{
"id": "water-7h",
"rule": "water",
"severity": 1,
"since_minutes": 430,
"want_any": ["вод", "попей", "пить", "выпей", "напит", "water", "drink"],
"tags": ["care"],
"note": "Long overdue. Severity is unchanged, so the phrasing must not escalate into alarm."
},
{
"id": "water-busy",
"rule": "water",
"severity": 1,
"since_minutes": 240,
"calendar_busy": true,
"want_any": ["вод", "попей", "пить", "выпей", "напит", "water", "drink"],
"tags": ["care", "bad-moment"],
"note": "Mid-meeting. A bad moment invites an apology, which is the check that should catch it."
},
{
"id": "meal-7h",
"rule": "meal",
"severity": 1,
"since_minutes": 420,
"want_any": ["ешь", "еда", "еды", "поешь", "перекус", "обед", "ужин", "покуш", "food", "eat"],
"tags": ["care"]
},
{
"id": "meal-11h-quiet",
"rule": "meal",
"severity": 1,
"since_minutes": 660,
"quiet_hours": true,
"want_any": ["ешь", "еда", "еды", "поешь", "перекус", "обед", "ужин", "покуш", "food", "eat"],
"tags": ["care", "bad-moment"],
"note": "Quiet hours suppresses care nudges in the gate, so this one only reaches the phraser via an explicit override. Included because it is the shape most likely to draw a hedge."
},
{
"id": "break-90m",
"rule": "break",
"severity": 2,
"since_minutes": 95,
"want_any": ["перерыв", "разомн", "встань", "отдохн", "пауз", "отойд", "размин", "break", "step away"],
"tags": ["care"]
},
{
"id": "break-4h",
"rule": "break",
"severity": 2,
"since_minutes": 240,
"want_any": ["перерыв", "разомн", "встань", "отдохн", "пауз", "отойд", "размин", "break", "step away"],
"tags": ["care"]
},
{
"id": "break-busy",
"rule": "break",
"severity": 2,
"since_minutes": 150,
"calendar_busy": true,
"want_any": ["перерыв", "разомн", "встань", "отдохн", "пауз", "отойд", "размин", "break", "step away"],
"tags": ["care", "bad-moment"]
},
{
"id": "service-down",
"rule": "service_down",
"severity": 4,
"since_minutes": 3,
"fact_key": "vaultwarden",
"fact_value": "\"down\"",
"fact_source": "poll:uptimekuma",
"want_any": ["vaultwarden", "сервис", "упал", "не отвеч", "лежит", "недоступ", "down", "service"],
"tags": ["ops"],
"note": "The service name is in the fact key, not the value. A nudge that says 'a service' without naming it is on-topic but useless — the on-topic check cannot catch that, a human reading Messages() can."
},
{
"id": "service-down-night",
"rule": "service_down",
"severity": 4,
"since_minutes": 2,
"quiet_hours": true,
"fact_key": "nextcloud",
"fact_value": "\"down\"",
"fact_source": "poll:uptimekuma",
"want_any": ["nextcloud", "сервис", "упал", "не отвеч", "лежит", "недоступ", "down", "service"],
"tags": ["ops", "bad-moment"],
"note": "03:00-shaped. Sev4 outranks quiet hours by design, so she speaks — plainly, without softening it into a maybe."
},
{
"id": "netdata-disk",
"rule": "netdata_critical",
"severity": 3,
"since_minutes": 5,
"fact_key": "netdata_alarm",
"fact_value": "\"critical\"",
"fact_source": "poll:netdata",
"want_any": ["netdata", "диск", "критич", "алярм", "аларм", "тревог", "место", "памят", "critical", "alarm", "disk"],
"tags": ["ops"]
},
{
"id": "netdata-busy",
"rule": "netdata_critical",
"severity": 3,
"since_minutes": 12,
"calendar_busy": true,
"fact_key": "netdata_alarm",
"fact_value": "\"critical\"",
"fact_source": "poll:netdata",
"want_any": ["netdata", "диск", "критич", "алярм", "аларм", "тревог", "место", "памят", "critical", "alarm", "disk"],
"tags": ["ops", "bad-moment"]
},
{
"id": "routine-pills",
"rule": "routine:таблетки",
"severity": 2,
"since_minutes": 0,
"want_any": ["таблетк", "приня", "лекарств", "pill"],
"tags": ["routine"],
"note": "A routine: rule has no fact of its own, so there is no duration to name. The rule name is the only context."
},
{
"id": "routine-stretch",
"rule": "routine:зарядка",
"severity": 1,
"since_minutes": 0,
"want_any": ["зарядк", "размин", "упражн", "потянис", "разомн", "stretch", "exercise"],
"tags": ["routine"]
},
{
"id": "morning-checklist",
"rule": "morning:утро",
"severity": 2,
"since_minutes": 0,
"want_any": ["утр", "чеклист", "список", "не сделан", "осталось", "morning"],
"tags": ["routine"],
"note": "In production cmd/mavend/tick.go phrases morning routines deterministically and never calls the LLM. Scored anyway: the phraser is reachable with this rule name, and a fallback that garbles it is still a bug."
}
]
}
+16
View File
@@ -67,6 +67,22 @@ func NewLLMPhraser(ctx context.Context, cfg Config) (*LLMPhraser, error) {
return p, nil
}
// NewLLMPhraserAt wires a phraser to a llama-server that someone else started
// and owns. It spawns nothing, so Close does not kill anything.
//
// This exists for the phrasing scorer (internal/phraser/eval), which must
// measure the phrasing against a shared llama-server without taking the model
// load hit per run or killing a server another process depends on. The daemon
// still uses NewLLMPhraser and still owns its own child process.
func NewLLMPhraserAt(baseURL string, cfg Config) *LLMPhraser {
return &LLMPhraser{
cfg: cfg,
client: &http.Client{Timeout: cfg.Timeout},
port: strings.TrimSuffix(baseURL, "/"),
cancel: func() {},
}
}
func (p *LLMPhraser) start(ctx context.Context) error {
args := []string{
"-m", p.cfg.ModelPath,
+5 -1
View File
@@ -92,7 +92,11 @@ func (s *Stub) Close() error { return nil }
// the predicate fire (the same State the predicate saw).
func (s *Stub) PhraseNudge(_ context.Context, c loop.Candidate) (delivery.PhrasedNudge, error) {
body, summary := phraseNudge(c)
return delivery.PhrasedNudge{Candidate: c, Body: body, Summary: summary}, nil
// "neutral" rather than empty: Mood is part of the documented output
// contract and the Stub is a production fallback, so it must satisfy the
// contract too. Template phrasing has no tone to report, and neutral is the
// enum's own default.
return delivery.PhrasedNudge{Candidate: c, Body: body, Summary: summary, Mood: "neutral"}, nil
}
// PhraseReminder — extracts the user's text from the reminder payload (raw
+4 -4
View File
@@ -26,10 +26,10 @@ func TestPythonDateParser(t *testing.T) {
ctx := context.Background()
tests := []struct {
name string
text string
wantOK bool
checkT func(t *testing.T, got, now time.Time)
name string
text string
wantOK bool
checkT func(t *testing.T, got, now time.Time)
}{
{
name: "ru relative — через час",
+31
View File
@@ -19,6 +19,37 @@ type Embedder interface {
Close() error
}
// AsymmetricEmbedder — an embedder that wants to know whether a text is a
// search query or a stored passage. Recall is asymmetric: a short question
// goes in, a longer note comes out. The e5 family is trained for exactly that
// and needs the side written into the text ("query: " / "passage: ").
//
// Optional on purpose: HashEmbedder has no such notion, so callers go through
// EmbedQuery and EmbedPassage below, which fall back to plain Embed.
type AsymmetricEmbedder interface {
Embedder
EmbedQuery(ctx context.Context, text string) ([]float32, error)
EmbedPassage(ctx context.Context, text string) ([]float32, error)
}
// EmbedQuery embeds text that is being searched WITH — a question.
func EmbedQuery(ctx context.Context, e Embedder, text string) ([]float32, error) {
if a, ok := e.(AsymmetricEmbedder); ok {
return a.EmbedQuery(ctx, text)
}
return e.Embed(ctx, text)
}
// EmbedPassage embeds text that is being searched FOR — a note or a fact on
// its way into the store. Store and lookup must use these two calls, not one
// of them twice, or the asymmetry buys nothing.
func EmbedPassage(ctx context.Context, e Embedder, text string) ([]float32, error) {
if a, ok := e.(AsymmetricEmbedder); ok {
return a.EmbedPassage(ctx, text)
}
return e.Embed(ctx, text)
}
// HashEmbedder — a deterministic bag-of-words embedder used for tests and as a
// non-zero default floor. NOT semantically meaningful across languages; the
// real classifier swaps in the multilingual ONNX model wholesale.
+57
View File
@@ -37,3 +37,60 @@ func TestHashEmbedderCyrillic(t *testing.T) {
t.Fatalf("cosine(shared)=%.3f not > cosine(disjoint)=%.3f", cosine(a, b), cosine(a, c))
}
}
// recordingEmbedder — an asymmetric embedder that only remembers which side
// was asked for. Enough to pin the dispatch; real vectors need the model.
type recordingEmbedder struct{ calls []string }
func (r *recordingEmbedder) Dim() int { return 2 }
func (r *recordingEmbedder) Close() error { return nil }
func (r *recordingEmbedder) Embed(_ context.Context, _ string) ([]float32, error) {
r.calls = append(r.calls, "embed")
return []float32{1, 0}, nil
}
func (r *recordingEmbedder) EmbedQuery(_ context.Context, _ string) ([]float32, error) {
r.calls = append(r.calls, "query")
return []float32{1, 0}, nil
}
func (r *recordingEmbedder) EmbedPassage(_ context.Context, _ string) ([]float32, error) {
r.calls = append(r.calls, "passage")
return []float32{0, 1}, nil
}
// TestEmbedQueryAndPassageSplit — a question and a stored note must not take
// the same path. If both ended up on the same call the asymmetric model buys
// nothing, which is the whole reason for the swap.
func TestEmbedQueryAndPassageSplit(t *testing.T) {
rec := &recordingEmbedder{}
if _, err := EmbedQuery(context.Background(), rec, "где логи?"); err != nil {
t.Fatalf("EmbedQuery: %v", err)
}
if _, err := EmbedPassage(context.Background(), rec, "логи в /var/log"); err != nil {
t.Fatalf("EmbedPassage: %v", err)
}
if len(rec.calls) != 2 || rec.calls[0] != "query" || rec.calls[1] != "passage" {
t.Errorf("calls %v, want [query passage]", rec.calls)
}
}
// TestEmbedFallsBackToPlainEmbed — HashEmbedder has no sides, so both helpers
// must still work and give the same vector.
func TestEmbedFallsBackToPlainEmbed(t *testing.T) {
h := NewHashEmbedder(64)
q, err := EmbedQuery(context.Background(), h, "text")
if err != nil {
t.Fatalf("EmbedQuery: %v", err)
}
p, err := EmbedPassage(context.Background(), h, "text")
if err != nil {
t.Fatalf("EmbedPassage: %v", err)
}
for i := range q {
if q[i] != p[i] {
t.Fatalf("hash embedder gave two different vectors for the same text")
}
}
}
+2 -2
View File
@@ -185,8 +185,8 @@ func TestONNXBaseline(t *testing.T) {
if lib == "" {
t.Skip("MAVEN_ONNX_LIB unset — see AGENTS.md § Embedder model for intent routing")
}
model := filepath.Join("../../..", "models/embedder/model.onnx")
tok := filepath.Join("../../..", "models/embedder/tokenizer.json")
model := filepath.Join("../../..", "models/embedder/multilingual-e5-small/model_quantized.onnx")
tok := filepath.Join("../../..", "models/embedder/multilingual-e5-small/tokenizer.json")
for _, p := range []string{lib, model, tok} {
if _, err := os.Stat(p); err != nil {
t.Skipf("missing %s: %v", p, err)
+16 -4
View File
@@ -23,7 +23,11 @@ import (
//
// llama-server -m /mnt/hdd1/llms/qwen3.5/Qwen3.5-0.8B.Q4_K_M.gguf \
// --host 127.0.0.1 --port 18099 -c 2048 -ngl 99
// MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-router
// MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-models
//
// Every report name carries the model llama-server reports over /v1/models, so
// a bake-off across checkpoints (#278, #250) produces tables you can tell
// apart. Point the variable at one server at a time.
//
// Three configurations, because "the LLM router" is ambiguous and the three
// numbers answer different questions:
@@ -53,6 +57,14 @@ func TestLLMRouterBaseline(t *testing.T) {
}
ctx := context.Background()
model, err := ModelID(ctx, base)
if err != nil {
// Not fatal: an unlabelled score is still a score. But say so loudly,
// because an unlabelled row in a bake-off table is worthless.
t.Logf("could not read model id from %s: %v — reports will say %q", base, err, "unknown-model")
model = "unknown-model"
}
t.Logf("scoring model %s at %s", model, base)
lr := router.NewLLMRouter(client)
// llm-only: the LLM stage in isolation. Route returns (Decision, ok, err);
@@ -68,7 +80,7 @@ func TestLLMRouterBaseline(t *testing.T) {
}
return d, nil
})
repLLM, err := Score(ctx, "llm-only (0.8B, as deployed)", llmOnly, f)
repLLM, err := Score(ctx, "llm-only ("+model+", as deployed)", llmOnly, f)
if err != nil {
t.Fatalf("Score llm-only: %v", err)
}
@@ -77,7 +89,7 @@ func TestLLMRouterBaseline(t *testing.T) {
// cascade+llm: stage-0 grammar → LLM → classifier fallback, the wiring #320
// proposes. Hash embedder for the fallback so the classifier contribution is
// the deterministic floor and any lift is attributable to the model.
repCascade, err := Score(ctx, "cascade+llm (0.8B) + hash fallback",
repCascade, err := Score(ctx, "cascade+llm ("+model+") + hash fallback",
newBaselineRouter(t, router.NewHashEmbedder(1024), lr), f)
if err != nil {
t.Fatalf("Score cascade: %v", err)
@@ -96,7 +108,7 @@ func TestLLMRouterBaseline(t *testing.T) {
// either way. Kept so the question stays answered instead of being
// re-asked, and so internal/llm does NOT grow a chat_template_kwargs field
// for a problem that does not exist.
repNoThink, err := Score(ctx, "llm-only (0.8B, thinking off) [diagnostic]",
repNoThink, err := Score(ctx, "llm-only ("+model+", thinking off) [diagnostic]",
RouterFunc(func(ctx context.Context, u string, now time.Time) (router.Decision, error) {
d, ok, err := router.NewLLMRouter(&noThinkCompleter{base: base, http: &http.Client{Timeout: 60 * time.Second}}).Route(ctx, u, now)
if err != nil {
+51
View File
@@ -0,0 +1,51 @@
package eval
import (
"context"
"encoding/json"
"fmt"
"net/http"
"strings"
)
// ModelID asks llama-server which model it has loaded, so a scoring run can
// label itself. Without this a bake-off between two models produces two tables
// that look identical, and the operator has to remember which server was up.
//
// Read from the server rather than passed in on purpose: a hand-typed label
// goes stale the moment someone restarts the server with a different -m.
func ModelID(ctx context.Context, base string) (string, error) {
req, err := http.NewRequestWithContext(ctx, "GET", strings.TrimSuffix(base, "/")+"/v1/models", nil)
if err != nil {
return "", err
}
resp, err := http.DefaultClient.Do(req)
if err != nil {
return "", err
}
defer resp.Body.Close()
if resp.StatusCode != 200 {
return "", fmt.Errorf("models: status %d", resp.StatusCode)
}
var out struct {
Data []struct {
ID string `json:"id"`
} `json:"data"`
}
if err := json.NewDecoder(resp.Body).Decode(&out); err != nil {
return "", err
}
if len(out.Data) == 0 {
return "", fmt.Errorf("models: empty list")
}
return shortModelID(out.Data[0].ID), nil
}
// shortModelID trims the path and the .gguf suffix — llama-server reports the
// file name it was started with, which is too long for a table header.
func shortModelID(id string) string {
if i := strings.LastIndexAny(id, "/\\"); i >= 0 {
id = id[i+1:]
}
return strings.TrimSuffix(id, ".gguf")
}
+33 -3
View File
@@ -29,7 +29,7 @@ func NewLLMRouter(c Completer) *LLMRouter { return &LLMRouter{c: c} }
const routeGrammar = `
root ::= "[" ws action ("," ws action)* ws "]"
action ::= "{" ws "\"intent\"" ws ":" ws intent ("," ws field)* ws "}"
intent ::= "\"fact\"" | "\"reminder\"" | "\"note\"" | "\"query\"" | "\"act\"" | "\"chat\"" | "\"system\""
intent ::= "\"fact\"" | "\"reminder\"" | "\"note\"" | "\"query\"" | "\"act\"" | "\"chat\"" | "\"system\"" | "\"unknown\""
field ::= key ws ":" ws string
key ::= "\"key\"" | "\"value\"" | "\"text\"" | "\"verb\""
string ::= "\"" ([^"\\] | "\\" .){0,120} "\""
@@ -41,12 +41,17 @@ ws ::= [ \t\n]*
// question naming a fact key ("сколько воды я выпил с утра") matched the fact
// rule first and was stored as an assertion — 15 of 76 fixture cases.
//
// Changed again 31-07-2026: added the "unknown" escape hatch so the model can
// admit it cannot route (Vikunja #359).
//
// The training workspace keeps its own copy of this prompt for relabelling, and
// `llm/check_prompt_parity.py` there compares the two. That copy is in another
// repo and was not touched, so parity will fail until it gets the same edit.
// repo and was not touched, so parity will fail until it gets the same edits —
// both the rule reorder and the "unknown" wording (Vikunja #362).
const routeSystem = `Классифицируй ровно одно сообщение пользователя. Верни ОДИН JSON-массив действий.
Ровно одно намерение: fact, reminder, note, query, act, chat, system.
Есть восьмое значение unknown только для случаев, когда просьбу невозможно понять.
Классифицируй по цели пользователя. Порядок решения:
1. Хочет напоминание в будущем reminder
@@ -56,12 +61,14 @@ const routeSystem = `Классифицируй ровно одно сообще
5. Утверждает: сообщает или обновляет текущее состояние/событие fact
6. Просит выполнить работу act
7. Про ассистента, настройки или память system
8. Иначе chat
8. Реплика обрывок или указание на неназванное («это», «то», «потом»), и без него непонятно, что именно нужно сделать unknown
9. Иначе chat
Различия:
- note сохранить информацию, без напоминания. text = суть.
- reminder уведомить позже. text = что напомнить.
- fact неявное обновление: пользователь сообщает, что что-то в мире изменилось (текущее/изменённое состояние, случившееся событие). key/value.
- unknown редкий случай. Ставь его, только если в самой реплике нет ни предмета, ни действия. Короткая, простая или незнакомая тема это не причина для unknown: приветствие и болтовня это chat, вопрос на любую тему это query, просьба сделать что-то названное это act.
- query против fact решает форма реплики, а не тема. Вопрос о состоянии это query, даже если названо то же самое, что бывает в fact. Только утверждение это fact.
Примеры:
@@ -76,6 +83,13 @@ const routeSystem = `Классифицируй ровно одно сообще
"напиши письмо" {"intent":"act","verb":"написать письмо"}
"очисти память" {"intent":"system"}
"привет" {"intent":"chat","text":"привет"}
"сделай это" {"intent":"unknown"}
"ну это" {"intent":"unknown"}
"потом" {"intent":"unknown"}
Но не путай здесь unknown не нужен:
"сделай кофе" {"intent":"act","verb":"сделать кофе"}
"что такое кватернион?" {"intent":"query","text":"что такое кватернион"}
"ага" {"intent":"chat","text":"ага"}
Ответ JSON-массив: по одному объекту на каждую просьбу. Обычно один. Если в реплике несколько просьб по объекту на каждую. "напомни купить молоко, и запиши что кофе кончился" [{"intent":"reminder","text":"купить молоко"},{"intent":"note","text":"кофе кончился"}]. Только JSON, без пояснений.`
@@ -84,6 +98,11 @@ const routeSystem = `Классифицируй ровно одно сообще
// the loop without hurting short slot values.
const routeRepeatPenalty = 1.15
// routeIntentUnknown — the model's way of saying "I could not route this".
// It is a wire value only: it never becomes a router.Intent, it just makes
// Route return ok=false so the caller drops to the classifier cascade.
const routeIntentUnknown = "unknown"
type routeAction struct {
Intent string `json:"intent"`
Key string `json:"key"`
@@ -92,6 +111,10 @@ type routeAction struct {
Verb string `json:"verb"`
}
// Route asks the model for one decision. The bool is false when there is no
// decision to use: either the model failed (err set) or it refused with the
// "unknown" intent (err nil). Both mean the same thing to the caller — use the
// classifier instead.
func (lr *LLMRouter) Route(ctx context.Context, utterance string, now time.Time) (Decision, bool, error) {
raw, err := lr.c.Complete(ctx, llm.Req{
System: routeSystem,
@@ -115,6 +138,13 @@ func (lr *LLMRouter) Route(ctx context.Context, utterance string, now time.Time)
// with the engine turn-on (Router.Route → []Decision, both voice.go handlers
// loop). Until then only the first ask is honored.
a := acts[0]
// The model refused. Report "no decision" without an error, which is the
// same fall-through the caller already uses for a parse failure — the
// classifier cascade gets the turn and its own confidence gate decides
// whether to ask. Better a slower second opinion than a confident guess.
if a.Intent == routeIntentUnknown {
return Decision{}, false, nil
}
d := Decision{Utterance: utterance, Stage: 1, Confidence: 1.0}
switch Intent(a.Intent) {
case IntentFact:
+142 -1
View File
@@ -100,8 +100,10 @@ func TestLLMRouterReminderMapping(t *testing.T) {
}
}
// An intent name that is not in the contract at all (as opposed to "unknown",
// which is a real refusal) still defaults to chat.
func TestLLMRouterChatFallback(t *testing.T) {
lr := NewLLMRouter(mockLLM{out: `{"intent":"unknown"}`})
lr := NewLLMRouter(mockLLM{out: `{"intent":"banana"}`})
d, ok, err := lr.Route(context.Background(), "как дела?", time.Now())
if err != nil || !ok {
t.Fatalf("ok=%v err=%v", ok, err)
@@ -111,6 +113,61 @@ func TestLLMRouterChatFallback(t *testing.T) {
}
}
// The model must be able to say "I could not route this".
func TestRouteGrammarAllowsUnknown(t *testing.T) {
if !strings.Contains(routeGrammar, `"\"unknown\""`) {
t.Fatal("grammar cannot express a refusal")
}
}
// If the prompt does not tell the model when to refuse, it never will.
func TestRoutePromptExplainsUnknown(t *testing.T) {
if !strings.Contains(routeSystem, "unknown") {
t.Fatal("prompt never mentions the unknown intent")
}
if !strings.Contains(routeSystem, `"сделай это" → {"intent":"unknown"}`) {
t.Fatal("prompt lost its worked refusal example")
}
// A refusal-only router is useless, so the prompt must also show cases that
// look ambiguous but are not.
if !strings.Contains(routeSystem, "здесь unknown не нужен") {
t.Fatal("prompt lost its counter-examples")
}
}
// A refusal is not an error. It reports "no decision" so the cascade moves on.
func TestLLMRouterUnknownRefuses(t *testing.T) {
lr := NewLLMRouter(mockLLM{out: `{"intent":"unknown"}`})
_, ok, err := lr.Route(context.Background(), "сделай это", time.Now())
if ok {
t.Fatal("a refusal must not produce a usable decision")
}
if err != nil {
t.Fatalf("a refusal is not an error, got %v", err)
}
}
// The whole point of the refusal: the turn keeps going on the classifier, the
// same way it does when the model returns garbage.
func TestRouterFallsBackWhenLLMRefuses(t *testing.T) {
c := NewClassifier(NewHashEmbedder(1024))
seedClassifier(t, c)
r := New(Config{
Classifier: c,
Extractor: Extractor{Time: StubDateTimeParser{}, Facts: DefaultFactParser{}},
Threshold: 0.4,
LLM: NewLLMRouter(mockLLM{out: `{"intent":"unknown"}`}),
})
d, err := r.Route(context.Background(), "напомни позвонить маме", refNow())
if err != nil {
t.Fatalf("route: %v", err)
}
// Stage 1 is the LLM's own answer; the classifier lands on stage 2 or 3.
if d.Stage < 2 {
t.Fatalf("want the classifier to decide, got stage %d (%+v)", d.Stage, d)
}
}
func TestLLMRouterLLMError(t *testing.T) {
lr := NewLLMRouter(mockLLM{out: "", err: fmt.Errorf("llm down")})
_, ok, err := lr.Route(context.Background(), "x", time.Now())
@@ -118,3 +175,87 @@ func TestLLMRouterLLMError(t *testing.T) {
t.Fatal("want ok=false, err!=nil on llm error")
}
}
// --- slot extraction on top of an LLM decision --------------------------------
// newLLMTestRouter — a router whose route always comes from the mock model.
func newLLMTestRouter(t *testing.T, out string) *Router {
t.Helper()
c := NewClassifier(NewHashEmbedder(1024))
seedClassifier(t, c)
acts := DefaultActMatcher{Fns: []string{"restart", "stop", "run", "backup"}}
return New(Config{
Classifier: c,
Extractor: Extractor{Time: StubDateTimeParser{}, Acts: acts, Facts: DefaultFactParser{}},
Threshold: 0.4,
LLM: NewLLMRouter(mockLLM{out: out}),
})
}
// The model cannot produce a fire time, so without extraction every LLM-routed
// reminder was dropped as "no time".
func TestLLMDecisionGetsReminderTime(t *testing.T) {
r := newLLMTestRouter(t, `{"intent":"reminder","text":"позвонить маме"}`)
d, err := r.Route(context.Background(), "напомни позвонить маме через 2 часа", refNow())
if err != nil {
t.Fatalf("route: %v", err)
}
if d.Intent != IntentReminder {
t.Fatalf("want reminder, got %v", d.Intent)
}
if !d.Slots.HasTime || !d.Slots.Time.Equal(refNow().Add(2*time.Hour)) {
t.Fatalf("want time now+2h, got %+v", d.Slots)
}
if d.Slots.Text != "позвонить маме" {
t.Fatalf("extraction overwrote the model's text: %q", d.Slots.Text)
}
}
// No time in the utterance ⇒ no time in the slots. Do not invent one; the
// daemon says it could not read the time.
func TestLLMReminderWithoutTimeStaysEmpty(t *testing.T) {
r := newLLMTestRouter(t, `{"intent":"reminder","text":"позвонить маме"}`)
d, err := r.Route(context.Background(), "напомни позвонить маме", refNow())
if err != nil {
t.Fatalf("route: %v", err)
}
if d.Slots.HasTime {
t.Fatalf("invented a time: %v", d.Slots.Time)
}
}
// An act decision arrived with no Fn, so the tool never ran.
func TestLLMDecisionGetsActFn(t *testing.T) {
r := newLLMTestRouter(t, `{"intent":"act","verb":"restart nginx"}`)
d, err := r.Route(context.Background(), "слушай, restart nginx пожалуйста", refNow())
if err != nil {
t.Fatalf("route: %v", err)
}
if !d.Slots.HasFn || d.Slots.Fn != "restart" || len(d.Slots.Args) != 1 || d.Slots.Args[0] != "nginx" {
t.Fatalf("want fn=restart args=[nginx], got %+v", d.Slots)
}
}
// The model's own slots win; extraction only fills gaps.
func TestLLMSlotsWinOverExtraction(t *testing.T) {
r := newLLMTestRouter(t, `{"intent":"fact","key":"hydration","value":"выпил"}`)
d, err := r.Route(context.Background(), "я выпил воду", refNow())
if err != nil {
t.Fatalf("route: %v", err)
}
if d.Slots.Key != "hydration" {
t.Fatalf("extraction overwrote the model's key: %q", d.Slots.Key)
}
}
// A fact the model left keyless still gets one from the parser.
func TestLLMFactGetsKeyFromParser(t *testing.T) {
r := newLLMTestRouter(t, `{"intent":"fact","text":"я выпил воду"}`)
d, err := r.Route(context.Background(), "я выпил воду", refNow())
if err != nil {
t.Fatalf("route: %v", err)
}
if !d.Slots.HasKey || d.Slots.Key != "water" {
t.Fatalf("want key=water, got %+v", d.Slots)
}
}
+28 -1
View File
@@ -12,6 +12,15 @@ import (
"golang.org/x/text/unicode/norm"
)
// The deployed model is multilingual-e5-small. e5 was trained with these two
// words glued to the front of every text, and it scores badly without them —
// they are part of the model, not a style choice. Swapping back to a symmetric
// paraphrase model means dropping them again.
const (
queryPrefix = "query: "
passagePrefix = "passage: "
)
const (
padTokenID = 1
unkTokenID = 3
@@ -54,7 +63,25 @@ func NewONNXEmbedder(modelPath, tokenizerPath, libPath string) (*onnxEmbedder, e
func (e *onnxEmbedder) Dim() int { return embedDim }
// Embed treats the text as a query. The classifier compares one short
// utterance to another short seed phrase, so both sides get the same prefix
// and the comparison stays fair. The recall path must call EmbedQuery and
// EmbedPassage instead.
func (e *onnxEmbedder) Embed(ctx context.Context, text string) ([]float32, error) {
return e.embed(ctx, queryPrefix+text)
}
// EmbedQuery — the question the user just asked.
func (e *onnxEmbedder) EmbedQuery(ctx context.Context, text string) ([]float32, error) {
return e.embed(ctx, queryPrefix+text)
}
// EmbedPassage — a note or fact being stored, or re-scored at lookup time.
func (e *onnxEmbedder) EmbedPassage(ctx context.Context, text string) ([]float32, error) {
return e.embed(ctx, passagePrefix+text)
}
func (e *onnxEmbedder) embed(ctx context.Context, text string) ([]float32, error) {
inputIDs, attentionMask, _ := e.tokenizer.Encode(text)
inputShape := ort.NewShape(1, int64(maxLength))
@@ -313,4 +340,4 @@ func preTokenize(text string) []string {
return out
}
var _ Embedder = (*onnxEmbedder)(nil)
var _ AsymmetricEmbedder = (*onnxEmbedder)(nil)
+34
View File
@@ -88,6 +88,7 @@ func (r *Router) Route(ctx context.Context, utterance string, now time.Time) (De
if r.llm != nil {
if d, ok, err := r.llm.Route(ctx, utterance, now); err == nil && ok {
d.Utterance = utterance
r.fillSlots(ctx, &d, now)
return d, nil
} else if err != nil {
log.Printf("router: llm route fell back to classifier: %v", err)
@@ -118,6 +119,39 @@ func (r *Router) Route(ctx context.Context, utterance string, now time.Time) (De
return d, nil
}
// fillSlots — run stage-2 extraction on an LLM decision and fill only the slots
// the model left empty. The LLM wins where it answered: it saw the sentence, the
// parsers are keyword tables. Extraction covers what the model cannot produce at
// all — a parsed reminder time and an allowlist fn.
//
// If a reminder still has no time, leave it missing. The daemon then says it
// could not read the time; inventing one would set a wrong alarm.
func (r *Router) fillSlots(ctx context.Context, d *Decision, now time.Time) {
ex := r.extractor.Extract(ctx, d.Intent, d.Utterance, now)
if !d.Slots.HasTime && ex.HasTime {
d.Slots.Time, d.Slots.HasTime = ex.Time, ex.HasTime
}
if !d.Slots.HasKey && ex.HasKey {
d.Slots.Key, d.Slots.Value, d.Slots.HasKey = ex.Key, ex.Value, ex.HasKey
}
if !d.Slots.HasFn && ex.HasFn {
d.Slots.Fn, d.Slots.Args, d.Slots.HasFn = ex.Fn, ex.Args, ex.HasFn
}
// For an act the model returns the verb in Text ("restart nginx"), which is
// often cleaner than the raw utterance ("maven, could you restart nginx").
// Try it too when the utterance did not match the allowlist.
if d.Intent == IntentAct && !d.Slots.HasFn && r.extractor.Acts != nil &&
d.Slots.Text != "" && d.Slots.Text != d.Utterance {
if fn, args, ok := r.extractor.Acts.Match(d.Slots.Text); ok {
d.Slots.Fn, d.Slots.Args, d.Slots.HasFn = fn, args, true
}
}
if d.Slots.Text == "" {
d.Slots.Text = ex.Text
}
// Stage stays 1: it says who decided the route, and that was the LLM.
}
// CorrectMisroute — the user corrected a bad classification. Appends a new
// example for the corrected intent (append-only — grows the classifier, no
// retrain). Same shape as nudges.outcome tuning cooldowns: more reliable over
+41
View File
@@ -55,6 +55,47 @@ func Validate(routines []Routine) error {
return nil
}
// Accepted — an accepted routine proposal as the tick driver sees it. This is a
// different shape from Routine: the schedule is a plain interval the pattern
// detector measured, not an operator-written cron expression. Accepted is when
// the human said yes; LastFired is nil until the first nudge.
type Accepted struct {
ID int64
Name string
IntervalDays float64
Accepted time.Time
LastFired *time.Time
}
// DueAccepted returns the accepted routines whose interval has passed. It does
// not mutate anything — the caller persists the new last-fired time, because
// that has to survive a restart (unlike Due's in-memory map).
//
// The clock starts at LastFired, or at Accepted for a routine that has never
// nudged. A routine with a non-positive interval never fires: a bad interval
// should mean silence, not a nudge every tick.
//
// One occurrence per call, no catch-up: the caller stamps the fire time as now,
// so a routine that was silent for a month nudges once and then waits a full
// interval. Never a backlog.
func DueAccepted(rs []Accepted, now time.Time) []Accepted {
var out []Accepted
for _, r := range rs {
if r.IntervalDays <= 0 {
continue
}
since := r.Accepted
if r.LastFired != nil {
since = *r.LastFired
}
gap := time.Duration(r.IntervalDays * 24 * float64(time.Hour))
if !now.Before(since.Add(gap)) {
out = append(out, r)
}
}
return out
}
// Due returns the routines whose schedule crossed since their last fire and
// records now as the new last-fire time for each one returned. The caller owns
// `last` (the tick driver holds it across ticks); Due mutates it in place.
+34
View File
@@ -5,6 +5,40 @@ import (
"time"
)
func TestDueAcceptedFiresOncePerInterval(t *testing.T) {
accepted := time.Date(2026, 7, 1, 9, 0, 0, 0, time.UTC)
fired := accepted.Add(3 * 24 * time.Hour)
rs := []Accepted{
{ID: 1, Name: "полить цветы", IntervalDays: 3, Accepted: accepted},
{ID: 2, Name: "покормить рыб", IntervalDays: 3, Accepted: accepted, LastFired: &fired},
{ID: 3, Name: "битый интервал", IntervalDays: 0, Accepted: accepted},
}
// One day in: nothing has waited a full interval.
if got := DueAccepted(rs, accepted.Add(24*time.Hour)); len(got) != 0 {
t.Fatalf("want nothing due after 1 day, got %+v", got)
}
// Three days in: the never-fired one is due. The one that already fired at
// day 3 starts its next three days from there. A zero interval never fires.
got := DueAccepted(rs, fired)
if len(got) != 1 || got[0].ID != 1 {
t.Fatalf("want only routine 1 due at day 3, got %+v", got)
}
// Six days in: both real routines are due.
if got := DueAccepted(rs, accepted.Add(6*24*time.Hour)); len(got) != 2 {
t.Fatalf("want both routines due at day 6, got %+v", got)
}
// A month later the zero-interval routine is still silent.
for _, r := range DueAccepted(rs, accepted.Add(30*24*time.Hour)) {
if r.ID == 3 {
t.Fatal("a routine with a zero interval must never fire")
}
}
}
func TestValidate(t *testing.T) {
ok := []Routine{{Name: "morning", Cron: "0 8 * * *", Body: "доброе утро"}}
if err := Validate(ok); err != nil {
+2
View File
@@ -72,6 +72,8 @@ ALTER TABLE reminders ADD COLUMN next_fire_ts INTEGER;`, // #2
CREATE INDEX IF NOT EXISTS idx_facts_resolution_pending ON facts (resolution_state) WHERE resolution_state = 'pending';`, // #7 — entity-aware memory (Vikunja #279): facts about a subject get resolved to a Nexus entity_id async
`CREATE INDEX IF NOT EXISTS idx_nudges_snoozed ON nudges (outcome_ts) WHERE outcome = 'snoozed';`, // #8 — SnoozedUntil runs every tick; keep it off a full scan (Vikunja #364)
`ALTER TABLE proposed_routines ADD COLUMN accepted_ts INTEGER;
ALTER TABLE proposed_routines ADD COLUMN last_fired_ts INTEGER;`, // #9 — accepted routines keep firing (Vikunja #366): the tick loop needs to know when a routine was accepted and when it last nudged
}
// migrate applies every migration with a number greater than the DB's current
+70 -20
View File
@@ -16,10 +16,15 @@ const (
RoutineDismissed = "dismissed"
)
// ProposedRoutine — a detected pattern the system wants to turn into a
// recurring reminder. Status 'proposed' means awaiting human confirmation;
// 'accepted' means the human confirmed and a reminder was created (reminder_id
// set); 'dismissed' means the human declined and we won't re-propose.
// ProposedRoutine — a detected pattern the system wants to nudge about on a
// repeating interval. Status 'proposed' means awaiting human confirmation;
// 'accepted' means the human confirmed and the tick loop now owns the schedule;
// 'dismissed' means the human declined and we won't re-propose.
//
// AcceptedTs is when the human said yes; it is the clock start for the first
// nudge. LastFiredTs is when the last nudge went out, nil until the first one.
// ReminderID is only set on rows accepted before Vikunja #366, when accepting
// created a one-shot reminder instead.
type ProposedRoutine struct {
ID int64
Action string
@@ -27,7 +32,9 @@ type ProposedRoutine struct {
IntervalDays float64
Status string // proposed | accepted | dismissed
CreatedTs time.Time
ReminderID *int64 // set when accepted
ReminderID *int64
AcceptedTs *time.Time
LastFiredTs *time.Time
}
var (
@@ -74,7 +81,7 @@ func (s *Store) CreateProposedRoutine(ctx context.Context, action, object string
// nil (no error) when no row exists.
func (s *Store) LookupProposedRoutine(ctx context.Context, action, object string) (*ProposedRoutine, error) {
row := s.db.QueryRowContext(ctx, `
SELECT id, action, object, interval_days, status, created_ts, reminder_id
SELECT id, action, object, interval_days, status, created_ts, reminder_id, accepted_ts, last_fired_ts
FROM proposed_routines
WHERE action = ? AND object = ?`, action, object)
r, err := scanProposedRoutine(row)
@@ -95,12 +102,8 @@ func (s *Store) ListProposedRoutines(ctx context.Context) ([]ProposedRoutine, er
// ListProposedRoutinesByStatus returns routines in one status, newest first.
// An empty status returns every row.
//
// TODO(vikunja#46): the tick loop should read the accepted ones from here so a
// routine the human said yes to has a home the loop can see, instead of only
// the reminder row that accepting happened to create.
func (s *Store) ListProposedRoutinesByStatus(ctx context.Context, status string) ([]ProposedRoutine, error) {
q := `SELECT id, action, object, interval_days, status, created_ts, reminder_id
q := `SELECT id, action, object, interval_days, status, created_ts, reminder_id, accepted_ts, last_fired_ts
FROM proposed_routines`
var args []any
if status != "" {
@@ -133,15 +136,14 @@ func (s *Store) ListProposedRoutinesByStatus(ctx context.Context, status string)
// `AND status = 'proposed'` makes the move one-way: an answered proposal can
// never be answered again.
//
// AcceptProposedRoutine flips status to 'accepted', links a reminder_id.
// Returns error if not in 'proposed' status.
//
// TODO(vikunja#46): the /routines page calls this through ipc to flip status
// from the authed surface.
func (s *Store) AcceptProposedRoutine(ctx context.Context, id, reminderID int64) error {
// AcceptProposedRoutine flips status to 'accepted' and records when. From that
// timestamp the tick loop owns the schedule: it re-reads accepted rows every
// tick and nudges when the interval has passed. Returns an error if the row is
// not in 'proposed' status.
func (s *Store) AcceptProposedRoutine(ctx context.Context, id int64, ts time.Time) error {
res, err := s.db.ExecContext(ctx,
`UPDATE proposed_routines SET status = 'accepted', reminder_id = ? WHERE id = ? AND status = 'proposed'`,
reminderID, id)
`UPDATE proposed_routines SET status = 'accepted', accepted_ts = ? WHERE id = ? AND status = 'proposed'`,
ts.UnixMilli(), id)
if err != nil {
return fmt.Errorf("accept proposed routine: %w", err)
}
@@ -152,6 +154,42 @@ func (s *Store) AcceptProposedRoutine(ctx context.Context, id, reminderID int64)
return nil
}
// ListAcceptedRoutines returns every accepted routine, oldest first. The tick
// loop reads this each tick and decides which ones are due.
func (s *Store) ListAcceptedRoutines(ctx context.Context) ([]ProposedRoutine, error) {
rows, err := s.db.QueryContext(ctx, `
SELECT id, action, object, interval_days, status, created_ts, reminder_id, accepted_ts, last_fired_ts
FROM proposed_routines
WHERE status = 'accepted'
ORDER BY id`)
if err != nil {
return nil, fmt.Errorf("list accepted routines: %w", err)
}
defer rows.Close()
var out []ProposedRoutine
for rows.Next() {
r, err := scanProposedRoutine(rows)
if err != nil {
return nil, err
}
out = append(out, r)
}
return out, rows.Err()
}
// MarkRoutineFired records that a routine just nudged. The stored time is the
// nudge time, not the time it was theoretically due, so a routine that was
// silent for a while starts its next interval from now — missed occurrences are
// dropped, never replayed as a backlog.
func (s *Store) MarkRoutineFired(ctx context.Context, id int64, ts time.Time) error {
if _, err := s.db.ExecContext(ctx,
`UPDATE proposed_routines SET last_fired_ts = ? WHERE id = ?`,
ts.UnixMilli(), id); err != nil {
return fmt.Errorf("mark routine fired: %w", err)
}
return nil
}
// DismissProposedRoutine flips status to 'dismissed'. Idempotent.
func (s *Store) DismissProposedRoutine(ctx context.Context, id int64) error {
_, err := s.db.ExecContext(ctx,
@@ -168,12 +206,24 @@ func scanProposedRoutine(sc scanner) (ProposedRoutine, error) {
var r ProposedRoutine
var created int64
var reminderID sql.NullInt64
if err := sc.Scan(&r.ID, &r.Action, &r.Object, &r.IntervalDays, &r.Status, &created, &reminderID); err != nil {
var accepted, lastFired sql.NullInt64
if err := sc.Scan(&r.ID, &r.Action, &r.Object, &r.IntervalDays, &r.Status, &created, &reminderID, &accepted, &lastFired); err != nil {
return ProposedRoutine{}, err
}
r.CreatedTs = time.UnixMilli(created).UTC()
if reminderID.Valid {
r.ReminderID = &reminderID.Int64
}
r.AcceptedTs = millisToTime(accepted)
r.LastFiredTs = millisToTime(lastFired)
return r, nil
}
// millisToTime turns a nullable unix-millis column into a *time.Time.
func millisToTime(v sql.NullInt64) *time.Time {
if !v.Valid {
return nil
}
t := time.UnixMilli(v.Int64).UTC()
return &t
}
+47 -14
View File
@@ -43,12 +43,7 @@ func TestCreateAndAcceptProposedRoutine(t *testing.T) {
}
// Accept
// First create a reminder to link
remID, err := s.CreateReminder(ctx, now.Add(7*24*time.Hour), `{"text":"refill cat water"}`, "0 10 * * 0")
if err != nil {
t.Fatalf("CreateReminder: %v", err)
}
if err := s.AcceptProposedRoutine(ctx, id, remID); err != nil {
if err := s.AcceptProposedRoutine(ctx, id, now); err != nil {
t.Fatalf("AcceptProposedRoutine: %v", err)
}
@@ -60,8 +55,50 @@ func TestCreateAndAcceptProposedRoutine(t *testing.T) {
if r.Status != "accepted" {
t.Fatalf("want status=accepted, got %s", r.Status)
}
if r.ReminderID == nil || *r.ReminderID != remID {
t.Fatalf("want reminder_id=%d, got %v", remID, r.ReminderID)
if r.AcceptedTs == nil || !r.AcceptedTs.Equal(now.Truncate(time.Millisecond)) {
t.Fatalf("want accepted_ts=%v, got %v", now, r.AcceptedTs)
}
if r.LastFiredTs != nil {
t.Fatalf("a freshly accepted routine has not fired yet, got %v", r.LastFiredTs)
}
}
// TestAcceptedRoutineFiredTimestamp — the tick loop's two reads: the accepted
// list, and the last-fired stamp it writes back after a nudge.
func TestAcceptedRoutineFiredTimestamp(t *testing.T) {
s := newTestStore(t)
ctx := context.Background()
now := time.Now().UTC().Truncate(time.Millisecond)
id, err := s.CreateProposedRoutine(ctx, "полить", "цветы", 3.0, now)
if err != nil {
t.Fatalf("CreateProposedRoutine: %v", err)
}
if err := s.AcceptProposedRoutine(ctx, id, now); err != nil {
t.Fatalf("AcceptProposedRoutine: %v", err)
}
list, err := s.ListAcceptedRoutines(ctx)
if err != nil {
t.Fatalf("ListAcceptedRoutines: %v", err)
}
if len(list) != 1 || list[0].ID != id {
t.Fatalf("want the one accepted routine, got %+v", list)
}
if list[0].IntervalDays != 3.0 {
t.Fatalf("want interval_days=3, got %v", list[0].IntervalDays)
}
fired := now.Add(3 * 24 * time.Hour)
if err := s.MarkRoutineFired(ctx, id, fired); err != nil {
t.Fatalf("MarkRoutineFired: %v", err)
}
list, err = s.ListAcceptedRoutines(ctx)
if err != nil {
t.Fatalf("ListAcceptedRoutines: %v", err)
}
if list[0].LastFiredTs == nil || !list[0].LastFiredTs.Equal(fired) {
t.Fatalf("want last_fired_ts=%v, got %v", fired, list[0].LastFiredTs)
}
}
@@ -165,7 +202,7 @@ func TestDismissedProposedRoutineStaysDismissed(t *testing.T) {
if err := s.DismissProposedRoutine(ctx, id); err != nil {
t.Fatalf("second DismissProposedRoutine: %v", err)
}
if err := s.AcceptProposedRoutine(ctx, id, 1); !errors.Is(err, ErrProposedRoutineNotFound) {
if err := s.AcceptProposedRoutine(ctx, id, time.Now().UTC()); !errors.Is(err, ErrProposedRoutineNotFound) {
t.Fatalf("want ErrProposedRoutineNotFound accepting a dismissed routine, got %v", err)
}
r, err := s.LookupProposedRoutine(ctx, "clean", "litter_box")
@@ -190,11 +227,7 @@ func TestListProposedRoutinesByStatus(t *testing.T) {
if err != nil {
t.Fatalf("CreateProposedRoutine: %v", err)
}
remID, err := s.CreateReminder(ctx, now.Add(4*24*time.Hour), `{"text":"water plants"}`, "")
if err != nil {
t.Fatalf("CreateReminder: %v", err)
}
if err := s.AcceptProposedRoutine(ctx, keep, remID); err != nil {
if err := s.AcceptProposedRoutine(ctx, keep, now); err != nil {
t.Fatalf("AcceptProposedRoutine: %v", err)
}
if err := s.DismissProposedRoutine(ctx, drop); err != nil {
+6 -1
View File
@@ -55,4 +55,9 @@ func spokenDate(dd, mm, yyyy string) string {
}
func mustInt(s string) int { n, _ := strconv.Atoi(s); return n }
func gap(y string) string { if y == "" { return "" }; return " " + y }
func gap(y string) string {
if y == "" {
return ""
}
return " " + y
}