Compare commits

...

10 Commits

Author SHA1 Message Date
kami 891136c65d gofmt the kiwix client and rewrite test
PR #41 and #44 landed these two files unformatted, so the gofmt gate that
PR #12 added to `make test` failed as soon as both were on one branch.
Struct-tag and comment alignment only, no semantic change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 21:34:03 +04:00
kami 41c7c13f42 Merge remote-tracking branch 'origin/overnight/snooze-works' into integration/jul31
# Conflicts:
#	internal/store/migrations.go
2026-07-31 21:32:18 +04:00
kami a324e8f624 Merge remote-tracking branch 'origin/overnight/eval-writeup' into integration/jul31 2026-07-31 21:31:36 +04:00
kami 51805e7f35 Merge remote-tracking branch 'origin/overnight/kiwix-rewrite' into integration/jul31 2026-07-31 21:31:36 +04:00
kami 742b2ad1d7 Score the Russian-to-keywords rewrite end to end (#403)
Same 9 cases as the retrieval eval, so the numbers compare directly:
hand-written keywords hit 8 of 8, this is what the model reaches on its
own. Reports the hand-written query next to the model's for every case,
because where the phrasing differs is the useful part.

Opt-in on MAVEN_KIWIX_URL + MAVEN_LLM_URL, like the other evals.

Result on Qwen3.5-0.8B: 3 of 8, identical on all three runs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:19:56 +04:00
kami b300ac5c70 Rewrite a Russian question into English Kiwix keywords (#403)
Kiwix ranks by keyword, not meaning, so a translated question finds song
and TV titles. This asks the resident model for the TOPIC instead: a short
English noun phrase, like a Wikipedia article title.

Locked down three ways, because a wrong query is silently wrong:
- A GBNF grammar, same idea as routeGrammar and responseGrammar. The
  reply must be {"query":"..."} with Latin words only. The JSON wrapper
  matters: this model always thinks out loud and this llama-server build
  ignores the thinking switch, so a bare word-list grammar just captured
  "Let me analyze this request carefully" for every question.
- max_tokens 32, since the answer is a few words.
- CleanQuery, which throws away empty, Russian and prose replies rather
  than passing them to Kiwix, and drops question words like "why" and
  "how much" that a keyword ranker cannot use anyway.

Client side only. Nothing is wired into the daemon or any config.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:19:42 +04:00
kami 0b90952e55 Write down every conversational eval score from tonight
Records all four configurations on the 27-case talk fixture, three runs
each: no grammar, plus grammar, plus Russian prompts, plus the truncation
fix. Composite, per-path and per-check, with the reproduce command.

The short version is that the plumbing got fixed and the score barely
moved. Grammar was the real win. Russian prompts helped a little and cut
latency by 5x. The truncation fix was necessary and bought nothing.

Also writes down three things that are easy to lose:

- The truncation cause was the grammar's 400-character bound, not the
  token cap. Measured at three caps, same 400 characters every time.
- Then I set the bound to 1000 against a 768-token cap and made it worse.
  The two limits have to agree.
- One run is contaminated and marked void: I ran an agent against the same
  llama-server, and the report still claimed zero errors while a third of
  the fixture silently answered "не знаю.". That is #397 and it is worse
  than filed — a busy server is indistinguishable from bad phrasing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 18:18:19 +04:00
kami ddb658ffbb Add a Kiwix search client and score retrieval (Vikunja #403)
Step one of letting Maven read instead of recall. No LLM yet.

internal/kiwix/client.go: search a local Kiwix server, parse the RSS
reply, hand back title + path + plain-text snippet + word count. The
snippet is the unit of context; a full article is ~100KB of HTML and
will not fit a 4096 token window.

internal/kiwix/retrieval_eval.go plus knowledge_v1.json: the 9 knowledge
questions from the phrasing fixture, each with hand-written English
keywords, scored on whether a wanted article comes back in the top 5.
Opt-in via MAVEN_KIWIX_URL, since CI has no Kiwix. No pass bar, the
number is the finding.

Result on the live mirror: 8/8 answerable questions hit, 7 of them at
rank 1. Retrieval works. Keywords are written by hand on purpose, since
Kiwix ranks by keyword and not by meaning, so a natural question fails.
A query-rewrite step is the next piece of work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 17:43:44 +04:00
kami ee3e6a9eaf Wire the snooze read into the Gatherer and honour it for reminders (#364)
The Gatherer now fills State.SnoozeUntil from store.SnoozedUntil instead
of nil, so a snooze finally reaches the gate. RemindDecisions gains the
one restraint check that applies to a reminder — quiet hours, presence
and cooldown are still bypassed, so "wake me 7" is unchanged. Reviewer:
the two tests in internal/loop/gate_test.go are the contract.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:33:51 +04:00
kami 8acb8a97c6 Read the recorded snooze outcomes back out of the nudges table (#364)
The gate honours State.SnoozeUntil but nothing ever filled it. New
store.SnoozedUntil returns, per rule, when the newest snooze runs out.
Reviewer: the fixed 2h SnoozeDuration and its reasoning in nudges.go —
nothing upstream can supply a per-nudge length, so no new column.
Expired snoozes are dropped in SQL, so silence can never be permanent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 02:33:42 +04:00
9 changed files with 971 additions and 0 deletions
+150
View File
@@ -0,0 +1,150 @@
# Conversational phrasing eval — 31-07-2026
Every score measured tonight, on the three paths the nudge eval never touched:
chat, query-with-notes, and general knowledge.
**Short version: the plumbing got fixed and the score barely moved.** Grammar and
Russian prompts together took the composite from ~9 to ~14 of 27. Everything
still failing is the model not knowing things or not holding a constraint, and
prompting is out of levers. Settles the measurement half of Vikunja #395 / #398 /
#400.
## How to reproduce
```sh
# llama-server: -c 4096 -ngl 99 -t 6, model /mnt/hdd1/llms/qwen3.5/Qwen3.5-0.8B.Q4_K_M.gguf
MAVEN_LLM_URL=http://127.0.0.1:18099 no_proxy=127.0.0.1,localhost \
deps/go/go/bin/go test -count=1 -timeout 40m \
-run TestLLMTalkBaseline ./internal/phraser/eval/ -v
```
Three runs per configuration, always. The fixture is 27 cases, so one reply
changing moves the composite by 3.7 points — a single run cannot tell a real
change from sampling noise. This was learned the expensive way: an earlier claim
that "one nudge case fails every run" turned out to be three different cases
across three runs.
**Run the box otherwise idle.** See the contamination note at the bottom.
## Composite, per configuration
| config | overall /27 | chat /9 | query /9 | knowledge /9 | canned fallbacks |
|---|---|---|---|---|---|
| baseline, no grammar | 7, 12, 7 | 1, 1, 0 | 2, 4, 2 | 4, 7, 5 | 0, 0, 0 |
| + GBNF grammar (#398) | 14, 15, 8 | 1, 3, 0 | 5, 6, 3 | 8, 6, 5 | 0, 0, 0 |
| + Russian prompts (#400) | 11, 17, 15 | 1, 5, 3 | 5, 6, 8 | 5, 6, 4 | 0, 0, 0 |
| + truncation fix, 1000ch/768tok | 12, 13, 10 | 2, 2, 1 | 7, 7, 5 | 3, 4, 4 | 3, 3, 6 |
| + rebalanced, 600ch/1024tok | **void — contaminated** | | | | |
"Canned fallbacks" counts replies that came back as the hardcoded `"не знаю."`
or `"поговорили."`. It is not a check, it is a health signal: those strings mean
the phraser gave up, and the eval scores them as ordinary bad replies.
## Per-check
| check | no grammar | + grammar | + RU prompts | + truncation fix |
|---|---|---|---|---|
| nonempty | 27, 27, 27 | 27, 27, 27 | 27, 27, 27 | 27, 27, 27 |
| ellipsis | 20, 19, 23 | 27, 27, 27 | 27, 27, 27 | 27, 27, 27 |
| lang | 13, 16, 15 | 23, 26, 26 | 25, 26, 25 | 26, 27, 27 |
| feminine | — | — | 25, 24, 26 | 25, 25, 27 |
| address | — | — | 21, 22, 22 | 22, 21, 22 |
| ontopic | — | — | 17, 24, 18 | 17, 19, 14 |
`nonempty` reading 27/27 everywhere is not good news — it was a broken check.
It tested for a non-blank string, so replies of literally `{` and `"15-16"`
passed it. Fixed on `overnight/fix-truncation`; it needs a letter now.
## What each change actually bought
**GBNF grammar (#398) — the biggest single win.** Qwen3.5-0.8B writes
`Thinking Process:` as plain text with no tags, `stripThink` only handles
`</think>`, so the JSON never closed and the plain-text fallback shipped the
literal reasoning. `ellipsis` went 20→27 and `lang` 13→26. The router had been
using a grammar for ages; the phraser asking nicely in the prompt was the
oversight.
**Russian prompts (#400) — modest, plus a large latency win.** Chat 1.3→3.0
average, query 4.7→6.3, knowledge 6.3→5.0. All inside the run-to-run spread, so
"probably better on the paths it targeted, not provable in three runs". p50
latency dropped from ~11.5s to ~2.3s and that part is consistent across all
three runs — shorter prompts, and she stopped emitting English reasoning first.
**Truncation fix — necessary, and did not help the score.** Two real bugs
(replies of `{`, and a `nonempty` check that passed them), both fixed, and the
composite went nowhere. A complete rambling wrong answer fails the same checks a
truncated one did. Worth doing anyway: the daemon was shipping `{` to a
text-to-speech voice.
## The truncation bug, since the cause was counter-intuitive
The grammar's `string ::= ... {0,400}` rule was the cause, not the token cap.
Measured against Qwen3.5-0.8B at three caps — 256, 768 and 2048 — the reply came
back **exactly 400 characters every time, cut mid-word** (`"Нужно записать и,"`).
Then I raised the bound to 1000 while the cap was 768 tokens and made it worse:
Russian runs ~1.5 characters per token here, so generation died on the *token*
cap instead, mid-object, and the new guard correctly refused it and shipped
`"не знаю."` — 3, 3 and 6 fallbacks per run, from zero. **The two limits have to
agree.** 600 characters needs ~400 tokens; the cap is 1024.
## Where the remaining failures live
`address` is stuck at 21-22 of 27 and `ontopic` at 14-19. Both resist prompting.
**The prompt now explicitly forbids exactly what she does.** It says never "вы",
use the singular — and she writes `вашей`, `подождите`, `делаете`, `хотите`,
`напишите`. Telling a 0.8B "never do X" does not work. Same for
`feminine`: `я готов`, `я понял`, `я нашел`, `я заметил`, `я сказал`.
**Some of `ontopic` is the fixture, not the model.** `chat-how-are-you` got
`"Привет! Я здесь, чтобы поговорить. Как дела сегодня?"` — a fine reply that
fails because `want_any` is `[норм, хорош, порядк, тут, работ]`. It fails in
every run, so it inflates the count. The `ontopic` column currently measures the
fixture as much as the model. Not fixed yet, deliberately: changing it would
break comparability with the runs above.
**Two replies worth reading, because they are not fixable by prompting:**
- Thunder and lightning: *"Скорость молнии — 8-10 тысяч километров в секунду, но
звук — 300 метров в секунду, что делает молнию громче."* Confidently wrong,
and it concludes lightning is *louder* rather than sound being *slower*.
- "расскажи обо мне": *"Ты — прекрасное существо, с душой и вниманием… Спасибо за
твою улыбку… О тебе — заповедь любви."* Sycophantic filler, zero information,
and precisely the "not a relationship" non-goal.
- Boiling an egg: `"15-16"` one run, `"1"` another. No unit, wrong number.
The first argues for reading instead of recalling (#403 — Kiwix retrieval scores
8/8 on the same questions given English keywords). The second and third argue
for templates on the paths where correctness matters (#392).
## Contamination note — how the last row got voided
I started the query-rewrite agent against the same llama-server the sweep was
using, and assumed contention would only affect latency. It did not. The
knowledge path collapsed to 0 of 9 with eight canned `"не знаю."` replies, p95
tripled to 23.7s, and **the report still said "0 errors"**.
That is Vikunja #397, and it is worse than filed: a merely *busy* server
produces a clean-looking report with a third of the fixture silently answering
`"не знаю."`. `PhraseChat` and `PhraseQuery` swallow every failure and return a
hardcoded string, so infrastructure trouble is indistinguishable from bad
phrasing in the score. The talk test guards the *start* and *end* of a run with
a model check, which catches a dead server but not a loaded one.
**Until #397 is fixed, treat any run made on a busy box as void.**
## Next
- Re-run 600ch/1024tok clean, to fill the void row.
- Score `Qwen3.5-2B-UD-Q4_K_XL` (already at `/mnt/hdd1/llms/qwen3.5/`, never
measured) on this fixture and the router fixture. Not the 4B — too big for
this box, owner's call.
- Newer sub-500M candidates (LFM2.5 200M/300M) are worth a run for routing.
Note `MODEL-BAKEOFF-31-07-2026.md` found LFM2.5-**1.2B** worse than
Qwen3.5-0.8B at Russian routing and 2.4× slower — but those are a different,
older generation, so that result does not predict the small ones.
- Fix `chat-how-are-you`'s `want_any`, and re-baseline once, so `ontopic`
measures the model.
- #397 first if anything, since it decides whether any of the above is
trustworthy.
+116
View File
@@ -0,0 +1,116 @@
// Package kiwix reads a local Kiwix server (offline Wikipedia and friends).
//
// Why: the resident model is a 0.8B and invents facts. Letting her read a local
// article snippet beats letting her recall. Nothing here talks to the internet;
// the Kiwix server is on the same box.
//
// This is search only. Full articles are ~100KB of HTML, far too big for a 4096
// token context, so the unit of context is the search snippet (~500 chars).
package kiwix
import (
"context"
"encoding/xml"
"fmt"
"html"
"io"
"net/http"
"net/url"
"regexp"
"strconv"
"strings"
"time"
)
// Result is one search hit.
type Result struct {
Title string // article title, e.g. "Rayleigh scattering"
Path string // e.g. /content/wikipedia_en_all_maxi_2026-02/Rayleigh_scattering
Snippet string // plain text, tags stripped, entities decoded
WordCount int // 0 if the server did not say
}
// Client is a Kiwix HTTP client. Boring on purpose: no retries, no cache.
type Client struct {
base string
http *http.Client
}
// New makes a client for a Kiwix base URL like http://127.0.0.1:8034.
func New(baseURL string) *Client {
return &Client{
base: strings.TrimRight(baseURL, "/"),
http: &http.Client{Timeout: 10 * time.Second},
}
}
// Search runs a keyword search in one ZIM (book) and returns up to limit hits.
//
// Ranking is keyword based, not semantic: "Rayleigh scattering" finds the right
// article, "why is the sky blue" finds a TV episode. Pass keywords, not questions.
func (c *Client) Search(ctx context.Context, pattern, book string, limit int) ([]Result, error) {
if limit <= 0 {
limit = 5
}
q := url.Values{}
q.Set("pattern", pattern)
q.Set("books.name", book)
q.Set("format", "xml")
q.Set("pageLength", strconv.Itoa(limit))
req, err := http.NewRequestWithContext(ctx, http.MethodGet, c.base+"/search?"+q.Encode(), nil)
if err != nil {
return nil, err
}
resp, err := c.http.Do(req)
if err != nil {
return nil, err
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
return nil, fmt.Errorf("kiwix search: http %d", resp.StatusCode)
}
return ParseSearchRSS(resp.Body)
}
// rss mirrors just the bits of the RSS 2.0 reply we use.
type rss struct {
Items []struct {
Title string `xml:"title"`
Link string `xml:"link"`
// innerxml keeps the <b> match markers so we can strip them ourselves.
Description struct {
Inner string `xml:",innerxml"`
} `xml:"description"`
WordCount string `xml:"wordCount"`
} `xml:"channel>item"`
}
var tagRE = regexp.MustCompile(`<[^>]*>`)
// ParseSearchRSS turns a Kiwix search reply into results. Exported so the parser
// is testable from a captured response, with no server running.
func ParseSearchRSS(r io.Reader) ([]Result, error) {
var doc rss
if err := xml.NewDecoder(r).Decode(&doc); err != nil {
return nil, fmt.Errorf("kiwix search: bad xml: %w", err)
}
out := make([]Result, 0, len(doc.Items))
for _, it := range doc.Items {
n, _ := strconv.Atoi(strings.ReplaceAll(it.WordCount, ",", ""))
out = append(out, Result{
Title: strings.TrimSpace(it.Title),
Path: strings.TrimSpace(it.Link),
Snippet: plainText(it.Description.Inner),
WordCount: n,
})
}
return out, nil
}
// plainText drops markup and decodes entities, leaving text a model can read.
func plainText(s string) string {
s = tagRE.ReplaceAllString(s, "")
s = html.UnescapeString(s)
return strings.TrimSpace(strings.Join(strings.Fields(s), " "))
}
+90
View File
@@ -0,0 +1,90 @@
package kiwix
import (
"context"
"os"
"strings"
"testing"
"time"
)
// A real reply from the live server, trimmed to two items.
const sampleRSS = `<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:opensearch="http://a9.com/-/spec/opensearch/1.1/">
<channel>
<title>Search: Rayleigh scattering</title>
<opensearch:totalResults>800</opensearch:totalResults>
<item>
<title>Rayleigh scattering</title>
<link>/content/wikipedia_en_all_maxi_2026-02/Rayleigh_scattering</link>
<description><b>Rayleigh</b> scattering causes the blue color of the sky &amp; yellow colors near the Sun.[1]</description>
<book><title>Wikipedia</title></book>
<wordCount>2,818</wordCount>
</item>
<item>
<title>HyperRayleigh scattering</title>
<link>/content/wikipedia_en_all_maxi_2026-02/Hyper%E2%80%93Rayleigh_scattering</link>
<description>...<b>Rayleigh</b> scattering" is a nonlinear optical counterpart.</description>
<book><title>Wikipedia</title></book>
<wordCount>914</wordCount>
</item>
</channel>
</rss>`
func TestParseSearchRSS(t *testing.T) {
got, err := ParseSearchRSS(strings.NewReader(sampleRSS))
if err != nil {
t.Fatalf("parse: %v", err)
}
if len(got) != 2 {
t.Fatalf("want 2 results, got %d", len(got))
}
if got[0].Title != "Rayleigh scattering" {
t.Errorf("title = %q", got[0].Title)
}
if got[0].Path != "/content/wikipedia_en_all_maxi_2026-02/Rayleigh_scattering" {
t.Errorf("path = %q", got[0].Path)
}
if got[0].WordCount != 2818 {
t.Errorf("wordCount = %d", got[0].WordCount)
}
want := "Rayleigh scattering causes the blue color of the sky & yellow colors near the Sun.[1]"
if got[0].Snippet != want {
t.Errorf("snippet = %q, want %q", got[0].Snippet, want)
}
if strings.Contains(got[1].Snippet, "<b>") {
t.Errorf("second snippet still has tags: %q", got[1].Snippet)
}
}
func TestParseSearchRSSBadXML(t *testing.T) {
if _, err := ParseSearchRSS(strings.NewReader("not xml at all")); err == nil {
t.Fatal("want an error on junk input")
}
}
// Opt-in: needs a live Kiwix server. CI has none.
// MAVEN_KIWIX_URL=http://127.0.0.1:8034 no_proxy=127.0.0.1,localhost go test -run Retrieval -v ./internal/kiwix/
func TestRetrievalEval(t *testing.T) {
base := os.Getenv("MAVEN_KIWIX_URL")
if base == "" {
t.Skip("set MAVEN_KIWIX_URL to run the retrieval eval")
}
noProxyLoopback(t)
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute)
defer cancel()
rep, err := RunRetrievalEval(ctx, New(base), 5)
if err != nil {
t.Fatalf("eval: %v", err)
}
// No pass bar on purpose: the number is the finding.
t.Log("\n" + rep.String() + rep.Detail())
}
// noProxyLoopback stops the box's SOCKS bridge from eating loopback requests.
func noProxyLoopback(t *testing.T) {
t.Setenv("no_proxy", "127.0.0.1,localhost")
t.Setenv("NO_PROXY", "127.0.0.1,localhost")
}
+63
View File
@@ -0,0 +1,63 @@
{
"name": "kiwix-knowledge-v1",
"book": "wikipedia_en_all_maxi_2026-02",
"note": "The 9 knowledge cases from internal/phraser/eval/talk_v1.json. Queries are hand-written English keywords on purpose: Kiwix ranks by keyword, not meaning, so a natural question fails. Writing them by hand separates 'retrieval is broken' from 'the model writes bad queries'.",
"cases": [
{
"id": "know-sky-blue",
"question": "почему небо синее?",
"query": "Rayleigh scattering sky blue",
"want_titles": ["Rayleigh scattering", "Diffuse sky radiation"]
},
{
"id": "know-boil-egg",
"question": "сколько варить яйцо вкрутую?",
"query": "boiled egg cooking",
"want_titles": ["Boiled egg", "Egg as food"]
},
{
"id": "know-ssd-vs-hdd",
"question": "чем ssd отличается от hdd?",
"query": "solid-state drive",
"want_titles": ["Solid-state drive", "Hard disk drive"]
},
{
"id": "know-cat-purr",
"question": "почему кошки мурчат?",
"query": "cat purr",
"want_titles": ["Purr", "Cat communication"]
},
{
"id": "know-hiccups",
"question": "как быстро избавиться от икоты?",
"query": "hiccup",
"want_titles": ["Hiccup"]
},
{
"id": "know-polite-form",
"question": "не могли бы вы объяснить, что такое vpn?",
"query": "virtual private network",
"want_titles": ["Virtual private network"]
},
{
"id": "know-dont-know",
"question": "как зовут моего соседа снизу?",
"query": "name of my downstairs neighbour",
"want_titles": [],
"expect_miss": true,
"note": "Unanswerable by design. Retrieval SHOULD find nothing useful. Counted as a hit only when nothing relevant comes back."
},
{
"id": "know-water-per-day",
"question": "сколько воды в день надо пить?",
"query": "human daily water requirement drinking",
"want_titles": ["Drinking water", "Water", "Dehydration", "Hydration"]
},
{
"id": "know-thunder-delay",
"question": "почему гром слышно позже молнии?",
"query": "thunder speed of sound lightning",
"want_titles": ["Thunder", "Lightning"]
}
]
}
+135
View File
@@ -0,0 +1,135 @@
package kiwix
// This scores retrieval alone: no LLM. For each general-knowledge question we
// hand-write English keywords and ask whether the article that would answer it
// comes back in the top N hits. If this score is low, reading Wikipedia cannot
// help the model no matter how good the prompt is.
//
// The unanswerable case (know-dont-know) is not scored. Whether the junk it
// returns is "nothing useful" is a human judgement, so the report just prints
// the titles and leaves the score to the 8 answerable cases.
import (
"context"
_ "embed"
"encoding/json"
"fmt"
"strings"
)
//go:embed knowledge_v1.json
var knowledgeFixtureJSON []byte
// EvalCase — one question with hand-written keywords.
type EvalCase struct {
ID string `json:"id"`
Question string `json:"question"`
Query string `json:"query"`
WantTitles []string `json:"want_titles"`
ExpectMiss bool `json:"expect_miss"`
}
type fixture struct {
Name string `json:"name"`
Book string `json:"book"`
Cases []EvalCase `json:"cases"`
}
// Outcome — what one case retrieved.
type Outcome struct {
Case EvalCase
Titles []string // titles of the top N hits, in rank order
Rank int // 1-based rank of the first wanted title, 0 if none
Err error
}
// Hit is true when a wanted title came back.
func (o Outcome) Hit() bool { return o.Rank > 0 }
// Report — the score plus per-case detail.
type Report struct {
Name string
Book string
TopN int
Scored int // answerable cases
Hits int
Errors int
Outcomes []Outcome
}
// Accuracy over the answerable cases.
func (r Report) Accuracy() float64 {
if r.Scored == 0 {
return 0
}
return float64(r.Hits) / float64(r.Scored)
}
// RunRetrievalEval searches for every fixture case.
func RunRetrievalEval(ctx context.Context, c *Client, topN int) (Report, error) {
var f fixture
if err := json.Unmarshal(knowledgeFixtureJSON, &f); err != nil {
return Report{}, err
}
rep := Report{Name: f.Name, Book: f.Book, TopN: topN}
for _, cs := range f.Cases {
res, err := c.Search(ctx, cs.Query, f.Book, topN)
o := Outcome{Case: cs, Err: err}
if err != nil {
rep.Errors++
}
for i, hit := range res {
o.Titles = append(o.Titles, hit.Title)
if o.Rank == 0 && matches(cs.WantTitles, hit.Title) {
o.Rank = i + 1
}
}
if !cs.ExpectMiss {
rep.Scored++
if o.Hit() {
rep.Hits++
}
}
rep.Outcomes = append(rep.Outcomes, o)
}
return rep, nil
}
func matches(want []string, title string) bool {
for _, w := range want {
if strings.EqualFold(strings.TrimSpace(title), w) {
return true
}
}
return false
}
// String — the headline number.
func (r Report) String() string {
var b strings.Builder
fmt.Fprintf(&b, "%s: %d/%d answerable questions retrieve a wanted article in top %d (%.1f%%), %d errors\n",
r.Name, r.Hits, r.Scored, r.TopN, 100*r.Accuracy(), r.Errors)
fmt.Fprintf(&b, " book: %s\n", r.Book)
return b.String()
}
// Detail — per case: what was asked, what was searched, what came back.
func (r Report) Detail() string {
var b strings.Builder
for _, o := range r.Outcomes {
mark := "MISS"
switch {
case o.Case.ExpectMiss:
mark = "n/a "
case o.Hit():
mark = fmt.Sprintf("hit@%d", o.Rank)
}
fmt.Fprintf(&b, " %-6s %-20s q=%q\n", mark, o.Case.ID, o.Case.Query)
if o.Err != nil {
fmt.Fprintf(&b, " error: %v\n", o.Err)
continue
}
fmt.Fprintf(&b, " got: %s\n", strings.Join(o.Titles, " | "))
}
return b.String()
}
+183
View File
@@ -0,0 +1,183 @@
package kiwix
// Turning a Russian question into an English Kiwix search.
//
// Kiwix ranks by keyword, not by meaning. "why is the sky blue" returns a TV
// episode; "Rayleigh scattering sky blue" returns the right article. So the
// model's job here is NOT translation — it is naming the English article the
// answer lives in.
//
// The output space is a handful of words, so it is worth locking down hard: a
// GBNF grammar for the shape, a tiny token cap, and a cleanup pass that throws
// away anything odd rather than handing junk to Kiwix.
import (
"context"
"encoding/json"
"fmt"
"strings"
"unicode"
"github.com/kami/maven/internal/llm"
)
// Completer — the LLM seam, so tests can fake it. *llm.Client satisfies it.
type Completer interface {
Complete(ctx context.Context, r llm.Req) (string, error)
}
// queryGrammar — one JSON object holding 1..6 keyword words. Latin letters,
// digits and hyphens only, so the model physically cannot answer the question
// or reply in Russian.
//
// Why the JSON wrapper: this model always thinks out loud and this llama-server
// build ignores the thinking switch (see ROUTING-EVAL-31-07-2026.md). A bare
// word-list grammar just captured the reasoning — every case came back as
// "Let me analyze this request carefully". Demanding JSON, like routeGrammar and
// responseGrammar already do, gives the reasoning nowhere to go.
const queryGrammar = `
root ::= "{" ws "\"query\"" ws ":" ws "\"" word (" " word){0,5} "\"" ws "}"
word ::= [A-Za-z0-9] [A-Za-z0-9-]{0,23}
ws ::= [ \t\n]*
`
// rewriteSystem — asks for search keywords, not an answer and not a translation.
const rewriteSystem = `You turn a question into a search query for English Wikipedia.
Rules:
- Output ONLY English search keywords. Never an answer, never an explanation.
- Do NOT translate the sentence. Name the thing the answer is about.
- The output must be a noun phrase, like a Wikipedia article title.
- Never use question words: no why, how, what, when, which, "how much",
"how long", "how to", "vs", "reason", "difference".
- 2 to 4 words.
Reply with JSON: {"query":"<keywords>"}
Good:
"почему листья желтеют осенью?" -> {"query":"leaf senescence autumn"}
"как работает микроволновка?" -> {"query":"microwave oven"}
"не могли бы вы объяснить, что такое блокчейн?" -> {"query":"blockchain"}
"сколько живут собаки?" -> {"query":"dog lifespan"}
"как избавиться от комаров в квартире?" -> {"query":"mosquito control"}
"чем чай отличается от кофе?" -> {"query":"tea"}
Only JSON, no explanation.`
// maxQueryTokens — the output is a few words plus the JSON wrapper. A tight cap
// is the cheapest guard against the model rambling into an answer.
const maxQueryTokens = 32
// Rewriter asks the resident model for English search keywords.
type Rewriter struct{ c Completer }
func NewRewriter(c Completer) *Rewriter { return &Rewriter{c: c} }
// Rewrite returns English keywords for a question in any language.
// It errors rather than returning something Kiwix should not see.
func (r *Rewriter) Rewrite(ctx context.Context, question string) (string, error) {
raw, err := r.c.Complete(ctx, llm.Req{
System: rewriteSystem,
User: strings.TrimSpace(question),
Grammar: queryGrammar,
MaxTokens: maxQueryTokens,
RepeatPenalty: 1.15,
})
if err != nil {
return "", err
}
return CleanQuery(unwrapJSON(raw))
}
// unwrapJSON pulls the query out of {"query":"..."}. If the reply is not that
// shape it is returned as-is, and CleanQuery decides whether it is usable.
func unwrapJSON(raw string) string {
s := strings.TrimSpace(raw)
if !strings.HasPrefix(s, "{") {
return s
}
var got struct{ Query string }
if err := json.Unmarshal([]byte(s), &got); err != nil {
return s
}
return got.Query
}
// maxQueryWords matches the grammar's bound. Anything longer is prose.
const maxQueryWords = 6
// CleanQuery checks and tidies whatever the model produced. The grammar makes
// bad output unlikely, not impossible (a server without grammar support, a
// different model), so this is the real gate in front of Kiwix.
//
// Exported so it can be tested without a model.
func CleanQuery(raw string) (string, error) {
s := strings.TrimSpace(raw)
// Models like to wrap answers in quotes. Drop surrounding ones.
s = strings.Trim(s, "\"'`")
// Keep the first line only: everything after it is prose.
if i := strings.IndexAny(s, "\r\n"); i >= 0 {
s = s[:i]
}
// Keep letters, digits, spaces and hyphens; anything else becomes a space.
var b strings.Builder
for _, ru := range s {
switch {
case unicode.IsLetter(ru) || unicode.IsDigit(ru) || ru == '-':
b.WriteRune(ru)
default:
b.WriteRune(' ')
}
}
words := strings.Fields(b.String())
if len(words) == 0 {
return "", fmt.Errorf("kiwix rewrite: empty query")
}
if len(words) > maxQueryWords {
return "", fmt.Errorf("kiwix rewrite: %d words, want at most %d (looks like prose)", len(words), maxQueryWords)
}
words = dropStopWords(words)
out := strings.Join(words, " ")
// The ZIMs are English. Non-Latin letters mean the model ignored the ask.
for _, ru := range out {
if unicode.IsLetter(ru) && !isLatin(ru) {
return "", fmt.Errorf("kiwix rewrite: query is not English: %q", out)
}
}
return out, nil
}
// stopWords — question words and filler. The model keeps writing question-shaped
// queries ("why is the sky blue", "how much water to drink daily") no matter how
// the prompt is worded, and Kiwix ranks on every word, so those words drag in
// song and episode titles. Dropping them in code is not a style preference: a
// keyword ranker gets nothing from them.
var stopWords = map[string]bool{
"a": true, "an": true, "the": true, "is": true, "are": true, "was": true,
"do": true, "does": true, "did": true, "to": true, "of": true, "in": true,
"on": true, "for": true, "and": true, "or": true, "my": true, "me": true,
"i": true, "it": true, "its": true, "be": true, "been": true, "get": true,
"how": true, "why": true, "what": true, "when": true, "which": true,
"who": true, "where": true, "much": true, "many": true, "long": true,
"vs": true, "than": true, "rid": true, "from": true, "about": true,
}
// dropStopWords removes filler, but never everything: if the query was nothing
// but stop words there is nothing better to search, so the original is kept and
// the caller sees whatever Kiwix makes of it.
func dropStopWords(words []string) []string {
kept := make([]string, 0, len(words))
for _, w := range words {
if !stopWords[strings.ToLower(w)] {
kept = append(kept, w)
}
}
if len(kept) == 0 {
return words
}
return kept
}
func isLatin(ru rune) bool {
return (ru >= 'a' && ru <= 'z') || (ru >= 'A' && ru <= 'Z')
}
+96
View File
@@ -0,0 +1,96 @@
package kiwix
// End-to-end score: Russian question -> model rewrite -> Kiwix search -> did a
// wanted article come back. Same 9 cases as the retrieval eval, so the two
// numbers are directly comparable: retrieval with hand-written keywords is the
// ceiling, this is what the model actually reaches.
import (
"context"
"encoding/json"
"fmt"
"strings"
)
// RewriteOutcome — one case, end to end.
type RewriteOutcome struct {
Outcome
ModelQuery string // what the model asked for ("" if it failed)
RewriteErr error
}
// RunRewriteEval rewrites every question with the model, then searches.
func RunRewriteEval(ctx context.Context, c *Client, rw *Rewriter, topN int) (RewriteReport, error) {
var f fixture
if err := json.Unmarshal(knowledgeFixtureJSON, &f); err != nil {
return RewriteReport{}, err
}
rep := RewriteReport{Report: Report{Name: f.Name + "-rewrite", Book: f.Book, TopN: topN}}
for _, cs := range f.Cases {
out := RewriteOutcome{Outcome: Outcome{Case: cs}}
q, err := rw.Rewrite(ctx, cs.Question)
out.ModelQuery, out.RewriteErr = q, err
if err == nil {
res, serr := c.Search(ctx, q, f.Book, topN)
out.Err = serr
for i, hit := range res {
out.Titles = append(out.Titles, hit.Title)
if out.Rank == 0 && matches(cs.WantTitles, hit.Title) {
out.Rank = i + 1
}
}
}
if out.RewriteErr != nil || out.Err != nil {
rep.Errors++
}
if !cs.ExpectMiss {
rep.Scored++
if out.Hit() {
rep.Hits++
}
}
rep.Cases = append(rep.Cases, out)
}
return rep, nil
}
// RewriteReport — the score plus per-case detail.
type RewriteReport struct {
Report
Cases []RewriteOutcome
}
// String — the headline number.
func (r RewriteReport) String() string {
return fmt.Sprintf("%s: %d/%d answerable questions retrieve a wanted article in top %d (%.1f%%), %d errors\n book: %s\n",
r.Name, r.Hits, r.Scored, r.TopN, 100*r.Accuracy(), r.Errors, r.Book)
}
// Detail — per case: hand-written query next to the model's, and what came back.
// The point is seeing WHERE the model's phrasing differs, not just the score.
func (r RewriteReport) Detail() string {
var b strings.Builder
for _, o := range r.Cases {
mark := "MISS"
switch {
case o.Case.ExpectMiss:
mark = "n/a "
case o.Hit():
mark = fmt.Sprintf("hit@%d", o.Rank)
}
fmt.Fprintf(&b, " %-6s %-20s\n", mark, o.Case.ID)
fmt.Fprintf(&b, " asked: %s\n", o.Case.Question)
fmt.Fprintf(&b, " hand: %q\n", o.Case.Query)
fmt.Fprintf(&b, " model: %q\n", o.ModelQuery)
if o.RewriteErr != nil {
fmt.Fprintf(&b, " rewrite rejected: %v\n", o.RewriteErr)
continue
}
if o.Err != nil {
fmt.Fprintf(&b, " search error: %v\n", o.Err)
continue
}
fmt.Fprintf(&b, " got: %s\n", strings.Join(o.Titles, " | "))
}
return b.String()
}
+33
View File
@@ -0,0 +1,33 @@
package kiwix
import (
"context"
"os"
"testing"
"time"
"github.com/kami/maven/internal/llm"
)
// Opt-in: needs a live Kiwix server AND a live llama-server.
// MAVEN_KIWIX_URL=http://127.0.0.1:8034 MAVEN_LLM_URL=http://127.0.0.1:18099 \
//
// no_proxy=127.0.0.1,localhost go test -run RewriteEval -v ./internal/kiwix/
func TestRewriteEval(t *testing.T) {
kbase, lbase := os.Getenv("MAVEN_KIWIX_URL"), os.Getenv("MAVEN_LLM_URL")
if kbase == "" || lbase == "" {
t.Skip("set MAVEN_KIWIX_URL and MAVEN_LLM_URL to run the rewrite eval")
}
noProxyLoopback(t)
ctx, cancel := context.WithTimeout(context.Background(), 15*time.Minute)
defer cancel()
rw := NewRewriter(llm.New(lbase, 3*time.Minute))
rep, err := RunRewriteEval(ctx, New(kbase), rw, 5)
if err != nil {
t.Fatalf("eval: %v", err)
}
// No pass bar on purpose: the number is the finding.
t.Log("\n" + rep.String() + rep.Detail())
}
+105
View File
@@ -0,0 +1,105 @@
package kiwix
import (
"context"
"testing"
"github.com/kami/maven/internal/llm"
)
// Bad model output must never reach Kiwix. No model needed for this.
func TestCleanQueryRejectsJunk(t *testing.T) {
bad := []struct{ name, raw string }{
{"empty", ""},
{"blank", " \n "},
{"russian came back", "почему небо синее"},
{"mixed russian", "sky синее scattering"},
{"full sentence", "The sky looks blue because of the scattering of sunlight by air molecules"},
{"prose with quotes", `Sure! Here is a good search query: "Rayleigh scattering", which explains it.`},
}
for _, c := range bad {
if got, err := CleanQuery(c.raw); err == nil {
t.Errorf("%s: want rejection, got %q", c.name, got)
}
}
}
func TestCleanQueryCleans(t *testing.T) {
ok := []struct{ raw, want string }{
{"Rayleigh scattering sky", "Rayleigh scattering sky"},
{" boiled egg cooking \n", "boiled egg cooking"},
{`"virtual private network"`, "virtual private network"},
{"solid-state drive", "solid-state drive"},
{"cat purr.", "cat purr"},
{"hiccup\nAlso: hiccough", "hiccup"},
// Question words are filler to a keyword ranker, so they go.
{"why is the sky blue", "sky blue"},
{"how much water to drink daily", "water drink daily"},
{"SSD vs HDD comparison", "SSD HDD comparison"},
// Nothing but filler: keep it rather than return nothing.
{"what is it", "what is it"},
}
for _, c := range ok {
got, err := CleanQuery(c.raw)
if err != nil {
t.Errorf("%q: %v", c.raw, err)
continue
}
if got != c.want {
t.Errorf("%q -> %q, want %q", c.raw, got, c.want)
}
}
}
type fakeCompleter struct {
out string
req llm.Req
}
func (f *fakeCompleter) Complete(_ context.Context, r llm.Req) (string, error) {
f.req = r
return f.out, nil
}
func TestRewriteConstrainsTheCall(t *testing.T) {
f := &fakeCompleter{out: `{"query":"Rayleigh scattering sky"}`}
got, err := NewRewriter(f).Rewrite(context.Background(), "почему небо синее?")
if err != nil {
t.Fatalf("rewrite: %v", err)
}
if got != "Rayleigh scattering sky" {
t.Errorf("query = %q", got)
}
if f.req.Grammar == "" {
t.Error("no grammar sent")
}
if f.req.MaxTokens == 0 || f.req.MaxTokens > 32 {
t.Errorf("max_tokens = %d, want a small cap", f.req.MaxTokens)
}
}
func TestRewriteRejectsBadModelOutput(t *testing.T) {
bad := []string{
`{"query":"почему небо синее"}`, // never translated
`{"query":""}`, // empty
`{"query":"the sky is blue because sunlight is scattered by air"}`, // an answer
// Note: a SHORT English prose fragment ("Let me analyze this request")
// is under the word cap and cannot be caught here. The grammar is what
// stops that one.
}
for _, out := range bad {
f := &fakeCompleter{out: out}
if got, err := NewRewriter(f).Rewrite(context.Background(), "почему небо синее?"); err == nil {
t.Errorf("%s: want rejection, got %q", out, got)
}
}
}
// A reply that is not the JSON shape but is still usable keywords should pass.
func TestRewriteFallsBackToPlainText(t *testing.T) {
f := &fakeCompleter{out: "Rayleigh scattering sky"}
got, err := NewRewriter(f).Rewrite(context.Background(), "почему небо синее?")
if err != nil || got != "Rayleigh scattering sky" {
t.Errorf("got %q, %v", got, err)
}
}