Merge the e5-small routing plan (#175)

This commit is contained in:
2026-08-05 16:23:22 +04:00
2 changed files with 159 additions and 0 deletions
+11
View File
@@ -171,6 +171,17 @@ p50 329ms** — better than the resident model and about 2.5× faster (`docs/eva
Vikunja #485). The workstation is never assumed up, so both sets of numbers are live. Judge a
routing change against the classifier and the resident model, since those are what always answer.
**The intended third engine is not a generative model** (owner's call, 05-08-2026, V-546,
`docs/plans/18-routing-heads-on-e5-small.md`). Routing has a bounded output space, so it is
classification, and the 118M multilingual-e5-small is already resident. Three heads on one
forward pass: intent, mood, and BIO slot tags. Roughly 5e15 FLOPs to train, so 10 to 30
minutes on the workstation. A 100M decoder from scratch is 10 to 20 GPU hours. Two things
it buys that a decoder cannot. No grammar is needed, because a softmax cannot emit a value
that does not exist. And max softmax is a calibratable confidence, where `Confidence: 1.0`
was a hardcode. **Fine-tune a copy of the weights.** The resident embedder backs memory
recall. Training it in place couples routing accuracy to recall@1, with nothing in the
suite to name the trade.
`Confidence: 1.0` used to be hardcoded in `llmrouter.go`, so the LLM
path could never ask for clarification (6/6 refusal cases missed on the fixture) — Vikunja
#359. Fixed 31-07-2026 with structural signal (single-token utterance, keyless fact, act with
+148
View File
@@ -0,0 +1,148 @@
# Plan: route with heads on e5-small, not with a generative model
**Owner's call, 05-08-2026. Vikunja #546.**
**Verdict: the routing model is the 118M multilingual-e5-small already resident on
homesrv.** It gets one classification head per output. No LoRA on a decoder, no 100M model
trained from scratch. Routing has a bounded output space, so it is classification. A model
that generates is being asked to do the wrong job.
Last verified: 05-08-2026 @ 52fd218
## The two options this rules out
**A LoRA on Qwen3-0.6B or 1.7B.** About 1 to 4 GPU hours on the workstation, for 20k
examples over three epochs. It works, and it still generates. So the output still needs a
GBNF grammar in front of it. The confidence still has to be rebuilt from structure, the way
`gateLLMDecision` does today.
**A 100M decoder from scratch.** It needs roughly 2B tokens to be a usable language model.
That is 6 × 1e8 × 2e9, about 1.2e18 FLOPs. A 16GB card
does that in 10 to 20 GPU hours at its effective throughput. It also needs a Russian tokenizer built and a corpus assembled.
What it buys is a small model generating Russian, and `CLAUDE.md` already records that as
the thing that does not work. LFM2.5-350M routes at 5.2% and answered "столица Франции?"
with the invented non-word "Сторзит".
That finding is about generation, not about size. A 118M encoder classifying Russian is a
different job with a bounded output space. The measured recall of e5-small on this box is
the evidence it reads the language well enough.
## What the model becomes
The encoder body stays as it is. Three heads sit on top of one forward pass:
| Output | Head | Reads |
|---|---|---|
| intent | `Linear(384, 7)` | the mean-pooled vector |
| mood | `Linear(384, 5)` | the mean-pooled vector |
| slots | `Linear(384, 9)` per token | `last_hidden_state` |
Seven intents are the existing enum: fact, reminder, note, query, act, chat, system. Five
moods are the existing enum: neutral, happy, thinking, tired, confused. Nine slot tags are
BIO over `key`, `value`, `when` and `fn`, plus outside.
Total head size is about 12k parameters. That number decides how this is served. See the
serving section below.
## Two things this buys that the current router cannot
**Constrained output stops being a grammar problem.** There is no free generation, so the
heads can only emit values that exist. Four mechanisms exist today because a decoder can
write anything. The GBNF grammar, the JSON parse, the fallback to plain text, the legacy
`{"body","summary"}` path. A softmax cannot write anything.
**Confidence becomes a real number.** `Confidence: 1.0` was hardcoded in `llmrouter.go`,
so the router could never ask for clarification. V-359 had to rebuild a signal out of
structure: single-token utterance, keyless fact, act with no allowlisted fn. Max softmax
over the intent head is calibratable against the fixture. `r.threshold` and the stage-3
gate would read a probability instead of a proxy. Two false clarifies survived the V-359
fix, both in the act-with-no-allowlisted-fn arm. That is the arm a calibrated score
replaces.
## Cost
About 20k examples at 64 tokens over five epochs is 6.4M tokens. So 6 × 1.2e8 × 6.4e6,
roughly 5e15 FLOPs. **10 to 30 minutes on the workstation. Under 2GB of VRAM.** It also
finishes overnight on homesrv's CPU when the card is busy. That matters, because the
workstation is never assumed up.
Freeze the embedding table. The XLM-R vocabulary is about 96M of the 118M parameters, and
it is the part that overfits 20k examples. Train the twelve layers and the heads, at 2e-5
on the body and 1e-3 on the heads, batch 32, sequence 64. Loss is cross-entropy on intent
plus mood plus per-token tags, with the tag term down-weighted.
## The trap: fine-tune a copy
The resident embedder backs memory recall. `docs/evals/2026-08-04-recall-e5-small.md`
measured ten points of recall@1 above MiniLM, at 2.5× the speed. `CLAUDE.md` says it stays
on homesrv permanently, because it backs the floor.
Training it in place couples routing accuracy to recall. Say a run gains four points of
intent accuracy and quietly loses six of recall@1. It would look like a win, and nothing
in the test suite would name the trade. So the routing weights are a second file. About
45MB extra, quantized. `modelIDFromPath` already derives the DB marker from the filename.
So two files means two ids, and no ambiguity about which model wrote an embedding.
## Serving: the heads do not need a runtime
`internal/router/onnxembedder.go` already asks the ONNX session for `last_hidden_state` at
`[1, 128, 384]` and mean-pools in Go. Both tensors the heads need already cross into Go on
every call.
At 12k parameters the heads are three dot products. They can be plain Go over a weights
file rather than a second ONNX graph. Then the export step covers the encoder only, and
the head weights are data. That keeps the whole thing inside the existing session, the
existing vendored tokenizer and the existing `TestONNX*` tests.
Wire it where the router sits. `pickLLMRouter` becomes a three-way choice. The classifier
stays underneath as the floor. The rule that keeps it there now still holds: a turn must
never break on a model. The resident model keeps chat, world answers and phrasing. None of that
is classification, and this cannot do it.
## The labeled set is the whole project
There are 77 RU routing cases and 30 Praxis cases today. That is a test set, not a
training set.
The stage 0 grammars are high-precision label functions. `AgendaQueryGrammars`,
`NarrativeQueryGrammar` and `PraxisGrammars` each decide a shape deterministically, so
running them over the turn history self-labels it. Training on their output distils the
rules into the model. That is the point rather than a compromise: the model generalizes
past a regex, where Go's `\b` never fires after a Cyrillic letter.
Two rules for the data:
- **The fixtures stay out of training.** Otherwise the measurement reads the rules and
reports them as the model.
- **Keep cases the grammars do not cover.** A set labeled only by the rules teaches only
the rules. The hard cases carry no interrogative and no question mark, which is why
V-498 existed.
Expanding to 20k runs through gemma-4-12b on the workstation. `docs/evals/2026-08-02-workstation-gemma4-12b.md`
measured 329ms per call. So that is 2 to 4 hours of the card, plus a day of the owner
reading it. That is the real cost of this plan.
## How it gets judged
The 77-case RU routing fixture, against the two numbers that always answer:
| Path | Full | Intent-only | p50 |
|---|---|---|---|
| classifier | 68.8% | — | 16.6µs |
| resident Qwen3-1.7B, through the cascade | 72.7% | 77.9% | 0.80 to 1.04s |
| heads on e5-small | to measure | to measure | to measure |
Latency should land near the classifier rather than near the router. This is one encoder
pass and three dot products, against the classifier's one encoder pass and a
nearest-neighbour scan. If it does, V-464 answers itself, and the latency trade the LLM
router asks for stops being a trade.
## Not decided here
- Whether mood belongs on this model at all. Mood is a phrasing property, and the router
emitting it is a convenience. A fifth head is cheap, so this is a question about where
the value is read, not about cost.
- Whether the slot head replaces the stage 0 grammars or sits behind them. Cascade order
is a separate measurement, and the rules are currently faster and exact.
- The confidence calibration method. Temperature scaling on a held-out split is the
obvious first try, and it has not been measured.