From fa67dd82fe60b01c6361a0e250dfe8710b169220 Mon Sep 17 00:00:00 2001 From: claude Date: Wed, 5 Aug 2026 16:23:13 +0400 Subject: [PATCH] plan: the third routing engine is heads on e5-small, not a small decoder (V-546) His call, written down so the rig can be prepared. The question was what it costs in GPU hours to train a small routing model. The answer is that the question has the wrong shape: routing emits one of 7 intents, one of 5 moods and a few spans, so it is classification, and a model that generates is being asked to do the wrong job. The model already exists on the box. multilingual-e5-small is 118M parameters, trained on Russian, quantized and resident. It gets three heads on one forward pass. Intent and mood read the mean-pooled vector, slots read last_hidden_state as BIO tags. That is about 12k parameters of head, which is why the serving side needs no second runtime: onnxembedder.go already pulls last_hidden_state at [1, 128, 384] into Go and pools it there, so the heads are three dot products over a weights file. Cost is 10 to 30 minutes on the workstation, under 2GB of VRAM, and it also finishes overnight on the homesrv CPU. A 100M decoder from scratch is 10 to 20 GPU hours plus a tokenizer plus a corpus, for a worse result. A LoRA on 0.6B is 1 to 4 hours and still generates, so it still needs the grammar and still has no real confidence. Two things this buys that no decoder can. Constrained output stops being a grammar problem, because a softmax cannot emit a value that does not exist. And max softmax is a calibratable confidence, where Confidence: 1.0 was a hardcode and V-359 had to rebuild the signal out of structure. The trap is in the plan twice because it is the one that silently costs something. Fine-tune a COPY. The resident embedder backs memory recall at ten points above MiniLM, and training it in place couples routing accuracy to recall@1 with nothing in the suite to name the trade. The real cost is the labeled set. 77 routing cases and 30 Praxis cases are a test set. The stage 0 grammars can self-label the turn history, which distils the rules into the model, but the fixtures stay out of training or the measurement reads the rules and reports them as the model. --- CLAUDE.md | 11 ++ docs/plans/18-routing-heads-on-e5-small.md | 148 +++++++++++++++++++++ 2 files changed, 159 insertions(+) create mode 100644 docs/plans/18-routing-heads-on-e5-small.md diff --git a/CLAUDE.md b/CLAUDE.md index 5bec859..93d86db 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -171,6 +171,17 @@ p50 329ms** — better than the resident model and about 2.5× faster (`docs/eva Vikunja #485). The workstation is never assumed up, so both sets of numbers are live. Judge a routing change against the classifier and the resident model, since those are what always answer. +**The intended third engine is not a generative model** (owner's call, 05-08-2026, V-546, +`docs/plans/18-routing-heads-on-e5-small.md`). Routing has a bounded output space, so it is +classification, and the 118M multilingual-e5-small is already resident. Three heads on one +forward pass: intent, mood, and BIO slot tags. Roughly 5e15 FLOPs to train, so 10 to 30 +minutes on the workstation. A 100M decoder from scratch is 10 to 20 GPU hours. Two things +it buys that a decoder cannot. No grammar is needed, because a softmax cannot emit a value +that does not exist. And max softmax is a calibratable confidence, where `Confidence: 1.0` +was a hardcode. **Fine-tune a copy of the weights.** The resident embedder backs memory +recall. Training it in place couples routing accuracy to recall@1, with nothing in the +suite to name the trade. + `Confidence: 1.0` used to be hardcoded in `llmrouter.go`, so the LLM path could never ask for clarification (6/6 refusal cases missed on the fixture) — Vikunja #359. Fixed 31-07-2026 with structural signal (single-token utterance, keyless fact, act with diff --git a/docs/plans/18-routing-heads-on-e5-small.md b/docs/plans/18-routing-heads-on-e5-small.md new file mode 100644 index 0000000..9c3b774 --- /dev/null +++ b/docs/plans/18-routing-heads-on-e5-small.md @@ -0,0 +1,148 @@ +# Plan: route with heads on e5-small, not with a generative model + +**Owner's call, 05-08-2026. Vikunja #546.** + +**Verdict: the routing model is the 118M multilingual-e5-small already resident on +homesrv.** It gets one classification head per output. No LoRA on a decoder, no 100M model +trained from scratch. Routing has a bounded output space, so it is classification. A model +that generates is being asked to do the wrong job. + +Last verified: 05-08-2026 @ 52fd218 + +## The two options this rules out + +**A LoRA on Qwen3-0.6B or 1.7B.** About 1 to 4 GPU hours on the workstation, for 20k +examples over three epochs. It works, and it still generates. So the output still needs a +GBNF grammar in front of it. The confidence still has to be rebuilt from structure, the way +`gateLLMDecision` does today. + +**A 100M decoder from scratch.** It needs roughly 2B tokens to be a usable language model. +That is 6 × 1e8 × 2e9, about 1.2e18 FLOPs. A 16GB card +does that in 10 to 20 GPU hours at its effective throughput. It also needs a Russian tokenizer built and a corpus assembled. +What it buys is a small model generating Russian, and `CLAUDE.md` already records that as +the thing that does not work. LFM2.5-350M routes at 5.2% and answered "столица Франции?" +with the invented non-word "Сторзит". + +That finding is about generation, not about size. A 118M encoder classifying Russian is a +different job with a bounded output space. The measured recall of e5-small on this box is +the evidence it reads the language well enough. + +## What the model becomes + +The encoder body stays as it is. Three heads sit on top of one forward pass: + +| Output | Head | Reads | +|---|---|---| +| intent | `Linear(384, 7)` | the mean-pooled vector | +| mood | `Linear(384, 5)` | the mean-pooled vector | +| slots | `Linear(384, 9)` per token | `last_hidden_state` | + +Seven intents are the existing enum: fact, reminder, note, query, act, chat, system. Five +moods are the existing enum: neutral, happy, thinking, tired, confused. Nine slot tags are +BIO over `key`, `value`, `when` and `fn`, plus outside. + +Total head size is about 12k parameters. That number decides how this is served. See the +serving section below. + +## Two things this buys that the current router cannot + +**Constrained output stops being a grammar problem.** There is no free generation, so the +heads can only emit values that exist. Four mechanisms exist today because a decoder can +write anything. The GBNF grammar, the JSON parse, the fallback to plain text, the legacy +`{"body","summary"}` path. A softmax cannot write anything. + +**Confidence becomes a real number.** `Confidence: 1.0` was hardcoded in `llmrouter.go`, +so the router could never ask for clarification. V-359 had to rebuild a signal out of +structure: single-token utterance, keyless fact, act with no allowlisted fn. Max softmax +over the intent head is calibratable against the fixture. `r.threshold` and the stage-3 +gate would read a probability instead of a proxy. Two false clarifies survived the V-359 +fix, both in the act-with-no-allowlisted-fn arm. That is the arm a calibrated score +replaces. + +## Cost + +About 20k examples at 64 tokens over five epochs is 6.4M tokens. So 6 × 1.2e8 × 6.4e6, +roughly 5e15 FLOPs. **10 to 30 minutes on the workstation. Under 2GB of VRAM.** It also +finishes overnight on homesrv's CPU when the card is busy. That matters, because the +workstation is never assumed up. + +Freeze the embedding table. The XLM-R vocabulary is about 96M of the 118M parameters, and +it is the part that overfits 20k examples. Train the twelve layers and the heads, at 2e-5 +on the body and 1e-3 on the heads, batch 32, sequence 64. Loss is cross-entropy on intent +plus mood plus per-token tags, with the tag term down-weighted. + +## The trap: fine-tune a copy + +The resident embedder backs memory recall. `docs/evals/2026-08-04-recall-e5-small.md` +measured ten points of recall@1 above MiniLM, at 2.5× the speed. `CLAUDE.md` says it stays +on homesrv permanently, because it backs the floor. + +Training it in place couples routing accuracy to recall. Say a run gains four points of +intent accuracy and quietly loses six of recall@1. It would look like a win, and nothing +in the test suite would name the trade. So the routing weights are a second file. About +45MB extra, quantized. `modelIDFromPath` already derives the DB marker from the filename. +So two files means two ids, and no ambiguity about which model wrote an embedding. + +## Serving: the heads do not need a runtime + +`internal/router/onnxembedder.go` already asks the ONNX session for `last_hidden_state` at +`[1, 128, 384]` and mean-pools in Go. Both tensors the heads need already cross into Go on +every call. + +At 12k parameters the heads are three dot products. They can be plain Go over a weights +file rather than a second ONNX graph. Then the export step covers the encoder only, and +the head weights are data. That keeps the whole thing inside the existing session, the +existing vendored tokenizer and the existing `TestONNX*` tests. + +Wire it where the router sits. `pickLLMRouter` becomes a three-way choice. The classifier +stays underneath as the floor. The rule that keeps it there now still holds: a turn must +never break on a model. The resident model keeps chat, world answers and phrasing. None of that +is classification, and this cannot do it. + +## The labeled set is the whole project + +There are 77 RU routing cases and 30 Praxis cases today. That is a test set, not a +training set. + +The stage 0 grammars are high-precision label functions. `AgendaQueryGrammars`, +`NarrativeQueryGrammar` and `PraxisGrammars` each decide a shape deterministically, so +running them over the turn history self-labels it. Training on their output distils the +rules into the model. That is the point rather than a compromise: the model generalizes +past a regex, where Go's `\b` never fires after a Cyrillic letter. + +Two rules for the data: + +- **The fixtures stay out of training.** Otherwise the measurement reads the rules and + reports them as the model. +- **Keep cases the grammars do not cover.** A set labeled only by the rules teaches only + the rules. The hard cases carry no interrogative and no question mark, which is why + V-498 existed. + +Expanding to 20k runs through gemma-4-12b on the workstation. `docs/evals/2026-08-02-workstation-gemma4-12b.md` +measured 329ms per call. So that is 2 to 4 hours of the card, plus a day of the owner +reading it. That is the real cost of this plan. + +## How it gets judged + +The 77-case RU routing fixture, against the two numbers that always answer: + +| Path | Full | Intent-only | p50 | +|---|---|---|---| +| classifier | 68.8% | — | 16.6µs | +| resident Qwen3-1.7B, through the cascade | 72.7% | 77.9% | 0.80 to 1.04s | +| heads on e5-small | to measure | to measure | to measure | + +Latency should land near the classifier rather than near the router. This is one encoder +pass and three dot products, against the classifier's one encoder pass and a +nearest-neighbour scan. If it does, V-464 answers itself, and the latency trade the LLM +router asks for stops being a trade. + +## Not decided here + +- Whether mood belongs on this model at all. Mood is a phrasing property, and the router + emitting it is a convenience. A fifth head is cheap, so this is a question about where + the value is read, not about cost. +- Whether the slot head replaces the stage 0 grammars or sits behind them. Cascade order + is a separate measurement, and the rules are currently faster and exact. +- The confidence calibration method. Temperature scaling on a held-out split is the + obvious first try, and it has not been measured.