plan: the third routing engine is heads on e5-small, not a small decoder (V-546)
His call, written down so the rig can be prepared. The question was what it costs in GPU hours to train a small routing model. The answer is that the question has the wrong shape: routing emits one of 7 intents, one of 5 moods and a few spans, so it is classification, and a model that generates is being asked to do the wrong job. The model already exists on the box. multilingual-e5-small is 118M parameters, trained on Russian, quantized and resident. It gets three heads on one forward pass. Intent and mood read the mean-pooled vector, slots read last_hidden_state as BIO tags. That is about 12k parameters of head, which is why the serving side needs no second runtime: onnxembedder.go already pulls last_hidden_state at [1, 128, 384] into Go and pools it there, so the heads are three dot products over a weights file. Cost is 10 to 30 minutes on the workstation, under 2GB of VRAM, and it also finishes overnight on the homesrv CPU. A 100M decoder from scratch is 10 to 20 GPU hours plus a tokenizer plus a corpus, for a worse result. A LoRA on 0.6B is 1 to 4 hours and still generates, so it still needs the grammar and still has no real confidence. Two things this buys that no decoder can. Constrained output stops being a grammar problem, because a softmax cannot emit a value that does not exist. And max softmax is a calibratable confidence, where Confidence: 1.0 was a hardcode and V-359 had to rebuild the signal out of structure. The trap is in the plan twice because it is the one that silently costs something. Fine-tune a COPY. The resident embedder backs memory recall at ten points above MiniLM, and training it in place couples routing accuracy to recall@1 with nothing in the suite to name the trade. The real cost is the labeled set. 77 routing cases and 30 Praxis cases are a test set. The stage 0 grammars can self-label the turn history, which distils the rules into the model, but the fixtures stay out of training or the measurement reads the rules and reports them as the model.
This commit is contained in:
@@ -171,6 +171,17 @@ p50 329ms** — better than the resident model and about 2.5× faster (`docs/eva
|
||||
Vikunja #485). The workstation is never assumed up, so both sets of numbers are live. Judge a
|
||||
routing change against the classifier and the resident model, since those are what always answer.
|
||||
|
||||
**The intended third engine is not a generative model** (owner's call, 05-08-2026, V-546,
|
||||
`docs/plans/18-routing-heads-on-e5-small.md`). Routing has a bounded output space, so it is
|
||||
classification, and the 118M multilingual-e5-small is already resident. Three heads on one
|
||||
forward pass: intent, mood, and BIO slot tags. Roughly 5e15 FLOPs to train, so 10 to 30
|
||||
minutes on the workstation. A 100M decoder from scratch is 10 to 20 GPU hours. Two things
|
||||
it buys that a decoder cannot. No grammar is needed, because a softmax cannot emit a value
|
||||
that does not exist. And max softmax is a calibratable confidence, where `Confidence: 1.0`
|
||||
was a hardcode. **Fine-tune a copy of the weights.** The resident embedder backs memory
|
||||
recall. Training it in place couples routing accuracy to recall@1, with nothing in the
|
||||
suite to name the trade.
|
||||
|
||||
`Confidence: 1.0` used to be hardcoded in `llmrouter.go`, so the LLM
|
||||
path could never ask for clarification (6/6 refusal cases missed on the fixture) — Vikunja
|
||||
#359. Fixed 31-07-2026 with structural signal (single-token utterance, keyless fact, act with
|
||||
|
||||
Reference in New Issue
Block a user