fa67dd82fe
His call, written down so the rig can be prepared. The question was what it costs in GPU hours to train a small routing model. The answer is that the question has the wrong shape: routing emits one of 7 intents, one of 5 moods and a few spans, so it is classification, and a model that generates is being asked to do the wrong job. The model already exists on the box. multilingual-e5-small is 118M parameters, trained on Russian, quantized and resident. It gets three heads on one forward pass. Intent and mood read the mean-pooled vector, slots read last_hidden_state as BIO tags. That is about 12k parameters of head, which is why the serving side needs no second runtime: onnxembedder.go already pulls last_hidden_state at [1, 128, 384] into Go and pools it there, so the heads are three dot products over a weights file. Cost is 10 to 30 minutes on the workstation, under 2GB of VRAM, and it also finishes overnight on the homesrv CPU. A 100M decoder from scratch is 10 to 20 GPU hours plus a tokenizer plus a corpus, for a worse result. A LoRA on 0.6B is 1 to 4 hours and still generates, so it still needs the grammar and still has no real confidence. Two things this buys that no decoder can. Constrained output stops being a grammar problem, because a softmax cannot emit a value that does not exist. And max softmax is a calibratable confidence, where Confidence: 1.0 was a hardcode and V-359 had to rebuild the signal out of structure. The trap is in the plan twice because it is the one that silently costs something. Fine-tune a COPY. The resident embedder backs memory recall at ten points above MiniLM, and training it in place couples routing accuracy to recall@1 with nothing in the suite to name the trade. The real cost is the labeled set. 77 routing cases and 30 Praxis cases are a test set. The stage 0 grammars can self-label the turn history, which distils the rules into the model, but the fixtures stay out of training or the measurement reads the rules and reports them as the model.