246db4e609
The swap itself already landed: deploy loads models/embedder/multilingual-e5-small/model_quantized.onnx, and onnxembedder.go grew EmbedQuery/EmbedPassage with the query:/passage: prefixes the model was trained with. What was missing is the half of #371 that says "re-run make eval-recall and compare against the recorded numbers", so nothing in the repo says whether it worked. It worked, on every axis at once. recall@1 60.0% → 70.4%, recall@3 80.0% → 85.2%, answered after the gate 48.0% → 63.0%, false recall 1/5 → 0/5, and latency p50 59ms → 23ms because the quantized file is 118MB against the 470MB fp32 one the old config loaded. The guitar-chords note no longer beats the docker-logs note. One premise of the task did not come true and the new doc says so. #371 expected a better retriever to separate the score distributions and make query_min_score tunable. It did not: right-first top-1 runs 0.791-0.890 and must-stay-silent runs 0.795-0.835, still overlapping, just higher and tighter. The margin separates them instead — 0.024 median against 0.002 — and 0.008 is the knee where all five silent cases are silenced at no cost. The score gate is close to inert now; the margin is the live dial. Neither is changed here, since #412 is where a sweep belongs. docs/evals/2026-08-04-recall-e5-small.md is the dated measurement. rearchitecture.md's "upgrade MiniLM → bge-m3 later" is now done and says so, CLAUDE.md names the retriever and the prefix rule where it already promises the embedder never leaves homesrv, and the Makefile comment points at this eval instead of the one that asked for the swap. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>