4f59ba78c6bc7c5c658a272eeee7d0b1e5c44b19
Numbers behind the resident-model change, plus the answer to "could a 230-350M model do this instead" — no, and the reason is worth keeping: LFM2.5's published instruction-following scores beat Qwen3.5-0.8B, and every one of those benchmarks except Multi-IF is English. In Russian the 350M invents non-words and the 230M answers in Spanish. Also fills the row TALK-EVAL-31-07-2026.md had to void for contamination, and corrects a wrong call I nearly made: the 1.7B's 16s p95 looked like the reasoning trace, but the 0.8B sits at 17s in every run and the 1.7B beat it twice out of three. The long tail is shared and is not the Thinking block. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
Description
No description provided
Languages
Go
97.1%
HTML
0.9%
Shell
0.6%
CSS
0.5%
Makefile
0.3%
Other
0.6%