kami 4f59ba78c6 Write down the five-model sweep and why the 1.7B won
Numbers behind the resident-model change, plus the answer to "could a 230-350M
model do this instead" — no, and the reason is worth keeping: LFM2.5's published
instruction-following scores beat Qwen3.5-0.8B, and every one of those benchmarks
except Multi-IF is English. In Russian the 350M invents non-words and the 230M
answers in Spanish.

Also fills the row TALK-EVAL-31-07-2026.md had to void for contamination, and
corrects a wrong call I nearly made: the 1.7B's 16s p95 looked like the reasoning
trace, but the 0.8B sits at 17s in every run and the 1.7B beat it twice out of
three. The long tail is shared and is not the Thinking block.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
2026-07-31 19:08:37 +04:00
2026-07-31 18:58:20 +04:00
2026-07-31 02:32:44 +04:00
2026-07-31 18:58:20 +04:00
2026-07-03 00:32:48 +02:00
S
Description
No description provided
70 MiB
Languages
Go 97.1%
HTML 0.9%
Shell 0.6%
CSS 0.5%
Makefile 0.3%
Other 0.6%