Make Qwen3-1.7B the resident model #46

Closed
claude wants to merge 3 commits from overnight/resident-1.7b into overnight/nudge-templates
Contributor

One line of config, and it is the biggest single score change tonight.

Stock Qwen3-1.7B — not the CPT'd one, that training is still running.
Measured on an otherwise idle box against both fixtures we have.

Routing, 77 Russian cases:

model on disk intent-only cascade
LFM2.5-230M 246 MB 33.8% 36.4%
LFM2.5-350M 379 MB 5.2% 20.8%
Qwen3.5-0.8B (deployed) 527 MB 59.7% 61.0%
Qwen3.5-2B 1.34 GB 62.3% 63.6%
Qwen3-1.7B stock 1.13 GB 67.5% 72.7%

Talk fixture, 27 cases, three runs each:

0.8B 1.7B
composite 13, 11, 8 20, 21, 18
address 21, 18, 18 26, 25, 23
ontopic 16, 19, 19 22, 23, 23
canned fallbacks 8, 5, 6 0, 2, 0

address is the point. It sat at 18-22 of 27 on the 0.8B no matter how the prompt
was worded — the prompt forbids "вы" and the model writes подождите, делаете
anyway. I read that as "prompting is out of levers". It was really "0.8B is out of
capacity". And 5-8 of 27 turns on the 0.8B ended in a hardcoded "не знаю.", so it
was failing to emit parseable JSON about a quarter of the time. The 1.7B: 0-2.

n_ctx goes 2048 -> 4096 in the same commit — Thinking variant, reasoning needs the
room, and 4096 is the context every score above was measured at.

The routing number needs the LLM router wired on to appear. Still nil, so this
buys the phrasing improvement today and the routing improvement when that lands.

Also documents a dead end so nobody retries it: sub-500M is not close.
LFM2.5-350M routes at 5.2%, worse than guessing among 7 intents, and answers
"столица Франции?" with "Сторзит", which is not a Russian word. The 230M answers
Russian in Spanish. Their published IFEval and BFCL scores are strong and are
entirely English — every benchmark in that table except Multi-IF.

Two commits: the swap, then the numbers. The second also fills the row
TALK-EVAL-31-07-2026.md had to void for contamination (0.8B at 600ch/1024tok =
13, 11, 8) and corrects a latency call I nearly got wrong — the 1.7B's 16s p95
looked like the reasoning trace, but the 0.8B sits at 17s every run.

One line of config, and it is the biggest single score change tonight. Stock Qwen3-1.7B — **not** the CPT'd one, that training is still running. Measured on an otherwise idle box against both fixtures we have. **Routing, 77 Russian cases:** | model | on disk | intent-only | cascade | |---|---|---|---| | LFM2.5-230M | 246 MB | 33.8% | 36.4% | | LFM2.5-350M | 379 MB | 5.2% | 20.8% | | Qwen3.5-0.8B (deployed) | 527 MB | 59.7% | 61.0% | | Qwen3.5-2B | 1.34 GB | 62.3% | 63.6% | | **Qwen3-1.7B stock** | 1.13 GB | **67.5%** | **72.7%** | **Talk fixture, 27 cases, three runs each:** | | 0.8B | 1.7B | |---|---|---| | composite | 13, 11, 8 | **20, 21, 18** | | address | 21, 18, 18 | **26, 25, 23** | | ontopic | 16, 19, 19 | **22, 23, 23** | | canned fallbacks | 8, 5, 6 | **0, 2, 0** | `address` is the point. It sat at 18-22 of 27 on the 0.8B no matter how the prompt was worded — the prompt forbids "вы" and the model writes `подождите`, `делаете` anyway. I read that as "prompting is out of levers". It was really "0.8B is out of capacity". And 5-8 of 27 turns on the 0.8B ended in a hardcoded `"не знаю."`, so it was failing to emit parseable JSON about a quarter of the time. The 1.7B: 0-2. `n_ctx` goes 2048 -> 4096 in the same commit — Thinking variant, reasoning needs the room, and 4096 is the context every score above was measured at. **The routing number needs the LLM router wired on to appear.** Still `nil`, so this buys the phrasing improvement today and the routing improvement when that lands. Also documents a dead end so nobody retries it: **sub-500M is not close.** LFM2.5-350M routes at 5.2%, worse than guessing among 7 intents, and answers "столица Франции?" with "Сторзит", which is not a Russian word. The 230M answers Russian in Spanish. Their published IFEval and BFCL scores are strong and are entirely English — every benchmark in that table except Multi-IF. Two commits: the swap, then the numbers. The second also fills the row `TALK-EVAL-31-07-2026.md` had to void for contamination (0.8B at 600ch/1024tok = 13, 11, 8) and corrects a latency call I nearly got wrong — the 1.7B's 16s p95 looked like the reasoning trace, but the 0.8B sits at 17s every run.
claude added 1 commit 2026-07-31 16:58:45 +02:00
Stock Qwen3-1.7B, not the CPT'd one — that training is still running. It won
on both fixtures we have, measured tonight on an otherwise idle box:

  routing, 77 RU cases, intent-only:  67.5%  vs  59.7%  for Qwen3.5-0.8B
  talk fixture, 27 cases:             20/27  vs  11-17/27

It also beat Qwen3.5-2B, which is 20% larger, on every routing column.

Two other things came with it:

n_ctx goes 2048 -> 4096. This is a Thinking variant, so reasoning tokens need
the room, and 4096 is the context every score above was measured at. Shipping
2048 would ship something nobody measured.

The doc now says not to bother with sub-500M models, because I checked and they
are not close. LFM2.5-350M routes at 5.2% — worse than guessing among 7 intents
— and answers "столица Франции?" with "Сторзит", which is not a word. The 230M
replies to Russian in Spanish. Their published IFEval and BFCL numbers are good
and they are all English.

Note the routing gain needs the LLM router actually wired on to show up. It is
still nil, so this commit buys the phrasing improvement today and the routing
improvement when that lands.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
kami added 1 commit 2026-07-31 17:08:59 +02:00
Numbers behind the resident-model change, plus the answer to "could a 230-350M
model do this instead" — no, and the reason is worth keeping: LFM2.5's published
instruction-following scores beat Qwen3.5-0.8B, and every one of those benchmarks
except Multi-IF is English. In Russian the 350M invents non-words and the 230M
answers in Spanish.

Also fills the row TALK-EVAL-31-07-2026.md had to void for contamination, and
corrects a wrong call I nearly made: the 1.7B's 16s p95 looked like the reasoning
trace, but the 0.8B sits at 17s in every run and the 1.7B beat it twice out of
three. The long tail is shared and is not the Thinking block.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
kami added 1 commit 2026-07-31 19:17:16 +02:00
The file ran two sweeps and the second one changed the resident model, but
the lede still opened with "Recommendation: keep Qwen3.5-0.8B". Anyone
landing on the file read the wrong conclusion and had to scroll 100 lines
to find that it had been replaced — and it contradicted CLAUDE.md, which
already says the resident model is Qwen3-1.7B.

Both sweeps are accurate, so nothing is rewritten. The lede now states the
outcome and the first sweep's verdict is scoped to what it actually tested:
it rejects LFM2.5-1.2B, which still holds. It never was a case for keeping
0.8B as the resident model.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
Owner

Superseded by #47, which landed this whole stack on master as one reviewed integration merge. This PR head is an ancestor of master — its commits are in, nothing here is lost. Closing as merged-by-proxy rather than merged, since the merge came in through #47.

Review threads on this PR were answered or acted on before the merge; the Russian wording fixes went in as #48.

Superseded by #47, which landed this whole stack on master as one reviewed integration merge. This PR head is an ancestor of master — its commits are in, nothing here is lost. Closing as merged-by-proxy rather than merged, since the merge came in through #47. Review threads on this PR were answered or acted on before the merge; the Russian wording fixes went in as #48.
kami closed this pull request 2026-07-31 20:22:48 +02:00

Pull request closed

Sign in to join this conversation.
No Reviewers
No Label
2 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: kami/Maven#46