The 67.1% "thinking off" column in ROUTING-EVAL-31-07-2026.md was an
artefact. It came from a hand-rolled HTTP client in the eval test that
did not send repeat_penalty, so it differed from the reference run on two
axes and the penalty was the one that mattered.
Re-scored back to back on an idle box with everything else held equal:
thinking off is identical to thinking on, case for case, same confusion
matrix, same three unparseable replies. A direct probe of the running
llama-server shows enable_thinking, thinking and reasoning_budget are all
ignored for this model on this build, so there was nothing to turn off.
No defaults changed. The misleading third configuration is removed from
internal/router/eval so its table cannot be quoted again.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
The earlier before/after was taken while another eval shared
llama-server. This run had the box to itself.
Intent accuracy 61.8% llm-only, 63.2% cascade, 67.1% with thinking off.
The prompt fix holds. note→fact shows up here too, so it is real.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
Settles the measurement half of Vikunja #319: the resident 0.8B routes
better than the deployed classifier (50.0% vs 36.8% intent-only) at ~27x
the latency, and neither path can refuse an ambiguous utterance.
Records the four-configuration comparison, the per-case evidence behind
each finding (query->fact x15 traced to routeSystem's rule order, the
five missed-clarify cosines, the two text-field repetition truncations),
the two hypotheses that were tested and closed (thinking mode, grammar
array runaway), and a next-steps list mapped to #319/#320/#359.
#320 should not flip as-is: it would remove the refusal lane rather than
improve it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X5JApcrCRVGmqrxnhynSik