0ed386eca61b73491a5bd5b0dec47278a96c4a1c
The earlier before/after was taken while another eval shared llama-server. This run had the box to itself. Intent accuracy 61.8% llm-only, 63.2% cascade, 67.1% with thinking off. The prompt fix holds. note→fact shows up here too, so it is real. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
Description
No description provided
Languages
Go
97.1%
HTML
0.9%
Shell
0.6%
CSS
0.5%
Makefile
0.3%
Other
0.6%