a103708a08f22916515b13358c42b7b964abb9e8
The budget was cut to 60s because a five-minute evaluation held the single llama-server slot, and a voice turn arriving mid-evaluation waited behind it. That collision is now solved where it belongs: the background client yields the slot while a turn is in flight. With the gate in place the short budget only truncates a Thinking model mid-synthesis, which costs an observation and saves no latency on any real turn. Kami's call, 2026-08-02. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BFaeSbLMEVG5ey8tejU3y2
Description
No description provided
Languages
Go
97.1%
HTML
0.9%
Shell
0.6%
CSS
0.5%
Makefile
0.3%
Other
0.6%