llm: give voice turns priority on the single llama-server slot

llama-server is started without -np, so it serves one request at a time and
everything else queues. Mail extraction is allowed two minutes on a Thinking
1.7B, and the reader hands core up to 25 messages back to back. A turn arriving
mid-extraction therefore waited for whatever was left of that budget: the router
timed out into the classifier cascade and its 36.8% floor, and the phraser, which
has no floor, simply waited. Memory evaluation had the same shape with a five
minute budget.

llm.Gate is the bound. Foreground requests never wait. Background requests run
one at a time and yield while a foreground request is in flight, plus a quiet
window after it that covers the gap between the router call and the phraser call
of one turn. Clients get their priority from llmClientFor or
llmBackgroundClientFor, so which side a caller is on is decided at wiring time.
It gates only what goes through those clients, which the comment on Gate says.

mail intake: the extraction timeout no longer wraps the capture writes. A model
answering at 119 seconds of a 120 second budget left the first CaptureTask one
second and the third none, so candidates the model had already produced were
dropped with a deadline error. The mailbox name is validated before it becomes
provenance, since "email:" is not a source and neither is an arbitrary string
posted at the socket. The enable log prints the normalised candidate bound
rather than the configured one, which said "max 0" and then wrote three.
Found in review of #64.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TrVSBKe3RFDF4fGYKWYQnX
This commit is contained in:
kami
2026-08-01 14:05:07 +04:00
parent 6c81df17ec
commit aee20a6abc
9 changed files with 443 additions and 11 deletions
+59
View File
@@ -180,3 +180,62 @@ func TestNewMailIntakeOffWithoutConfig(t *testing.T) {
t.Error("without a llama-server phraser there is nothing to extract with")
}
}
// The mailbox name becomes the provenance string, which is the vocabulary the
// loop's rules trust. "email:" is not a source and neither is "email:anything
// he could post at the socket".
func TestIngestRejectsBadMailbox(t *testing.T) {
for _, name := range []string{"", " ", "IN BOX", "IN\nBOX", "IN\x00BOX", strings.Repeat("щ", maxMailboxChars+1)} {
mi, st, fake := newTestIntake(t, `[{"text":"дело","due":""}]`)
req := ingestReq()
req.Mailbox = name
if _, err := mi.ingest(context.Background(), req); err == nil {
t.Errorf("mailbox %q was accepted", name)
}
if fake.calls != 0 {
t.Errorf("mailbox %q reached the model", name)
}
if tasks, _ := st.ListTasks(context.Background(), ""); len(tasks) != 0 {
t.Errorf("mailbox %q wrote %d tasks", name, len(tasks))
}
}
}
// slowLLM burns most of the extraction budget before answering, the way a
// Thinking 1.7B does on a long mail.
type slowLLM struct {
reply string
delay time.Duration
}
func (s *slowLLM) Complete(ctx context.Context, _ llm.Req) (string, error) {
select {
case <-time.After(s.delay):
return s.reply, nil
case <-ctx.Done():
return "", ctx.Err()
}
}
// The extraction budget must not also bound the writes. It used to be one
// context, so a model answering near the deadline lost the candidates it had
// just produced.
func TestIngestCapturesAfterASlowExtraction(t *testing.T) {
st := newTestStore(t)
mi := &mailIntake{
st: st,
ex: email.NewExtractor(&slowLLM{reply: `[{"text":"оплатить счёт","due":""}]`, delay: 90 * time.Millisecond}, 0, nil),
timeout: 100 * time.Millisecond,
now: func() time.Time { return time.Date(2026, 8, 1, 10, 0, 0, 0, time.UTC) },
}
resp, err := mi.ingest(context.Background(), ingestReq())
if err != nil {
t.Fatalf("ingest: %v", err)
}
if resp.Created != 1 {
t.Fatalf("resp = %+v, want the candidate captured", resp)
}
if tasks, _ := st.ListTasks(context.Background(), ""); len(tasks) != 1 {
t.Errorf("got %d tasks, want 1", len(tasks))
}
}