diff --git a/docs/evals/2026-08-05-temperature-sweep.md b/docs/evals/2026-08-05-temperature-sweep.md new file mode 100644 index 0000000..c266d73 --- /dev/null +++ b/docs/evals/2026-08-05-temperature-sweep.md @@ -0,0 +1,58 @@ +# Talk temperature sweep: Qwen3-1.7B, 4 temperatures × 3 runs + +Date: 05-08-2026. Model: Qwen3-1.7B-UD-Q4_K_XL, the resident model, on homesrv. +Harness: `TestTalkTemperatureSweep` (`internal/phraser/eval/temperature_test.go`), +gated on `MAVEN_LLM_URL` + `MAVEN_TEMP_SWEEP`. Fixture: the 36-case talk set. +Wall clock: 3394s for all twelve runs. Vikunja #402. + +## What was asked + +Whether 0.7 is the right sampling temperature for phrasing, and whether a lower +one buys persona compliance. + +## Numbers + +| temp | run 1 | run 2 | run 3 | mean | errors | +|---|---|---|---|---|---| +| 0.70 | 24/36 | 24/36 | 22/36 | 23.3 (64.8%) | 7, 4, 6 | +| 0.40 | 25/36 | 25/36 | 26/36 | 25.3 (70.4%) | 3, 4, 5 | +| 0.20 | 22/36 | 23/36 | 26/36 | 23.7 (65.7%) | 6, 5, 4 | +| 0.05 | 23/36 | 25/36 | 23/36 | 23.7 (65.7%) | 5, 5, 6 | + +## What it says + +**The sweep does not separate the temperatures.** 0.40 leads by 5.6 points on +the mean. The spread inside a single temperature is 11 points: 0.20 ranges 22 to +26 across three runs of the same setting. Three runs cannot tell a 5.6-point +effect from that noise. Lowering the temperature to 0.05 does not help either. +That is the result that would have been most useful if it had. + +**So the default stays 0.7.** `Config.Temperature` is now a config field, so +setting it is a one-line change. No measurement here justifies moving it. Anyone +re-running this needs more runs per setting, not more settings. + +## The finding that is not about temperature + +Sixty of the failures across twelve runs are one error: +`phraser: model output starts as JSON but does not parse`. The case fails with +an empty string, so it costs a whole case rather than one check. + +They are not spread evenly. Every one lands in the `reply` family, and the +distribution is: + +| case | runs failed (of 12) | +|---|---| +| reply-reminder-tomorrow | 12 | +| reply-reminder-evening | 12 | +| reply-question-bait | 11 | +| reply-formality-bait | 9 | +| reply-note-router | 8 | +| reply-fact-weight | 6 | + +Two cases fail in every single run at every temperature. That is not sampling +noise, and no temperature will fix it. It is a defect in the reply phrasing +path. It caps the talk fixture at 30/36 before persona is scored at all. Filed +as Vikunja #537. + +The talk score of 27/36 recorded on 2026-08-04 went through a different call +path. It is not comparable to the numbers above.