# Talk temperature sweep: Qwen3-1.7B, 4 temperatures × 3 runs Date: 05-08-2026. Model: Qwen3-1.7B-UD-Q4_K_XL, the resident model, on homesrv. Harness: `TestTalkTemperatureSweep` (`internal/phraser/eval/temperature_test.go`), gated on `MAVEN_LLM_URL` + `MAVEN_TEMP_SWEEP`. Fixture: the 36-case talk set. Wall clock: 3394s for all twelve runs. Vikunja #402. ## What was asked Whether 0.7 is the right sampling temperature for phrasing, and whether a lower one buys persona compliance. ## Numbers | temp | run 1 | run 2 | run 3 | mean | errors | |---|---|---|---|---|---| | 0.70 | 24/36 | 24/36 | 22/36 | 23.3 (64.8%) | 7, 4, 6 | | 0.40 | 25/36 | 25/36 | 26/36 | 25.3 (70.4%) | 3, 4, 5 | | 0.20 | 22/36 | 23/36 | 26/36 | 23.7 (65.7%) | 6, 5, 4 | | 0.05 | 23/36 | 25/36 | 23/36 | 23.7 (65.7%) | 5, 5, 6 | ## What it says **The sweep does not separate the temperatures.** 0.40 leads by 5.6 points on the mean. The spread inside a single temperature is 11 points: 0.20 ranges 22 to 26 across three runs of the same setting. Three runs cannot tell a 5.6-point effect from that noise. Lowering the temperature to 0.05 does not help either. That is the result that would have been most useful if it had. **So the default stays 0.7.** `Config.Temperature` is now a config field, so setting it is a one-line change. No measurement here justifies moving it. Anyone re-running this needs more runs per setting, not more settings. ## The finding that is not about temperature Sixty of the failures across twelve runs are one error: `phraser: model output starts as JSON but does not parse`. The case fails with an empty string, so it costs a whole case rather than one check. They are not spread evenly. Every one lands in the `reply` family, and the distribution is: | case | runs failed (of 12) | |---|---| | reply-reminder-tomorrow | 12 | | reply-reminder-evening | 12 | | reply-question-bait | 11 | | reply-formality-bait | 9 | | reply-note-router | 8 | | reply-fact-weight | 6 | Two cases fail in every single run at every temperature. That is not sampling noise, and no temperature will fix it. It is a defect in the reply phrasing path. It caps the talk fixture at 30/36 before persona is scored at all. Filed as Vikunja #537. The talk score of 27/36 recorded on 2026-08-04 went through a different call path. It is not comparable to the numbers above.