Merge task/402-sweep-the-sampling-temperature (V-merge)
This commit is contained in:
@@ -0,0 +1,58 @@
|
||||
# Talk temperature sweep: Qwen3-1.7B, 4 temperatures × 3 runs
|
||||
|
||||
Date: 05-08-2026. Model: Qwen3-1.7B-UD-Q4_K_XL, the resident model, on homesrv.
|
||||
Harness: `TestTalkTemperatureSweep` (`internal/phraser/eval/temperature_test.go`),
|
||||
gated on `MAVEN_LLM_URL` + `MAVEN_TEMP_SWEEP`. Fixture: the 36-case talk set.
|
||||
Wall clock: 3394s for all twelve runs. Vikunja #402.
|
||||
|
||||
## What was asked
|
||||
|
||||
Whether 0.7 is the right sampling temperature for phrasing, and whether a lower
|
||||
one buys persona compliance.
|
||||
|
||||
## Numbers
|
||||
|
||||
| temp | run 1 | run 2 | run 3 | mean | errors |
|
||||
|---|---|---|---|---|---|
|
||||
| 0.70 | 24/36 | 24/36 | 22/36 | 23.3 (64.8%) | 7, 4, 6 |
|
||||
| 0.40 | 25/36 | 25/36 | 26/36 | 25.3 (70.4%) | 3, 4, 5 |
|
||||
| 0.20 | 22/36 | 23/36 | 26/36 | 23.7 (65.7%) | 6, 5, 4 |
|
||||
| 0.05 | 23/36 | 25/36 | 23/36 | 23.7 (65.7%) | 5, 5, 6 |
|
||||
|
||||
## What it says
|
||||
|
||||
**The sweep does not separate the temperatures.** 0.40 leads by 5.6 points on
|
||||
the mean. The spread inside a single temperature is 11 points: 0.20 ranges 22 to
|
||||
26 across three runs of the same setting. Three runs cannot tell a 5.6-point
|
||||
effect from that noise. Lowering the temperature to 0.05 does not help either.
|
||||
That is the result that would have been most useful if it had.
|
||||
|
||||
**So the default stays 0.7.** `Config.Temperature` is now a config field, so
|
||||
setting it is a one-line change. No measurement here justifies moving it. Anyone
|
||||
re-running this needs more runs per setting, not more settings.
|
||||
|
||||
## The finding that is not about temperature
|
||||
|
||||
Sixty of the failures across twelve runs are one error:
|
||||
`phraser: model output starts as JSON but does not parse`. The case fails with
|
||||
an empty string, so it costs a whole case rather than one check.
|
||||
|
||||
They are not spread evenly. Every one lands in the `reply` family, and the
|
||||
distribution is:
|
||||
|
||||
| case | runs failed (of 12) |
|
||||
|---|---|
|
||||
| reply-reminder-tomorrow | 12 |
|
||||
| reply-reminder-evening | 12 |
|
||||
| reply-question-bait | 11 |
|
||||
| reply-formality-bait | 9 |
|
||||
| reply-note-router | 8 |
|
||||
| reply-fact-weight | 6 |
|
||||
|
||||
Two cases fail in every single run at every temperature. That is not sampling
|
||||
noise, and no temperature will fix it. It is a defect in the reply phrasing
|
||||
path. It caps the talk fixture at 30/36 before persona is scored at all. Filed
|
||||
as Vikunja #537.
|
||||
|
||||
The talk score of 27/36 recorded on 2026-08-04 went through a different call
|
||||
path. It is not comparable to the numbers above.
|
||||
Reference in New Issue
Block a user