934a28d67c202e5e9768836722a9501cb1eddcd1
Two reporting tests over the 91-case RU fixture, no ratchet: a ratchet here would freeze a number nobody has decided to hold. TestStage0Contention runs the 21 grammars one at a time instead of stopping at the first match. One case of 91 draws two, ru-query-019, where calendar-query beats agenda-query by list position alone. TestONNXClaimConfidenceDistribution buckets the reported confidence by the layer that produced it. Stage 0 is 20/20 at a hardcoded 1.0. The classifier scores 62% below its median and 62% above, across a cosine range of 0.859 to 0.942, with a top-two margin of p50 0.009. The float is not a confidence. newBaselineClassifier and baselineGrammars split out of newBaselineRouter so the measurement runs the same rules the daemon runs. TestONNXBaseline is unchanged at 64/91.
Description
No description provided
Languages
Go
97.1%
HTML
0.9%
Shell
0.6%
CSS
0.5%
Makefile
0.3%
Other
0.6%