docs/evals/
One file per measurement, named for the day it was taken. Never edited after
that day. A newer number is a new file, never an edit to an old one.
Reasoning does not belong here. It belongs in the subsystem's living doc under
docs/. An eval holds the setup, the numbers and what they rule out.
Rules for this directory
- Filename is
YYYY-MM-DD-<what-was-measured>.md. The date is the day it ran.
- The first line is the claim, not the topic. A reader picks a file from this
index without opening it, so the H1 has to carry the finding.
- Head the file with the date, the task id, the box and the build.
- A file that replaces an older number says so. It names the file it
replaces. This index then marks that one superseded.
- A superseded file is not deleted and not edited. It records what was believed
that day. A living doc may still cite it for the run itself.
Reading a number out of here
Check the state column before citing a row. A number in a living doc must
cite a live file. The 2026-08-11 classifier baseline exists for that reason.
A pair in docs/routing.md went stale unnoticed. Its source predated the
encodeWord fix at feabf9f. Nothing failed (V-704).
Routing: the arms and the cascade
| measurement |
state |
| Routing evaluation |
superseded |
| Resident model bake-off |
live |
| gemma-4-12b on the workstation, against the resident Qwen3-1.7B |
live |
| Routing from audio: four paths, one fixture |
live |
| Routing with the resident model, re-measured |
live |
| The routing trajectory, and the number that is missing |
live |
| Gemma as a label function, and what it found in the seeds |
live |
| Moving the seed files onto the router prompt's boundaries |
live |
| The first destination number |
superseded |
| The destination, with a model that can name one |
live |
| The routing heads, running in Go |
live |
| Two heads on e5-small, and the first destination the router did not need a model for |
live |
| A slot head, and the corpus that did not exist this morning |
live |
| A clarify head, and a confidence that is not a hardcode |
live |
| MASSIVE Russian warm-start for the routing heads |
live |
| gemma-4-E4B against gemma-4-12B on the routing fixture |
live |
| The classifier baseline after the tokenizer fix |
live |
| The CPT+SFT Qwen3-1.7B routes better and cannot hold a sentence |
live |
docs/routing.md holds the arm table these feed. Cite from there, not from here.
Stage 0, slots and the turn
Ecosystem
Recall and memory
Phrasing and talk
World: search and Kiwix
Speech in and out
Runtime and storage
Whole-system runs
Each run pairs a write-up with its raw transcript. The transcript is the
evidence and is not summarised anywhere else.
Audit
The 2026-08-26 baseline scores all 146 v1 DoD criteria in docs/spec.md and is
cited by every verified cell in docs/capabilities/ledger.yaml. It measures
maven-instruct-b2-Q4_K_XL, not the Qwen3-1.7B CLAUDE.md names, and it
expires when the model or deploy/mavend.json moves.
Its open findings live in docs/caveats/, one entry each with a revisit
trigger. Read the index there, not this file, for what is still broken.