Files
Maven/docs/capabilities/README.md
T
claude 8153e5eaa5 Classify every gap and rank the work (V-725)
Session 2 step 3, and the end of the explanation half.

gaps.md compares responsibilities and never package names. Eight classes. The
four capability classes are derived from the ledger's gap_class field and
rebuild with build_ledger.py. The four architecture classes are read from
findings.md and invariants.md, and every entry names the capability or invariant
it affects. An entry naming neither is marked non-blocking cleanup in those
words, which is the whole of class 8 and its eleven rows.

Of 46 v1 capabilities: 5 missing, 21 partial and reachable, 9 built and
unreachable, 11 reachable and unverified.

The nine unreachable ones are seven config blocks and two compose entries. Not
one is a code defect.

The ranked list puts phrasing first: speak-as-herself fails all three criteria,
and everything that asks the resident model to write a Russian sentence inherits
that. His own name not being stored is second. Nine capabilities one config
change from reachable is fourth, and it is the highest ratio of capability to
work in the list.

Items 10, 12 and 13 stall on unresolved invariants and are the owner's call, not
work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 12:44:15 +04:00

5.4 KiB

docs/capabilities/

Generated. Regenerated from docs/spec.md plus a named eval. Not hand-edited.

docs/spec.md says what Maven should do. This directory says how much of that exists, measured rather than asserted. The predecessor audit had a green test suite while 22 of 39 capabilities were not live, which is the failure mode the whole directory is built against.

The files

file what it is hand-edited
ledger.yaml the ledger: 51 capabilities, 156 DoD criteria, one verdict per criterion, seven implementation dimensions per capability no
build_ledger.py extracts the ledger from docs/spec.md and joins the two inputs yes, it is the source
domains.yaml the domain axis, one of the two judgment calls in the extraction yes
implementation.yaml capability to component mapping, the other judgment call yes
verdicts.json one verdict per criterion id, produced by scoring a probe run no, scored
probes_field.json 25 multi-turn probes: the owner's real week yes
probes_dod.json probes derived from the ledger's criteria no, generated
run_probes.py drives a probe file through the deployed stack yes
store_counts.py row counts per store, over IPC yes
invariants.md the twelve cross-cutting rules the 51 capabilities imply yes
gaps.md eight gap classes and the one ranked priority list yes, except classes 1-4
out/ raw probe output, one JSON object per line no

Rebuilding

python3 docs/capabilities/build_ledger.py      # spec.md + domains.yaml + implementation.yaml + verdicts.json + maven-architecture.json -> ledger.yaml

The generator is also the checker. It fails, loudly and non-zero, on a capability with no DoD criteria, a capability with no State line, a criterion id collision, a capability with no domain or more than two, an unknown domain name, a domains.yaml row naming a capability that does not exist, a verdicts.json row scoring a criterion that does not exist, a verdict word outside the five, and a reason outside the plan's list. It caught the domain reconciler silently dropping recall from its 51.

It also fails on a capability missing from implementation.yaml, a component id that docs/architecture/maven-architecture.json does not carry, and a component status the dimension table does not know. A capability absent from the mapping would read no on every dimension, which is indistinguishable from a capability nothing carries.

The seven dimensions

Never one implemented boolean. Coded and unwired, wired and unconfigured, and configured and undeployed are three different pieces of work.

designed, code_present, wired, configured, deployed and reachable come from the status field of every component mapped to the capability, rolled up as all yes, none no, otherwise partial. verified comes from the criteria verdicts: yes when every one passes, partial when some do.

designed answers a narrower question, because docs/spec.md states all 51 of them. It reads yes when a living doc owns the subsystem, and spec-only when the State line says no living doc or no package.

A component the mapping does not use is reported by name at the end of a build. Shared infrastructure is excluded on purpose: mapping core.reactive_handler to everything would give all 51 rows the same status and say nothing.

Running the probes

The probes run on homesrv, where mavweb is on 127.0.0.1:9201 and the mavend socket is reachable from inside the container.

# Build the probe binary. It needs CGO and the target's glibc, so build it in a
# trixie container: both the golang image and the mavend image are trixie.
docker run --rm -v "$PWD":/src -w /src \
  -e CGO_ENABLED=1 -e GOFLAGS=-mod=vendor \
  -e GOCACHE=/src/.cache/gocache -e GOPATH=/src/.cache/gopath \
  golang:1.25-trixie go build -buildvcs=false -o /src/.cache/e2eprobe ./cmd/e2eprobe
docker cp .cache/e2eprobe maven-mavend-1:/tmp/e2eprobe

python3 docs/capabilities/store_counts.py                          # before
python3 docs/capabilities/run_probes.py probes_field.json > out/field.raw.jsonl
python3 docs/capabilities/store_counts.py                          # after

Two things the harness learned the hard way

Probes must be isolated. mavweb hardcodes one conversation id for the whole web reach, so a clarify parked by one probe is still parked for the next. The first run measured the previous probe, not the current one: the park set at turn 8 appended Сейчас 01:07. В какой день? to turns 9 through 13, five unrelated turns in a row, including plain statements. run_probes.py now sends отмена before every probe. The contaminated run is kept at out/field.contaminated.jsonl, because the leak is a finding and not only an artifact.

Readback is the contract, not the file. The plaintext database copy at /dev/shm/maven-plain.db would answer every question faster and would bypass the IPC contract the ledger exists to measure. The mavweb pages are a second-hand rendering of the same thing.

What a verdict means

pass comes only from live evidence: the deployed build, the real model, real store rows. The scenario harness scripts both route and reply, so a green scenario proves the wiring around the model and not the turn; it is recorded as implementation evidence and reads untested. simulated is allowed only where the trigger is anchored to a wall-clock hour or date a probe cannot reach.