From bc19f9e9fd225e832d61271956bff0ebbef929b7 Mon Sep 17 00:00:00 2001 From: kami Date: Thu, 13 Aug 2026 23:32:05 +0400 Subject: [PATCH] Measure the registry against the run it describes ARCHITECTURE.md step 0. eval/chapter-truth.json holds the 38 occurrences the user walked crop by crop on the 19:44 run of 2026-08-12, and audit_registry.py now prints purity per cluster and fragmentation per person against it. All six baseline numbers reproduce. Rows key on page-space geometry, purity is a share, and fragmentation is a count of ids, so nothing in the file names a panel_id or a character_id. The fifth cycle re-crops and calls /characters/reset, and the file survives both. That was the ordering trap in the handoff. NEXT.md said 2 of woman B's 9 crops were really woman A and never said which. They are panel_order 31 and 33, identified from p030 and p032. Four fixes to the audit itself, all pre-existing: - 20 characters counted where 14 are live and 6 are merge losers kept on purpose - the assignment spread keyed on name, so the two Seonhos summed into one line - the default worked example was panel_index 7, a panel vision skips. NEXT.md's "panel 7" is panel_order 7, one lower - nothing about skipped panels. 41 of 116 are skip=True, four checked and all four correct, and they hold 28 of the chapter's 122 dialogue lines That last count is the measured case for an offscreen speaker_ref kind: 23% of dialogue sits on panels with no character to attribute to. Checks: audit_registry.py --selftest covers the IoU match, the greedy tie-break and the purity maths with no database. ruff check . exits 0. Co-Authored-By: Claude Opus 5 --- AGENTS.md | 22 ++ ARCHITECTURE.md | 17 +- JOURNAL.md | 45 +++ NEXT.md | 26 +- audit_registry.py | 177 ++++++++- decisions/CLAUDE.md | 3 + decisions/measurement-spine.md | 56 +++ eval/chapter-truth.json | 669 +++++++++++++++++++++++++++++++++ 8 files changed, 1001 insertions(+), 14 deletions(-) create mode 100644 decisions/measurement-spine.md create mode 100644 eval/chapter-truth.json diff --git a/AGENTS.md b/AGENTS.md index beff1b4..4b2e6f8 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -21,6 +21,28 @@ The orchestrator's equivalent is `docker compose up -d --build orchestrator` on bakes its source. That recreates the container, which drops any `docker cp`'d file such as `audit_registry.py`. +## Measuring the registry + +`audit_registry.py` reads `/data/manga.db`, so it runs inside the container and needs two files copied +in. Both are dropped by a rebuild. From this repo, on workpc: + +```bash +tar cf - audit_registry.py eval/chapter-truth.json | /usr/bin/ssh kami@192.168.1.104 \ + "cat > /tmp/audit.tar && docker cp /tmp/audit.tar manga-orchestrator:/tmp/ \ + && docker exec manga-orchestrator tar xf /tmp/audit.tar -C /app \ + && docker exec manga-orchestrator python3 /app/audit_registry.py" + +.venv/bin/python audit_registry.py --selftest # the IoU match and purity maths, no db needed +``` + +It prints purity and fragmentation against `eval/chapter-truth.json`, which is the 19:44 run of +2026-08-12 walked crop by crop (`decisions/measurement-spine.md`). Every baseline line says +`= baseline` or the delta, so a regression is visible without reading the numbers. + +Traps. The audit's second argument is `panel_index`, and `NEXT.md`'s "panel 7" is `panel_order` 7, one +lower. A truth row matches on page-space geometry at IoU 0.5. A row printed as `unmatched` means the +detection moved more than half its area, not that the character was lost. + ## Video analysis When diagnosing render motion, transitions, timing, or visual artifacts, use diff --git a/ARCHITECTURE.md b/ARCHITECTURE.md index 67def2e..799c785 100644 --- a/ARCHITECTURE.md +++ b/ARCHITECTURE.md @@ -30,7 +30,9 @@ the DoD is the accuracy number rather than the mechanism. ## The measurement spine -Nothing below can be called done without this, and it does not exist yet. +Nothing below can be called done without this. **The character half is built, 2026-08-13**, and +reproduces the table below on the 19:44 run. The dialogue half is still empty +(`decisions/measurement-spine.md#truth-scope-is-38`). The only ground truth in the project is the eyeball pass over the 19:44 run of 2026-08-12. It lives in prose in `NEXT.md`. Write it to `eval/chapter-truth.json` against chapter @@ -56,6 +58,17 @@ it to print purity and fragmentation against the truth file. Do not write an eva **Done when:** `audit_registry.py ` prints purity per labelled character and fragmentation per labelled person, and reproduces the table above on the 19:44 run. +**Met for the character half.** All 38 labelled occurrences match an assignment at IoU 0.5, and all six +numbers print `= baseline`. The definitions had to be pinned down first. Purity is a share and +fragmentation is a count of ids, so neither needs a `character_id` from the truth file +(`decisions/measurement-spine.md#purity-is-id-free`). Run it with: + +```bash +docker cp audit_registry.py manga-orchestrator:/app/ +docker cp eval/chapter-truth.json manga-orchestrator:/app/eval/ +docker exec manga-orchestrator python3 /app/audit_registry.py +``` + ## 1. The page is a region graph, not a list of panels ``` @@ -328,7 +341,7 @@ Detection, vision and character embeddings already exist. The order below is cho falsifiable by the step's own DoD before the next one starts. ``` -0. eval/chapter-truth.json + purity and fragmentation in audit_registry.py +0. eval/chapter-truth.json + purity and fragmentation in audit_registry.py DONE, characters only 1. plane + species per detection -> section 1 DoD 2. face-plus-person embedding -> section 4 embedding check 3. chapter-wide clustering, weak cannot-link on plane -> section 4 purity DoD diff --git a/JOURNAL.md b/JOURNAL.md index f4c2a40..7114e24 100644 --- a/JOURNAL.md +++ b/JOURNAL.md @@ -756,3 +756,48 @@ graph edge with a typed union, narrative plane for art-in-art, and a persistent section carries what exists against it today. Nothing in it is built. Nothing ran on a GPU after the dialogue stage. + +## 2026-08-13 — the measurement spine, built and checked against the run it describes + +`ARCHITECTURE.md` step 0. No GPU, no stage run. The 19:44 run of 2026-08-12 was still in +`/data/manga.db`, which is what made this possible before the fifth cycle destroys it. + +`eval/chapter-truth.json` holds 38 labelled occurrences over the three characters the user walked crop by +crop. `audit_registry.py` now matches them to the run's assignments and prints purity per cluster and +fragmentation per person: + +``` +truth: 38 labelled occurrences, 38 matched an assignment at IoU >= 0.5, 0 unmatched + LIM SEONHO [b1dd5659] lead 14/16 = 0.88 = baseline (wrong: photo_of_lead 1, chibi_of_lead 1) + character_2b1b12a1 woman_a 13/13 = 1.00 = baseline + Seonho [f0d4e901] woman_b 7/9 = 0.78 = baseline (wrong: woman_a 2) + fragmentation lead 1, woman_a 2, woman_b 1 all = baseline +``` + +`NEXT.md` recorded that 2 of woman B's 9 crops were really woman A, and never said which. They are +`panel_order` 31 and 33, identified by pulling `p030` and `p032` from `s3://panels` and looking. Woman A +wears a cream top with black trim and a pearl necklace, and both crops are unmistakably her at zoom. +That closes the 7-and-2 arithmetic and is why her fragmentation is 2. + +The design is in `decisions/measurement-spine.md`. Rows key on page-space geometry rather than +`panel_id`, purity is a share and fragmentation a count of ids rather than anything naming a +`character_id`. So the file survives both the re-crop and `/characters/reset` that the fifth cycle runs. +`audit_registry.py --selftest` covers the IoU match, the greedy tie-break and the purity maths without a +database. + +Four things the audit itself was getting wrong, all pre-existing: + +- it counted 20 characters where 14 are live and 6 are merge losers kept on purpose +- it keyed the assignment spread on `name`, so the two characters both named `Seonho` summed to one line +- its default worked example was `panel_index` 7, a balloon-only panel vision skips. `NEXT.md`'s + "panel 7" is `panel_order` 7, one lower, the office shot with 6 detections +- it said nothing about skipped panels. 41 of 116 are `skip=True`, so 119 detections come from 75 + panels. Four skips were checked by eye and all four are correct + +The skip count carries one new fact worth more than the audit fix. Those 41 panels hold 28 of the +chapter's 122 dialogue lines, and a skipped panel has no character at all. 23% of the dialogue has no +visible speaker to attribute to under any linker. That is the case for `offscreen` as a fourth +`speaker_ref` kind, now measured rather than assumed. + +Step 0's dialogue half is not written. It needs 30 lines watched panel by panel, and `NEXT.md` holds no +per-line speaker truth to transcribe. diff --git a/NEXT.md b/NEXT.md index 64ba814..b87d458 100644 --- a/NEXT.md +++ b/NEXT.md @@ -37,6 +37,21 @@ weakness is that its biggest character has no name. Coverage fell because gemma now clears 12 crops instead of naming them from a shortlist that did not contain the right person. Coverage is still the `has_face` gate plus those refusals, and nothing else. +Four numbers the audit was misreporting, found while building the truth file on 2026-08-13: + +- **41 of 116 panels are `skip=True`**, so all 119 detections come from 75 panels. Four were checked by + eye. All four are right: a balloon-only panel, the chapter title card, a traffic light, an SFX word. + The skip rate is not a defect. It does mean the coverage number describes two thirds of the chapter. +- **Those skipped panels hold 28 of the chapter's 122 dialogue lines**, and a skipped panel has no + character at all. So 23% of dialogue can never resolve to a visible speaker, whatever the linker + does. This is the evidence for `offscreen` as a fourth `speaker_ref` kind, and it is now measured. +- **"panel 7" in this file means `panel_order` 7, which is `panel_index` 6**, the wide office shot with + 6 detections. `panel_index` 7 is the balloon-only panel vision skips. `audit_registry.py` defaulted + to 7 and printed an empty worked example. +- **The registry holds 20 rows, of which 14 are live** and 6 are merge losers kept on purpose. The + audit was printing 20 as the cast size. It also keyed the spread on `name`, so the two characters + both named `Seonho` were summed into one line of 11. + ## Next 1. **Re-crop the chapter and run the fifth cycle.** Four changes are written and tested since the fourth @@ -75,6 +90,14 @@ contain the right person. Coverage is still the `has_face` gate plus those refus carries a DoD. Its build order starts at step 0, `eval/chapter-truth.json` plus purity and fragmentation printed by `audit_registry.py`. Nothing below it is measurable without that file. + **Step 0's character half is done, 2026-08-13.** `eval/chapter-truth.json` holds 38 labelled + occurrences keyed on page-space geometry, and all six baseline numbers print `= baseline` + (`decisions/measurement-spine.md`). The ordering trap is closed. The file survives both the re-crop + and the reset, so the fifth cycle can be measured against it. Run the audit before and after. + + The dialogue half is not written and needs the chapter watched panel by panel, 30 lines with their + true speaker. It is the only part of step 0 still open. + Steps 1 and 2 are independent and cheap: `plane` and `species` per detection, and the face-plus-person embedding already queued in item 1 above. The weak cannot-link is step 3 and depends on step 2, not on `plane` alone. Co-presence as a hard constraint is load-bearing while cosine cannot @@ -173,7 +196,8 @@ contain the right person. Coverage is still the `has_face` gate plus those refus (`caveats/speaker-attribution.md#tail-is-not-geometry`) is now behind item 5. With no name to attach, geometry buys nothing. 11. Resolve a speaker answer across the whole dialogue window, not just the answering panel. The last 3 - unresolved refs describe a neighbouring panel in the same 8-panel call. + unresolved refs describe a neighbouring panel in the same 8-panel call. The window is also the only + way to reach the 28 lines that sit on skipped panels, measured above. 12. Start Phase 2 from `ROADMAP.md`. Set SQLite `busy_timeout` before any concurrency work (`caveats/audit-open.md#sqlite-locking`). diff --git a/audit_registry.py b/audit_registry.py index c3e1948..2d793a1 100644 --- a/audit_registry.py +++ b/audit_registry.py @@ -6,25 +6,121 @@ worth running at all: 1. did the bbox fix land — are stored boxes pixels, or still gemma's 0-1000 grid, 2. how many characters did the rebaseline mint, and did one of them absorb the chapter again, - 3. what happened on panel 7, the worked example. + 3. purity and fragmentation against `eval/chapter-truth.json`, the measurement spine, + 4. what happened on panel 7, the worked example. -Usage: docker exec manga-orchestrator python3 /app/audit_registry.py [chapter_id] [panel_index] +Usage: docker exec manga-orchestrator python3 /app/audit_registry.py [chapter_id] [panel_index] [truth] + python3 audit_registry.py --selftest # scoring only, no db + +Both files have to be inside the container, and `docker compose up --build` drops them: + docker cp audit_registry.py manga-orchestrator:/app/ + docker cp eval/chapter-truth.json manga-orchestrator:/app/eval/ """ import collections import json +import os import sqlite3 import sys +IOU_MIN = 0.5 # ponytail: a fixed floor. Boxes move a little between runs, people do not. + + +def _iou(a, b): + ix = max(0, min(a[2], b[2]) - max(a[0], b[0])) + iy = max(0, min(a[3], b[3]) - max(a[1], b[1])) + inter = ix * iy + if not inter: + return 0.0 + ua = (a[2] - a[0]) * (a[3] - a[1]) + (b[2] - b[0]) * (b[3] - b[1]) - inter + return inter / ua if ua > 0 else 0.0 + + +def _score(truth, found): + """Match truth rows to this run's assignments by page-space IoU, then cluster them. + + truth: [{page, box, person}], found: [{page, box, character_id}]. Greedy on best IoU, each row + and each assignment used once. Returns (per_cluster, fragmentation, unmatched_truth). + + Purity is the largest share of one true person inside a cluster, so it needs no character_id + from the truth file. It has to stay that way: /characters/reset remints every id. + """ + pairs = sorted( + ((_iou(t["box"], f["box"]), ti, fi) + for ti, t in enumerate(truth) for fi, f in enumerate(found) + if t["page"] == f["page"] and _iou(t["box"], f["box"]) >= IOU_MIN), + key=lambda p: -p[0]) + used_t, used_f, per_cluster, holders = set(), set(), {}, {} + for _, ti, fi in pairs: + if ti in used_t or fi in used_f: + continue + used_t.add(ti) + used_f.add(fi) + person, cid = truth[ti]["person"], found[fi]["character_id"] + per_cluster.setdefault(cid, collections.Counter())[person] += 1 + holders.setdefault(person, set()).add(cid) + frag = {p: len(ids) for p, ids in holders.items()} + return per_cluster, frag, [t for i, t in enumerate(truth) if i not in used_t] + + +def _selftest(): + truth = [ + {"page": 0, "box": [0, 0, 100, 100], "person": "a"}, + {"page": 0, "box": [200, 0, 300, 100], "person": "a"}, + {"page": 0, "box": [400, 0, 500, 100], "person": "b"}, + {"page": 0, "box": [600, 0, 700, 100], "person": "art"}, + {"page": 1, "box": [0, 0, 100, 100], "person": "a"}, + ] + found = [ + {"page": 0, "box": [4, 4, 104, 104], "character_id": "c1"}, # shifted by a re-crop + {"page": 0, "box": [200, 0, 300, 100], "character_id": "c2"}, # a again, on a second id + {"page": 0, "box": [400, 0, 500, 100], "character_id": "c1"}, # b, folded into c1 + {"page": 0, "box": [600, 0, 700, 100], "character_id": "c1"}, # the art, also c1 + {"page": 1, "box": [0, 0, 100, 100], "character_id": "c9"}, # same box, other page + ] + per_cluster, frag, unmatched = _score(truth, found) + assert dict(per_cluster["c1"]) == {"a": 1, "b": 1, "art": 1}, per_cluster + assert max(per_cluster["c1"].values()) / sum(per_cluster["c1"].values()) == 1 / 3 + assert frag == {"a": 3, "b": 1, "art": 1}, frag # a is split over c1, c2 and c9 + assert not unmatched, unmatched + # two candidates for one row: the tighter box wins and the looser one is left out + per_cluster, _, unmatched = _score( + truth[:1], found[:1] + [{"page": 0, "box": [0, 0, 100, 100], "character_id": "c8"}]) + assert list(per_cluster) == ["c8"], per_cluster + assert not unmatched, unmatched + # a row nothing overlaps stays unmatched rather than snapping to the nearest box + _, _, unmatched = _score(truth, [{"page": 0, "box": [0, 0, 20, 20], "character_id": "c1"}]) + assert len(unmatched) == 5, unmatched + print("selftest ok") + + +if "--selftest" in sys.argv: + _selftest() + raise SystemExit + CHAPTER = sys.argv[1] if len(sys.argv) > 1 else "7c944dd4-e972-42c7-ba60-9f6939548e80" -WORKED_EXAMPLE = int(sys.argv[2]) if len(sys.argv) > 2 else 7 +# panel_index, the db column. NEXT.md's "panel 7" is panel_order 7, which is this chapter's +# panel_index 6, the wide office shot. panel_index 7 is a balloon-only panel vision skips. +WORKED_EXAMPLE = int(sys.argv[2]) if len(sys.argv) > 2 else 6 +TRUTH = sys.argv[3] if len(sys.argv) > 3 else "/app/eval/chapter-truth.json" +if not os.path.exists(TRUTH): + TRUTH = os.path.join(os.path.dirname(os.path.abspath(__file__)), "eval/chapter-truth.json") c = sqlite3.connect("/data/manga.db") c.row_factory = sqlite3.Row manga_id = c.execute("SELECT manga_id FROM chapters WHERE chapter_id=?", (CHAPTER,)).fetchone()[0] reg = {r["character_id"]: dict(r) for r in c.execute( - "SELECT character_id, name, aliases, gender, ref_image_uris, embedding_uri " + "SELECT character_id, name, aliases, gender, ref_image_uris, embedding_uri, merged_into " "FROM characters WHERE manga_id=?", (manga_id,))} +# a merge keeps the losing row and sets merged_into, so a raw COUNT overstates the cast. +live = {k: r for k, r in reg.items() if not r["merged_into"]} + + +def _label(cid): + """Two characters can carry the same name — the naming defect mints duplicates — so a label + that is only the name silently sums two clusters into one line.""" + name = (reg.get(cid, {}).get("name") or "").strip() + return f"{name} [{cid[10:18]}]" if name else cid[:20] panels = c.execute( 'SELECT panel_id, panel_index, page_index, bbox FROM panels WHERE chapter_id=? ORDER BY panel_order', @@ -39,6 +135,8 @@ max_ratio = 0.0 assigned_total = 0 per_char = collections.Counter() worked = None +found = [] # every assignment as a page-space box, for the truth match below +skipped = lines = lines_skipped = 0 for p in panels: row = c.execute("SELECT result_json FROM vision_results WHERE panel_id=?", (p["panel_id"],)).fetchone() @@ -49,27 +147,42 @@ for p in panels: # [x, y, w, h], not corners — panel 3 of this chapter is [0, 615, 900, 106]. pb = json.loads(p["bbox"] or "null") pw, ph = (pb[2], pb[3]) if pb and len(pb) == 4 else (None, None) + ox, oy = (pb[0], pb[1]) if pb and len(pb) == 4 else (0, 0) assigns = {a["local_id"]: (a["character_id"], a["confidence"]) for a in c.execute( "SELECT local_id, character_id, confidence FROM identity_assignments WHERE panel_id=?", (p["panel_id"],))} assigned_total += len(assigns) for cid, _ in assigns.values(): - per_char[reg.get(cid, {}).get("name") or cid[:16]] += 1 + per_char[_label(cid)] += 1 people = [ch for ch in (v.get("characters") or []) if ch.get("bbox")] detections += len(people) + n_lines = len(v.get("dialogue") or []) + lines += n_lines + if v.get("skip"): + skipped += 1 + lines_skipped += n_lines for ch in people: x1, y1, x2, y2 = ch["bbox"] past_1000 += 1 if max(x2, y2) > 1000 else 0 at_1000 += 1 if 1000 in (x2, y2) else 0 if pw and ph: max_ratio = max(max_ratio, x2 / pw, y2 / ph) + cid = assigns.get(ch["local_id"], (None,))[0] + if cid: + found.append({"page": p["page_index"], "character_id": cid, + "box": [x1 + ox, y1 + oy, x2 + ox, y2 + oy]}) if p["panel_index"] == WORKED_EXAMPLE: worked = (p, v, people, assigns, pw, ph) -named = [r for r in reg.values() if (r["name"] or "").strip()] -print(f"registry: {len(reg)} characters, {len(named)} named -> {sorted((r['name'] or '') for r in named)}") +named = [r for r in live.values() if (r["name"] or "").strip()] +print(f"registry: {len(live)} live characters ({len(reg) - len(live)} merged away), " + f"{len(named)} named -> {sorted((r['name'] or '') for r in named)}") print(f"detections: {detections} assignments: {assigned_total} " f"= {100*assigned_total/max(detections,1):.0f}% coverage") +# a skipped panel is one vision judged to hold no character. It keeps its dialogue, so those lines +# have no visible speaker to attribute to and are the case for an `offscreen` speaker_ref kind. +print(f"vision skipped {skipped}/{len(panels)} panels, holding " + f"{lines_skipped}/{lines} dialogue lines") if per_char: top, n = per_char.most_common(1)[0] print(f"assignment spread: {dict(per_char.most_common(8))}") @@ -78,16 +191,58 @@ if per_char: print(f"bbox space: {past_1000}/{detections} boxes exceed 1000, {at_1000} sit exactly on 1000, " f"largest coord/panel-dimension = {max_ratio:.2f}") print(f" verdict: {'PIXELS' if past_1000 or max_ratio > 0.02 and at_1000 == 0 else 'STILL 0-1000 GRID'}") -missing_refs = [k for k, r in reg.items() if not r["ref_image_uris"] or not r["embedding_uri"]] -print(f"characters missing a ref crop or embedding: {len(missing_refs)}") +missing_refs = [k for k, r in live.items() if not r["ref_image_uris"] or not r["embedding_uri"]] +print(f"live characters missing a ref crop or embedding: {len(missing_refs)}") + +# 3. purity and fragmentation against the eyeball pass. Purity is per cluster, fragmentation is per +# real person. Both are computed off page-space geometry, so a re-crop and a /characters/reset do not +# invalidate the truth file. +truth = json.load(open(TRUTH)) if os.path.exists(TRUTH) else None +if not truth: + print(f"\nno truth file at {TRUTH}, skipping purity and fragmentation") +elif truth["chapter_id"] != CHAPTER: + print(f"\ntruth file is for chapter {truth['chapter_id'][:8]}, not this one. skipping.") +else: + rows, base = truth["occurrences"], truth.get("baseline", {}) + per_cluster, frag, unmatched = _score(rows, found) + matched = len(rows) - len(unmatched) + print(f"\ntruth: {len(rows)} labelled occurrences, {matched} matched an assignment " + f"at IoU >= {IOU_MIN}, {len(unmatched)} unmatched") + for u in unmatched: + src = u.get("from_19_44", {}) + print(f" unmatched {u['person']:14} was ord {src.get('panel_order')} {src.get('local_id')} " + f"box {u['box']}") + print("purity per cluster, dominant person first:") + for cid, cnt in sorted(per_cluster.items(), key=lambda kv: -sum(kv[1].values())): + person, correct = cnt.most_common(1)[0] + total = sum(cnt.values()) + want = (base.get("clusters") or {}).get(person) + delta = "" + if want: + delta = (" = baseline" if (want["assignments"], want["correct"]) == (total, correct) + else f" vs baseline {want['correct']}/{want['assignments']} = {want['purity']:.2f}") + print(f" {_label(cid):22} {person:14} {correct}/{total} = {correct/total:.2f}{delta}") + for other, n in cnt.most_common()[1:]: + print(f" {'':22} {'wrong: ' + other:14} {n}") + print("fragmentation per person, ids holding their occurrences:") + for person in truth.get("people", {}): + want = (base.get("fragmentation") or {}).get(person) + got = frag.get(person, 0) + print(f" {person:14} {got}" + ("" if want is None else + (" = baseline" if got == want else f" vs baseline {want}"))) + scope_only = len(found) - matched + print(f"assignments outside the truth's scope: {scope_only}/{len(found)}, never checked by eye") if worked: p, v, people, assigns, pw, ph = worked print(f"\npanel_index {WORKED_EXAMPLE} ({p['panel_id']}), {pw}x{ph}:") + if not people: + print(f" no boxed detection. skip={v.get('skip')!r} " + f"dialogue_status={v.get('dialogue_status')!r}, " + f"{len(v.get('characters') or [])} unboxed character entries") for ch in people: cid, conf = assigns.get(ch["local_id"], (None, None)) - name = reg.get(cid, {}).get("name") or (cid[:16] if cid else "-- none --") - print(f" {ch['local_id']:10} {ch['bbox']!s:28} {name:22} " + print(f" {ch['local_id']:10} {ch['bbox']!s:28} {_label(cid) if cid else '-- none --':22} " f"{'' if conf is None else f'{conf:.2f}'}") else: print(f"\npanel_index {WORKED_EXAMPLE} not found in this chapter") diff --git a/decisions/CLAUDE.md b/decisions/CLAUDE.md index 918048e..048da74 100644 --- a/decisions/CLAUDE.md +++ b/decisions/CLAUDE.md @@ -54,3 +54,6 @@ still live belongs in `caveats/`. | [A name is a word set, not a string](identity-naming.md#alias-grouping) | closed, GPU pending | | [A confident caption names a character on its own](identity-naming.md#caption-is-strong) | closed, GPU pending | | [A name belongs to one character](identity-naming.md#one-name-one-character) | closed, GPU pending | +| [The truth file keys on page-space geometry](measurement-spine.md#truth-is-page-geometry) | closed | +| [Purity is a share, so it never names an id](measurement-spine.md#purity-is-id-free) | closed | +| [The truth is 38 assignments, and holds no dialogue yet](measurement-spine.md#truth-scope-is-38) | open | diff --git a/decisions/measurement-spine.md b/decisions/measurement-spine.md new file mode 100644 index 0000000..46cdaf4 --- /dev/null +++ b/decisions/measurement-spine.md @@ -0,0 +1,56 @@ +# The measurement spine + +`eval/chapter-truth.json` plus the purity and fragmentation section of `audit_registry.py`. This is +`ARCHITECTURE.md` step 0, and nothing below it in that build order is measurable without it. + +## The truth file keys on page-space geometry {#truth-is-page-geometry} + +**Closed, 2026-08-13.** Reproduces the baseline on the 19:44 run of 2026-08-12. + +A truth row cannot name a `panel_id` or a `character_id`. A re-crop remints panel ids, and +`/characters/reset` remints character ids. The fifth cycle does both, so a truth file keyed on either +would die the moment it was first used. + +Each row carries `[x1, y1, x2, y2]` in page space, which is the panel's `bbox` offset plus the +detection box. The page strip is the fetched artifact and does not change when panels are re-sliced, so +the key survives. `audit_registry.py:_score` matches a row to an assignment on the same page at +`IOU_MIN = 0.5`, greedily, best overlap first, each row and each assignment used once. + +The 19:44 `panel_order`, `local_id` and `character_id` are kept per row under `from_19_44`, for +debugging only. Nothing matches on them. + +## Purity is a share, so it never names an id {#purity-is-id-free} + +**Closed, 2026-08-13.** + +Purity of a cluster is the largest share of one true person inside it. The lead's 16 assignments hold +14 of him, so 0.88, and the truth file never has to say which id he was. Fragmentation of a person is +the number of clusters holding their occurrences. Woman A has 15 occurrences over 2 ids, so 2. + +So both numbers survive a reset. The definitions in `ARCHITECTURE.md` were not explicit about that. +A cluster is reported under its dominant person. Every other person in it prints as a `wrong:` line, +which is where the photograph and the chibi show up on the lead. + +The `baseline` block in the truth file holds the 19:44 numbers, and each printed line says +`= baseline` or the delta. That makes step 0's DoD a check rather than an eyeball. + +## The truth is 38 assignments, and holds no dialogue yet {#truth-scope-is-38} + +**Open, 2026-08-13.** The character half is done, the dialogue half is not written. + +The eyeball pass covered three characters, 38 of the run's 60 assignments. The other 10 characters were +never checked and carry no truth row. The audit prints them as out of scope rather than as correct. +Anything claiming chapter-wide accuracy off this file is overreaching. + +The prose in `NEXT.md` left two rows unnamed. They were identified on 2026-08-13 by re-cropping `p030` +and `p032`: `panel_order` 31 and 33 on woman B's id are both woman A. That closes the recorded +arithmetic of 7 correct and 2 wrong, and it is why woman A's fragmentation is 2. + +One judgement is recorded in the file rather than hidden. The chibi of woman B at `panel_order` 28 +counts as her, because it is a stylised rendering of the subject. The chibi of the lead at +`panel_order` 20 does not, because it is a drawing inside the scene. `not_people` names the second kind. + +`ARCHITECTURE.md` also asks for 30 dialogue lines with their true speaker, typed +`visible | offscreen | narrator | unknown`. `dialogue` is an empty array. No such pass exists, because +`NEXT.md` records no per-line speaker truth. Writing it means watching the chapter panel by panel. The +audit prints nothing about speakers until it is filled. diff --git a/eval/chapter-truth.json b/eval/chapter-truth.json new file mode 100644 index 0000000..54a5180 --- /dev/null +++ b/eval/chapter-truth.json @@ -0,0 +1,669 @@ +{ + "chapter_id": "7c944dd4-e972-42c7-ba60-9f6939548e80", + "manga_id": "ef105a86-4b7e-4ac4-b45c-b7d83b8f5b5e", + "source": "the eyeball pass over the 19:44 UTC run of 2026-08-12, recorded as prose in NEXT.md. The two woman_a crops sitting on woman_b's id were unnamed there and were identified on 2026-08-13 by re-cropping p030 and p032 from s3://panels.", + "scope": "three characters walked crop by crop, 38 of the run's 60 assignments. The other 10 characters were never checked by eye and carry no truth row.", + "keyed_on": "page-space pixel box, [x1,y1,x2,y2] = the panel's bbox offset plus the detection box, matched by IoU. Not panel_id and not character_id: a re-crop remints panel ids and /characters/reset remints character ids, and the fifth cycle does both.", + "people": { + "lead": { + "label": "the lead, named LIM SEONHO in the 19:44 registry", + "gender": "m" + }, + "woman_a": { + "label": "woman A, black bob, cream top with black trim, pearl necklace. A main character the 19:44 registry never named. Expected to become Choi Haeseon from the p040 caption.", + "gender": "f" + }, + "woman_b": { + "label": "woman B, the green-dress woman, brown side ponytail with a scrunchie. Wrongly named Seonho in the 19:44 registry.", + "gender": "f" + } + }, + "not_people": { + "photo_of_lead": "a photograph of the lead inside a panel", + "chibi_of_lead": "a chibi drawing of the lead on a monitor" + }, + "baseline": { + "run": "2026-08-12 19:44-19:52 UTC, on a registry reset beforehand", + "detections": 119, + "assignments": 60, + "clusters": { + "lead": { + "assignments": 16, + "correct": 14, + "purity": 0.88 + }, + "woman_a": { + "assignments": 13, + "correct": 13, + "purity": 1.0 + }, + "woman_b": { + "assignments": 9, + "correct": 7, + "purity": 0.78 + } + }, + "fragmentation": { + "lead": 1, + "woman_a": 2, + "woman_b": 1 + } + }, + "occurrences": [ + { + "page": 0, + "box": [ + 0, + 53814, + 810, + 54790 + ], + "person": "woman_a", + "from_19_44": { + "panel_order": 38, + "local_id": "person_1", + "character_id": "character_2b1b12a1", + "confidence": 1.0 + } + }, + { + "page": 0, + "box": [ + 0, + 56215, + 862, + 57486 + ], + "person": "woman_a", + "from_19_44": { + "panel_order": 40, + "local_id": "person_1", + "character_id": "character_2b1b12a1", + "confidence": 1.0 + } + }, + { + "page": 0, + "box": [ + 0, + 60385, + 900, + 61496 + ], + "person": "woman_a", + "from_19_44": { + "panel_order": 43, + "local_id": "person_1", + "character_id": "character_2b1b12a1", + "confidence": 1.0 + } + }, + { + "page": 0, + "box": [ + 496, + 66848, + 620, + 67051 + ], + "person": "woman_a", + "from_19_44": { + "panel_order": 48, + "local_id": "person_2", + "character_id": "character_2b1b12a1", + "confidence": 1.0 + } + }, + { + "page": 0, + "box": [ + 58, + 68579, + 760, + 69375 + ], + "person": "woman_a", + "from_19_44": { + "panel_order": 49, + "local_id": "person_1", + "character_id": "character_2b1b12a1", + "confidence": 1.0 + } + }, + { + "page": 0, + "box": [ + 285, + 78844, + 409, + 79211 + ], + "person": "woman_a", + "from_19_44": { + "panel_order": 57, + "local_id": "person_1", + "character_id": "character_2b1b12a1", + "confidence": 0.9 + } + }, + { + "page": 0, + "box": [ + 0, + 82703, + 757, + 83432 + ], + "person": "woman_a", + "from_19_44": { + "panel_order": 59, + "local_id": "person_1", + "character_id": "character_2b1b12a1", + "confidence": 0.9 + } + }, + { + "page": 0, + "box": [ + 310, + 86176, + 688, + 87211 + ], + "person": "woman_a", + "from_19_44": { + "panel_order": 61, + "local_id": "person_1", + "character_id": "character_2b1b12a1", + "confidence": 0.9 + } + }, + { + "page": 0, + "box": [ + 628, + 90916, + 898, + 91743 + ], + "person": "woman_a", + "from_19_44": { + "panel_order": 65, + "local_id": "person_2", + "character_id": "character_2b1b12a1", + "confidence": 0.9 + } + }, + { + "page": 0, + "box": [ + 0, + 94674, + 844, + 95892 + ], + "person": "woman_a", + "from_19_44": { + "panel_order": 67, + "local_id": "person_1", + "character_id": "character_2b1b12a1", + "confidence": 1.0 + } + }, + { + "page": 0, + "box": [ + 0, + 120322, + 898, + 121699 + ], + "person": "woman_a", + "from_19_44": { + "panel_order": 86, + "local_id": "person_1", + "character_id": "character_2b1b12a1", + "confidence": 1.0 + } + }, + { + "page": 0, + "box": [ + 0, + 151177, + 539, + 152023 + ], + "person": "woman_a", + "from_19_44": { + "panel_order": 108, + "local_id": "person_1", + "character_id": "character_2b1b12a1", + "confidence": 1.0 + } + }, + { + "page": 0, + "box": [ + 0, + 157221, + 862, + 158723 + ], + "person": "woman_a", + "from_19_44": { + "panel_order": 111, + "local_id": "person_1", + "character_id": "character_2b1b12a1", + "confidence": 0.897 + } + }, + { + "page": 0, + "box": [ + 0, + 8758, + 900, + 10118 + ], + "person": "lead", + "from_19_44": { + "panel_order": 10, + "local_id": "person_1", + "character_id": "character_b1dd5659", + "confidence": 0.0 + } + }, + { + "page": 0, + "box": [ + 412, + 11199, + 900, + 11986 + ], + "person": "lead", + "from_19_44": { + "panel_order": 11, + "local_id": "person_2", + "character_id": "character_b1dd5659", + "confidence": 0.9 + } + }, + { + "page": 0, + "box": [ + 523, + 16210, + 716, + 16532 + ], + "person": "lead", + "from_19_44": { + "panel_order": 13, + "local_id": "person_2", + "character_id": "character_b1dd5659", + "confidence": 0.826 + } + }, + { + "page": 0, + "box": [ + 454, + 20833, + 715, + 21006 + ], + "person": "photo_of_lead", + "from_19_44": { + "panel_order": 17, + "local_id": "person_1", + "character_id": "character_b1dd5659", + "confidence": 0.9 + }, + "note": "a photograph of him inside the panel, not him" + }, + { + "page": 0, + "box": [ + 202, + 22420, + 900, + 23566 + ], + "person": "lead", + "from_19_44": { + "panel_order": 18, + "local_id": "person_1", + "character_id": "character_b1dd5659", + "confidence": 1.0 + } + }, + { + "page": 0, + "box": [ + 0, + 24339, + 899, + 25041 + ], + "person": "lead", + "from_19_44": { + "panel_order": 19, + "local_id": "person_1", + "character_id": "character_b1dd5659", + "confidence": 1.0 + } + }, + { + "page": 0, + "box": [ + 30, + 26777, + 202, + 27204 + ], + "person": "chibi_of_lead", + "from_19_44": { + "panel_order": 20, + "local_id": "person_3", + "character_id": "character_b1dd5659", + "confidence": 1.0 + }, + "note": "a chibi of him on a monitor, art inside the panel" + }, + { + "page": 0, + "box": [ + 106, + 51044, + 741, + 51945 + ], + "person": "lead", + "from_19_44": { + "panel_order": 36, + "local_id": "person_1", + "character_id": "character_b1dd5659", + "confidence": 0.9 + } + }, + { + "page": 0, + "box": [ + 26, + 64499, + 730, + 65435 + ], + "person": "lead", + "from_19_44": { + "panel_order": 47, + "local_id": "person_1", + "character_id": "character_b1dd5659", + "confidence": 1.0 + } + }, + { + "page": 0, + "box": [ + 382, + 66803, + 464, + 67054 + ], + "person": "lead", + "from_19_44": { + "panel_order": 48, + "local_id": "person_1", + "character_id": "character_b1dd5659", + "confidence": 0.888 + } + }, + { + "page": 0, + "box": [ + 97, + 71438, + 619, + 72061 + ], + "person": "lead", + "from_19_44": { + "panel_order": 51, + "local_id": "person_1", + "character_id": "character_b1dd5659", + "confidence": 1.0 + } + }, + { + "page": 0, + "box": [ + 463, + 91111, + 675, + 91494 + ], + "person": "lead", + "from_19_44": { + "panel_order": 65, + "local_id": "person_1", + "character_id": "character_b1dd5659", + "confidence": 1.0 + } + }, + { + "page": 0, + "box": [ + 0, + 97794, + 900, + 98950 + ], + "person": "lead", + "from_19_44": { + "panel_order": 69, + "local_id": "person_1", + "character_id": "character_b1dd5659", + "confidence": 1.0 + } + }, + { + "page": 0, + "box": [ + 251, + 110663, + 788, + 112008 + ], + "person": "lead", + "from_19_44": { + "panel_order": 80, + "local_id": "person_1", + "character_id": "character_b1dd5659", + "confidence": 1.0 + } + }, + { + "page": 0, + "box": [ + 316, + 114988, + 653, + 115534 + ], + "person": "lead", + "from_19_44": { + "panel_order": 82, + "local_id": "person_1", + "character_id": "character_b1dd5659", + "confidence": 1.0 + } + }, + { + "page": 0, + "box": [ + 0, + 123335, + 899, + 124244 + ], + "person": "lead", + "from_19_44": { + "panel_order": 88, + "local_id": "person_1", + "character_id": "character_b1dd5659", + "confidence": 1.0 + } + }, + { + "page": 0, + "box": [ + 648, + 4620, + 767, + 4830 + ], + "person": "woman_b", + "from_19_44": { + "panel_order": 7, + "local_id": "person_6", + "character_id": "character_f0d4e901", + "confidence": 1.0 + } + }, + { + "page": 0, + "box": [ + 94, + 11408, + 540, + 11986 + ], + "person": "woman_b", + "from_19_44": { + "panel_order": 11, + "local_id": "person_1", + "character_id": "character_f0d4e901", + "confidence": 0.0 + } + }, + { + "page": 0, + "box": [ + 350, + 15975, + 437, + 16256 + ], + "person": "woman_b", + "from_19_44": { + "panel_order": 13, + "local_id": "person_1", + "character_id": "character_f0d4e901", + "confidence": 1.0 + } + }, + { + "page": 0, + "box": [ + 0, + 35632, + 421, + 36197 + ], + "person": "woman_b", + "from_19_44": { + "panel_order": 26, + "local_id": "person_1", + "character_id": "character_f0d4e901", + "confidence": 0.933 + } + }, + { + "page": 0, + "box": [ + 79, + 38398, + 390, + 38652 + ], + "person": "woman_b", + "from_19_44": { + "panel_order": 28, + "local_id": "person_1", + "character_id": "character_f0d4e901", + "confidence": 1.0 + }, + "note": "a chibi of her as an emote, counted as her. Distinct from chibi_of_lead, which is a drawing inside the scene rather than a stylised rendering of the subject." + }, + { + "page": 0, + "box": [ + 0, + 41191, + 435, + 42178 + ], + "person": "woman_b", + "from_19_44": { + "panel_order": 29, + "local_id": "person_1", + "character_id": "character_f0d4e901", + "confidence": 1.0 + } + }, + { + "page": 0, + "box": [ + 322, + 43342, + 886, + 44191 + ], + "person": "woman_b", + "from_19_44": { + "panel_order": 30, + "local_id": "person_1", + "character_id": "character_f0d4e901", + "confidence": 0.943 + } + }, + { + "page": 0, + "box": [ + 471, + 45263, + 525, + 45385 + ], + "person": "woman_a", + "from_19_44": { + "panel_order": 31, + "local_id": "person_3", + "character_id": "character_f0d4e901", + "confidence": 1.0 + }, + "note": "woman_a in her cream top, verified 2026-08-13 on a p030 zoom" + }, + { + "page": 0, + "box": [ + 81, + 47159, + 899, + 48564 + ], + "person": "woman_a", + "from_19_44": { + "panel_order": 33, + "local_id": "person_1", + "character_id": "character_f0d4e901", + "confidence": 0.853 + }, + "note": "woman_a, unmistakable at full panel width" + } + ], + "dialogue": [], + "dialogue_todo": "ARCHITECTURE.md's spine also asks for 30 dialogue lines with their true speaker, typed visible | offscreen | narrator | unknown. No such eyeball pass has been done: NEXT.md records no per-line speaker truth, and the dialogue stage is cleared, so there is nothing to transcribe. It has to be watched panel by panel and written here." +}