Files
manga-recap-pipeline/decisions/identity-bbox.md
T
kami ca4661763c Stop a background extra's action reaching narration
build_scene already dropped an unassigned detection from `characters` and
`present`, so an extra never reached the cast list. Its ACTION did. `actions` was
built from every detection, and that list is what the script prompt renders and
what the verifier uses as evidence, so "standing at the window" arrived as a fact
about the panel with no character attached and the verifier confirmed it, because
the action really was in the blob.

Skip has_face is False, the same gate and the same fail-open semantics as
enrollment. Self-check covers all three cases: a real cast member's action
survives, a faceless one's does not, and a detection from a panel where the
detector never ran keeps its action.

decisions/identity-bbox.md#extras-gate-consumers

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 19:16:30 +04:00

11 KiB

identity-bbox

The coordinate space of a vision character box, and what reads it.

A vision bbox arrives on gemma's 0-1000 grid, and /vision converts it to pixels

Closed, 2026-08-12.

build_detect_prompt asks for a "pixel bounding box". The model answers on its own normalized grid regardless. Measured over all 113 detections of job 778297bc, read straight from /review/identity:

test result
boxes with x2 past the 900px panel width 47 of 113
boxes with y2 past 1000, on panels 1257 to 2307px tall 0 of 113
boxes clamped at exactly 1000 21 in x, 5 in y
coordinate range over every box [0, 1000]

Pixels cannot behave that way. A person standing in the lower half of a 2307px panel needs y2 near 2000, and it never once exceeds 1000.

Consumed as pixels the boxes collapse into the top-left corner of the panel. Four consumers were reading them:

  • worker_identity._crop_bbox at worker_identity.py:200, which embeds the crop. This is why a crop of a speech balloon's edge matched Choi Haeseon at 0.9.
  • _pair_faces_to_present in worker_vision.py, which compares real detector face boxes, in pixels, against these. The gate could almost never pass, which is the mechanism behind the 7 unknown results out of 7 som_face lines already recorded at worker_vision.py:169.
  • the set-of-mark boxes drawn for attribution.
  • the review UI, which crops client-side off the panel PNG.

/vision now calls _bbox_to_pixels(characters, w, h) before returning, so all four see pixels and no consumer needs to know the grid existed. Verified by drawing the converted boxes on panel 7. Five of six land on their subject, including person_5, who is Seonho in the foreground and had no identity. person_1 still frames a window mullion with nobody in it, which is #extras-as-cast, not this.

The prompt text still says "pixel bounding box". Rewording it changes what the model emits and needs a GPU run to re-verify, so the boundary converts instead. The ponytail: note on _bbox_to_pixels records that. It also records the trap: a model that really answered in pixels would be scaled down here.

Consequence: every assignment in the registry came from a wrong crop. The existing embeddings and ref_image_uris are enrolled on balloons and window frames. Re-running identity is what makes the registry mean anything. The anonymous-identity and extra-versus-cast work cannot be judged until that rerun happens.

A stage result proves nothing until the worker is newer than the edit

Closed, 2026-08-12. The first rerun after the bbox fix reproduced the defect exactly: 46 of 110 boxes past the 900px panel width, coordinates clamping at 1000, y2 never once past 1000 on panels up to 2307px tall. The same fingerprint as #bbox-is-normalized measured before the fix.

The fix was not wrong. It was not loaded.

vision worker process started   12:00:09
worker_vision.py modified       12:11:35
8113bdf, which contains _bbox_to_pixels, committed  12:16:22

Python binds a module once, at process start. ./start_workers.sh had launched the worker eleven minutes before the file changed, so /vision served pre-fix code for the whole run and returned raw grid boxes. Nothing in the result said so. The stage reported completed 116/116, the orchestrator recorded no error, and identity and reconcile ran to completion on top of it. Cost: one full vision + identity + reconcile cycle, plus a registry reset to undo the 8 characters it minted.

The rerun against a restarted worker gives the opposite reading over the same 110 detections: 0 boxes past the width, 0 past the height, one coordinate on 1000 which is now a real pixel value, and a deepest box reaching 100% down its panel with max y2 = 2307. Boxes track the panel, so they are pixels.

check_stale.sh compares every running worker's process start against its module's mtime and exits non-zero if any is stale. This failure mode was already known as advice — the render worker "must be restarted by hand to pick up an edit" — and advice did not stop it happening. Run the check before any stage run that is meant to prove a code change.

Forbids: citing a stage result as evidence about a code change without establishing that the worker serving it postdates the change.

A detection with no detected face never enrolls or binds

Closed, 2026-08-12. Written and self-checked, not yet proven on a GPU run.

Fixing the coordinate space made the extras problem worse, not better. With the boxes finally landing on their subjects, panel 7's four background extras became four good crops of four irrelevant people, and one of them bound to Seonho at confidence 1.00. Before the fix the same detection was a crop of scenery and matched nothing much. Correct geometry turned a harmless failure into a poisoned reference set for the lead.

The measured panel 7 outcome, converted boxes, against the art:

box who assigned
[457, 657, 642, 937] Seonho, foreground Seonho
[669, 591, 763, 822] the colleague, unnamed in the story character_f7a4fd, anonymous
[428, 386, 496, 526] background extra none
[34, 414, 122, 564] background extra none
[498, 386, 568, 533] background extra character_d72710 at 0.94
[31, 554, 94, 728] background extra Seonho at 1.00

/vision now stamps has_face on every character by running face_detect.detect_faces on the panel and reusing _pair_faces_to_present for containment, so the gate uses the same margin and the same global shortest-first assignment as the speaker path. worker_identity.py skips a character with has_face is False before it crops, embeds, matches or mints.

Two properties are deliberate. It fails open: a missing or raising detector marks every character True, because dropping a whole panel's cast is worse than the over-detection the gate exists to trim. And it gates on is False, not falsiness, so a vision blob written before this change (no key) behaves as it did rather than silently dropping every character.

Cost: a cast member drawn from behind, or in a style the detector misses, now takes no identity on that panel. That is the abstain this pipeline already prefers to a wrong bind (caveats/speaker-attribution.md#no-anonymous-identity).

Forbids: enrolling a reference crop, or binding a character, from a region no face detector confirms.

A resolver NONE mints an anonymous character, it does not clear the crop

Closed, 2026-08-12. Written and self-checked, not yet proven on a GPU run.

caveats/speaker-attribution.md#no-anonymous-identity asked whether the Tier-2 gemma resolver can answer "none of these". It can, and it always could. /vision/resolve at worker_vision.py:1071 maps choice: 0 to state="new", an out-of-range index to state="unresolved", and a parse failure to unresolved as well. The abstain path was never the defect.

The defect was one branch on the other side of the contract. The orchestrator read only v.get("character_id") and treated every falsy value the same way: unassign_identity on every crop of the tracklet. So a deliberate "this is a real person the roster does not hold" and a hallucinated index both produced nothing, and the unnamed colleague was unknown on every panel she appeared on. The stale ponytail: comment above that block named the reason nobody fixed it, and the reason was real: minting a character needs an embedding_uri, and the orchestrator cannot compute one. siglip is resident in the identity worker, gemma is resident in the vision worker, and session_manager forbids both at once.

What removes the blocker is carrying the embedding, not a third GPU pass. /identity/resolve already computes an embedding per crop and already uploads the crop to s3://manga/{manga}/characters/_crops/{panel}_{local}.png. It now writes the embedding to the same key with a .npy suffix and returns emb_uri in each shortlist entry. The mint is then a local create_character(manga_id, None, appearance, [crop_uri], emb_uri, gender), and the existing per-tracklet assign loop binds every member to it.

tracklets.resolve_outcome holds the three-way decision as a pure function, so the branch that runs is the branch the self-check covers: known on a named answer, mint on state="new" with an emb_uri, clear on unresolved, on a NONE with no embedding, and on an older worker that sends no state at all.

Two limits are deliberate. A tracklet's candidate gallery is built before the loop mints anything, so one person split across two unlinked tracklets still gets two anonymous ids; run_stage_reconcile merges unnamed twins on appearance overlap and is what folds them. And an anonymous character's text sheet (worker_vision._sheet) carries no name, so gemma re-recognising it on a later panel leans on the reference images rather than the description.

Contract: shortlists[].emb_uri is new in the /identity/resolve response. Invariant 7 — both repos changed in the same session.

Forbids: treating an absent character_id as one outcome. A resolver that answered and a resolver that failed are different facts.

The extras gate runs at enrollment and at narration, not at the speaker prompt

Closed, 2026-08-12. Written and self-checked, not yet proven on a GPU run.

#face-gates-enrollment stops a faceless detection taking an identity. It does not stop the detection being narrated, because three places read the raw vision character list and only one of them consults an assignment. Panel 7 is the worked example: six detections, four of them extras or scenery.

worker_scene.build_scene already drops an unassigned detection from characters and present (worker_scene.py:63), so an extra never reached the cast list. Its action did. actions was built from every detection, and that list is what the script prompt renders and what the verifier uses as evidence. So "standing at the window" arrived as a fact about the panel with no character attached, and the verifier confirmed it because the action really was in the blob. Both now skip has_face is False.

service._beat picks the cinematographer's "who" from the first three detections, falling back to a detection's action when it has no name. An extra could take a slot and steer the camera. Also gated.

service._present_characters is deliberately not gated. It builds the dialogue stage's candidate speaker list and the set-of-mark boxes. Two reasons. The failure is already contained: an extra chosen as the speaker has no identity assignment, so normalize_dialogue resolves it to unknown rather than to a wrong name. And the gate's own cost lands hardest here, because a character drawn from behind has no face box, so gating would delete a real speaker from the only list that can attribute their line.

All three gates test is False, never falsiness. A vision blob written before the gate existed carries no has_face key, and a panel whose detector failed is marked True by the fail-open path. Both keep their previous behaviour.

Forbids: adding a fourth consumer of vision["characters"] without deciding which side of this line it is on. The blob keeps every detection on purpose, so the audit can still see what was gated.