this is a v3 delta spec: it lists what changes on top of the working v2 pipeline (see spec-v2.md). v2 is validated end-to-end (a full chapter renders to one narrated mp4). v3 is about quality of the narration and the look of the video, plus panel curation. structured for hand-off to an agent. # manga recap v3 — quality + look ## 0. what v2 does today (baseline) - one narrated `chapter.mp4` per chapter: fit-on-black panels, uniform ken-burns zoom, burned subtitles, hard cuts, single default TTS voice. - identity assigns a stable `character_id` per body (siglip2), but characters are unnamed, so the script prints "Person A/B". the script worker sees one panel at a time (no memory). - every panel is narrated (no curation). --- ## A. cross-panel identity & continuity [biggest quality lever] [DONE] SHIPPED: - descriptive frozen labels: worker_scene emits `label` per character (name -> frozen registry appearance via `_describe`, "the one with black hair and glasses" -> Person X fallback); worker_script `_name_map` consumes it. no "Person A" when appearance is known, char_id preserved. - chapter_context populated: service.py `_chapter_context` feeds a cast+entities roster to every script call so the narrator has memory (kills "Unknown voice", item 11). - name acquisition + backfill (item 12): vision emits `name` only when a char is explicitly named; service.py scene stage binds it via `set_character_name` (first name wins, only if unnamed); script stage resolves display names from the FINAL registry (`names_by_id`) so a name learned in ANY panel backfills to every panel's narration. ceiling: vision name-inference can misattribute; first confident name sticks. upgrade path = roster seed for higher precision. STILL TODO (small, deferred): - true per-panel story-state (only what's been seen so far) for "establishing shot once / don't re-narrate repeats" (the C-todo) — currently whole-manga roster. --- original plan --- one feature: carry a running **story-state** through a chapter so narration has memory. covers former items 4, 11, 12. - **stable anchor, mutable label.** the `character_id` is never lost (already stable via siglip2). only its *display label* changes: - name known -> the name. - name unknown -> a **descriptive phrase** ("guy with black hair and glasses", "girl with short hair holding a phone") derived ONCE from vision's appearance data and frozen per character_id, reused across all panels. never regenerate per-panel (labels would drift). - **name acquisition + backfill.** when dialogue addresses a character by name and speaker attribution is confident, bind name -> character_id and switch the label to the name from then on (optionally backfill earlier references). - **story-state threaded per chapter.** accumulate {cast seen (id,label,name?), entities/props seen (the cat)} and feed it into scene+script (populate the currently-empty `chapter_context`). this is what lets the narrator say "he calls out to his cat" instead of leaking a bare name (item 12), and removes "Unknown voice" (item 11). - **optional roster seed (not required).** a per-manga character list (names, maybe ref art) makes naming + cross-panel matching more reliable than dialogue inference. graceful-degrade without it; add when available. touches: worker_vision (appearance -> label), worker_scene / orchestrator (story-state store + speaker continuity), worker_script (consume context), orchestrator characters table (label, name). ## B. narration voice & style (items 2, 3) [DONE] - **less descriptive gemma** (item 2): DONE. prompt now demands TIGHT recap (1-2 sentences, ~35 words, "no purple prose"); max_tokens 512->160, temp 0.7->0.6 (worker_script). - **one fixed voice** (item 3): DONE. dots.tts samples a RANDOM speaker per call with no reference, so the voice drifted. worker_tts now bootstraps one seeded reference clip (`~/.cache/manga-tts/ narrator_ref.wav`) and clones from it every synth. delete the file to reroll; swap REF_WAV/REF_TEXT for a hand-picked voice. - **multi-voice per character: future**, not v3 (user: "later") — same mechanism, per-character ref clips. ## C. panel curation (item 10) [DONE — basic] - skip panels with no narrative content (pure scenery / transition art); show an establishing shot once, don't re-narrate repeats. vision flags "scenery-only / skip". big pacing + cost win; do early because it shrinks every downstream stage. - SHIPPED: vision emits `skip:true` for no-narrative panels (worker_vision prompt+schema); the script stage honors it (service.py run_stage_script) -> no narration -> cascades to no audio, no clip. Single chokepoint, no new db columns (skip rides in the vision result_json blob). - STILL TODO (needs cross-panel memory, folds into A): "establishing shot once, don't re-narrate repeats" — requires dedup of repeated scenery across panels. deferred to story-state. ## D. render / look overhaul (items 1, 5, 6, 7, 8) all in worker_render.py (self-contained ffmpeg builder), except #6. - **#1 dynamic clip duration** — DONE. `_audio_dur` probes narration length via ffprobe; scene_timing kept only as fallback. no more 4.0s sawtooth / vanishing subs. - **#5 blurred-panel background fill** — DONE. filter_complex: cover-scaled + boxblur copy behind the fitted panel (replaces black bars). - **#7 subtitles** — DONE. boxed caption: bold white on translucent black box (BorderStyle 3), lower-center, DejaVu Sans 52. (user chose boxed over outline/minimal.) - **#8 content-aware motion** — DONE. vision picks a camera.effect per panel from the fixed vocab (+ to=[x,y] for dolly_to_subject); it rides the vision result_json -> scene_graph.camera -> render. worker_render `_motion(camera, frames)` maps the vocab to a zoompan z/x/y expression (static/zoom/pan/dolly/shake/orbit); default zoom_in = old ken burns. orchestrator render stage reads scene_graph.camera. self-checked (ffmpeg renders pan + dolly clips). - **#6 multi-panel composites + non-generic transitions** [the ambitious one, sequenced last] - **non-generic transitions — DONE.** vision emits `transition` (cut|crossfade|dissolve| fade_black|fade_white|wipe_left|wipe_right|push) per panel -> scene_graph.transition -> orchestrator assemble collects a per-clip list -> worker_render `_xfade_chain` builds an xfade+acrossfade graph (audio kept in sync, overlap clamped to clip length). all-cut/empty keeps the fast stream-copy concat (no re-encode); any real transition re-encodes the chapter. self-checked (two clips fade_white into one mp4). transition = OUT of that panel. - **multi-panel composites — DONE (basic).** grouping source = new pass (`grouping.py`, deterministic): consecutive SMALL panels (area < 0.7*page-max, same source page, capped at MAX_GROUP=3) share ONE shot. render + assemble both call plan_groups so their grouping matches. a composite = vertical stack of the members on a W×H still; each panel keeps its own script+tts (per-panel stages unchanged); the audios play in sequence while a moving highlight marks the active row ("swap the front panel while audio plays"). clip keyed by the group LEADER; assemble takes one clip per group and the group's exit transition = last member's. solo groups render exactly as before, so blast radius is only genuinely-small adjacent panels. knobs: MAX_GROUP, SMALL_FRAC. ceilings: vertical-stack layout only (no grid/overlay/front-swap animation beyond the highlight), size+adjacency grouping (no narrative-beat grouping), black gutters (no blurred fill in composites). upgrade paths noted in code. ## D2. beat pacing, timed cues, subtitle redesign, collage (vikunja 185/187/188/183) scene-level narration made the beat (not the panel) the render unit; these refine how a beat plays. all in worker_render.py + collage.py, orchestrator run_stage_render. - **185 content-aware panel timing — DONE.** a beat's one narration used to split EQUALLY across its member panels (`D/n`). now the orchestrator scores each member from its scene graph (dialogue amount, action density, camera emphasis: `_beat_weights`) and passes bounded weights to render; worker `_beat_slices(D, n, weights)` distributes screen-time by weight with a per-panel floor (no flashing) and a deterministic equal-split fallback (degenerate/short weights). self-checked. - **187 timed subtitle cues — DONE.** the whole narration paragraph used to show from frame 1, leaking later-panel lines early. `cue_plan` splits the flowing narration into sentence cues, times them on the SAME [0,D] timeline as the panel slices (length-weighted + readable floor, reused `_beat_slices`), maps each cue to the panel on screen at its start, and clamps a cue never to precede its panel. per-cue ASS events (not one paragraph). cue plan returned + persisted to `clips.cues_json` for the review UI / collage. no-TTS-alignment fallback = punctuation/length estimate. self-checked ("Hold up." can't appear before its panel). - **188 subtitle redesign — DONE.** `_ass_multi` now takes mode (off|minimal|boxed, env `SUB_MODE`, default minimal) + orientation presets (portrait/landscape safe-areas). smaller type (40/32 vs 52), minimal = thin outline + soft shadow (no box), boxed = restrained ~50% box. `_wrap2` caps every cue at ≤2 short lines (ellipsis truncation) so it never regrows into a paragraph. `_sub_align` dodges the subject vertically using the director focus point (camera.to.y low -> caption to top). ceiling: no real face/bubble CV yet — camera focus is the only subject signal wired. - **183 animated manga collage — DONE (core + wired behind `COLLAGE` flag).** `collage.py plan_layout` is a pure deterministic planner: ordered panel aspects + RTL + dominant + frame -> template (centered_wide / vertical_pair / stacked_wides / strip_over_dominant / supporting_left_dominant_right / quad), aspect-correct resting rects (no stretch), entrance vectors (RTL slide / dominant scale), z-order, hold/transition timing. `worker_render collage_cmd` renders it: blurred plate of the dominant panel + sharp aspect-fit panels with restrained drop shadows, non-dominant panels sliding into place over ~0.45s, crisp holds; reuses 185 slices + 187 cues + 188 subtitles. orchestrator run_stage_render calls it when `COLLAGE=1` (else the ken-burns montage). ceilings (QA-tuned against the reference video, no still to eyeball here): transition-only directional/radial MOTION BLUR not applied (clean slide/scale); per-panel hold DRIFT omitted (static hold); beat-replace "streak" rides the assemble xfade. planner + renderer both self-checked; visual acceptance is the QA run. ## D3. identity tracklets + review gates (vikunja 160/136) - **160 tracklet spine — DONE (code, QA-pending).** the Tier-2 gemma decider (identity phase-2) used to resolve identity once PER CROP, so one person got adjudicated repeatedly and could get two answers on two panels. `tracklets.py link_tracklets` (pure, union-find, self-checked) first LINKS the per-panel siglip shortlists into per-person tracklets over a locality window — hard gender gate + appearance-token overlap + shared-candidate — then the resolver runs gemma ONCE per tracklet against the UNION of the members' shortlists (the person's gallery, consensus-ordered) and applies that one character_id to every crop in the tracklet. fewer gemma calls, one answer per person. ceiling: gemma NONE (cosine over-merged onto an existing char) still isn't split to a fresh id — needs a 3rd siglip re-embed pass; add if the NONE-rate is high. - **136 review gates — DONE (code, QA-pending), off by default (`GATES=1`).** two human sign-off gates in the pipeline loop: gate `script` pauses BEFORE tts/render, gate `audiovisual` pauses BEFORE assemble. When enabled and unapproved the job goes `awaiting_review` and the loop returns; `/review/approve` records the gate (`gates` table) and relaunches `_run_pipeline` (skips completed stages → resumes at the gated stage). Gate-1: `/review/scripts` shows narration per panel grouped by beat + cheap textual flags (`review_flags.py`: generic-handle / repeated / empty, self-checked; semantic flags — spoilers, altered meaning, attribution — left to the human eye, ceiling noted). Gate-2: `/review/audio` serves the per-panel wav so the reviewer actually listens; `/review/retts` re-synthesizes ONE panel's TTS + re-renders just its beat clip (delete_clip forces it) and clears the AV gate. web UI (index.html review tab): gate-approve buttons + status, per-panel audio player, flags badge, re-synthesize button. reuses existing /review/panels + /review/script + /review/preview. ceiling: no panel-to-cue drag-edit UI (cues persisted in clips.cues_json, edit endpoint not built); flags are textual only. ## E. reference generic manga-recap youtube style (single narrator, static panels w/ slow pan/zoom, minimal text, dramatic pacing) — but with **less generic transitions** (feeds #6). #6 goes beyond the reference. --- ## sequencing (proposed) 1. **C panel curation** — shrinks everything downstream first. 2. **A identity/continuity** — most changes how it *feels* (names + memory). 3. **B voice+style** — small, high impact. 4. **D basics (#1,#5,#7,#8)** — most changes how it *looks*, low risk. 5. **D#6 compositing** — last; highest risk/effort, isolated to render. ## direction schema (decided — enables #8 and #6) narration (text->TTS) and direction (camera/transition/layers->render) are separate concerns from the SAME scene graph. VISION picks the effect (it sees the panel + has char bboxes for focus points) and emits a direction block into the scene_graph JSON; RENDER just executes it. fixed vocab: - camera.effect: static | zoom_in | zoom_out | pan_left | pan_right | pan_up | pan_down | dolly_to_subject | orbit | shake | hold (+ from/to [x,y] normalized, easing) - transition.type: cut | crossfade | fade_black | fade_white | wipe_left | wipe_right | push | dissolve (+ duration) - layers: {artifact, motion} (parallax, when the depth renderer is wired) not built yet; lands before #8 (needs camera) and #6 (needs transition). ## open design points - roster seed format if/when available. - B#3: which fixed dots.tts reference voice to pin for the whole video (needs a pick). --- # v3.1 delta — identity is broken; add human-in-loop review ## evidence (last job, 116-panel Teto X Egen, from manga.db) - **53 character rows for ~8 actual people.** "Seonho" stored 5× (ids 85ad36b0/811bf0d7/0033ff7f/ bbddf20c/90269899), Egen Girl 2×. ~40 rows have a single identity_assignment (one-time NPCs). - **root cause: siglip2 over-splits.** same person across panels (pose/expression/B&W line art) doesn't clear cosine 0.85, so nearly every appearance spawns a new character_id. - **cascade:** names never propagate (panel 10 names id A, panel 51's Seonho is id C -> `names_by_id` can't map it -> narration falls back to "the woman with straight dark hair", flips gender to "her"). re-introductions + DB/S3 bloat all follow from the split. - **concrete bug:** `worker_identity.py` create call hardcodes `"name": None` — vision reads the name but identity discards it at registration. `db.create_character` is a blind INSERT, zero dedup. ## E1. pipeline decomposition [DECIDED — the vision mega-call splits] Reason: one gemma4 call doing 4 jobs (detect + dialogue + direction + scene) does each worse, and there was nowhere to run a real cross-panel merge. Split the jobs; insert a merge stage where the model actually compares two crops. **Final stage order** (OCR removed — gemma4 transcribes bubble text image-natively; a `roster` prepass now runs after crop, see below): ``` fetch → crop → roster → vision → identity → reconcile → dialogue → direct → scene → script → tts → layers → render → assemble ``` - **`vision`** (per panel, image call) — character DETECT only: `local_id`, `bbox`, `appearance` (hair/clothing/features), `gender` (m|f|unknown), `emotion`, `action`. plus `skip` + `scene` (location/time). the only stage that must precede identity (identity crops its bboxes). - **`identity`** (per panel) — siglip assigns `local_id → character_id`. on a collision (siglip ambiguous OR hair+clothing appearance overlap OR name collision) it records a **merge-candidate pair** instead of silently minting a twin. still creates provisional ids; reconcile prunes them. - **`reconcile`** (NEW, once per chapter, after all panels) — the second vision pass. for each flagged candidate pair, `POST worker_vision /vision/same` with the two `ref_image_uris` crops -> gemma4 "same character? {yes|no, confidence}". confirmed -> `db.merge_characters(loser->keeper)`: repoint `identity_assignments`, union `name`/`aliases`/`ref_image_uris`, delete loser row. loser ref crops are PRESERVED on the keeper as gallery views (not deleted from S3 — deleting them left the keeper's own uris dangling). cast is CLEAN + named from here on. only flagged pairs get a call. lives as `run_stage_reconcile` in service.py + `/vision/same` on worker_vision (reuses warm gemma4). - **`dialogue`** (per panel, image call) — gemma reads bubble text straight from the panel image (no external OCR): bubble classify (speech/thought/shout/narration/sfx), merge split bubbles, text cleanup (sentence case, fix names/proper nouns, sfx literal), speaker attribution (`speaker=local_id` + `confidence`), entities. prompt gets the CLEAN named cast present in the panel as context -> better attribution ("Seonho (m) vs Haeseon (f)"). - **`direct`** (per panel, off scene understanding) — `camera.effect` (+`to`) and `transition` from the fixed direction vocab (unchanged schema, section D/direction-schema). runs after identity so it can weight who's present. may glance at the image or work from the scene graph — cheap. new workers/endpoints: `worker_vision` becomes detect + `/vision/same`; new `worker_dialogue.py` and `worker_direct.py` (or `/dialogue` + `/direct` endpoints — same warm gemma4, decide at build like reconcile). db: `STAGES` gains `reconcile`,`dialogue`,`direct`; `run_stage_*` for each; session_proxy call_* wrappers. the old omni-`vision` prompt/schema is partitioned across the four prompts. ## F. identity overhaul [rides on E1] Goal: 53 rows -> ~12; named recurring cast stays one id; NPCs never persist. - **F1 merge = the `reconcile` stage** (see E1). the fix for the hardcoded `"name": None` bug: vision/ identity now carry `name`+`gender`+`appearance` into `create_character` so provisional rows are named/typed; reconcile does the actual dedup with the model comparing crops (not a blind string match). `db.merge_characters` + candidate-flagging in identity. - **F2 gender field.** add `gender TEXT` to characters schema. `vision` emits `gender`. worker_script `_name_map` feeds it to the model ("Seonho (he)") for correct pronouns — kills "her bold personality" flips even when the name resolves. - **F3 confirm-before-persist.** hold unnamed provisional embeddings in a per-session cache; only write a DB row + S3 crop/npy when an embedding is seen >=2x OR is named. the ~40 one-timers never touch the bucket. session_id already flows into identity. (complements reconcile: reconcile merges the twins that WERE created; F3 stops most from being created at all.) - **F4 centroid + re-rank matching.** store multiple refs per char, match against the accumulated centroid not one arbitrary embedding; break siglip ties by name/gender agreement. ceiling: perfect auto-ID on B&W is not achievable — reconcile+F4 get ~80%, the UI (G) closes the rest. don't over-invest. - **DECIDED: start clean.** no migration script. wipe manga.db + the minio `manga` bucket before the first v3.1 run (destructive ops step, run at build time). fixed logic applies to fresh data. ## G. review / edit UI (replaces the auto CoT-strip idea, item C-bleed) Rationale: the human reviews the final video and knows exactly what's wrong and where. Give them a synced editor instead of chasing every failure through the pipeline. ~80% of infra already exists in the orchestrator: `static/index.html` dashboard, `/job/video` (streams mp4), `/db/tables` + `/db/table`, `/stage/run` + `/stage/clear` (re-render), `set_character_name`. - **timestamp -> panel is trivial:** clips concat in `panel_order`, so panel start offset = cumsum of prior panels' `audio.duration`. store the offset at assemble time (or compute on the fly) as an offset table the page binary-searches on `