Working tree (including .git) was lost to an rm. Rebuilt by replaying Write/Edit/ Read/attachment events from 25 Claude sessions and 22 successful codex apply_patch blocks into one timestamp-ordered timeline. Verified against ground truth recorded in the transcripts: wc -l on 10 files and ls -l on 5 files at 2026-07-18T13:13:44Z both match exactly; 18 files are byte-identical to their newest ~/.claude/file-history blob. See HANDOFF.md for sources, gaps, and how to rebuild .venv. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
26 KiB
this is a v3 delta spec: it lists what changes on top of the working v2 pipeline (see spec-v2.md). v2 is validated end-to-end (a full chapter renders to one narrated mp4). v3 is about quality of the narration and the look of the video, plus panel curation. structured for hand-off to an agent.
manga recap v3 — quality + look
0. what v2 does today (baseline)
- one narrated
chapter.mp4per chapter: fit-on-black panels, uniform ken-burns zoom, burned subtitles, hard cuts, single default TTS voice. - identity assigns a stable
character_idper body (siglip2), but characters are unnamed, so the script prints "Person A/B". the script worker sees one panel at a time (no memory). - every panel is narrated (no curation).
A. cross-panel identity & continuity [biggest quality lever] [DONE]
SHIPPED:
- descriptive frozen labels: worker_scene emits
labelper character (name -> frozen registry appearance via_describe, "the one with black hair and glasses" -> Person X fallback); worker_script_name_mapconsumes it. no "Person A" when appearance is known, char_id preserved. - chapter_context populated: service.py
_chapter_contextfeeds a cast+entities roster to every script call so the narrator has memory (kills "Unknown voice", item 11). - name acquisition + backfill (item 12): vision emits
nameonly when a char is explicitly named; service.py scene stage binds it viaset_character_name(first name wins, only if unnamed); script stage resolves display names from the FINAL registry (names_by_id) so a name learned in ANY panel backfills to every panel's narration. ceiling: vision name-inference can misattribute; first confident name sticks. upgrade path = roster seed for higher precision. STILL TODO (small, deferred): - true per-panel story-state (only what's been seen so far) for "establishing shot once / don't re-narrate repeats" (the C-todo) — currently whole-manga roster.
--- original plan --- one feature: carry a running story-state through a chapter so narration has memory. covers former items 4, 11, 12.
- stable anchor, mutable label. the
character_idis never lost (already stable via siglip2). only its display label changes:- name known -> the name.
- name unknown -> a descriptive phrase ("guy with black hair and glasses", "girl with short hair holding a phone") derived ONCE from vision's appearance data and frozen per character_id, reused across all panels. never regenerate per-panel (labels would drift).
- name acquisition + backfill. when dialogue addresses a character by name and speaker attribution is confident, bind name -> character_id and switch the label to the name from then on (optionally backfill earlier references).
- story-state threaded per chapter. accumulate {cast seen (id,label,name?), entities/props seen
(the cat)} and feed it into scene+script (populate the currently-empty
chapter_context). this is what lets the narrator say "he calls out to his cat" instead of leaking a bare name (item 12), and removes "Unknown voice" (item 11). - optional roster seed (not required). a per-manga character list (names, maybe ref art) makes naming + cross-panel matching more reliable than dialogue inference. graceful-degrade without it; add when available.
touches: worker_vision (appearance -> label), worker_scene / orchestrator (story-state store + speaker continuity), worker_script (consume context), orchestrator characters table (label, name).
B. narration voice & style (items 2, 3) [DONE]
- less descriptive gemma (item 2): DONE. prompt now demands TIGHT recap (1-2 sentences, ~35 words, "no purple prose"); max_tokens 512->160, temp 0.7->0.6 (worker_script).
- one fixed voice (item 3): DONE. dots.tts samples a RANDOM speaker per call with no reference,
so the voice drifted. worker_tts now bootstraps one seeded reference clip (
~/.cache/manga-tts/ narrator_ref.wav) and clones from it every synth. delete the file to reroll; swap REF_WAV/REF_TEXT for a hand-picked voice. - multi-voice per character: future, not v3 (user: "later") — same mechanism, per-character ref clips.
C. panel curation (item 10) [DONE — basic]
- skip panels with no narrative content (pure scenery / transition art); show an establishing shot once, don't re-narrate repeats. vision flags "scenery-only / skip". big pacing + cost win; do early because it shrinks every downstream stage.
- SHIPPED: vision emits
skip:truefor no-narrative panels (worker_vision prompt+schema); the script stage honors it (service.py run_stage_script) -> no narration -> cascades to no audio, no clip. Single chokepoint, no new db columns (skip rides in the vision result_json blob). - STILL TODO (needs cross-panel memory, folds into A): "establishing shot once, don't re-narrate repeats" — requires dedup of repeated scenery across panels. deferred to story-state.
D. render / look overhaul (items 1, 5, 6, 7, 8)
all in worker_render.py (self-contained ffmpeg builder), except #6.
- #1 dynamic clip duration — DONE.
_audio_durprobes narration length via ffprobe; scene_timing kept only as fallback. no more 4.0s sawtooth / vanishing subs. - #5 blurred-panel background fill — DONE. filter_complex: cover-scaled + boxblur copy behind the fitted panel (replaces black bars).
- #7 subtitles — DONE. boxed caption: bold white on translucent black box (BorderStyle 3), lower-center, DejaVu Sans 52. (user chose boxed over outline/minimal.)
- #8 content-aware motion — DONE. vision picks a camera.effect per panel from the fixed vocab
(+ to=[x,y] for dolly_to_subject); it rides the vision result_json -> scene_graph.camera ->
render. worker_render
_motion(camera, frames)maps the vocab to a zoompan z/x/y expression (static/zoom/pan/dolly/shake/orbit); default zoom_in = old ken burns. orchestrator render stage reads scene_graph.camera. self-checked (ffmpeg renders pan + dolly clips). - #6 multi-panel composites + non-generic transitions [the ambitious one, sequenced last]
- non-generic transitions — DONE. vision emits
transition(cut|crossfade|dissolve| fade_black|fade_white|wipe_left|wipe_right|push) per panel -> scene_graph.transition -> orchestrator assemble collects a per-clip list -> worker_render_xfade_chainbuilds an xfade+acrossfade graph (audio kept in sync, overlap clamped to clip length). all-cut/empty keeps the fast stream-copy concat (no re-encode); any real transition re-encodes the chapter. self-checked (two clips fade_white into one mp4). transition = OUT of that panel. - multi-panel composites — DONE (basic). grouping source = new pass (
grouping.py, deterministic): consecutive SMALL panels (area < 0.7*page-max, same source page, capped at MAX_GROUP=3) share ONE shot. render + assemble both call plan_groups so their grouping matches. a composite = vertical stack of the members on a W×H still; each panel keeps its own script+tts (per-panel stages unchanged); the audios play in sequence while a moving highlight marks the active row ("swap the front panel while audio plays"). clip keyed by the group LEADER; assemble takes one clip per group and the group's exit transition = last member's. solo groups render exactly as before, so blast radius is only genuinely-small adjacent panels. knobs: MAX_GROUP, SMALL_FRAC. ceilings: vertical-stack layout only (no grid/overlay/front-swap animation beyond the highlight), size+adjacency grouping (no narrative-beat grouping), black gutters (no blurred fill in composites). upgrade paths noted in code.
- non-generic transitions — DONE. vision emits
D2. beat pacing, timed cues, subtitle redesign, collage (vikunja 185/187/188/183)
scene-level narration made the beat (not the panel) the render unit; these refine how a beat plays. all in worker_render.py + collage.py, orchestrator run_stage_render.
- 185 content-aware panel timing — DONE. a beat's one narration used to split EQUALLY across its
member panels (
D/n). now the orchestrator scores each member from its scene graph (dialogue amount, action density, camera emphasis:_beat_weights) and passes bounded weights to render; worker_beat_slices(D, n, weights)distributes screen-time by weight with a per-panel floor (no flashing) and a deterministic equal-split fallback (degenerate/short weights). self-checked. - 187 timed subtitle cues — DONE. the whole narration paragraph used to show from frame 1, leaking
later-panel lines early.
cue_plansplits the flowing narration into sentence cues, times them on the SAME [0,D] timeline as the panel slices (length-weighted + readable floor, reused_beat_slices), maps each cue to the panel on screen at its start, and clamps a cue never to precede its panel. per-cue ASS events (not one paragraph). cue plan returned + persisted toclips.cues_jsonfor the review UI / collage. no-TTS-alignment fallback = punctuation/length estimate. self-checked ("Hold up." can't appear before its panel). - 188 subtitle redesign — DONE.
_ass_multinow takes mode (off|minimal|boxed, envSUB_MODE, default minimal) + orientation presets (portrait/landscape safe-areas). smaller type (40/32 vs 52), minimal = thin outline + soft shadow (no box), boxed = restrained ~50% box._wrap2caps every cue at ≤2 short lines (ellipsis truncation) so it never regrows into a paragraph._sub_aligndodges the subject vertically using the director focus point (camera.to.y low -> caption to top). ceiling: no real face/bubble CV yet — camera focus is the only subject signal wired. - 183 animated manga collage — DONE (core + wired behind
COLLAGEflag).collage.py plan_layoutis a pure deterministic planner: ordered panel aspects + RTL + dominant + frame -> template (centered_wide / vertical_pair / stacked_wides / strip_over_dominant / supporting_left_dominant_right / quad), aspect-correct resting rects (no stretch), entrance vectors (RTL slide / dominant scale), z-order, hold/transition timing.worker_render collage_cmdrenders it: blurred plate of the dominant panel + sharp aspect-fit panels with restrained drop shadows, non-dominant panels sliding into place over ~0.45s, crisp holds; reuses 185 slices + 187 cues + 188 subtitles. orchestrator run_stage_render calls it whenCOLLAGE=1(else the ken-burns montage). ceilings (QA-tuned against the reference video, no still to eyeball here): transition-only directional/radial MOTION BLUR not applied (clean slide/scale); per-panel hold DRIFT omitted (static hold); beat-replace "streak" rides the assemble xfade. planner + renderer both self-checked; visual acceptance is the QA run.
D3. identity tracklets + review gates (vikunja 160/136)
- 160 tracklet spine — DONE (code, QA-pending). the Tier-2 gemma decider (identity phase-2) used to
resolve identity once PER CROP, so one person got adjudicated repeatedly and could get two answers on
two panels.
tracklets.py link_tracklets(pure, union-find, self-checked) first LINKS the per-panel siglip shortlists into per-person tracklets over a locality window — hard gender gate + appearance-token overlap + shared-candidate — then the resolver runs gemma ONCE per tracklet against the UNION of the members' shortlists (the person's gallery, consensus-ordered) and applies that one character_id to every crop in the tracklet. fewer gemma calls, one answer per person. ceiling: gemma NONE (cosine over-merged onto an existing char) still isn't split to a fresh id — needs a 3rd siglip re-embed pass; add if the NONE-rate is high. - 136 review gates — DONE (code, QA-pending), off by default (
GATES=1). two human sign-off gates in the pipeline loop: gatescriptpauses BEFORE tts/render, gateaudiovisualpauses BEFORE assemble. When enabled and unapproved the job goesawaiting_reviewand the loop returns;/review/approverecords the gate (gatestable) and relaunches_run_pipeline(skips completed stages → resumes at the gated stage). Gate-1:/review/scriptsshows narration per panel grouped by beat + cheap textual flags (review_flags.py: generic-handle / repeated / empty, self-checked; semantic flags — spoilers, altered meaning, attribution — left to the human eye, ceiling noted). Gate-2:/review/audioserves the per-panel wav so the reviewer actually listens;/review/rettsre-synthesizes ONE panel's TTS + re-renders just its beat clip (delete_clip forces it) and clears the AV gate. web UI (index.html review tab): gate-approve buttons + status, per-panel audio player, flags badge, re-synthesize button. reuses existing /review/panels + /review/script + /review/preview. ceiling: no panel-to-cue drag-edit UI (cues persisted in clips.cues_json, edit endpoint not built); flags are textual only.
E. reference
generic manga-recap youtube style (single narrator, static panels w/ slow pan/zoom, minimal text, dramatic pacing) — but with less generic transitions (feeds #6). #6 goes beyond the reference.
sequencing (proposed)
- C panel curation — shrinks everything downstream first.
- A identity/continuity — most changes how it feels (names + memory).
- B voice+style — small, high impact.
- D basics (#1,#5,#7,#8) — most changes how it looks, low risk.
- D#6 compositing — last; highest risk/effort, isolated to render.
direction schema (decided — enables #8 and #6)
narration (text->TTS) and direction (camera/transition/layers->render) are separate concerns from the SAME scene graph. VISION picks the effect (it sees the panel + has char bboxes for focus points) and emits a direction block into the scene_graph JSON; RENDER just executes it. fixed vocab:
- camera.effect: static | zoom_in | zoom_out | pan_left | pan_right | pan_up | pan_down | dolly_to_subject | orbit | shake | hold (+ from/to [x,y] normalized, easing)
- transition.type: cut | crossfade | fade_black | fade_white | wipe_left | wipe_right | push | dissolve (+ duration)
- layers: {artifact, motion} (parallax, when the depth renderer is wired) not built yet; lands before #8 (needs camera) and #6 (needs transition).
open design points
- roster seed format if/when available.
- B#3: which fixed dots.tts reference voice to pin for the whole video (needs a pick).
v3.1 delta — identity is broken; add human-in-loop review
evidence (last job, 116-panel Teto X Egen, from manga.db)
- 53 character rows for ~8 actual people. "Seonho" stored 5× (ids 85ad36b0/811bf0d7/0033ff7f/ bbddf20c/90269899), Egen Girl 2×. ~40 rows have a single identity_assignment (one-time NPCs).
- root cause: siglip2 over-splits. same person across panels (pose/expression/B&W line art) doesn't clear cosine 0.85, so nearly every appearance spawns a new character_id.
- cascade: names never propagate (panel 10 names id A, panel 51's Seonho is id C ->
names_by_idcan't map it -> narration falls back to "the woman with straight dark hair", flips gender to "her"). re-introductions + DB/S3 bloat all follow from the split. - concrete bug:
worker_identity.pycreate call hardcodes"name": None— vision reads the name but identity discards it at registration.db.create_characteris a blind INSERT, zero dedup.
E1. pipeline decomposition [DECIDED — the vision mega-call splits]
Reason: one gemma4 call doing 4 jobs (detect + dialogue + direction + scene) does each worse, and there was nowhere to run a real cross-panel merge. Split the jobs; insert a merge stage where the model actually compares two crops.
Final stage order (OCR removed — gemma4 transcribes bubble text image-natively; a roster
prepass now runs after crop, see below):
fetch → crop → roster → vision → identity → reconcile → dialogue → direct → scene → script → tts → layers → render → assemble
vision(per panel, image call) — character DETECT only:local_id,bbox,appearance(hair/clothing/features),gender(m|f|unknown),emotion,action. plusskip+scene(location/time). the only stage that must precede identity (identity crops its bboxes).identity(per panel) — siglip assignslocal_id → character_id. on a collision (siglip ambiguous OR hair+clothing appearance overlap OR name collision) it records a merge-candidate pair instead of silently minting a twin. still creates provisional ids; reconcile prunes them.reconcile(NEW, once per chapter, after all panels) — the second vision pass. for each flagged candidate pair,POST worker_vision /vision/samewith the tworef_image_uriscrops -> gemma4 "same character? {yes|no, confidence}". confirmed ->db.merge_characters(loser->keeper): repointidentity_assignments, unionname/aliases/ref_image_uris, delete loser row. loser ref crops are PRESERVED on the keeper as gallery views (not deleted from S3 — deleting them left the keeper's own uris dangling). cast is CLEAN + named from here on. only flagged pairs get a call. lives asrun_stage_reconcilein service.py +/vision/sameon worker_vision (reuses warm gemma4).dialogue(per panel, image call) — gemma reads bubble text straight from the panel image (no external OCR): bubble classify (speech/thought/shout/narration/sfx), merge split bubbles, text cleanup (sentence case, fix names/proper nouns, sfx literal), speaker attribution (speaker=local_id+confidence), entities. prompt gets the CLEAN named cast present in the panel as context -> better attribution ("Seonho (m) vs Haeseon (f)").direct(per panel, off scene understanding) —camera.effect(+to) andtransitionfrom the fixed direction vocab (unchanged schema, section D/direction-schema). runs after identity so it can weight who's present. may glance at the image or work from the scene graph — cheap.
new workers/endpoints: worker_vision becomes detect + /vision/same; new worker_dialogue.py and
worker_direct.py (or /dialogue + /direct endpoints — same warm gemma4, decide at build like
reconcile). db: STAGES gains reconcile,dialogue,direct; run_stage_* for each; session_proxy
call_* wrappers. the old omni-vision prompt/schema is partitioned across the four prompts.
F. identity overhaul [rides on E1]
Goal: 53 rows -> ~12; named recurring cast stays one id; NPCs never persist.
- F1 merge = the
reconcilestage (see E1). the fix for the hardcoded"name": Nonebug: vision/ identity now carryname+gender+appearanceintocreate_characterso provisional rows are named/typed; reconcile does the actual dedup with the model comparing crops (not a blind string match).db.merge_characters+ candidate-flagging in identity. - F2 gender field. add
gender TEXTto characters schema.visionemitsgender. worker_script_name_mapfeeds it to the model ("Seonho (he)") for correct pronouns — kills "her bold personality" flips even when the name resolves. - F3 confirm-before-persist. hold unnamed provisional embeddings in a per-session cache; only write a DB row + S3 crop/npy when an embedding is seen >=2x OR is named. the ~40 one-timers never touch the bucket. session_id already flows into identity. (complements reconcile: reconcile merges the twins that WERE created; F3 stops most from being created at all.)
- F4 centroid + re-rank matching. store multiple refs per char, match against the accumulated centroid not one arbitrary embedding; break siglip ties by name/gender agreement. ceiling: perfect auto-ID on B&W is not achievable — reconcile+F4 get ~80%, the UI (G) closes the rest. don't over-invest.
- DECIDED: start clean. no migration script. wipe manga.db + the minio
mangabucket before the first v3.1 run (destructive ops step, run at build time). fixed logic applies to fresh data.
G. review / edit UI (replaces the auto CoT-strip idea, item C-bleed)
Rationale: the human reviews the final video and knows exactly what's wrong and where. Give them a
synced editor instead of chasing every failure through the pipeline. ~80% of infra already exists in
the orchestrator: static/index.html dashboard, /job/video (streams mp4), /db/tables + /db/table,
/stage/run + /stage/clear (re-render), set_character_name.
- timestamp -> panel is trivial: clips concat in
panel_order, so panel start offset = cumsum of prior panels'audio.duration. store the offset at assemble time (or compute on the fly) as an offset table the page binary-searches on<video>timeupdate. - view:
<video>+ side panel that, for the current panel, shows editable: script text, assigned characters (name / gender / appearance), camera.effect, transition. - edits -> small endpoints (most exist): add
set_gender,merge_characters(ids[]->keep),set_script(panel_id,text),set_camera(panel_id,effect). merge_characters is where the Seonho mess gets fixed in seconds — and the merge is ground truth that feeds F. - re-render: "re-render from stage X" button = existing
/stage/clear+/stage/runon the affected panels/clips. effect preview = dropdown of_motionvocab +XFADEmap, re-render one clip. - lazy ceiling: no DAW timeline, no non-destructive edit graph. "video + synced editable form + re-render button" over existing data/endpoints. ~1 day. highest-leverage addition: turns every future quality bug into a 30-second manual fix.
v3.1 sequencing
- E1 decomposition — split vision into
vision/dialogue/direct, addreconcile+direct+dialogueto STAGES with stage runners. Foundation; F1 lives inside it.- 1a.
vision-> detect-only (+gender, F2). 1b.dialoguestage. 1c.directstage. - 1d.
identitycandidate-flagging +db.merge_characters. 1e.reconcilestage +/vision/same.
- 1a.
- F3 (confirm-before-persist) — kills NPC bloat.
- G (review UI) — builds on the cleaner DB; closes the residual reconcile/F4 can't.
- F4 (centroid/re-rank) — only if G shows auto-ID still too noisy to be worth hand-fixing.
Build order note: each of 1a–1e is independently deployable/testable (workpc workers reload via tmux; orchestrator needs a container rebuild). Do a full clean re-run after 1e before starting G.
D#6 revisited — camera-traversal composite [DONE 2026-07-14]
the vstack composite was the wrong look; target = camera moving across panels in their REAL page
positions (motion-comic page traversal, partial-neighbor reveals). geometry already exists: panels
store bbox + page_index + panel_order, source page is retained. composite feeds a constant 9:16
crop-window that HOLDS on each member then quick-eases (0.5s) to the next (traverse_cmd,
_pan_expr in worker_render.py). page uri is reconstructed from page_key(...) (deterministic — no
schema change). vstack kept as the fallback when bbox/page_uri are missing.
orchestrator TODO: run_stage_render composite branch must pass each member's bbox + one
page_uri into call_render_composite.
E2 — chapter roster / context pre-pass [SPEC — not built]
problem. every stage sees one panel at a time, so nobody has a chapter-level view. symptoms in
the wild: the MC is "the man in the yellow shirt" at panel 7 but named at 10 (name learned too late
to backfill cleanly); unnamed extras get re-described every panel; bubble-only panels (p005) have no
cast to attribute a line to; the narrator's tone/naming drifts. _chapter_context today is just the
DB registry accumulated so far — partial and late.
idea. one cheap gemma4 pass up front that reads the pages (text image-natively) and produces a chapter view that grounds the per-panel stages. it does NOT replace vision/identity — it's coarse context only, threaded back as NAME hints (never identity evidence).
stage. new roster stage, placed after crop, before vision. names surface in captions/
dialogue on the page images themselves (e.g. "Lim Seonho, Designer at Everyday"). one artifact per
chapter, stored on chapters.roster_json, resumable (skip if present), non-blocking on failure.
method (implemented). one multi-image /roster call over an even spread of ~6 chapter pages
(ROSTER_SAMPLE_PAGES) -> {premise, characters:[{name,aliases,gender,species,description}]}.
threaded into: vision detect (known_characters hints, registry wins), dialogue (off-panel cast for
bubble-only panels), script (_chapter_context premise). ceiling: sampled pages; widen the sample if
recall bites.
output (stored per chapter, e.g. a chapter_roster blob):
{"premise": "3-5 sentence chapter setup, present tense, no spoilers ahead",
"roster": [{"name":"Lim Seonho","aliases":["Seonho"],"gender":"m",
"appearance":"short brown hair, glasses, yellow shirt",
"role":"designer at Everyday; the MC"}]}
consumers (all already have the hook, just feed them the roster instead of the thin registry):
vision(detect): pass roster asknown_charactersso detections carry a provisional name.identity/reconcile: seed name labels — a face gets its roster name on FIRST sighting, killing "named at 10, described at 7". embeddings stay authoritative for who-is-who; roster only supplies the label when appearance matches.dialogue: roster = the known cast for off-panel attribution (the p005 case now has names/handles to point at instead of "".)script: seedbrieffrompremiseat panel 1 (no cold start); roster names ->names_by_id+introducedfrom the start (stable handles, no re-description).
caveats. roster names are a COARSE read → HINTS, not ground truth: apply a roster name only when a detected/identified character's appearance matches a roster entry; never override an embedding match. one extra pass per chapter (latency/cost) — acceptable, it's once per chapter, warm session.