Files
manga-recap-pipeline/manga-recap-pipeline-spec.md
kami ff6a512630 Reconstruct repo from Claude Code + codex transcripts
Working tree (including .git) was lost to an rm. Rebuilt by replaying Write/Edit/
Read/attachment events from 25 Claude sessions and 22 successful codex apply_patch
blocks into one timestamp-ordered timeline.

Verified against ground truth recorded in the transcripts: wc -l on 10 files and
ls -l on 5 files at 2026-07-18T13:13:44Z both match exactly; 18 files are
byte-identical to their newest ~/.claude/file-history blob.

See HANDOFF.md for sources, gaps, and how to rebuild .venv.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 02:42:41 +04:00

11 KiB
Raw Permalink Blame History

Manga/Manhwa Recap Pipeline — Technical Spec

Status: draft v1 Target: fully local, single-GPU, resumable batch pipeline. No publishing (copyright out of scope). Core constraint: no new PyTorch / ROCm wheels. Existing llama.cpp (gemma) stays. Everything else runs on onnxruntime + opencv + numpy + scikit-learn.


1. Goal

Given raw manga pages or manhwa strips for a series, automatically produce a narrated recap video with burned-in subtitles, with minimal human involvement (one-time character-bank setup per series).

This spec replaces a dependency on the magi/magiv2 model by decomposing its subtasks into swappable, low-dependency components.


2. Design principles

  • Dependency isolation. gemma runs on llama.cpp (unchanged). All CV/ML runs via ONNX Runtime. Detection models are small enough to run on CPU EP; ROCm EP optional. No transformers, no second torch install.
  • Specialists over generalist. Do not ask one VLM to detect + identify + attribute. Each subtask is a dumb specialized component; a clean data bus (JSON files) connects them.
  • Resumable. Every stage reads its input artifact from disk and writes its output artifact to disk. A stage is skipped if its output exists and is newer than its input (unless --force).
  • Supervised shortcut for identity. Replace magi's unsupervised character clustering with a per-series character bank (exemplar crops + names). Simpler, more robust, and how magiv2 gets names anyway.
  • Fail loud, fail per-panel. Low-confidence panels are flagged, not silently guessed. Confidence thresholds are config, not magic numbers in code.

3. Architecture overview

ingest → [1] panels → [2] detect → [3] identity → [4] text
                                                      │
                          [5] speaker-bind ←──────────┤
                                                      │
                          [6] filter essential ←──────┘
                                    │
                          [7] transcript (ordered)
                                    │
              ┌─────────────────────┴───────────────────┐
        [8] scene-action (gemma)                   (transcript.json)
              └─────────────────────┬───────────────────┘
                                    │
                          [9] script gen (LLM, chapter + rolling summary)
                                    │
        [10] TTS (dots.tts) ──► [11] visual assembly (ffmpeg) ──► [12] subs (faster-whisper) ──► [13] mux

Stages 27 are the magi replacement. Stages 813 are the existing recap backend.


4. Filesystem layout

work/
  <series>/
    source/            # input pages/strips
    ch<NN>/
      panels/          # panel crops, ordered
      panels.json      # [1]
      detections.json  # [2]
      identities.json  # [3]
      texts.json       # [4]
      transcript.json  # [5][6][7] merged
      script.json      # [9]
      audio/           # [10] wav per segment
      video.mp4        # [13] final
    bank/
      bank.json        # character bank (per series, hand-built once)
      crops/           # exemplar images

Artifact-per-stage = resumability. Delete an artifact to re-run that stage forward.


5. Stage specs

[1] Panel extraction

  • Manga (page-based): kumiko → ordered panel polygons. Pure OpenCV.
  • Manhwa (vertical strip): slice at horizontal whitespace bands. OpenCV: row-wise background uniformity → cut points. Produces pseudo-panels.
  • In: source/* · Out: panels/, panels.json
  • Dep: kumiko, opencv. No torch.
// panels.json
{
  "type": "manga",            // or "manhwa"
  "reading_order": "rtl",     // rtl | ltr | ttb
  "panels": [
    { "id": "p001", "page": 1, "bbox": [x,y,w,h], "file": "panels/p001.png", "order": 0 }
  ]
}

[2] Detection (text + balloons + characters)

  • Text + balloons: comic-text-detector (ships ONNX). Returns text regions + balloon masks. This is what kills the OCR problem — you never OCR a full page again.
  • Character boxes: YOLOv8 anime face/person model exported to ONNX. CPU-fine.
  • In: panels/ · Out: detections.json
  • Dep: onnxruntime, opencv, numpy.
// detections.json  (per panel)
{
  "p001": {
    "balloons": [
      { "id": "b0", "mask_poly": [[x,y],...], "centroid": [x,y], "bbox": [x,y,w,h] }
    ],
    "text_regions": [
      { "id": "t0", "bbox": [x,y,w,h], "in_balloon": "b0" }   // null if floating (sfx/sign)
    ],
    "chars": [
      { "id": "c0", "bbox": [x,y,w,h], "crop": "..." }
    ]
  }
}

[3] Character identity (bank match)

  • Embed every chars[*] crop with SigLIP or an anime ArcFace model (ONNX).
  • Cosine-match against bank.json entries. Above threshold → assign name; else "unknown".
  • Bank is built once per series by hand (10 min): a few exemplar crops + a name each. This is the only required human touch.
  • In: detections.json, bank/ · Out: identities.json
  • Dep: onnxruntime, numpy, sklearn (cosine / nearest-neighbour).
  • Note: this is the accuracy-critical stage. Generic embeddings + same-face syndrome = the weakest link vs magi. The bank is what rescues it. Do NOT attempt unsupervised clustering as primary — it merges lookalikes.
// bank.json
{ "characters": [
    { "name": "Aria", "exemplars": ["bank/crops/aria_0.png","bank/crops/aria_1.png"] }
]}
// identities.json  (per panel)
{ "p001": { "c0": { "name": "Aria", "score": 0.82 }, "c1": { "name": "unknown", "score": 0.41 } } }

[4] Text extraction

  • For each in_balloon text region: crop the balloon, feed the clean crop to gemma (llama.cpp) → read text. Isolated crops read far better than full pages.
  • In: detections.json, panels/ · Out: texts.json
  • Dep: llama.cpp (existing). No torch.
// texts.json
{ "p001": { "t0": "We can't stay here." } }

[5] Speaker binding

Two paths — start geometry, fall back to gemma on low confidence.

  • A — tail geometry (default, pure CV): from the balloon mask, find the tail = sharpest protrusion off the centroid. Vector centroid→tip; nearest chars[*] box along that ray = speaker. Hand-rolled magiv2 tail logic.
  • B — set-of-mark + gemma (fallback): draw numbered boxes on chars + balloons on the panel image, ask gemma "balloon 2 → face #3 or #5?". Grounding via drawn marks beats free-form spatial reasoning. No new deps.
  • Trigger fallback when: >2 candidate chars, ambiguous/absent tail, or geometry confidence < threshold.
  • In: detections.json, identities.json, texts.json · Out: merged into transcript.json.

[6] Essential vs non-essential filter

  • Pure geometry: text_region.in_balloon != null → dialogue. Floating on raw art → sfx/sign → drop. Kills "THUD" and street-sign garbage with no classifier.

[7] Transcript assembly (reading order)

  • Panel order from panels.json. Within panel: manga = sort (top → right-to-left); manhwa = top-down.
  • Emit ordered speaker+line list.
// transcript.json
{ "chapter": 12, "lines": [
    { "panel": "p001", "speaker": "Aria", "line": "We can't stay here.", "conf": 0.79, "flagged": false }
]}

[8] Scene-action description

  • gemma describes physical action per panel ("she draws a sword") — the one thing magi does NOT do and gemma is good at. Runs in parallel with 27.
  • Out: scene field per panel, merged for the script stage.

[9] Script generation

  • LLM input = one chapter of transcript + scene-actions + a ~500-token rolling summary of prior chapters. Do NOT token-max the context; attention degrades mid-window and coherence/attribution drop.
  • Prompt: compress and narrate, not transcribe. Output narration segments each mapped to source panel IDs (needed for visual timing).
  • Out: script.json
// script.json
{ "segments": [
    { "id": "s0", "text": "Cornered in the ruins, Aria makes her choice...", "panels": ["p001","p002"] }
]}

[10] TTS

  • dots.tts, wav per segment. Keep durations — they drive visual timing.

[11] Visual assembly

  • ffmpeg ken-burns (pan/zoom) per panel. Panel display time = proportional to its segment's audio length. moviepy to orchestrate or raw filtergraphs for lean/fast.

[12] Subtitles

  • faster-whisper on generated audio → timestamped SRT → burn in with ffmpeg. Easier than aligning from the script side.

[13] Mux

  • Combine video + audio + burned subs → video.mp4.

6. Dependency matrix

Stage Tooling New torch?
1 panels kumiko, opencv no
2 detect comic-text-detector (onnx), yolo-anime (onnx) no
3 identity siglip/arcface (onnx), sklearn no
4 text llama.cpp gemma no
5 bind numpy geometry + gemma fallback no
6 filter numpy geometry no
7 order numpy no
8 scene llama.cpp gemma no
9 script local LLM no
10 tts dots.tts (its own env)
11 video ffmpeg, moviepy no
12 subs faster-whisper (ctranslate2, not torch)
13 mux ffmpeg no

Net: the magi replacement (27) adds only ONNX Runtime + a couple of onnx model files. No ROCm wheel churn.


7. Orchestration

  • Bash driver calls python stage scripts; artifacts handed off as files.
  • Each stage: stage_N.py --series X --chapter NN [--force].
  • Skip logic: if output exists and mtime(output) > mtime(input) and not --force, skip.
  • Model loading: sequence stages so you never hold two large models in VRAM at once (gemma vs detectors vs faster-whisper). One card handles all, serially.
  • Do NOT wrap in a Spring service. This is a batch job, not request/response — a service is pure added attack surface and state for zero benefit.

8. Hardware notes

  • gemma (mmproj) is the VRAM heavyweight; detectors are CPU-viable.
  • Sequence model loads; never co-resident. Target: single decent GPU, staged.
  • Detection/embedding on CPU EP is fine and frees VRAM for gemma.

9. Known risks / caveats

  • Re-ID accuracy is the weak link vs magi's end-to-end association. Mitigation: character bank (mandatory), tune cosine threshold per series, flag unknown rather than guess.
  • Manhwa layout: comic-text-detector and the anime detectors are manga/anime-trained. Vertical webtoon art + non-japanese layout = degraded results. Expect tuning on the whitespace slicer and lower binding confidence.
  • Tail geometry fails on off-panel speakers and thought bubbles. Fallback to gemma SoM; if still ambiguous, flag the line.
  • Rolling summary drift: long series accumulate summary error. Periodically re-anchor the summary from a canonical synopsis if available.

10. Pre-build check (do this first)

Before building any of this: magi's HF weights may load through your existing ROCm torch directly, ignoring its pinned requirements.txt (the version pain is usually the wrapper deps, not torch). 5-minute test. If it instantiates and runs, you skip this entire rebuild. If it OOMs or the arch won't load, decompose per this spec.