No description
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
2026-09-15 13:59:35 +02:00
pvl_spotting chore: split into multiple files and add config file 2026-09-15 13:59:35 +02:00
samples init 2026-09-14 17:54:34 +02:00
.gitignore chore: split into multiple files and add config file 2026-09-15 13:59:35 +02:00
config.example.toml chore: split into multiple files and add config file 2026-09-15 13:59:35 +02:00
pyproject.toml chore: split into multiple files and add config file 2026-09-15 13:59:35 +02:00
README.md chore: split into multiple files and add config file 2026-09-15 13:59:35 +02:00
server.py chore: split into multiple files and add config file 2026-09-15 13:59:35 +02:00
uv.lock chore: split into multiple files and add config file 2026-09-15 13:59:35 +02:00

Manga text spotting — PaddleOCR-VL-1.6 (+ Qwen3-VL SFX pass)

PaddleOCR-VL-1.6 detects and transcribes all main text on a manga/comic page as quad polygons. Furigana is excluded. Stylized sound-effect lettering, which PaddleOCR-VL cannot recognize, is captured by a second pass through Qwen3-VL-30B-A3B on llama.cpp (Vulkan, experts in RAM) and merged into the results in orange.

Run

The main entry point is the API server: it loads every model once at startup and keeps them resident, so pages are processed without per-call load time.

uv run server.py       # host/port from config.toml; first start downloads models
  • GET /health — {"status": "ok"} once the models are loaded.
  • POST /spot — multipart file (a page image) plus optional query params:
    • sfx (bool) — detect stylized sound effects via Qwen3-VL.
    • direction (vertical|horizontal) — overrides config; vertical manga columns vs horizontal Western-comic lines (affects paragraph merging).
    • order (rtl|ltr) — column reading order for vertical text.
    • overlay (bool) — adds overlay_png (base64) with the regions drawn.
  • Returns {"image_size": [w, h], "regions": [...]} — type: text | sfx, polygons/boxes, transcriptions and per-line polygons (same format the old overlay.json used).
  • Interactive docs: http://<host>:<port>/docs.

config.toml controls the server host/port, device, model repos/paths (including which models are auto-downloaded on first start), and the algorithm thresholds (bubble conf, merge pad, upscale threshold, ...).

A CLI is available for debugging:

uv run python -m pvl_spotting --help

First start downloads PaddleOCR-VL-1.6 (~2 GB), SAM 2.1 (~1 GB), the speech-bubble YOLO weights and (unless download_ggufs = false) the ~18 GB SFX GGUFs into the HuggingFace cache. Local copies under models/ (and the optional paddle_dir / sam_dir snapshots) are preferred when present. The SFX pass auto-starts the Qwen3-VL llama-server (port 8792, ~2 min load; keep it running between pages — subsequent SFX passes are ~80 s, the spotting pass ~30–140 s).

Files

  • server.py — FastAPI server (/health, /spot), config-driven.
  • config.example.toml — tracked configuration template; copy it to config.toml (untracked) and adjust: ports, model repos/local paths, algorithm thresholds.
  • pvl_spotting/ — the pipeline package:
    • pipeline.py — per-page orchestration (+ create_pipeline).
    • spotting.py — PaddleOCR-VL spotting pass + quad parsing.
    • segmentation.py — YOLO + SAM 2 bubble masks.
    • postprocess.py — furigana filter, paragraph merging, bubble grouping.
    • sfx.py — Qwen3-VL SFX pass (llama-server management, clustering, overlap arbitration).
    • download.py — first-start model download/local-file checks.
    • models.py — model loading (incl. the rope config patch).
    • geometry.py, overlay.py, config.py — helpers.
    • __main__.py — single-page CLI.
  • llamacpp/ — llama.cpp Vulkan build (llama-server for the SFX pass).
  • models/ — local model copies (bubble_seg YOLO weights, Qwen3-VL-30B GGUFs); used when present, otherwise fetched from the hub into the HF cache.
  • samples/ — test pages (test_jp, test_en, test_sfx).

Requirements

  • AMD GPU with ROCm 7.2.1 support (RX 9070 XT tested) + 26.2.2 graphics driver.
  • Python 3.12 (pinned; the ROCm torch wheels are cp312 only).
  • uv sync resolves everything, including the ROCm SDK wheels from repo.radeon.com (see [tool.uv.sources] in pyproject.toml).

Pipeline notes

  • Config patch: transformers 5.x has native paddleocr_vl support but the HF repo's config.json still uses pre-v5 rope fields — the script injects rope_parameters (text: default/500000/mrope [16,24,24]; vision: axial/10000) before loading, otherwise init crashes.
  • Furigana filter: two layers — the prompt asks for main lines only, and a geometric filter drops regions that are thin AND adjacent to a ≥1.4× thicker same-orientation region.
  • Paragraph merging: per the user-supplied --direction, adjacent lines/columns merge into one region when the perpendicular gap is ≤ 0.75× the line thickness and they overlap ≥ 30% along the text flow. Merged regions get a straight bounding box, text in reading order (--order for vertical columns), and keep the original line polygons in lines.
  • SFX pass (Qwen3-VL via llama-server): two prompts (SFX-only and general spotting), union minus regions covered by the main pass, then proximity clustering (SFX are often single characters scattered diagonally, e.g. one big あははははは). SFX takes over overlapping main-text regions and inherits their transcription; regions with no text get a crop-OCR fallback.