- Python 100%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
| pvl_spotting | ||
| samples | ||
| .gitignore | ||
| config.example.toml | ||
| pyproject.toml | ||
| README.md | ||
| server.py | ||
| uv.lock | ||
Manga text spotting — PaddleOCR-VL-1.6 (+ Qwen3-VL SFX pass)
PaddleOCR-VL-1.6 detects and transcribes all main text on a manga/comic page as quad polygons. Furigana is excluded. Stylized sound-effect lettering, which PaddleOCR-VL cannot recognize, is captured by a second pass through Qwen3-VL-30B-A3B on llama.cpp (Vulkan, experts in RAM) and merged into the results in orange.
Run
The main entry point is the API server: it loads every model once at startup and keeps them resident, so pages are processed without per-call load time.
uv run server.py # host/port from config.toml; first start downloads models
GET /health—{"status": "ok"}once the models are loaded.POST /spot— multipartfile(a page image) plus optional query params:sfx(bool) — detect stylized sound effects via Qwen3-VL.direction(vertical|horizontal) — overrides config; vertical manga columns vs horizontal Western-comic lines (affects paragraph merging).order(rtl|ltr) — column reading order for vertical text.overlay(bool) — addsoverlay_png(base64) with the regions drawn.
- Returns
{"image_size": [w, h], "regions": [...]}—type: text | sfx, polygons/boxes, transcriptions and per-line polygons (same format the oldoverlay.jsonused). - Interactive docs:
http://<host>:<port>/docs.
config.toml controls the server host/port, device, model repos/paths
(including which models are auto-downloaded on first start), and the
algorithm thresholds (bubble conf, merge pad, upscale threshold, ...).
A CLI is available for debugging:
uv run python -m pvl_spotting --help
First start downloads PaddleOCR-VL-1.6 (~2 GB), SAM 2.1 (~1 GB), the
speech-bubble YOLO weights and (unless download_ggufs = false) the ~18 GB
SFX GGUFs into the HuggingFace cache. Local copies under models/ (and the
optional paddle_dir / sam_dir snapshots) are preferred when present. The SFX pass auto-starts the Qwen3-VL llama-server (port
8792, ~2 min load; keep it running between pages — subsequent SFX passes are
~80 s, the spotting pass ~30–140 s).
Files
server.py— FastAPI server (/health,/spot), config-driven.config.example.toml— tracked configuration template; copy it toconfig.toml(untracked) and adjust: ports, model repos/local paths, algorithm thresholds.pvl_spotting/— the pipeline package:pipeline.py— per-page orchestration (+create_pipeline).spotting.py— PaddleOCR-VL spotting pass + quad parsing.segmentation.py— YOLO + SAM 2 bubble masks.postprocess.py— furigana filter, paragraph merging, bubble grouping.sfx.py— Qwen3-VL SFX pass (llama-server management, clustering, overlap arbitration).download.py— first-start model download/local-file checks.models.py— model loading (incl. the rope config patch).geometry.py,overlay.py,config.py— helpers.__main__.py— single-page CLI.
llamacpp/— llama.cpp Vulkan build (llama-server for the SFX pass).models/— local model copies (bubble_seg YOLO weights, Qwen3-VL-30B GGUFs); used when present, otherwise fetched from the hub into the HF cache.samples/— test pages (test_jp, test_en, test_sfx).
Requirements
- AMD GPU with ROCm 7.2.1 support (RX 9070 XT tested) + 26.2.2 graphics driver.
- Python 3.12 (pinned; the ROCm torch wheels are cp312 only).
uv syncresolves everything, including the ROCm SDK wheels from repo.radeon.com (see[tool.uv.sources]in pyproject.toml).
Pipeline notes
- Config patch: transformers 5.x has native
paddleocr_vlsupport but the HF repo's config.json still uses pre-v5 rope fields — the script injectsrope_parameters(text: default/500000/mrope [16,24,24]; vision: axial/10000) before loading, otherwise init crashes. - Furigana filter: two layers — the prompt asks for main lines only, and a geometric filter drops regions that are thin AND adjacent to a ≥1.4× thicker same-orientation region.
- Paragraph merging: per the user-supplied
--direction, adjacent lines/columns merge into one region when the perpendicular gap is ≤ 0.75× the line thickness and they overlap ≥ 30% along the text flow. Merged regions get a straight bounding box, text in reading order (--orderfor vertical columns), and keep the original line polygons inlines. - SFX pass (Qwen3-VL via llama-server): two prompts (SFX-only and general spotting), union minus regions covered by the main pass, then proximity clustering (SFX are often single characters scattered diagonally, e.g. one big あははははは). SFX takes over overlapping main-text regions and inherits their transcription; regions with no text get a crop-OCR fallback.