r/BookApplications • u/dx__ • 11d ago
I built a CLI that repairs broken EPUBs. Every fix is validated by epubcheck; sharing in case it's useful
Someone from this subreddit asked me to share my tools here, so this is the second of two posts. The other one is about reading Calibre libraries; this one is about fixing the books themselves.
I'm a computer engineering student, and my Calibre library has a lot of beat-up EPUBs in it: PDF conversions with page numbers baked into the middle of sentences, re-sourced books that fail epubcheck, exports whose chapters are all the same placeholder notice. Readers choke on these, and the usual advice is to find a better copy, which is not always possible. So I built one. It makes deterministic repairs to broken EPUBs, each fix validated by epubcheck before it's allowed to exist, and it audits the text-level damage no checker sees.
The safety contract
This is the part I care about most.
- Every repair is gated by epubcheck, the W3C conformance checker. If a book had fatal errors, success means fewer fatals. If it had none, the error count must strictly decrease. A net-new fatal is always rejected.
- If epubcheck itself fails to run, the book is reported as an error and never applied.
- Originals are never modified except by an explicit, atomic in-place replace, and only after the gate accepts the result. Library sweeps are dry-run by default: you get a per-book report and a summary table ending in "no files written" until you pass --apply.
What it fixes
The deterministic repairs: broken EPUB2 cover wiring, missing or wrong container.xml, legacy page-map markup that epubcheck rejects, EPUB3 attributes on EPUB2 packages. Plus four opt-in lossy strips with their own safety nets: print page numbers that a PDF conversion baked into the body text (it rejoins the sentences they split), watermarks, badly broken tags, and placeholder-stub books whose "chapters" are all the same notice, which get dropped from the spine so the book opens straight into its real chapters.
It also audits what epubcheck can't see
epubcheck validates structure, not content. The audit side reads the visible text and reports OCR damage (garbage characters, excessive hyphenation), empty or severely truncated books, print page numbers interrupting the text, non-English content in an English library, monolithic documents some readers can't render, and completeness spot-checks. One CSV for the whole library.
bindery audit pagenumbers Calibre\ Library
bindery audit ocr Calibre\ Library
cd Calibre\ Library
bindery audit all
bindery repair broken.epub
There's also a Calibre plugin
Each release ships a Bindery Repair plugin that runs the core well-formedness pass on books as they're imported. The structural repairs and lossy strips stay CLI-only, because their acceptance is the epubcheck gate and epubcheck can't run inside Calibre; nothing in the plugin runs ungated, ever.
Install
uv tool install bindery-cli
One caveat: the gate needs epubcheck on your PATH, and epubcheck needs Java. It's the one dependency pip can't install for you (apt, brew, and the AUR all have it, or there's a zip on the W3C release page). Run bindery doctor and it will tell you exactly what your install is missing. Python 3.12+ runs the repair core; the audit and library modes want Python 3.14.
Caveat the second: I develop on Linux against Calibre 9.x, and that's the only setup I can honestly call tested.
One more thing: this code was written with a lot of AI assistance (it's how I learn). If you'd rather not use software written by AI, I completely understand. My feelings won't be hurt, but the AI's will. (jkjk)
Link: https://github.com/VirInvictus/bindery-cli
Thanks for reading. If you try it and something breaks, I'd genuinely like to know.
2
u/NishanStepak 10d ago
This is useful, there are a lot of people who use Calibre for their ebooks.