EPUB to Markdown: Getting a Whole Book Into Your LLM
Say you want to ask an AI questions about an entire book — a technical manual, a reference text, a public-domain classic you're studying. You've got the EPUB. But you can't just hand a model an .epub file; it wants text, and an EPUB is a package, not a plain document. The good news is that of all the formats, ebooks are among the friendliest to convert — and the interesting problem isn't the conversion at all. It's the size.
Here's how EPUB to Markdown works, why it's cleaner than you'd expect, and what to do about the fact that a book is enormous by LLM standards.
Why EPUBs convert cleanly
Peek inside an .epub and you'll find something reassuring: it's a zip of HTML files, one or several per chapter, plus a manifest describing the reading order. That matters because HTML is already structured text — headings, paragraphs, lists, emphasis, all explicitly marked up. So converting an EPUB to Markdown is mostly a matter of walking the chapters in order and mapping each one's HTML to Markdown, the same near-lossless translation that makes web pages easy.
The result is a single, clean Markdown document: chapter headings intact, paragraphs flowing, the book's structure preserved from cover to back. MarkPrep's EPUB converter reads the package, follows the manifest's reading order, and stitches the chapters into one Markdown file.
The real challenge: a book is huge
This is where ebooks get interesting. A full-length book can run 80,000–150,000 words — which is well over 100,000 tokens. Even with today's large context windows, dropping an entire book into a single prompt is often impractical, expensive, or flatly over the limit. So the moment you've converted it, the question stops being "did it convert" and becomes "how do I feed something this big to a model?"
The answer is almost always retrieval, not brute force. Rather than paste the whole book, you:
- Split it into sensible chunks — ideally on chapter and section boundaries, which the clean heading structure you just preserved makes easy.
- Embed those chunks into a vector store.
- Retrieve only the relevant passages when you ask a question.
This is exactly the RAG pattern, and it's why the chunk plan matters so much for books: splitting on headings keeps each chunk a coherent passage instead of slicing mid-argument. A book is the format where careless fixed-size chunking hurts most, because the whole point is asking questions that span sections.
Formatting the conversion should respect
A couple of book-specific things worth checking in the output:
- Chapter boundaries should map to headings, so the structure is chunk-able.
- Footnotes and endnotes should survive in some readable form rather than vanishing.
- Front matter (title page, table of contents) is usually noise for AI purposes and fine to trim.
A note on where the file is — and what it is
Two honest caveats. First, respect copyright: converting a book you own for your own use is one thing; redistributing an author's work is another — the tool doesn't change what you're allowed to do with the content. Second, as with every format here, the conversion runs locally in your browser, so the file never uploads. For DRM-protected commercial ebooks, note that DRM is a separate lock the converter won't (and shouldn't) break.
In summary
Ebooks convert to Markdown more cleanly than almost anything, because underneath they're already structured HTML. The work isn't the conversion — it's respecting the size: convert the EPUB, keep the chapter structure, and then chunk and retrieve rather than trying to cram an entire book into one prompt.
Try it on your own file
Convert a document and watch the token counter — free, no account, nothing uploaded.