A Field Guide to Document Formats for LLMs (What Converts Cleanly, What Fights You)
Every knowledge base is a junk drawer of formats. PDFs from finance, Word docs from legal, spreadsheets from ops, a few scanned contracts, some scraped web pages, the odd PowerPoint nobody will admit to making. When you decide to make all of it usable by an AI, you quickly learn that the format matters more than you expected — some files become clean Markdown almost for free, and some fight you the whole way.
This is a field guide to that spread: which document formats convert cleanly for LLMs, which need extra care, and why the destination is almost always Markdown regardless of where you start.
Why Markdown is the destination, not the source
First, the target. You convert to Markdown because it's the format LLMs handle best: it encodes structure (headings, lists, tables) in plain text, with none of the visual baggage that formats like PDF and DOCX carry. It's compact, unambiguous, and it chunks cleanly for retrieval. So the real question is never "what format does the model want" — it's Markdown — but "how much of a fight is my format to get there."
Here's the terrain, roughly from easiest to hardest.
The easy ones: already structured text
HTML, Markdown, plain text, CSV, JSON. These are text with (mostly) explicit structure, so conversion is close to lossless. CSV becomes a clean table trivially. HTML maps heading-for-heading — the only work is stripping the navigation, ads, and scripts that surround the actual content. These formats rarely surprise you.
The middle: office documents
Word (.docx), Excel (.xlsx), PowerPoint (.pptx). These carry real structure underneath a lot of decoration. The conversion is reliable but the quality depends on the converter: does it map Word's heading styles to Markdown heading levels, reconstruct Excel's grid as a proper table, pull the text and notes out of PowerPoint slides? Done well, these are clean. Done lazily, you get bold-everything and collapsed tables. Format-specific converters exist for exactly this reason — a dedicated Word-to-Markdown or Excel-to-Markdown path handles each format's quirks properly.
The hard part: PDFs
Digital PDFs are fine — they have a text layer, so they behave like the easy group. Scanned PDFs are the genuinely hard case: they're images, with no text to extract, so they need OCR to reconstruct the words from pixels before any Markdown is possible. Accuracy there depends on scan quality, and the output needs spot-checking. PDF is the one format where "it's a PDF" tells you almost nothing — a clean digital PDF and a phone photo of a page are worlds apart in effort.
The wildcards: images and the web
Images (screenshots, photos of documents) are pure OCR — same story as scans. Web pages by URL are HTML underneath, but wrapped in so much cruft that the real skill is isolating the article from the page furniture. Both are doable; both need a converter that knows what to throw away.
The rule that cuts across all of it
Whatever you start with, two things stay true once you reach Markdown:
- Check the size in tokens, not pages. Format affects density — a table-heavy spreadsheet and a prose report of the "same length" have wildly different token counts. Pages lie; tokens don't.
- Keep sensitive files local. The more confidential the document, the more it matters where the conversion runs. A converter that works in your browser never uploads any of these formats, which is the only version of this that's safe for contracts, records, and internal data.
The takeaway
You don't need a different strategy per format so much as an honest map of the terrain: text-like formats are nearly free, office documents are reliable with a good converter, digital PDFs are easy and scans are work, images and web pages need cruft-removal. Aim everything at clean Markdown, measure the result in tokens, and — if the files matter — do it all without them leaving your machine.
Try it on your own file
Convert a document and watch the token counter — free, no account, nothing uploaded.