Scanned PDF to Markdown: Getting Text Out of an Image
Here's a quick way to tell what kind of PDF you're holding: try to select a sentence with your cursor. If the text highlights, it's a digital PDF with a real text layer, and converting it is easy. If your cursor just draws a box over what looks like text — nothing highlights — you've got a scan: a stack of images that happen to look like a document. There's no text in there to copy, because to the computer it's a photograph.
That's why scanned PDFs defeat ordinary converters. You can't extract text that isn't there. To turn a scanned PDF into Markdown, something first has to read the pixels and work out what letters they form — that's OCR, optical character recognition — and only then can the result become clean Markdown. Here's how that works, where it's imperfect, and how to do it without shipping your scan to a server.
OCR is the step everything hinges on
For a digital PDF, conversion is basically transcription — the text is already there. For a scan, OCR has to reconstruct the text from an image: separating characters from background, coping with skew, smudges, coffee stains, and the slight blur of a phone photo of a page. It's genuinely hard work, and it's why scanned conversion is slower and less perfect than digital conversion.
Modern OCR is good — on a clean, straight, high-resolution scan it's remarkably accurate. But it degrades honestly with input quality: a crisp 300-DPI scan reads beautifully; a dim photo of a crumpled receipt does not. No tool escapes this, and any that claims perfect accuracy on bad inputs is lying to you.
What to check in the output
Because OCR can misread, you don't trust a scanned conversion blindly — you spot-check it:
- Numbers and names first. A misread
8for a3in a contract or invoice matters far more than a typo in a paragraph. Scan the figures. - Tables. Columns in a scan are the hardest thing to reconstruct; check they didn't collapse.
- Low-confidence pages. A good converter flags the pages it wasn't sure about, so you know where to look instead of re-reading everything.
MarkPrep's converter runs OCR on scanned pages automatically and surfaces where the confidence was low — so you catch the doubtful bits before they flow downstream into a model or a search index.
The privacy catch most people miss
Scans are often the most sensitive documents there are — signed contracts, medical forms, ID pages, handwritten notes. And OCR has historically been a cloud service: you upload the image, someone's server reads it, you get text back. For a confidential scan, that's a real exposure.
It no longer has to be. OCR can run as WebAssembly inside your browser, on your own machine — the scan is read locally, the pixels never travel anywhere. It's a little slower than a GPU server farm, but for the enormous category of "documents I can't upload," running the OCR locally is the difference between usable and off-limits. Watch the network tab if you don't believe it: during conversion, nothing carrying your file should leave the page.
Once it's text, it's like any other document
The moment OCR has turned your scan into Markdown, it behaves like any converted document: you can check its token count, see whether it fits your model, and chunk it if it doesn't. The hard part was always getting out of the image — after that, you're back on solid ground.
In short
A scanned PDF isn't a document a computer can read until OCR makes it one. Convert it with OCR that runs locally in your browser, spot-check the numbers and tables, heed the low-confidence flags, and you'll turn a stack of page-images into clean, private, model-ready Markdown.
Try it on your own file
Convert a document and watch the token counter — free, no account, nothing uploaded.