Turning a Pile of Research Papers Into a RAG Knowledge Base
A researcher's folder of PDFs is a specific kind of problem. There might be two hundred papers in there, and the dream is obvious: ask questions across all of them, find the paper that made a particular argument, have an assistant that's actually read your field. That's a textbook use for retrieval-augmented generation. But academic PDFs are also just about the hardest documents to convert cleanly — and if the conversion is bad, the whole knowledge base is built on sand.
Here's what makes research papers difficult, how to convert them to Markdown so they're usable, and how to think about the RAG pipeline on top.
Why papers fight back
A typical journal PDF throws several problems at a converter at once:
- Two-column layouts. Read naïvely, a two-column page interleaves the columns — you get the first line of the left column, then the first line of the right, then back — producing text that's technically present and completely scrambled. Column order is the single most common way paper conversion goes wrong.
- Dense tables. Results tables are the point of many papers, and they're often intricate. Flattened or mis-aligned, they become useless.
- Equations. Mathematical notation rarely survives as clean text; expect it to be approximate.
- Citations, headers, and footers. Running heads, page numbers, and reference markers clutter the extracted text if they're not stripped.
A converter that reconstructs the reading order correctly — following each column top to bottom before moving on — clears the biggest hurdle. MarkPrep's converter clusters text into proper reading order and strips the repeated page furniture, which is exactly what a scrambled two-column extraction needs.
Clean the text before you embed
This is the step people skip, and it quietly wrecks retrieval. When you embed a chunk of text into a vector store, everything in that chunk becomes part of its meaning — including the header that repeats "J. Neural Methods, Vol 12" on every page, and the page numbers, and the citation cruft. That noise dilutes the embedding and makes retrieval fuzzier. Stripping headers, footers, and page numbers before embedding isn't cosmetic; it measurably sharpens what comes back. The clean-up that happens during conversion is doing real work for your retrieval quality.
Chunk on structure, not by the character
Once each paper is clean Markdown, the heading structure — Abstract, Introduction, Methods, Results, Discussion — is a gift. Chunk on those boundaries and each chunk is a coherent, self-contained passage, which is precisely what you want a retriever to return. Slice every 500 characters instead and you'll cut through the middle of a Methods description and strand half a Results table, so a query gets back a fragment that answers nothing. Seeing the chunk plan before you commit lets you keep sections intact and attach the heading path to each chunk, so a retrieved passage carries the context of where in the paper it came from.
The privacy angle researchers actually care about
Plenty of what's in that folder is unpublished — your own drafts, papers under review, confidential preprints, industry data. Uploading those to a cloud converter to build your knowledge base is exactly the wrong move. Converting locally, in the browser, means your unpublished work stays yours while you build the index. For anyone working with pre-publication material, that's not a preference, it's the difference between being able to use these tools at all and not.
Putting it together
The pipeline for a paper-based knowledge base:
- Convert each PDF to Markdown with correct column reading order.
- Clean — strip headers, footers, page numbers.
- Chunk on section headings, attaching the heading path as metadata.
- Embed and retrieve the resulting passages.
Get step one wrong — scrambled columns, cruft everywhere — and no amount of clever retrieval saves it. Get the conversion right and the rest of RAG becomes the well-trodden part.
Bottom line
Research papers are the stress test of document conversion: columns, tables, equations, and citation noise all at once. Convert them to Markdown with real reading-order reconstruction, clean the page furniture before embedding, chunk on the paper's own sections, and keep unpublished work on your own machine while you do it. That's a knowledge base you can trust.
Try it on your own file
Convert a document and watch the token counter — free, no account, nothing uploaded.