MarkPrep
All posts
2026-09-17 · 6 min

HTML to Markdown: Getting the Article Without the Page Around It

A web page is mostly not the thing you came for. Wrapped around the article you actually want sits the navigation, the cookie banner, a "subscribe to our newsletter" box, three related-article headlines, the ads, and the footer — easily nine-tenths of the page. Your eye filters all of it out without trying. An LLM has no such filter: give it a web page and it reads the furniture along with the content, and it charges you tokens for every word of it.

Converting HTML to Markdown is really about one thing: keeping the article and throwing the page away. Done right, you hand the model clean, structured content — headings, paragraphs, lists, links — and none of the chrome. Here's how it works and what separates a clean conversion from a messy one.

HTML already has structure — that's the good news

Unlike a PDF, HTML tells you what everything is. An <h1> is a heading, a <ul> is a list, a <table> is a table, an <a> is a link. So the mapping to Markdown is direct and near-lossless: headings become #, lists become -, links become [text](url). This is why HTML sits among the easiest formats to convert — the structure is explicit, no reconstruction required.

The catch is that the structure of the content is tangled up with the structure of the page. The nav bar is also a <ul>. The cookie banner is also a <div> full of text. So a good HTML-to-Markdown conversion isn't just "map the tags" — it's "map the tags and drop the ones that aren't content."

What a clean conversion removes

The difference between usable and unusable output is almost entirely about subtraction. A good converter strips:

  • Navigation, headers, footers — the same on every page, meaningless to a model.
  • Scripts and styles — code, not content.
  • Cookie/consent banners, newsletter prompts, share buttons — interface, not information.
  • Ads and "related content" widgets — actively misleading if they end up in your context.

What survives is the article: its heading structure, its paragraphs, its real links and tables. MarkPrep's converter does this subtraction — it drops the script, nav, and footer elements and keeps the content — and it also handles the common case of pointing at a page by URL, fetching the HTML and converting it in one step.

Tables and links survive (and that matters)

Two things people often lose in a sloppy HTML conversion are worth protecting. Tables should come out as Markdown pipe tables, not flattened text — the same rule that matters for spreadsheets. And links are worth keeping as real Markdown links: for an LLM building context, knowing that a phrase pointed somewhere is often useful signal, not noise.

The one privacy note for URLs

Most HTML conversion is of public web pages, so privacy is less charged than with your own contracts. The one honest exception: if a converter fetches a URL for you, that fetch is a network request — the tool has to go get the page. That's expected and fine for public pages. It's worth knowing the distinction: converting HTML you already have is fully local; fetching a URL means one outbound request to retrieve the page, which any tool must make.

After the cruft is gone

Stripping the page furniture does something nice beyond cleanliness: it shrinks the token count, often dramatically, because most of what you removed was tokens the model would have been charged for and learned nothing from. Once you've got just the article, a quick token check tells you whether it fits — and clean web content usually does, now that the 90% of noise is gone.

The point

An LLM can read a clean article and reason about it well. It reads a full web page — banners, nav, ads and all — badly and expensively. Convert the HTML to Markdown first, drop everything that was page rather than content, keep the tables and links, and give the model the 800 words you actually meant.

Try it on your own file

Convert a document and watch the token counter — free, no account, nothing uploaded.