How Many Tokens Is My Document? (And Why You Should Care)
The error always shows up at the worst possible moment. You've pasted a long document into an LLM to actually get something done, and instead of an answer you get a wall of red text about exceeding some token limit — a number you have no way to check against a document you never measured. Nobody teaches you to look at a file and think "that's about 60,000 tokens." So you find out the hard way, mid-task, when it's most annoying.
The fix is boring and it's this: know the token count before you paste. Not the word count — the token count. They're not the same thing, and the gap between them is where people get caught out. Here's what a token actually is, how to count them in a document, and why the number matters more than the page count you've been eyeballing.
What a token actually is
A token is the unit a language model reads in. It isn't a word and it isn't a character — it's a chunk somewhere in between. Common words are often a single token; longer or rarer words get split into several; punctuation and even the space before a word affect the split.
OpenAI's own guidance gives the rule of thumb everyone uses: roughly "4 characters" per token, or about "three-quarters of a word" for English text — so 100 tokens is about 75 words. It's an estimate, not a law: the same text tokenizes differently across models and languages. But as a mental yardstick it's close enough to be useful.
The practical translation: multiply your word count by about 1.3 and you're in the right ballpark for tokens. A 10,000-word report is roughly 13,000 tokens. A dense 90-page PDF can easily clear 100,000.
Why word count lies
Here's the trap. You look at a document, see "12 pages," and assume it's small. But pages are a layout accident — font size, margins, images, and whitespace decide how much text lands on a page, and none of that tells a model anything. Two "10-page" PDFs can differ by 5× in actual token load.
Worse, token count isn't even a fixed property of the text — it depends on the model's tokenizer. The same paragraph is a different number of tokens to GPT, Claude, and Gemini, because they encode text differently. Code, tables, unusual formatting, and non-English text all tokenize less efficiently than clean prose, quietly inflating the count. So "how long is this document" has no single answer until you tie it to a specific model.
This is why guessing fails and why the length errors feel so arbitrary. You're estimating a number that (a) doesn't match the word count you can see and (b) changes depending on which model you're pasting into.
Why the number matters: the context window
A token count on its own is trivia. It becomes useful the moment you compare it to your model's context window — the maximum number of tokens the model can hold at once (your input plus its output).
If your document fits, great, paste away. If it doesn't, the model does one of a few unhelpful things: rejects it, truncates it silently, or forces you to hack it into pieces by hand. Modern context windows have grown enormously and keep changing model to model and month to month — so the only number that matters is your model's current limit, checked against this document's token count. Chasing a memorized figure is pointless when both sides of the comparison move.
That comparison — "does this specific document fit this specific model?" — is the actual question. It's also the one almost nothing in the usual workflow answers for you.
How do I count the tokens in a document?
A few options, roughly in order of effort:
- Estimate: word count × ~1.3. Fine for a gut check; wrong enough to burn you near a limit.
- A tokenizer tool: paste text in and get an exact count for a given model. Accurate, but you have to extract the text first and do it per model.
- Count at conversion time: if you're turning a PDF or Word file into Markdown anyway, the ideal is to get the token count as part of that step, per model, without a second tool. That's exactly what MarkPrep's token counter does — it counts as it converts and tells you, live, whether the result fits your model's window or needs splitting.
The reason we built it into the converter is that the token count is useless after you've already pasted and hit the wall. You want it at the moment you have the document in hand, not after the error.
What to do when it doesn't fit
Knowing the number is half the value; the other half is knowing your options when it's too big:
- Chunk it deliberately — split on heading boundaries, not by slicing every N characters through the middle of a table. Seeing the chunk plan before you split is the difference between retrieval that returns whole sections and retrieval that returns fragments.
- Trim the dead weight — repeated headers, footers, and page numbers add tokens and teach your model nothing. Stripping them shrinks the count and improves quality at the same time.
- Pick the right model — if a document is close, a model with a larger current window may swallow it whole. But check the number first, don't assume.
Stop guessing
The whole problem is that you've been judging document size by the one metric — page count — that has nothing to do with what a model reads. Token count is the real measure, word count is a loose proxy for it, and page count is noise.
Next time you're about to feed a document to an LLM, get its token count first. If you're converting it to Markdown anyway, do both in one step — clean Markdown out, exact token count per model, and a straight answer on whether it'll fit before you paste a single character.
Try it on your own file
Convert a document and watch the token counter — free, no account, nothing uploaded.