LLM Token Counter and Prompt Reducer
Count tokens exactly, then cut the ones you are paying for nothing.
Runs entirely in your browser
Loading the tool…
- —Tokens
- —Characters
- —Words
- —Characters per token
Each level includes the ones before it.
- —Tokens before
- —Tokens after
- —Saved
- —Reduction
Read this before you use the output. Removing a word can change what a model does, and the sentences that compress worst are the ones carrying your constraints.
Show the tokens themselves
Every coloured block is one token. This is the split the model sees, which is why a long number costs more than a long word.
Paste a prompt and the count above is the number the model will actually
charge you for. Not an estimate and not a word count multiplied by 1.3: this
page carries OpenAI's o200k_base vocabulary, all 199,998
entries of it, and runs the same byte-pair encoding their tokeniser runs.
It is checked against tiktoken itself over 1,587 cases, so
"exact" is a test result rather than a claim.
The second half is the reason the first half is useful. Most prompts carry 30 to 40 per cent that costs money and does nothing: I would like you to please, it is very important that you should try to, in order to. The reducer takes those out in four graded steps and shows you a word-level diff of what left, because a compressor whose changes you cannot see is one you cannot put near a prompt that matters.
Everything happens in this tab. The vocabulary is inlined into the page, the encoder is WebAssembly compiled from Rust, and the page is served with a header forbidding it to open a connection to anywhere. Your prompts are frequently the most sensitive text you own, and this is the rare token counter you can use on one without handing it to a server first.
How to use it
- Paste your prompt. The token count updates as you type.
- Pick a reduction level. Each one includes the ones before it.
- Read the diff, then copy the reduced prompt if you are happy with it.
Questions
Is this the real tokeniser, or an estimate?
The real one. This page ships OpenAI's published o200k_base vocabulary and implements the same byte-pair encoding, including the pre-tokenisation rules, which is the part most reimplementations get wrong. It is tested against tiktoken over 1,587 cases covering contractions, CJK, combining marks, emoji with zero-width joiners, Windows line endings and runs of whitespace, and every one must match token for token or the build fails.
Which models does this count for?
The ones using o200k_base: GPT-5, GPT-4o, GPT-4o mini, GPT-4.1, o1 and o3. GPT-4, GPT-4 Turbo, GPT-3.5-turbo and the embedding models use the older cl100k_base vocabulary and are counted on the GPT-4 token counter instead. Using the wrong one gives a wrong number rather than an error, which is why they are separate pages.
Does this work for Claude, Gemini or Llama?
No, and it would be dishonest to pretend otherwise. Anthropic does not publish a tokeniser for current Claude models: the only accurate count is their own count_tokens endpoint, which is a network call. Claude 4.7 and later use a newer tokeniser producing roughly 30 per cent more tokens than earlier ones for the same text, so an offline guess would be wrong by a different amount depending on the model. Most sites offering a "Claude token counter" are running a GPT tokeniser and showing you the wrong number. Llama and Mistral publish theirs and could be added; Claude cannot be.
Is the token count the same as my bill?
Close, but not identical, and the gap is worth knowing. A chat request adds a few tokens of structure per message, and system prompts, tool definitions, images and documents all add more. What this counts is the tokens in the text you pasted. For a system prompt you are about to send a million times that is the number that matters; for an exact invoice, only the API's own usage figures are authoritative.
Will the reducer change what the model does?
It can, and the higher levels are likelier to. Whitespace cannot. Removing "please" and "I would like you to" almost never does. Caveman, which drops articles and copulas, sometimes does: "the refund is approved by a manager, not by the agent" becomes "refund approved manager, not agent", which lists the right nouns in the wrong relationship. That level is off by default and the diff is there so you can see it happen. Re-test anything you reduce.
What does it refuse to touch?
Fenced code blocks, inline code, quoted strings and {placeholders}. Those are the parts of a prompt where an edit is a bug rather than a rewording, and no rule is allowed inside them. Few-shot examples cannot be detected reliably, so read the diff if your prompt contains any.
Why is a long number more tokens than a long word?
Because the pre-tokeniser cuts digits into runs of at most three before encoding begins, so 1234567 is three tokens where a seven-letter word is usually one. Open "Show the tokens themselves" to watch it happen. It is worth knowing when a prompt carries tables of figures: those cost far more than their length suggests.
Does my prompt get uploaded?
No. The vocabulary and the encoder are both inside this page, and the page is served with a Content-Security-Policy whose connect-src is none, so the browser will not let it open a connection even if it tried. Load the page once, go offline, and it still counts. Prompts often contain things you would not paste into a form on somebody else's server, which is most of the point.
Where the tokens actually go
Two rules explain most surprises. Digits are cut into runs of at most three
before encoding, and a leading space belongs to the word after it rather than
the word before. That second one is why " the" is one token and
"the" at the start of a line is another.
| Text | Tokens | Why |
|---|---|---|
hello world | 2 | Common words, one each, the space riding with the second |
1234567 | 3 | Digits cut into runs of three |
antidisestablishmentarianism | 6 | Rare, so it falls back to pieces |
中文测试 | 4 to 8 | CJK costs more per character than Latin text |
🎉 | 2 | One emoji is several bytes, and bytes are what BPE merges |
indented | 2 | Runs of spaces have tokens of their own |
The practical consequence for a prompt you send often: tables of numbers, deep JSON and non-Latin text all cost more than their length suggests, and polite framing costs exactly as much as instruction does.
The other vocabulary
Which vocabulary a model uses is not a detail: counting with the wrong one gives a wrong number and no warning. The older models have a page of their own:
Your data stays on your device
Everything above runs inside your browser as WebAssembly compiled from Rust. Nothing you type is uploaded, logged or stored on a server. You can load this page once, go offline, and it still works.
This page makes no requests at all, to anywhere. That is not a promise in the copy: it is a Content-Security-Policy header your browser enforces, and connect-src on it is none. Open the network tab and watch nothing happen.