Qwen Token Counter
Exact counts for Qwen 2.5, from its own tokeniser, in your browser.
Runs entirely in your browser
Loading the tool…
- —Tokens
- —Characters
- —Words
- —Characters per token
Each level includes the ones before it.
- —Tokens before
- —Tokens after
- —Saved
- —Reduction
Read this before you use the output. Removing a word can change what a model does, and the sentences that compress worst are the ones carrying your constraints.
Show the tokens themselves
Every coloured block is one token. This is the split the model sees, which is why a long number costs more than a long word.
Qwen 2.5 tokenises with a byte-level BPE of its own, 151,643 entries, published under Apache 2.0. This page carries it and runs the same algorithm, so the count is the count, not a conversion from a GPT number.
It differs from OpenAI's encodings in two ways that change the
arithmetic. Qwen splits digits one at a time, where GPT takes them
in runs of three, so 1234567 is seven tokens here and three
there. And it normalises text to NFC first, which OpenAI's tokenisers do
not: decomposed characters, which is most accented text off a Mac and a
good deal of CJK input, are composed before counting. Both differences are
implemented rather than approximated, and the build checks all 1,587 test
cases against the tokeniser Qwen ships.
The reducer is the same one as on the other counters, and the saving it reports is measured with this vocabulary rather than a GPT one.
How to use it
- Paste your prompt. The token count updates as you type.
- Pick a reduction level. Each one includes the ones before it.
- Read the diff, then copy the reduced prompt if you are happy with it.
Questions
Which Qwen models does this cover?
The Qwen 2.5 family, whose tokeniser this is. Qwen 3 and later may ship a different vocabulary; where they do, this count would be close but not exact, and close is the thing this page exists to avoid. Check the tokeniser your model actually loads if the number matters.
Why is a number so many more tokens here than on GPT?
Because Qwen's pre-tokeniser cuts digits one at a time and OpenAI's cuts them in runs of up to three. A ten-digit number is ten tokens here and four there. If your prompts carry tables of figures, that difference is most of your bill and it is worth measuring rather than assuming.
What is NFC normalisation and why does it matter?
Unicode can write the same accented character two ways: as one code point, or as a base letter followed by a combining mark. Qwen's tokeniser composes them into the single form before counting; OpenAI's tokenisers do not. Text copied off a Mac is frequently in the decomposed form, so counting it with the wrong rule gives a number that is too high. This page applies NFC on this vocabulary and not on the others, which is what each tokeniser actually does.
Can I count Llama or Mistral with this?
No. Llama's tokeniser sits behind a gated repository and a licence with terms; Mistral's uses a different algorithm again, SentencePiece with a Metaspace pre-tokeniser rather than byte-level BPE, so it is a separate implementation rather than another vocabulary. Neither is here yet, and guessing with the wrong tokeniser is exactly the error this family of pages exists to avoid.
The newer vocabulary
The same counter and the same reducer, carrying the vocabulary the current models use.
Your data stays on your device
Everything above runs inside your browser as WebAssembly compiled from Rust. Nothing you type is uploaded, logged or stored on a server. You can load this page once, go offline, and it still works.
This page makes no requests at all, to anywhere. That is not a promise in the copy: it is a Content-Security-Policy header your browser enforces, and connect-src on it is none. Open the network tab and watch nothing happen.