The Same Model Charges You 3× More Per Character — Depending on What You Paste
Every token calculator on the internet uses some characters-per-token ratio, usually 4. We wanted
to know how wrong that is, so we built a fixed public corpus — prose, TypeScript, Python, JSON/YAML,
markdown, URLs, and CJK text, 5,681 characters, CC0-licensed with a published hash — and ran it
through four vendors' own tokenizer.json files. Not approximations of their tokenizers: the
actual files they publish, executed exactly.
The result that reframed our roadmap
| Content | Qwen3 | Tekken v3 (Mistral) | GLM | DeepSeek | | :----------------- | ----: | ------------------: | ---: | -------: | | Prose | 4.99 | 4.88 | 4.99 | 4.99 | | Markdown | 4.50 | 4.44 | 4.50 | 4.36 | | Code | 4.05 | 3.87 | 4.09 | 3.85 | | JSON / YAML | 2.50 | 2.45 | 2.74 | 2.70 | | URLs & identifiers | 1.67 | 1.63 | 1.93 | 1.99 | | CJK | 1.47 | 1.23 | 1.47 | 1.62 |
Read it vertically: one tokenizer spans roughly a 4× range from clean prose down to CJK, with JSON at about half of prose and raw identifiers at a third. Read it horizontally: four independently built tokenizers, on identical text, land within about 4% of each other for most classes.
Content type moves your token count roughly 25× more than model choice does.
The control that made us trust the harness
DeepSeek publishes its own ratios — the only vendor that states both an English and a Chinese figure. Our measured Chinese ratio (1.62 chars/token) reproduces their published 1.667 to within 1%, which is what says the harness works. Our measured English ratio (4.99 on prose) does not match their published 3.33 — because a single published "English" number is a conservative planning figure blended across content, not a measurement of prose. That gap is the whole story of this post.
What it means for a real paste
A realistic deploy runbook — markdown headings, a fenced TypeScript block, a fenced JSON config, a dashboard URL — costs 25–45% more tokens than a flat 4-chars/token estimate says. Our token calculator now classifies what you paste and prices each class at its own measured ratio, showing you the split instead of one blended number.
The corpus, its hash, and the full method are published so the numbers are reproducible by someone who does not trust us. That is the standard every number on this site is supposed to meet.