See how the same text tokenizes differently across model families, token by token. The counts diverge more than most people expect, and the divergence is what makes cross-provider cost comparison misleading.
Paste any prompt. See token counts and total cost across the seven frontier and workhorse models side by side. Tokenizer estimates are empirical (Anthropic ~3.7 chars/token, OpenAI ~3.9, Google ~4.1).
| Model | Input tokens | Input $ | Output $ | Total $ | Note |
|---|---|---|---|---|---|
| Claude Opus 4.7 | 58 | $0.000870 | $0.037500 | $0.038370 | Anthropic frontier |
| Claude Sonnet 4.6 | 58 | $0.000174 | $0.007500 | $0.007674 | Anthropic workhorse |
| Claude Haiku 4.5 | 58 | $0.000046 | $0.002000 | $0.002046 | Anthropic fast |
| GPT-5 | 55 | $0.000660 | $0.030000 | $0.030660 | OpenAI frontier |
| GPT-5 mini | 55 | $0.000138 | $0.006000 | $0.006137 | OpenAI workhorse |
| Gemini 2.5 Pro | 53 | $0.000066 | $0.005000 | $0.005066 | Google frontier |
| Gemini 2.5 Flash | 53 | $0.000008 | $0.000300 | $0.000308 | Google fast |
Comparing providers on price per token is only meaningful if a token means the same thing, and it does not. The same input can produce meaningfully different counts under different tokenizers, so a provider with a lower headline rate can be more expensive for your actual workload. The only honest comparison is cost for the same input, not cost per token.
The gap widens sharply for code and non-English text. Tokenizers are trained predominantly on English prose, so code fragments into more tokens than its length suggests, and languages using non-Latin scripts can cost several times more per character. If your workload is multilingual, this is not a rounding error. It is a budget line.