What a token is
Language models don't read letters or words: they read tokens, chunks of text that may be a whole word, part of one or a symbol. In English a token averages about 4 characters; in Spanish and Portuguese tokens are shorter, because tokenizers were trained on more English text.
How we estimate
The counter divides the number of characters by each tokenizer's average characters per token for the language:
| Tokenizer | Portuguese | Spanish | English | Code |
|---|---|---|---|---|
| GPT-5.6 | 3.6 | 3.7 | 4.2 | 3.3 |
| Claude 5.x (Opus, Sonnet, Fable) | 2.55 | 2.6 | 2.9 | 2.4 |
| Claude Haiku 4.5 | 3.3 | 3.4 | 3.8 | 3.1 |
According to Anthropic, Claude models from 4.7 on use a new tokenizer that produces roughly 30% more tokens for the same text. That's why the same prompt "weighs" more on Opus 5.5 than on Haiku 4.5.
When you need the exact number
The estimate is within about ±15% for ordinary text. Both APIs return the real count with every response (the usage field), and Anthropic offers a free token-counting endpoint you can call before sending.
Text you paste here never leaves your browser.