How API billing works
AI APIs bill per million tokens (MTok), with separate prices for what you send (input) and what the model writes (output). Output costs 5 to 6 times more than input, so long answers weigh more than long prompts.
Example with Claude Sonnet 5.5 ($2 input and $10 output per million): 2,000 input tokens and 500 output tokens cost $0.009 per request. At 1,000 requests a day for 30 days, that's $270 a month.
Current prices per 1 million tokens
| Model | Input /1M | Cached /1M | Output /1M | Context |
|---|---|---|---|---|
| Claude Fable 5.1 | $10.00 | $0.25 | $50.00 | 1,000,000 |
| Claude Opus 5.5 | $4.00 | $0.20 | $20.00 | 1,000,000 |
| Claude Sonnet 5.5 | $2.00 | $0.20 | $10.00 | 1,000,000 |
| Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 | 200,000 |
| GPT-5.6 Sol | $5.00 | $0.50 | $30.00 | 1,050,000 |
| GPT-5.6 Terra | $2.00 | $0.20 | $12.00 | 1,050,000 |
| GPT-5.6 Luna | $0.20 | $0.0200 | $1.20 | 1,050,000 |
Public list prices as of October 2026, in US dollars, excluding taxes.
Three ways to pay less
- Prompt caching: the part that repeats on every call (instructions, documents) is read from cache for a fraction of the price. Use the "Input read from cache" field or the caching calculator.
- Batch API: if results can wait up to 24 hours, Anthropic and OpenAI give 50% off input and output.
- Right model per task: classification, extraction and short answers run fine on cheap models (Haiku, Luna); keep Opus, Fable and Sol for work that needs reasoning.
What the estimate leaves out
Sales tax or VAT, and the cost of writing to Claude's cache (1.25× input), which the caching calculator does include.