Why each message costs more than the last
The API doesn't remember the conversation. On every customer message your system resends the system prompt, the full history and the new message. Over a conversation of n messages, total input grows quadratically:
With 6 messages, a 3,000-token prompt is read 6 times: 18,000 tokens of instructions alone. That's why caching the system prompt is usually the easiest saving.
How to cut the cost
- Cache the system prompt and knowledge base.
- Summarize the history once a conversation passes 10 to 15 messages.
- Start on a cheap model (Haiku 4.5, GPT-5.6 Luna) and escalate to a bigger one only for hard questions.