Why agents burn so many tokens
Every agent step (reading a file, running a command, editing code) is a request that resends instructions, history and the files already read. A 40-step session with 60K tokens of context reads 2.4 million input tokens.
Caching is what saves you: with 90% of the context read from cache, each step on Claude Opus 5.5 costs a few cents.
Plan or API?
For heavy daily use, Max plans usually come out cheaper than the API, as long as the weekly limits hold. For occasional use, or in CI and automation, the API with a budget is more predictable.
How to spend less
- Start a fresh session for each new task instead of dragging a huge history along.
- Use a smaller model for mechanical work (tests, formatting).
- Keep the project instructions file short and stable so it stays cached.