What the context window is
It's the most tokens the model sees in one request: instructions, history, documents and the answer itself. Claude Opus 5.5, Sonnet 5.5 and Fable 5.1 have 1 million tokens; Haiku 4.5 has 200K. GPT-5.6 takes 1.05 million, with up to 922K of input.
Big context, big bill
Filling 1 million input tokens on Claude Opus 5.5 costs $4 per request. If the same document goes into every call, use prompt caching. On GPT-5.6, above 272K input tokens the whole request costs 2× on input and 1.5× on output.
Max output
Output has its own limit, far below the context: 128K tokens on Claude 5.x models and GPT-5.6, and 64K on Haiku 4.5.