01
Memory: from half a board paper to a full year of them
In one sentenceThe flagship context window went from 8,192 tokens in March 2023 to 1,050,000 in September 2026, but the usable part is smaller than the advertised part.
In March 2023 GPT-4 arrived with a context window of 8,192 tokens, a figure OpenAI still publishes on its model documentation page. Take a twenty page board paper as roughly 9,000 words, which is about 12,000 tokens. That is the assumption this report uses throughout, and it means GPT-4 could not hold one board paper in view at once. It held about two thirds of one.
What this means for the board
- Fit is no longer the constraint. A year of minutes, the annual report and the budget will fit in one conversation with any current flagship, so stop scoping AI work around document size.
- Ask for the usable number, not the advertised one. Published research shows accuracy falling well before the stated limit, so treat the window as a ceiling and require evidence at the length you actually use.
Context window of OpenAI's flagship, March 2023 to September 2026
| Point | series |
|---|---|
| 2023 H1 | 8.19K |
| 2023 H2 | 128K |
| 2024 H1 | 128K |
| 2024 H2 | 128K |
| 2025 H1 | 1.05M |
| 2025 H2 | 400K |
| 2026 H1 | 1.05M |
| 2026 H2 | 1.05M |
Board papers: 1 = 12,000 tokens.
How this was counted
One lineage is plotted, OpenAI's flagship, because it is the single line most boards have actually used and every point is on OpenAI's own model pages. Rival flagships tracked closely and are named in the text. The board paper series divides the window by 12,000 tokens, being a twenty page paper of about 9,000 words at roughly 1.33 tokens per word. Token counts differ per tokenizer, so read the paper counts as an order of magnitude, not a measurement.
Evidence that argues the other way
Of the models claiming a context window of 32,000 tokens or more, only half maintained satisfactory performance at 32,000 tokens.
RULER: What's the Real Context Size of Your Long-Context Language Models? (arXiv:2404.06654)
On the NoLiMa evaluation, which removes literal word overlap between question and answer, GPT-4o fell from 99.3 percent at short context to 69.7 percent at 32,000 tokens, and eleven of thirteen models tested dropped below half their own short-context score.
NoLiMa: Long-Context Evaluation Beyond Literal Matching (arXiv:2502.05167)
The largest generally available context window has not moved since 27 June 2024, when Google opened two million tokens on Gemini 1.5 Pro to all developers. Every verified frontier flagship in September 2026 sits at one million to 1.05 million.
Google, new features for the Gemini API and Google AI Studio. OpenAI model documentation. Anthropic model overview. Gemini API: Gemini 3.8 Flash