AI for management · updated 8 September 2026
Which AI can a management team trust?
- Part of the AI Board research programme
Memory: the advertised window
The largest published context window went from 4.1K to 10M tokens, so document length is rarely the limit any more.
| Point | series |
|---|---|
| 2022 | 4.1K |
| 2023 | 8.19K |
| 100K | |
| 128K | |
| 200K | |
| 2024 | 2M |
| 2025 | 10M |
Practical memory versus advertised
The advertised window is not working memory: the strongest model in the NoLiMa study lost 29.6 percentage points of accuracy by 32,000 tokens.
- Short text99.3%
- At 32,000 tokens69.7%
| Label | Value |
|---|---|
| Short text | 99.3% |
| At 32,000 tokens | 69.7% |
Confident mistakes per generation
Between GPT-4o and GPT-5 in thinking mode, the share of answers carrying a major factual error fell by a factor of 4.3.
- GPT-4o (May 2024)20.6%
- o3 (Apr 2025)22.0%
- GPT-5 main (Aug 2025)11.6%
- GPT-5 thinking (Aug 2025)4.8%
| Label | Value |
|---|---|
| GPT-4o (May 2024) | 20.6% |
| o3 (Apr 2025) | 22.0% |
| GPT-5 main (Aug 2025) | 11.6% |
| GPT-5 thinking (Aug 2025) | 4.8% |
EU residency among flagships
8 of the 13 current flagship models have a documented route to running with data residency inside the EU.
- Documented EU route8
- No documented EU route5
| Segment | Value | Share |
|---|---|---|
| Documented EU route | 8 | 62% |
| No documented EU route | 5 | 38% |
Cost per executive per working week
At public list prices, one executive's week of heavy use costs between US$0.89 and US$11.88, before any discount.
- GPT-6 AstraUS$11.88
- Claude Fable 5.1US$10.62
- Claude Opus 5US$5.94
- Qwen3.8-MaxUS$5.16
- Mistral Medium 3.5US$4.05
- Grok 4.6US$2.64
- GPT-5.6 TerraUS$2.50
- Claude Sonnet 5US$2.38
- DeepSeek-V4-ProUS$1.26
- Gemini 3.8 FlashUS$0.89
| Label | Value |
|---|---|
| GPT-6 Astra | US$11.88 |
| Claude Fable 5.1 | US$10.62 |
| Claude Opus 5 | US$5.94 |
| Qwen3.8-Max | US$5.16 |
| Mistral Medium 3.5 | US$4.05 |
| Grok 4.6 | US$2.64 |
| GPT-5.6 Terra | US$2.50 |
| Claude Sonnet 5 | US$2.38 |
| DeepSeek-V4-Pro | US$1.26 |
| Gemini 3.8 Flash | US$0.89 |
What benchmarks do not answer
Five questions a management team actually asks, and no public benchmark answers any of them on your own board pack.
- Memory4
- Factuality3
- Honesty2
- Citation2
- Discipline3
| Label | Value |
|---|---|
| Memory | 4 |
| Factuality | 3 |
| Honesty | 2 |
| Citation | 2 |
| Discipline | 3 |
02Landscape now
The current flagship models, side by side
Every current flagship in the dataset, with the facts a management team decides on: how much it can read, whether it can run in the EU, and what it costs per million tokens.
| Lab | Model | Memory | EU residency | In / M | Out / M |
|---|---|---|---|---|---|
| OpenAI | GPT-6 Astra | 1.05M tokens | No | US$10.00 | US$50.00 |
| Anthropic | Claude Fable 5.1 | 1M tokens | YesEU multi-region | US$10.00 | US$50.00 |
| Anthropic | Claude Mythos 5.1 | 1M tokens | No | US$10.00 | US$50.00 |
| Google (DeepMind) | Gemini 3.8 Flash | 1.05M tokens | YesEU multi-region | US$0.75 | US$3.75 |
| xAI | Grok 4.6 | 500K tokens | No | US$2.00 | US$6.00 |
| Meta | Muse Spark 1.3 | 1M tokens | No | US$1.25 | US$4.25 |
| DeepSeek | DeepSeek-V4-Pro | 1M tokens | Yesself-hosted | US$1.32 | US$3.96 |
| Mistral AI | Mistral Medium 3.5 | 256K tokens | YesEU region | US$1.50 | US$7.50 |
| Alibaba (Qwen) | Qwen3.8-Max | 1M tokens | YesEU region | US$2.00 | US$6.00 |
| Z.ai | GLM-5.3 | 1M tokens | No | US$1.40 | US$4.40 |
| Moonshot AI | Kimi K3 | 1.05M tokens | Yesself-hosted | US$3.00 | US$15.00 |
| Cohere | Command A+ | 128K tokens | Yesself-hosted | not published | not published |
| NVIDIA | Nemotron 3 Ultra | 1M tokens | Yesself-hosted | not published | not published |
See the full landscape, including the tier below and the models being retired
03The research behind the numbers
Three pages, one question
Nobody reads a research report end to end. The charts above are the summary; these pages are where each number is argued, sourced and dated.
AI Board Model Index, first edition
The ranking lands on this page. There are no results yet and no scores anywhere on this site, because nothing has been measured. When the first edition is published, the leaderboard opens here and the method page becomes its appendix.
04For your role
What this means at your seat
The same landscape looks different from three chairs. Each role has a section in the guide and a page on what an AI assistant does in that job.
CEO
Synthesis across a messy pile of documents, and an answer you can repeat in a meeting without checking it first.
CFO
Citation discipline. A number without a document and a page number is an opinion, however confident it sounds.
CTO
Residency, licences and the exit route. The generation you can host in Europe is not always the newest one.
05Method
How we work
This page follows the same discipline as the rest of the site: nothing is published that cannot be checked.
- Primary sources only
- Every fact about a model comes from vendor documentation, an official launch post, a system card, a peer-reviewed paper, or a cloud provider's own region list. Never from a summary of a summary.
- Facts as of a stated date
- Model specifications change weekly. Every chart prints the date its figures were verified, so you can see how old a claim is before you act on it.
- What we cannot verify, we leave out
- If a figure has no primary source, it does not appear on the page, not as an estimate and not as a range. A shorter page is the price of a checkable one.
- No vendor money
- No model vendor pays for a place in this research, and none of them sees a page before it is published.
- Corrections
- Found something that is wrong or out of date? Tell us and we correct the page, with the change dated. Contact details are on the about page.
Become AI-native before your competition does
Ride the AI wave instead of swimming behind it. Request access and we'll schedule your install.
Your company brain lives on your laptop. You choose if a question goes to a cloud model or stays fully local.