Skip to content
AI Board

AI for management · updated 8 September 2026

Which AI can a management team trust?

  • Part of the AI Board research programme

Memory: the advertised window

The largest published context window went from 4.1K to 10M tokens, so document length is rarely the limit any more.

The largest vendor-stated context window rose from 4.1K tokens in November 2022 to 10M tokens in April 2025, drawn on a logarithmic scale.
Pointseries
20224.1K
20238.19K
100K
128K
200K
20242M
202510M
Running maximum of the vendor-stated context window, November 2022 to April 2025. The last step is Meta's own claim for Llama 4 Scout, an advertised window rather than measured memory. Logarithmic scale. Source: vendor documentation and launch posts, read 8 September 2026.

Read the analysis

Practical memory versus advertised

The advertised window is not working memory: the strongest model in the NoLiMa study lost 29.6 percentage points of accuracy by 32,000 tokens.

  • Short text99.3%
  • At 32,000 tokens69.7%
In the NoLiMa study GPT-4o scored 99.3% on short text and 69.7% once the context reached 32,000 tokens.
LabelValue
Short text99.3%
At 32,000 tokens69.7%
NoLiMa (Modarressi et al., ICML 2025), GPT-4o. At 32,000 tokens 11 of the 13 models tested, all of which advertise at least 128,000 tokens, fell below half of their own short-text score.

Read the analysis

Confident mistakes per generation

Between GPT-4o and GPT-5 in thinking mode, the share of answers carrying a major factual error fell by a factor of 4.3.

  • GPT-4o (May 2024)20.6%
  • o3 (Apr 2025)22.0%
  • GPT-5 main (Aug 2025)11.6%
  • GPT-5 thinking (Aug 2025)4.8%
Share of answers containing at least one major factual error across 4 model generations, from 22.0% at worst to 4.8% at best. Lower is better.
LabelValue
GPT-4o (May 2024)20.6%
o3 (Apr 2025)22.0%
GPT-5 main (Aug 2025)11.6%
GPT-5 thinking (Aug 2025)4.8%
OpenAI GPT-5 system card, 13 August 2025, production-representative prompts with browsing enabled. Vendor figures on a vendor evaluation, so read them as direction. Counter-example from the same lab: on PersonQA, o3 hallucinated twice as often as the older o1.

Read the analysis

EU residency among flagships

8 of the 13 current flagship models have a documented route to running with data residency inside the EU.

8of flagships
  • Documented EU route8
  • No documented EU route5
8 of 13 current flagship models have a documented EU route, and 5 have none.
SegmentValueShare
Documented EU route862%
No documented EU route538%
A route means an EU cloud region, a European provider, or self-hosting published weights. The tier below flagship is better covered: of 39 current models, 28 have a documented route. Source: vendor and cloud provider documentation, read 8 September 2026.

Read the analysis

Cost per executive per working week

At public list prices, one executive's week of heavy use costs between US$0.89 and US$11.88, before any discount.

  • GPT-6 AstraUS$11.88
  • Claude Fable 5.1US$10.62
  • Claude Opus 5US$5.94
  • Qwen3.8-MaxUS$5.16
  • Mistral Medium 3.5US$4.05
  • Grok 4.6US$2.64
  • GPT-5.6 TerraUS$2.50
  • Claude Sonnet 5US$2.38
  • DeepSeek-V4-ProUS$1.26
  • Gemini 3.8 FlashUS$0.89
Estimated interface cost of one executive's working week across 10 models, from US$11.88 for the dearest to US$0.89 for the cheapest.
LabelValue
GPT-6 AstraUS$11.88
Claude Fable 5.1US$10.62
Claude Opus 5US$5.94
Qwen3.8-MaxUS$5.16
Mistral Medium 3.5US$4.05
Grok 4.6US$2.64
GPT-5.6 TerraUS$2.50
Claude Sonnet 5US$2.38
DeepSeek-V4-ProUS$1.26
Gemini 3.8 FlashUS$0.89
Estimate on public list prices in US dollars: 40 questions a week, 60,000 input tokens per question, 1,500 output tokens per answer, cached input where the lab publishes a cache price. An estimate, not a quote. Source: vendor pricing pages, read 8 September 2026.

Read the analysis

What benchmarks do not answer

Five questions a management team actually asks, and no public benchmark answers any of them on your own board pack.

  • Memory4
  • Factuality3
  • Honesty2
  • Citation2
  • Discipline3
Number of public benchmarks that come closest per board question: Memory 4, Factuality 3, Honesty 2, Citation 2, Discipline 3. None of them uses a real board pack.
LabelValue
Memory4
Factuality3
Honesty2
Citation2
Discipline3
Count of the public benchmarks that come closest per question, from our benchmark review of 8 September 2026. All of them measure on public or single-document material, which is the gap the Model Index is built to close.

Read the analysis

02Landscape now

The current flagship models, side by side

Every current flagship in the dataset, with the facts a management team decides on: how much it can read, whether it can run in the EU, and what it costs per million tokens.

13 current flagship models. Memory is the vendor-stated context window, not measured working memory. Prices are API list prices per million tokens in US dollars. Source: vendor documentation, read 8 September 2026.
LabModelMemoryEU residencyIn / MOut / M
OpenAIGPT-6 Astra1.05M tokensNoUS$10.00US$50.00
AnthropicClaude Fable 5.11M tokensYesEU multi-regionUS$10.00US$50.00
AnthropicClaude Mythos 5.11M tokensNoUS$10.00US$50.00
Google (DeepMind)Gemini 3.8 Flash1.05M tokensYesEU multi-regionUS$0.75US$3.75
xAIGrok 4.6500K tokensNoUS$2.00US$6.00
MetaMuse Spark 1.31M tokensNoUS$1.25US$4.25
DeepSeekDeepSeek-V4-Pro1M tokensYesself-hostedUS$1.32US$3.96
Mistral AIMistral Medium 3.5256K tokensYesEU regionUS$1.50US$7.50
Alibaba (Qwen)Qwen3.8-Max1M tokensYesEU regionUS$2.00US$6.00
Z.aiGLM-5.31M tokensNoUS$1.40US$4.40
Moonshot AIKimi K31.05M tokensYesself-hostedUS$3.00US$15.00
CohereCommand A+128K tokensYesself-hostednot publishednot published
NVIDIANemotron 3 Ultra1M tokensYesself-hostednot publishednot published

See the full landscape, including the tier below and the models being retired

03The research behind the numbers

Three pages, one question

Nobody reads a research report end to end. The charts above are the summary; these pages are where each number is argued, sourced and dated.

Planned November 2026 / no results yet

AI Board Model Index, first edition

The ranking lands on this page. There are no results yet and no scores anywhere on this site, because nothing has been measured. When the first edition is published, the leaderboard opens here and the method page becomes its appendix.

04For your role

What this means at your seat

The same landscape looks different from three chairs. Each role has a section in the guide and a page on what an AI assistant does in that job.

CEO

Synthesis across a messy pile of documents, and an answer you can repeat in a meeting without checking it first.

What a CEO should pick

AI CEO

CFO

Citation discipline. A number without a document and a page number is an opinion, however confident it sounds.

What a CFO should pick

AI CFO

CTO

Residency, licences and the exit route. The generation you can host in Europe is not always the newest one.

What a CTO should pick

AI CTO

05Method

How we work

This page follows the same discipline as the rest of the site: nothing is published that cannot be checked.

Primary sources only
Every fact about a model comes from vendor documentation, an official launch post, a system card, a peer-reviewed paper, or a cloud provider's own region list. Never from a summary of a summary.
Facts as of a stated date
Model specifications change weekly. Every chart prints the date its figures were verified, so you can see how old a claim is before you act on it.
What we cannot verify, we leave out
If a figure has no primary source, it does not appear on the page, not as an estimate and not as a range. A shorter page is the price of a checkable one.
No vendor money
No model vendor pays for a place in this research, and none of them sees a page before it is published.
Corrections
Found something that is wrong or out of date? Tell us and we correct the page, with the change dated. Contact details are on the about page.

Become AI-native before your competition does

Ride the AI wave instead of swimming behind it. Request access and we'll schedule your install.

Your company brain lives on your laptop. You choose if a question goes to a cloud model or stays fully local.