Skip to content
AI Board

Research · September 2026

AI Models for Executives 2026

Every current model, the four things that changed since 2023, and the honest answer on European hosting. Written for the person who has to decide whether an AI may read the board pack, not for the person who has to deploy it.

  • Published
  • 22 min read
  • Part of the AI Board research programme

Published context windows, 2022 to 2026

Running maximum of the vendor-stated context window, logarithmic scale. Sources: vendor documentation and launch posts, accessed 8 September 2026.

Final point is Meta's own claim for Llama 4 Scout, supported by a single retrieval-style test, not an independent measurement.

Every step in the largest published context window between November 2022 and September 2026.
DateModelPublished context window
30 Nov 2022ChatGPT (GPT-3.5)4.1K
14 Mar 2023GPT-48.19K
11 May 2023Claude (100K)100K
6 Nov 2023GPT-4 Turbo128K
21 Nov 2023Claude 2.1200K
27 Jun 2024Gemini 1.5 Pro (2M)2M
5 Apr 2025Llama 4 Scout10M
8labs shipping a current flagship
39current models tracked here
28with a route to EU data residency

EU figure carries residency caveats. Section 06 works through them model by model.

01Introduction

Why this report exists

Every week another model is announced, and every announcement arrives with a benchmark table that means very little to a director. This report does something else. It maps the model landscape as it stood on 8 September 2026, in the language of someone who has to decide whether an AI may read the board pack.

Three terms carry the whole report, so they are defined once, here. A token is the unit a model reads text in, roughly three quarters of an English word. A context window is how much text a model can hold in view at one time, counted in tokens; throughout this report one twenty page board paper is taken as about 12,000 tokens, an assumption stated so it can be argued with. A flagship is the strongest model a lab sells as of September 2026, as distinct from the cheaper tiers underneath it that most organisations actually use day to day.

It answers four questions.

  1. Which models exist, which are current, and which are being retired under you?
  2. What changed between 2023 and 2026 that genuinely matters for executive work?
  3. Which models can be run with data residency inside the EU, and at what cost?
  4. Which model belongs in which situation, for a chief executive, a finance director and a technology lead?

The scored companion to this map is the AI Board Model Index, first edition November 2026. If you only need the practical answer for one role, start with which AI model for which role.

The one-paragraph summary

In September 2026, 8 labs profiled in this report ship a current flagship model: OpenAI, Anthropic, Google, xAI, Meta, DeepSeek, Mistral and Alibaba. Four more, Cohere, Moonshot AI, NVIDIA and Z.ai, appear in the model list without a full lab profile. Between them this report tracks 39 current models. Three things changed since 2023 that a board should know about, and one thing did not.

First, memory. GPT-4 arrived in March 2023 with a context window of 8,192 tokens, less than one board paper. As of 8 September 2026 OpenAI documents GPT-6 Astra at 1,050,000 tokens, Anthropic documents Claude Fable 5.1 at 1,000,000 and Google documents Gemini 3.8 Flash at 1,048,576. On this report's assumption that is roughly 87 board papers held in view at once. Two qualifications belong beside that number. The ceiling has not moved since 27 June 2024, when Google opened a two million token window on Gemini 1.5 Pro to every developer. And published research is blunt about the gap between advertised and usable: the RULER study (Hsieh and colleagues, 2024) found that only half of the models claiming 32,000 tokens or more still performed satisfactorily at 32,000.

Second, reasoning. Since September 2024 every lab has shipped models that work through a problem before answering, and by 2026 that behaviour is built into the flagships rather than sold as a separate product. Vendors report substantial falls in confident factual errors between their own generations. Those are vendor figures on vendor prompts, so read them as direction rather than as measurement. Section 03 sets out exactly what each lab published, and what it did not.

Third, sovereignty. For the first time, models strong enough for board work can be run with data residency inside the European Union: 28 of the current models tracked here have a route the vendor documents, counting self-hosting and European providers. The detail that matters is which flagships do not. GPT-6 Astra has no EU route, Claude Fable 5.1 and Gemini 3.8 Flash have one each, and the tier just below them is where most European organisations will land. Section 06 works through the routes and the EU AI Act dates, which moved in 2026.

What did not change: every model still occasionally states something that is not in your documents, confidently and in your house style. No release in four years has removed that. It is why this report ends in a test you can run rather than a brand we recommend, and why measurement, not marketing, should decide the question.

02The landscape at a glance

Every current flagship, September 2026

13 flagship models from 12 organisations, as their own documentation described them on 8 September 2026. Of those, 8 have a documented route to running inside the European Union, 5 publish their weights, and every single one reasons before it answers.

Stated memory per flagship
  • GPT-6 Astra1.05M tokens, about 87 board papers
  • Gemini 3.8 Flash1.05M tokens, about 87 board papers
  • Kimi K31.05M tokens, about 87 board papers
  • Claude Fable 5.11M tokens, about 83 board papers
  • Claude Mythos 5.11M tokens, about 83 board papers
  • Muse Spark 1.31M tokens, about 83 board papers
  • DeepSeek-V4-Pro1M tokens, about 83 board papers
  • Qwen3.8-Max1M tokens, about 83 board papers
  • GLM-5.31M tokens, about 83 board papers
  • Nemotron 3 Ultra1M tokens, about 83 board papers
  • Grok 4.6500K tokens, about 41 board papers
  • Mistral Medium 3.5256K tokens, about 21 board papers
  • Command A+128K tokens, about 10 board papers
Vendor-stated context windows for the current flagship models. The largest is GPT-6 Astra at 1,050,000 tokens, about 87 board papers. The smallest is Command A+ at 128,000 tokens.
LabelValue
GPT-6 Astra1.05M tokens, about 87 board papers
Gemini 3.8 Flash1.05M tokens, about 87 board papers
Kimi K31.05M tokens, about 87 board papers
Claude Fable 5.11M tokens, about 83 board papers
Claude Mythos 5.11M tokens, about 83 board papers
Muse Spark 1.31M tokens, about 83 board papers
DeepSeek-V4-Pro1M tokens, about 83 board papers
Qwen3.8-Max1M tokens, about 83 board papers
GLM-5.31M tokens, about 83 board papers
Nemotron 3 Ultra1M tokens, about 83 board papers
Grok 4.6500K tokens, about 41 board papers
Mistral Medium 3.5256K tokens, about 21 board papers
Command A+128K tokens, about 10 board papers
Vendor-stated context window per current flagship, in tokens. Source: each lab's own model documentation, accessed 8 September 2026. This is a capacity, not a measurement of how well the model uses it.
List price per million input tokens
  • Gemini 3.8 Flash$0.75 in, $3.75 out
  • Muse Spark 1.3$1.25 in, $4.25 out
  • DeepSeek-V4-Pro$1.32 in, $3.96 out
  • GLM-5.3$1.40 in, $4.40 out
  • Mistral Medium 3.5$1.50 in, $7.50 out
  • Grok 4.6$2 in, $6 out
  • Qwen3.8-Max$2 in, $6 out
  • Kimi K3$3 in, $15 out
  • GPT-6 Astra$10 in, $50 out
  • Claude Fable 5.1$10 in, $50 out
  • Claude Mythos 5.1$10 in, $50 out
List prices for a million input tokens run from $0.75 for Gemini 3.8 Flash to $10 for the most expensive flagships, a spread of about 13 times.
LabelValue
Gemini 3.8 Flash$0.75 in, $3.75 out
Muse Spark 1.3$1.25 in, $4.25 out
DeepSeek-V4-Pro$1.32 in, $3.96 out
GLM-5.3$1.40 in, $4.40 out
Mistral Medium 3.5$1.50 in, $7.50 out
Grok 4.6$2 in, $6 out
Qwen3.8-Max$2 in, $6 out
Kimi K3$3 in, $15 out
GPT-6 Astra$10 in, $50 out
Claude Fable 5.1$10 in, $50 out
Claude Mythos 5.1$10 in, $50 out
Published API list price for one million input tokens, in US dollars, sorted from cheapest. Command A+, Nemotron 3 Ultra: no per-token list price published. Introductory rates, off-peak rates and long-context surcharges are in the per-model notes in section 04. Source: each lab's own pricing page, accessed 8 September 2026.

The table

Current flagship models per lab, as of 8 September 2026. Every field comes from the lab's own documentation. Nothing in this table is scored.
LabFlagshipReleasedStated memoryReasoningEU residencyOpen weightsList price
OpenAIGPT-6 Astra3 Sept 20261,050,000 tokensabout 87 board papersYesNoLower tiers with a documented EU route: GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna.No$10 in$50 outper million tokens
AnthropicClaude Fable 5.11 Sept 20261,000,000 tokensabout 83 board papersYesYesDocumented routes: Google Cloud Agent Platform (Vertex AI), EU multi-region endpointNo$10 in$50 outper million tokens
AnthropicClaude Mythos 5.11 Sept 20261,000,000 tokensabout 83 board papersYesNoLower tiers with a documented EU route: Claude Fable 5.1, Claude Opus 5, Claude Sonnet 5.No$10 in$50 outper million tokens
Google (DeepMind)Gemini 3.8 Flash2 Sept 20261,048,576 tokensabout 87 board papersYesYesDocumented routes: Google Cloud Vertex AI, eu multi-region only (no EU regional endpoint)No$0.75 in$3.75 outper million tokens
xAIGrok 4.612 Aug 2026500,000 tokensabout 41 board papersYesNoNo model in this lab's current line-up has a documented EU route.No$2 in$6 outper million tokens
MetaMuse Spark 1.32 Sept 20261,000,000 tokensabout 83 board papersYesNoLower tiers with a documented EU route: Muse Glimmer 30B.No$1.25 in$4.25 outper million tokens
DeepSeekDeepSeek-V4-Pro13 Aug 20261,000,000 tokensabout 83 board papersYesYesDocumented routes: Self-hosted on your own or a rented GPU cluster (weights are public under MIT); European GPU rental with a data processing agreement, for example an EU region of a sovereign cloudYesMIT$1.32 in$3.96 outper million tokens
Mistral AIMistral Medium 3.5Apr 2026256,000 tokensabout 21 board papersYesYesDocumented routes: Mistral La Plateforme with Mistral Regional Endpoints set to Europe; Scaleway Generative APIs (France), served as mistral-medium-3.5-128b; Self-hosted with vLLM, TensorRT-LLM or TGI on your own hardware; Microsoft Azure AI, AWS Bedrock, Google Cloud Vertex AI, IBM watsonx.ai and OutscalePartialModified MIT License: weights are downloadable, but commercial use carries exceptions for companies above a revenue threshold, so it is not a plain open-source licence$1.50 in$7.50 outper million tokens
Alibaba (Qwen)Qwen3.8-Max3 Aug 20261,000,000 tokensabout 83 board papersYesYesDocumented routes: Alibaba Cloud Model Studio, Germany (Frankfurt) region eu-central-1No$2 in$6 outper million tokens
Z.aiGLM-5.318 Aug 20261,000,000 tokensabout 83 board papersYesNoLower tiers with a documented EU route: GLM-5.2.No$1.40 in$4.40 outper million tokens
Moonshot AIKimi K3Jul 20261,048,576 tokensabout 87 board papersYesYesDocumented routes: Self-hosted on your own or a rented GPU cluster, under the Kimi K3 LicenseYesKimi K3 License (custom; commercial inference above roughly USD 20 million of annual revenue carries revenue-sharing obligations)$3 in$15 outper million tokens
CohereCommand A+20 May 2026128,000 tokensabout 10 board papersYesYesDocumented routes: Cohere private deployment inside your own virtual private cloud; On-premises deployment on 1 x B200 or 2 x H100 at four-bit weights and activations; Self-hosted from the Apache-2.0 weights on Hugging FaceYesApache-2.0No per-token list price published
NVIDIANemotron 3 Ultra4 Jun 20261,000,000 tokensabout 83 board papersYesYesDocumented routes: Self-hosted on your own or a rented GPU cluster (OpenMDW-1.1); Any European GPU provider willing to run open weights for youYesOpenMDW-1.1 (Open Model, Weights and Data License, version 1.1)No per-token list price published

How to read the table

Every row is the model a lab put at the top of its own range on 8 September 2026, taken from that lab's own documentation. Nothing here is scored. A specification table tells you what a model is allowed to do, not how well it does it, and the distance between those two things is the reason the AI Board Model Index exists.

Why stated memory is not working memory

Memory is the most quoted number in this market and the least reliable guide to what a model will do with your board pack. It is a capacity, not a competence: the vendor is telling you how much text the model will accept, not how much of it the model is still using correctly at page seven hundred.

The evidence for that gap is published research, not marketing. According to RULER (Hsieh et al., COLM 2024), of seventeen models that all claimed a window of thirty-two thousand tokens or more, only half still performed acceptably at exactly that length. NoLiMa (Modarressi et al., ICML 2025) is blunter: of thirteen models advertising at least a hundred and twenty-eight thousand tokens, eleven scored below half of their own short-text result once the input reached thirty-two thousand, and GPT-4o, one of the strongest in that study, still fell from 99.3 percent to 69.7 percent. Liu et al. (Transactions of the ACL, 2024) add why the failures feel arbitrary: accuracy is highest when the relevant passage sits at the very start or the very end of the input, and drops in the middle.

The EU column is the one most often misread

Three questions hide inside the words EU hosting. Where the computation physically happens. Who the contracting party is, because an EU region of an American cloud satisfies residency and still leaves an American parent company in the chain. And availability per model, which is the only one this column tracks.

What changed in the last ninety days

Most of this table did not exist ninety days ago. Anthropic shipped Sonnet 5 on 30 June 2026 with the flagship's million-token window at a fifth of the price, Opus 5 on 24 July, and Fable 5.1 on 1 September. OpenAI replaced version numbers with named tiers on 9 July and unveiled GPT-6 Astra on 3 September. DeepSeek shipped V4-Pro with public weights under an MIT licence on 13 August, and Google made Gemini 3.8 Flash generally available on 2 September while its Pro slot stayed in preview.

03The four shifts

What changed, 2023 to 2026: the four shifts that matter

Four numbers moved between March 2023 and September 2026, and each one changes a different decision: memory, reasoning, sovereignty and price. Every chart below carries the rule used to select its data and the published evidence that argues against it.

Token and context window are defined once, in the introduction. Where this section turns tokens into pages it uses one stated assumption: a twenty page board paper of about 9,000 words is 12,000 tokens. Read the four shifts in order. They are not four versions of the same good news.

01

Memory: from half a board paper to a full year of them

In one sentenceThe flagship context window went from 8,192 tokens in March 2023 to 1,050,000 in September 2026, but the usable part is smaller than the advertised part.

In March 2023 GPT-4 arrived with a context window of 8,192 tokens, a figure OpenAI still publishes on its model documentation page. Take a twenty page board paper as roughly 9,000 words, which is about 12,000 tokens. That is the assumption this report uses throughout, and it means GPT-4 could not hold one board paper in view at once. It held about two thirds of one.

What this means for the board

  • Fit is no longer the constraint. A year of minutes, the annual report and the budget will fit in one conversation with any current flagship, so stop scoping AI work around document size.
  • Ask for the usable number, not the advertised one. Published research shows accuracy falling well before the stated limit, so treat the window as a ceiling and require evidence at the length you actually use.

Context window of OpenAI's flagship, March 2023 to September 2026

Logarithmic line chart. The context window of OpenAI's flagship model rose from 8,192 tokens in March 2023 to 1,050,000 tokens in September 2026, with a fall to 400,000 tokens for GPT-5 in August 2025.
Pointseries
2023 H18.19K
2023 H2128K
2024 H1128K
2024 H2128K
2025 H11.05M
2025 H2400K
2026 H11.05M
2026 H21.05M
Logarithmic scale. Points, in order: GPT-4, GPT-4 Turbo, GPT-4o, GPT-4o, GPT-4.1, GPT-5, GPT-5.6 Sol, GPT-6 Astra. Source: OpenAI model documentation, accessed 8 September 2026.

Board papers: 1 = 12,000 tokens.

How this was counted

One lineage is plotted, OpenAI's flagship, because it is the single line most boards have actually used and every point is on OpenAI's own model pages. Rival flagships tracked closely and are named in the text. The board paper series divides the window by 12,000 tokens, being a twenty page paper of about 9,000 words at roughly 1.33 tokens per word. Token counts differ per tokenizer, so read the paper counts as an order of magnitude, not a measurement.

Evidence that argues the other way

  • Of the models claiming a context window of 32,000 tokens or more, only half maintained satisfactory performance at 32,000 tokens.

    RULER: What's the Real Context Size of Your Long-Context Language Models? (arXiv:2404.06654)

  • On the NoLiMa evaluation, which removes literal word overlap between question and answer, GPT-4o fell from 99.3 percent at short context to 69.7 percent at 32,000 tokens, and eleven of thirteen models tested dropped below half their own short-context score.

    NoLiMa: Long-Context Evaluation Beyond Literal Matching (arXiv:2502.05167)

  • The largest generally available context window has not moved since 27 June 2024, when Google opened two million tokens on Gemini 1.5 Pro to all developers. Every verified frontier flagship in September 2026 sits at one million to 1.05 million.

    Google, new features for the Gemini API and Google AI Studio. OpenAI model documentation. Anthropic model overview. Gemini API: Gemini 3.8 Flash

02

Reasoning: fewer confident mistakes, not zero mistakes

In one sentenceOn OpenAI's own production-traffic evaluation the share of answers containing a major factual error fell from 20.6 percent for GPT-4o to 4.8 percent for GPT-5 in thinking mode, while some reasoning models moved the other way.

One term first. A reasoning model is one that produces a chain of internal working before it answers, rather than answering in one pass. OpenAI shipped the first widely available one, o1, in December 2024. By 2026 the behaviour is built into every flagship and can be dialled up or down per question.

What this means for the board

  • Turn thinking on for anything that goes into a board pack. The same family of models differs by a factor of four in major-error rate between its fast mode and its thinking mode.
  • Test the model on your own documents before you trust it on them. Reasoning that helps on open questions can hurt on grounded summarising, and only your own material will tell you which case you are in.

Answers containing at least one major factual error

  • GPT-4o (May 2024)20.6%
  • o3 (Apr 2025)22.0%
  • GPT-5 main (Aug 2025)11.6%
  • GPT-5 thinking (Aug 2025)4.8%
Bar chart. On OpenAI's production-representative prompts with browsing enabled, the share of answers containing at least one major factual error was 20.6 percent for GPT-4o, 22.0 percent for o3, 11.6 percent for GPT-5 and 4.8 percent for GPT-5 in thinking mode.
LabelValue
GPT-4o (May 2024)20.6%
o3 (Apr 2025)22.0%
GPT-5 main (Aug 2025)11.6%
GPT-5 thinking (Aug 2025)4.8%
Lower is better. One evaluation, one document, so the four bars are comparable with each other. Source: OpenAI, GPT-5 System Card, 13 August 2025, section 3.7.

The counter-example: PersonQA hallucination rate across OpenAI's reasoning models

  • o1 (Dec 2024)16%
  • o3 (Apr 2025)33%
  • o4-mini (Apr 2025)48%
Bar chart. On PersonQA the hallucination rate rose from 16 percent for o1 to 33 percent for o3 and 48 percent for o4-mini, so a newer reasoning model was the more hallucinatory one.
LabelValue
o1 (Dec 2024)16%
o3 (Apr 2025)33%
o4-mini (Apr 2025)48%
Lower is better. A different evaluation from the chart above, so the two are not comparable. Source: OpenAI, o3 and o4-mini System Card, 16 April 2025, table 4, and the o1 System Card of 5 December 2024.

How this was counted

The main series comes from one evaluation in one document, the GPT-5 system card, so the four bars are comparable with each other. Numbers from different vendors or different evaluations are not comparable and are kept in separate series. The Vectara leaderboard measures grounded summarising only, and its own methodology changed between 2023 and 2026, so it is shown as a snapshot of one date rather than as a trend line.

Evidence that argues the other way

  • On PersonQA, OpenAI's o3 hallucinated at 0.33 and o4-mini at 0.48 against 0.16 for the older o1. o3 was the more accurate model and the more hallucinatory one at the same time.

    OpenAI, o3 and o4-mini System Card (16 April 2025). OpenAI, o1 System Card (5 December 2024)

  • Anthropic reports that Claude Opus 5 is 11 percent more accurate than Claude Opus 4.8 and that its rate of hallucinations is 6 percent higher.

    Anthropic, Claude Opus 5 System Card (24 July 2026)

  • Extending the reasoning length of a reasoning model can make its performance worse, an inverse scaling relationship reported across five distinct failure modes.

    Inverse Scaling in Test-Time Compute

  • On Google DeepMind's FACTS benchmark suite every model evaluated scored below 70 percent overall, including the best of its own generation.

    Google DeepMind, FACTS benchmark suite

03

Sovereignty: from one EU option to eight, with the newest model still outside

In one sentenceIn 2023 one frontier-class model could be run with EU data residency. In September 2026 eight can, but OpenAI's newest flagship, GPT-6 Astra, is not one of them.

As of September 2026 the count is eight: five hosted with a residency guarantee, being Claude Fable 5.1, Claude Opus 5, Claude Sonnet 5, GPT-5.6 Sol and Mistral Medium 3.5, and three self-hostable, being Mistral Large 3, DeepSeek-V4-Pro and the open Qwen3.8 release. Then the finding that changes a decision. GPT-6 Astra, released on 3 September 2026 and OpenAI's most capable model, appears in Microsoft's Americas and Global tables and in none of the Europe tables. GPT-5.6 Sol is in nine European regions. If your policy is EU residency, your OpenAI ceiling as of September 2026 is the previous generation. Grok 4.6 and Meta's Muse Spark 1.3 have no verified EU route at all.

What this means for the board

  • EU residency is now a real choice, not a compromise you have to explain. Eight frontier-class models can be run inside the EU, four of them from the top of a major lab's range.
  • Check the specific model, not the vendor. OpenAI's newest flagship is not available with EU residency as of September 2026 while its previous generation is in nine European regions, and the same gap has appeared before with Claude 3 Opus.

Frontier-class models that can be run inside the EU, per year

  • 20231
  • 20244
  • 20256
  • 20268
Bar chart. The number of frontier-class models that can be run inside the EU rose from one in 2023 to four in 2024, six in 2025 and eight in September 2026.
LabelValue
20231
20244
20256
20268
Counted against the definition below. Every model behind every number is named in the roster, so the count can be checked. Google Gemini is counted in no year: no Google source names its EU regions with the models available in them.

Route one: hosted with an EU data residency guarantee

  • 20231
  • 20243
  • 20254
  • 20265
Bar chart. Models hosted with an EU data residency guarantee rose from one in 2023 to three in 2024, four in 2025 and five in September 2026.
LabelValue
20231
20243
20254
20265
The vendor or a cloud provider commits to processing inside EU regions.

Route two: open weights, self-hostable on EU hardware

  • 20230
  • 20241
  • 20253
  • 20263
Bar chart. Models with open weights that a company can run on its own EU hardware rose from none in 2023 to one in 2024 and three in both 2025 and September 2026.
LabelValue
20230
20241
20253
20263
The two routes do not add up to the total: a model available by both routes, such as Mistral Large 3 in 2025 and 2026, is counted once in the total.

How this was counted

The two definitions are set out beside the chart, and every counted model is named here so the count can be checked. 2023, hosted: GPT-4 (Azure OpenAI, EU regions). 2024, hosted: GPT-4o (Azure OpenAI, EU Data Zone), Claude 3.5 Sonnet (AWS Bedrock, Europe Frankfurt), Mistral Large 2 (Mistral infrastructure in Europe); self-hostable: Llama 3.1 405B (open weights). 2025, hosted: GPT-5 (Azure OpenAI, EU Data Zone), Claude 3.7 Sonnet (AWS Bedrock, EU cross-region inference), Claude Sonnet 4.5 (AWS Bedrock, regional endpoints), Mistral Large 3 (Mistral infrastructure in Europe); self-hostable: Mistral Large 3 (Apache 2.0), DeepSeek-R1 (MIT licence), Qwen 3 (Apache 2.0). 2026, hosted: Claude Fable 5.1 (Google Cloud eu multi-region only), Claude Opus 5 (AWS Bedrock, Frankfurt and Ireland), Claude Sonnet 5 (AWS Bedrock, Frankfurt and Ireland), GPT-5.6 Sol (Azure OpenAI, EU Data Zone), Mistral Medium 3.5 (Mistral regional endpoints); self-hostable: Mistral Large 3 (Apache 2.0), DeepSeek-V4-Pro (published weights), Qwen3.8-2.4T-A95B (open release under the Qwen licence). A model counted in both routes, such as Mistral Large 3, is counted once in the total. Claude 2 on Amazon Bedrock in Frankfurt from October 2023 is excluded because the AWS announcement names the vendor and not the model version. Google Gemini is excluded from every year because Google's residency documentation names the models with EU processing but not the date each one got it, so no Gemini model can be placed in a year

Evidence that argues the other way

  • GPT-6 Astra, OpenAI's most capable model as of September 2026, does not appear in any Europe data zone table in Microsoft's model region availability documentation, while GPT-5.6 Sol appears in nine European regions.

    Microsoft Foundry: models sold directly by Azure. OpenAI model documentation

  • A lab's flagship is not automatically the best model an EU customer can reach. Claude 3 Opus, Anthropic's 2024 flagship, never reached Amazon Bedrock in Frankfurt; only Claude 3 Sonnet and Haiku did.

    AWS, Claude 3.5 Sonnet and Haiku available in more regions. Anthropic, Claude in Amazon Bedrock

  • Data residency describes where data is processed, not who can be compelled to produce it. The EU regions of Microsoft, Amazon and Google are operated by United States companies.

    Microsoft, announcing Azure OpenAI Data Zones. Anthropic, Claude in Amazon Bedrock

  • The high-risk obligations of the AI Act were deferred by Regulation (EU) 2026/1744, in force 27 July 2026, to 2 December 2027 for Annex III systems and 2 August 2028 for Annex I systems. The 2 August 2026 general application date was not moved.

    Regulation (EU) 2026/1744, the Digital Omnibus on AI. Regulation (EU) 2024/1689, the AI Act. European Commission, the Commission starts enforcing AI Act rules

04

Price: a fixed capability got 300 times cheaper, the frontier got dearer again

In one sentenceGPT-4 quality cost $30 per million input tokens in March 2023 and $0.10 by February 2025, while the newest flagship of September 2026 costs $10 again.

There are two price stories and they run in opposite directions, so a single line would be a false claim. The first story is the cost of a fixed capability. GPT-4 launched in March 2023 at $30 per million input tokens, a figure still on OpenAI's own model page and on the archived pricing page of 1 April 2023. By November 2023 GPT-4 Turbo delivered comparable quality at $10. Epoch AI, which tracks the cheapest model reaching each capability milestone, places Gemini 1.5 Pro at that level once Google repriced it to $1.25, and Gemini 2.0 Flash at $0.10 in February 2025. That is a fall of about 300 times in under two years for the same job, or roughly five to six times a year.

What this means for the board

  • Budget the job, not the model. The price of doing a known task fell about 300 times between March 2023 and February 2025, so a workload that was uneconomic two years ago probably is not now.
  • Do not write always latest model into a policy. The newest flagship cost eight times more per million input tokens in September 2026 than the flagship of a year earlier.

What GPT-4 class quality cost, per million input tokens

Logarithmic line chart. The cheapest published price for GPT-4 class quality fell from 30 US dollars per million input tokens in March 2023 to 10 dollars in November 2023, 1.25 dollars in 2024 and 0.10 dollars in February 2025.
Pointseries
Mar 2023US$30.00
Nov 2023US$10.00
2024US$1.25
2025US$0.10
Logarithmic scale. Points, in order: GPT-4, GPT-4 Turbo, Gemini 1.5 Pro, Gemini 2.0 Flash. List prices per million input tokens, standard tier, no batch or cache discount.

Why the line stops in 2025

No 2026 point is plotted on the fixed capability line. The cheapest 2026 candidates, GPT-5.6 Luna at $0.20 and DeepSeek V4-Flash at $0.22 per million input tokens, are documented prices, but we could not verify against a primary source that either matches or beats the March 2023 GPT-4 on the benchmark Epoch AI used. Publishing that point would be a guess, so it is left out.

The other direction: OpenAI's flagship at the end of each year

  • 2023 · GPT-4 TurboUS$10.00
  • 2024 · GPT-4oUS$2.50
  • 2025 · GPT-5US$1.25
  • 2026 · GPT-6 AstraUS$10.00
Bar chart. OpenAI's flagship input price ran 10 US dollars per million tokens in 2023, 2.50 in 2024, 1.25 in 2025 and 10 again in September 2026.
LabelValue
2023 · GPT-4 TurboUS$10.00
2024 · GPT-4oUS$2.50
2025 · GPT-5US$1.25
2026 · GPT-6 AstraUS$10.00
Anthropic's top of range moved the same way, from 8 dollars for Claude 2.1 to 15 for Claude 3 Opus, 5 for Claude Opus 4.5 and 10 for Claude Fable 5.1. Google's Pro tier ran 3.50, then 1.25, then 2.00.

How this was counted

All values are published list prices per million input tokens, standard tier, no batch or cache discount. Output prices are excluded because a board pack workload is almost all input. Per lab the rule is the model the lab presented as the top of its range at the end of that year: for Anthropic that is the Claude 2, Opus and Fable line, for Google the Pro tier rather than the Flash tier. The fixed capability line follows Epoch AI's selection of the cheapest model reaching GPT-4's benchmark score, with each price re-verified against the vendor or archived vendor page and restated as an input price rather than Epoch's blended input and output price.

Evidence that argues the other way

  • OpenAI's flagship input price rose from $1.25 per million tokens for GPT-5 in August 2025 to $10 for GPT-6 Astra in September 2026, an eightfold increase.

    OpenAI API pricing. OpenAI model documentation, GPT-5

  • Anthropic states that the tokenizer used from Claude Opus 4.7 onward produces about 30 percent more tokens for the same text, so an unchanged price per token is a higher price per document.

    Anthropic pricing. Anthropic model overview

  • Published estimates of how fast a fixed capability gets cheaper range from 9 to 900 times per year depending on which benchmark is held fixed, so any single multiple should be treated as an illustration.

    Epoch AI, LLM inference price trends. Stanford HAI, 2025 AI Index Report

What has not changed

Four things moved. Six did not. Each statement below was true of the models of 2023 and is still true of the models of September 2026, and each is measured rather than asserted.

Models still state things that are not in your documents. On OpenAI's own production-representative evaluation, the best configuration of GPT-5 still produced a major factual error in 4.8 percent of answers, and Anthropic wrote in the Claude Opus 4.5 system card that it is still far from removing factual hallucinations in the absence of external tools.

OpenAI, GPT-5 System Card (13 August 2025). Anthropic, Claude Opus 4.5 System Card (November 2025)

A newer model is not automatically a more truthful one. OpenAI measured a PersonQA hallucination rate of 0.16 for o1 and 0.33 for the later o3, and Anthropic reports that Claude Opus 5 is 11 percent more accurate than Claude Opus 4.8 while hallucinating 6 percent more often.

OpenAI, o1 System Card (5 December 2024). OpenAI, o3 and o4-mini System Card (16 April 2025). Anthropic, Claude Opus 5 System Card (24 July 2026)

Instructions drift over a long conversation. Across more than 200,000 simulated conversations with fifteen models, researchers at Microsoft and Salesforce measured an average 39 percent drop in performance in multi-turn conversations against the same tasks asked in one turn, and found that once a model takes a wrong turn it does not recover.

Laban, Hayashi, Zhou & Neville, LLMs Get Lost In Multi-Turn Conversation (arXiv:2505.06120)

The same question does not reliably give the same answer. In a controlled test, 1,000 identical requests at temperature zero produced 80 different completions, with the divergence starting at token 103. OpenAI's own API reference says its seed parameter is best effort and that determinism is not guaranteed, and Anthropic has removed the temperature, top-p and top-k controls entirely from Claude Opus 4.7 and later.

Thinking Machines Lab, Defeating Nondeterminism in LLM Inference. OpenAI API reference, seed parameter. Anthropic model deprecations

Vendors retire the model you standardised on. Named production models have run 12 to 22 months before shutdown: Claude 3.7 Sonnet 12 months, Claude Opus 4.1 12 months, Claude Sonnet 4 and Opus 4 13 months, Claude 3.5 Sonnet 16 months, GPT-5 about 16 months to its 11 December 2026 shutdown, Gemini 2.0 Flash about 16 months. Anthropic commits to at least 60 days of notice and OpenAI has been giving about six months.

Anthropic model deprecations. OpenAI model deprecations. Google, Gemini API model deprecations and changelog

A bigger context window is still not a bigger usable memory. Three independent studies since 2024 have found accuracy falling well before the advertised limit, and no vendor publishes a usable-length figure to sit beside the advertised one.

RULER: What's the Real Context Size of Your Long-Context Language Models? (arXiv:2404.06654). NoLiMa: Long-Context Evaluation Beyond Literal Matching (arXiv:2502.05167). Chroma, Context Rot: How Increasing Input Tokens Impacts LLM Performance

04The vendors

Eight labs, lab by lab

Eight labs ship a model a management team could reasonably point at its own documents. Here is what each one actually sells in September 2026, what it costs, and where it may be run.

Read every block the same way. What the lab is, how its line of models has moved, and the current line-up as a table with the vendor's own note beside every EU route and every price. The order runs from the labs most European boards already use to the ones they are most likely to be arguing about.

OpenAI

San Francisco, California · US

The lab behind ChatGPT and the GPT model line, and the reason most boards have an opinion about AI at all. Its models reach executives through ChatGPT workspaces, the OpenAI API and Microsoft Foundry, which is where most European data-residency questions end up.

Generations

GPT-4 family14 Mar 2023 to 23 Oct 2026 · 5 modelso-series reasoning models17 Dec 2024 to 11 Dec 2026 · 6 modelsGPT-5 family7 Aug 2025 to now · 11 modelsGPT-6 family3 Sept 2026 to now · 1 model

Models, oldest to newest

  1. GPT-413 Jun 2023end of service 23 Oct 2026Memory: not published
  2. GPT-3.5 Turbo25 Jan 2024end of service 23 Oct 2026Memory: not published
  3. o117 Dec 2024end of service 23 Oct 2026Memory: not published
  4. o316 Apr 2025end of service 11 Dec 2026Memory: not published
  5. o4-mini16 Apr 2025end of service 23 Oct 2026Memory: not published
  6. GPT-57 Aug 2025end of service 11 Dec 2026Memory: not published
  7. GPT-5.6 Sol9 Jul 2026Memory: 1.05M tokens
  8. GPT-5.6 Terra9 Jul 2026Memory: 1.05M tokens
  9. GPT-5.6 Luna9 Jul 2026Memory: 1.05M tokens
  10. GPT-6 Astra3 Sept 2026Memory: 1.05M tokens
  • current
  • retired or retiring
  • unverified
Model lineage for OpenAI: 4 current, 6 retired or retiring, each shown with the context window its vendor publishes.
ModelReleasedRoleMemory
GPT-413 Jun 2023retired or retiringnot published
GPT-3.5 Turbo25 Jan 2024retired or retiringnot published
o117 Dec 2024retired or retiringnot published
o316 Apr 2025retired or retiringnot published
o4-mini16 Apr 2025retired or retiringnot published
GPT-57 Aug 2025retired or retiringnot published
GPT-5.6 Sol9 Jul 2026current1.05M tokens
GPT-5.6 Terra9 Jul 2026current1.05M tokens
GPT-5.6 Luna9 Jul 2026current1.05M tokens
GPT-6 Astra3 Sept 2026current1.05M tokens
Bar length is the published context window, on one scale across all eight labs. Source: vendor documentation, read 8 September 2026.
Current line-up: OpenAI, 8 Sept 2026.
ModelRoleReleasedMemoryEU residencyPrice per 1M tokens (in / out)
GPT-6 AstrareasoningFlagship3 Sept 20261.05M tokensNo EU routeMicrosoft Foundry offers Astra in Global and US Data Zone deployments only. Microsoft's own launch post lists no EU Data Zone, so an organisation that requires European inference stays on an older tier for now.US$10 / US$50OpenAI's current flagship. Prompts over 272K tokens are billed at a long-context rate of $20 input and $75 output per million.
GPT-5.6 SolreasoningMid tier9 Jul 20261.05M tokensEU routeAzure OpenAI / Microsoft Foundry, EU Data Zone (9 EU regions)Microsoft's Data Zone Standard availability table (revision of 3 September 2026) lists the GPT-5.6 tiers in nine European regions, so prompts and responses can be processed inside the Azure EU Data Boundary. That table does not list GPT-6 Astra.US$4 / US$20The previous flagship, still served. OpenAI states this is promotional pricing available at least through 21 November 2026.
GPT-5.6 TerrareasoningMid tier9 Jul 20261.05M tokensEU routeAzure OpenAI / Microsoft Foundry, EU Data Zone (9 EU regions)Microsoft's Data Zone Standard availability table (revision of 3 September 2026) lists the GPT-5.6 tiers in nine European regions, so prompts and responses can be processed inside the Azure EU Data Boundary. That table does not list GPT-6 Astra.US$2 / US$12The balanced tier, the default choice for most document work.
GPT-5.6 LunareasoningSmall and fast9 Jul 20261.05M tokensEU routeAzure OpenAI / Microsoft Foundry, EU Data Zone (9 EU regions)Microsoft's Data Zone Standard availability table (revision of 3 September 2026) lists the GPT-5.6 tiers in nine European regions, so prompts and responses can be processed inside the Azure EU Data Boundary. That table does not list GPT-6 Astra.US$0.20 / US$1.20The fast, cheap tier. Fine for drafting and summarising, not for figures that go to the board without a check.

Official model documentation: OpenAI

Anthropic

San Francisco, California · US

The lab behind Claude, founded by former OpenAI researchers and known for publishing unusually detailed model, pricing and retirement documentation. Its models are sold directly and through Amazon Bedrock, Google Cloud and Microsoft Foundry, which gives a European buyer more than one contracting route.

Generations

Claude 2 generation11 Jul 2023 to 21 Jul 2025 · 2 modelsClaude 3 generation4 Mar 2024 to 20 Apr 2026 · 6 modelsClaude 4 generation14 May 2025 to now · 10 modelsClaude 5 generation9 Jun 2026 to now · 4 models

Models, oldest to newest

  1. Claude Opus 329 Feb 2024end of service 5 Jan 2026Memory: 200K tokens
  2. Claude Haiku 37 Mar 2024end of service 20 Apr 2026Memory: 200K tokens
  3. Claude Haiku 3.522 Oct 2024end of service 19 Feb 2026Memory: 200K tokens
  4. Claude Sonnet 3.719 Feb 2025end of service 19 Feb 2026Memory: 200K tokens
  5. Claude Opus 414 May 2025end of service 15 Jun 2026Memory: 200K tokens
  6. Claude Sonnet 414 May 2025end of service 15 Jun 2026Memory: 200K tokens
  7. Claude Opus 4.15 Aug 2025end of service 5 Aug 2026Memory: 200K tokens
  8. Claude Haiku 4.515 Oct 2025Memory: 200K tokens
  9. Claude Opus 4.828 May 2026Memory: 1M tokens
  10. Claude Fable 59 Jun 2026Memory: 1M tokens
  11. Claude Sonnet 530 Jun 2026Memory: 1M tokens
  12. Claude Opus 524 Jul 2026Memory: 1M tokens
  13. Claude Fable 5.11 Sept 2026Memory: 1M tokens
  14. Claude Mythos 5.11 Sept 2026Memory: 1M tokens
  • current
  • retired or retiring
  • unverified
Model lineage for Anthropic: 7 current, 7 retired or retiring, each shown with the context window its vendor publishes.
ModelReleasedRoleMemory
Claude Opus 329 Feb 2024retired or retiring200K tokens
Claude Haiku 37 Mar 2024retired or retiring200K tokens
Claude Haiku 3.522 Oct 2024retired or retiring200K tokens
Claude Sonnet 3.719 Feb 2025retired or retiring200K tokens
Claude Opus 414 May 2025retired or retiring200K tokens
Claude Sonnet 414 May 2025retired or retiring200K tokens
Claude Opus 4.15 Aug 2025retired or retiring200K tokens
Claude Haiku 4.515 Oct 2025current200K tokens
Claude Opus 4.828 May 2026current1M tokens
Claude Fable 59 Jun 2026current1M tokens
Claude Sonnet 530 Jun 2026current1M tokens
Claude Opus 524 Jul 2026current1M tokens
Claude Fable 5.11 Sept 2026current1M tokens
Claude Mythos 5.11 Sept 2026current1M tokens
Bar length is the published context window, on one scale across all eight labs. Source: vendor documentation, read 8 September 2026.
Current line-up: Anthropic, 8 Sept 2026.
ModelRoleReleasedMemoryEU residencyPrice per 1M tokens (in / out)
Claude Fable 5.1reasoningFlagship1 Sept 20261M tokensEU routeGoogle Cloud Agent Platform (Vertex AI), EU multi-region endpointThe EU route runs through Google Cloud, where a multi-region 'eu' endpoint keeps inference inside the European Union at a ten percent premium. As of 8 September 2026 Amazon Bedrock offers Fable 5.1 regional endpoints in us-east-1 only, Microsoft Foundry offers Global and US Data Zone deployments, and Anthropic's own API supports only 'us' and 'global' inference. One route, not four.US$10 / US$50Anthropic's current flagship, 1M-token window. Cache reads cost 2.5 percent of the input price, the lowest cache rate any lab publishes, which matters when the same document set is re-read every day.
Claude Mythos 5.1reasoningFlagship1 Sept 20261M tokensNo EU routeAccess is limited to verified programmes and, at launch, to organisations in the United States. Not a route a European management team can plan around.US$10 / US$50
Claude Opus 5reasoningMid tier24 Jul 20261M tokensEU routeGoogle Cloud Agent Platform (Vertex AI), EU multi-region endpoint; Amazon Bedrock EU inference profile (Frankfurt, Ireland, Paris, Milan, Madrid, Stockholm, Zurich, London)Broader European coverage than the flagship, because Bedrock's EU inference profile and Google Cloud's EU multi-region endpoint both serve it. Regional and multi-region routing carries a ten percent premium on both clouds.US$5 / US$25Opus-tier head, 1M-token window, half the token price of Fable 5.1.
Claude Sonnet 5reasoningMid tier30 Jun 20261M tokensEU routeGoogle Cloud Agent Platform (Vertex AI), EU multi-region endpoint; Amazon Bedrock EU inference profileSame European routing options as Claude Opus 5, at a fifth of the price.US$2 / US$10Mid tier with the same 1M-token window as the flagships. Anthropic confirmed in its pricing notes that the scheduled increase to $3/$15 on 1 September 2026 will not happen.
Claude Fable 5reasoningMid tier9 Jun 20261M tokensEU routeGoogle Cloud Agent Platform (Vertex AI), EU multi-region endpointUS$10 / US$50
Claude Opus 4.8reasoningMid tier28 May 20261M tokensEU routeAmazon Bedrock EU inference profile; Google Cloud Agent Platform (Vertex AI), EU multi-region endpointUS$5 / US$25
Claude Haiku 4.5reasoningSmall and fast15 Oct 2025200K tokensEU routeAmazon Bedrock EU inference profile; Google Cloud Agent Platform (Vertex AI), EU multi-region endpointUS$1 / US$5Small and fast, 200K-token window.

Official model documentation: Anthropic

Google (DeepMind)

Mountain View, California · US

Alphabet's AI arm, formed around the DeepMind research group in London, ships the Gemini line. Gemini has the widest distribution of any model family because it is built into Search, Workspace, Android and Chrome, so many organisations are already using it before anyone signs a contract.

Generations

Gemini 1.5 generation15 Feb 2024 to 29 Sept 2025 · 3 modelsGemini 2 generation5 Feb 2025 to 1 Jun 2026 · 4 modelsGemini 3 generation18 Nov 2025 to now · 7 models

Models, oldest to newest

  1. Gemini 2.0 FlashDec 2024end of service 1 Jun 2026Memory: not published(unverified)
  2. Gemini 3.5 Flash-Lite2026Memory: not published(unverified)
  3. Gemini 3.1 Pro19 Feb 2026Memory: 1.05M tokens
  4. Gemini 3.7 FlashAug 2026Memory: not published(unverified)
  5. Gemini 3.8 Flash2 Sept 2026Memory: 1.05M tokens
  • current
  • retired or retiring
  • unverified
Model lineage for Google (DeepMind): 4 current, 1 retired or retiring, each shown with the context window its vendor publishes.
ModelReleasedRoleMemory
Gemini 2.0 FlashDec 2024retired or retiringnot published
Gemini 3.5 Flash-Lite2026currentnot published
Gemini 3.1 Pro19 Feb 2026current1.05M tokens
Gemini 3.7 FlashAug 2026currentnot published
Gemini 3.8 Flash2 Sept 2026current1.05M tokens
Bar length is the published context window, on one scale across all eight labs. Source: vendor documentation, read 8 September 2026.
Current line-up: Google (DeepMind), 8 Sept 2026.
ModelRoleReleasedMemoryEU residencyPrice per 1M tokens (in / out)
Gemini 3.8 FlashreasoningFlagship2 Sept 20261.05M tokensEU routeGoogle Cloud Vertex AI, eu multi-region only (no EU regional endpoint)Google Cloud's locations documentation lists Gemini 3.8 Flash on the eu multi-region endpoint, which keeps machine learning processing inside EU member states, and on no EU regional endpoint. If your policy names one country rather than the Union, this model does not meet it yet.US$0.75 / US$3.75Google's generally available flagship is a Flash model. This is introductory pricing through 31 December 2026; from 1 January 2027 it doubles to $1.50 input and $7.50 output. Budget for the higher number.
Gemini 3.7 FlashunverifiedreasoningMid tierAug 2026not publishedNo EU routeWe could not confirm an EU data-residency deployment for this model in the vendor's own documentation on 8 September 2026. Treat European routing as something to confirm in writing before any board material is sent.US$0.75 / US$3.75
Gemini 3.1 ProreasoningMid tier19 Feb 20261.05M tokensNo EU routeGoogle Cloud's locations documentation records neither at-rest data residency nor machine learning processing in the EU for Gemini 3.1 Pro preview. The model is still published as a preview endpoint, which is worth noting in a change-management plan.US$2 / US$12
Gemini 3.5 Flash-LiteunverifiedreasoningSmall and fast2026not publishedNo EU routeWe could not confirm an EU data-residency deployment for this model in the vendor's own documentation on 8 September 2026. Treat European routing as something to confirm in writing before any board material is sent.US$0.30 / US$2.50Cheap tier in the current generation.

Official model documentation: Google (DeepMind)

Meta

Menlo Park, California · US

Meta built the Llama line that made open weights normal, and now ships the Muse models: Muse Spark through a paid API and Muse Glimmer as downloadable open weights under Apache 2.0. That split makes Meta the most difficult lab to summarise in one sentence for a procurement file.

Models, oldest to newest

  1. Muse Spark 1.1Jul 2026Memory: 1M tokens(unverified)
  2. Muse Glimmer 30BAug 2026Memory: 131K tokens
  3. Muse Spark 1.32 Sept 2026Memory: 1M tokens
  • current
  • retired or retiring
  • unverified
Model lineage for Meta: 3 current, 0 retired or retiring, each shown with the context window its vendor publishes.
ModelReleasedRoleMemory
Muse Spark 1.1Jul 2026current1M tokens
Muse Glimmer 30BAug 2026current131K tokens
Muse Spark 1.32 Sept 2026current1M tokens
Bar length is the published context window, on one scale across all eight labs. Source: vendor documentation, read 8 September 2026.
Current line-up: Meta, 8 Sept 2026.
ModelRoleReleasedMemoryEU residencyPrice per 1M tokens (in / out)
Muse Spark 1.3reasoningFlagship2 Sept 20261M tokensNo EU routeMeta publishes no European data-residency option for the Meta Model API. Until it does, this is an American API endpoint and should be treated as one in a data-protection assessment.US$1.25 / US$4.25
Muse Glimmer 30Bopen gewichtenreasoningApache-2.0Open gewichtenAug 2026131K tokensEU routeSelf-hosted on your own or an EU provider's hardwareOpen weights under Apache 2.0, small enough to run on a single high-end consumer graphics card once quantised. Hosting is therefore your own decision, which is the cleanest possible answer to a data-residency question.not published
Muse Spark 1.1unverifiedreasoningMid tierJul 20261M tokensNo EU routeWe could not confirm an EU data-residency deployment for this model in the vendor's own documentation on 8 September 2026. Treat European routing as something to confirm in writing before any board material is sent.US$1.25 / US$4.25

Official model documentation: Meta

xAI

Palo Alto, California · US

Founded by Elon Musk in 2023, xAI builds the Grok models and distributes them through the X platform and its own API. It moves fast on price and speed, and its documented serving regions are in the United States, which shapes what a European board can do with it.

Models, oldest to newest

  1. Grok 4.3Apr 2026Memory: 1M tokens(unverified)
  2. Grok 4.5Jul 2026Memory: 500K tokens(unverified)
  3. Grok 4.612 Aug 2026Memory: 500K tokens
  • current
  • retired or retiring
  • unverified
Model lineage for xAI: 3 current, 0 retired or retiring, each shown with the context window its vendor publishes.
ModelReleasedRoleMemory
Grok 4.3Apr 2026current1M tokens
Grok 4.5Jul 2026current500K tokens
Grok 4.612 Aug 2026current500K tokens
Bar length is the published context window, on one scale across all eight labs. Source: vendor documentation, read 8 September 2026.
Current line-up: xAI, 8 Sept 2026.
ModelRoleReleasedMemoryEU residencyPrice per 1M tokens (in / out)
Grok 4.6reasoningFlagship12 Aug 2026500K tokensNo EU routexAI documents this model as available in us-east-1 and us-west-2. There is no documented European serving region, which is the single most important fact for a European board considering Grok.US$2 / US$6500K-token window. Prices shown are for prompts under 200K tokens; above that the rate doubles to $4 input and $12 output per million.
Grok 4.5unverifiedreasoningMid tierJul 2026500K tokensNo EU routeWe could not confirm an EU data-residency deployment for this model in the vendor's own documentation on 8 September 2026. Treat European routing as something to confirm in writing before any board material is sent.US$2 / US$6
Grok 4.3unverifiedreasoningMid tierApr 20261M tokensNo EU routeWe could not confirm an EU data-residency deployment for this model in the vendor's own documentation on 8 September 2026. Treat European routing as something to confirm in writing before any board material is sent.US$1.25 / US$2.501M-token window, cheaper than 4.6. Prices shown are for prompts under 200K tokens.

Official model documentation: xAI

Mistral AI

Paris, France · FR

The only European lab in this list, with a high release cadence and a catalogue that mixes its own models with open models it hosts. For organisations whose sovereignty requirement is a European contracting party rather than only a European region, Mistral is the shortest path.

Models, oldest to newest

  1. Mistral Large 2.1Nov 2024end of service 31 May 2026Memory: not published
  2. Mistral Medium 3.1Aug 2025end of service 31 Aug 2026Memory: not published
  3. Mistral Large 3Dec 2025Memory: 256K tokens
  4. Mistral Small 4Mar 2026Memory: not published(unverified)
  5. Mistral Medium 3.5Apr 2026Memory: 256K tokens
  • current
  • retired or retiring
  • unverified
Model lineage for Mistral AI: 3 current, 2 retired or retiring, each shown with the context window its vendor publishes.
ModelReleasedRoleMemory
Mistral Large 2.1Nov 2024retired or retiringnot published
Mistral Medium 3.1Aug 2025retired or retiringnot published
Mistral Large 3Dec 2025current256K tokens
Mistral Small 4Mar 2026currentnot published
Mistral Medium 3.5Apr 2026current256K tokens
Bar length is the published context window, on one scale across all eight labs. Source: vendor documentation, read 8 September 2026.
Current line-up: Mistral AI, 8 Sept 2026.
ModelRoleReleasedMemoryEU residencyPrice per 1M tokens (in / out)
Mistral Medium 3.5open gewichtenreasoningModified MIT License: weights are downl...FlagshipApr 2026256K tokensEU routeMistral La Plateforme with Mistral Regional Endpoints set to Europe; Scaleway Generative APIs (France), served as mistral-medium-3.5-128b; Self-hosted with vLLM, TensorRT-LLM or TGI on your own hardware; Microsoft Azure AI, AWS Bedrock, Google Cloud Vertex AI, IBM watsonx.ai and OutscaleThis is the only model in this file where the lab itself is European and offers in-region European inference as a contracted product. Mistral made Regional Endpoints generally available on 11 August 2026, letting you pin inference to Europe or the US.US$1.50 / US$7.50The strongest model from the only European frontier lab. Mistral publishes no cached-input rate, so a caching discount cannot be assumed in a budget.
Mistral Small 4unverifiedopen gewichtenreasoningApache-2.0Small and fastMar 2026not publishedEU routeMistral La Plateforme with Mistral Regional Endpoints set to Europe; Self-hosted on your own hardware (Apache-2.0)US$0.15 / US$0.60Small tier for routine drafting and classification.
Mistral Large 3open gewichtenApache-2.0Open gewichtenDec 2025256K tokensEU routeMistral La Plateforme with Mistral Regional Endpoints set to Europe; Self-hosted on your own hardware (Apache-2.0)US$0.50 / US$1.50Cheaper than Medium 3.5 despite the name. Read the model card before assuming Large means stronger.

Official model documentation: Mistral AI

DeepSeek

Hangzhou, China · CN

A Chinese lab that made frontier-level models dramatically cheaper and publishes its weights, so the model can be downloaded and run on hardware you control. For a European board that means the interesting question is not whether to use the Chinese API but whether to self-host the weights.

Models, oldest to newest

  1. DeepSeek legacy endpoints (deepseek-chat, deepseek-reasoner)Dec 2024end of service 24 Jul 2026Memory: not published
  2. DeepSeek-V4-Flash31 Jul 2026Memory: 1M tokens
  3. DeepSeek-V4-Pro13 Aug 2026Memory: 1M tokens
  4. DeepSeek-V4-Flash-Vision-Exp21 Aug 2026Memory: 1M tokens
  • current
  • retired or retiring
  • unverified
Model lineage for DeepSeek: 3 current, 1 retired or retiring, each shown with the context window its vendor publishes.
ModelReleasedRoleMemory
DeepSeek legacy endpoints (deepseek-chat, deepseek-reasoner)Dec 2024retired or retiringnot published
DeepSeek-V4-Flash31 Jul 2026current1M tokens
DeepSeek-V4-Pro13 Aug 2026current1M tokens
DeepSeek-V4-Flash-Vision-Exp21 Aug 2026current1M tokens
Bar length is the published context window, on one scale across all eight labs. Source: vendor documentation, read 8 September 2026.
Current line-up: DeepSeek, 8 Sept 2026.
ModelRoleReleasedMemoryEU residencyPrice per 1M tokens (in / out)
DeepSeek-V4-Flash-Vision-Expopen gewichtenreasoningMid tier21 Aug 20261M tokensEU routeSelf-hosted, subject to weights being published for this experimental buildDeepSeek labels this build experimental. Treat it as a pilot option for scanned documents, not as something to standardise on.US$0.44 / US$1.32
DeepSeek-V4-Proopen gewichtenreasoningMITFlagship13 Aug 20261M tokensEU routeSelf-hosted on your own or a rented GPU cluster (weights are public under MIT); European GPU rental with a data processing agreement, for example an EU region of a sovereign cloudThere is no managed EU endpoint for V4-Pro from a large European provider that we could confirm on 8 September 2026. Amazon Bedrock lists DeepSeek-V3.1 and R1, not V4. The realistic EU route for this specific model is running the weights yourself.US$1.32 / US$3.96Peak-hours list price. Off-peak, which DeepSeek defines as every hour outside 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, the same tokens cost half: $0.66 input, $1.98 output, $0.022 cached.
DeepSeek-V4-Flashopen gewichtenreasoningMITMid tier31 Jul 20261M tokensEU routeScaleway Generative APIs (France), served as deepseek-v4-flash-0731; Self-hosted on your own or a rented GPU cluster (weights are public under MIT)Scaleway serves this model from European data centres at a 256,000 token context window, which is a quarter of the window DeepSeek offers on its own API. Check the served window, not the model's maximum, before you promise a use case.US$0.44 / US$1.32Peak-hours list price; off-peak is half. The cheapest frontier-adjacent API on this list.

Official model documentation: DeepSeek

Alibaba (Qwen)

Hangzhou, China · CN

Alibaba Cloud ships the Qwen family and also resells other labs' models through the same cloud, which makes its changelog one of the busiest in the industry. Part of the Qwen range is released as open weights, usually under Alibaba's own licence rather than a standard open-source one.

Models, oldest to newest

  1. Qwen3.8-Max3 Aug 2026Memory: 1M tokens
  2. Qwen3.8-2.4T-A95B (open build)13 Aug 2026Memory: 262K tokens
  3. Qwen3.8-27B19 Aug 2026Memory: 262K tokens
  4. Qwen3.8-Flash26 Aug 2026Memory: 1M tokens
  • current
  • retired or retiring
  • unverified
Model lineage for Alibaba (Qwen): 4 current, 0 retired or retiring, each shown with the context window its vendor publishes.
ModelReleasedRoleMemory
Qwen3.8-Max3 Aug 2026current1M tokens
Qwen3.8-2.4T-A95B (open build)13 Aug 2026current262K tokens
Qwen3.8-27B19 Aug 2026current262K tokens
Qwen3.8-Flash26 Aug 2026current1M tokens
Bar length is the published context window, on one scale across all eight labs. Source: vendor documentation, read 8 September 2026.
Current line-up: Alibaba (Qwen), 8 Sept 2026.
ModelRoleReleasedMemoryEU residencyPrice per 1M tokens (in / out)
Qwen3.8-FlashreasoningMid tier26 Aug 20261M tokensEU routeAlibaba Cloud Model Studio, Germany (Frankfurt) region eu-central-1Same caveat as the flagship: Alibaba states that supported models differ per region, so confirm Frankfurt availability before you commit.US$0.15 / US$0.47Small tier on the international endpoint.
Qwen3.8-27Bopen gewichtenreasoningApache-2.0Open gewichten19 Aug 2026262K tokensEU routeOVHcloud AI Endpoints (France); IONOS AI Model Hub (Germany); Self-hosted on a single modern data centre GPUThis is the most practical sovereign option in this file for a company without a GPU cluster: two European providers serve it, and the weights are Apache-2.0 so you can move between them or bring it in house.not published
Qwen3.8-2.4T-A95B (open build)open gewichtenreasoningQwen3.8-Max LicenseOpen gewichten13 Aug 2026262K tokensEU routeSelf-hosted on your own or a rented GPU clusterAt 2.4 trillion parameters this is a multi node deployment. Realistically it is for organisations that already operate a GPU cluster, or for a managed sovereign platform running it on your behalf.not published
Qwen3.8-MaxreasoningFlagship3 Aug 20261M tokensEU routeAlibaba Cloud Model Studio, Germany (Frankfurt) region eu-central-1Alibaba Cloud Model Studio does run a Germany (Frankfurt) region, but Alibaba states in its own documentation that regions differ in supported models, features and pricing. Confirm in the Frankfurt console that this specific model is offered there before you rely on it.US$2 / US$6Price on Alibaba's international (Singapore) endpoint. The China-domestic endpoint is a different service with its own prices and its own legal position.

Official model documentation: Alibaba (Qwen)

05Four years in one line

Four years in one line: releases, retirements, rules

Releases are the news. Retirements are the work. The European rules are the calendar. This is all three on one rail, and every date on it is confirmed against a primary source.

A timeline earns its place on a page like this only if every line in it is a date on which something an organisation depends on either appeared or stopped existing. That is the filter used here. The rail below carries the releases that actually moved the frontier, the shutdowns that broke somebody's integration, and the dates the European Union has written into law.

On the rail

  • release: 11
  • retired: 11
  • milestone: 3
  • regulation: 7

Horizontal timeline. It scrolls sideways as you scroll down, and becomes a plain vertical list on a phone or when your system asks for reduced motion.

  1. 30 Nov 2022OpenAI / release

    ChatGPT opens to the public

    OpenAI released ChatGPT as a free research preview built on GPT-3.5, with a window of roughly four thousand tokens, about six pages.

  2. 14 Mar 2023OpenAI / release

    GPT-4 arrives with an eight thousand token window

    GPT-4 was the first model most executives took seriously, and it held about twelve pages of text at a time.

  3. 11 Jul 2023Anthropic / release

    Claude 2 makes one hundred thousand tokens normal

    Anthropic released Claude 2 with the one hundred thousand token window it had opened in May as standard, more than ten times the GPT-4 of March.

  4. 6 Nov 2023OpenAI / release

    GPT-4 Turbo raises the window to 128,000 tokens

    At its first developer day OpenAI raised the window sixteenfold and cut prices.

  5. 27 Jun 2024Google / release

    Two million tokens becomes generally available

    Google made the two million token window on Gemini 1.5 Pro available to every developer, after announcing it at its May developer conference.

  6. 1 Aug 2024European Union / regulation

    The EU AI Act enters into force

    The AI Act became law across the European Union, with its obligations phased in over the following four years.

  7. 12 Sept 2024OpenAI / milestone

    o1-preview is the first model trained to think first

    OpenAI released o1-preview and o1-mini, models trained to reason step by step before answering rather than producing the first plausible sentence.

  8. 6 Nov 2024Anthropic / retired

    Anthropic retires the entire Claude 1 line

    Claude 1.0 through 1.3 and the Claude Instant models were switched off, two months after the deprecation notice.

  9. 20 Jan 2025DeepSeek / milestone

    DeepSeek-R1 makes reasoning cheap and open

    DeepSeek launched R1, an openly licensed reasoning model that performed close to o1 at a fraction of the price.

  10. 2 Feb 2025European Union / regulation

    Prohibited AI practices and AI literacy start applying

    The first binding obligations of the AI Act took effect: a list of banned uses such as social scoring and untargeted facial image scraping, plus a duty to ensure staff who use AI have sufficient AI literacy.

  11. 14 Jul 2025OpenAI / retired

    GPT-4.5 is switched off in the API

    Four and a half months after its research preview, GPT-4.5 was removed from the API with three months notice.

  12. 2 Aug 2025European Union / regulation

    Obligations for general-purpose AI models start applying

    Providers of general-purpose AI models became subject to transparency, documentation, copyright policy and, for the largest models, systemic risk obligations.

  13. 7 Aug 2025OpenAI / release

    GPT-5 replaces the model picker with a router

    OpenAI released GPT-5 as a single system that decides per question whether to answer quickly or think first, with a four hundred thousand token window.

  14. 29 Sept 2025Google / retired

    Gemini 1.5 is shut down in the Gemini API

    Google switched off Gemini 1.5 Pro, 1.5 Flash and 1.5 Flash-8B in the Gemini API, with the Vertex AI versions retiring on 24 May and 24 September 2025.

  15. 5 Jan 2026Anthropic / retired

    Claude 3 Opus is retired

    The model that was Anthropic's flagship in March 2024 was switched off, twenty two months after release and six months after the notice.

  16. 19 Feb 2026Anthropic / retired

    Claude 3.7 Sonnet and Claude 3.5 Haiku are retired

    Anthropic switched off Claude 3.7 Sonnet, the model that introduced extended thinking, one year after its release.

  17. 9 Mar 2026Google / retired

    Gemini 3 Pro Preview is shut down

    Google shut down Gemini 3 Pro Preview less than four months after release and redirected calls to Gemini 3.1 Pro Preview.

  18. 24 Apr 2026DeepSeek / release

    DeepSeek V4 reaches one million tokens under an MIT licence

    DeepSeek launched V4-Pro and V4-Flash with a one million token window and public weights under an MIT licence.

  19. 9 Jun 2026Anthropic / release

    Claude Fable 5 opens the Claude 5 generation

    Anthropic released Claude Fable 5 and Claude Mythos 5: the same underlying model at two levels of safeguarding, with Fable generally available and Mythos restricted to vetted organisations.

  20. 12 Jun 2026Anthropic / milestone

    A US export directive suspends Fable 5 worldwide

    According to Anthropic's own post, Anthropic suspended Claude Fable 5 and Mythos 5 for all users following a United States government export-control directive, restoring service on 1 July 2026.

  21. 15 Jun 2026Anthropic / retired

    Claude Opus 4 and Sonnet 4 are retired

    The original Claude 4 models were switched off thirteen months after release, two months after the deprecation notice.

  22. 9 Jul 2026OpenAI / release

    GPT-5.6 replaces version numbers with named tiers

    OpenAI released GPT-5.6 as three durable tiers, Sol, Terra and Luna, all with a 1,050,000 token window.

  23. 27 Jul 2026European Union / regulation

    The Digital Omnibus on AI enters into force

    Regulation (EU) 2026/1744 of 8 July 2026 was published on 24 July and entered into force on 27 July, amending the AI Act.

  24. 2 Aug 2026European Union / regulation

    The AI Act becomes generally applicable

    The Article 50 transparency duties started applying: people must be told when they are talking to an AI, AI-generated content must be marked, and deepfakes must be disclosed.

  25. 1 Sept 2026Anthropic / release

    Claude Fable 5.1 becomes Anthropic's flagship

    Anthropic released Fable 5.1 and Mythos 5.1 with a one million token window, 128,000 tokens of output, always-on adaptive thinking, and cache reads cut by seventy five percent.

  26. 3 Sept 2026OpenAI / release

    GPT-6 Astra is unveiled

    OpenAI unveiled GPT-6 Astra with a 1,050,000 token window, 128,000 tokens of output and a knowledge cutoff of 30 April 2026, at ten dollars in and fifty dollars out per million tokens.

  27. 28 Sept 2026OpenAI / retired

    The last GPT-3.5 era models shut down

    gpt-3.5-turbo-instruct, babbage-002, davinci-002 and gpt-3.5-turbo-1106 are switched off, a year after the notice.

  28. 20 Oct 2026Google / retired

    Gemini 2.5 Pro, Flash and Flash-Lite retire

    Google has announced the retirement of the whole Gemini 2.5 line, sixteen months after it reached general availability.

  29. 23 Oct 2026OpenAI / retired

    GPT-4, GPT-4 Turbo, o1, o3-mini and o4-mini shut down

    OpenAI switches off the entire legacy family in one step: gpt-3.5-turbo-0125, gpt-4-0613, gpt-4-turbo, gpt-4o-2024-05-13, o1, o1-pro, o3-mini and o4-mini, including fine-tuned variants.

  30. 11 Dec 2026OpenAI / retired

    The original GPT-5 and o3 shut down

    gpt-5, gpt-5-mini, gpt-5-nano, gpt-5-pro, o3 and o3-pro are switched off sixteen months after GPT-5 launched.

  31. 2 Dec 2027European Union / regulation

    High-risk obligations apply to Annex III systems

    The obligations for stand-alone high-risk systems, covering areas such as employment, education, credit, biometrics and critical infrastructure, start applying.

  32. 2 Aug 2028European Union / regulation

    High-risk obligations apply to AI inside regulated products

    AI embedded in products already covered by EU product safety law, such as machinery, medical devices and lifts, comes under the high-risk regime.

The retirement problem

Two words do most of the damage here, and they do not mean the same thing. A model is deprecated when the vendor announces that it will stop being available and puts a date on it. Nothing breaks that day. The model still answers. A model is retired, or shut down, on the day the endpoint stops answering and every call to it fails. The gap between those two moments is the notice period, and it is the only window in which a migration is cheap. All three of the large American laboratories publish both dates: Anthropic keeps a deprecation register, OpenAI keeps one, and Google publishes lifecycle tables for Vertex AI alongside a dated changelog for the Gemini API.

For planning purposes the honest range is six to eighteen months of useful life per model version, with two to twelve months of notice before the end of it. Write that number into the plan rather than into a risk register nobody reads: a model your organisation standardises on in September 2026 will most likely be gone within a year and a half, and its replacement will behave differently on your own documents. Not worse, usually better, but differently. Prompts that were tuned around one model's habits stop landing, output formats shift, and a summarisation step that never used to invent a heading suddenly does.

A model exit clause, in five lines

None of this is exotic to buy. Ask for it at signature, when it costs nothing, rather than after a shutdown date is announced, when it is close to impossible to obtain. Your own counsel will phrase it better, but the substance is five points:

  1. The vendor's published notice period, in writing, with the address of the register where deprecations appear. If the answer is not a URL, there is no notice period.
  2. A named successor model, and a commitment that your data-processing terms, your hosting region and your retention settings carry over to it unchanged.
  3. The right to test that successor on your own documents, with your own prompts and your own evaluation set, before the switch is forced, and a stated window in which to do it.
  4. An exit route if the successor fails that test: export of your prompts, your evaluation results and your data in a usable format, and no penalty for leaving mid-term.
  5. For regulated work, a written statement of where the successor is processed. A successor in a different region is a new supplier assessment, not a version bump.

Every verified retirement

Retired and announced shutdowns from the model dataset, most recent first. Sources: the vendors' own deprecation registers and lifecycle tables, accessed 8 September 2026. Only entries confirmed against a primary source are listed.
ModelReleasedShutdownSuccessor
GPT-57 Aug 202511 Dec 2026announcedGPT-5.6 Sol
o316 Apr 202511 Dec 2026announcedGPT-5.6 Sol
o4-mini16 Apr 202523 Oct 2026announcedGPT-5.6 Terra
o117 Dec 202423 Oct 2026announcedGPT-5.6 Sol
GPT-413 Jun 202323 Oct 2026announcedGPT-5.6 Sol
GPT-3.5 Turbo25 Jan 202423 Oct 2026announcedGPT-5.6 Terra
Mistral Medium 3.1Aug 202531 Aug 2026goneMistral Medium 3.5
Claude Opus 4.15 Aug 20255 Aug 2026goneClaude Opus 4.8
DeepSeek legacy endpoints (deepseek-chat, deepseek-reasoner)Dec 202424 Jul 2026goneDeepSeek-V4-Flash
Claude Opus 414 May 202515 Jun 2026goneClaude Opus 4.8
Claude Sonnet 414 May 202515 Jun 2026goneClaude Sonnet 5
Mistral Large 2.1Nov 202431 May 2026goneMistral Medium 3.5
Claude Haiku 37 Mar 202420 Apr 2026goneClaude Haiku 4.5
Claude Sonnet 3.719 Feb 202519 Feb 2026goneClaude Sonnet 5
Claude Haiku 3.522 Oct 202419 Feb 2026goneClaude Haiku 4.5
Claude Opus 329 Feb 20245 Jan 2026goneClaude Opus 4.8

Releases against retirements, per year

Two honest notes on reading these bars. The 2026 release count runs to 3 September 2026 only, so the year is not finished. The 2026 retirement count is higher than the others partly because it includes shutdowns that have been announced but not yet executed, which is precisely why they are useful: those dates are still in front of you.

Verified releases per year

  • 20221 event
  • 202310 events
  • 202410 events
  • 202516 events
  • 202621 events
Verified model releases per calendar year in the research dataset, 2022 to 2026.
LabelValue
20221 event
202310 events
202410 events
202516 events
202621 events

Verified retirements per year

  • 20220 events
  • 20230 events
  • 20241 event
  • 20256 events
  • 202612 events
Verified model retirements and announced shutdowns per calendar year in the research dataset, 2022 to 2026.
LabelValue
20220 events
20230 events
20241 event
20256 events
202612 events

Counted from the verified events in the research timeline dataset. 2026 releases run to 3 September 2026; 2026 retirements include shutdowns already announced for the rest of the year.

Three habits cover most of the risk, and none of them needs a project. Keep a list of every model name your organisation actually calls, with one named owner per line, including the ones buried in automations and spreadsheets. Subscribe those owners to the vendors' deprecation registers, which are public web pages that change a handful of times a year. And when a successor arrives, re-run the same test you used to choose the model in the first place, on your own documents, before the old one goes dark.

The test itself, written out per executive role with the questions we use, is on Which AI model for which role. The person who usually ends up owning the migration calendar has a page of their own: AI CTO.

06Sovereignty

EU hosting: what is actually available inside the Union

Seven routes, one finding that surprises most boards, and a set of dates that moved. As of September 2026, running a strong model in Europe is a solved problem for the tier below the frontier and a narrow one for the frontier itself.

Most boards start this conversation with a benchmark table. Start with the hosting requirement instead: a model your organisation is not allowed to use scores zero on every metric that matters. Work out which models survive the requirement, then compare quality inside that shortlist.

Seven words, defined once

Data residency
A promise about geography: your documents are stored, and your questions answered, inside a stated region. It says nothing about who owns the company that runs the machines.
Data sovereignty
A promise about jurisdiction and control: which country's law reaches the data, and which people can touch it. Residency is a subset of sovereignty, not a synonym for it.
Inferentie (inferentie)
The moment the model reads your question and writes an answer. This is where your board pack actually travels, so it is the step whose location matters most.
Data zone and inference profile
Cloud settings that pin inference to a group of regions. Microsoft calls it a Data Zone deployment, AWS calls it a cross-region inference profile, Google calls it machine learning processing in a region.
Open weights
The model file itself is published and you may download and run it. Open weights is not the same as open source: the training data and code usually stay closed, and the licence may carry conditions.
Zero data retention
A contract setting in which the provider keeps no copy of your prompts or answers after the request finishes, not even in the short-term logs kept to detect abuse.
US CLOUD Act exposure
The question of whether a US-headquartered provider can be compelled by US law to hand over data it holds abroad. An EU region does not by itself answer that question; a separate EU legal entity is what providers offer against it.

Frontier-class means a model its lab presented as its most capable general-purpose model and that sat in the leading group on published general capability at the time of release. It is a judgement, so every model counted is named below and can be checked. Inside the EU means one of two things, counted separately: hosted with a data residency guarantee, where the vendor or a cloud provider commits to processing inside EU regions, or self-hostable, where the weights are published under a licence that allows a company to run the model on its own EU hardware. Data residency is not the same as jurisdiction: the EU regions of Microsoft, Amazon and Google are operated by US companies, so residency describes where the data is processed and not who can be compelled to produce it.

Which model, on which route, with EU data residency

EU region, US cloudLab's own EU endpointEuropean model vendorEuropean cloud, open modelsSovereign cloudYour own hardwareConsumer apps
Claude Fable 5.1
GPT-6 Astra
Gemini 3.8 Flash
Claude Opus 5
Claude Sonnet 5
Claude Haiku 4.5
GPT-5.6 Sol
GPT-5.5
Gemini 3.5 Flash
Gemini 2.5 Pro
Mistral Medium 3.5
DeepSeek-V4-Pro
Amazon Nova Pro
no EU residency documentedEU residency documented
Model by hosting route matrix. Of the three newest US flagships, Claude Fable 5.1 and Gemini 3.8 Flash have documented EU data residency on Google Cloud's eu multi-region only and GPT-6 Astra on no route, while Claude Opus 5, Claude Sonnet 5, GPT-5.6 Sol, Gemini 3.5 Flash and Gemini 2.5 Pro have it on the EU region of a US cloud, and Mistral Medium 3.5, DeepSeek-V4-Pro and Amazon Nova Pro can be run inside the EU by a European operator.
EU region, US cloudLab's own EU endpointEuropean model vendorEuropean cloud, open modelsSovereign cloudYour own hardwareConsumer apps
Claude Fable 5.1100%0%0%0%0%0%0%
GPT-6 Astra0%0%0%0%0%0%0%
Gemini 3.8 Flash100%0%0%0%0%0%0%
Claude Opus 5100%0%0%0%0%0%0%
Claude Sonnet 5100%0%0%0%0%0%0%
Claude Haiku 4.5100%0%0%0%0%0%0%
GPT-5.6 Sol100%0%0%0%0%0%0%
GPT-5.5100%0%0%0%0%0%0%
Gemini 3.5 Flash100%0%0%0%0%0%0%
Gemini 2.5 Pro100%0%0%0%0%0%0%
Mistral Medium 3.50%0%100%50%0%100%0%
DeepSeek-V4-Pro0%0%0%50%0%100%0%
Amazon Nova Pro0%0%0%0%100%0%0%
Sources: vendor and cloud provider documentation, accessed 8 September 2026. In the data table behind this grid, 100 percent means documented EU data residency, 50 percent means available on the route with an unconfirmed catalogue, and 0 percent means no EU residency documented or not offered on that route.

Reading key

  • Documented EU data residency for this model on this route.
  • Technically available on this route, but the provider's current catalogue was not confirmed. Ask for it in writing.
  • Offered on this route without documented EU data residency.
  • Not offered on this route.

The seven routes

Ordered from the route most mid-market companies take to the one that answers every jurisdiction question and costs the most to run. Read the unsettled line of each first: that is where the procurement work sits.

01 EU region, US cloud

EU region of a US cloud

processing in the EUOperational burden: low

You rent a frontier model from Amazon, Microsoft or Google and pin the processing to European data centres.

Named offerings

  • Amazon Bedrock, EU cross-region inference profile
  • Microsoft Foundry, Data Zone Standard (European Union)
  • Google Cloud, EU multi-region with data residency at rest and machine learning processing
  • OpenAI API regional processing on eu.api.openai.com

What this route does not settle

Three things. The contracting party is still a US-headquartered company, so US CLOUD Act exposure is a legal judgement your counsel makes, not a setting you switch on. Support and engineering staff outside the EU may still be able to reach systems, which is why the sovereign tiers below exist. And the setting does not cover every newest flagship: GPT-6 Astra has no EU route at all, and Claude Fable 5.1 and Gemini 3.8 Flash have Google Cloud's eu multi-region only, so the residency requirement quietly narrows your model choice.

02 Lab's own EU endpoint

The lab's own European endpoint

processing in the EU, with contractual controlsOperational burden: low

You buy directly from the model lab and it offers a European address that keeps storage and processing in Europe.

Named offerings

  • OpenAI API regional processing, eu.api.openai.com
  • Anthropic Claude in Amazon Bedrock, zero operator access
  • Mistral La Plateforme and Le Chat Enterprise

What this route does not settle

The vendor is still American, so the CLOUD Act question is unchanged. Per-model eligibility on the European endpoint is not published in the documentation we could read, so treat the model list as something to confirm in writing during procurement rather than as a public fact.

03 European model vendor

A European model vendor

operated by an EU legal entityOperational burden: low

You buy the model from a company headquartered in the EU, which removes the American jurisdiction question from the contract entirely.

Named offerings

  • Mistral Le Chat Enterprise, Mistral cloud or your own cloud
  • Mistral La Plateforme API

What this route does not settle

If you take the managed Mistral cloud rather than self-hosting, confirm in the contract which subprocessors are used and in which countries, because a European vendor can still buy compute abroad. We could not read Mistral's current data processing agreement during this research, so treat subprocessor geography as an open question for procurement.

04 European cloud, open models

A European cloud running open models

operated by an EU legal entityOperational burden: low

A European provider hosts published model weights on European hardware and sells you the inference as a normal API.

Named offerings

  • IONOS AI Model Hub
  • Scaleway Generative APIs
  • OVHcloud AI Endpoints
  • STACKIT AI services
  • T-Systems sovereign AI services

What this route does not settle

Model choice is limited to what is published, and a published model can be Chinese in origin even when the hardware is German. If the board's objection is to Chinese vendors rather than to American ones, read the model card before reading the data centre address.

05 Sovereign cloud

A hyperscaler's sovereign cloud

operated by an EU legal entityOperational burden: medium

The same three US clouds, but sold through a separate European legal structure with European staff and no operational control from outside the EU.

Named offerings

  • AWS European Sovereign Cloud, region eusc-de-east-1 (Brandenburg)
  • Microsoft EU Data Boundary for the Microsoft Cloud
  • Google Cloud Data Boundary with Assured Workloads
  • Google Cloud Dedicated with S3NS (France) and T-Systems (Germany)
  • Google Distributed Cloud air-gapped

What this route does not settle

Model availability, and whether a separate EU legal entity of an American group satisfies your interpretation of sovereignty. That second question is a board decision informed by counsel, not a technical fact, and reasonable European boards answer it differently.

06 Your own hardware

Open weights on hardware you control

your hardware, your weightsOperational burden: high

You download a published model and run it on your own servers, or on rented servers you fully control.

Named offerings

  • Your own data centre
  • Rented dedicated GPU capacity at a European provider
  • A managed private deployment of published weights

What this route does not settle

Nothing about jurisdiction. Everything about capability and continuity: you are now responsible for security patching, for model upgrades, and for the fact that the published state of the art moves every few months while your cluster does not.

07 Consumer apps

The route nobody chose: consumer apps

processing in the EUOperational burden: low

Staff paste board material into a free or personal chat account, and the organisation's data policy never sees it.

Named offerings

  • Free and personal tiers of consumer chat apps
  • Browser extensions and note apps with a built-in assistant

What this route does not settle

Everything. Treat this as the baseline risk you are actually mitigating, and measure your policy by how much of this traffic it removes.

From requirement to route

Start from the hosting requirement, not from the benchmark table. A model you are not allowed to use scores zero. Work down these six branches until one matches what your board, your works council and your largest customer's contract actually require, then pick the best model on that branch and test it on your own documents.

Must your data stay in the EU?
No residency requirement
Must the operator be an EU legal entity?
EU region, US vendor allowed
Must the model vendor be European too?
EU operator required
Must it run on hardware you control?
European vendor
Our own hardware
Each outcome maps onto one of the routes above.

The regulatory calendar, after the Digital Omnibus

  1. 1 Aug 2024regulation

    The AI Act enters into force

    Regulation (EU) 2024/1689 enters into force. Nothing is required of a company yet; the clock starts on a staged set of deadlines that runs to 2028.

  2. 2 Feb 2025regulation

    Prohibited practices and AI literacy apply

    The bans on a short list of AI uses, such as social scoring and untargeted facial-image scraping, start to apply. So does Article 4, the AI literacy obligation, which reaches every organisation that provides or deploys an AI system, not only the ones building models.

  3. 2 Aug 2025regulation

    Obligations for general-purpose AI models apply

    The obligations on providers of general-purpose AI models, meaning the labs that train the models themselves, start to apply, together with the governance and confidentiality rules. Member States must have designated competent authorities and set out penalties by this date.

  4. 27 Jul 2026regulation

    The Digital Omnibus on AI enters into force

    Regulation (EU) 2026/1744 amends the AI Act. It defers the high-risk obligations, softens the wording of the AI literacy duty from ensuring a sufficient level to supporting the development of AI literacy, allows simplified technical documentation for SMEs and small mid-caps, and adds prohibitions on systems generating non-consensual intimate material and child sexual abuse material.

  5. 2 Aug 2026regulation

    The rest of the AI Act applies, and enforcement begins

    The remainder of the AI Act becomes applicable, including the Article 50 transparency duties: tell people when they are talking to an AI system, and label synthetic content. National market surveillance authorities supervise and enforce from this date, which is the first moment the literacy duty has a visible enforcer behind it.

  6. 2 Dec 2026regulation

    Marking of synthetic content and the new prohibitions

    Synthetic audio, image, video and text produced by systems placed on the market before 2 August 2026 must comply with the marking obligations. The omnibus prohibitions on non-consensual intimate material and child sexual abuse material also take effect.

  7. 2 Aug 2027regulation

    Older general-purpose models must comply

    Providers of general-purpose AI models placed on the market before 2 August 2025 must be in full compliance. In practice this closes the grandfathering window for the model generation your organisation may still be running.

  8. 2 Dec 2027regulation

    High-risk systems under Annex III

    The Chapter III requirements apply to stand-alone high-risk systems: biometrics, critical infrastructure, education, employment, migration, asylum and border control. Deferred from 2 August 2026 by the Digital Omnibus. For most mid-market boards the relevant one is employment: AI used in recruitment or in decisions about staff.

  9. 2 Aug 2028regulation

    High-risk AI inside regulated products

    The Chapter III requirements apply to AI embedded in products that already carry EU product legislation, such as lifts, toys and machinery. Deferred from 2 August 2027 by the Digital Omnibus.

What a mid-market board actually has to do

AI literacy: Article 4

Providers and deployers must take measures to support the development of AI literacy among staff and others who operate their AI systems on their behalf. There is no mandated training format and no certification. What a board should be able to show is that the people using the assistant know roughly how it fails, what it must not be fed, and who to ask. The Commission and Member States are now formally tasked with supporting this effort, in particular for SMEs.

Applies from 2 Feb 2025

Transparency: Article 50

People must be told when they are interacting with an AI system, and synthetic audio, image, video and text must be marked as such. For most companies this lands on the website chatbot and on anything published that a model wrote. These duties were not deferred by the Digital Omnibus.

Applies from 2 Aug 2026

Check once whether you are high risk

The high-risk regime now bites on 2 December 2027 for stand-alone systems under Annex III. The category that catches ordinary companies is employment: AI used to filter applicants or to inform decisions about staff. The deferral is time to prepare, not a reason to skip the check, and the check itself is a one-hour exercise for most boards.

Applies from 2 Dec 2027

Write down which model, where, and on what data

No article requires a mid-market deployer to keep a model register, but every conversation with an auditor, an insurer or a works council starts with the same three questions. A single page listing the model, the hosting route, the data it may see and the person accountable answers all of them and takes an afternoon to produce.

Applies from 2 Aug 2026

Training, retention and residency, per vendor

Data handling per vendor and service, from the providers' own documentation, accessed 8 September 2026. Cloud rows describe the hosting platform; the model vendor's own terms govern training on that platform.
Vendor and serviceTrains on your data by default?Retention and residency controlsSource
OpenAI APINo, unless you explicitly opt in to share data.Zero Data Retention removes customer content from abuse-monitoring logs, otherwise kept up to 30 days. European regional storage and processing run on a separate endpoint.
Anthropic, commercial productsNo, by default, for Claude for Work, the Anthropic API and Claude Gov.Anthropic states zero operator access for Claude in Amazon Bedrock: its own personnel cannot reach the inference infrastructure.
Anthropic, consumer ClaudeOnly if the user opts in, if a conversation is flagged for safety review, or in an explicit programme.Incognito chats are excluded from model improvement even when the setting is on. Consumer terms, agreed by the individual, not by your organisation.
Microsoft FoundryNot stated in the documentation we checked. The model vendor's terms govern this.Data Zone deployments process prompts and responses only within the stated zone; the European Union zone follows the Azure EU Data Boundary. The GPT-5.6 line is listed in nine European regions; Claude on Azure has a US data zone and no EU one.
Google Cloud, GeminiNot stated in the documentation we checked. The model vendor's terms govern this.Per model: at-rest data residency plus machine learning processing in the EU multi-region for Gemini 3.5 Flash and Gemini 2.5 Pro, the eu multi-region only for Gemini 3.8 Flash, and neither for Gemini 3.1 Pro preview.
AWS European Sovereign CloudNot stated in the documentation we checked. The model vendor's terms govern this.A separate European legal structure, operated exclusively by EU residents, physically and logically separate from other AWS regions. Amazon Nova Pro is the only foundation model listed for the region.
IONOS AI Model HubNo. IONOS states that prompts and outputs are not used to train, fine-tune or otherwise improve any model.Inference in German data centres and nowhere else, never logged and never accessed, not disclosed to third parties, within the scope of BSI C5.

Three of these rows are cloud platforms rather than model vendors, so the training question belongs to whoever built the model you deploy on them. Ask that vendor, and file the answer with the hosting decision.

What to do with this

Write down which route you are on, which model that leaves you, and who signed for it. That single page answers the first three questions of every auditor, insurer, works council and customer procurement team, and it is the artefact the AI Act's documentation culture rewards.

The CTO view of this decision / How we handle security questions

This section describes third-party offerings only, checked against provider documentation on 8 September 2026. It makes no claim about where AI Board itself runs or which certifications it holds.

07Which model for which situation

Six situations, and what to start with in each

Positioning, not measurement. Every model below is named for an attribute its vendor publishes, stated in the same breath as the name.

A board does not have a model problem. It has half a dozen recurring situations, and each narrows the field before quality enters the conversation. As the sovereignty chapter puts it: a model you are not allowed to use scores zero.

You are a CEO who wants one assistant that reads everything

Start with
Claude Fable 5.1 and GPT-6 Astra: both publish a window of roughly one million tokens, over eighty board papers in one conversation at this report's conversion, and both publish a reasoning mode. Our data records an EU route on Google Cloud for Fable 5.1 and none for GPT-6 Astra.
Test first
Hand it your last four board packs and ask which three decisions changed between the first and the fourth, with the page behind each.

AI for the CEO

You are a CFO and every number has to be traceable

Start with
The mid tier, tested before you pay for the flagship. Claude Sonnet 5 and GPT-5.6 Terra publish a window of the same order as the flagships and a reasoning mode, yet cost about a fifth as much on the model below. Claude Opus 5 sits between them, with a documented EU inference route on Amazon Bedrock.
Test first
Twenty questions whose answers each sit in exactly one appendix. Count separately how many answers name a document and page, and how many name the right one.

AI for the CFO

You are a CTO who has to sign off on where the data lives

Start with
The EU-region branch of this report's decision tree, where three clouds document a route: Claude Opus 5, Sonnet 5 and Haiku 4.5 on the Amazon Bedrock EU inference profile, the GPT-5.6 tiers in nine European regions on Azure Data Zone Standard, and Gemini in European regions on Vertex AI. That list runs a generation behind the newest flagships, and GPT-6 Astra is not on it.
Test first
Have each vendor name the processing region and the contracting entity in writing, then run your ten hardest questions on the EU endpoint and on the global one.

AI for the CTO

You want the lowest defensible cost per executive per week

Start with
The chart below, which applies one stated set of assumptions to the current models in our pricing data. Gemini 3.8 Flash comes out cheapest and DeepSeek-V4-Pro next, with Claude Sonnet 5 and GPT-5.6 Terra close behind. Two caveats: Google calls the Gemini price introductory through 31 December 2026, doubling on 1 January 2027, and the DeepSeek figure is the peak-hours rate.
Test first
Run one week of genuine questions on a cheap tier and on a flagship, and count how many cheap-tier answers you had to check by hand. Cheap is only cheap if checking is free.

Which model for which executive role

The model has to run on hardware you control

Start with
DeepSeek-V4-Pro, published under the MIT License with a one-million-token window, is the strongest genuinely permissive option in our data. Below it sit Mistral Large 3, Qwen3.8-27B and Muse Glimmer 30B, all Apache-2.0, and GLM-5.2 under MIT. Mistral Medium 3.5 also publishes its weights, but under a Modified MIT licence with exceptions above a revenue threshold, so it belongs on this list only after your lawyer has read that licence.
Test first
Before anyone buys a GPU, rent one for a week and run your twenty hardest questions on it. Measure cost per useful answer over three years, not the licence fee.

AI for the CTO

The assistant should be the one already in the suite you pay for

Start with
What the suite vendors publish themselves. Microsoft documents OpenAI GPT models and, since January 2026, Anthropic as a subprocessor, naming Claude Fable 5.0 and 5.1, and lets a user pick Claude in Researcher. Google Workspace runs on Gemini only, with a Flash tier for everyday work and a Pro tier still carrying a preview model id.
Test first
Ask the same five questions in the suite assistant and in a direct run of the model you would otherwise buy. If the suite answer is good enough, the procurement is over.

Which model for which executive role

The published attributes behind those starting points

Starting points per situation, with the published attributes behind them. Source: vendor documentation, read 8 September 2026. Cost is an estimate on stated assumptions, not a quote.
SituationStarting pointContext windowEU residency routeOpen weightsCost per working week
Board pack synthesisClaude Fable 5.11M tokensDocumentedNot publishedUS$10.62
Traceable finance answersClaude Sonnet 51M tokensDocumentedNot publishedUS$2.38
EU residency firstClaude Opus 51M tokensDocumentedNot publishedUS$5.94
Lowest cost per working weekGemini 3.8 Flash1.05M tokensDocumentedNot publishedUS$0.89
Fully self-hostedDeepSeek-V4-Pro1M tokensDocumentedPublishedMITUS$1.26

The sixth situation has no row on purpose: Microsoft and Google publish the family behind their assistants, not the version, so no window, route or token price honestly belongs in these columns.

Cost per executive per working week

Price per million tokens is the wrong unit for a board, because nobody buys a million tokens. The unit a director can argue with is one executive for one working week, built from the assumptions beside the chart. Where a lab publishes no cached-input price no discount is applied, which is why Mistral and Alibaba sit higher here than their headline prices suggest.

  • GPT-6 AstraUS$11.88
  • Claude Fable 5.1US$10.62
  • Claude Opus 5US$5.94
  • Qwen3.8-MaxUS$5.16
  • Mistral Medium 3.5US$4.05
  • Grok 4.6US$2.64
  • GPT-5.6 TerraUS$2.50
  • Claude Sonnet 5US$2.38
  • DeepSeek-V4-ProUS$1.26
  • Gemini 3.8 FlashUS$0.89
Estimated cost in US dollars for one executive for one working week, per model. The dearest model in the list costs more than thirteen times the cheapest.
LabelValue
GPT-6 AstraUS$11.88
Claude Fable 5.1US$10.62
Claude Opus 5US$5.94
Qwen3.8-MaxUS$5.16
Mistral Medium 3.5US$4.05
Grok 4.6US$2.64
GPT-5.6 TerraUS$2.50
Claude Sonnet 5US$2.38
DeepSeek-V4-ProUS$1.26
Gemini 3.8 FlashUS$0.89
Estimate on published API list prices, read 8 September 2026. Excludes VAT, volume discounts, batch and off-peak rates, and the retrieval platform.

The assumptions, in full

  • One executive asks the assistant 40 questions in a working week, which is eight a day.
  • Each question carries 60,000 input tokens of context, roughly 90 pages of board papers, minutes and appendices pulled in alongside the question itself.
  • Each answer is 1,500 output tokens, about two pages.
  • Where the lab publishes a cached-input price, 70 percent of those input tokens are served from cache, because the same document set is re-read all week. Where the lab publishes no cache price, no discount is applied.
  • All figures are public list prices in US dollars, excluding VAT, excluding volume discounts, excluding the batch and off-peak rates several labs offer, and excluding the cost of the platform that does the retrieving.

What comes next

This report is a map of what the vendors publish. It carries no scores, because scoring models properly means running the same documents and questions past all of them and publishing both.

The first edition of the AI Board Model Index is planned for November 2026: the current models in this report, put through one fixed set of board documents and questions, scored on things a director can argue with. It does not exist yet. There are no results on this page and none on the Model Index page, and any scorecard shown there is labelled an illustrative example of the format, with no model name beside it.

What we re-check every quarter

  • List prices per million tokens, cached-input rates, and when promotional pricing ends.
  • Published context windows, and whether the vendor still publishes one at all.
  • EU residency routes per model and cloud, including models excluded from a vendor's own EU boundary.
  • Model status: what is retiring or retired, and which successor the vendor names.
  • Licences on published weights, which can change at the next release.
  • Which models the office suites name, and whether they publish versions rather than families.

The honest limits of this report

Every fact here was read on 8 September 2026 and is stated as of that date. Vendor pages change without notice and without a changelog, and some will have changed before you read this. Where one page contradicted another, we say so rather than pick the tidier number.

Some things we could not confirm: Microsoft publishes no per-user price for the enterprise Copilot add-on, Alibaba publishes a price for Qwen3.8-Max on its international endpoint only, not for the China-domestic service, and several European providers serve open models at a smaller window than the model's maximum. Those gaps are printed as gaps. AI Board Studio has no model of its own and takes no payment from vendors.

Read about the Model Index and its method. Which AI model for which executive role.

08Method and limits

How this report was made, and what it leaves out

Every fact above was read off a vendor page on 8 September 2026. This closing section states the method, the deliberate omissions, the correction route, and the sources in full.

How these facts were gathered

Every model fact in this report comes from the company that ships the model: its own documentation, release notes, pricing page or announcement, opened on 8 September 2026. Primary sources only. A specification counts when the vendor publishes it, not when a news article, a leaderboard or a reseller repeats it. Where a vendor page and an external tracker disagreed, the vendor page won. Where two pages from the same vendor disagreed with each other, the item was left out rather than resolved by guesswork.

What is deliberately not in this report

Two categories were left out on purpose. The first is vendor benchmark percentages that could not be traced back to a named test, a named model version and a named comparison. Launch posts are full of them. A claim that one model is a given percentage better than another, with no test and no baseline, is an advertisement wearing the clothes of a measurement, so it was not reprinted here.

The second is scores. This report contains no measured ranking of any kind. Judging models against one fixed document set and one fixed list of questions is the job of the AI Board Model Index, whose first edition is planned for November 2026. Until that measurement exists, a number next to a model name would be exactly the behaviour this report criticises in vendors. Where the report says which model suits which situation, that is a reading of published capabilities, labelled as positioning.

How to report a correction

We would rather be corrected than quoted wrongly. If a figure above is out of date, or was wrong on the day it was published, tell us and send the vendor page that proves it. A correction backed by a primary source is applied, and the accessed date next to that source moves with it. The contact route is on the about page, linked below.

Contact details are on the about page

What happens next to this page

The intention is to re-check this report every quarter: reopen every source, refresh the prices, the regions and the retirement dates, and restate the date at the top. That is an intention rather than a guarantee, and we would rather say so plainly than promise a schedule a small team may miss. If a lab ships something that changes the picture sooner, the intention is to update sooner.

Questions a board asks about this

Twelve questions that come up in almost every management team when this subject reaches the agenda. The answers stay deliberately free of model specifications, because those change faster than an FAQ can.

Which AI model is safest for board documents?
There is no single safe model, and any vendor who says otherwise is selling something. Safety for board work is four separate things: does the model stay inside your own documents, does it say so when an answer is not there, does it point to the document and page it used, and can it run where your data is allowed to live. The current flagship models from the large labs are the strongest starting point on the first three, and the report above shows which of them can run under European data residency. The honest answer is that you test a shortlist on your own documents before anything sensitive goes near it.
Can ChatGPT be hosted in the EU?
The consumer app is a product, not a hosting choice. The models behind it can be reached through cloud platforms that offer European regions, which is how most organisations satisfy a data residency requirement without leaving the model their people already know. Two caveats belong in the minutes. Data residency in an EU region is not the same as sovereignty, because an American provider remains subject to American law. And the available regions and terms differ per model and change over time, so verify them in the vendor documentation at the moment you sign. The hosting section of the report above lists the routes per lab.
What does a 1 million token context window mean in pages?
A token is a fragment of text, on average a little less than a word. The context window is how much text a model can hold in one conversation, your documents and its answer together. As a rule of thumb, a window of a million tokens is several hundred thousand words, which is the order of magnitude of a year of board papers plus the annual report. Two warnings. Vendors state the maximum the model accepts, not the point where quality holds up, and a model that can hold a document set does not always use it correctly deep inside it. The Model Index measures that practical limit: /en/research/model-index.
Is an open-weight model good enough for finance?
Open-weight models can be downloaded and run on hardware you control, which is why finance and legal teams look at them first. In 2026 the strongest open models are genuinely usable for reading, summarising and drafting. The gap to the closed flagships shows up in the hard cases: synthesis across many documents, and the discipline to say that a figure is simply not in the set. For finance work, citation quality decides it, not general intelligence. Run the open candidate and a closed flagship on the same twenty questions from your own accounts before you commit to either. The role guide sets out that test for the CFO: /en/ai-cfo.
We tested an AI model in 2023 and it invented numbers. Has that changed?
Partly, and the part that changed matters. The models of 2023 had a small memory, no reasoning step and weak source discipline. The current generation holds far more text, works through a problem before answering, and is markedly better at pointing to the document it used. Vendors report large reductions in confident errors between generations; those are their own figures, but the direction is not in dispute. What has not changed is that every model can still state something that is not in your documents. So the 2023 verdict is out of date and the 2023 habit of checking anything that leaves the building is not. The report above lists which models are retired.
What is a reasoning model, and does a board need one?
A reasoning model works through a problem in steps before it answers, instead of producing the first plausible response. That extra step costs seconds and money, and it buys accuracy on exactly the work a management team cares about: comparing documents, checking a number against its source, noticing that two papers contradict each other. In the current flagships reasoning is built in and can be turned up or down per question. For a quick summary of one memo, keep it low. For anything that ends up in a board pack, turn it up and accept the wait. The report above describes how this shift happened.
Do the vendors' own benchmark figures mean anything?
They mean something as direction and very little as proof. A vendor figure comes from a test the vendor chose, on prompts the vendor wrote, against the comparison the vendor picked. That is marketing with real measurement inside it. Independent leaderboards help, but they mostly measure exam questions, coding and research puzzles, not whether a model invents a figure in a board memo or bluffs when a quarter is missing. The only number that decides your case is the one you produce on your own documents. That is why the Model Index runs one fixed test across every model: /en/research/model-index.
Does the EU AI Act mean we cannot use an American model?
No. The Act regulates how AI is used and which obligations follow from the risk of that use, not the nationality of the supplier. Using a general model to read and summarise your own documents sits at the light end, with transparency, oversight and documentation duties rather than prohibitions. Obligations tighten sharply when AI touches hiring, credit, safety or citizens. The practical board task is to write down which uses you allow, who checks the output before it is acted on, and where the data lives, and to have counsel confirm the classification. The regulatory timeline in the report above lists the dates.
Is a model from a Chinese or American lab a risk for a European company?
The risk is not the nationality of the lab, it is where your documents travel and who can compel access to them. Sending board material to any interface outside your control is a decision worth recording. For models published as open weights the question largely disappears, because you can run the weights on your own hardware or with a European provider and nothing leaves your perimeter. For closed models, a European region limits exposure but does not remove foreign legal reach. Decide the hosting requirement first, then choose from what satisfies it. Our own security approach is here: /en/security.
What happens when the model we standardised on is retired?
New versions arrive every few weeks and older ones are switched off with a few months of notice, which is short for an organisation that has built a process around one of them. Protect yourself by keeping the model replaceable. Keep your documents, prompts and house rules in a system you own rather than inside one vendor's product, and keep a written test you can rerun on a new model in an afternoon. Then a retirement notice is a scheduled task instead of a project. The report above lists which models are already retired or scheduled for shutdown, with the successor to move to.
What is the difference between the model, the assistant and the product we buy?
The model is the engine: a trained system that turns text going in into text coming out. The assistant is everything around it: your documents, the permissions, the memory, the house rules, the connections to mail and files. The product is the packaging of both. The distinction is commercial, not academic, because the model is the part you can swap and the assistant is where your investment accumulates. AI Board is that assistant layer and stays model independent, which is also why we can measure engines without having one to defend. See what the assistant layer holds: /en/company-brain.
Where do the facts in this report come from, and how current are they?
Every model fact here comes from the vendor's own documentation or announcement, checked shortly before publication, with the source listed next to the claim. Prices, regions and availability change quickly, so treat them as correct on the date stated and verify before you sign anything. Where a figure is a vendor's own benchmark we label it as such rather than presenting it as an independent result. We take no payment from model vendors and run no affiliate links. Corrections are genuinely welcome and reach us through the contact details on /en/about.

Sources

  1. 01OpenAI model documentationGPT-4 at 8,192 tokens and GPT-6 Astra at 1,050,000 tokens.Accessed 8 Sept 2026
  2. 02Microsoft Azure blog: GPT-6 Astra in Microsoft FoundryDeployment options for GPT-6 Astra: Global and US Data Zone geographies only.Accessed 8 Sept 2026
  3. 03Anthropic model overviewClaude Fable 5.1 context window of 1,000,000 tokens.Accessed 8 Sept 2026
  4. 04Anthropic: Claude on Google CloudGlobal, multi-region (us and eu) and regional endpoints for Claude.Accessed 8 Sept 2026
  5. 05Gemini API: Gemini 3.8 FlashInput window of 1,048,576 tokens.Accessed 8 Sept 2026
  6. 06Google, new features for the Gemini API and Google AI Studio27 June 2024: the two million token window on Gemini 1.5 Pro opened to all developers.Accessed 8 Sept 2026
  7. 07Meta AI, The Llama 4 herdMeta states a ten million token window for Llama 4 Scout, evidenced by a retrieval-style evaluation.Accessed 8 Sept 2026
  8. 08RULER: What's the Real Context Size of Your Long-Context Language Models? (arXiv:2404.06654)Hsieh et al., 2024. Only half of the models claiming 32K or more maintain satisfactory performance at 32K.Accessed 8 Sept 2026
  9. 09OpenAI API pricingList prices per million tokens, including promotional pricing notes.Accessed 8 Sept 2026
  10. 10Microsoft Foundry: models sold directly by AzurePer-model release dates, context windows and knowledge cutoffs.Accessed 8 Sept 2026
  11. 11Anthropic pricingAccessed 8 Sept 2026
  12. 12Anthropic API release notesAnnouncement dates per model.Accessed 8 Sept 2026
  13. 13Anthropic: introducing Claude Fable 5.1 and Claude Mythos 5.1Accessed 8 Sept 2026
  14. 14Anthropic, Claude in Amazon BedrockRegion table, global versus regional endpoints, EU inference profiles.Accessed 8 Sept 2026
  15. 15Anthropic: Claude in Microsoft FoundryGlobal Standard and US Data Zone Standard deployment types only.Accessed 8 Sept 2026
  16. 16Anthropic: data residencyinference_geo supports only 'us' and 'global'; workspace geo only 'us'.Accessed 8 Sept 2026
  17. 17Gemini API model listAccessed 8 Sept 2026
  18. 18Gemini API pricingIntroductory rates through 31 December 2026.Accessed 8 Sept 2026
  19. 19xAI model documentation: grok-4.6500K context window, pricing tiers, serving regions us-east-1 and us-west-2.Accessed 8 Sept 2026
  20. 20xAI model documentationContext windows and tiered input and output pricing.Accessed 8 Sept 2026
  21. 21Meta AI Research: introducing Muse Spark 1.3Released 2 September 2026 in Muse Code and the Meta Model API.Accessed 8 Sept 2026
  22. 22Meta: Muse Spark model page1M token context window, standard and contributor endpoints.Accessed 8 Sept 2026
  23. 23Meta AI for developersAccessed 8 Sept 2026
  24. 24DeepSeek API docs, changelogDated release notes: V4-Flash on 31 July 2026, V4-Pro general availability on 13 August 2026, vision preview on 21 August 2026, legacy endpoint shutdown on 24 July 2026.Accessed 8 Sept 2026
  25. 25DeepSeek API docs, Models and PricingContext length, max output and the peak and off-peak token prices for every current DeepSeek model.Accessed 8 Sept 2026
  26. 26Hugging Face, deepseek-ai/DeepSeek-V4-Pro model card1.6T parameters with 49B activated, one million token context, MIT License, FP8 weights.Accessed 8 Sept 2026
  27. 27Scaleway, Generative APIs supported modelsEuropean serverless inference catalogue: DeepSeek-V4-Flash-0731, Mistral Medium 3.5, GLM-5.2 and Qwen models with their served context windows.Accessed 8 Sept 2026
  28. 28AWS, DeepSeek models on Amazon BedrockBedrock lists DeepSeek-V3.1 and DeepSeek-R1. No V4 model was listed on 8 September 2026.Accessed 8 Sept 2026
  29. 29Mistral AI docs, models overviewCurrent premier and open models with API names and version stamps, plus the deprecation and retirement table.Accessed 8 Sept 2026
  30. 30Mistral AI, API pricingList price per million input and output tokens on La Plateforme.Accessed 8 Sept 2026
  31. 31Mistral AI, In-region inference, open models, and new European infrastructure for sovereign AIAnnouncement of 11 August 2026: Mistral Regional Endpoints generally available with a choice between Europe and the US, a priority tier in preview, and European Compute Units.Accessed 8 Sept 2026
  32. 32Hugging Face, mistralai/Mistral-Medium-3.5-128B model cardDense 128B parameters, 256k context, text and image input, reasoning effort none or high, Modified MIT License.Accessed 8 Sept 2026
  33. 33Mistral AI docs, deployment indexSelf-deployment via vLLM, TensorRT-LLM, TGI and SkyPilot; cloud availability on AWS Bedrock, Azure AI, Google Vertex AI, IBM watsonx.ai, Snowflake Cortex and Outscale.Accessed 8 Sept 2026
  34. 34Qwen Cloud docs, model changelogDated 2026 model entries for the Qwen3.8, Qwen3.7 and Qwen3.6 series.Accessed 8 Sept 2026
  35. 35Alibaba Cloud press room, Alibaba Unveils Qwen3.8-Max3 August 2026 launch, 2.4 trillion parameters with 95 billion activated, up to one million tokens of context, sparse MoE with hybrid attention, multimodal.Accessed 8 Sept 2026
  36. 36Alibaba Cloud, What is Model StudioSix Model Studio regions including Germany (Frankfurt, eu-central-1). Alibaba states that regions differ in endpoints, API keys, supported models, platform features and pricing.Accessed 8 Sept 2026
  37. 37South China Morning Post, Alibaba's Qwen3.8-Max made widely accessible ahead of open-weights releaseSecondary source, cited only for the reported list price of 2.00 USD per million input tokens and 6.00 USD per million output tokens.Accessed 8 Sept 2026
  38. 38Z.ai docs, release notesDated releases: GLM-5 on 12 February 2026, GLM-5.2 on 16 June 2026, GLM-5.3 on 18 August 2026, GLM-5.3-Flash on 26 August 2026.Accessed 8 Sept 2026
  39. 39Z.ai docs, pricingList price per million tokens in USD for every GLM model.Accessed 8 Sept 2026
  40. 40Z.ai docs, GLM-5.3 model guide1M token context, 128K max output, reasoning always enabled with low, high and max effort, text-only input, same base model as GLM-5.2 with post-training improvements.Accessed 8 Sept 2026
  41. 41Hugging Face, moonshotai/Kimi-K3 model cardJuly 2026, 2.8T total with 104B activated, 1,048,576 token context, text, image and video, reasoning always enabled with low, high and max effort, Kimi K3 License.Accessed 8 Sept 2026
  42. 42Kimi platform docs, chat pricing for Kimi K31,048,576 token context, USD 3.00 per million input tokens on cache miss, USD 15.00 per million output tokens.Accessed 8 Sept 2026
  43. 43Cohere docs, modelsCurrent Cohere model ids with context window and maximum output tokens.Accessed 8 Sept 2026
  44. 44Cohere docs, Command A+Released 20 May 2026, sparse mixture of experts, 218B total with 25B active, text and image input, 48 languages, Apache 2.0 weights on Hugging Face, runs on 1 x B200 or 2 x H100 at W4A4.Accessed 8 Sept 2026
  45. 45Cohere, pricingNo per-token list price is published for Command A+ on 8 September 2026; only legacy Command models are priced publicly, with newer models quoted as custom enterprise pricing.Accessed 8 Sept 2026
  46. 46Hugging Face, nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 model cardReleased 4 June 2026, 550B total with 55B active, context up to 1M tokens, OpenMDW-1.1 licence, text in and out, reasoning toggled with enable_thinking.Accessed 8 Sept 2026
  47. 47NVIDIA Technical Blog, Nemotron 3 Ultra powers faster, more efficient reasoning for long-running agentsNVIDIA's own positioning of the model for long-running agents and one-million-token context.Accessed 8 Sept 2026
  48. 48NoLiMa: Long-Context Evaluation Beyond Literal Matching (arXiv:2502.05167)Abstract: 'At 32K, for instance, 11 models drop below 50% of their strong short-length baselines.'Accessed 8 Sept 2026
  49. 49ICML 2025 proceedings entryAccessed 8 Sept 2026
  50. 50Lost in the Middle: How Language Models Use Long Contexts (TACL 2024)Abstract: performance is often highest at the beginning or end of the input context and significantly degrades in the middle.Accessed 8 Sept 2026
  51. 51Chroma, Context Rot: How Increasing Input Tokens Impacts LLM Performance18 models across four labs; performance degrades consistently with input length.Accessed 8 Sept 2026
  52. 52Google, Gemini API long context guideLong context limitations: with multiple needles 'the model does not perform with the same accuracy'.Accessed 8 Sept 2026
  53. 53Vectara, Introducing the next generation of Vectara's Hallucination LeaderboardOver 7,700 articles, up to 32K tokens; 'hallucination rates are generally higher under the new benchmark'.Accessed 8 Sept 2026
  54. 54OpenAI API docs, your dataNo training on API data by default, 30-day abuse-monitoring retention, Zero Data Retention, eu.api.openai.com regional processing.Accessed 8 Sept 2026
  55. 55Anthropic, Redeploying Fable 5Accessed 8 Sept 2026
  56. 56Anthropic, Introducing Claude Sonnet 5Accessed 8 Sept 2026
  57. 57OpenAI model deprecationsOpenAI's own list of announced deprecations and shutdown dates.Accessed 8 Sept 2026
  58. 58Meta AI, Introducing Muse Spark and the Meta Model APIAccessed 8 Sept 2026
  59. 59Anthropic, Introducing Claude Opus 5Accessed 8 Sept 2026
  60. 60Regulation (EU) 2026/1744 (Digital Omnibus on AI), Official JournalAccessed 8 Sept 2026
  61. 61European Commission, regulatory framework for AIAccessed 8 Sept 2026
  62. 62European Commission, transparency obligations under Article 50Accessed 8 Sept 2026
  63. 63Gemini API changelogGoogle's dated release and shutdown log for the Gemini API.Accessed 8 Sept 2026
  64. 64Vertex AI model versions and lifecycleGoogle Cloud's release and retirement dates per model version.Accessed 8 Sept 2026
  65. 65Anthropic, What is new in Claude Fable 5.1Accessed 8 Sept 2026
  66. 66OpenAI, GPT-6 AstraAccessed 8 Sept 2026
  67. 67OpenAI model documentation, GPT-48,192 context window; $30 per million input tokens, $60 per million output tokens.Accessed 8 Sept 2026
  68. 68OpenAI model documentation, GPT-4 Turbo128,000 context window; $10 per million input tokens.Accessed 8 Sept 2026
  69. 69OpenAI model documentation, GPT-4o128,000 context window; $2.50 per million input tokens.Accessed 8 Sept 2026
  70. 70OpenAI model documentation, GPT-4.11,047,576 context window; $2 per million input tokens.Accessed 8 Sept 2026
  71. 71OpenAI model documentation, GPT-5400,000 context window, 272,000 maximum input tokens; $1.25 per million input tokens.Accessed 8 Sept 2026
  72. 72OpenAI model documentation, GPT-5.6 Sol1,050,000 context window; $4 per million input tokens.Accessed 8 Sept 2026
  73. 73Anthropic, Introducing 100K context windows11 May 2023: context window expanded from 9K to 100K tokens.Accessed 8 Sept 2026
  74. 74Anthropic, Introducing Claude 2.121 November 2023: 200K token context window, described as roughly 150,000 words.Accessed 8 Sept 2026
  75. 75OpenAI, GPT-5 System Card (13 August 2025)Section 3.7. Production-representative prompts, browsing enabled: share of responses containing at least one major factual error.Accessed 8 Sept 2026
  76. 76OpenAI, o1 System Card (5 December 2024)PersonQA hallucination rate 0.30 for GPT-4o and 0.20 for o1.Accessed 8 Sept 2026
  77. 77OpenAI, o3 and o4-mini System Card (16 April 2025)Table 4. PersonQA hallucination rate: o1 0.16, o3 0.33, o4-mini 0.48.Accessed 8 Sept 2026
  78. 78Anthropic, Claude Opus 4.5 System Card (November 2025)Section 4.1: we are still far from removing factual hallucinations in the absence of external tools.Accessed 8 Sept 2026
  79. 79Anthropic, Claude Opus 5 System Card (24 July 2026)Section 6.5.1: accuracy 11 percent higher than Opus 4.8, hallucination rate 6 percent higher.Accessed 8 Sept 2026
  80. 80Google DeepMind, FACTS benchmark suite9 December 2025. Error rate cut 55 percent on FACTS Search and 35 percent on FACTS Parametric from Gemini 2.5 Pro to Gemini 3 Pro; every model tested scored below 70 percent overall.Accessed 8 Sept 2026
  81. 81Vectara hallucination leaderboard (HHEM-2.3)Updated 11 May 2026. Share of summaries that introduce facts not in the source document.Accessed 8 Sept 2026
  82. 82Inverse Scaling in Test-Time ComputeGema et al., Transactions on Machine Learning Research, December 2025. Longer reasoning can lower accuracy.Accessed 8 Sept 2026
  83. 83Microsoft, announcing Azure OpenAI Data Zones6 November 2024. EU Data Zone processes data within any EU member nation.Accessed 8 Sept 2026
  84. 84AWS, Amazon Bedrock is now available in the Europe (Frankfurt) region19 October 2023. Names participating vendors but no Claude version.Accessed 8 Sept 2026
  85. 85AWS, Claude 3.5 Sonnet and Haiku available in more regions7 August 2024. Claude 3.5 Sonnet in Europe (Frankfurt).Accessed 8 Sept 2026
  86. 86AWS, Claude 3.7 Sonnet now available in EuropeApril 2025. Cross-region inference across Ireland, Paris, Frankfurt and Stockholm.Accessed 8 Sept 2026
  87. 87Mistral AI, Au Large26 February 2024. Mistral Large safely hosted on Mistral infrastructure in Europe.Accessed 8 Sept 2026
  88. 88Regulation (EU) 2024/1689, the AI ActPublished in the Official Journal 12 July 2024, in force 1 August 2024. Article 113 sets the staged application dates.Accessed 8 Sept 2026
  89. 89Regulation (EU) 2026/1744, the Digital Omnibus on AIIn force 27 July 2026. Defers the high-risk obligations to 2 December 2027 for Annex III systems and 2 August 2028 for Annex I systems.Accessed 8 Sept 2026
  90. 90European Commission, the Commission starts enforcing AI Act rules31 July 2026. From 2 August 2026 the AI Office and national authorities begin enforcement.Accessed 8 Sept 2026
  91. 91OpenAI pricing page, archived 1 April 2023GPT-4 8K at $0.03 per 1K prompt tokens and $0.06 per 1K completion tokens, that is $30 and $60 per million.Accessed 8 Sept 2026
  92. 92OpenAI pricing page, archived 21 November 2023GPT-4 Turbo at $0.01 per 1K input tokens, that is $10 per million.Accessed 8 Sept 2026
  93. 93OpenAI API pricing page, archived 2 October 2024GPT-4o at $2.50 per million input tokens after the August 2024 update; GPT-4o mini at $0.15.Accessed 8 Sept 2026
  94. 94Anthropic pricing page, archived 20 June 2024Claude 2 and 2.1 at $8 per million input tokens; Claude 3 Opus at $15; Claude 3.5 Sonnet at $3.Accessed 8 Sept 2026
  95. 95Google Gemini API pricing page, archived 3 June 2024Gemini 1.5 Pro at $3.50 per million input tokens for prompts up to 128K tokens.Accessed 8 Sept 2026
  96. 96Google Gemini API pricing page, archived 16 February 2025Gemini 1.5 Pro repriced to $1.25 per million input tokens; Gemini 2.0 Flash at $0.10.Accessed 8 Sept 2026
  97. 97Epoch AI, LLM inference price trendsThe price of GPT-4 level performance on PhD level science questions fell about 40 times per year; across six benchmarks the range is 9 to 900 times per year.Accessed 8 Sept 2026
  98. 98Stanford HAI, 2025 AI Index ReportInference cost for a system performing at the level of GPT-3.5 fell more than 280 times between November 2022 and October 2024.Accessed 8 Sept 2026
  99. 99Anthropic, Reasoning models don't always say what they thinkAverage across hint types: Claude 3.7 Sonnet 25%, DeepSeek R1 39%.Accessed 8 Sept 2026
  100. 100Shojaee et al., The Illusion of Thinking (arXiv:2506.06941)Three regimes: standard models better at low complexity, reasoning models better at medium, both collapse at high complexity.Accessed 8 Sept 2026
  101. 101Laban, Hayashi, Zhou & Neville, LLMs Get Lost In Multi-Turn Conversation (arXiv:2505.06120)Laban, Hayashi, Zhou and Neville, 2025. Average 39 percent drop across six task types, 200,000 simulated conversations, 15 models.Accessed 8 Sept 2026
  102. 102Thinking Machines Lab, Defeating Nondeterminism in LLM Inference10 September 2025. 1,000 identical requests at temperature 0 produced 80 different completions.Accessed 8 Sept 2026
  103. 103OpenAI API reference, seed parameterBest effort sampling only. Determinism is not guaranteed.Accessed 8 Sept 2026
  104. 104Anthropic model deprecationsRetirement dates per Claude model and the 60 day notice commitment.Accessed 8 Sept 2026
  105. 105Google, Gemini API model deprecations and changelogGemini 1.0 Pro removed 18 February 2025; Gemini 1.5 shut down 29 September 2025; Gemini 2.0 Flash shut down 1 June 2026.Accessed 8 Sept 2026
  106. 106DeepSeek, model documentationOfficial model documentation for the lab.Accessed 8 Sept 2026
  107. 107Mistral AI, model documentationOfficial model documentation for the lab.Accessed 8 Sept 2026
  108. 108Google DeepMind model card: Gemini 3.1 ProRelease date 19 February 2026, 1M input tokens, 64K output tokens.Accessed 8 Sept 2026
  109. 109Muse Glimmer 30B model card (Hugging Face)Released August 2026, Apache 2.0, 131,072+ token context, text and image input.Accessed 8 Sept 2026
  110. 110Meta: Muse Glimmer model page30B parameters, Apache 2.0, built for local agents.Accessed 8 Sept 2026
  111. 111Hugging Face, deepseek-ai/DeepSeek-V4-Flash model card284B parameters with 13B activated, one million token context, MIT License, three reasoning effort modes.Accessed 8 Sept 2026
  112. 112Hugging Face, mistralai/Mistral-Large-3-675B-Instruct-2512 model card675B total with 41B active, 2.5B vision encoder, 256,000 token context, Apache 2.0.Accessed 8 Sept 2026
  113. 113Hugging Face, Qwen/Qwen3.8-2.4T-A95B model card262,144 tokens natively, extensible to 1,010,000; text only; thinking cannot be disabled; final response up to 131,072 tokens; licence qwen3.8-max.Accessed 8 Sept 2026
  114. 114Hugging Face, Qwen3.8-2.4T-A95B LICENSE fileCommercial use permitted; naming obligation above 100 million monthly active users or USD 20 million monthly revenue; separate licence required for model-as-a-service businesses above USD 50 million of revenue over twelve consecutive months.Accessed 8 Sept 2026
  115. 115Hugging Face, Qwen/Qwen3.8-27B model card27B dense, native vision-language, 262,144 token context extensible to about 1,000,000, 131,072 recommended output, Apache 2.0, thinking on by default with effort levels.Accessed 8 Sept 2026
  116. 116OVHcloud, AI Endpoints model catalogueQwen3.8-27B listed in the catalogue alongside other open-weight models.Accessed 8 Sept 2026
  117. 117IONOS Cloud docs, AI Model Hub modelsQwen 3.8-27B listed with a 262k context window; IONOS documents the hosting country per model on each model page.Accessed 8 Sept 2026
  118. 118Mistral, Medium 3.5 model cardAccessed 8 Sept 2026
  119. 119Mistral platform changelogAccessed 8 Sept 2026
  120. 120xAI, Grok 4.6Accessed 8 Sept 2026
  121. 121OpenAI, Introducing ChatGPTAccessed 8 Sept 2026
  122. 122OpenAI, GPT-4Accessed 8 Sept 2026
  123. 123Anthropic, Claude 2Accessed 8 Sept 2026
  124. 124OpenAI DevDay 2023 announcementsAccessed 8 Sept 2026
  125. 125Regulation (EU) 2024/1689 (AI Act), Official JournalAccessed 8 Sept 2026
  126. 126OpenAI, Introducing OpenAI o1-previewAccessed 8 Sept 2026
  127. 127OpenAI, Introducing GPT-5Accessed 8 Sept 2026
  128. 128Anthropic, Introducing Claude Fable 5 and Claude Mythos 5Accessed 8 Sept 2026
  129. 129AWS press release, AWS European Sovereign Cloud launchLaunch of the first region in Brandenburg, EU parent company and subsidiaries, more than 90 services.Accessed 8 Sept 2026
  130. 130AWS European Sovereign Cloud User Guide, Amazon BedrockAccessed 8 Sept 2026
  131. 131AWS European Sovereign Cloud, model support by Regioneusc-de-east-1 lists Amazon Nova Pro, in-region inference only.Accessed 8 Sept 2026
  132. 132Anthropic Privacy Center, model training on commercial productsAccessed 8 Sept 2026
  133. 133Anthropic Privacy Center, model training on consumer productsAccessed 8 Sept 2026
  134. 134Microsoft Learn, Claude models in Microsoft FoundryPage dated 2026-09-01. Data Zone Standard (US) only; no EU data zone listed for Claude.Accessed 8 Sept 2026
  135. 135Microsoft Learn, deployment types in Microsoft Foundry ModelsPage dated 2026-08-06. Definition of Global, Data Zone and Standard processing.Accessed 8 Sept 2026
  136. 136Microsoft Learn, region availability for Foundry Models sold by AzurePage dated 2026-09-03. Data Zone Standard, Europe tab: gpt-5.6 tiers listed, gpt-6-astra not listed.Accessed 8 Sept 2026
  137. 137Microsoft, EU Data Boundary completionAccessed 8 Sept 2026
  138. 138Microsoft Learn, learn about the EU Data BoundaryAccessed 8 Sept 2026
  139. 139Google Cloud docs, data residency for Gemini EnterpriseEU multi-region: at-rest DRZ and MLP supported for Gemini 3.5 Flash and Gemini 2.5 Pro; Gemini 3.8 Flash and the Claude Fable 5 line (Claude Fable 5.1 as of September 2026) on the eu multi-region endpoint only (no EU regional endpoint); Gemini 3.1 Pro preview no EU DRZ or MLP. Re-read 8 Sep 2026.Accessed 8 Sept 2026
  140. 140Google Cloud blog, Europe's path to open digital sovereigntyPublished 2026-06-18. Three sovereign tiers and named European partners.Accessed 8 Sept 2026
  141. 141IONOS AI Model Hub docs, data handlingInference in German data centres only; prompts never logged; not used for training.Accessed 8 Sept 2026
  142. 142Mistral AI, Le Chat Enterprise announcementPublished 2025-05-07. Self-hosted, private cloud or Mistral cloud deployment.Accessed 8 Sept 2026
  143. 143Hugging Face, deepseek-ai/DeepSeek-V4-Pro-0813 model card1.6T parameters; repository and weights under the MIT License.Accessed 8 Sept 2026
  144. 144Hugging Face, mistralai organisation pageAccessed 8 Sept 2026
  145. 145Hugging Face, deepseek-ai organisation pageAccessed 8 Sept 2026
  146. 146EU AI Act Explorer, implementation timelineAccessed 8 Sept 2026
  147. 147EU AI Act Explorer, Article 4 AI literacyAccessed 8 Sept 2026
  148. 148EU AI Act Explorer, Digital Omnibus on AIRegulation (EU) 2026/1744, in force 27 July 2026.Accessed 8 Sept 2026
  149. 149Cheng et al., ELEPHANT: Measuring and understanding social sycophancy in LLMs (arXiv:2505.13995)11 models; face preservation 45 percentage points above humans; models affirm whichever side the user adopts in 48% of cases.Accessed 8 Sept 2026
  150. 150Magesh et al., Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, Journal of Empirical Legal Studies 22(2), 2025Abstract: Lexis+ AI and the Thomson Reuters tools 'each hallucinate between 17% and 33% of the time'; providers' claims 'are overstated'.Accessed 8 Sept 2026
  151. 151Hugging Face, zai-org/GLM-5.2 model card753B parameters, one million token context, MIT licence, multiple thinking effort levels.Accessed 8 Sept 2026
  152. 152xAI model documentationGrok prices per million tokens, split below and above a 200K-token prompt.Accessed 8 Sept 2026
  153. 153Alibaba Cloud Model Studio billingUSD prices per million tokens on the international (Singapore) endpoint.Accessed 8 Sept 2026
  154. 154Microsoft Learn, Data, privacy and security for Microsoft CopilotPrompts, responses and Microsoft Graph data are not used to train foundation models. EU traffic stays inside the EU Data Boundary, except for Anthropic models. Prompts and responses are stored as Copilot activity history and are governed by Microsoft Purview retention policies.Accessed 8 Sept 2026
  155. 155Microsoft Learn, Anthropic models in Microsoft Online ServicesPage dated 2 September 2026. Anthropic is a Microsoft subprocessor; Anthropic models are excluded from the EU Data Boundary and are off by default in the EU, EFTA and the UK. Names Fable 5.0 and Fable 5.1 as Fable-class models.Accessed 8 Sept 2026
  156. 156Microsoft Learn, OpenAI as a subprocessor in Microsoft Online ServicesPage dated 5 August 2026. OpenAI added to the subprocessor list on 23 June 2026, usable from 9 July 2026, enabled for all users of eligible commercial customers from 24 July 2026. OpenAI-operated models are inside the EU Data Boundary.Accessed 8 Sept 2026
  157. 157Microsoft Learn, Understanding AI functionality and models in Microsoft Online ServicesDistinguishes models hosted and operated by Microsoft, AI subprocessors, and AI independent processors. For Microsoft-hosted models, data does not leave Microsoft.Accessed 8 Sept 2026
  158. 158Microsoft 365 blog, Expanding model choice in Microsoft 365 Copilot24 September 2025. First announcement of Anthropic models in the Researcher agent and in Copilot Studio. Superseded on the hosting question by the September 2026 subprocessor documentation.Accessed 8 Sept 2026
  159. 159Microsoft Learn, Microsoft Foundry Models overviewPage dated 28 July 2026. Catalogue of over 10,000 models from Microsoft, Azure OpenAI, Anthropic, DeepSeek, Meta, Mistral, Cohere and Hugging Face, split into models sold by Azure and models from partners and community.Accessed 8 Sept 2026
  160. 160Microsoft 365 Copilot pricingCopilot Business: USD 18 per user per month with an annual commitment (promotional to 31 December 2026), list USD 21, USD 25.20 month to month, on top of a qualifying Microsoft 365 plan, up to 300 users. Verified in pricing.ts; the enterprise add-on price was not readable.Accessed 8 Sept 2026
  161. 161Google Workspace, Google Workspace with GeminiWorkspace plans include the Gemini app, Gemini Notebook and Gemini in Gmail, Docs and Meet. Submissions are not used to train models and are not reviewed by humans.Accessed 8 Sept 2026
  162. 162Google Workspace, Generative AI in Google Workspace Privacy HubLast updated 14 August 2026. Workspace does not use customer data to train models without the customer's prior permission or instruction; content is not human reviewed outside the domain. Gemini in Workspace prompt retention is admin controlled; Gemini Notebook content is not retained after the session ends.Accessed 8 Sept 2026
  163. 163Google, Gemini Apps release updates and improvementsGemini 3.6 Flash released to the Gemini app on 21 July 2026; the app's picker offers Flash tiers for everyday work and Pro for harder reasoning.Accessed 8 Sept 2026
  164. 164Google Workspace pricingEuro list prices read on 8 September 2026: Business Starter EUR 6.80, Standard EUR 13.60, Plus EUR 21.10 per user per month, Enterprise on request. Promotional prices were shown alongside. An AI Expanded Access add-on is offered without a price on the page.Accessed 8 Sept 2026
  165. 165Google Cloud, Use Anthropic Claude models on Vertex AIVertex AI lists Anthropic partner models including Claude Fable 5.1, Claude Opus 5 and Claude Sonnet 5. Vertex AI is a Google Cloud developer platform, not the Workspace assistant.Accessed 8 Sept 2026

The landscape above will look different in a quarter, and that is the point: the organisations that come out ahead are not the ones that picked the cleverest model in September 2026, they are the ones that built a way of reading, testing and replacing models as a routine. Read the report, run your own twenty questions on your own documents, and keep the assistant layer around the model in your own hands.

AI Board Studio is model independent. We have no model of our own, we receive no payment from model vendors, and we use no affiliate links.

Become AI-native before your competition does

Ride the AI wave instead of swimming behind it. Request access and we'll schedule your install.

Your company brain lives on your laptop. You choose if a question goes to a cloud model or stays fully local.