The state of AI 2026 comes down to three things that happened in the first eight months of the year — and almost nobody has put them in the same sentence.
A model crossed 96% on SWE-bench Verified. Not 96% on a toy benchmark — 96% on real GitHub issues from Django and Flask and scikit-learn, the benchmark that was supposed to take years. It is now effectively finished as a discriminator, and the three models at the top of it are separated by less than a percentage point.
The best available evidence on whether AI coding tools make experienced developers faster still says no. METR’s randomised controlled trial found a 19% slowdown. Their own attempt to re-run it in 2026 collapsed because developers refused to work without AI, even at $50 an hour. The people using these tools believe they are three times faster. Measured, they are not.
And Microsoft shipped a 14-billion-parameter reasoning model into Windows itself, which is a sentence that would have been science fiction in 2024 and is now a footnote in a Build keynote.
Those three facts are the whole state of AI 2026 in miniature. The models got extraordinary. The measured productivity gain did not follow. And the interesting frontier quietly moved off the datacentre and onto the device in your hand.
This is the long version of the state of AI 2026. It is written for people who have to make decisions with money attached — which model to route to, whether to let agents write production code, what to buy, and what to tell a client or a data protection officer. I have separated what shipped from what is reported from what is rumour, because the single most expensive mistake in this field is planning around a model that does not exist yet.
Last updated 3 September 2026. It will need updating again by Christmas.
TL;DR — The State of AI 2026 in 90 Seconds
The frontier is a four-horse race and Anthropic is currently winning it on the aggregate. As of 2 September 2026, Artificial Analysis’s Intelligence Index v4.1.1 has Claude Fable 5.1 at 65.7, Claude Opus 5 at 63.0, Grok 4.6 at 60.9 and GPT-5.6 Sol at 58.9. But the aggregate hides the interesting part: GPT-5.6 Sol leads BrowseComp at 90.4%, Gemini 3.8 Flash delivers 58.7 at $0.75 per million input tokens, and the top ten models span just 8.2 points — a range narrow enough that harness and prompt differences can move a model several places.
Open weights are roughly one generation behind and closing. GLM-5.3 (753B, open) sits at 59.5 — sixth overall, ahead of GPT-5.6 Sol. DeepSeek V4 Pro shipped 1.6 trillion parameters under an actual MIT licence. Ornith-1.5-397B hit 86 on SWE-bench Verified and edged past Claude Opus 4.8 on Terminal-Bench 2.1, 86.1 to 85.0. The capability gap is now about six months. The licence gap is the one that matters.
“Vibe coding” won the adoption argument and lost the evidence argument. 90% of developers use at least one AI coding tool at work. The enterprise coding-agent market is running at roughly $10 billion annualised. And developer trust in AI output fell from 40% to 29% in two years. The tools are universal, and the people using them trust them less every year. Both of those things are rational.
The security bill has arrived and it is itemised. Veracode found 45% of AI-generated code samples introduce an OWASP Top 10 vulnerability. Apiiro measured AI-assisted developers producing commits three to four times faster and security findings ten times faster. Roughly 20% of AI-generated code references packages that do not exist. This is not an argument against the tools. It is an argument about where you put the gate.
Offline AI became genuinely useful this year, on hardware people already own. A 2-billion-parameter model runs at 61 tokens per second on an iPhone. Microsoft put a 14B reasoner into Windows behind a 40-TOPS NPU requirement. Google gates its best on-device model behind 12 GB of RAM, and so does Apple. The specification that decides what you can run offline is memory, not compute, and memory is the thing that got scarce.
The honest one-liner on the state of AI 2026: this is the year the models stopped being the bottleneck. Verification became the bottleneck. Every important decision in this article — which model, which harness, which device, which gate — is downstream of that one change.

Quick Comparison: The Frontier, September 2026
Artificial Analysis Intelligence Index v4.1.1 aggregates nine evaluations: GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience and AA-LCR. Scores below are the 2 September 2026 snapshot.
| Model | Lab | AA Index | Weights | Input / Output per 1M |
|---|---|---|---|---|
| Claude Fable 5.1 | Anthropic | 65.7 | Closed | $10.00 / $50.00 |
| Claude Opus 5 | Anthropic | 63.0 | Closed | $5.00 / $25.00 |
| Claude Fable 5 | Anthropic | 62.1 | Closed | $10.00 / $50.00 |
| Grok 4.6 | xAI | 60.9 | Closed | — |
| Kimi K3 | Moonshot | 59.7 | Open weight | — |
| GLM-5.3 | Z.ai | 59.5 | Open weight | $1.40 / $4.40 |
| GPT-5.6 Sol | OpenAI | 58.9 | Closed | $5.00 / $30.00 |
| Gemini 3.8 Flash | 58.7 | Closed | $0.75 / $3.75 | |
| Qwen3.8 Max Preview | Alibaba | 58.1 | Closed | — |
| GLM-5.3-Flash | Z.ai | 57.5 | Open weight | $0.15 / $0.47 |
| Claude Opus 4.8 | Anthropic | 57.3 | Closed | $5.00 / $25.00 |
| Muse Spark 1.2 | Meta | 56.8 | Closed | — |
| GPT-5.5 | OpenAI | 56.3 | Closed | — |
Read the fourth column before the third. Two of the top ten models have downloadable weights, and one of them costs $0.15 per million input tokens.
Quick Comparison: The Open-Weight Frontier
| Model | Total / Active | Context | Licence | Released |
|---|---|---|---|---|
| Kimi K3 | 2.8T / 104B | 1M | Revenue-tiered (custom) | July 2026 |
| Qwen3.8-2.4T-A95B | 2.4T / 95B | 262K (1M via YaRN) | Apache 2.0 | 14 Aug 2026 |
| DeepSeek V4 Pro | 1.6T / 49B | 1M | MIT | 12 Aug 2026 |
| GLM-5.3 | 753B / ~40B | 1M | Custom ($10B gate) | Aug 2026 |
| MiniMax M3 | 428B / 23B | 1M | Custom (attribution) | 1 June 2026 |
| Ornith-1.5-397B | 397B MoE | 262K (1M via YaRN) | MIT | Aug 2026 |
| Gemma 4 31B | 31B dense | 256K | Apache 2.0 | 6 Apr 2026 |
| Qwen3.8-27B | 27B dense | 262K | Apache 2.0 | 14 Aug 2026 |
| Gemma 4 E4B | ~8B | 128K | Apache 2.0 | 6 Apr 2026 |
Three of those nine are genuinely OSI-clean. That is the real story of open weights in 2026, and I will come back to it in Part Three.
ACT I — THE MODELS
Part One: What Actually Shipped in 2026
Before any argument, the ledger. I am separating this into three tiers because the difference between them is where money gets wasted.
Confirmed and shipping
- 6 April 2026 — Google releases Gemma 4 under Apache 2.0, a licence change from Gemma 3’s bespoke terms. Four variants: 31B dense, 26B with 4B active (MoE), E4B at ~8B, and E2B at 5.1B total with 2.3B active. 256K context on the large models, 128K on the E-series. All support text, image and video input; E2B and E4B add native audio.
- 1 June 2026 — MiniMax M3. 428B total / 23B active, 1M context, native multimodality, a new sparse attention operator, weights on Hugging Face. Covered in depth in my earlier piece on local AI hardware.
- 2 June 2026 — Microsoft announces Aion 1.0 at Build. A family of on-device models shipped into Windows 11 itself: Aion 1.0 Instruct (CPU-capable, in Edge Canary from 150.0.4070) and Aion 1.0 Plan, a 14-billion-parameter reasoning and tool-calling model with a 32K context window.
- 9 June 2026 — Claude Fable 5. Anthropic’s top tier, $10/$50 per million, 1M context, 128K max output.
- 9 July 2026 — OpenAI ships GPT-5.6 in three variants: Sol ($5/$30), Terra ($2.50/$15) and Luna ($1/$6). Cyber capability is built into Sol rather than sold as a separate model, with enhanced access via OpenAI’s Trusted Access for Cyber programme.
- July 2026 — Moonshot opens Kimi K3’s weights. 2.8 trillion parameters, 1M context, native MXFP4, under a revenue-tiered licence.
- 24 July 2026 — Claude Opus 5. $5/$25, unchanged from Opus 4.8. Anthropic claims it “more than doubles Opus 4.8’s performance” on Frontier-Bench v0.1 and scores three times the next-best model on ARC-AGI 3.
- 12 August 2026 — DeepSeek V4 Pro exits preview with MIT-licensed weights: 1.6T total (some coverage rounds this to 1.7T), 49B active, 1M context, 893 GB of BF16 and FP8 tensors, plus a DSpark speculative decoding module. Terminal-Bench 2.1 at 87.9, HLE-with-tools at 60.0, CyberGym at 83.3. DeepSeek raised API prices at the same time.
- 14 August 2026 — Qwen3.8 open weights under Apache 2.0: a 27B dense multimodal model and a 2.4T-A95B giant, both on Hugging Face and ModelScope, 262K native context extensible to 1M with YaRN.
- August 2026 — GLM-5.3 weights land on Hugging Face roughly two weeks after the API launch. 753B MoE, 1M context, 128K output, BF16 and FP8.
- August 2026 — Ornith-1.5. Three scales (397B MoE, 35B MoE with 3B active, 9B dense), built on Qwen3.5 and Gemma 4 with continued pretraining, MIT licensed, and claiming 86 on SWE-bench Verified.
- 1 September 2026 — Claude Fable 5.1 and Mythos 5.1. Adaptive thinking on by default, 1M context, 128K output. Base pricing unchanged at $10/$50, but cache reads drop 75%, from $1.00 to $0.25 per million.
- 2 September 2026 — Gemini 3.8 Flash at $0.75/$3.75, 1M input context, 66K output, March 2026 knowledge cutoff, with text, image, video, audio and PDF input.
- 2 September 2026 — Meta ships Muse Spark 1.3 and the Muse Code CLI.
Reported, credible, not confirmed
- Meta will open Muse Spark’s weights “soon.” That is Mark Zuckerberg on X, with no date, no parameter count, no architecture detail and no licence text. Meta has not disclosed the model’s size at any version.
- Gemini Nano v4 is described as forthcoming for Android, with Nano v3 currently restricted to the Pixel 10 series.
- Aion 1.0 Plan is “coming months” for general availability; only Instruct is in preview.
Rumour, single source, treat accordingly
- Everything you have read about the next Anthropic, OpenAI or Google flagship. In a market where four labs shipped a top-ten model in a nine-week window, forward rumours have a half-life of about a fortnight.
One more thing worth flagging, because it is new and almost nobody outside security is discussing it. Claude Mythos 5.1 scores 60.9 on Terminal-Bench 4.0 — higher than Fable 5.1’s 55.8 — and 95.5% on SWE-bench Verified. You cannot buy it.
Anthropic describes Mythos as its “most capable model for cybersecurity and biology research”, and it is available only to vetted cyberdefenders and life scientists through two trusted-access programmes: the Cyber Verification Program and the Life Sciences Verification Program. Anthropic’s own wording is that they are “only able to make it available to a set of US organizations, though we’re working to expand access.” Project Glasswing, the access-expansion effort, had reached roughly 150 organisations in more than fifteen countries as of June 2026.
The reasoning is dual-use — a model that is excellent at finding vulnerabilities is excellent at finding them for either side — and I think it is defensible. But note the structural consequence, because it is genuinely new: the highest-scoring model on a public coding benchmark is now one that most of the world cannot access at any price. Capability gating on safety grounds rather than commercial grounds is going to become more common, not less, and if you are outside the US it is an argument for caring about open weights that has nothing to do with cost.
Show Image Nine months, twelve significant releases, four of them open weight. Solid bars are shipped with weights or API access you can buy today; dotted outlines are announced-but-not-available.
Part Two: What Actually Separates The Frontier Models Now
Here is the uncomfortable truth about the top of that leaderboard: on most work you will actually do, the top eight models are interchangeable.
I do not say that to be contrarian. I say it because the aggregate index spread between Claude Fable 5.1 at 65.7 and GLM-5.3-Flash at 57.5 is eight points, and GLM-5.3-Flash costs $0.15 per million input tokens against Fable’s $10.00. That is a 67× price difference for a 12% capability difference on a composite score, and composite scores compress exactly the differences that matter to you while inflating the ones that do not.
So the useful question is not “which model is best.” It is “what is each of them actually best at, and does that overlap with my week.”
The four things that genuinely differ
1. Long-horizon agentic reliability. This is the real 2026 differentiator and it is badly measured. Anthropic’s Opus 5 is explicitly positioned as “a step change improvement for the Opus tier powering long-running agents” — not as a smarter model, as a model that does not fall over on hour six. The evidence for this class of improvement lives in agentic benchmarks (Terminal-Bench, OSWorld, Frontier-Bench) rather than in knowledge benchmarks, and those benchmarks are young, noisy and scaffold-sensitive.
2. Cache economics. Fable 5.1’s headline change was not a capability jump. It was cutting cache reads 75%, from $1.00 to $0.25 per million tokens — which MarkTechPost reckons is roughly 25% lower cost for typical workflows and up to 45% for agentic ones. For any agent that re-reads a large system prompt or codebase on every turn, cache pricing is a bigger lever on your monthly bill than the model choice. Nobody puts that on a leaderboard.
3. Browsing and tool use. GPT-5.6 Sol scores 90.4% on BrowseComp — 92.2% at the Ultra reasoning setting. That is a specific, published strength, and if your workload is research-and-synthesis rather than code-and-verify, it is worth more than two points of aggregate index.
4. Price-per-adequate-answer. Gemini 3.8 Flash at 58.7 index for $0.75 input is the most obviously mispriced thing on the board relative to the closed frontier. GLM-5.3-Flash at 57.5 for $0.15 is the most obviously mispriced thing full stop.
The pricing table, because it is the actual decision
| Model | Input /1M | Output /1M | Cached input /1M | Context |
|---|---|---|---|---|
| Claude Fable 5.1 | $10.00 | $50.00 | $0.25 | 1M / 128K out |
| Claude Opus 5 | $5.00 | $25.00 | — | 1M / 128K out |
| GPT-5.6 Sol | $5.00 | $30.00 | — | — |
| GPT-5.6 Terra | $2.50 | $15.00 | — | — |
| Claude Sonnet 5 | $2.00 | $10.00 | — | 1M / 128K out |
| GLM-5.3 | $1.40 | $4.40 | — | 1M / 128K out |
| GPT-5.6 Luna | $1.00 | $6.00 | — | — |
| Gemini 3.8 Flash | $0.75 | $3.75 | — | 1M / 66K out |
| GLM-5.3-Flash | $0.15 | $0.47 | — | — |
| MiniMax M3 | $0.30 | $1.20 | — | 1M |
| Muse Spark (Contributors) | $0.10 | $0.20 | $0.002 | — |
That last row deserves a second look. Meta’s Contributors tier is $0.10 in and $0.20 out — with potential access to your data for training. That is not a price; it is a trade. It is also, as far as I can tell, the most honest expression of the actual business model underneath a lot of cheap inference, and I would rather have it stated in a pricing table than buried in a DPA.
Show Image Intelligence on one axis, price on a log scale on the other. Eight points of index separate the top of the chart from the bottom; 67× separates the prices. The efficient frontier is the bottom-right, and almost nobody buys there.

The API breakage nobody warned you about
Fable 5.1 shipped with three breaking changes: forced tool use is gone, thinking blocks are now model-specific, and mid-conversation edits to system messages or tools now produce errors rather than being silently accepted.
That third one will bite people. A large number of agent harnesses mutate the system prompt or the tool list between turns — adding a tool when a subtask starts, trimming instructions to save context. On Fable 5.1 that is now an error. If you upgrade a model ID in a config file and your agent starts throwing 400s on turn four, that is why.
The general lesson, which is worth more than the specific one: in 2026 a model upgrade is an API migration. Write your model IDs so that swapping one is a single-line change, and pin them. You will be making that change again inside a quarter.
Part Three: The Open-Weight Frontier, And The Licence Trap
This is the part of the 2026 story I find genuinely exciting, and also the part where the marketing is most misleading.
The capability gap is roughly six months
Look at what open weights achieved between June and August 2026:
- DeepSeek V4 Pro: 1.6T total, 49B active, MIT licence, Terminal-Bench 2.1 at 87.9. That is a higher Terminal-Bench figure than most closed models published this year.
- Ornith-1.5-397B: 86 on SWE-bench Verified, 86.1 on Terminal-Bench 2.1 against Claude Opus 4.8’s 85.0, and 92.8 on GPQA Diamond. MIT licensed, 262K native context extensible to about 1M with YaRN. Built on top of Qwen3.5 and Gemma 4 with continued pretraining and post-training.
- GLM-5.3: sixth on the Artificial Analysis index at 59.5, ahead of GPT-5.6 Sol.
- Qwen3.8: a 2.4-trillion-parameter model under Apache 2.0, plus a 27B dense multimodal sibling that Alibaba says beats the larger Qwen3.7-Plus on coding and office tasks.
Six months ago the honest framing was “open weights are a generation behind and cheaper.” That framing is now wrong. The correct framing is: for coding and agentic work specifically, the best open-weight models are inside the error bars of last generation’s closed frontier, and they are 10× to 60× cheaper.
Ornith-1.5 is the most interesting model of the year and almost nobody has heard of it
I want to dwell on this one, because it is a genuinely novel thing rather than a scaling result.
Ornith describes 1.5 as extending “self-scaffolding into an end-to-end self-improvement loop.” In plain language: the system proposes its own new training tasks, generates task-specific scaffolds for them, produces solution rollouts, and optimises all three jointly with GRPO. Stronger models let it propose harder tasks; better scaffolds discover better strategies; better outputs give a cleaner learning signal.
It is trained on top of Qwen3.5 and Gemma 4 — that is, on other people’s open weights — and it beats Claude Opus 4.8 on Terminal-Bench 2.1 while being 397B parameters and MIT licensed.
If that loop generalises, it is the most consequential result in this article, because it means the open ecosystem can improve without any lab spending frontier-scale money on a pretraining run. If it does not generalise — if what we are seeing is benchmark-specific overfitting from a self-generated curriculum — then it is a cautionary tale about exactly the same thing. I do not know which. Neither, as far as I can tell, does anyone outside Ornith. What I would want before betting on it: an independent evaluation on tasks that were not in the self-generated curriculum’s neighbourhood, and a held-out human baseline. Neither exists yet.
Now the licences, because “open source” is doing a lot of dishonest work
Here is what those weights actually cost you in obligations:
| Model | Licence | The catch |
|---|---|---|
| DeepSeek V4 Pro | MIT | None. Genuinely open source. |
| Qwen3.8 | Apache 2.0 | None. Genuinely open source. |
| Gemma 4 | Apache 2.0 | None — and a real improvement on Gemma 3’s bespoke terms. |
| Ornith-1.5 | MIT | None stated. Verify the model card yourself before shipping. |
| GLM-5.3 | Custom | Companies over $10B annual revenue must pass Z.ai’s security review before any commercial use. Z.ai moved away from MIT for this release. |
| Kimi K3 | Custom | Revenue-tiered; prominent attribution required for high-volume providers. |
| MiniMax M3 | Custom | Commercial use requires prominent “Built with MiniMax M3” attribution; separate written authorisation above $20M annual revenue. |
| Muse Spark | Unknown | Announced as “coming soon.” No licence text exists. |
Three observations. I went through the fine print of the first three of these in detail in Kimi K3 vs Fable 5.1 vs GLM-5.3, so here I will stick to what changed since.
First, the gates are aimed at hyperscalers, not at you. A $10 billion revenue threshold excludes essentially every reader of this article. If you are a German Mittelstand business, an agency, a consultancy or a solo developer, the practical answer for GLM-5.3 and Kimi K3 is: you are fine.
Second, “you are fine” is not the same as “your legal review is fine.” A bespoke licence means somebody bills you to read it. An MIT or Apache 2.0 licence means somebody reads it in ninety seconds and moves on. For enterprise procurement that difference is worth more than four points of benchmark, and it is the single strongest argument for DeepSeek V4 Pro and Qwen3.8 over GLM-5.3 in a commercial product.
Third, watch the direction of travel. Z.ai went from MIT on GLM-5.2 to a proprietary licence with a security-review gate on GLM-5.3. DeepSeek went the other way and kept MIT while raising API prices — which is a coherent strategy: charge for convenience, give away the weights, let inference providers compete on serving. Open weights are not a movement. They are a commercial tactic, and tactics change per release. Do not build a product architecture that assumes next year’s version will carry this year’s licence.
The word “open source” should be retired for models
None of these are open source in the sense the OSI means, except the MIT and Apache 2.0 ones — and even those release weights without releasing training data or training code, which is a meaningfully different thing from open-source software.
The vocabulary I use, and would suggest:
- Open weight — you can download and run the parameters. Says nothing about terms.
- Openly licensed — MIT, Apache 2.0, or equivalent. No revenue gate, no attribution mandate, no field-of-use restriction.
- Open source — weights, training code and data composition, under an OSI-approved licence. Almost nothing at the frontier qualifies.
If a vendor says “open source model” and means the first one, that is not a lie exactly, but it is the kind of imprecision that ends up in a procurement document and then in a dispute.
Show Image Capability is converging. Licences are diverging. For most commercial buyers the right-hand column is the one that decides the purchase, not the score.
Part Four: The Benchmark Problem
SWE-bench Verified now looks like this:
| Rank | Model | Score |
|---|---|---|
| 1 | Claude Opus 5 | 96% |
| 2 | Claude Mythos 5 | 95.5% |
| 3 | Claude Fable 5 | 95% |
| 4 | Claude Opus 4.8 | 88.6% |
| 5 | Claude Opus 4.7 (Adaptive) | 87.6% |
| 6 | Ornith-1.5-397B | 86% |
| 7 | Claude Sonnet 5 | 85.2% |
| 8 | GPT-5.3 Codex | 85% |
Three models within one percentage point at the top. That benchmark is done. It cannot tell you anything useful about the difference between Opus 5, Mythos 5 and Fable 5, and anyone quoting it as a purchasing input in September 2026 is quoting a saturated instrument.
This is the general condition of AI evaluation right now, and it has four failure modes worth naming:
Saturation. When the top scores cluster above 90%, remaining differences are dominated by label noise and harness details rather than capability. SWE-bench Verified is there. GPQA Diamond is close.
Scaffold sensitivity. The same model scores differently depending on the agent loop wrapped around it — retry policy, tool set, context management, how many turns it gets. Terminal-Bench and OSWorld results are especially sensitive to this. Two labs reporting the same benchmark are frequently not running the same experiment.
Contamination. Every public benchmark eventually leaks into training corpora. The response has been an arms race of fresher, harder, private benchmarks — Frontier-Bench, Agents’ Last Exam, GDPval, CritPt, Terminal-Bench-Science — which solves contamination and creates a new problem: you cannot independently verify a private benchmark.
Selection. Labs publish the benchmarks they win. Anthropic led with Frontier-Bench and ARC-AGI-3 for Opus 5. OpenAI led with Agents’ Last Exam and a Coding Agent Index for GPT-5.6. Z.ai led with CyberGym. None of them are lying. All of them are choosing.
What I actually trust in September 2026
Aggregates over singles. Artificial Analysis’s index has the virtue of being run by a third party, on a fixed harness, across nine evaluations, on 190 models. Its absolute numbers mean little; its ordering is the most defensible public signal available.
Cost-normalised results. Anthropic’s own Opus 5 claims are the good example of how to report this: “within 0.5% of Fable 5’s peak score at half the cost” on CursorBench 3.2, and beating Fable 5 on OSWorld 2.0 “at one-third the cost.” That is a useful claim because it names the trade.
Token efficiency. Artificial Analysis noted that Gemma 4 31B used just 39 million output tokens to complete the entire Intelligence Index. For an agentic workload, a model that reaches the same answer in a third of the tokens is a third of the bill and a third of the latency. This is under-reported and it is one of the few metrics that translates directly into money.
Your own evals. Twenty tasks from your actual backlog, run against three models, scored by you. It takes an afternoon. It will tell you more than every leaderboard in this article combined, and it is the only evaluation that is definitionally uncontaminated.
And one thing I do not trust: any single tokens-per-second or benchmark number quoted without a harness, a quantisation and a date. Treat those as order-of-magnitude claims.
Part Five: The Number That Actually Moved — Agentic Time Horizon
If you only take one metric from this article, take this one.
METR measures the length of task — in human hours — that a model can complete autonomously with 50% success. It is a better proxy for “can this thing do my job” than any knowledge benchmark, because it captures the thing that actually fails in practice: not knowing the answer, but staying coherent for long enough to get there.
The trajectory, from METR’s own work and the AI Digest’s tracking of it:
| Period | 50%-success time horizon |
|---|---|
| 2022 (ChatGPT launch era) | ~30 seconds |
| Mid-2024 (GPT-4o class) | ~4 minutes |
| Late 2025 (Claude Opus 4.5) | 320 minutes (~5.3 hours), CI [170, 729] |
| 2026 frontier | 14+ hours |
The doubling time is the contested part. METR’s Time Horizon 1.1 release on 29 January 2026 — 228 tasks, 31 of them estimated at eight or more hours of human effort, migrated onto the UK AI Security Institute’s Inspect framework — reports 196 days (7 months) on a hybrid of old and new data, but 131 days when restricted to post-2023 models, which is about 20% faster than the 165 days the previous methodology gave.
Extrapolating an exponential is how people embarrass themselves, so let me flag the caveats loudly before I use the number.
The caveats are large. METR’s own confidence intervals are enormous — Opus 4.5’s 320-minute horizon has a range of 170 to 729 minutes. Only 5 of their 31 long tasks have measured human baselines; the rest are estimates. Task composition changes the measured trend. And “50% success on a well-specified task with a clean environment” is a long way from “50% success on your codebase with your flaky test suite and your undocumented deployment process.”
With all of that said, the direction is not in doubt and the magnitude is roughly right. On the extrapolation, frontier agents reach about one working day of autonomous task length in 2027 and about one working week in 2028.
Why this is the number that matters
Because it is the only metric that predicts what changes about your job.
A model with a 4-minute horizon is autocomplete. You stay in the loop continuously, and the quality of the model determines how good the suggestions are.
A model with a 5-hour horizon is a colleague you brief in the morning and review at lunch. You are not in the loop; you are at the end of it. The quality of the model determines how often the review is painful.
A model with a 40-hour horizon is something we do not have a management structure for. You would be reviewing a week of work you did not watch happen, in a codebase that changed while you were not looking, against a specification you wrote before you knew what the problems were.
Every argument in the rest of this article — about vibe coding, about verification, about where the bottleneck sits — is really an argument about what happens as that number climbs. The models got long-horizon. Our review processes did not.
Show Image The 50%-success time horizon, with METR’s confidence intervals shown honestly. The intervals are wide enough to matter — and the trend survives them anyway.
ACT II — THE FUTURE OF VIBE CODING
Part Six: What “Vibe Coding” Means Now
Andrej Karpathy coined the term in February 2025 to describe something specific and slightly mischievous: giving in to the vibes, forgetting the code exists, accepting diffs without reading them, and letting the model drive. It was a description of a mode, offered half in jest, for throwaway weekend projects.
Eighteen months later the phrase has been through the full lifecycle. It became a movement, then a job title, then a pejorative, then a Collins word of the year, then a category of venture funding, and now — in the only development that actually matters — a set of engineering practices that no longer resemble what Karpathy described.
I want to be precise about the split, because the two things get conflated constantly and they have opposite risk profiles.
Vibe coding, original sense. Prompt, accept, run, prompt again. Do not read the diff. The artefact is disposable. This is genuinely great, and I do it several times a week — for a script that reshapes a CSV, a one-off scraper, a visualisation I will look at once, a prototype whose only job is to make a conversation concrete. The defining feature is not the speed. It is that nobody will ever depend on it.
Agentic engineering, current sense. Write a specification. Give the agent a test suite, a linter, a type checker and a sandbox. Let it run for an hour. Review the diff properly. Merge behind the same gates as a human PR. The defining feature is that the model’s output is treated as a junior colleague’s output — useful, fast, and not trusted.
These share a technology and share almost nothing else. If you want the practitioner-level version of the second one, that is what my earlier piece on vibe coding in 2026 is about; this section is about where the practice is heading. The catastrophic outcomes in Part Eight all come from applying the first mode’s practices to the second mode’s stakes.
The adoption numbers are enormous and mostly not in dispute
- 90% of developers use at least one AI tool at work, per JetBrains’ AI Pulse survey (January 2026). Stack Overflow’s 2025 survey had 84% using or planning to use.
- Nearly 80% of new GitHub developers adopt Copilot within their first week (GitHub Octoverse 2025).
- 59% of developers use three or more AI coding tools in parallel. That statistic tells you more than the adoption number does — the market has not consolidated, it has fragmented into task-specific tools.
The market numbers
| Tool | Metric | Source / date |
|---|---|---|
| GitHub Copilot | ~42% share by user base; 20M+ users | Gartner / vendor, 2026 |
| Cursor | $2B+ ARR; 1M+ daily active users; ~20–25% of market by revenue | Reported, February 2026 |
| Claude Code | ~$2.5B run-rate; 46% “most loved” vs Cursor’s 19% | Reported, April 2026 |
| Enterprise coding-agent market | ~$9.8–11B annualised | Gartner, April 2026 |
Two things stand out. Copilot has the users and the lowest satisfaction of the top three. Claude Code has the smallest install base and by far the highest affection. That is the signature of a market in the middle of a transition — the incumbent won distribution, the challenger won the workflow, and the users have not finished moving yet.
And the number that undercuts all of it
Developer trust in AI output fell from about 40% in 2024 to 29% in 2026. A separate SonarSource survey found 96% of developers do not fully trust AI-generated code to be functionally correct.
Sit with that. Usage went up. Trust went down. Those are not contradictory; they are what mature tool adoption looks like. Nobody “trusts” a chainsaw. They use it constantly, with both hands, wearing protection, and they never turn their back on it.
The people who understand these tools best are the ones most careful with them. That is a healthy signal, and it is completely absent from most vendor marketing.
Part Seven: Does It Actually Make You Faster? The Evidence, Honestly
This is where I lose some readers, and I would rather lose them here than mislead them.
The randomised trial says no
METR ran the study everyone cites, in July 2025: experienced open-source developers, working on their own mature repositories (1M+ lines), randomised at the task level to use or not use early-2025 AI tools.
Result: 19% slower with AI.
And the finding that matters more than the headline: the developers believed they had been about 20% faster. They overestimated AI’s effect on their time by roughly 40 percentage points.
The 2026 follow-up did not resolve it — it revealed something worse
METR tried to re-run the experiment and, in February 2026, published an unusually candid post about why the design broke.
- Recruitment failed. Developers increasingly refuse to work without AI access, even at $50 an hour.
- Task selection was contaminated. 30–50% of developers admitted avoiding the tasks they expected AI to handle well. One participant put it perfectly: “I avoid issues like AI can finish things in just 2 hours, but I have to spend 20 hours.”
- Time tracking stopped working once developers were running multiple agents at once.
Their raw 2026 interim numbers ranged from −18% to −4% — still slower, less dramatically so — but METR themselves describe this as “only very weak evidence” given the selection bias.
Read the second bullet again, because it is the most interesting finding in the whole literature and it is buried in a methodology note. If developers systematically route the AI-friendly tasks away from a study measuring AI’s benefit, the study will understate the benefit. It also implies that in real life, developers are already doing the routing — which is precisely the behaviour that makes the tools valuable and makes the RCT hard to run.
The surveys say yes, by a lot, and they are inflated
METR’s May 2026 survey (n=349: 87 software engineers, 71 researchers, 129 academics and PhD students, 48 founders and managers; average 12 years’ programming experience, 19 months of AI tool use) did something clever. It separated value from speed.
The median respondent reported a 3× speed multiplier but only a 1.4–2× change in the value of their work. Retrospectively they estimated 1.3× value in March 2025; they forecast 2.5× by March 2027.
METR are explicit that value is the number survey designers actually care about, and that respondents naturally think in speed — which is exactly the bias the 2025 RCT quantified at 40 percentage points.
So: self-report says 3×. Self-report corrected for what people mean says 1.4–2×. The one controlled measurement we have says 0.81×.
The organisational data says “it depends on you, not the tool”
DORA’s 2026 report (Google Cloud, May 2026) is the most useful thing published on this all year, and it is useful because it stops arguing about the tool.
- Modelled first-year ROI of 39% for a 500-person engineering organisation, with roughly an 8-month payback — $11.6M of value against $8.4M of investment.
- A J-curve: productivity dips before it rises, driven by workflow adaptation, code review verification cost, and downstream changes to testing and approval gates.
- Change failure rate rising from 5% to 6% post-adoption, costing $344,000 in downtime in their sample calculation.
- And the thesis: “The greatest returns on AI investment come not from the tools themselves but from a strategic focus on the underlying organizational system.”
That last line is the 2026 consensus if there is one. AI amplifies whatever your engineering practice already is. Good tests, real code review, fast rollback, clear ownership — AI makes those organisations meaningfully faster. Weak tests, rubber-stamp review, manual deploys, unclear ownership — AI makes those organisations fail faster and at higher volume.
How to reconcile all of this
I do not think these findings actually conflict. Here is the reconciliation I have arrived at, and I hold it loosely:
AI coding tools are a large win on unfamiliar work and a small loss on deeply familiar work. On a codebase you know intimately, in a language you are fluent in, on a problem you have solved before, you are faster than the review-and-correct loop. On an unfamiliar framework, a language you use twice a year, boilerplate, test scaffolding, migrations, or a codebase you have not opened in eight months, the model is transformative.
METR studied experienced developers on their own mature repositories — the single scenario where you would predict the smallest gain. That does not make the finding wrong. It makes it specific, and it should make you suspicious of anyone who cites it as a general verdict, in either direction.
The second reconciliation: the gains are real but they land somewhere other than where people look for them. Not “this feature took four hours instead of six.” More like: the migration that was never going to get prioritised got done; the test coverage that was always going to stay at 40% went to 75%; the documentation exists. Those do not show up in a task-completion time study, and they are most of the actual value.
Part Eight: The Security Bill Has Arrived, And It Is Itemised
If Part Seven was uncomfortable, this part is the one people skip. Please do not skip it.
The measurements
Veracode tested over 100 large language models across 80 coding tasks in Java, Python, C# and JavaScript:
- 45% of AI-generated code samples introduce an OWASP Top 10 vulnerability. The pass rate was essentially unchanged through early 2026 — models got much smarter and no more secure.
- Java was worst at a 72% failure rate.
- 86% failed cross-site scripting defence. 88% were vulnerable to log injection.
Apiiro ran its Deep Code Analysis engine across Fortune 50 enterprise repositories from December 2024 to June 2025:
- AI-assisted developers produce commits at three to four times the rate of their peers.
- They introduce security findings at ten times the rate.
- Privilege escalation paths rose 322%. Architectural design flaws rose 153%.
USENIX Security 2025 analysed 576,000 AI-generated code samples:
- Roughly 20% reference packages that do not exist.
- 43% of hallucinated package names reproduce consistently across similar prompts — which is the part that turns a bug into an attack surface. A predictable hallucination is a slot an attacker can pre-register on npm or PyPI.
Georgia Tech’s Vibe Security Radar traced CVEs back through git history to attribute them to AI tooling:
| Month (2026) | CVEs attributed |
|---|---|
| January | 6 |
| February | 15 |
| March | 35 |
74 confirmed in total, with researchers estimating the true figure across open source is five to ten times higher.
RedAccess indexed roughly 380,000 publicly accessible vibe-coded applications across Lovable, Base44, Replit and Netlify. About 5,000 were leaking sensitive corporate or personal data.
The incidents, dated
| Date | Incident |
|---|---|
| Jul 2025 | Replit AI deletes a production database; 1,200 executive records affected, 4,000 fabricated accounts generated |
| Feb 2026 | Lovable-built EdTech app exposes 18,697 user records including 4,538 student accounts |
| Apr 2026 | Lovable BOLA vulnerability — every pre-November 2025 project reachable in five API calls |
| 26 Apr 2026 | Cursor “9-second wipe”: a company database deleted along with its backups |
| 30 Apr 2026 | Gemini CLI CVSS-10 remote code execution in CI workflows, patched |
The Amazon case, which I am going to handle carefully
You will have seen the Amazon story. Four Sev-1 incidents between December 2025 and March 2026 — a 13-hour AWS Cost Explorer outage in China, incorrect delivery times in shopping carts on 2 March costing a reported ~120,000 orders, and a 6-hour incident on 5 March where North American orders reportedly dropped 99%.
Reporting from the Financial Times, Fortune and The Register connected internal Amazon documents to what were described as “Gen-AI assisted changes,” with warnings that rapid code generation was exposing vulnerabilities. Amazon has disputed direct causation, stating in some responses that none of the incidents involved AI-written code.
I am flagging three things about this story rather than repeating it as fact:
- The most-cited impact figure (6.3 million lost orders) traces to a Medium post, not to Amazon or to a named outlet’s own reporting. Treat it as an estimate, not a number.
- The 5 March deployment reportedly went out without formal documentation or approval. That is a process failure. AI velocity made it possible to happen faster; it did not make the approval gate optional.
- That distinction is the whole point, and it is why I include the story at all. The failure mode is almost never “the AI wrote bad code.” It is “the AI wrote code faster than the process that was supposed to check it.”
What this actually implies for practice
Not “stop using the tools.” Nobody serious is arguing that, and the CSA’s own research note — which collates most of the above — leads with governance rather than abstinence.
The practical shape, in the order I would implement it:
- Treat any AI tool with credential access as a privileged system. It is one. Scope its tokens accordingly.
- SAST, dependency scanning and secret detection on every commit. Not nightly. Every commit. The commit rate went up three to four times; a nightly scan is now a three-day feedback loop in effective terms.
- Software composition analysis, specifically for hallucinated packages. This is the cheapest, highest-value control on the list because the 20% figure is so large and the failure is so silent.
- A written policy that AI assistance is not used unreviewed on authentication, cryptography, authorisation or payment code. These are the four areas where a subtle error is both likely and catastrophic, and where a reviewer’s eye is worth the most.
- Extend your SBOM to record AI tool provenance. When the next Georgia-Tech-style attribution study lands on your codebase, you want to be able to answer the question yourself.
- Treat third-party vibe-coding platforms as SaaS vendors. The RedAccess numbers are what happens when nobody does.
None of that is exotic. Most of it is what a competent shop was doing in 2019. The change is that the volume went up an order of magnitude, so the controls have to be automatic rather than cultural.
Show Image The security data, from Veracode, Apiiro, USENIX and Georgia Tech. Note the shape: commit velocity went up 3–4×, security findings went up 10×. The gap between those two multipliers is the entire problem.
Part Nine: What Replaced Vibe Coding
The practice that actually works in 2026 has a name — several competing names, which is how you know it is real — and it is roughly the opposite of vibes.
Call it spec-driven development, agentic engineering, or context engineering. The common structure:
1. The specification is the source code now
Not a Jira ticket. An actual document: what the thing does, what it must not do, what the interfaces are, what “done” looks like, what the failure modes are. Committed to the repository, versioned, reviewed.
This feels like a regression to 1998 and it is not. The difference is that in 1998 the spec was an input to a human who would reinterpret it, and it rotted the moment coding started. In 2026 the spec is an input to a process that regenerates the implementation cheaply — which means when requirements change you change the spec and re-derive, rather than patching code that no longer matches any written description of itself.
The economics flipped. When implementation was expensive and specification was cheap, you specified loosely and invested in the code. When implementation is cheap and specification is the bottleneck, you invest in the spec.
2. The test suite is the trust boundary
An agent that can run your tests and iterate against them is a fundamentally different tool from one that cannot. This is the single highest-leverage change most teams can make, and it is not about AI at all — it is that AI finally made the ROI on test coverage obvious to people who were never going to be persuaded by craft arguments.
Practical version: the agent gets a sandbox, a test command, a lint command and a type check, and it is not done until all four are green. Everything else is negotiable.
3. Review moved from lines to diffs to behaviour
Line-by-line review does not scale to a thousand-line agent-generated diff, and pretending otherwise is how rubber-stamping starts.
What works instead, in rough order of value:
- Review the spec, not the implementation. If the spec is right and the tests pass, the implementation is a detail.
- Review the diff’s shape — what files were touched, what dependencies were added, what got deleted. Most agent disasters are visible at this level: an unexpected file, a new dependency, a deletion nobody asked for.
- Review the security-sensitive lines specifically. Auth, crypto, authz, payments, anything touching PII. This is a small fraction of any diff.
- Do not review the boilerplate. You were never really reviewing it anyway.
4. The harness matters as much as the model
This is under-appreciated. The same model, in two different agent loops, produces work of wildly different quality — which is exactly why benchmark comparisons between labs are so slippery.
The pieces that matter: how context is managed as the run gets long, whether the agent can see its own test failures, whether it can browse the repository or only what you paste, whether it has a plan step separated from an execution step, and whether it knows when to stop and ask.
Meta’s Muse Spark 1.3 release notes are interesting on exactly this point — they describe the model as now asking clarifying questions on ambiguous prompts, invoking help when stuck, and confirming before consequential actions. Those are not intelligence improvements. They are manners, and manners are what make a long-horizon agent survivable.
5. Model routing is now standard practice
Nobody serious runs one model for everything — I walked through a two-model version of this in GLM-5.3-Flash vs MiniMax M3 in OpenCode. The pattern that has settled:
json
{
"$schema": "https://opencode.ai/config.json",
"model": "anthropic/claude-opus-5",
"small_model": "ollama/qwen3.8:27b",
"agent": {
"plan": {
"description": "Architecture, root cause, anything where being wrong is expensive",
"model": "anthropic/claude-fable-5-1"
},
"build": {
"description": "The bulk of implementation — long-horizon, cost-sensitive",
"model": "anthropic/claude-opus-5"
},
"bulk": {
"description": "High-volume mechanical work where price dominates",
"model": "zai/glm-5-3-flash"
},
"private": {
"description": "Anything touching client data — runs on hardware we own",
"model": "mlx/gpt-oss-120b"
},
"grunt": {
"description": "Renames, docstrings, lint fixes, test scaffolding",
"model": "ollama/gemma-4-26b-a4b"
}
}
}
The reasoning, line by line:
- Planning goes to the most capable model you can afford. A bad plan costs more downstream tokens than the plan ever cost.
- Building goes to the best long-horizon model at a sane price. Opus 5 at $5/$25 is explicitly positioned for this and is half Fable’s price.
- Bulk mechanical work goes to GLM-5.3-Flash at $0.15/$0.47, because at that price the question is not whether it is as good, it is whether it is good enough for the specific task. Frequently it is.
- Anything private goes local, on a model that fits your machine.
- Grunt work goes to a small local model, because renaming a symbol across forty files does not need frontier reasoning and does not need to leave your laptop.
Write the model IDs so swapping one is a one-line change. Given that four labs shipped a top-ten model between June and September, you will be making that change again within the quarter.
Part Ten: The Future of Vibe Coding — Six Predictions, With Confidence Attached
Predictions without confidence levels are entertainment. Here are mine, with how much I would bet.
1. The word dies; the practice splits permanently. High confidence.
“Vibe coding” is already becoming a term of abuse in professional settings and a term of pride nowhere except marketing. What replaces it is a clean two-tier split: disposable AI-generated artefacts where nobody reads the code (fine, correct, do more of it), and reviewed agentic engineering where the model’s output goes through the same gates as a human’s (also fine, and where all the enterprise money is).
The dangerous middle — production code, unread — is what the 2026 incident list is made of, and it will get regulated, insured or audited out of existence in commercial contexts. Not because anyone bans it, but because a security questionnaire eventually asks.
2. Verification becomes the bottleneck, and the tooling money follows it. High confidence.
We are already there and the market has not repriced yet. If agents can produce a week of work, the constraint is a human’s capacity to establish that a week of work is correct.
Expect the interesting products of 2027 to be about checking rather than generating: property-based test generation, semantic diff tools that describe behaviour change rather than line change, automated spec-conformance checking, runtime verification of agent actions, and — this is the one I would build — tools that tell you which 5% of a large diff a human actually needs to read.
3. Formal methods have their moment, twenty years late. Medium confidence.
The historical objection to formal specification was that writing the spec cost more than writing the code. That objection is now dead, because writing code is nearly free and writing specs is the bottleneck anyway.
Type systems, contracts, invariants, property tests and lightweight formal verification suddenly have a use case that pays for itself: they are machine-checkable statements of intent that an agent can iterate against without a human in the loop. I would expect Rust, TypeScript’s stricter modes, and contract libraries in Python to be disproportionate beneficiaries.
The reason my confidence is only medium: the field has predicted a formal methods renaissance roughly every seven years since 1985.
4. Local models eat the grunt tier entirely. Medium-high confidence.
Look at what a 27B model at Apache 2.0 does now, add another year, and put it on a machine with 64 GB of unified memory. Renames, docstrings, test scaffolding, commit messages, lint fixes, first-draft translations, log analysis — none of that needs a frontier model, all of it is high volume, and all of it is exactly the work you would rather not send to a third party from a client’s repository.
The frontier stays in the cloud for planning and hard reasoning. The volume moves local. Part Eleven onward is about why that is now physically possible.
5. The junior developer role changes shape rather than disappearing. Medium confidence, and I hold this one loosest.
The pessimistic case is straightforward: if agents do the work juniors used to do, nobody hires juniors, and in ten years there are no seniors.
I think the more likely outcome is that the entry-level job changes from “write the code” to “verify the code and own the spec” — which is a harder job to start in, not an easier one, and which is going to require the industry to rebuild how it trains people. The teams that figure this out will have an enormous hiring advantage over the teams that simply stop hiring.
The honest caveat: this is the prediction where my incentives and my analysis point the same direction, which is exactly when to be suspicious of yourself.
6. Somebody has a genuinely catastrophic, publicly attributed AI-coding failure. High confidence on the event, low on the timing.
Not a leaked database. Something with a regulator, a financial impact in nine figures, and a root cause analysis that says an agent did something no human reviewed.
The CVE attribution curve — 6, then 15, then 35 in successive months of 2026, with researchers estimating the true rate five to ten times higher — is not a curve that flattens on its own. The question is whether the industry tightens its own gates first or waits for the incident that forces it.
What I would do about all six, in one sentence: invest in tests, specs and review tooling rather than in prompt technique, because prompt technique depreciates every time a model ships and verification infrastructure compounds.
Show Image The two practices share a technology and nothing else. The left-hand loop is correct for disposable work and catastrophic for production; the right-hand loop is what the incident list in Part Eight was missing.
ACT III — OFFLINE AI ON CONSUMER DEVICES
Part Eleven: Offline AI Stopped Being A Hobby This Year
For about two years, “run it locally” was a principled compromise. You could do it. The honest pitch was: here is a 7-billion-parameter model that is worse than the free tier of everything, on hardware you already own, and the reason to do it is principle.
That changed in 2026, and it changed for three separate reasons that arrived at the same time.
Small models got genuinely good. Gemma 4’s E4B is roughly 8 billion parameters, supports 128K of context, handles text, images, video and native audio, and is Apache 2.0. The 31B dense variant scores 39 on the Artificial Analysis index — inside the range that was frontier eighteen months ago — and did the entire index run on 39 million output tokens, which is a fraction of what comparable models burn. Qwen3.8-27B is a dense multimodal model under Apache 2.0 with 262K native context that Alibaba claims beats the much larger Qwen3.7-Plus on coding and office tasks.
The operating systems shipped it. This is the change that actually matters and it happened quietly. Microsoft put Aion 1.0 into Windows 11. Apple ships a 3-billion-parameter dense model plus a 20-billion-parameter sparse model on iPhones. Google ships Gemini Nano on Pixel and Samsung devices. On-device AI stopped being something you install and became something that is already there, which is the difference between a technology enthusiasts use and a technology everyone uses.
And the hardware finally has the memory. Not the compute — the memory. I will come back to why that distinction is the whole story.
What “offline AI” actually buys you
Let me be concrete, because this gets sold badly.
Latency. A local 4B model responds in tens of milliseconds. A cloud round trip is hundreds at best. For anything interactive and continuous — dictation, live translation, autocomplete, summarising as you type — local wins on feel, not on quality, and feel is what people notice.
Privacy that is architectural rather than contractual. With an API, every prompt leaves your infrastructure and you need a processor agreement, a transfer impact assessment, a data flow map and an answer for your data protection officer. With a model running on the device, the cross-border transfer question does not arise, because there is no transfer. For German and EU work this is frequently the entire argument, and it closes deals that price never would.
Availability. On a plane. On a site with no coverage. In a facility where network access is the thing that is forbidden. During your provider’s incident.
Cost predictability. Not cost saving — I will do that maths in Part Fifteen and it does not say what people want it to say. But a fixed hardware cost with zero marginal cost per token is a very different budget line from a usage-based bill that scales with an agent’s enthusiasm.
Version pinning. A closed API model gets updated behind a stable model ID and behaviour drifts. A local checkpoint is byte-identical in a year. If you have ever had a working prompt quietly stop working, you know what that is worth.
And what it does not buy you
Frontier capability. The gap between a 4B model on your phone and Claude Fable 5.1 is not small and will not close. Anyone telling you otherwise is selling something.
Speed on large models. A local 400B model generates tokens slower than the same model on a vendor’s API, because the vendor runs it on hardware with four to six times your memory bandwidth. You trade throughput for control.
Part Twelve: The Phone In Your Pocket
The specification that decides everything is RAM
Every major platform has now drawn a memory line, and the line is 12 GB.
| Platform | On-device model | Memory gate |
|---|---|---|
| Apple | AFM 3 Core (3B dense) + AFM 3 Core Advanced (20B sparse, 1–4B active) | 12 GB for two iOS 27 features |
| Gemini Nano v3 (Pixel 10 only); Nano v2 on Pixel 9 and older; Nano v4 forthcoming | 12 GB minimum for Gemini Intelligence | |
| Qualcomm / Samsung | Snapdragon 8 Elite Gen 5 for Galaxy — 3rd-gen Oryon CPU, Hexagon NPU, +39% NPU capacity | Device-dependent |
Apple’s arrangement is the technically interesting one. AFM 3 Core Advanced is a 20-billion-parameter sparse model that lives in flash storage. Shared experts stay resident in memory; routed experts are swapped into DRAM only when a request needs them. That is how you fit a 20B model onto a phone at all — you trade NAND read bandwidth for DRAM capacity.
But swapping experts into DRAM still requires somewhere to swap them into, on top of the OS, the running app and the camera pipeline. The gap between 9 GB and 12 GB is not three gigabytes of market segmentation. It is the working set for expert swapping under real memory pressure. Ming-Chi Kuo reports that two iOS 27 features require 12 GB: Siri voice expressiveness and pace customisation, and a substantial speech-to-text accuracy improvement. The base iPhone 18, delayed to spring 2027 with a reported 9 GB, does not clear the bar.
Google has the same physics and a worse public-relations problem. Gemini Intelligence needs Nano v3, Nano v3 is currently Pixel 10 only, and Pixel 9 devices — sold under a seven-year update promise — are on Nano v2. The promise was about updates. It was never about features, and the distinction that made sense to a product manager in 2023 makes no sense at all to a customer in 2026.
Expect this to be the consumer AI story of the next two years: a long-support promise that delivers security patches to a device that cannot run the features the patches are shipped alongside.
What a phone can actually do, measured
Independent benchmarking on an iPhone 17 Pro (A19 Pro) across four runtimes:
| Model | Runtime | Decode tok/s | Peak memory |
|---|---|---|---|
| Gemma 4 E2B (4-bit) | LiteRT-LM | 55.4 | 641 MB |
| Gemma 4 E2B (4-bit) | MLX | 47.5 | 2,900 MB |
| Gemma 4 E2B (4-bit) | llama.cpp | 37.8 | 3,156 MB |
| Gemma 4 E2B (4-bit) | CoreML / ANE | 33.4 | 1,187 MB |
| Qwen 3.5 2B (4-bit) | MLX | 61.2 | 1,279 MB |
| Qwen 3.5 2B (4-bit) | llama.cpp | 39.1 | 1,479 MB |
| Qwen 3.5 2B (4-bit) | CoreML / ANE | 27.9 | 241 MB |
Two findings jump out.
The Neural Engine is the slowest path and by far the most memory-efficient. 27.9 tok/s at 241 MB against MLX’s 61.2 tok/s at 1,279 MB — five times the speed for five times the memory. That trade is exactly why Apple uses the ANE for always-on system features and why third-party apps skip it for interactive chat.
Sixty tokens per second on a 2B model in your pocket is faster than most people read. Which means the constraint on phone-local AI is not speed. It is capability, capability is a function of parameters, parameters are a function of memory — and memory is the thing every vendor just gated.
Part Thirteen: The PC, And Microsoft’s Quiet Bet
Microsoft did something at Build 2026 on 2 June that I think will look more significant in retrospect than it did on the day: it shipped models into Windows itself.
Aion 1.0 Instruct is a small language model for summarisation, rewriting and extraction. It needs a modern CPU and Windows 11 — no NPU, no discrete GPU. It is available now in Edge Canary from build 150.0.4070, exposed to web developers through the Prompt, Summarizer, Writer and Rewriter APIs, and its open weights landed on Hugging Face in July 2026.
Aion 1.0 Plan is the interesting one: a 14-billion-parameter reasoning and tool-calling model with a 32K context window, built for agentic workflows, file management and sub-agent orchestration. “Coming months” at time of writing.
The framing Microsoft used was “unmetered intelligence to every home and every desk.” Set the marketing aside and look at the architecture decision: a JavaScript API in the browser that calls a local model, with no key, no bill and no network. If that becomes normal, an enormous category of small AI features stops being a SaaS product and becomes a platform capability. That is a much bigger deal for the AI industry’s economics than any individual model release this year.
The hardware floor is real and it excludes a lot of machines
The NPU path requires a Copilot+ PC with a minimum 40-TOPS NPU:
| Silicon | NPU | Qualifies? |
|---|---|---|
| Qualcomm Snapdragon X Elite | 45 TOPS | ✅ |
| Intel Lunar Lake | 45–48 TOPS | ✅ |
| AMD Ryzen AI 300 / Max+ 395 | 50 TOPS | ✅ hardware; software “coming later” |
| Intel Meteor Lake | ~10–11 TOPS | ❌ |
And here is the number that should reset expectations: on a 45-TOPS NPU, an 8B model generates roughly 5 tokens per second. The 14B Plan model will be slower.
Five tokens per second is not a chat experience. It is fine for a background summarisation, a rewrite, a classification, an extraction — the things Aion 1.0 Instruct is actually for. It is not fine for anything you sit and watch.
Why so slow on 45 TOPS? Because token generation is bound by memory bandwidth, not by raw compute. The NPU can do the matrix multiplications; the problem is streaming the weights out of system memory for every single token. This is the same wall that governs every device in this article, and it is why “TOPS” is a nearly useless number for shopping.
The desktop boxes
If you want a machine specifically for local inference, the 128 GB class has settled into two options and one outlier:
| NVIDIA DGX Spark | AMD Ryzen AI Max+ 395 (Strix Halo) | Mac Studio M5 Ultra | |
|---|---|---|---|
| Memory | 128 GB unified | 128 GB LPDDR5X-8000 | up to 512 GB unified |
| Bandwidth | ~273 GB/s | ~256 GB/s | 1.2 TB/s |
| Price | ~$3,999 | $2,300–$2,500 | from $5,499 |
| Notes | CUDA ecosystem | Strong CPU (1.6 TFLOPS Linpack vs Spark’s 708 GFLOPS) | 512 GB tier late Oct, unpriced |
Those prices and the Spark and Strix Halo figures come from testing published in late December 2025, so treat them as indicative rather than current — memory pricing has moved since. The DGX Spark and Strix Halo are much closer on token generation than their compute figures suggest — because both are stuck around 256–273 GB/s of bandwidth and that is the binding constraint. The Spark wins decisively on image generation (~120 TFLOPS vs ~46 on FLUX.1 Dev BF16) and on CUDA compatibility. Strix Halo wins on price and CPU.
The Mac Studio M5 Ultra is in a different category entirely on both memory and bandwidth, and I wrote about it at length in the local AI hardware piece. The short version: 512 GB at 1.2 TB/s is the only mainstream machine that can hold a 400B-class model in memory, and the price for that privilege is somewhere around $16,000–$20,000 configured.
The models to actually run
| Model | Params | Licence | Runs comfortably in | Good for |
|---|---|---|---|---|
| Gemma 4 E2B | 5.1B / 2.3B active | Apache 2.0 | 4 GB | Phones, always-on, audio input |
| Gemma 4 E4B | ~8B | Apache 2.0 | 8 GB | Laptop grunt work, summarisation |
| Aion 1.0 Instruct | small SLM | Open weights (HF, Jul 2026) | CPU only | Windows-native text tasks |
| Qwen3.8-27B | 27B dense | Apache 2.0 | 24–32 GB | Best general local model under 32 GB |
| Gemma 4 26B-A4B | 26B / 4B active | Apache 2.0 | 24 GB | Fast MoE, low active cost |
| Gemma 4 31B | 31B dense | Apache 2.0 | 32 GB | Highest local quality under 32 GB |
| Ornith-1.5 35B | 35B / 3B active | MIT | 32 GB | Agentic coding, tiny active footprint |
| gpt-oss-120b | 120B / 5.1B active | Apache 2.0 | 128 GB | The 128 GB sweet spot |
| Ornith-1.5-397B | 397B MoE | MIT | 256 GB+ | Frontier-adjacent, needs a big box |
| MiniMax M3 | 428B / 23B active | Custom | 512 GB | Only fits on an M5 Ultra |
| DeepSeek V4 Pro | 1.6T / 49B active | MIT | Cluster | Not a consumer device model |
| Kimi K3 | 2.8T / 104B active | Custom | Cluster | Fits on no Mac Apple sells |
If you are sizing an Apple machine specifically, I went through the memory tiers in detail in the MacBook Pro M5 Pro 64GB local AI guide — the short version is that 64 GB is the point where the compromises stop being annoying.
If you want one recommendation: on a 32 GB machine, run Qwen3.8-27B or Gemma 4 31B. Both Apache 2.0, both multimodal, both genuinely useful, both fit. That is the 2026 default and it is a much better default than 2025 had.
Show Image What actually fits, by memory tier. The vertical rules are the memory sizes consumers can actually buy; the models that overflow the right-hand rule are the ceiling of local frontier AI in 2026.
Part Fourteen: The Memory Wall Is The Whole Story
Everything in Act III comes back to one equation, so here it is explicitly.
Will it fit?
weights (GB) ≈ total_params (B) × bytes_per_param
4-bit → × 0.5
8-bit → × 1.0
BF16 → × 2.0
usable ≈ (system memory × 0.75) − OS and app overhead − KV cache
How fast will it generate?
tokens/sec ≈ (memory_bandwidth × efficiency) ÷ active_bytes_per_token
Note that the second formula uses active parameters, not total. This is why mixture-of-experts models are so disproportionately good on consumer hardware. Gemma 4’s 26B-A4B has 26 billion parameters of capability and 4 billion parameters of per-token cost. Ornith-1.5’s 35B activates 3B. MiniMax M3 has 428B total and 23B active.
Efficiency in that formula is roughly 0.5 to 0.8 in practice depending on runtime, quantisation and how much of the memory bandwidth the rest of the system is using.
Now apply it. A 45-TOPS NPU laptop with LPDDR5X at, say, 120 GB/s, running an 8B model at 4-bit (4 GB of active weights): 120 × 0.6 ÷ 4 ≈ 18 tokens/second in theory. The measured figure is about 5. The difference is scheduling, memory contention, prefill costs and the fact that NPU toolchains are immature. That gap between the arithmetic and the measurement is where most disappointment in this space lives.
The DRAM shortage is the hidden variable behind every number in this article
Memory manufacturers spent 2025 and 2026 reallocating capacity toward AI datacentre customers, who sign larger contracts at better margins than consumer device makers. The consequences show up everywhere:
- Apple raised prices across Macs and iPads in June 2026.
- The 512 GB Mac Studio tier slipped to late October on supply grounds and remains unpriced.
- Apple’s M6 caps at 32 GB of unified memory, and the M6 Pro, Max and Ultra were cancelled outright.
- Ming-Chi Kuo reported Apple cutting 2026 hardware shipments over memory supply.
Every ceiling in this article traces to the same shortage. The 12 GB phone gate. The 32 GB M6 cap. The 128 GB laptop ceiling. The delayed 512 GB tier. Not one of them is a technical limit of the silicon; all of them are allocation decisions in a market where memory is scarce and expensive.
If you have been waiting for prices to normalise before buying a machine for local inference: nobody credible is forecasting when that happens.
Part Fifteen: Setting It Up
The practical section. Everything here I have run; where I am reporting rather than running, I say so.
The easiest path: LM Studio
A GUI, MLX-native on Apple Silicon, model browser built in, OpenAI-compatible server with one toggle. If you want to be running a model in ten minutes, start here. It is also the tool Apple benchmarks against in its own press releases, which tells you something about how mainstream this has become.
The most flexible path: Ollama
bash
# Install
curl -fsSL https://ollama.com/install.sh | sh
# Pull and run the 2026 default
ollama run qwen3.8:27b
# The context default is too small for real code work — raise it
OLLAMA_CONTEXT_LENGTH=65536 ollama serve
That second command is the one people miss. Ollama’s default context is set conservatively; on a code task it will silently truncate and you will blame the model.
The fastest path on Apple Silicon: MLX
bash
pip install --upgrade mlx-lm
# Generate directly
mlx_lm.generate \
--model mlx-community/Qwen3.8-27B-4bit \
--prompt "Explain mixture-of-experts routing in two paragraphs." \
--max-tokens 512
# Or serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Qwen3.8-27B-4bit --port 8080
Point any OpenAI-compatible client at http://localhost:8080/v1.
Raise the wired memory limit before you complain about crashes
macOS will not let the GPU wire all of your unified memory by default. On a large-memory machine this is the difference between a model loading and a model failing:
bash
# Check the current limit
sysctl iogpu.wired_limit_mb
# Example for a 128 GB machine, leaving ~14 GB for the OS
sudo sysctl iogpu.wired_limit_mb=116736
# Persist it
echo "iogpu.wired_limit_mb=116736" | sudo tee -a /etc/sysctl.conf
Do not set this to your full memory size. Leave the OS 12–16 GB on a large machine or you will get kernel panics rather than fast inference. And re-check it after every macOS update, because updates reset it.
On Windows, use the OS models first
Before installing anything, check what Windows already gives you. The Prompt, Summarizer, Writer and Rewriter APIs in Edge call Aion 1.0 Instruct locally with no key and no bill. For summarisation, rewriting and extraction inside a web app, that is a complete solution with zero infrastructure.
javascript
// Edge Canary 150.0.4070+ — runs entirely on device
const summarizer = await Summarizer.create({ type: "key-points" });
const summary = await summarizer.summarize(longText);
If you need more than that, LM Studio and Ollama both run fine on Windows, and on a machine without a qualifying NPU they will use the GPU or CPU, which for models in the 8B class is frequently faster than the NPU path anyway.
A sane hybrid setup
The architecture I would actually build, and roughly what I run:
- Tier 1 — on device. A small model for routing, classification, extraction, summarisation, dictation. Gemma 4 E4B or Aion 1.0 Instruct. Always available, zero marginal cost.
- Tier 2 — your own hardware. A 27B–120B open-weight model on a machine you own, for anything touching client data that must not leave the jurisdiction. Qwen3.8-27B on 32 GB,
gpt-oss-120bon 128 GB. - Tier 3 — frontier via API. Planning, hard reasoning, long agentic runs, anything where being wrong is expensive.
The debate is never “local or cloud.” It is where you draw the lines, and the lines moved a long way this year.
Part Sixteen: The Cost And Compliance Maths
The savings argument does not survive a calculator
Take a realistic small-business workload: a working month of moderate agentic coding, say 40 million input and 3 million output tokens.
| Option | Monthly cost |
|---|---|
| GLM-5.3-Flash API ($0.15 / $0.47) | ~$7.40 |
| Gemini 3.8 Flash API ($0.75 / $3.75) | ~$41 |
| Claude Opus 5 API ($5 / $25) | ~$275 |
| Claude Fable 5.1 API ($10 / $50, uncached) | ~$550 |
| A $2,400 Strix Halo box, amortised over 3 years | ~$67/month + electricity |
Note what that table actually says. Local hardware is cheaper than the top of the frontier and more expensive than the bottom of it. Against GLM-5.3-Flash, a local box never breaks even on tokens. Against Fable 5.1 running uncached, it pays for itself in five months — but you are not comparing like with like, because Fable 5.1 is a much better model than anything that fits in 128 GB.
So do not buy local hardware to save money. Buy it for one of these five reasons, all of which are real:
- Data residency that is an architecture decision rather than a legal project.
- No rate limits, no quotas, no capacity incidents. Your machine does not have a busy Tuesday.
- Version pinning forever. A local checkpoint is byte-identical in a year.
- Offline operation.
- You needed the workstation anyway. This is the most common case where buying is straightforwardly correct — if you were already going to spend €2,500 on a machine, the marginal cost of the AI capability is the memory upgrade, not the whole box.
The EU compliance picture, because it changed in 2026
For readers in Germany and the EU, the regulatory ground moved this year and the dates are now fixed rather than conditional.
- 2 August 2025 — GPAI model providers took on documentation, training-content-summary and copyright obligations.
- 2 August 2026 — enforcement powers and Article 50 transparency duties took effect. Fines are now applicable.
- 27 July 2026 — the Digital Omnibus, Regulation (EU) 2026/1744, entered into force, replacing the previous conditional mechanism for high-risk systems with fixed dates.
- 2 December 2026 — end of the grace period for machine-readable watermarking on existing systems.
- 2 August 2027 — the two-year transition window closes for GPAI models already on the market.
- 2 December 2027 — stand-alone high-risk systems under Annex III (hiring, credit scoring, education, critical infrastructure).
- 2 August 2028 — high-risk AI embedded in regulated products under Annex I (medical devices, machinery).
The Digital Omnibus also added prohibitions on AI-generated non-consensual intimate imagery and CSAM, and expanded the EU AI Office’s supervisory powers.
What this means practically for the readers of this article. If you are building on top of models rather than training them, most of the GPAI obligations sit with the provider, not with you. What lands on you is transparency — telling people when they are interacting with AI, and marking synthetic content — plus the high-risk classification question if your system touches hiring, credit, education or critical infrastructure. The December 2027 date is far enough away to plan for and close enough that “we’ll look at it later” is no longer a strategy.
And the point most relevant to Act III: an on-device model materially simplifies several of these conversations at once. No transfer, no processor, a much shorter data flow map. That is not a loophole — the transparency obligations still apply — but it removes an entire category of work.
I am not a lawyer and this is not legal advice. For anything with money attached, get a German or EU-qualified adviser to look at your specific deployment.
Part Seventeen: The Honest Case Against All Of This
If I stopped at Part Sixteen this would be an advert.
The measured productivity evidence still does not support the enthusiasm. One randomised controlled trial, 19% slower. A follow-up that could not be run cleanly. Self-reports inflated by 40 percentage points. Everyone in this industry, myself included, is operating substantially on vibes about vibes.
The security data is bad and not improving. Veracode’s 45% OWASP failure rate was essentially unchanged through early 2026 while models got dramatically smarter. That is not a transitional problem that scale fixes. Capability and security are not the same axis, and nobody is optimising hard for the second one.
Benchmarks are saturated at the top and unverifiable in the middle. SWE-bench Verified has three models within a point. The benchmarks replacing it are increasingly private, which solves contamination by making independent verification impossible.
Open weights are one licence revision from not being open. Z.ai moved GLM from MIT to a proprietary licence with a security-review gate between 5.2 and 5.3. Meta has promised open weights for Muse Spark with no date, no size and no licence text. If your architecture assumes downloadable weights next year, that assumption is a vendor’s business decision, not a law of nature.
The frontier is still not local, and increasingly not purchasable either. Mythos 5.1, the highest scorer on Terminal-Bench 4.0, is restricted to vetted cyberdefenders and life scientists, currently mostly in the US. Kimi K3 and DeepSeek V4 Pro do not fit on any consumer machine. What fits on your desk is an excellent second-tier model, which is a genuinely useful thing to be and is not the frontier.
On-device performance is worse than the marketing implies. Five tokens per second for an 8B model on a 45-TOPS NPU. A 12 GB memory gate that excludes most phones in circulation. A seven-year update promise that does not cover features. The gap between “AI PC” branding and what the machine can actually run is currently enormous.
And geopolitics is a procurement input whether or not you think it should be. Several of the best open-weight models — Kimi, GLM, DeepSeek, Qwen, MiniMax — come from Chinese labs. For public-sector, defence-adjacent and some regulated clients that is a hard blocker regardless of where the weights run or what the licence says. Meanwhile the most capable Anthropic coding model is gated to vetted US organisations. The open ecosystem is being pulled apart along national lines from both ends, and if you are in Europe you are on neither end of it.
Part Eighteen: The Decision Framework for the State of AI 2026
Show Image Seven situations, seven answers. Find the row that sounds like your year.
Seven situations, mapped onto the state of AI 2026 as it actually stands today.
1. “I want the best model and cost is not the issue.”
Claude Fable 5.1 at 65.7 on the aggregate index, $10/$50, 1M context — and turn caching on, because the 75% cache read cut is worth 25–45% of your bill. Keep GPT-5.6 Sol available for browsing-heavy research at 90.4% BrowseComp. Do not agonise; the gap between them is smaller than the gap between your best and worst prompt.
2. “I run a lot of agents and the bill is getting silly.”
Claude Opus 5 at $5/$25 for the long-horizon work, GLM-5.3-Flash at $0.15/$0.47 for bulk mechanical work, and a local model for grunt. Routing is the single biggest cost lever available to you and it is a config file, not a project. Measure your token split by agent role before you optimise anything — most people discover 80% of their spend is on tasks that did not need a frontier model.
3. “GDPR and data residency are hard requirements for client work.”
Qwen3.8-27B or Gemma 4 31B on hardware you own, both Apache 2.0, both multimodal. Step up to gpt-oss-120b on a 128 GB machine if the work needs it. The point is not speed — it is that no prompt leaves the building and your legal review becomes a licence read instead of a transfer impact assessment. Pick permissively licensed models specifically so that read is short.
4. “I want to try local AI without spending much.”
Whatever you already own, plus LM Studio, plus Gemma 4 E4B. Eight gigabytes of memory is enough to be genuinely useful. Find out whether local AI is for you before you spend money on it — most people discover they want tier 1 and tier 3 and never needed tier 2 at all.
5. “I’m buying a machine specifically for local inference.”
Buy on memory bandwidth and capacity, in that order, and ignore TOPS. In the 128 GB class, Strix Halo at ~$2,400 is the value pick and DGX Spark at ~$3,999 buys you CUDA and much better image generation. Above that, the Mac Studio M5 Ultra is the only thing that holds a 400B model, at a price that only makes sense if you needed the workstation anyway.
6. “My team is adopting AI coding tools and I own the risk.”
Do the boring things, in this order: SAST and secret detection on every commit; software composition analysis for hallucinated packages; a written policy excluding unreviewed AI code from auth, crypto, authz and payments; agent credentials scoped as privileged; and an SBOM that records tool provenance. Then expect a J-curve — DORA measured a real dip before the gain — and do not panic in month three.
7. “Which phone should I buy for AI?”
Any 12 GB model. That is the line every platform has drawn: iPhone 17 Pro and 18 Pro at 12 GB, Pixel 10 for Nano v3, Galaxy S26 Ultra on Snapdragon 8 Elite Gen 5. The base iPhone 18 at a reported 9 GB and Pixel 9 and older will run the phone fine and will not get the features. If you are on a 12 GB device already, this year’s upgrade is incremental for AI specifically.
Part Nineteen: What I Am Watching Next Quarter
Whether Meta actually ships Muse Spark’s weights, and under what licence. “Soon” from a CEO on X, with no parameter count and no licence text, is the weakest form of commitment in this article. If they ship under Apache 2.0 it materially changes the open ecosystem. If they ship under a bespoke licence with a revenue gate, that is now the industry default and open weights become a marketing category rather than a property.
Whether the Ornith-1.5 self-improvement result replicates. An independent evaluation on tasks outside the self-generated curriculum is the thing that would settle it. This is the highest-variance item on the list: it is either the most important result of 2026 or a cautionary tale about self-generated benchmarks, and I do not currently know which.
Whether Aion 1.0 Plan actually ships and what it runs at. A 14B reasoning model in Windows is a genuinely new category. If it runs at 3 tokens per second on qualifying hardware, it is a demo. If it runs at 20, it changes what a desktop application can assume.
The CVE attribution curve. 6, 15, 35 in successive months. If April through August continued that trajectory, the Georgia Tech data will be the most consequential AI security publication of the year when it lands.
Whether anyone builds the review tooling. The gap between what agents can produce and what humans can verify is the defining problem of 2027, and as of today almost all the venture money is still going into generation rather than verification. Whoever ships the tool that reliably tells you which 5% of a diff needs human eyes will be very hard to compete with.
Whether the memory shortage eases. It is the hidden variable behind the 12 GB phone gate, the 32 GB M6 cap, the delayed 512 GB Mac tier and the June Apple price rise. Nobody credible is forecasting relief.
And whether safety-gated frontier models stay an exception. Mythos 5.1 is gated for dual-use reasons that I find defensible, and Project Glasswing is actively widening access — roughly 150 organisations across more than fifteen countries by June 2026. But the precedent is set: the top scorer on a public coding benchmark is now something most organisations cannot buy at any price. If that becomes a pattern rather than a special case, “can my organisation get vetted” turns into a model selection criterion, and the argument for open weights outside the US stops being about cost entirely.
Frequently Asked Questions
What is the best LLM in September 2026?
On the aggregate, Claude Fable 5.1 leads Artificial Analysis’s Intelligence Index v4.1.1 at 65.7, ahead of Claude Opus 5 at 63.0, Claude Fable 5 at 62.1, Grok 4.6 at 60.9 and GPT-5.6 Sol at 58.9. But the aggregate compresses the differences that matter. GPT-5.6 Sol leads browsing at 90.4% on BrowseComp; Gemini 3.8 Flash delivers 58.7 at $0.75 per million input tokens; GLM-5.3-Flash delivers 57.5 at $0.15. For most real workloads the top eight models are interchangeable and the decision should be made on price, latency, cache economics and licence rather than on ranking.
What is the best open-weight model right now?
It depends on your hardware and your licence tolerance. GLM-5.3 (753B) scores highest at 59.5 but carries a custom licence with a $10 billion revenue security-review gate. DeepSeek V4 Pro (1.6T total, 49B active) is MIT licensed and scores 87.9 on Terminal-Bench 2.1, but needs a cluster. For a machine you actually own, Qwen3.8-27B or Gemma 4 31B — both Apache 2.0 — are the best models that fit in 32 GB.
Is vibe coding dead in 2026?
The word is dying; the practice split in two. Vibe coding in the original sense — prompt, accept, do not read the diff — remains correct for disposable work: scripts, prototypes, one-off analyses. What replaced it for anything that ships is spec-driven agentic engineering: a written specification, a sandbox, a test suite the agent must pass, and a real review of the resulting diff. The dangerous middle — production code, unreviewed — is where every documented incident of 2026 came from.
Does AI coding actually make developers faster?
The only randomised controlled trial says no. METR found experienced developers were 19% slower with early-2025 AI tools on mature codebases, while believing they had been 20% faster — a 40-percentage-point overestimate. Their 2026 follow-up could not be run cleanly because developers refused to work without AI. Surveys report a 3× speed multiplier but only a 1.4–2× change in the value of work. My reading: large gains on unfamiliar work, boilerplate, tests and migrations; small or negative gains on codebases you know intimately.
How risky is AI-generated code?
Measurably risky and improving slowly. Veracode found 45% of AI-generated samples introduce an OWASP Top 10 vulnerability, with Java worst at 72%. Apiiro measured AI-assisted developers producing commits 3–4× faster and security findings 10× faster, with privilege escalation paths up 322%. Roughly 20% of AI-generated code references packages that do not exist, and 43% of those hallucinated names recur consistently — which is a supply chain attack surface. The fix is automated gates on every commit, not abstinence.
Can I run a good LLM offline on my own computer?
Yes, and 2026 is the first year that is unambiguously true. On 8 GB of memory, Gemma 4 E4B is genuinely useful. On 32 GB, Qwen3.8-27B or Gemma 4 31B are strong multimodal models under Apache 2.0. On 128 GB, gpt-oss-120b is the sweet spot. What does not fit on any consumer machine is the actual frontier: Kimi K3 at 2.8 trillion parameters and DeepSeek V4 Pro at 1.6 trillion both need a cluster.
What is Microsoft Aion 1.0?
A family of on-device models announced at Build on 2 June 2026 and shipped into Windows 11. Aion 1.0 Instruct is a small model for summarisation, rewriting and extraction that runs on a modern CPU with no NPU required, exposed through Edge’s Prompt, Summarizer, Writer and Rewriter JavaScript APIs, with open weights on Hugging Face since July 2026. Aion 1.0 Plan is a 14-billion-parameter reasoning and tool-calling model with a 32K context window for agentic workflows, requiring a Copilot+ PC with a 40-TOPS minimum NPU, and was still “coming months” at the start of September 2026.
Why do phones need 12 GB of RAM for AI features?
Because on-device models are memory-bound, not compute-bound. Apple’s AFM 3 Core Advanced is a 20-billion-parameter sparse model that lives in flash storage and swaps routed experts into DRAM on demand — which still needs somewhere to swap them into, on top of the OS, the app and the camera pipeline. Ming-Chi Kuo reports two iOS 27 features requiring 12 GB. Google’s Gemini Intelligence has the same 12 GB minimum. The gap between 9 GB and 12 GB is the working set, not market segmentation.
Is an NPU worth having for local AI?
Less than the marketing implies. On a 45-TOPS NPU, an 8B model generates roughly 5 tokens per second, because token generation is limited by memory bandwidth rather than compute. NPUs are excellent for always-on, low-power, short-burst tasks — the ones an operating system runs constantly. For interactive chat on a laptop, the GPU or even the CPU path is frequently faster. Buy on memory bandwidth and capacity; treat TOPS as close to meaningless for shopping.
Are open-weight models actually open source?
Mostly not. DeepSeek V4 Pro (MIT), Qwen3.8 (Apache 2.0), Gemma 4 (Apache 2.0) and Ornith-1.5 (MIT) carry genuinely permissive licences. GLM-5.3 requires companies above $10 billion revenue to pass Z.ai’s security review; Kimi K3 uses a revenue-tiered licence; MiniMax M3 mandates prominent “Built with MiniMax M3” attribution for any commercial use and written authorisation above $20 million revenue. None of them release training data or training code. “Open weight” is the accurate term; “open source” is doing dishonest work.
Is it cheaper to run models locally than to use an API?
Almost never on token cost alone. A moderate month of agentic work — 40M input and 3M output tokens — costs about $7 on GLM-5.3-Flash and about $41 on Gemini 3.8 Flash. A $2,400 local box amortised over three years is roughly $67 a month before electricity. Local hardware beats the top of the frontier on cost and loses to the bottom of it. Buy local for data residency, rate-limit freedom, version pinning and offline operation — or because you needed the workstation anyway.
What does the EU AI Act require in 2026?
As of 2 August 2026, enforcement powers and Article 50 transparency duties apply to general-purpose AI, and fines are now available. The Digital Omnibus, Regulation (EU) 2026/1744, entered into force on 27 July 2026 and set fixed dates: 2 December 2026 ends the watermarking grace period; 2 August 2027 closes the transition for GPAI models already on the market; 2 December 2027 applies to stand-alone high-risk systems under Annex III; 2 August 2028 covers high-risk AI embedded in regulated products. If you build on models rather than train them, most GPAI obligations sit with the provider — transparency and high-risk classification are what land on you. This is not legal advice.
How long can an AI agent work unsupervised?
METR measures the task length a model completes autonomously with 50% success. That figure went from about 30 seconds in 2022 to over 14 hours for frontier systems in 2026, with a doubling time of roughly 196 days on METR’s hybrid fit, or 131 days when restricted to post-2023 models. The confidence intervals are very wide — Claude Opus 4.5’s 320-minute horizon ranges from 170 to 729 minutes — and only 5 of METR’s 31 long tasks use measured human baselines. Extrapolated, that is about one working day of autonomous task length in 2027.
Which model should I use for coding specifically?
Route rather than choose. Planning and architecture to the most capable model you can afford (Claude Fable 5.1); the bulk of long-horizon implementation to Claude Opus 5 at half the price; high-volume mechanical work to GLM-5.3-Flash at $0.15/$0.47; anything touching client data to a local open-weight model; and renames, docstrings and lint fixes to a small local model. Write the model IDs so that swapping one is a single-line change — four labs shipped a top-ten model between June and September 2026.
The Bottom Line
Twelve months ago, the interesting question in AI was which model was smartest. The state of AI 2026 is that this question has become close to uninteresting. The top ten models span 8.2 points on the best available aggregate, and the cheapest of them costs one sixty-seventh of the most expensive.
What replaced it is more useful and less exciting: can you verify what these things produce, fast enough to keep up with them?
That single question explains every finding in this article. It explains why SWE-bench saturated at 96% and stopped being informative. It explains why adoption of AI coding tools hit 90% while trust fell to 29%. It explains why Veracode’s security failure rate did not budge while capability soared, because capability and verifiability are different axes and only one of them is being optimised. It explains why DORA found that the returns come from the organisation rather than the tool. And it explains why the practice that actually works in 2026 — specification, sandbox, test gate, real review — looks so much less magical than the demos.
Four things I would take away.
The frontier commoditised faster than anyone expected, and the open-weight gap is now about six months. A 397B MIT-licensed model beats last generation’s closed frontier on SWE-bench. A 753B open-weight model outranks GPT-5.6 Sol on the aggregate index. Plan for a world where the model is not your differentiator, because it very nearly is not one already.
Read the licence before the benchmark. Capability is converging; terms are diverging. Z.ai moved away from MIT between releases. Meta has promised weights with no licence text. Anthropic’s best coding model is gated by nationality. For a commercial buyer, four points of benchmark is worth less than a licence your lawyer can read in ninety seconds.
Verification is the whole game now, and almost nobody is funding it. Agents can produce a day of work; we review it with tools built for reviewing an hour of work. If you are choosing where to spend engineering effort in the next twelve months, spend it on tests, specifications and review tooling rather than on prompt technique. Prompt technique depreciates with every model release. Verification infrastructure compounds.
And offline AI got real, on a memory budget that just got expensive. A 27-billion-parameter Apache-licensed multimodal model runs on a machine you can buy for €2,500. Windows ships a model in the operating system. Your phone runs a 20-billion-parameter sparse model out of flash. Every ceiling on that — the 12 GB phone gate, the 32 GB laptop cap, the 40-TOPS floor — is a memory allocation decision in a market where DRAM is scarce, not a limit of the silicon.
The setup I would actually build today: Gemma 4 E4B on the device for the always-on tier; Qwen3.8-27B on your own hardware for anything that must not leave the building; GLM-5.3-Flash at $0.15/$0.47 for the high-volume middle; and Claude Opus 5 or Fable 5.1 for planning and the genuinely hard problems, with caching switched on. Route between them in a config file, and expect to rewrite that file in three months.
Revisit this in December. Meta will either have shipped weights or not. Aion 1.0 Plan will either be real or vapour. The Ornith self-improvement result will have replicated or quietly disappeared. And the CVE attribution curve will have told us whether 2026 was the year the industry tightened its own gates or the year it waited for someone else to.
Are you running open-weight models in production in the EU, or shipping agent-written code through a review process you actually trust? I am particularly interested in two things: measured token-cost splits by agent role, and anything that works for reviewing thousand-line agent diffs. Corrections and counter-evidence welcome — this article gets updated when the picture changes.
Last updated: 3 September 2026.



