Zum Inhalt springen
DONNERSTAG, SEPTEMBER 3, 2026UNABHÄNGIGE TECH-NACHRICHTEN, TESTBERICHTE UND PRAKTISCHE ANLEITUNGEN
ÜBER
AKTUELLES /

Kimi K3 vs. Fable 5.1 vs. GLM-5.3: Die Ära der offenen Gewichtsklassen ist angebrochen

LESEN SIE DEN LEITFADEN
KI-AGENTENGM / 782

Kimi K3 vs. Fable 5.1 vs. GLM-5.3: Die Ära der offenen Gewichtsklassen ist angebrochen

Dark comparison card for Kimi K3, Claude Fable 5.1 and GLM-5.3 showing Terminal-Bench 2.1 scores of 88.3 against 88.0 for open versus closed models

Vor sechs Wochen hätte ich diesen Artikel anders geschrieben.

Im Juli stellte Moonshot AI ein Modell der Spitzenklasse mit 2,8 Billionen Parametern auf Hugging Face bereit und ermöglichte es jedem, es herunterzuladen. Am 28. August veröffentlichte Z.ai schließlich die GLM-5.3-Gewichte, die es zwei Wochen lang zurückgehalten hatte – nachdem festgestellt worden war, dass das eigene Modell Exploits auf eine Weise miteinander verketten konnte, um die niemand gebeten hatte. Und am 1. September brachte Anthropic „Claude Fable 5.1“ auf den Markt und senkte die Kosten für agentische Aufgaben um bis zu 451 TP3T, ohne den Listenpreis anzutasten.

Drei Modelle. Drei völlig unterschiedliche Antworten auf die Frage, was ein KI-Labor der Öffentlichkeit schuldig ist. Und zum ersten Mal seit Beginn dieser ganzen Angelegenheit sind die Optionen mit offener Gewichtung nicht die Kompromisslösung.

Das ist die eigentliche Geschichte von Ende 2026, und sie ist etwa achttausend Wörter wert, denn in den Details stecken das ganze Geld und alle Fehler.

Ich habe die Architekturunterlagen, die Modellkarten und die Lizenztexte herausgesucht – sorgfältig, Zeile für Zeile, denn zwei der drei entsprechen nicht den Bezeichnungen, die man im allgemeinen Sprachgebrauch dafür verwendet –, dazu die Hersteller-Benchmarks, die Bewertungen neutraler Aggregatoren, die tatsächlichen Preise pro Token und Ollamas brandneues Preisblatt. Anschließend habe ich die Kostenberechnungen für einen normalen Arbeitsmonat durchgeführt, anstatt von einem Szenario am Tag der Markteinführung auszugehen.

Das ist dabei herausgekommen.


TL;DR – Das 60-Sekunden-Urteil

GLM-5.3 bietet derzeit das beste Preis-Leistungs-Verhältnis im Bereich der KI – und das mit großem Abstand. Laut Artificial Analysis liegt es nur einen halben Punkt hinter Kimi K3 und etwa dreieinhalb Punkte hinter Claude Opus 5, mit $1,40 beim Einlesen und $4,40 beim Auslesen pro Million Tokens. Bei derselben agentenbasierten Aufgabe kostet es etwa ein Sechstel dessen, was Fable 5.1 kostet. Die Gewichte stehen zum Download bereit. Wenn Sie ein Modell suchen, auf das Sie Ihren Agenten ausrichten können, und Sie kein Unternehmen mit einem Umsatz von $10 Milliarden betreiben, dann ist dies genau das Richtige für Sie.

Kimi K3 ist das Leistungsstärkste, was man legal herunterladen kann. 2,8 Billionen Parameter, 104 Milliarden aktive Parameter, ein echter 1-Mio.-Kontext, native Bildverarbeitung, der beste bisher veröffentlichte Wert für agentisches Browsen (BrowseComp 91,2) und – die Zahl, die der Debatte um den “Rückstand offener Modelle” ein Ende setzen sollte – 88,3 auf Terminal-Bench 2.1, gegenüber 88,0 bei Claude Fable 5 und 88,8 bei GPT-5.6 Sol auf derselben Version. Außerdem kostet es $3/$15 – mehr als doppelt so viel wie GLM-5.3 – und für seine 1,5 TB an Gewichten ist ein Rack erforderlich, keine Workstation. Es ist das Modell, das man einsetzt, wenn die Aufgabe schwierig und langwierig ist, und das Modell, auf das man verweist, wenn jemand behauptet, offene Gewichte seien eine Generation im Rückstand.

Der Claude Fable 5.1 ist nach wie vor das Modell der Wahl, wenn Fehler teuer zu stehen kommen. Es ist führend bei langanhaltender, stundenlanger agentischer Arbeit und bei der chaotischen Fehlerbehebung, die in einer Rangliste nie sauber zum Ausdruck kommt. Anthropic hat am 1. September die Cache-Lesezugriffe um 75% reduziert, wodurch lange Agenten-Sitzungen deutlich kostengünstiger geworden sind als zuvor. Es ist zudem sechsmal so teuer wie GLM-5.3 pro Aufgabe (Closed) und unterliegt nun Anti-Destillations-Einschränkungen, die einige bestehende Integrationen unbrauchbar machen.

Der ehrliche Einzeiler: Die Open-Weight-Modelle haben die Leistungslücke auf etwa drei bis fünf Monate verringert und den Preisunterschied gänzlich beseitigt. Was Sie im September 2026 tatsächlich von Anthropic kaufen, ist Zuverlässigkeit bei den schwierigsten 10%-Aufgaben – und das Recht, sich über all das keine Gedanken machen zu müssen.

Und die Wendung, die niemand laut aussprechen will: Weder Kimi K3 noch GLM-5.3 werden noch unter einer Open-Source-Lizenz vertrieben. Beide sind diesen Sommer auf maßgeschneiderte, umsatzabhängige Bedingungen umgestiegen. Offene Gewichtsangaben, ja. Open Source, nein. Lesen Sie den Abschnitt über Lizenzen, bevor es Ihre Rechtsabteilung tut.

Comparison table of licence terms for Kimi K3, GLM-5.3, GLM-5.3-Flash and Claude Fable 5.1, showing revenue thresholds, attribution rules and which models are genuinely open source
Herunterladbare Gewichte, drei verschiedene Saitensätze. Nur GLM-5.3-Flash unterliegt weiterhin der MIT-Lizenz.

Übersichtstabelle

Kimi K3GLM-5.3Claude Fable 5.1
HerstellerMoonshot AIZ.ai (Zhipu AI)Anthropic
Veröffentlicht16. Juli 2026 (API) · 27. Juli (Gewichte)14. August 2026 (API) · 28. August (Gewichte)1. September 2026
Verfügbare GewichteJaJaNein
LizenzKimi K3-Lizenz (maßgeschneidert, umsatzabhängig)glm-5.3-Lizenz (maßgeschneidert, umsatzabhängig)Urheberrechtlich geschützt
Gesamtparameter2,8 T753BNicht bekannt
Aktiv pro Token104B (16 von 896 Experten)~40 Mrd. (geschätzt)Nicht bekannt
ArchitekturKDA + Attention-Residuen, Stable LatentMoE, 69 KDA- und 24 Gated-MLA-SchichtenMoE, gleiche Basis wie GLM-5.2, profitiert ausschließlich vom Post-TrainingNicht bekannt
Kontextfenster1,048,5761,000,0001,000,000
Maximale Leistung131.072 Standard (bis zu 1.048.576)128,000128,000
Multimodale EingabeText, Bild, Video (MoonViT-V2)Nur TextText, Bild
NachdenkenImmer eingeschaltetImmer eingeschaltet, logische_Überlegung niedrig/hoch/maxStändig aktiv, anpassungsfähig, auf einzelne Nachrichten zugeschnitten (Beta)
Einkaufspreis / 1 Mio.$3.00$1.40$10.00
Zwischengespeicherte Eingaben / 1 Mio.$0.30$0.26$0.25
Verkaufspreis / 1 Mio.$15.00$4.40$50.00
AA-Intelligenzindex59.759.5Noch keine Bewertung (Fable 5: 62,1)
Kosten pro AA-Bewertungsaufgabe$0.84$0.68— (Opus 5: $2.34)
Platzbedarf bei eigenem Hosting~1.5 TB (native MXFP4)~1.5 TB BF16 / ~750 GB FP8K.A.
On Ollamakimi-k3:cloudglm-5.3:cloudVia Claude Code / Claude Desktop as a client

All pricing is first-party list pricing. Artificial Analysis figures are from the September 2026 index snapshot; Fable 5.1 launched on 1 September and had not been independently scored at the time of writing.


Why This Three-Way Is The Comparison That Matters

For two years the open-versus-closed conversation had a comfortable shape: closed models were better, open models were cheaper, and you picked your spot on that trade-off. Everyone knew where they stood.

That shape broke this summer.

Nathan Lambert, who tracks this more carefully than almost anyone, puts the current capability gap between frontier open-weight and frontier closed models at three to five months — down from the six to nine months people were quoting a year ago. At the time he wrote, Kimi K3 sat at #2 on the Vals AI index and near the top of Artificial Analysis’s, beaten only by Claude Fable and GPT-5.6 Sol Max. The September index snapshot I use later in this article has it fourth, behind Grok 4.6. Either way: that is not “good for an open model.” That is a podium finish.

Meanwhile the usage data has gone somewhere genuinely surprising. Chinese-origin models captured roughly 61% of all tokens routed through OpenRouter by May 2026. The US share of that traffic collapsed from about 70% to about 30% over twelve months. Meta’s Llama — the model that started the open-weight wave — fell below 1% of routed volume. Google’s share went from roughly 37% to 13%. OpenRouter’s own 100-trillion-token study with a16z found open-weight models accounting for about a third of all token volume on the platform.

You can argue about what OpenRouter traffic represents. It over-indexes on developers, hobbyists and cost-sensitive workloads, and it under-counts enterprise contracts that never touch a router. Fine. But it is the largest public dataset we have on what people actually choose when they can choose freely, and the direction is unambiguous.

So the question in September 2026 is no longer “are open models good enough yet.” It is much sharper:

Given three models that all clear the bar, which one do you point your agent at — and what does being wrong actually cost you?

That is a question about benchmarks, prices, licences and infrastructure, in roughly that order of how much people think about them and exactly the reverse order of how much they matter.


What Kimi K3 Actually Is

Moonshot AI announced Kimi K3 on 16 July 2026 and released the full weights on 27 July. It is, by parameter count, the largest open-weight model anyone has ever published — 2.8 trillion total parameters, roughly 75% larger than DeepSeek V4 Pro.

That number is less impressive than it sounds and more impressive than it sounds, in that order.

Die Architektur, in einfachen Worten

Kimi K3 is a mixture-of-experts model that activates 104 billion parameters per token, selecting 16 experts from a pool of 896. So while the total is 2.8T, any given forward pass touches under 4% of the weights. That is where the economics come from: enormous stored knowledge, moderate inference cost.

The interesting engineering is in the attention stack. K3 uses 93 layers built from 69 KDA (Kimi Delta Attention) layers and 24 Gated MLA layers — a hybrid linear-to-full attention design at roughly a 3:1 interleave ratio. KDA is a linear-attention variant that runs in linear time relative to sequence length rather than quadratic, which is what makes the 1M context economically viable instead of merely advertised. Moonshot reports up to a 75% KV-cache reduction and up to 6× decode throughput at 1M context as a result.

Layered on top is Attention Residuals (AttnRes), which Moonshot describes as a drop-in replacement for standard residual connections, and a Stable LatentMoE framework for routing. The combination is claimed to deliver roughly a 2.5× improvement in overall scaling efficiency versus Kimi K2.

Vision comes from MoonViT-V2, a 401M-parameter encoder, and it is native rather than bolted on — K3 handles text, images and video in one model.

One more detail that matters for anyone thinking about self-hosting: K3 was trained with native MXFP4 quantisation-aware training, with experts in MXFP4 and activations in MXFP8. This is not a post-hoc quantisation. The published checkpoint is the quantised model, which is why 2.8 trillion parameters land at roughly 1.5 TB on disk rather than the ~5.6 TB a BF16 checkpoint of that size would need.

The benchmark story

Moonshot’s published numbers are strong across the board, and unusually, several have held up under independent evaluation:

  • Terminal-Bench 2.1: 88.3 — the highest published score on that version of the benchmark
  • FrontierSWE: 81.2
  • DeepSWE: 67.5
  • BrowseComp: 91.2 — state of the art for agentic web research
  • GPQA Diamond: 93.5
  • MMMU-Pro: 81.6 / 83.4 (vision)
  • SWE Marathon: 91.0
  • OfficeQA Pro: 81.3
  • AutomationBench: 30.8
  • LMArena Frontend Code Arena: 1,679 Elo — first place, ahead of Claude Fable 5 (1,631) and GPT-5.6 Sol (1,618), winning six of seven frontend domains

That last one deserves a pause. An open-weight model you can download for free took first place on a human-preference frontend coding arena, ahead of the flagship closed models from the two best-funded labs in the world. Whatever you think about arena benchmarks as a methodology, that is a headline that would have been unthinkable in 2025.

On the aggregate indexes it lands slightly lower: 59.7 on the Artificial Analysis Intelligence Index, third overall, behind Claude Opus 5 (63.0) and Claude Fable 5 (62.1). Moonshot’s own showcase numbers — GDPval-AA v2 at 1,687 und AA-Briefcase at 1,527 — put it third and second respectively in those knowledge-work evaluations. On the Vals Index v2 it lands at 57.8%, against Claude Opus 5’s 67.2%, Claude Fable 5’s 66.0% and GPT-5.6 Sol’s 63.7% — the widest open-versus-closed gap in any index I found, and worth holding in mind against the coding numbers.

Bar chart of the cost of one 60-turn agentic coding task showing Claude Fable 5.1 at 11.94 dollars, Kimi K3 at 4.07, GLM-5.3 at 1.85 and GLM-5.3-Flash at 0.21
The same 60-turn agentic task at each vendor’s list pricing. The vertical axis is the entire argument for open weights.

The demos, and how to read them

Moonshot published two long-horizon autonomy demonstrations that got a lot of attention: a 48-hour autonomous run completing a full chip design pipeline (a 4mm design closing timing at 100 MHz, simulating at 8,700+ tokens/second), and a reproduction of the astrophysics I-Love-Q relation in about two hours, against the one-to-two weeks a senior researcher would typically need.

These are genuinely impressive and I would not build a purchasing decision on them. Vendor-run, vendor-scaffolded, vendor-selected demonstrations tell you what a model can do on a good day with expert supervision, not what it does in your repository on a Tuesday. Treat them as an existence proof for long-horizon capability, not as a spec.


Show Image Two published architectures and one black box. The open models tell you exactly how they work; the closed one tells you what it scores.

What GLM-5.3 Actually Is

GLM-5.3 has the strangest release story of the three, and it is the one worth telling properly, because it is the first time a major lab has visibly hesitated over publishing weights for reasons other than commercial ones.

Z.ai announced GLM-5.3 on 14 August 2026 with the tagline “Built to Code. Ready for Cyber Defense.” The model is a 753-billion-parameter mixture of experts, with roughly 40 billion active per token according to community estimates derived from the GLM-5.2 configuration. It carries a 1M-token context window and a 128K max output, and — unlike GLM-5.3-Flash, which is a completely different model — it is text-only.

“Scaling post-training is all we did”

The most interesting technical fact about GLM-5.3 is that there is no new base model. It reuses the same pretrained foundation as GLM-5.2. Every gain came from dramatically extended post-training, which Z.ai summarised as: “Scaling post-training is all we did for GLM-5.3.”

The deltas are not small:

BenchmarkGLM-5.2GLM-5.3
Terminal-Bench 3.04.628.3
DeepSWE v1.146.266.9
CyberGym77.2%84.5%

A six-fold improvement on Terminal-Bench and a twenty-point jump on DeepSWE, from the same base weights, is a genuinely important result. It says the frontier still has a lot of headroom in post-training, which is by far the cheaper half of the training bill. That is precisely why Chinese labs operating with, as Lambert puts it, “orders of magnitude less capital” than American ones are keeping pace.

The two weeks Z.ai spent not shipping

Z.ai said at launch that the weights would follow “in stages following rigorous safety evaluations,” with a target of roughly two weeks. Its Hugging Face repository listed 28 August. That date came, went by a matter of hours, and then the weights appeared alongside a technical writeup explaining the delay.

The reason was cyber capability. During evaluation, Z.ai reported that GLM-5.3 identified 2,436 vulnerabilities across 269 open-source projects, of which 1,097 were rated critical or high severity. More to the point, the model demonstrated multi-stage exploit-chain reasoning — autonomously chaining exploitation steps together — which Z.ai characterised as behaviour that was “not fully intended” and that “continued compounding as training scaled.”

I want to be careful here, because this is easy to read as either marketing or alarmism and it is probably neither. A model that is good at finding vulnerabilities is good at finding vulnerabilities; the defensive and offensive uses are the same capability pointed in different directions. Z.ai’s benchmark spread reflects that honestly: CyberGym 84.5% (vulnerability discovery, where it leads) against ExploitBench 54.4% (exploitation, where Claude Fable 5 scores 78.0%). The lab deliberately optimised for the defensive half and it shows.

What is genuinely notable is that a lab took a two-week hold on a flagship release and published its reasoning. Whether you find the reasoning convincing, the precedent is new.

Where GLM-5.3 actually lands

Z.ai’s published comparison chart does not claim a clean sweep, which is refreshing:

BenchmarkGLM-5.3Best comparator
AutomationBench48.2%GLM-5.3 leads
CyberGym84.5%GLM-5.3 leads
GDPval-AA v21,769GLM-5.3 leads
Terminal-Bench 3.028.3%GPT-5.6 Sol: 34.6%
DeepSWE v1.166.9%GPT-5.6 Sol: 72.7%
ExploitBench54.4%Claude Fable 5: 78.0%
Code Bench (50K tokens)31.4%

Independently, Artificial Analysis puts it at 59.5 on the Intelligence Index — statistically indistinguishable from Kimi K3’s 59.7, from a model with roughly a quarter of the parameters.

One caveat from that same evaluation that nobody in the marketing mentions: GLM-5.3 generated 170 million output tokens across the Artificial Analysis suite against a 72 million median for its class. It is verbose. logische_Überlegung defaults to max and thinking cannot be disabled, so you are paying for a lot of deliberation whether the task earns it or not. Even so, the total cost to run the full evaluation came to $0.68 per task, against $0.84 for Kimi K3 and $2.34 for Claude Opus 5 — so the verbosity does not erase the price advantage. It just narrows it.


What Claude Fable 5.1 Actually Is

Anthropic released Claude Fable 5.1 on 1. September 2026, alongside Claude Mythos 5.1 — the same model with different safeguard configuration, available only through trusted-access programmes for cybersecurity and life-sciences work.

The positioning is specific and, I think, correct: this is “a model built for work that does not finish in one prompt.” Anthropic is not claiming a general intelligence leap. It is claiming that sustained, multi-hour agentic sessions go better.

The numbers

BenchmarkFable 5Fable 5.1Mythos 5.1
Terminal-Bench 4.042.0%55.8%60.9%
Terminal-Bench-Science 0.124.7%52.6%
AutomationBench17.1%31.4%
CursorBench 3.2.070.5%73.4%
OSWorld 2.0 (partial)72.9%77.9%
Humanity’s Last Exam (with tools)63.8%65.0%
GDPval-AA v21,853

For external context on the same ruler: GPT-5.6 Sol scores 52.3% on Terminal-Bench 4.0, so Fable 5.1’s 55.8% takes the lead and Mythos 5.1’s 60.9% extends it.

The Terminal-Bench-Science jump — 24.7% to 52.6%, a clean doubling — is the one I would weight most heavily if your work involves scientific computing or data pipelines. And on Browserbase’s hardest computer-use benchmark, Fable 5.1 completed 82% of tasks against Claude Opus 5’s 74%.

Anthropic also reports roughly a 60% reduction in cyber-safety false positives — the model refusing legitimate security work because it pattern-matched to something dangerous. If you have ever had a coding assistant decline to help you write a rate limiter, you will understand why that matters.

The pricing change is the actual news

Headline rates did not move: $10 per million input tokens, $50 per million output. What changed is the cache read multiplier, and it changed a lot.

Fable 5Fable 5.1
Eingabe$10.00$10.00
Ausgabe$50.00$50.00
Cache write (5 min)$12.50$12.50
Cache write (1 hour)$20.00$20.00
Cache-Lesezugriff$1.00$0.25

That is a 75% cut, achieved by moving the cache-read multiplier from the standard 0.1× of base input to 0.025×. Anthropic quotes roughly 25% cheaper for typical workloads and up to 45% cheaper for highly agentic work — and that second number is real, because agentic loops are overwhelmingly cache reads. A 60-turn agent session re-reads the same accumulated context dozens of times. The Batch API halves input and output on top of that, landing at $5/$25.

This is a smart, targeted price cut. It makes the specific thing Fable 5.1 is best at — long sessions — meaningfully cheaper without devaluing the model generally.

Three breaking changes, one regression

Before you upgrade a production integration, read these properly.

1. Forced tool use is gone. Requests using tool_choice: "any" oder tool_choice: "tool" now return a 400 error. Thinking is always on and cannot be bypassed, and forcing a tool call would skip it. You must use tool_choice: "auto".

2. Thinking blocks are one-directional. Fable 5.1 can read thinking blocks produced by earlier Claude models, but earlier models cannot read Fable 5.1’s. Only Mythos 5.1 can. If your architecture switches models mid-conversation to save money, that fallback now forces a re-plan.

3. Editing history invalidates thinking blocks. For accounts created after 31 August 2026, editing prior turns, the system prompt, or the tool array invalidates all subsequent thinking blocks — errors or silent drops. This is an explicit anti-distillation measure and it breaks a common pattern where agent frameworks rewrite context between turns.

And one regression worth knowing: parallel tool calling is reportedly more variable in 5.1, sometimes issuing a single call per turn where 5 would batch several. On a long agent loop that costs wall-clock time.

There is also a tokenizer detail that catches people out on cost comparisons: Fable 5.1 uses the same tokenizer as Claude Opus 4.7, which produces roughly 30% more tokens than the tokenizer in older Claude models for the same text. If you are benchmarking cost against a 2025 baseline, adjust for that before concluding anything.


The Benchmark Problem: Three Models, Three Different Rulers

Here is the part most comparison articles get wrong, and it is worth being blunt about.

You cannot line these three models up on Terminal-Bench. Look:

  • Kimi K3: 88.3 on Terminal-Bench 2.1
  • GLM-5.3: 28.3 on Terminal-Bench 3.0
  • Fable 5.1: 55.8 on Terminal-Bench 4.0

Those are three different benchmarks that share a name. Each version got substantially harder than the last — that is the entire point of releasing a new version. Reading those three numbers as a ranking would tell you Kimi K3 is three times better than GLM-5.3 at terminal work, which is nonsense.

The same problem applies to AutomationBench, where version numbers are inconsistently reported, and to the SWE-bench family, where the Verified, Pro and Marathon variants measure genuinely different things. Labs are not doing this to deceive anyone — they run the evaluation that was current when they trained, and evaluations turn over every few months now. But the effect on a casual reader is the same as deception, so it is worth flagging loudly.

What you can legitimately compare, within a version:

Auf Terminal-Bench 2.1: Kimi K3 (88.3), GPT-5.6 Sol (88.8), Claude Fable 5 (88.0) and GLM-5.3-Flash (84.3). Four models, one ruler. This is the comparison that actually matters and I have given it its own table in the next section.

Auf Terminal-Bench 4.0: Fable 5.1 (55.8) against GPT-5.6 Sol (52.3). Fable wins.

Beyond that, be sceptical of anyone showing you a single bar chart with all three of this article’s headline models on a Terminal-Bench axis. It cannot be done honestly.


Where They Actually Meet: The Cross-Comparable Numbers

Strip out the benchmark-version mismatches and three comparisons survive.

1. Artificial Analysis Intelligence Index (September 2026)

Independently run, same methodology across every model, no vendor involvement in selection.

RankModellScoreFreihanteln
1Claude Opus 563.0Nein
2Claude Fable 562.1Nein
3Grok 4.660.9Nein
4Kimi K359.7Ja
5GLM-5.359.5Ja
6GPT-5.6 Sol58.9Nein
7Qwen3.8 Max Preview58.1Weights pending
8GLM-5.3-Flash57.5Yes (MIT)
31MiniMax M345.4Ja

Fable 5.1 had not been independently scored at the time of writing — it launched the day before. Given 5.1’s benchmark deltas over Fable 5, expect it to land at or above 62.1 when it is.

A note on precision: different snapshots of this index round differently — one September pull has Kimi K3 and GLM-5.3 both at 60 and GPT-5.6 Sol at 61. The ordering at the top is stable; the sub-point gaps are not. Do not build an argument on half a point.

Read that table carefully. The best open-weight model is 3.3 points behind the best closed model on the broadest neutral index available. A year ago that gap was double digits.

2. Terminal-Bench 2.1 — the one version where three of them meet

This is the single most useful table in the article, and I nearly missed it. Kimi K3, Claude Fable 5 and GPT-5.6 Sol were all scored on the same version of Terminal-Bench:

ModellTerminal-Bench 2.1Freihanteln
GPT-5.6 Sol88.8Nein
Kimi K388.3Ja
Claude Fable 588.0Nein
GLM-5.3-Flash84.3Yes (MIT)

Half a point separates a model you can download from the best closed model on the same benchmark, on the same version, on terminal-native agentic work.

That is the number to remember. Not the index aggregate, not the vendor chart — this one. On the specific task shape that agentic coding actually is, the gap is inside the noise.

One caveat, stated by the source and worth repeating: all of Kimi K3’s published figures are max-effort runs, at logische_Überlegung maximum and temperature 1.0. Different labs use different evaluation harnesses and different effort settings, so even same-version comparisons carry more uncertainty than a clean table suggests.

3. Cost to run the same evaluation suite

If the table above is the best capability comparison, this is the best value comparison — because it is the only figure that bundles capability, verbosity, reasoning overhead and price into a single number:

ModellCost per Intelligence Index task
GLM-5.3$0.68
Kimi K3$0.84
Claude Opus 5$2.34

Same tasks, same scoring, actual invoices. GLM-5.3 delivers 94% of Opus 5’s index score for 29% of the cost.

4. GDPval-AA v2 (knowledge work)

ModellScore
Claude Fable 5.11,853
Claude Fable 5 Max1,815
GLM-5.31,769
GPT-5.6 Sol Max1,747.8
Kimi K31,687

This one is worth caveating: it is assembled from vendor-published charts rather than a single independent run, and the “Max” suffixes indicate different effort settings. But the ordering is broadly consistent across sources, and it says something real — on general knowledge work, as opposed to coding, the closed models still hold a clearer lead than the coding benchmarks suggest.

That is the pattern across all three comparisons. The open models have essentially caught up on coding and agentic tasks. They are still a step behind on broad knowledge work. If your workload is the former, the case for paying closed-model prices is weak. If it is the latter, it is still defensible.


Show Image Three and a bit points separate the best closed model from the best downloadable one. That gap was double digits twelve months ago.

The Cost Maths, With Real Numbers

Benchmarks are abstract. Invoices are not.

Das Szenario

One medium feature, done agentically: roughly 60 Werkzeugwechselvorgänge, im Durchschnitt 45,000 input tokens per turn je mehr Kontext sich ansammelt, und 2,000 output tokens per turn.

  • Gesamteingabe: 2,7 Millionen Token
  • Gesamtleistung: 120.000 Token
  • Angenommene Trefferquote des Prompt-Caches: 80% (2,16 Mio. im Cache, 540.000 aktuell)

Claude Fable 5.1

KomponenteTokenBewertungKosten
Neue Informationen540.000$10.00/M$5.40
Zwischengespeicherte Eingabe2,16 Mio.$0.25/M$0.54
Ausgabe120.000$50.00/M$6.00
Gesamt≈ $11.94

Kimi K3

KomponenteTokenBewertungKosten
Neue Informationen540.000$3.00/M$1.62
Zwischengespeicherte Eingabe2,16 Mio.$0,30/M$0.65
Ausgabe120.000$15.00/M$1.80
Gesamt≈ $4.07

GLM-5.3

KomponenteTokenBewertungKosten
Neue Informationen540.000$1.40/M$0.76
Zwischengespeicherte Eingabe2,16 Mio.$0.26/M$0.56
Ausgabe120.000$4.40/M$0.53
Gesamt≈ $1.85

And for reference — GLM-5.3-Flash

KomponenteTokenBewertungKosten
Neue Informationen540.000$0,15/M$0.08
Zwischengespeicherte Eingabe2,16 Mio.$0.03/M$0.06
Ausgabe120.000$0,50/M$0.06
Gesamt≈ $0,21

The ratios

  • GLM-5.3 is 6.5× cheaper than Fable 5.1 for identical work
  • Kimi K3 is 2.9× cheaper than Fable 5.1
  • GLM-5.3 is 2.2× cheaper than Kimi K3
  • GLM-5.3-Flash is 58× cheaper than Fable 5.1, at 57.5 on the intelligence index against Fable 5’s 62.1

Scaling to a working month

Vier solcher Aufgaben pro Tag, zwanzig Arbeitstage – 80 agentische Durchläufe:

ModellMonthly token cost
Claude Fable 5.1≈ $955
Kimi K3≈ $326
GLM-5.3≈ $148
GLM-5.3-Flash≈ $17

Three honest caveats on these figures

Cache writes are excluded. Fable 5.1 charges $12.50/M for 5-minute cache writes, and a long session writes cache repeatedly. Including writes moves Fable’s real number up, not down. The 80% hit rate is also optimistic for short sessions.

Reasoning tokens are billed output tokens. All three models have always-on thinking. GLM-5.3 in particular is documented as verbose — 170M output tokens against a 72M class median across the AA suite — so its real-world output volume runs above what a naive turn count suggests. The $0.68-per-task figure already accounts for this, which is why I trust it more than my own model above.

Fable 5.1’s tokenizer produces ~30% more tokens than older Claude models for the same text. If you are comparing against a historical Claude bill, adjust.

Even after all three corrections, the ordering does not change and the magnitude barely does. Closed-frontier work costs roughly six times what the best open-weight model costs, per completed task.


Show Image The same 60-turn agentic task at each vendor’s list pricing. The vertical axis is the entire argument for open weights.

Licences: What “Open” Actually Buys You In 2026

This is the section I would most like people to read, because the vocabulary has drifted badly and it is going to cost somebody a lot of money.

Neither Kimi K3 nor GLM-5.3 is open source. Both publish downloadable weights under bespoke licences with revenue-triggered conditions. That is a meaningfully different thing from MIT or Apache 2.0, and the difference lands squarely on the kind of business that would most benefit from self-hosting.

The Kimi K3 License

You will see “Modified MIT” repeated all over the internet for K3. It is wrong — that described K2. K3 ships under a custom document with two distinct gates:

The Model-as-a-Service gate. If you provide third parties with access to model inference or fine-tuning — where those third parties control inputs, parameters or training data — and “the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars” over any consecutive 12 months, you must negotiate a separate commercial agreement with Moonshot. Note aggregate revenue, across all affiliates, not revenue attributable to K3. A €25M-turnover consultancy that resells inference is over the line even if the AI business is a rounding error.

The attribution gate. Products exceeding 100 million monthly active users or $20 million monthly revenue must display “Kimi K3” prominently in the product interface.

The exemption that matters most: purely internal use is unrestricted. Running K3 to support your own developers, researchers, legal team or general employee productivity does not trigger either gate. For the overwhelming majority of businesses — including essentially all of mine — that is the relevant clause, and the answer is that you are fine.

The clause that is new relative to K2 is the $20M MaaS gate. K2 only required attribution above its thresholds. K3 adds a much lower revenue gate aimed specifically at commercial inference resellers. Moonshot is not trying to stop you using the model; it is trying to stop cloud providers building a business on it for free.

The glm-5.3 License

Z.ai’s flagship went a different direction, and it is a bigger break with precedent than most coverage acknowledged.

GLM-5.2 shipped under MIT. GLM-5.3-Flash shipped under MIT. GLM-5.3 did not.

The custom licence permits individuals and ordinary businesses to run, deploy and fine-tune the model with no additional restrictions. The single gate is aimed at hyperscalers: an entity whose “aggregate revenue of the Licensee and its affiliates exceeds 10 billion US dollars (or the equivalent in other currencies) in total over any consecutive 12 months” muss pass Z.AI’s security review before using the model or its derivatives for any commercial purpose.

That is a $10 billion threshold. It excludes almost everyone. But two details in the fine print deserve attention:

First, the licence explicitly carves out app developers embedding the model in features, and resellers merely routing requests to other hosts. The target is narrow and clearly labelled: companies that want to host GLM-5.3 commercially at scale.

Second — and this is the genuinely awkward part — the licence publishes no criteria, timeline or appeal process for that security review. It states only that its “scope and method… shall be reasonably determined by Z.AI.” If you are a large enterprise, that is an unbounded dependency on a foreign vendor’s discretion, and your procurement team will say so.

Also worth noting: despite the two-week cyber-capability hold that preceded release, the licence contains no acceptable-use restrictions, no cyber carve-outs and no output-ownership claims. The safety concern shaped the release timing, not the release terms.

So what is actually still MIT?

GLM-5.3-Flash. 320B parameters, 18B active, natively multimodal, MIT-licensed, 57.5 on the intelligence index, and $0.15/$0.50 per million tokens — currently half that under a promotional discount running to 9 September 2026. If your requirement is genuine unrestricted permissive licensing rather than merely downloadable weights, that is where the frontier currently sits, and it is remarkably high.

The practical summary

Kimi K3GLM-5.3GLM-5.3-FlashFable 5.1
Weights downloadableJaJaJaNein
OSI-style open sourceNeinNeinYes (MIT)Nein
Restriction trigger$20M revenue (MaaS)$10B revenue (MaaS)KeineK.A.
Attribution requiredAbove 100M MAU / $20M-monthCopyright notice onlyCopyright notice onlyK.A.
Internal use restrictedNeinNeinNeinK.A.
Fine-tuning permittedJaJaJaNein
Acceptable-use policyStandardNone in licence textKeineAnthropic AUP applies

For a typical German Mittelstand business, an agency, or a consultancy running models internally or embedding them in a product: all three open options are usable without negotiating anything. The gates exist to catch AWS, not you. But “read the licence” has gone from pedantic advice to actual advice, and if you are above about €20M turnover and reselling inference in any form, get it in front of counsel before you deploy.


The Real Case For Open Models

I want to make this argument properly rather than as a slogan, because “open models are great” is not an argument and the actual reasons are more interesting than the cheerleading.

1. The price collapse is not a subsidy, it is structural

The instinctive objection to cheap Chinese models is that the pricing is loss-leading and will normalise upward. Some of it certainly is. But the structural point stands independently: when weights are public, inference becomes a commodity market. Twenty hosting providers compete to serve the same model, and none of them can charge a capability premium because none of them has a capability moat. That is why GLM-5.3 runs at $1.40/$4.40 and Fable 5.1 runs at $10/$50 for scores three points apart.

Closed-model pricing includes the R&D amortisation, the margin, and the fact that there is exactly one seller. Open-weight pricing includes electricity, hardware amortisation and a thin margin. Those are different businesses, and the gap will not close by open models getting more expensive.

The evidence that this is real rather than temporary: DeepSeek captured roughly 17% of token usage on Vercel by May 2026 while holding about 1% of revenue share. That is not a pricing anomaly. That is what commoditisation looks like on a chart.

2. Exit rights change your negotiating position even if you never exercise them

Most teams reading this will not self-host. The hardware section below explains why — Kimi K3 needs a rack. But the option has value regardless.

If Anthropic raises prices, deprecates the model you built on, changes its acceptable-use policy in a way that breaks your product, or has a bad quarter, your recourse with a closed model is a migration project. With an open-weight model your recourse is downloading a file you could have downloaded any time. You may never do it. The fact that you could is what makes the vendor relationship a commercial one rather than a dependency.

This is not theoretical. Fable 5.1 shipped with three breaking changes on 1 September, one of which invalidates thinking blocks when you edit conversation history for accounts created after 31 August. If that pattern is load-bearing in your architecture, you are re-engineering on Anthropic’s schedule, not yours.

3. Data residency stops being a project

For anyone operating under GDPR, this is the argument that actually closes deals.

With a closed API, every prompt leaves your infrastructure. You need a processor agreement, a transfer impact assessment if the processor is outside the EEA, a data-flow map, and an answer for your DPO about what happens to the data at rest. All of that is doable. It is also weeks of work per vendor, repeated whenever the vendor changes anything.

With open weights on infrastructure you control — your own hardware, or a German or EU GPU host — the cross-border transfer question does not arise, because there is no transfer. The model runs where your data already lives. Legal review goes from a transfer assessment to a licence read.

For German and EU clients this is frequently the deciding factor, and it is worth being precise about the caveat: using GLM-5.3 through Z.ai’s API does not give you this. You get it from running the weights yourself, or from a provider hosting them in your jurisdiction. The licence makes that legal; it does not make it automatic.

4. Inspectability is real, if underused

Open weights mean you can examine the architecture, run your own evaluations on the actual model rather than an endpoint that might be silently updated, quantise it to fit your hardware, fine-tune it on your domain, and — importantly — pin a version forever.

That last one is underrated. Closed APIs get updated behind stable model IDs. Behaviour drifts. Prompts that worked stop working, and you cannot diff the change because you cannot see it. A local checkpoint is byte-identical in a year’s time.

5. The competitive pressure benefits everyone, including closed-model users

Look at what happened this summer. Kimi K3 lands in July at $3/$15 with frontier-adjacent scores. GLM-5.3 lands in August at $1.40/$4.40 within half a point of it. And on 1 September, Anthropic cut cache reads by 75%.

I am not claiming direct causation — Anthropic does not publish its pricing rationale and cache-read cuts are a natural optimisation. But a market with credible cheap substitutes prices differently from one without them, and the substitutes got credible this year. If you use closed models exclusively, the open-weight ecosystem is still quietly making your bills smaller.

6. Post-training is where the gains are, and it is cheap

GLM-5.3’s headline result — a six-fold Terminal-Bench improvement from the same base weights — is the most strategically important number in this entire article, and it is not about GLM.

It says the expensive part of building a frontier model (pretraining) is increasingly a solved, commoditised input, and the differentiating part (post-training, RL environments, long-horizon task design) is comparatively affordable. That is why four Chinese labs with a combined valuation of about $159 billion are keeping pace with American labs valued at multiples of that. Lambert’s assessment is that Chinese labs are running with “orders of magnitude less capital.”

For anyone downstream, the implication is straightforward: the number of credible model vendors is going up, not down. Plan your architecture accordingly.

7. The ecosystem effects compound

Qwen passed one billion cumulative downloads on Hugging Face, overtaking Llama, and now anchors over 200,000 tagged models and 113,000+ derivatives — roughly 40% of all new LLM derivatives on the platform are Qwen-based. That is a tooling, quantisation, fine-tuning and deployment ecosystem that exists only because the weights are public.

You benefit from that ecosystem whether or not you contribute to it. GGUF conversions, vLLM kernels, LoRA adapters, quantisation recipes, evaluation harnesses — all of it exists because thousands of people could get their hands on the actual weights.

8. It is where the developers already went

Chinese-origin models at ~61% of OpenRouter tokens. US model share down from ~70% to ~30% in a year. Four of the five most-used models on the router are Chinese-origin. Xiaomi’s MiMo models alone account for roughly 21% of routed tokens and about 22% of all coding traffic.

Developer behaviour is a leading indicator. It was a leading indicator for Docker, for Postgres, for Linux. Betting against it has a poor historical record.

Horizontal bar chart of Terminal-Bench 2.1 scores showing GPT-5.6 Sol at 88.8, Kimi K3 at 88.3, Claude Fable 5 at 88.0 and GLM-5.3-Flash at 84.3, with open-weight models in teal
Four models, one version of the benchmark. Half a point separates the best model you can download from the best one you cannot.

…And The Honest Case Against

If I only wrote the section above, this would be an advert.

Open weights are not open source, and the vocabulary drift is a real problem. Two of the three models here ship under bespoke revenue-gated licences with no OSI approval. GLM-5.3’s security-review clause has no published criteria or appeal process. If your compliance framework requires OSI-approved licensing, the honest answer is that your options are GLM-5.3-Flash, DeepSeek’s MIT-licensed flagships, and a shrinking list of others.

Vendor benchmarks remain vendor benchmarks. Every headline number Moonshot and Z.ai published is self-reported and self-selected. Where neutral aggregators have measured, the numbers broadly hold up — which is genuinely to both labs’ credit — but “broadly hold up” is doing work in that sentence.

Self-hosting is a fantasy for most teams. Kimi K3 is 1.5 TB and Moonshot recommends 64+ accelerators. GLM-5.3 needs 8×H200 at FP8 as a bare minimum. The exit right is real; the exercise cost is a data centre.

Support is what you make it. When Fable 5.1 breaks, there is a company with an SLA. When your self-hosted GLM-5.3 deployment produces garbage at 300K context on a Friday night, there is a GitHub issue and your own competence.

Geopolitics is a real procurement input. All three open-weight options discussed here come from Chinese labs. For some clients — public sector, defence-adjacent, certain regulated industries — that is a hard blocker regardless of licence terms or where the weights run. It is not my job to tell you whether that concern is well-founded. It is my job to tell you it will come up in the meeting.

The closed models still lead on general knowledge work. GDPval-AA v2 has Fable 5.1 at 1,853 against GLM-5.3’s 1,769 and Kimi K3’s 1,687. On coding the gap is gone. On broad professional knowledge work it is not.


Show Image Herunterladbare Gewichte, drei verschiedene Saitensätze. Nur GLM-5.3-Flash unterliegt weiterhin der MIT-Lizenz.

Ollama In 2026: The Pricing Change That Actually Matters

Ollama has quietly become the most important piece of infrastructure in this conversation, and on 31 August 2026 it changed how it charges in a way that is worth understanding properly.

What changed

Ollama moved its Pro, Max and Team plans from GPU-time billing to industry-standard per-token pricing, with a pool of usage credits included in every plan. The company’s stated reason is refreshingly concrete: models like Kimi K3 made GPU-time metrics impossible to predict. When one model activates 104B parameters and another activates 3B, “an hour of GPU” stops meaning anything to the person paying.

The new plans

PlanPreisIncluded monthly usageConcurrency
Kostenlos$0Small monthly credit, starter models1 request
Pro$20/mo (or $200/yr)$60 of usage3 requests
Max$100/mo$300 of usage10 requests
Team$500/mo$1,000 shared, unlimited users10 requests
UnternehmenCustomCustomCustom

Read the Pro row again. $20 buys $60 of tokens. That is a 3× multiplier on included usage, and when the pool runs out you continue at exactly the same published per-token rate — no penalty tier, no service fee.

What was removed

This is the part developers actually noticed: no 5-hour resets and no weekly caps. Anyone who has hit a rolling usage window mid-refactor will understand why that mattered more than the price. The monthly pool refreshes on your subscription date and, importantly, does not roll over — so size your plan to your normal month, not your busiest one.

The terms that matter for EU work

Three commitments in the announcement are directly relevant if you are handling client data:

  • Zero data retention. Prompts and responses are never logged and never trained on.
  • Hosting in the US and Europe, plus Singapore for a limited set of Qwen models.
  • Per-request cost visibility in your account — you can see exactly what each call cost.

For a German consultancy that is a materially better compliance story than most gateway providers offer, though “hosted in Europe” is a routing statement rather than a contractual data-residency guarantee. If residency is a hard requirement rather than a preference, ask Ollama for it in writing before you assume it.

The actual per-token rates

This is where it gets interesting. Ollama publishes rates per model, and they track first-party pricing closely:

Model on OllamaEingabe / 1MCached / 1MAusgabe / 1M
kimi-k3$3.00$0.30$15.00
glm-5.3$1.40$0.26$4.40
mistral-large-3$0.50$0.50$1.50
gemma4$0.14$0.05$0.40
nemotron-3-super$0.015$0.015$0.60

Run the monthly maths from earlier through this. Eighty agentic runs a month on GLM-5.3 costs about $148 in tokens. A Max plan at $100/month covers $300 of usage — so the same workload that would cost roughly $955/month on Fable 5.1 fits comfortably inside a $100 Ollama subscription with headroom to spare.

That is the entire open-model economic argument compressed into one line on an invoice.

Show Image Eighty agentic runs a month on GLM-5.3 fit inside a $100 Ollama Max plan with change. The same work on Fable 5.1 is a $955 invoice.

What else Ollama shipped in 2026

The pricing change did not happen in isolation. The last few months have been busy:

  • 25 August — Claude Desktop support. Ollama now works as a third-party gateway provider for Claude Desktop, so you can drive open models through Anthropic’s own client.
  • 26 August — v0.33.0. The Claude Desktop gateway integration landed, prefill recovery was restored to the cache, and Ollama disabled Claude Code’s token-countdown system message, which was invalidating the KV cache on every turn. That last one is a quiet but significant performance fix.
  • 26 August — v0.33.1. MLX support for Qwen3.8 Flash Next, structured output, and a fix for GPU timeouts when loading models from slower storage.
  • 28 August — v0.33.2. Dark mode restored, macOS handoff fixed, and Claude Desktop proxy requests kept alive during model catalogue updates.
  • 20 August — v0.32.15. New desktop onboarding, and metadata caching between requests that cut time-to-first-token by roughly half.
  • 11 August — NVIDIA Nemotron 3.5 Lightning, a 30B model tuned for agentic workflows with tool calling on personal hardware.
  • 10 August — Meta’s Muse Glimmer, a 30B multimodal model under Apache 2.0, accelerated by Ollama’s MLX engine. Worth noting given Llama’s collapse in routed usage — Meta is still shipping genuinely permissive weights.
  • 9 July — $88M funding round, with Ollama reporting 8.9 million developers.
  • 29 June — Gemma 4 on MLX up to 90% faster via multi-token prediction, which disproportionately benefits coding agents on Apple Silicon.
  • 5 June — v0.30 added GGUF compatibility through llama.cpp, broadening hardware support well beyond Apple Silicon.

The through-line is that Ollama has stopped being “the easy way to run a small model on your laptop” and become a general-purpose gateway that happens to also run models locally. The cloud catalogue now includes kimi-k3:cloud, glm-5.3:cloud, glm-5.3-flash, deepseek-v4-pro, deepseek-v4-flash, minimax-m3, qwen3.5 across seven sizes, gpt-oss at 20b and 120b, gemma4, the Nemotron 3 family and mistral-large-3.


Setting It All Up: Copy-Paste Configs

Ollama + Claude Code (the fastest path)

Ollama ships a one-command launcher that handles the environment wiring for you:

bash

ollama launch claude

If you would rather do it manually — which you should if you are scripting it — install Claude Code first:

bash

# macOS / Linux
curl -fsSL https://claude.ai/install.sh | bash

# Windows (PowerShell)
irm https://claude.ai/install.ps1 | iex

Then point it at Ollama:

bash

export ANTHROPIC_AUTH_TOKEN=ollama
export ANTHROPIC_API_KEY=""
export ANTHROPIC_BASE_URL=http://localhost:11434

And run against whichever model you want:

bash

# A local model
claude --model qwen3.5

# A cloud model — note the :cloud suffix
claude --model kimi-k3:cloud

Or inline, without exporting anything globally:

bash

ANTHROPIC_AUTH_TOKEN=ollama \
ANTHROPIC_API_KEY="" \
ANTHROPIC_BASE_URL=http://localhost:11434 \
claude --model glm-5.3:cloud

Two things that will bite you.

Erstens, do not export ANTHROPIC_BASE_URL permanently in your shell profile. It is a global variable and it will silently redirect every Anthropic-speaking tool on your machine to Ollama, including the one you wanted talking to Anthropic. Scope it per-project or per-invocation.

Zweitens, set your context length to 64k or higher for anything working on a real repository. Ollama’s default is smaller, and a coding agent that quietly runs out of window will just start forgetting things rather than erroring.

For CI, Docker or scripted runs, --yes skips the interactive prompts:

bash

ollama launch claude --model glm-5.3:cloud --yes -- -p "how does this repository work?"

GLM-5.3 direct from Z.ai

If you want first-party routing rather than going through Ollama, Z.ai exposes three protocol endpoints:

ProtokollBasis-URL
Anthropische Botschaftenhttps://api.z.ai/api/anthropic
OpenAI-Chat-Vervollständigungenhttps://api.z.ai/api/coding/paas/v4
Antworten von OpenAIhttps://api.z.ai/api/v1

An OpenCode provider block for GLM-5.3:

json

{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "zai": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "Z.AI",
      "options": {
        "baseURL": "https://api.z.ai/api/coding/paas/v4",
        "apiKey": "{env:ZAI_API_KEY}"
      },
      "models": {
        "glm-5.3": {
          "name": "GLM-5.3",
          "limit": { "context": 1000000, "output": 128000 },
          "options": { "reasoning_effort": "high", "temperature": 1 }
        }
      }
    }
  },
  "model": "zai/glm-5.3"
}

Important upgrade note: GLM-5.3 requires reasoning to be enabled. If your existing configuration sets thinking.type: "disabled", that will now fail. Change it to "enabled" and set reasoning_effort: "low" if you want the old latency profile back.

logische_Überlegung is your main quality-versus-cost dial: niedrig for latency-sensitive calls, hoch for substantial work, max (the default) for problems that have genuinely defeated hoch. Given the documented verbosity, I would run hoch as an everyday setting rather than leaving max on and wondering where the output tokens went.

The routing setup I would actually recommend

None of this is a “pick one” decision. The correct answer for most teams is a three-tier routing table, and every serious agent harness supports per-agent models.

json

{
  "$schema": "https://opencode.ai/config.json",
  "model": "zai/glm-5.3",
  "small_model": "ollama/gemma4",
  "agent": {
    "plan": {
      "description": "Architecture, root-cause analysis, anything expensive to get wrong",
      "model": "anthropic/claude-fable-5-1"
    },
    "build": {
      "description": "The bulk of implementation work",
      "model": "zai/glm-5.3",
      "options": { "reasoning_effort": "high" }
    },
    "research": {
      "description": "Long-context reading, web research, multi-source synthesis",
      "model": "ollama/kimi-k3:cloud"
    },
    "grunt": {
      "description": "Renames, docstrings, test scaffolding, lint fixes",
      "model": "ollama/glm-5.3-flash"
    }
  }
}

The reasoning behind each line:

  • Planning gets Fable 5.1. Planning is where a bad decision costs you the most downstream tokens, and long-horizon coherence is precisely what 5.1 was built for. It is also where you spend the fewest tokens, so the price premium applies to the smallest slice of your bill.
  • Building gets GLM-5.3. Best capability-per-euro on the market, and the bulk of your token spend lands here.
  • Research gets Kimi K3. BrowseComp 91.2 and a genuine 1M context make it the best available option for reading a lot of things and synthesising them.
  • Mechanical work gets GLM-5.3-Flash. At $0.15/$0.50 it is effectively free, and renaming a symbol across 40 files does not need frontier reasoning.
  • kleines_Modell gets something tiny for the harness’s own housekeeping — title generation, summarisation, internal utility calls.

You are not choosing a winner. You are building a gearbox. And write the model IDs so that swapping one is a one-line change, because in this market you will be making that change again within the quarter.


What You Can Actually Run Locally (The Hardware Reality)

Let me kill an assumption before it costs somebody money.

You cannot run Kimi K3 or GLM-5.3 on a workstation. Not with a 5090. Not with two.

Kimi K3

  • ~1.5 TB of weights in native MXFP4 — and remember, that ist the quantised checkpoint, not a starting point for further compression
  • Moonshot recommends a minimum of 64 accelerators for competitive serving
  • Realistic self-hosting starts at multi-node clusters; reference deployments use GB300 NVL72 racks
  • There is no official Ollama library entry for local Kimi K3 and no consumer GGUF conversion worth pointing you at
  • Reports of it running on clusters of consumer RTX 5090s exist, but “a cluster of 5090s” is not a laptop and the throughput is not comparable

GLM-5.3

  • ~1.5 TB in BF16, roughly 750 GB at FP8
  • Minimum viable single node: 8× H200 (1,128 GB of GPU memory) at FP8, leaving around 375 GB for KV cache
  • BF16 needs two 8×H200 nodes, or a single 8×B300 node (2,304 GB)
  • A single 8×H200 node in BF16 is too small for the weights plus cache

Show Image The frontier open models need a rack. The 24 GB tier is where “runs on my machine” actually lives — and it has got very good.

So what does “local AI” actually mean in September 2026?

It means a different tier of model, and that tier has got genuinely good:

ModellVRAMWhat it is for
Qwen3.8-27B24 GB at Q4Best all-rounder on consumer hardware — 61.7% SWE-bench
gpt-oss:20b16 GBBest small model, adjustable reasoning effort
Gemma 4 E4B~6 GBVision plus tool calling on a modern laptop
Mistral 7B8 GBFastest general-purpose option, 40–60 tok/sec
DeepSeek-R1 7B5 GBChain-of-thought reasoning on a laptop GPU
Llama 4 Scout~55 GB at Q410M context, multimodal — workstation territory

A Qwen3.8-27B scoring 61.7% on SWE-bench, running entirely on a 24 GB consumer GPU with no network connection, is a remarkable thing that would have sounded like science fiction eighteen months ago. It is not GLM-5.3 and it is not pretending to be.

The honest framing: open weights at the frontier buy you sovereignty and price, not local execution. Open weights in the 7B–30B range buy you genuine local execution, at a real but acceptable capability cost, and that is the tier where “runs on my machine, sees no network” is an achievable requirement.

The architecture that works for most of my clients is exactly that split: a small local model for anything touching genuinely sensitive data, and a hosted open-weight frontier model for everything else — with the weights available as insurance rather than as a deployment plan.


Wo jedes Modell tatsächlich versagt

No hype. Here are the honest weaknesses.

Kimi K3 weaknesses

It is expensive for an open model. $3/$15 is 2.1× GLM-5.3’s input rate and 3.4× its output rate for essentially the same intelligence index score. If you are choosing K3 over GLM-5.3, be clear about what you are buying with that premium — usually it is the vision stack or the browsing performance.

The licence has the lower gate. A $20M aggregate-revenue MaaS threshold catches far more organisations than GLM-5.3’s $10B. If inference resale is anywhere in your business model, K3 is the more constrained of the two.

Self-hosting is out of reach for almost everybody. 1.5 TB and 64+ accelerators is a serious infrastructure commitment. The exit right is more theoretical here than with any other model in this comparison.

Broad knowledge work trails. GDPval-AA v2 at 1,687 puts it behind GLM-5.3, both Claude Fables and GPT-5.6 Sol Max. It is a coding and agentic specialist that happens to be enormous.

GLM-5.3 weaknesses

Text only. No image input, no video. If your workflow includes screenshot debugging, design-to-code or document vision, GLM-5.3 simply cannot do it and you need GLM-5.3-Flash, Kimi K3 or Fable 5.1 instead. This is the single most common configuration mistake I expect people to make, because the model IDs look related and are not.

It is verbose, and verbosity is billed. 170M output tokens against a 72M class median across the AA suite. logische_Überlegung defaults to max, and thinking cannot be turned off. Budget for more output tokens than your turn count implies.

The licence is no longer MIT, and the security-review clause is unbounded. For most readers the $10B threshold makes this academic. For anyone near it, “scope and method shall be reasonably determined by Z.AI” is not a clause your legal team will enjoy.

Terminal-Bench 3.0 at 28.3 trails GPT-5.6 Sol’s 34.6. On the specific benchmark closest to terminal-native agentic coding, it is behind the closed competition on the same ruler.

It is new to open weights. Released 28 August. The community has had days, not months. Long-tail deployment bugs have not surfaced yet.

Claude Fable 5.1 weaknesses

The price. Roughly 6.5× GLM-5.3 per completed agentic task, even after the 75% cache-read cut. For most work that gap is not defensible on capability grounds any more.

Three breaking changes. Forced tool use returns 400. Thinking blocks do not travel backwards to older models. Editing conversation history invalidates thinking blocks on accounts created after 31 August 2026. Any of these can break a working integration on upgrade.

Parallel tool calling regressed. Reports of one call per turn where Fable 5 batched several. On a long agent loop that is wall-clock time you are paying for twice.

Zero exit optionality. No weights, no self-hosting, no version pinning beyond what Anthropic offers, no inspection. When it changes, you adapt.

The tokenizer inflates comparisons. ~30% more tokens than older Claude models for the same text, which makes historical cost comparisons misleading in Anthropic’s favour if you are not careful.


The Decision Framework: Six Scenarios

Show Image Six scenarios, six answers. Find the row that sounds like your week.

1. “I want one model. Set it, forget it, keep the bill sane.”

GLM-5.3. Within half a point of Kimi K3 and roughly three points of the best closed model on the neutral index, at $1.40/$4.40. Run it through Ollama on a Pro or Max plan, set Denkaufwand: hoch, and get on with your work. The only thing that should push you off this answer is needing image input.

2. “My work is visual — UI, design-to-code, screenshot debugging, documents.”

Kimi K3 or GLM-5.3-Flash, not GLM-5.3. K3’s MoonViT-V2 handles text, images and video natively and scores 81.6/83.4 on MMMU-Pro. GLM-5.3-Flash is the budget option with native multimodality and an MIT licence. GLM-5.3 is text-only and will simply refuse the input.

3. “Long autonomous sessions where being wrong is expensive.”

Claude Fable 5.1. This is what it was built for and the benchmarks back it: Terminal-Bench-Science doubled, 82% on Browserbase’s hardest computer-use tasks against Opus 5’s 74%, and the largest published gains on multi-hour agentic work. The 75% cache-read cut makes exactly this workload up to 45% cheaper than it was. Pay the premium where a mistake costs more than the tokens.

4. “EU data residency is a hard requirement.”

GLM-5.3-Flash if you need MIT, GLM-5.3 if you need capability. Self-host on your own hardware or an EU GPU provider and the cross-border transfer question stops existing. Budget 8×H200 for GLM-5.3 at FP8; Flash is far more tractable at 320B/18B. If self-hosting is out of budget, Ollama’s Europe hosting with zero data retention is the pragmatic middle ground — but get the residency commitment in writing rather than inferring it from a marketing page.

5. “Small team, tight budget, coding all day.”

GLM-5.3-Flash as default, GLM-5.3 for hard problems, Ollama Pro at $20. Flash costs $0.15/$0.50 — currently half that until 9 September — and scores 57.5 on the intelligence index. Twenty dollars buys sixty dollars of tokens with no weekly caps. For a two-to-four person team this is close to unbeatable. The GLM Coding Plan at $18/month (Lite) is the alternative if you prefer a fixed quota to a credit pool.

GLM-5.3-Flash. It is the only model in this comparison under a standard OSI-approved licence (MIT), it is natively multimodal, it scores 57.5 on the neutral index, and the weights are on Hugging Face with no revenue gates, no attribution mandates and no security-review clause. When the question is “what can we defend in a contract review,” permissive licensing beats three points of benchmark every time.


Was ich im nächsten Quartal im Auge behalten werde

Whether Fable 5.1 lands above 62.1 on the Artificial Analysis index. It launched the day before this article and has not been independently scored. Its benchmark deltas over Fable 5 suggest it should, but “should” is not “did,” and the gap to Kimi K3’s 59.7 is the number the whole open-versus-closed argument turns on.

Whether the licence drift continues. In eight weeks we went from GLM-5.2 under MIT to GLM-5.3 under a bespoke licence with a discretionary security review, and from Kimi K2’s modified-MIT to K3’s revenue-gated terms. If GLM-6 and K4 tighten further, “open weights” becomes a marketing term rather than a meaningful category. GLM-5.3-Flash staying MIT is the counter-signal worth tracking.

Whether other labs adopt Z.ai’s staged-release pattern. A two-week hold with a published safety rationale is a new norm. If it holds, it is a good one. If it becomes a reason weights ship later and later, it is a soft path to not shipping them at all.

Independent replication of the cyber capability claims. 2,436 vulnerabilities across 269 projects is an extraordinary number and it is entirely self-reported. Somebody neutral needs to check it, because if it is accurate it reframes the entire open-weights safety conversation, and if it is not, it was effective marketing.

Ollama’s per-token rates six months from now. $20 for $60 of usage is an aggressive introductory posture from a company that raised $88M in July. Whether those multipliers survive contact with real unit economics is the single biggest variable in the “cheap open models” thesis for small teams.

Whether local models close on the 30B tier. Qwen3.8-27B at 61.7% SWE-bench on 24 GB is the most under-discussed result of the year. The frontier gets the headlines; the 24 GB tier is what changes what an ordinary business can do without an API key.


Häufig gestellte Fragen

Is Kimi K3 really open source? No. Kimi K3’s weights are freely downloadable from Hugging Face, but under a custom “Kimi K3 License,” not an OSI-approved open-source licence. Model-as-a-Service operators whose aggregate revenue exceeds $20 million over any consecutive 12 months must negotiate a separate commercial agreement with Moonshot, and products above 100 million monthly active users or $20 million monthly revenue must display “Kimi K3” in their interface. Purely internal use is unrestricted.

Is GLM-5.3 better than Kimi K3? They are effectively tied on the neutral aggregate index — 59.5 against 59.7 on Artificial Analysis. GLM-5.3 is 2.2× cheaper per agentic task and scores higher on knowledge work (GDPval-AA v2: 1,769 vs 1,687). Kimi K3 is natively multimodal, leads on agentic browsing (BrowseComp 91.2) and took first place on Arena.AI’s frontend code arena. Choose GLM-5.3 for cost and text-based coding; Kimi K3 for vision, browsing and long-horizon research.

How much cheaper are open models than Claude Fable 5.1? On a modelled 60-turn agentic task, GLM-5.3 costs about $1.85 against Fable 5.1’s $11.94 — roughly 6.5× cheaper. Kimi K3 costs about $4.07, roughly 2.9× cheaper. GLM-5.3-Flash costs about $0.21, roughly 58× cheaper. Over 80 runs a month that is $148 versus $955.

Can I run Kimi K3 or GLM-5.3 locally? Not on consumer hardware. Kimi K3 is roughly 1.5 TB even in its native MXFP4 format and Moonshot recommends 64+ accelerators. GLM-5.3 needs a minimum of 8×H200 (1,128 GB) at FP8. For genuine local execution, look at Qwen3.8-27B (24 GB at Q4), gpt-oss:20b (16 GB) or Gemma 4 E4B (~6 GB).

What changed in Claude Fable 5.1? Released 1 September 2026. Cache reads dropped 75% from $1.00 to $0.25 per million tokens, making typical workloads about 25% cheaper and agentic workloads up to 45% cheaper; input and output stayed at $10/$50. Terminal-Bench 4.0 rose from 42.0% to 55.8%, Terminal-Bench-Science from 24.7% to 52.6%, AutomationBench from 17.1% to 31.4%. Three breaking changes affect forced tool use, thinking-block portability and history editing.

What is Ollama’s new pricing? From 31 August 2026, Pro, Max and Team plans use per-token pricing with included credits: Pro $20/month for $60 of usage, Max $100 for $300, Team $500 for $1,000 shared across unlimited users. The 5-hour and weekly caps were removed entirely, there are no service fees, credits do not roll over, and all plans carry zero data retention with hosting in the US and Europe.

Which of these models is genuinely MIT-licensed? Only GLM-5.3-Flash — a separate 320B/18B natively multimodal model released 26 August 2026, scoring 57.5 on the Artificial Analysis index at $0.15/$0.50 per million tokens. GLM-5.3 and Kimi K3 both use bespoke revenue-gated licences; Claude Fable 5.1 is fully proprietary.

Why can’t I compare these models on Terminal-Bench? Because they were each evaluated on a different version. Kimi K3’s 88.3 is on Terminal-Bench 2.1, GLM-5.3’s 28.3 is on 3.0, and Fable 5.1’s 55.8 is on 4.0. Each version is substantially harder than the last, so the numbers are not on the same scale. Compare within a version only. On Terminal-Bench 2.1, for example, Kimi K3 scores 88.3 against Claude Fable 5’s 88.0, GPT-5.6 Sol’s 88.8 and GLM-5.3-Flash’s 84.3 — that comparison is valid, and it is the one worth quoting.

Should I use Ollama or go direct to the vendor? Ollama if you want one billing relationship, one API surface, easy model switching and EU/US hosting with zero data retention — its per-token rates track first-party pricing closely. Direct if you need first-party features like Z.ai’s logische_Überlegung controls at full fidelity, vendor SLAs, or subscription plans such as the GLM Coding Plan. Many teams run both and route by workload.

Does GLM-5.3 support images? No. GLM-5.3 is text-only. GLM-5.3-Flash — a completely different model despite the similar name — is natively multimodal and handles text, image, video and file input. Sending images to the wrong model ID is the most common early mistake with the GLM family.


Das Fazit

Twelve months ago the open-weight question was whether these models were usable. Six months ago it was whether they were competitive. In September 2026 it is genuinely: what are you still paying a closed-model premium for?

There is a real answer to that question, and it is narrower than it used to be. Claude Fable 5.1 leads on the hardest sustained agentic work, on general knowledge work, and on the class of debugging where a model needs to hold a messy problem in its head for hours without drifting. Anthropic’s 75% cache-read cut targets exactly that workload. If your failure cost exceeds your token cost, that premium is rational.

For everything else, the maths has moved. GLM-5.3 delivers 94% of Claude Opus 5’s index score at 29% of the cost. Kimi K3 scores 88.3 on Terminal-Bench 2.1 against Claude Fable 5’s 88.0 on the same version, and took first place on LMArena’s Frontend Code Arena ahead of both flagship closed models. Ollama will sell you $60 of tokens for $20 and remove the usage caps while doing it. That combination did not exist in the spring.

But do not let the enthusiasm skip the fine print, because there are two of them and both matter.

The first is licensing. “Open weights” and “open source” have quietly stopped meaning the same thing. Kimi K3 and GLM-5.3 both ship under bespoke, revenue-gated licences. The gates are high enough that most readers are unaffected — but “most readers are unaffected” is not the same as “unrestricted,” and GLM-5.3’s undefined security-review clause is a genuine procurement risk for large enterprises. GLM-5.3-Flash under MIT is the last fully permissive frontier-adjacent option, which is precisely why it deserves more attention than it gets.

The second is that self-hosting is mostly aspirational. 1.5 TB of weights and 64 accelerators is not an exit plan for a mid-sized business. What open weights buy you at this tier is a competitive inference market, price transparency, version pinning, jurisdiction choice, and a negotiating position. Those are worth a great deal. They are not the same as running the thing in your basement.

The setup I would actually build: GLM-5.3 as the default, Fable 5.1 on planning and the genuinely hard problems, Kimi K3 for vision and long-horizon research, GLM-5.3-Flash for the grunt work, and a 27B local model for anything that must never leave the building. Route by workload, not by loyalty. Write your configuration so the model IDs are a one-line change.

Because the only prediction I am confident about is that this article will need updating before Christmas.


Are you running any of these three in production? I am particularly interested in whether GLM-5.3’s verbosity shows up as a real cost problem at scale, and whether anyone has actually put the Kimi K3 licence in front of counsel and got a clear read on the MaaS definition. Get in touch — corrections and counter-evidence welcome, and this article gets updated when the picture changes.

Last updated: 2 September 2026.

Fügen Sie Ihr Signal hinzu.

Deine E-Mail-Adresse wird nicht veröffentlicht. Erforderliche Felder sind mit * markiert

de_DEDeutsch