Aller directement au contenu
JEUDI, SEPTEMBRE 3, 2026ACTUALITÉS TECHNIQUES INDÉPENDANTES, TESTS ET GUIDES PRATIQUES
À PROPOS
DERNIÈRES NOUVELLES /

Kimi K3, Fable 5.1 et GLM-5.3 : la nouvelle ère de la catégorie « Open-Weight » est lancée

LIRE LE GUIDE
AGENTS D'IAGM / 782

Kimi K3, Fable 5.1 et GLM-5.3 : la nouvelle ère de la catégorie « Open-Weight » est lancée

Dark comparison card for Kimi K3, Claude Fable 5.1 and GLM-5.3 showing Terminal-Bench 2.1 scores of 88.3 against 88.0 for open versus closed models

Il y a six semaines, j'aurais rédigé cet article différemment.

En juillet, Moonshot AI a mis en ligne sur Hugging Face un modèle de pointe comportant 2 800 milliards de paramètres et l'a rendu accessible à tous. Le 28 août, Z.ai a enfin publié les poids du GLM-5.3 qu’elle gardait sous le coude depuis deux semaines — après avoir découvert que son propre modèle pouvait enchaîner des exploits d’une manière que personne ne lui avait demandée. Et le 1er septembre, Anthropic a lancé Claude Fable 5.1 et a rendu le travail des agents jusqu’à 451 TP3T moins cher sans modifier le prix affiché.

Trois modèles. Trois réponses totalement différentes à la question de savoir ce qu’un laboratoire d’IA doit au public. Et pour la première fois depuis le début de toute cette affaire, les options « open-weight » ne constituent pas un choix de compromis.

Voilà en quoi consiste réellement l'histoire de la fin de l'année 2026, et elle mérite bien environ huit mille mots, car c'est dans les détails que se cachent tout l'argent et toutes les erreurs.

J'ai rassemblé les documents techniques, les fiches de spécifications, les textes des licences — en les examinant minutieusement, ligne par ligne, car deux de ces trois éléments ne correspondent pas à ce que les gens appellent communément ainsi —, les benchmarks des fournisseurs, les notes des agrégateurs indépendants, les taux réels par jeton, ainsi que la toute nouvelle grille tarifaire d'Ollama. J’ai ensuite effectué les calculs de coûts sur la base d’un mois de travail normal, plutôt que d’un scénario de jour de lancement.

Voici le résultat.


En bref — Le verdict en 60 secondes

Le GLM-5.3 est actuellement le meilleur rapport qualité-prix en matière d'IA, et de loin. Selon les résultats d'Artificial Analysis, ses performances se situent à moins d'un demi-point de celles de Kimi K3 et à environ trois points et demi derrière celles de Claude Opus 5, avec $1,40 en entrée et $4,40 en sortie par million de tokens. Pour la même tâche d’agent, son coût représente environ un sixième de celui de Fable 5.1. Les poids sont téléchargeables. Si vous cherchez un modèle sur lequel baser votre agent et que vous ne dirigez pas une entreprise générant $10 milliards de chiffre d’affaires, c’est celui-ci qu’il vous faut.

Kimi K3 est le logiciel le plus performant que vous puissiez télécharger légalement. 2,8 billions de paramètres, 104 milliards actifs, un véritable contexte de 1 million, une vision native, le meilleur score de navigation agentique jamais publié (BrowseComp 91,2), et — le chiffre qui devrait mettre fin au débat selon lequel “ les modèles ouverts sont à la traîne ” — 88,3 sur Terminal-Bench 2.1, contre 88,0 pour Claude Fable 5 et 88,8 pour GPT-5.6 Sol sur la même version. Il coûte également $3/$15 — soit plus du double du GLM-5.3 — et ses 1,5 To de poids nécessitent un rack, et non une station de travail. C'est le modèle que l'on utilise lorsque la tâche est difficile et longue, et celui que l'on cite lorsque quelqu'un vous dit que les poids ouverts ont une génération de retard.

Le Claude Fable 5.1 reste le modèle vers lequel on se tourne lorsque la moindre erreur peut coûter cher. Il fait figure de référence en matière de travail agentique soutenu pendant plusieurs heures et de débogage fastidieux qui n’apparaît jamais clairement dans un classement. Le 1er septembre, Anthropic a réduit les lectures de cache de 75%, ce qui rend les longues sessions d’agents nettement moins coûteuses qu’auparavant. Son prix par tâche est également six fois supérieur à celui du GLM-5.3, il s’agit d’une offre fermée, et elle est désormais soumise à des restrictions anti-distillation qui empêchent le fonctionnement de certaines intégrations existantes.

La phrase qui résume tout en un mot : Les modèles « open-weight » ont réduit l'écart de performances à environ trois à cinq mois, et ils ont complètement comblé l'écart de prix. Ce que vous achetez réellement chez Anthropic en septembre 2026, c'est la fiabilité sur les tâches les plus difficiles de type 10% — et le droit de ne pas avoir à vous soucier de tout cela.

Et le rebondissement que personne n'ose avouer à voix haute : Ni Kimi K3 ni GLM-5.3 ne sont désormais distribués sous licence open source. Les deux projets sont passés cet été à des conditions sur mesure, subordonnées à des conditions commerciales. Des poids ouverts, oui. De l'open source, non. Lisez la section sur les licences avant que votre équipe juridique ne s'en charge.

Comparison table of licence terms for Kimi K3, GLM-5.3, GLM-5.3-Flash and Claude Fable 5.1, showing revenue thresholds, attribution rules and which models are genuinely open source
Poids téléchargeables, trois jeux de cordes différents. Seul GLM-5.3-Flash est encore sous licence MIT.

Tableau comparatif rapide

Kimi K3GLM-5.3Claude Fable 5.1
CréateurMoonshot AIZ.ai (Zhipu AI)Anthropique
Publié16 juillet 2026 (API) · 27 juillet (pesée)14 août 2026 (API) · 28 août (pesée)1er septembre 2026
Poids disponiblesOuiOuiNon
LicenceLicence Kimi K3 (personnalisée, soumise à des conditions de revenus)Licence glm-5.3 (personnalisée, liée au chiffre d'affaires)Exclusif
Nombre total de paramètres2,8 T753BNon divulgué
Actif par jeton104B (16 sur 896 experts)environ 40 milliards (estimation)Non divulgué
ArchitectureKDA + résidus d'attention, LatentMoE stable, 69 couches KDA + 24 couches MLA à porteLe modèle MoE, qui repose sur la même base que le GLM-5.2, tire profit uniquement de l'apprentissage postérieur.Non divulgué
Fenêtre de contexte1,048,5761,000,0001,000,000
Puissance maximale131 072 en défaut (jusqu'à 1 048 576)128,000128,000
Saisie multimodaleTexte, image, vidéo (MoonViT-V2)Texte uniquementTexte, image
RéflexionToujours activéToujours actif, effort_de_raisonnement min/max/max.Fonctionnement permanent, adaptatif, effort par message (bêta)
Prix d'entrée / 1M$3.00$1.40$10.00
Données d'entrée mises en cache / 1 Mo$0.30$0.26$0.25
Prix de vente / 1 M$15.00$4.40$50.00
Indice d'intelligence AA59.759.5Pas encore noté (Fable 5 : 62,1)
Coût par tâche d'évaluation AA$0.84$0.68— (Opus 5 : $2.34)
Empreinte de l'hébergement autonomeenviron 1,5 To (MXFP4 natif)~1,5 To BF16 / ~750 Go FP8N/A
Sur Ollamakimi-k3:cloudglm-5.3 : cloudVia Claude Code / Claude Desktop en tant que client

Tous les prix indiqués sont les prix catalogue des fabricants. Les chiffres d'Artificial Analysis proviennent de l'instantané de l'indice de septembre 2026 ; Fable 5.1 a été lancé le 1er septembre et n'avait pas encore fait l'objet d'une évaluation indépendante au moment de la rédaction de cet article.


Pourquoi cette comparaison à trois est celle qui compte vraiment

Pendant deux ans, le débat « ouvert contre fermé » a suivi un schéma bien établi : les modèles fermés étaient meilleurs, les modèles ouverts étaient moins chers, et chacun choisissait sa position en fonction de ce compromis. Tout le monde savait où il en était.

Cette forme s'est cassée cet été.

Nathan Lambert, qui suit cette question de plus près que presque n'importe qui d'autre, estime que l'écart actuel de performances entre les modèles « frontier » en catégorie « open-weight » et ceux de la catégorie « closed » s'élève à trois à cinq mois — contre les six à neuf mois que l’on avançait il y a un an. Au moment où il a rédigé cet article, Kimi K3 occupait la #2 de l’indice Vals AI et figurait parmi les premiers de l’indice d’Artificial Analysis, devancé uniquement par Claude Fable et GPT-5.6 Sol Max. Le relevé de l’indice de septembre que j’utilise plus loin dans cet article le place en quatrième position, derrière Grok 4.6. Quoi qu’il en soit : ce n’est pas “ un bon résultat pour un modèle ouvert ”. C’est une place sur le podium.

Par ailleurs, les données d'utilisation ont révélé une tendance vraiment surprenante. Les modèles d'origine chinoise ont capté environ 61% de l'ensemble des jetons acheminés via OpenRouter d’ici mai 2026. La part des États-Unis dans ce trafic s’est effondrée, passant d’environ 70% à environ 30% en douze mois. Llama de Meta — le modèle à l’origine de la vague des modèles « open-weight » — est passé sous la barre des 1% de volume acheminé. La part de Google est passée d’environ 371 TP3T à 131 TP3T. L’étude menée par OpenRouter en collaboration avec a16z sur 100 000 milliards de jetons a révélé que les modèles « open-weight » représentaient environ un tiers du volume total de jetons sur la plateforme.

On peut débattre de ce que représente le trafic d’OpenRouter. Il surreprésente les développeurs, les amateurs et les charges de travail où le coût est un facteur déterminant, tandis qu’il sous-estime les contrats d’entreprise qui n’utilisent jamais de routeur. Soit. Mais il s’agit du plus grand ensemble de données public dont nous disposons sur ce que les gens choisissent réellement lorsqu’ils ont la liberté de choisir, et la tendance est sans ambiguïté.

Ainsi, en septembre 2026, la question ne sera plus “ les modèles ouverts sont-ils déjà suffisamment performants ? ”. Elle sera bien plus précise :

Parmi trois modèles qui répondent tous aux critères requis, lequel recommandez-vous à votre agent ? Et quel est le coût réel d'une erreur de votre part ?

C'est une question qui porte sur les indices de référence, les prix, les licences et les infrastructures, à peu près dans cet ordre selon l'importance que les gens leur accordent, mais exactement dans l'ordre inverse de leur importance réelle.


Qu'est-ce que la Kimi K3, au juste ?

Moonshot AI a annoncé le lancement de Kimi K3 le 16 juillet 2026 et a publié l’ensemble des poids le 27 juillet. Il s’agit, en nombre de paramètres, du plus grand modèle à poids ouverts jamais publié : 2,8 billions de paramètres au total, soit environ 75% de plus que DeepSeek V4 Pro.

Ce chiffre est à la fois moins impressionnant qu'il n'y paraît et plus impressionnant qu'il n'y paraît, dans cet ordre-là.

L'architecture, en termes simples

Kimi K3 est un modèle de type « mélange d'experts » qui active 104 milliards de paramètres par jeton, en sélectionnant 16 experts parmi un ensemble de 896. Ainsi, bien que le total s'élève à 2,8 T, chaque passage en avant ne touche que moins de 41 TP3T des poids. C'est là que réside l'intérêt économique : un immense capital de connaissances stockées, pour un coût d'inférence modéré.

C'est au niveau de la pile de gestion de l'attention que réside l'intérêt technique. K3 utilise 93 couches constituées de 69 couches KDA (Kimi Delta Attention) et de 24 couches Gated MLA — une architecture hybride combinant attention linéaire et attention complète, avec un rapport d’entrelacement d’environ 3:1. Le KDA est une variante de l’attention linéaire qui s’exécute en temps linéaire par rapport à la longueur de la séquence, plutôt qu’en temps quadratique, ce qui rend le contexte de 1 million de caractères économiquement viable, et non pas simplement théorique. Moonshot fait état d’une capacité pouvant atteindre 75% : réduction de la mémoire cache KV et jusqu'à Débit de décodage multiplié par 6 avec un contexte de 1 Mo en conséquence.

Par-dessus, on trouve Attention aux résidus (AttnRes), que Moonshot décrit comme un substitut direct aux connexions résiduelles standard, et un LatentMoE stable cadre de routage. Cette combinaison permettrait d'atteindre environ un Amélioration de 2,5 fois de l'efficacité globale de mise à l'échelle contre Kimi K2.

La vision vient de MoonViT-V2, un encodeur comportant 401 millions de paramètres, et qui est intégré de manière native plutôt que rajouté après coup — K3 traite le texte, les images et la vidéo au sein d’un seul modèle.

Un autre détail important pour tous ceux qui envisagent l'auto-hébergement : K3 a été entraîné avec Entraînement natif tenant compte de la quantification MXFP4, avec des experts dans MXFP4 et des activations dans MXFP8. Il ne s'agit pas d'une quantification a posteriori. Le point de contrôle publié correspond au modèle quantifié, ce qui explique pourquoi 2,8 billions de paramètres aboutissent à environ 1,5 To d'espace disque plutôt que les quelque 5,6 To qu'exigerait un point de contrôle BF16 de cette taille.

L'histoire de référence

Les chiffres publiés par Moonshot sont excellents dans tous les domaines et, fait inhabituel, plusieurs d'entre eux ont été confirmés par une évaluation indépendante :

  • Terminal-Bench 2.1 : 88,3 — le meilleur score publié pour cette version du test de performance
  • FrontierSWE : 81,2
  • DeepSWE : 67,5
  • BrowseComp : 91,2 — État de l’art de la recherche sur le Web agentique
  • GPQA Diamond : 93,5
  • MMMU-Pro : 81,6 / 83,4 (vision)
  • Marathon SWE : 91,0
  • OfficeQA Pro : 81,3
  • AutomationBench : 30,8
  • LMArena Frontend Code Arena : 1 679 Elo — première place, devant Claude Fable 5 (1 631) et GPT-5.6 Sol (1 618), remportant six des sept domaines front-end

Ce dernier point mérite qu'on s'y attarde. Un modèle à poids ouvert, téléchargeable gratuitement, a remporté la première place lors d’un concours de programmation de front-end axé sur les préférences humaines, devançant les modèles fermés phares des deux laboratoires les mieux financés au monde. Quelle que soit votre opinion sur les benchmarks de ce type en tant que méthodologie, c’est un titre qui aurait été impensable en 2025.

En ce qui concerne les indices globaux, il se situe légèrement en dessous : 59,7 sur l'indice d'intelligence artificielle analytique, troisième au classement général, derrière Claude Opus 5 (63,0) et Claude Fable 5 (62,1). Les résultats de Moonshot — GDPval-AA v2 à 1 687 et AA-Mallette à 1 527 — ce qui les place respectivement en troisième et deuxième position dans ces évaluations du travail intellectuel. Selon l'indice Vals v2, il se classe à la 57.8%, face au 67,21 TP3T de Claude Opus 5, les 66,01 TP3T de Claude Fable 5 et les 63,71 TP3T de GPT-5.6 Sol — l'écart le plus important entre ’ ouvert ’ et « fermé » que j'ai trouvé dans tous les indices, et qu'il convient de garder à l'esprit face aux chiffres relatifs au codage.

Bar chart of the cost of one 60-turn agentic coding task showing Claude Fable 5.1 at 11.94 dollars, Kimi K3 at 4.07, GLM-5.3 at 1.85 and GLM-5.3-Flash at 0.21
La même tâche agentique de 60 tours au prix catalogue de chaque fournisseur. L'axe vertical représente l'argument complet pour les poids ouverts.

Les démos et comment les interpréter

Moonshot a publié deux démonstrations d'autonomie à long terme qui ont suscité beaucoup d'intérêt : une Fonctionnement autonome de 48 heures permettant de mener à bien l'ensemble du processus de conception d'une puce (une conception de 4 mm fonctionnant à 100 MHz, simulant un débit de plus de 8 700 jetons par seconde), et une reproduction de l'astrophysique La relation I-Love-Q dans environ deux heures, contre une à deux semaines dont aurait généralement besoin un chercheur expérimenté.

Ces démonstrations sont certes impressionnantes, mais je ne m’appuierais pas sur elles pour prendre une décision d’achat. Les démonstrations organisées, mises en place et sélectionnées par les fournisseurs vous montrent ce dont un modèle est capable dans les meilleures conditions, sous la supervision d’experts, et non ce qu’il fait dans votre référentiel un mardi quelconque. Considérez-les comme une preuve de principe pour des capacités à long terme, et non comme un cahier des charges.


Afficher l'image Deux architectures publiées et une « boîte noire ». Les modèles ouverts vous expliquent exactement comment ils fonctionnent ; le modèle fermé vous indique le score obtenu.

Qu'est-ce que le GLM-5.3, au juste ?

La version GLM-5.3 est celle dont l'histoire de publication est la plus singulière parmi les trois, et c'est celle qui mérite d'être racontée en détail, car c'est la première fois qu'un grand laboratoire a manifestement hésité à publier des coefficients de pondération pour des raisons autres que commerciales.

Z.ai a annoncé la sortie de GLM-5.3 le 14 août 2026 avec le slogan “ Conçu pour respecter les normes. Prêt pour la cyberdéfense. ”. Ce modèle est un Mélange d'experts à 753 milliards de paramètres, avec environ 40 milliards de tokens actifs selon les estimations de la communauté issues de la configuration GLM-5.2. Il dispose d'une fenêtre de contexte de 1 million de tokens et d'une sortie maximale de 128 000, et — contrairement à GLM-5.3-Flash, qui est un modèle totalement différent — il est texte seul.

“ Nous n'avons fait que développer l'activité après la formation ”

L'aspect technique le plus intéressant concernant le GLM-5.3 est qu'il n'y a pas de nouveau modèle de base. Il réutilise la même base pré-entraînée que le GLM-5.2. Tous les progrès ont été obtenus grâce à un post-entraînement considérablement prolongé, que Z.ai a résumé ainsi : “ Pour la version GLM-5.3, nous nous sommes contentés d'optimiser la mise à l'échelle après l'entraînement. ”

Les écarts ne sont pas négligeables :

RéférenceGLM-5.2GLM-5.3
Terminal-Bench 3.04.628.3
DeepSWE v1.146.266.9
CyberGym77.2%84.5%

Une amélioration multipliée par six sur Terminal-Bench et un bond de vingt points sur DeepSWE, à partir des mêmes poids de base, constituent un résultat véritablement significatif. Cela montre que la frontière dispose encore d’une marge de progression importante en phase de post-entraînement, qui représente de loin la partie la moins coûteuse du coût total de l’entraînement. C’est précisément pour cette raison que les laboratoires chinois, qui fonctionnent, comme le dit Lambert, avec “ des capitaux inférieurs de plusieurs ordres de grandeur ” à ceux des laboratoires américains, parviennent à suivre le rythme.

Les deux semaines pendant lesquelles Z.ai n'a pas effectué de livraisons

Lors de son lancement, Z.ai avait indiqué que les paramètres seraient mis à disposition “ par étapes, après des évaluations de sécurité rigoureuses ”, avec un délai cible d’environ deux semaines. Son dépôt Hugging Face mentionnait la date du 28 août. Cette date est arrivée, s’est écoulée en l’espace de quelques heures, puis les paramètres ont été publiés, accompagnés d’une note technique expliquant ce retard.

La raison tenait aux capacités cybernétiques. Au cours de l'évaluation, Z.ai a indiqué que le GLM-5.3 avait identifié 2 436 vulnérabilités réparties sur 269 projets open source, dont 1 097 ont été classées comme critiques ou présentant un niveau de gravité élevé. Plus précisément, le modèle a démontré que raisonnement sur les chaînes d'exploitation à plusieurs étapes — en enchaînant de manière autonome les différentes étapes d’exploitation — ce que Z.ai a qualifié de comportement “ non entièrement prévu ” et qui “ s’est amplifié à mesure que l’entraînement prenait de l’ampleur ”.”

Je tiens à faire preuve de prudence ici, car on pourrait facilement interpréter cela comme du marketing ou de l’alarmisme, alors qu’il ne s’agit probablement ni de l’un ni de l’autre. Un modèle efficace pour détecter les vulnérabilités est efficace pour détecter les vulnérabilités ; les utilisations défensives et offensives relèvent de la même capacité, orientée dans des directions différentes. Les résultats des tests comparatifs de Z.ai reflètent honnêtement cette réalité : CyberGym 84.5% (la découverte de vulnérabilités, et ses conséquences) face à ExploitBench 54.4% (exploitation, où Claude Fable 5 totalise 78,01 TP3T). Le laboratoire a délibérément optimisé la moitié défensive, et cela se voit.

Ce qui est vraiment remarquable, c'est qu'un laboratoire ait décidé de retarder de deux semaines la sortie d'un produit phare et ait rendu public son raisonnement. Que ce raisonnement vous semble convaincant ou non, ce précédent est inédit.

Où le GLM-5.3 atterrit-il réellement ?

Le tableau comparatif publié par Z.ai ne prétend pas faire l'unanimité, ce qui est rafraîchissant :

RéférenceGLM-5.3Meilleur comparateur
AutomationBench48.2%Conducteurs GLM-5.3
CyberGym84.5%Conducteurs GLM-5.3
GDPval-AA v21,769Conducteurs GLM-5.3
Terminal-Bench 3.028.3%GPT-5.6 Résultat : 34,61 TP3T
DeepSWE v1.166.9%GPT-5.6 Résultat : 72,71 TP3T
ExploitBench54.4%Claude Fable 5 : 78,01 TP3T
Code Bench (50 000 jetons)31.4%

Par ailleurs, selon Artificial Analysis, ce chiffre s'élèverait à 59,5 sur l'indice d'intelligence — statistiquement impossible à distinguer du 59,7 obtenu par Kimi K3, à partir d’un modèle comportant environ un quart des paramètres.

Une mise en garde issue de cette même évaluation dont personne au service marketing ne parle : GLM-5.3 a généré 170 millions de jetons générés par la suite Artificial Analysis, contre une médiane de 72 millions pour les solutions comparables. C'est trop détaillé. effort_de_raisonnement La valeur par défaut est max Et comme on ne peut pas désactiver la réflexion, vous payez pour un long processus de réflexion, que la tâche le mérite ou non. Malgré tout, le coût total de la réalisation de l’évaluation complète s’est élevé à $0,68 par tâche, contre $0,84 pour la Kimi K3 et $2,34 pour la Claude Opus 5 — la complexité ne compense donc pas l'avantage en termes de prix. Elle ne fait que le réduire.


Qu'est-ce que « Claude Fable 5.1 » exactement ?

Anthropic a publié Claude Fable 5.1 le 1er septembre 2026, aux côtés du Claude Mythos 5.1 — le même modèle doté d’une configuration de sécurité différente, disponible uniquement dans le cadre de programmes d’accès sécurisé destinés aux travaux dans les domaines de la cybersécurité et des sciences de la vie.

Ce positionnement est précis et, à mon avis, juste : il s'agit de “ un modèle conçu pour des tâches qui ne s'achèvent pas en une seule instruction. ” Anthropic ne prétend pas avoir fait un bond en avant en matière d'intelligence générale. L'entreprise affirme simplement que les sessions d'agentique prolongées, s'étalant sur plusieurs heures, donnent de meilleurs résultats.

Les chiffres

RéférenceFable 5Fable 5.1Mythos 5.1
Terminal-Bench 4.042.0%55.8%60.9%
Terminal-Bench-Science 0.124.7%52.6%
AutomationBench17.1%31.4%
CursorBench 3.2.070.5%73.4%
OSWorld 2.0 (extrait)72.9%77.9%
Le dernier examen de l'humanité (avec des outils)63.8%65.0%
GDPval-AA v21,853

Pour plus d'informations sur ce même souverain : GPT-5.6 Sol obtient un score de 52,31 TP3T sur Terminal-Bench 4.0, si bien que le 55,81 TP3T de Fable 5.1 prend la tête et que le 60,91 TP3T de Mythos 5.1 creuse l'écart.

Le bond enregistré par Terminal-Bench-Science — passant de 24,71 TP3T à 52,61 TP3T, soit un doublement net — est celui auquel j’accorderais le plus d’importance si votre travail implique le calcul scientifique ou les pipelines de données. Et dans le test de performance le plus exigeant de Browserbase en matière d’utilisation de l’ordinateur, Fable 5.1 a terminé 82% de tâches contre les 74% de Claude Opus 5.

Anthropic indique également environ un Réduction de 60% des faux positifs en matière de cybersécurité — le modèle refusant une tâche de sécurité légitime au motif qu’elle correspondait à un modèle associé à un élément dangereux. Si vous avez déjà vu un assistant de programmation refuser de vous aider à écrire un limiteur de débit, vous comprendrez pourquoi cela a de l’importance.

C'est le changement de tarification qui fait l'actualité

Les taux de référence sont restés inchangés : $10 par million de jetons d'entrée, $50 par million de jetons de sortie. Ce qui a changé, c'est le multiplicateur de lecture du cache, et ce changement est considérable.

Fable 5Fable 5.1
Saisie$10.00$10.00
Résultat$50.00$50.00
Écriture dans le cache (5 min)$12.50$12.50
Écriture dans le cache (1 heure)$20.00$20.00
Lecture du cache$1.00$0.25

Il s'agit d'une réduction de 75%, obtenue en ramenant le multiplicateur de lecture du cache de la valeur standard de 0,1× de l'entrée de base à 0,025×. Anthropic avance un chiffre approximatif de 251 TP3T moins cher pour les charges de travail classiques et jusqu’à 451 TP3T moins cher pour les tâches hautement agentiques — et ce deuxième chiffre est bien réel, car les boucles d’agent consistent en très grande majorité en des lectures de cache. Une session d’agent de 60 itérations relit le même contexte accumulé des dizaines de fois. L’API Batch réduit de moitié les entrées et les sorties, ce qui donne $5/$25.

Il s'agit d'une baisse de prix intelligente et ciblée. Elle rend nettement plus abordable ce pour quoi la Fable 5.1 excelle — les longues sessions — sans pour autant dévaloriser le modèle dans son ensemble.

Trois modifications majeures, une régression

Avant de mettre à niveau une intégration en production, lisez attentivement ces informations.

1. L'utilisation forcée des outils n'existe plus. Demandes utilisant tool_choice : " any " ou choix_d'outil : " outil " renvoie maintenant un Erreur 400. La réflexion est toujours active et ne peut pas être contournée ; forcer l'appel d'un outil permettrait de la sauter. Vous devez utiliser tool_choice : " auto ".

2. Les blocages cognitifs sont unidirectionnels. Fable 5.1 peut lire les ’ thinking blocks » générés par les anciens modèles Claude, mais ces derniers ne peuvent pas lire ceux de Fable 5.1. Seul Mythos 5.1 en est capable. Si votre infrastructure change de modèle en cours de conversation pour réaliser des économies, ce repli vers une solution de secours impose désormais une nouvelle planification.

3. L'historique des modifications rend les « blocs de réflexion » caduques. Pour les comptes créés après le 31 août 2026, la modification des tours précédents, des invites du système ou de la palette d’outils invalide tous les blocs de réflexion suivants — qu’il s’agisse d’erreurs ou d’abandons silencieux. Il s’agit d’une mesure explicite de lutte contre la distillation, qui met fin à un schéma courant dans lequel les frameworks d’agents réécrivent le contexte entre deux tours.

Et voici une régression qu'il est bon de connaître : L'appel parallèle des outils serait plus variable dans la version 5.1 : il arrive parfois qu'un seul appel soit effectué par tour, alors que la version 5 en regroupait plusieurs. Sur une boucle d'agent longue, cela se traduit par une perte de temps réel.

Il y a également un détail concernant le tokeniseur qui peut prêter à confusion lors des comparaisons de coûts : Fable 5.1 utilise le même tokeniseur que Claude Opus 4.7, qui produit environ 30% : plus de jetons que le tokeniseur des anciens modèles Claude pour le même texte. Si vous comparez les coûts par rapport à une référence de 2025, tenez-en compte avant de tirer des conclusions.


Le problème de référence : trois modèles, trois règles différentes

Voici le point sur lequel la plupart des articles comparatifs se trompent, et il vaut mieux le dire sans détours.

Il n'est pas possible d'aligner ces trois modèles sur Terminal-Bench. Regarde :

  • Kimi K3 : 88.3 sur Terminal-Bench 2.1
  • GLM-5.3 : 28.3 sur Terminal-Bench 3.0
  • Fable 5.1 : 55.8 sur Terminal-Bench 4.0

Il s'agit de trois benchmarks différents qui portent le même nom. Chaque version est devenue nettement plus difficile que la précédente — c'est d'ailleurs tout l'intérêt de sortir une nouvelle version. Si l'on interprétait ces trois chiffres comme un classement, on en conclurait que Kimi K3 est trois fois plus performant que GLM-5.3 pour le travail terminal, ce qui est absurde.

Le même problème se pose avec AutomationBench, où les numéros de version ne sont pas indiqués de manière cohérente, ainsi qu’avec la famille SWE-bench, où les variantes « Verified », « Pro » et « Marathon » mesurent des éléments réellement différents. Les laboratoires ne font pas cela dans le but de tromper qui que ce soit : ils utilisent les critères d’évaluation en vigueur au moment de leur formation, et ces critères changent désormais tous les quelques mois. Mais pour un lecteur non averti, l’effet est le même que s’il s’agissait d’une tromperie ; il convient donc de le signaler haut et fort.

Ce que vous pouvez légitimement comparer, au sein d'une même version :

Sur Terminal-Bench 2.1: Kimi K3 (88,3), GPT-5.6 Sol (88,8), Claude Fable 5 (88,0) et GLM-5.3-Flash (84,3). Quatre modèles, une seule échelle de référence. C’est cette comparaison qui importe réellement, et je lui ai consacré un tableau à part entière dans la section suivante.

Sur Terminal-Bench 4.0: Fable 5.1 (55,8) contre GPT-5.6 Sol (52,3). Fable l'emporte.

Par ailleurs, méfiez-vous de quiconque vous présente un seul graphique à barres reprenant les trois principaux modèles évoqués dans cet article sur un axe « Terminal-Bench ». C'est tout simplement impossible à réaliser de manière honnête.


Where They Actually Meet: The Cross-Comparable Numbers

Strip out the benchmark-version mismatches and three comparisons survive.

1. Artificial Analysis Intelligence Index (September 2026)

Independently run, same methodology across every model, no vendor involvement in selection.

RankModèleScorePoids libres
1Claude, opus 563.0Non
2Claude Fable 562.1Non
3Grok 4.660.9Non
4Kimi K359.7Oui
5GLM-5.359.5Oui
6GPT-5.6 Sol58.9Non
7Qwen3.8 Max Preview58.1Weights pending
8GLM-5.3-Flash57.5Yes (MIT)
31MiniMax M345.4Oui

Fable 5.1 had not been independently scored at the time of writing — it launched the day before. Given 5.1’s benchmark deltas over Fable 5, expect it to land at or above 62.1 when it is.

A note on precision: different snapshots of this index round differently — one September pull has Kimi K3 and GLM-5.3 both at 60 and GPT-5.6 Sol at 61. The ordering at the top is stable; the sub-point gaps are not. Do not build an argument on half a point.

Read that table carefully. The best open-weight model is 3.3 points behind the best closed model on the broadest neutral index available. A year ago that gap was double digits.

2. Terminal-Bench 2.1 — the one version where three of them meet

This is the single most useful table in the article, and I nearly missed it. Kimi K3, Claude Fable 5 and GPT-5.6 Sol were all scored on the same version of Terminal-Bench:

ModèleTerminal-Bench 2.1Poids libres
GPT-5.6 Sol88.8Non
Kimi K388.3Oui
Claude Fable 588.0Non
GLM-5.3-Flash84.3Yes (MIT)

Half a point separates a model you can download from the best closed model on the same benchmark, on the same version, on terminal-native agentic work.

That is the number to remember. Not the index aggregate, not the vendor chart — this one. On the specific task shape that agentic coding actually is, the gap is inside the noise.

One caveat, stated by the source and worth repeating: all of Kimi K3’s published figures are max-effort runs, at effort_de_raisonnement maximum and temperature 1.0. Different labs use different evaluation harnesses and different effort settings, so even same-version comparisons carry more uncertainty than a clean table suggests.

3. Cost to run the same evaluation suite

If the table above is the best capability comparison, this is the best value comparison — because it is the only figure that bundles capability, verbosity, reasoning overhead and price into a single number:

ModèleCost per Intelligence Index task
GLM-5.3$0.68
Kimi K3$0.84
Claude, opus 5$2.34

Same tasks, same scoring, actual invoices. GLM-5.3 delivers 94% of Opus 5’s index score for 29% of the cost.

4. GDPval-AA v2 (knowledge work)

ModèleScore
Claude Fable 5.11,853
Claude Fable 5 Max1,815
GLM-5.31,769
GPT-5.6 Sol Max1,747.8
Kimi K31,687

This one is worth caveating: it is assembled from vendor-published charts rather than a single independent run, and the “Max” suffixes indicate different effort settings. But the ordering is broadly consistent across sources, and it says something real — on general knowledge work, as opposed to coding, the closed models still hold a clearer lead than the coding benchmarks suggest.

That is the pattern across all three comparisons. The open models have essentially caught up on coding and agentic tasks. They are still a step behind on broad knowledge work. If your workload is the former, the case for paying closed-model prices is weak. If it is the latter, it is still defensible.


Afficher l'image Three and a bit points separate the best closed model from the best downloadable one. That gap was double digits twelve months ago.

The Cost Maths, With Real Numbers

Benchmarks are abstract. Invoices are not.

Le scénario

One medium feature, done agentically: roughly 60 cycles d'appel d'outils, soit en moyenne 45,000 input tokens per turn à mesure que le contexte s'étoffe, et 2,000 output tokens per turn.

  • Total des données saisies : 2,7 millions de jetons
  • Production totale : 120 000 jetons
  • Taux supposé de réussite des accès au cache des invites : 80% (2,16 millions en cache, 540 000 récents)

Claude Fable 5.1

ComposantJetonsTauxCoût
Nouvelles informations540 000$10.00/M$5.40
Données mises en cache2,16 millions$0.25/M$0.54
Résultat120 000$50.00/M$6.00
Total≈ $11.94

Kimi K3

ComposantJetonsTauxCoût
Nouvelles informations540 000$3.00/M$1.62
Données mises en cache2,16 millions$0,30/M$0.65
Résultat120 000$15.00/M$1.80
Total≈ $4.07

GLM-5.3

ComposantJetonsTauxCoût
Nouvelles informations540 000$1.40/M$0.76
Données mises en cache2,16 millions$0.26/M$0.56
Résultat120 000$4.40/M$0.53
Total≈ $1.85

And for reference — GLM-5.3-Flash

ComposantJetonsTauxCoût
Nouvelles informations540 000$0,15/M$0.08
Données mises en cache2,16 millions$0.03/M$0.06
Résultat120 000$0,50/M$0.06
Total≈ $0,21

The ratios

  • GLM-5.3 is 6.5× cheaper than Fable 5.1 for identical work
  • Kimi K3 is 2.9× cheaper than Fable 5.1
  • GLM-5.3 is 2.2× cheaper than Kimi K3
  • GLM-5.3-Flash is 58× cheaper than Fable 5.1, at 57.5 on the intelligence index against Fable 5’s 62.1

Scaling to a working month

Quatre tâches de ce type par jour, vingt jours ouvrés — soit 80 exécutions :

ModèleMonthly token cost
Claude Fable 5.1≈ $955
Kimi K3≈ $326
GLM-5.3≈ $148
GLM-5.3-Flash≈ $17

Three honest caveats on these figures

Cache writes are excluded. Fable 5.1 charges $12.50/M for 5-minute cache writes, and a long session writes cache repeatedly. Including writes moves Fable’s real number up, not down. The 80% hit rate is also optimistic for short sessions.

Reasoning tokens are billed output tokens. All three models have always-on thinking. GLM-5.3 in particular is documented as verbose — 170M output tokens against a 72M class median across the AA suite — so its real-world output volume runs above what a naive turn count suggests. The $0.68-per-task figure already accounts for this, which is why I trust it more than my own model above.

Fable 5.1’s tokenizer produces ~30% more tokens than older Claude models for the same text. If you are comparing against a historical Claude bill, adjust.

Even after all three corrections, the ordering does not change and the magnitude barely does. Closed-frontier work costs roughly six times what the best open-weight model costs, per completed task.


Afficher l'image La même tâche agentique de 60 tours au prix catalogue de chaque fournisseur. L'axe vertical représente l'argument complet pour les poids ouverts.

Licences: What “Open” Actually Buys You In 2026

This is the section I would most like people to read, because the vocabulary has drifted badly and it is going to cost somebody a lot of money.

Neither Kimi K3 nor GLM-5.3 is open source. Both publish downloadable weights under bespoke licences with revenue-triggered conditions. That is a meaningfully different thing from MIT or Apache 2.0, and the difference lands squarely on the kind of business that would most benefit from self-hosting.

The Kimi K3 License

You will see “Modified MIT” repeated all over the internet for K3. It is wrong — that described K2. K3 ships under a custom document with two distinct gates:

The Model-as-a-Service gate. If you provide third parties with access to model inference or fine-tuning — where those third parties control inputs, parameters or training data — and “the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars” over any consecutive 12 months, you must negotiate a separate commercial agreement with Moonshot. Note aggregate revenue, across all affiliates, not revenue attributable to K3. A €25M-turnover consultancy that resells inference is over the line even if the AI business is a rounding error.

The attribution gate. Products exceeding 100 million monthly active users or $20 million monthly revenue must display “Kimi K3” prominently in the product interface.

The exemption that matters most: purely internal use is unrestricted. Running K3 to support your own developers, researchers, legal team or general employee productivity does not trigger either gate. For the overwhelming majority of businesses — including essentially all of mine — that is the relevant clause, and the answer is that you are fine.

The clause that is new relative to K2 is the $20M MaaS gate. K2 only required attribution above its thresholds. K3 adds a much lower revenue gate aimed specifically at commercial inference resellers. Moonshot is not trying to stop you using the model; it is trying to stop cloud providers building a business on it for free.

The glm-5.3 License

Z.ai’s flagship went a different direction, and it is a bigger break with precedent than most coverage acknowledged.

GLM-5.2 shipped under MIT. GLM-5.3-Flash shipped under MIT. GLM-5.3 did not.

The custom licence permits individuals and ordinary businesses to run, deploy and fine-tune the model with no additional restrictions. The single gate is aimed at hyperscalers: an entity whose “aggregate revenue of the Licensee and its affiliates exceeds 10 billion US dollars (or the equivalent in other currencies) in total over any consecutive 12 months” doit pass Z.AI’s security review before using the model or its derivatives for any commercial purpose.

That is a $10 billion threshold. It excludes almost everyone. But two details in the fine print deserve attention:

First, the licence explicitly carves out app developers embedding the model in features, and resellers merely routing requests to other hosts. The target is narrow and clearly labelled: companies that want to host GLM-5.3 commercially at scale.

Second — and this is the genuinely awkward part — the licence publishes no criteria, timeline or appeal process for that security review. It states only that its “scope and method… shall be reasonably determined by Z.AI.” If you are a large enterprise, that is an unbounded dependency on a foreign vendor’s discretion, and your procurement team will say so.

Also worth noting: despite the two-week cyber-capability hold that preceded release, the licence contains no acceptable-use restrictions, no cyber carve-outs and no output-ownership claims. The safety concern shaped the release timing, not the release terms.

So what is actually still MIT?

GLM-5.3-Flash. 320B parameters, 18B active, natively multimodal, MIT-licensed, 57.5 on the intelligence index, and $0.15/$0.50 per million tokens — currently half that under a promotional discount running to 9 September 2026. If your requirement is genuine unrestricted permissive licensing rather than merely downloadable weights, that is where the frontier currently sits, and it is remarkably high.

The practical summary

Kimi K3GLM-5.3GLM-5.3-FlashFable 5.1
Weights downloadableOuiOuiOuiNon
OSI-style open sourceNonNonYes (MIT)Non
Restriction trigger$20M revenue (MaaS)$10B revenue (MaaS)AucunN/A
Attribution requiredAbove 100M MAU / $20M-monthCopyright notice onlyCopyright notice onlyN/A
Internal use restrictedNonNonNonN/A
Fine-tuning permittedOuiOuiOuiNon
Acceptable-use policyStandardNone in licence textAucunAnthropic AUP applies

For a typical German Mittelstand business, an agency, or a consultancy running models internally or embedding them in a product: all three open options are usable without negotiating anything. The gates exist to catch AWS, not you. But “read the licence” has gone from pedantic advice to actual advice, and if you are above about €20M turnover and reselling inference in any form, get it in front of counsel before you deploy.


The Real Case For Open Models

I want to make this argument properly rather than as a slogan, because “open models are great” is not an argument and the actual reasons are more interesting than the cheerleading.

1. The price collapse is not a subsidy, it is structural

The instinctive objection to cheap Chinese models is that the pricing is loss-leading and will normalise upward. Some of it certainly is. But the structural point stands independently: when weights are public, inference becomes a commodity market. Twenty hosting providers compete to serve the same model, and none of them can charge a capability premium because none of them has a capability moat. That is why GLM-5.3 runs at $1.40/$4.40 and Fable 5.1 runs at $10/$50 for scores three points apart.

Closed-model pricing includes the R&D amortisation, the margin, and the fact that there is exactly one seller. Open-weight pricing includes electricity, hardware amortisation and a thin margin. Those are different businesses, and the gap will not close by open models getting more expensive.

The evidence that this is real rather than temporary: DeepSeek captured roughly 17% of token usage on Vercel by May 2026 while holding about 1% of revenue share. That is not a pricing anomaly. That is what commoditisation looks like on a chart.

2. Exit rights change your negotiating position even if you never exercise them

Most teams reading this will not self-host. The hardware section below explains why — Kimi K3 needs a rack. But the option has value regardless.

If Anthropic raises prices, deprecates the model you built on, changes its acceptable-use policy in a way that breaks your product, or has a bad quarter, your recourse with a closed model is a migration project. With an open-weight model your recourse is downloading a file you could have downloaded any time. You may never do it. The fact that you could is what makes the vendor relationship a commercial one rather than a dependency.

This is not theoretical. Fable 5.1 shipped with three breaking changes on 1 September, one of which invalidates thinking blocks when you edit conversation history for accounts created after 31 August. If that pattern is load-bearing in your architecture, you are re-engineering on Anthropic’s schedule, not yours.

3. Data residency stops being a project

For anyone operating under GDPR, this is the argument that actually closes deals.

With a closed API, every prompt leaves your infrastructure. You need a processor agreement, a transfer impact assessment if the processor is outside the EEA, a data-flow map, and an answer for your DPO about what happens to the data at rest. All of that is doable. It is also weeks of work per vendor, repeated whenever the vendor changes anything.

With open weights on infrastructure you control — your own hardware, or a German or EU GPU host — the cross-border transfer question does not arise, because there is no transfer. The model runs where your data already lives. Legal review goes from a transfer assessment to a licence read.

For German and EU clients this is frequently the deciding factor, and it is worth being precise about the caveat: using GLM-5.3 through Z.ai’s API does not give you this. You get it from running the weights yourself, or from a provider hosting them in your jurisdiction. The licence makes that legal; it does not make it automatic.

4. Inspectability is real, if underused

Open weights mean you can examine the architecture, run your own evaluations on the actual model rather than an endpoint that might be silently updated, quantise it to fit your hardware, fine-tune it on your domain, and — importantly — pin a version forever.

That last one is underrated. Closed APIs get updated behind stable model IDs. Behaviour drifts. Prompts that worked stop working, and you cannot diff the change because you cannot see it. A local checkpoint is byte-identical in a year’s time.

5. The competitive pressure benefits everyone, including closed-model users

Look at what happened this summer. Kimi K3 lands in July at $3/$15 with frontier-adjacent scores. GLM-5.3 lands in August at $1.40/$4.40 within half a point of it. And on 1 September, Anthropic cut cache reads by 75%.

I am not claiming direct causation — Anthropic does not publish its pricing rationale and cache-read cuts are a natural optimisation. But a market with credible cheap substitutes prices differently from one without them, and the substitutes got credible this year. If you use closed models exclusively, the open-weight ecosystem is still quietly making your bills smaller.

6. Post-training is where the gains are, and it is cheap

GLM-5.3’s headline result — a six-fold Terminal-Bench improvement from the same base weights — is the most strategically important number in this entire article, and it is not about GLM.

It says the expensive part of building a frontier model (pretraining) is increasingly a solved, commoditised input, and the differentiating part (post-training, RL environments, long-horizon task design) is comparatively affordable. That is why four Chinese labs with a combined valuation of about $159 billion are keeping pace with American labs valued at multiples of that. Lambert’s assessment is that Chinese labs are running with “orders of magnitude less capital.”

For anyone downstream, the implication is straightforward: the number of credible model vendors is going up, not down. Plan your architecture accordingly.

7. The ecosystem effects compound

Qwen passed one billion cumulative downloads on Hugging Face, overtaking Llama, and now anchors over 200,000 tagged models and 113,000+ derivatives — roughly 40% of all new LLM derivatives on the platform are Qwen-based. That is a tooling, quantisation, fine-tuning and deployment ecosystem that exists only because the weights are public.

You benefit from that ecosystem whether or not you contribute to it. GGUF conversions, vLLM kernels, LoRA adapters, quantisation recipes, evaluation harnesses — all of it exists because thousands of people could get their hands on the actual weights.

8. It is where the developers already went

Chinese-origin models at ~61% of OpenRouter tokens. US model share down from ~70% to ~30% in a year. Four of the five most-used models on the router are Chinese-origin. Xiaomi’s MiMo models alone account for roughly 21% of routed tokens and about 22% of all coding traffic.

Developer behaviour is a leading indicator. It was a leading indicator for Docker, for Postgres, for Linux. Betting against it has a poor historical record.

Horizontal bar chart of Terminal-Bench 2.1 scores showing GPT-5.6 Sol at 88.8, Kimi K3 at 88.3, Claude Fable 5 at 88.0 and GLM-5.3-Flash at 84.3, with open-weight models in teal
Four models, one version of the benchmark. Half a point separates the best model you can download from the best one you cannot.

…And The Honest Case Against

If I only wrote the section above, this would be an advert.

Open weights are not open source, and the vocabulary drift is a real problem. Two of the three models here ship under bespoke revenue-gated licences with no OSI approval. GLM-5.3’s security-review clause has no published criteria or appeal process. If your compliance framework requires OSI-approved licensing, the honest answer is that your options are GLM-5.3-Flash, DeepSeek’s MIT-licensed flagships, and a shrinking list of others.

Vendor benchmarks remain vendor benchmarks. Every headline number Moonshot and Z.ai published is self-reported and self-selected. Where neutral aggregators have measured, the numbers broadly hold up — which is genuinely to both labs’ credit — but “broadly hold up” is doing work in that sentence.

Self-hosting is a fantasy for most teams. Kimi K3 is 1.5 TB and Moonshot recommends 64+ accelerators. GLM-5.3 needs 8×H200 at FP8 as a bare minimum. The exit right is real; the exercise cost is a data centre.

Support is what you make it. When Fable 5.1 breaks, there is a company with an SLA. When your self-hosted GLM-5.3 deployment produces garbage at 300K context on a Friday night, there is a GitHub issue and your own competence.

Geopolitics is a real procurement input. All three open-weight options discussed here come from Chinese labs. For some clients — public sector, defence-adjacent, certain regulated industries — that is a hard blocker regardless of licence terms or where the weights run. It is not my job to tell you whether that concern is well-founded. It is my job to tell you it will come up in the meeting.

The closed models still lead on general knowledge work. GDPval-AA v2 has Fable 5.1 at 1,853 against GLM-5.3’s 1,769 and Kimi K3’s 1,687. On coding the gap is gone. On broad professional knowledge work it is not.


Afficher l'image Poids téléchargeables, trois jeux de cordes différents. Seul GLM-5.3-Flash est encore sous licence MIT.

Ollama In 2026: The Pricing Change That Actually Matters

Ollama has quietly become the most important piece of infrastructure in this conversation, and on 31 August 2026 it changed how it charges in a way that is worth understanding properly.

What changed

Ollama moved its Pro, Max and Team plans from GPU-time billing to industry-standard per-token pricing, with a pool of usage credits included in every plan. The company’s stated reason is refreshingly concrete: models like Kimi K3 made GPU-time metrics impossible to predict. When one model activates 104B parameters and another activates 3B, “an hour of GPU” stops meaning anything to the person paying.

The new plans

PlanPrixIncluded monthly usageConcurrency
Gratuit$0Small monthly credit, starter models1 request
Pro$20/mo (or $200/yr)$60 of usage3 requests
Max$100/mo$300 of usage10 requests
Team$500/mo$1,000 shared, unlimited users10 requests
EntrepriseCustomCustomCustom

Read the Pro row again. $20 buys $60 of tokens. That is a 3× multiplier on included usage, and when the pool runs out you continue at exactly the same published per-token rate — no penalty tier, no service fee.

What was removed

This is the part developers actually noticed: no 5-hour resets and no weekly caps. Anyone who has hit a rolling usage window mid-refactor will understand why that mattered more than the price. The monthly pool refreshes on your subscription date and, importantly, does not roll over — so size your plan to your normal month, not your busiest one.

The terms that matter for EU work

Three commitments in the announcement are directly relevant if you are handling client data:

  • Zero data retention. Prompts and responses are never logged and never trained on.
  • Hosting in the US and Europe, plus Singapore for a limited set of Qwen models.
  • Per-request cost visibility in your account — you can see exactly what each call cost.

For a German consultancy that is a materially better compliance story than most gateway providers offer, though “hosted in Europe” is a routing statement rather than a contractual data-residency guarantee. If residency is a hard requirement rather than a preference, ask Ollama for it in writing before you assume it.

The actual per-token rates

This is where it gets interesting. Ollama publishes rates per model, and they track first-party pricing closely:

Model on OllamaEntrée / 1MCached / 1MSortie / 1M
kimi-k3$3.00$0.30$15.00
glm-5.3$1.40$0.26$4.40
mistral-large-3$0.50$0.50$1.50
gemma4$0.14$0.05$0.40
nemotron-3-super$0.015$0.015$0.60

Run the monthly maths from earlier through this. Eighty agentic runs a month on GLM-5.3 costs about $148 in tokens. A Max plan at $100/month covers $300 of usage — so the same workload that would cost roughly $955/month on Fable 5.1 fits comfortably inside a $100 Ollama subscription with headroom to spare.

That is the entire open-model economic argument compressed into one line on an invoice.

Afficher l'image Eighty agentic runs a month on GLM-5.3 fit inside a $100 Ollama Max plan with change. The same work on Fable 5.1 is a $955 invoice.

What else Ollama shipped in 2026

The pricing change did not happen in isolation. The last few months have been busy:

  • 25 August — Claude Desktop support. Ollama now works as a third-party gateway provider for Claude Desktop, so you can drive open models through Anthropic’s own client.
  • 26 August — v0.33.0. The Claude Desktop gateway integration landed, prefill recovery was restored to the cache, and Ollama disabled Claude Code’s token-countdown system message, which was invalidating the KV cache on every turn. That last one is a quiet but significant performance fix.
  • 26 August — v0.33.1. MLX support for Qwen3.8 Flash Next, structured output, and a fix for GPU timeouts when loading models from slower storage.
  • 28 August — v0.33.2. Dark mode restored, macOS handoff fixed, and Claude Desktop proxy requests kept alive during model catalogue updates.
  • 20 August — v0.32.15. New desktop onboarding, and metadata caching between requests that cut time-to-first-token by roughly half.
  • 11 August — NVIDIA Nemotron 3.5 Lightning, a 30B model tuned for agentic workflows with tool calling on personal hardware.
  • 10 August — Meta’s Muse Glimmer, a 30B multimodal model under Apache 2.0, accelerated by Ollama’s MLX engine. Worth noting given Llama’s collapse in routed usage — Meta is still shipping genuinely permissive weights.
  • 9 July — $88M funding round, with Ollama reporting 8.9 million developers.
  • 29 June — Gemma 4 on MLX up to 90% faster via multi-token prediction, which disproportionately benefits coding agents on Apple Silicon.
  • 5 June — v0.30 added GGUF compatibility through llama.cpp, broadening hardware support well beyond Apple Silicon.

The through-line is that Ollama has stopped being “the easy way to run a small model on your laptop” and become a general-purpose gateway that happens to also run models locally. The cloud catalogue now includes kimi-k3:cloud, glm-5.3 : cloud, glm-5.3-flash, deepseek-v4-pro, deepseek-v4-flash, minimax-m3, qwen3.5 across seven sizes, gpt-oss at 20b and 120b, gemma4, the Nemotron 3 family and mistral-large-3.


Setting It All Up: Copy-Paste Configs

Ollama + Claude Code (the fastest path)

Ollama ships a one-command launcher that handles the environment wiring for you:

bash

ollama launch claude

If you would rather do it manually — which you should if you are scripting it — install Claude Code first:

bash

# macOS / Linux
curl -fsSL https://claude.ai/install.sh | bash

# Windows (PowerShell)
irm https://claude.ai/install.ps1 | iex

Then point it at Ollama:

bash

export ANTHROPIC_AUTH_TOKEN=ollama
export ANTHROPIC_API_KEY=""
export ANTHROPIC_BASE_URL=http://localhost:11434

And run against whichever model you want:

bash

# A local model
claude --model qwen3.5

# A cloud model — note the :cloud suffix
claude --model kimi-k3:cloud

Or inline, without exporting anything globally:

bash

ANTHROPIC_AUTH_TOKEN=ollama \
ANTHROPIC_API_KEY="" \
ANTHROPIC_BASE_URL=http://localhost:11434 \
claude --model glm-5.3:cloud

Two things that will bite you.

Tout d'abord, do not export ANTHROPIC_BASE_URL permanently in your shell profile. It is a global variable and it will silently redirect every Anthropic-speaking tool on your machine to Ollama, including the one you wanted talking to Anthropic. Scope it per-project or per-invocation.

Deuxièmement, set your context length to 64k or higher for anything working on a real repository. Ollama’s default is smaller, and a coding agent that quietly runs out of window will just start forgetting things rather than erroring.

For CI, Docker or scripted runs, --yes skips the interactive prompts:

bash

ollama launch claude --model glm-5.3:cloud --yes -- -p "how does this repository work?"

GLM-5.3 direct from Z.ai

If you want first-party routing rather than going through Ollama, Z.ai exposes three protocol endpoints:

ProtocoleURL de base
Messages anthropiqueshttps://api.z.ai/api/anthropic
Compléments de conversation OpenAIhttps://api.z.ai/api/coding/paas/v4
Réponses d'OpenAIhttps://api.z.ai/api/v1

An OpenCode provider block for GLM-5.3:

json

{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "zai": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "Z.AI",
      "options": {
        "baseURL": "https://api.z.ai/api/coding/paas/v4",
        "apiKey": "{env:ZAI_API_KEY}"
      },
      "models": {
        "glm-5.3": {
          "name": "GLM-5.3",
          "limit": { "context": 1000000, "output": 128000 },
          "options": { "reasoning_effort": "high", "temperature": 1 }
        }
      }
    }
  },
  "model": "zai/glm-5.3"
}

Important upgrade note: GLM-5.3 requires reasoning to be enabled. If your existing configuration sets thinking.type: "disabled", that will now fail. Change it to "enabled" and set reasoning_effort: "low" if you want the old latency profile back.

effort_de_raisonnement is your main quality-versus-cost dial: faible for latency-sensitive calls, élevé for substantial work, max (the default) for problems that have genuinely defeated élevé. Given the documented verbosity, I would run élevé as an everyday setting rather than leaving max on and wondering where the output tokens went.

The routing setup I would actually recommend

None of this is a “pick one” decision. The correct answer for most teams is a three-tier routing table, and every serious agent harness supports per-agent models.

json

{
  "$schema": "https://opencode.ai/config.json",
  "model": "zai/glm-5.3",
  "small_model": "ollama/gemma4",
  "agent": {
    "plan": {
      "description": "Architecture, root-cause analysis, anything expensive to get wrong",
      "model": "anthropic/claude-fable-5-1"
    },
    "build": {
      "description": "The bulk of implementation work",
      "model": "zai/glm-5.3",
      "options": { "reasoning_effort": "high" }
    },
    "research": {
      "description": "Long-context reading, web research, multi-source synthesis",
      "model": "ollama/kimi-k3:cloud"
    },
    "grunt": {
      "description": "Renames, docstrings, test scaffolding, lint fixes",
      "model": "ollama/glm-5.3-flash"
    }
  }
}

The reasoning behind each line:

  • Planning gets Fable 5.1. Planning is where a bad decision costs you the most downstream tokens, and long-horizon coherence is precisely what 5.1 was built for. It is also where you spend the fewest tokens, so the price premium applies to the smallest slice of your bill.
  • Building gets GLM-5.3. Best capability-per-euro on the market, and the bulk of your token spend lands here.
  • Research gets Kimi K3. BrowseComp 91.2 and a genuine 1M context make it the best available option for reading a lot of things and synthesising them.
  • Mechanical work gets GLM-5.3-Flash. At $0.15/$0.50 it is effectively free, and renaming a symbol across 40 files does not need frontier reasoning.
  • petit_modèle gets something tiny for the harness’s own housekeeping — title generation, summarisation, internal utility calls.

You are not choosing a winner. You are building a gearbox. And write the model IDs so that swapping one is a one-line change, because in this market you will be making that change again within the quarter.


What You Can Actually Run Locally (The Hardware Reality)

Let me kill an assumption before it costs somebody money.

You cannot run Kimi K3 or GLM-5.3 on a workstation. Not with a 5090. Not with two.

Kimi K3

  • ~1.5 TB of weights in native MXFP4 — and remember, that est the quantised checkpoint, not a starting point for further compression
  • Moonshot recommends a minimum of 64 accelerators for competitive serving
  • Realistic self-hosting starts at multi-node clusters; reference deployments use GB300 NVL72 racks
  • There is no official Ollama library entry for local Kimi K3 and no consumer GGUF conversion worth pointing you at
  • Reports of it running on clusters of consumer RTX 5090s exist, but “a cluster of 5090s” is not a laptop and the throughput is not comparable

GLM-5.3

  • ~1.5 TB in BF16, roughly 750 GB at FP8
  • Minimum viable single node: 8× H200 (1,128 GB of GPU memory) at FP8, leaving around 375 GB for KV cache
  • BF16 needs two 8×H200 nodes, or a single 8×B300 node (2,304 GB)
  • A single 8×H200 node in BF16 is too small for the weights plus cache

Afficher l'image The frontier open models need a rack. The 24 GB tier is where “runs on my machine” actually lives — and it has got very good.

So what does “local AI” actually mean in September 2026?

It means a different tier of model, and that tier has got genuinely good:

ModèleVRAMWhat it is for
Qwen3.8-27B24 GB at Q4Best all-rounder on consumer hardware — 61.7% SWE-bench
gpt-oss:20b16 GBBest small model, adjustable reasoning effort
Gemma 4 E4B~6 GBVision plus tool calling on a modern laptop
Mistral 7B8 GBFastest general-purpose option, 40–60 tok/sec
DeepSeek-R1 7B5 GBChain-of-thought reasoning on a laptop GPU
Llama 4 Scout~55 GB at Q410M context, multimodal — workstation territory

A Qwen3.8-27B scoring 61.7% on SWE-bench, running entirely on a 24 GB consumer GPU with no network connection, is a remarkable thing that would have sounded like science fiction eighteen months ago. It is not GLM-5.3 and it is not pretending to be.

The honest framing: open weights at the frontier buy you sovereignty and price, not local execution. Open weights in the 7B–30B range buy you genuine local execution, at a real but acceptable capability cost, and that is the tier where “runs on my machine, sees no network” is an achievable requirement.

The architecture that works for most of my clients is exactly that split: a small local model for anything touching genuinely sensitive data, and a hosted open-weight frontier model for everything else — with the weights available as insurance rather than as a deployment plan.


Où chaque modèle présente réellement des failles

No hype. Here are the honest weaknesses.

Kimi K3 weaknesses

It is expensive for an open model. $3/$15 is 2.1× GLM-5.3’s input rate and 3.4× its output rate for essentially the same intelligence index score. If you are choosing K3 over GLM-5.3, be clear about what you are buying with that premium — usually it is the vision stack or the browsing performance.

The licence has the lower gate. A $20M aggregate-revenue MaaS threshold catches far more organisations than GLM-5.3’s $10B. If inference resale is anywhere in your business model, K3 is the more constrained of the two.

Self-hosting is out of reach for almost everybody. 1.5 TB and 64+ accelerators is a serious infrastructure commitment. The exit right is more theoretical here than with any other model in this comparison.

Broad knowledge work trails. GDPval-AA v2 at 1,687 puts it behind GLM-5.3, both Claude Fables and GPT-5.6 Sol Max. It is a coding and agentic specialist that happens to be enormous.

GLM-5.3 weaknesses

Text only. No image input, no video. If your workflow includes screenshot debugging, design-to-code or document vision, GLM-5.3 simply cannot do it and you need GLM-5.3-Flash, Kimi K3 or Fable 5.1 instead. This is the single most common configuration mistake I expect people to make, because the model IDs look related and are not.

It is verbose, and verbosity is billed. 170M output tokens against a 72M class median across the AA suite. effort_de_raisonnement La valeur par défaut est max, and thinking cannot be turned off. Budget for more output tokens than your turn count implies.

The licence is no longer MIT, and the security-review clause is unbounded. For most readers the $10B threshold makes this academic. For anyone near it, “scope and method shall be reasonably determined by Z.AI” is not a clause your legal team will enjoy.

Terminal-Bench 3.0 at 28.3 trails GPT-5.6 Sol’s 34.6. On the specific benchmark closest to terminal-native agentic coding, it is behind the closed competition on the same ruler.

It is new to open weights. Released 28 August. The community has had days, not months. Long-tail deployment bugs have not surfaced yet.

Claude Fable 5.1 weaknesses

The price. Roughly 6.5× GLM-5.3 per completed agentic task, even after the 75% cache-read cut. For most work that gap is not defensible on capability grounds any more.

Three breaking changes. Forced tool use returns 400. Thinking blocks do not travel backwards to older models. Editing conversation history invalidates thinking blocks on accounts created after 31 August 2026. Any of these can break a working integration on upgrade.

Parallel tool calling regressed. Reports of one call per turn where Fable 5 batched several. On a long agent loop that is wall-clock time you are paying for twice.

Zero exit optionality. No weights, no self-hosting, no version pinning beyond what Anthropic offers, no inspection. When it changes, you adapt.

The tokenizer inflates comparisons. ~30% more tokens than older Claude models for the same text, which makes historical cost comparisons misleading in Anthropic’s favour if you are not careful.


The Decision Framework: Six Scenarios

Afficher l'image Six scenarios, six answers. Find the row that sounds like your week.

1. “I want one model. Set it, forget it, keep the bill sane.”

GLM-5.3. Within half a point of Kimi K3 and roughly three points of the best closed model on the neutral index, at $1.40/$4.40. Run it through Ollama on a Pro or Max plan, set effort_de_raisonnement : élevé, and get on with your work. The only thing that should push you off this answer is needing image input.

2. “My work is visual — UI, design-to-code, screenshot debugging, documents.”

Kimi K3 or GLM-5.3-Flash, not GLM-5.3. K3’s MoonViT-V2 handles text, images and video natively and scores 81.6/83.4 on MMMU-Pro. GLM-5.3-Flash is the budget option with native multimodality and an MIT licence. GLM-5.3 is text-only and will simply refuse the input.

3. “Long autonomous sessions where being wrong is expensive.”

Claude Fable 5.1. This is what it was built for and the benchmarks back it: Terminal-Bench-Science doubled, 82% on Browserbase’s hardest computer-use tasks against Opus 5’s 74%, and the largest published gains on multi-hour agentic work. The 75% cache-read cut makes exactly this workload up to 45% cheaper than it was. Pay the premium where a mistake costs more than the tokens.

4. “EU data residency is a hard requirement.”

GLM-5.3-Flash if you need MIT, GLM-5.3 if you need capability. Self-host on your own hardware or an EU GPU provider and the cross-border transfer question stops existing. Budget 8×H200 for GLM-5.3 at FP8; Flash is far more tractable at 320B/18B. If self-hosting is out of budget, Ollama’s Europe hosting with zero data retention is the pragmatic middle ground — but get the residency commitment in writing rather than inferring it from a marketing page.

5. “Small team, tight budget, coding all day.”

GLM-5.3-Flash as default, GLM-5.3 for hard problems, Ollama Pro at $20. Flash costs $0.15/$0.50 — currently half that until 9 September — and scores 57.5 on the intelligence index. Twenty dollars buys sixty dollars of tokens with no weekly caps. For a two-to-four person team this is close to unbeatable. The GLM Coding Plan at $18/month (Lite) is the alternative if you prefer a fixed quota to a credit pool.

GLM-5.3-Flash. It is the only model in this comparison under a standard OSI-approved licence (MIT), it is natively multimodal, it scores 57.5 on the neutral index, and the weights are on Hugging Face with no revenue gates, no attribution mandates and no security-review clause. When the question is “what can we defend in a contract review,” permissive licensing beats three points of benchmark every time.


Ce que je vais suivre au cours du prochain trimestre

Whether Fable 5.1 lands above 62.1 on the Artificial Analysis index. It launched the day before this article and has not been independently scored. Its benchmark deltas over Fable 5 suggest it should, but “should” is not “did,” and the gap to Kimi K3’s 59.7 is the number the whole open-versus-closed argument turns on.

Whether the licence drift continues. In eight weeks we went from GLM-5.2 under MIT to GLM-5.3 under a bespoke licence with a discretionary security review, and from Kimi K2’s modified-MIT to K3’s revenue-gated terms. If GLM-6 and K4 tighten further, “open weights” becomes a marketing term rather than a meaningful category. GLM-5.3-Flash staying MIT is the counter-signal worth tracking.

Whether other labs adopt Z.ai’s staged-release pattern. A two-week hold with a published safety rationale is a new norm. If it holds, it is a good one. If it becomes a reason weights ship later and later, it is a soft path to not shipping them at all.

Independent replication of the cyber capability claims. 2,436 vulnerabilities across 269 projects is an extraordinary number and it is entirely self-reported. Somebody neutral needs to check it, because if it is accurate it reframes the entire open-weights safety conversation, and if it is not, it was effective marketing.

Ollama’s per-token rates six months from now. $20 for $60 of usage is an aggressive introductory posture from a company that raised $88M in July. Whether those multipliers survive contact with real unit economics is the single biggest variable in the “cheap open models” thesis for small teams.

Whether local models close on the 30B tier. Qwen3.8-27B at 61.7% SWE-bench on 24 GB is the most under-discussed result of the year. The frontier gets the headlines; the 24 GB tier is what changes what an ordinary business can do without an API key.


Foire aux questions

Is Kimi K3 really open source? No. Kimi K3’s weights are freely downloadable from Hugging Face, but under a custom “Kimi K3 License,” not an OSI-approved open-source licence. Model-as-a-Service operators whose aggregate revenue exceeds $20 million over any consecutive 12 months must negotiate a separate commercial agreement with Moonshot, and products above 100 million monthly active users or $20 million monthly revenue must display “Kimi K3” in their interface. Purely internal use is unrestricted.

Is GLM-5.3 better than Kimi K3? They are effectively tied on the neutral aggregate index — 59.5 against 59.7 on Artificial Analysis. GLM-5.3 is 2.2× cheaper per agentic task and scores higher on knowledge work (GDPval-AA v2: 1,769 vs 1,687). Kimi K3 is natively multimodal, leads on agentic browsing (BrowseComp 91.2) and took first place on Arena.AI’s frontend code arena. Choose GLM-5.3 for cost and text-based coding; Kimi K3 for vision, browsing and long-horizon research.

How much cheaper are open models than Claude Fable 5.1? On a modelled 60-turn agentic task, GLM-5.3 costs about $1.85 against Fable 5.1’s $11.94 — roughly 6.5× cheaper. Kimi K3 costs about $4.07, roughly 2.9× cheaper. GLM-5.3-Flash costs about $0.21, roughly 58× cheaper. Over 80 runs a month that is $148 versus $955.

Can I run Kimi K3 or GLM-5.3 locally? Not on consumer hardware. Kimi K3 is roughly 1.5 TB even in its native MXFP4 format and Moonshot recommends 64+ accelerators. GLM-5.3 needs a minimum of 8×H200 (1,128 GB) at FP8. For genuine local execution, look at Qwen3.8-27B (24 GB at Q4), gpt-oss:20b (16 GB) or Gemma 4 E4B (~6 GB).

What changed in Claude Fable 5.1? Released 1 September 2026. Cache reads dropped 75% from $1.00 to $0.25 per million tokens, making typical workloads about 25% cheaper and agentic workloads up to 45% cheaper; input and output stayed at $10/$50. Terminal-Bench 4.0 rose from 42.0% to 55.8%, Terminal-Bench-Science from 24.7% to 52.6%, AutomationBench from 17.1% to 31.4%. Three breaking changes affect forced tool use, thinking-block portability and history editing.

What is Ollama’s new pricing? From 31 August 2026, Pro, Max and Team plans use per-token pricing with included credits: Pro $20/month for $60 of usage, Max $100 for $300, Team $500 for $1,000 shared across unlimited users. The 5-hour and weekly caps were removed entirely, there are no service fees, credits do not roll over, and all plans carry zero data retention with hosting in the US and Europe.

Which of these models is genuinely MIT-licensed? Only GLM-5.3-Flash — a separate 320B/18B natively multimodal model released 26 August 2026, scoring 57.5 on the Artificial Analysis index at $0.15/$0.50 per million tokens. GLM-5.3 and Kimi K3 both use bespoke revenue-gated licences; Claude Fable 5.1 is fully proprietary.

Why can’t I compare these models on Terminal-Bench? Because they were each evaluated on a different version. Kimi K3’s 88.3 is on Terminal-Bench 2.1, GLM-5.3’s 28.3 is on 3.0, and Fable 5.1’s 55.8 is on 4.0. Each version is substantially harder than the last, so the numbers are not on the same scale. Compare within a version only. On Terminal-Bench 2.1, for example, Kimi K3 scores 88.3 against Claude Fable 5’s 88.0, GPT-5.6 Sol’s 88.8 and GLM-5.3-Flash’s 84.3 — that comparison is valid, and it is the one worth quoting.

Should I use Ollama or go direct to the vendor? Ollama if you want one billing relationship, one API surface, easy model switching and EU/US hosting with zero data retention — its per-token rates track first-party pricing closely. Direct if you need first-party features like Z.ai’s effort_de_raisonnement controls at full fidelity, vendor SLAs, or subscription plans such as the GLM Coding Plan. Many teams run both and route by workload.

Does GLM-5.3 support images? No. GLM-5.3 is text-only. GLM-5.3-Flash — a completely different model despite the similar name — is natively multimodal and handles text, image, video and file input. Sending images to the wrong model ID is the most common early mistake with the GLM family.


En résumé

Twelve months ago the open-weight question was whether these models were usable. Six months ago it was whether they were competitive. In September 2026 it is genuinely: what are you still paying a closed-model premium for?

There is a real answer to that question, and it is narrower than it used to be. Claude Fable 5.1 leads on the hardest sustained agentic work, on general knowledge work, and on the class of debugging where a model needs to hold a messy problem in its head for hours without drifting. Anthropic’s 75% cache-read cut targets exactly that workload. If your failure cost exceeds your token cost, that premium is rational.

For everything else, the maths has moved. GLM-5.3 delivers 94% of Claude Opus 5’s index score at 29% of the cost. Kimi K3 scores 88.3 on Terminal-Bench 2.1 against Claude Fable 5’s 88.0 on the same version, and took first place on LMArena’s Frontend Code Arena ahead of both flagship closed models. Ollama will sell you $60 of tokens for $20 and remove the usage caps while doing it. That combination did not exist in the spring.

But do not let the enthusiasm skip the fine print, because there are two of them and both matter.

The first is licensing. “Open weights” and “open source” have quietly stopped meaning the same thing. Kimi K3 and GLM-5.3 both ship under bespoke, revenue-gated licences. The gates are high enough that most readers are unaffected — but “most readers are unaffected” is not the same as “unrestricted,” and GLM-5.3’s undefined security-review clause is a genuine procurement risk for large enterprises. GLM-5.3-Flash under MIT is the last fully permissive frontier-adjacent option, which is precisely why it deserves more attention than it gets.

The second is that self-hosting is mostly aspirational. 1.5 TB of weights and 64 accelerators is not an exit plan for a mid-sized business. What open weights buy you at this tier is a competitive inference market, price transparency, version pinning, jurisdiction choice, and a negotiating position. Those are worth a great deal. They are not the same as running the thing in your basement.

The setup I would actually build: GLM-5.3 as the default, Fable 5.1 on planning and the genuinely hard problems, Kimi K3 for vision and long-horizon research, GLM-5.3-Flash for the grunt work, and a 27B local model for anything that must never leave the building. Route by workload, not by loyalty. Write your configuration so the model IDs are a one-line change.

Because the only prediction I am confident about is that this article will need updating before Christmas.


Are you running any of these three in production? I am particularly interested in whether GLM-5.3’s verbosity shows up as a real cost problem at scale, and whether anyone has actually put the Kimi K3 licence in front of counsel and got a clear read on the MaaS definition. Get in touch — corrections and counter-evidence welcome, and this article gets updated when the picture changes.

Last updated: 2 September 2026.

Ajoutez votre signal.

Votre adresse e-mail ne sera pas publiée. Les champs obligatoires sont indiqués avec *

fr_FRFrançais