Sei settimane fa avrei scritto questo articolo in modo diverso.
A luglio, Moonshot AI ha pubblicato su Hugging Face un modello all’avanguardia con 2,8 trilioni di parametri, rendendolo disponibile per il download a chiunque. Il 28 agosto, Z.ai ha finalmente rilasciato i pesi del GLM-5.3 che teneva in serbo da due settimane, dopo aver scoperto che il proprio modello era in grado di concatenare exploit in modi che nessuno gli aveva chiesto di fare. E il 1° settembre, Anthropic ha lanciato Claude Fable 5.1, rendendo il lavoro degli agenti fino a 45% più economico senza modificare il prezzo di listino.
Tre modelli. Tre risposte completamente diverse alla domanda su cosa un laboratorio di intelligenza artificiale debba al pubblico. E, per la prima volta da quando è iniziata tutta questa vicenda, le opzioni "open-weight" non rappresentano una scelta di compromesso.
Questa è la storia vera della fine del 2026, e vale circa ottomila parole, perché è proprio nei dettagli che si nascondono tutti i soldi e tutti gli errori.
Ho esaminato i documenti relativi all’architettura, le schede dei modelli, i testi delle licenze — con attenzione, riga per riga, perché due dei tre non sono proprio ciò che la gente chiama comunemente con quei nomi — i benchmark dei fornitori, i punteggi degli aggregatori neutrali, i tassi reali per token e il nuovissimo listino prezzi di Ollama. Poi ho calcolato i costi basandomi su un normale mese lavorativo anziché su uno scenario relativo al giorno del lancio.
Ecco cosa ne è venuto fuori.
TL;DR — Il verdetto in 60 secondi
GLM-5.3 rappresenta attualmente il miglior rapporto qualità-prezzo nel campo dell'intelligenza artificiale, e non c'è proprio paragone. Secondo Artificial Analysis, il modello si colloca a meno di mezzo punto da Kimi K3 e a circa tre punti e mezzo da Claude Opus 5, con $1,40 in entrata e $4,40 in uscita per milione di token. Per lo stesso compito di tipo agentico, costa all’incirca un sesto di quanto costa Fable 5.1. I pesi sono scaricabili. Se cercate un modello su cui basare il vostro agente e non gestite un’azienda con un fatturato di $10 miliardi, eccolo qui.
Kimi K3 è il programma più potente che si possa scaricare legalmente. 2,8 trilioni di parametri, 104 miliardi attivi, un contesto autentico da 1 milione, visione nativa, il miglior punteggio di navigazione agentica mai pubblicato (BrowseComp 91,2) e — il dato che dovrebbe porre fine alla tesi secondo cui “i modelli aperti sono in ritardo” — 88,3 su Terminal-Bench 2.1, contro l’88,0 di Claude Fable 5 e l’88,8 di GPT-5.6 Sol sulla stessa versione. Inoltre costa $3/$15 — più del doppio rispetto al GLM-5.3 — e i suoi 1,5 TB di pesi richiedono un rack, non una workstation. È il modello che si utilizza quando il compito è complesso e di lunga durata, ed è il modello che si cita quando qualcuno sostiene che i pesi aperti siano di una generazione arretrata.
Claude Fable 5.1 rimane il modello a cui ricorrere quando un errore può costare caro. È all’avanguardia nel lavoro agentico prolungato, della durata di diverse ore, e nel debug complesso che non compare mai in modo chiaro nelle classifiche. Il 1° settembre Anthropic ha ridotto le letture dalla cache di 75%, rendendo le sessioni lunghe degli agenti notevolmente più economiche rispetto al passato. Il suo costo per task è inoltre sei volte superiore a quello di GLM-5.3, è in versione chiusa e ora prevede restrizioni anti-distillazione che compromettono alcune integrazioni esistenti.
La battuta schietta: I modelli open-weight hanno ridotto il divario in termini di capacità a circa tre-cinque mesi e hanno azzerato il divario di prezzo. Ciò che in realtà si acquista da Anthropic nel settembre 2026 è l’affidabilità nelle attività più complesse del 10% — e il diritto di non dover pensare a nulla di tutto ciò.
E il colpo di scena che nessuno vuole ammettere ad alta voce: Né Kimi K3 né GLM-5.3 sono più distribuiti con licenza open source. Entrambi sono passati quest’estate a condizioni personalizzate, subordinate al raggiungimento di determinati ricavi. Pesi aperti, sì. Open source, no. Leggete la sezione sulle licenze prima che lo faccia il vostro team legale.

Tabella di confronto rapido
| Kimi K3 | GLM-5.3 | Claude Fable 5.1 | |
|---|---|---|---|
| Produttore | Moonshot AI | Z.ai (Zhipu AI) | Antropico |
| Pubblicato | 16 luglio 2026 (API) · 27 luglio (pesi) | 14 agosto 2026 (API) · 28 agosto (pesi) | 1° settembre 2026 |
| Pesi disponibili | Sì | Sì | No |
| Licenza | Licenza Kimi K3 (personalizzata, a pagamento) | Licenza glm-5.3 (personalizzata, basata sui ricavi) | Esclusivo |
| Parametri totali | 2,8 T | 753B | Non divulgato |
| Attivo per token | 104B (16 su 896 esperti) | ~40 miliardi (stima) | Non divulgato |
| Architettura | KDA + residui di attenzione, LatentMoE stabile, 69 livelli KDA + 24 livelli MLA con gate | MoE, con la stessa base del GLM-5.2, trae vantaggio esclusivamente dal post-addestramento | Non divulgato |
| Finestra di contesto | 1,048,576 | 1,000,000 | 1,000,000 |
| Potenza massima | 131.072 in default (fino a 1.048.576) | 128,000 | 128,000 |
| Input multimodale | Testo, immagine, video (MoonViT-V2) | Solo testo | Testo, immagine |
| Riflessioni | Sempre attivo | Sempre attivo, sforzo_di_ragionamento basso/alto/massimo | Sempre attivo, adattivo, sforzo per singolo messaggio (beta) |
| Prezzo di acquisto / 1 milione | $3.00 | $1.40 | $10.00 |
| Dati in cache / 1M | $0.30 | $0.26 | $0.25 |
| Prezzo di vendita / 1 milione | $15.00 | $4.40 | $50.00 |
| Indice di intelligenza AA | 59.7 | 59.5 | Non è ancora stato valutato (Fable 5: 62,1) |
| Costo per attività di valutazione AA | $0.84 | $0.68 | — (Opus 5: $2.34) |
| Impronta di hosting autonomo | ~1,5 TB (MXFP4 nativo) | ~1,5 TB BF16 / ~750 GB FP8 | N/D |
| Su Ollama | kimi-k3:cloud | glm-5.3:cloud | Tramite Claude Code / Claude Desktop come client |
Tutti i prezzi indicati sono quelli di listino del produttore. I dati di Artificial Analysis si riferiscono all’istantanea dell’indice di settembre 2026; Fable 5.1 è stato lanciato il 1° settembre e, al momento della stesura del presente documento, non era ancora stato valutato da fonti indipendenti.
Perché questo confronto a tre è quello che conta davvero
Per due anni il dibattito tra modelli aperti e chiusi ha seguito uno schema ben definito: i modelli chiusi erano migliori, quelli aperti più economici, e ognuno sceglieva la propria posizione in base a quel compromesso. Tutti sapevano da che parte stare.
Quella forma si è rotta quest’estate.
Nathan Lambert, che segue la questione con maggiore attenzione rispetto a quasi chiunque altro, stima che l'attuale divario di capacità tra i modelli "frontier open-weight" e quelli "frontier closed" sia pari a da tre a cinque mesi — in calo rispetto ai sei-nove mesi che si ipotizzavano un anno fa. Al momento in cui scriveva, Kimi K3 si collocava al #2 nell’indice Vals AI e vicino alla vetta di quello di Artificial Analysis, superato solo da Claude Fable e GPT-5.6 Sol Max. L’istantanea dell’indice di settembre che utilizzerò più avanti in questo articolo lo colloca al quarto posto, dietro a Grok 4.6. In ogni caso: non si tratta di un risultato “buono per un modello aperto”. È un piazzamento sul podio.
Nel frattempo, i dati sull’utilizzo hanno registrato un andamento davvero sorprendente. I modelli di origine cinese hanno registrato circa 61% di tutti i token instradati tramite OpenRouter entro maggio 2026. La quota statunitense di tale traffico è crollata da circa 70% a circa 30% nell’arco di dodici mesi. Llama di Meta — il modello che ha dato il via all’ondata dei modelli open-weight — è scesa al di sotto degli 1% di volume instradato. La quota di Google è passata da circa 37% a 13%. Lo studio condotto da OpenRouter in collaborazione con a16z su 100 trilioni di token ha rilevato che i modelli open-weight rappresentano circa un terzo del volume totale di token sulla piattaforma.
Si può discutere su cosa rappresenti il traffico di OpenRouter. Esso sovrastima la presenza di sviluppatori, appassionati e carichi di lavoro sensibili ai costi, mentre sottostima i contratti aziendali che non prevedono mai l’utilizzo di un router. Va bene. Ma si tratta del più ampio set di dati pubblico di cui disponiamo su ciò che le persone scelgono effettivamente quando possono scegliere liberamente, e la tendenza è inequivocabile.
Quindi, nel settembre 2026, la domanda non sarà più “i modelli aperti sono già abbastanza validi?”. Sarà molto più precisa:
Se ci sono tre modelli che soddisfano tutti i requisiti, quale indicheresti al tuo agente? E quanto ti costerebbe effettivamente sbagliare?
Si tratta di una questione che riguarda i benchmark, i prezzi, le licenze e le infrastrutture, più o meno nell’ordine in cui la gente li considera e esattamente nell’ordine inverso rispetto alla loro importanza.
Che cos’è in realtà la Kimi K3
Moonshot AI ha annunciato Kimi K3 il 16 luglio 2026 e ne ha reso noti i pesi completi il 27 luglio. Si tratta, in termini di numero di parametri, del modello a peso aperto più grande mai pubblicato: 2,8 trilioni di parametri in totale, circa 75% in più rispetto a DeepSeek V4 Pro.
Quel numero è meno impressionante di quanto sembri e più impressionante di quanto sembri, proprio in quest’ordine.
L'architettura, in parole semplici
Kimi K3 è un modello “mixture-of-experts” che attiva 104 miliardi di parametri per token, selezionando 16 esperti da un gruppo di 896. Quindi, sebbene il totale sia pari a 2,8T, ogni singolo passaggio in avanti coinvolge meno di 4% dei pesi. È da qui che deriva il vantaggio economico: un’enorme quantità di conoscenza immagazzinata, a fronte di un costo di inferenza moderato.
L'aspetto ingegneristico più interessante risiede nello stack di attenzione. K3 utilizza 93 livelli costituiti da 69 livelli KDA (Kimi Delta Attention) e 24 livelli Gated MLA — un modello ibrido che combina attenzione lineare e attenzione completa con un rapporto di interleaving di circa 3:1. Il KDA è una variante dell’attenzione lineare che opera in tempo lineare rispetto alla lunghezza della sequenza anziché in tempo quadratico, ed è proprio questo che rende il contesto da 1 milione di simboli economicamente fattibile, anziché solo teorico. Moonshot riporta valori fino a un 75% Riduzione della cache KV e fino a Velocità di decodifica 6 volte superiore con 1 milione di contesti di conseguenza.
Sopra c’è Attenzione ai residui (AttnRes), che Moonshot descrive come un sostituto diretto delle connessioni residue standard, e un LatentMoE stabile struttura per l’instradamento. Si sostiene che tale combinazione garantisca all’incirca un Miglioramento di 2,5 volte dell'efficienza complessiva di scalabilità contro Kimi K2.
La visione nasce da MoonViT-V2, un encoder con 401M parametri, integrato nativamente anziché aggiunto a posteriori: K3 gestisce testo, immagini e video in un unico modello.
Un altro dettaglio importante per chiunque stia valutando l’opzione dell’hosting autonomo: K3 è stato addestrato con addestramento nativo MXFP4 che tiene conto della quantizzazione, con esperti in MXFP4 e attivazioni in MXFP8. Non si tratta di una quantizzazione post hoc. Il checkpoint pubblicato è il modello quantizzato, motivo per cui 2,8 trilioni di parametri si attestano a circa 1,5 TB su disco anziché i circa 5,6 TB che un checkpoint BF16 di quelle dimensioni richiederebbe.
La storia del benchmark
I dati pubblicati da Moonshot sono ottimi su tutta la linea e, cosa insolita, molti di essi hanno superato una valutazione indipendente:
- Terminal-Bench 2.1: 88,3 — il punteggio più alto pubblicato per quella versione del benchmark
- FrontierSWE: 81,2
- DeepSWE: 67,5
- BrowseComp: 91,2 — lo stato dell’arte della ricerca sul web agenziale
- GPQA Diamond: 93,5
- MMMU-Pro: 81,6 / 83,4 (visione)
- Maratona SWE: 91,0
- OfficeQA Pro: 81,3
- AutomationBench: 30,8
- LMArena Frontend Code Arena: 1.679 Elo — primo posto, davanti a Claude Fable 5 (1.631) e GPT-5.6 Sol (1.618), aggiudicandosi sei dei sette domini frontend
Quest’ultimo dato merita una riflessione. Un modello a peso aperto scaricabile gratuitamente si è classificato al primo posto in una competizione di programmazione di front-end basata sulle preferenze umane, superando i modelli chiusi di punta dei due laboratori più finanziati al mondo. A prescindere da ciò che si pensi dei benchmark di questo tipo come metodologia, si tratta di una notizia che sarebbe stata impensabile nel 2025.
Per quanto riguarda gli indici aggregati, si attesta a un livello leggermente inferiore: 59,7 nell'Indice di Intelligenza Analitica Artificiale, terzo in classifica generale, dietro a Claude Opus 5 (63,0) e Claude Fable 5 (62,1). I dati di riferimento di Moonshot — GDPval-AA v2 a 1.687 e AA-Briefcase a 1.527 — lo collocano rispettivamente al terzo e al secondo posto in quelle valutazioni sul lavoro intellettuale. Nell’Indice di Vals v2 si colloca al 57.8%, rispetto al 67,21 TP3T di Claude Opus 5, il 66,01 TP3T di Claude Fable 5 e il 63,71 TP3T di GPT-5.6 Sol — il divario più ampio tra ’aperto’ e "chiuso" in qualsiasi indice che ho trovato, e che vale la pena tenere a mente quando si analizzano i dati di codifica.

Le demo e come interpretarle
Moonshot ha pubblicato due dimostrazioni di autonomia a lungo termine che hanno suscitato grande interesse: una Esecuzione autonoma di 48 ore che completa l'intero processo di progettazione di un chip (un progetto da 4 mm con frequenza di chiusura a 100 MHz, simulato a oltre 8.700 token al secondo) e una riproduzione dell'astrofisica Relazione I-Love-Q tra circa due ore, contro le una o due settimane di cui avrebbe normalmente bisogno un ricercatore senior.
Sono davvero impressionanti, ma non baserei su di esse una decisione di acquisto. Le dimostrazioni organizzate, preparate e selezionate dal fornitore mostrano ciò che un modello è in grado di fare in una giornata ideale sotto la supervisione di un esperto, non ciò che fa nel vostro repository un martedì qualsiasi. Consideratele una prova di fattibilità per capacità a lungo termine, non come specifiche tecniche.
Mostra immagine Due architetture rese pubbliche e una “scatola nera”. I modelli aperti spiegano esattamente come funzionano; quello chiuso indica il punteggio ottenuto.
Che cos’è in realtà GLM-5.3
La versione GLM-5.3 ha la storia di rilascio più singolare delle tre, ed è quella che vale la pena raccontare per bene, perché è la prima volta che un importante laboratorio ha visibilmente esitato a pubblicare i pesi per ragioni diverse da quelle commerciali.
Z.ai ha annunciato GLM-5.3 il 14 agosto 2026 con lo slogan “Built to Code. Ready for Cyber Defense.” Il modello è un Miscela di esperti con 753 miliardi di parametri, con circa 40 miliardi di token attivi secondo le stime della comunità ricavate dalla configurazione GLM-5.2. Presenta una finestra di contesto di 1 milione di token e un output massimo di 128K e — a differenza di GLM-5.3-Flash, che è un modello completamente diverso — è solo testo.
“Ci siamo limitati a scalare dopo l'allenamento”
L'aspetto tecnico più interessante del GLM-5.3 è che non esiste un nuovo modello di base. Esso riutilizza la stessa base preaddestrata del GLM-5.2. Tutti i miglioramenti sono stati ottenuti grazie a un post-addestramento notevolmente esteso, che Z.ai ha sintetizzato come segue: “Per la versione GLM-5.3 ci siamo limitati a ottimizzare le prestazioni dopo la fase di addestramento.”
I delta non sono trascurabili:
| Punto di riferimento | GLM-5.2 | GLM-5.3 |
|---|---|---|
| Terminal-Bench 3.0 | 4.6 | 28.3 |
| DeepSWE v1.1 | 46.2 | 66.9 |
| CyberGym | 77.2% | 84.5% |
Un miglioramento di sei volte rispetto a Terminal-Bench e un balzo di venti punti su DeepSWE, partendo dagli stessi pesi di base, è un risultato davvero significativo. Ciò dimostra che la frontiera ha ancora un ampio margine di miglioramento nella fase post-addestramento, che rappresenta di gran lunga la parte più economica del costo complessivo dell’addestramento. È proprio per questo che i laboratori cinesi, che operano con, come afferma Lambert, “ordin di grandezza inferiori di capitale” rispetto a quelli americani, riescono a tenere il passo.
Le due settimane in cui Z.ai non ha effettuato spedizioni
Z.ai aveva dichiarato al momento del lancio che i pesi sarebbero stati resi disponibili “in più fasi, al termine di rigorose valutazioni di sicurezza”, con un obiettivo di circa due settimane. Il suo repository su Hugging Face indicava la data del 28 agosto. Quella data è arrivata, è trascorsa nel giro di poche ore e poi i pesi sono stati pubblicati insieme a una nota tecnica che spiegava il ritardo.
Il motivo era legato alle capacità informatiche. Durante la valutazione, Z.ai ha riferito che GLM-5.3 ha identificato 2.436 vulnerabilità in 269 progetti open source, di cui 1.097 sono state classificate come critiche o di gravità elevata. Più precisamente, il modello ha dimostrato che ragionamento sulla catena di exploit a più fasi — concatenando in modo autonomo le fasi di sfruttamento — cosa che Z.ai ha definito come un comportamento “non del tutto intenzionale” e che “continuava ad aggravarsi man mano che l’addestramento si ampliava”.”
Voglio essere cauto su questo punto, perché è facile interpretare quanto detto come una mossa di marketing o come allarmismo, mentre probabilmente non è né l’uno né l’altro. Un modello che è efficace nell’individuare le vulnerabilità è efficace nell’individuare le vulnerabilità; gli usi difensivi e quelli offensivi sono la stessa capacità orientata in direzioni diverse. Il divario nei benchmark di Z.ai riflette onestamente questo aspetto: CyberGym 84.5% (individuazione delle vulnerabilità, con tutte le sue conseguenze) contro ExploitBench 54.4% (sfruttamento, dove Claude Fable 5 totalizza 78,01 TP3T). Il laboratorio ha ottimizzato deliberatamente la metà campo difensiva e i risultati lo dimostrano.
Ciò che è davvero degno di nota è che un laboratorio abbia ritardato di due settimane il lancio di un prodotto di punta e abbia reso pubbliche le proprie motivazioni. A prescindere dal fatto che si ritenga o meno convincenti tali motivazioni, si tratta comunque di un precedente senza precedenti.
Dove atterra effettivamente il GLM-5.3
La tabella comparativa pubblicata da Z.ai non pretende di essere esaustiva, il che è una boccata d’aria fresca:
| Punto di riferimento | GLM-5.3 | Il miglior comparatore |
|---|---|---|
| AutomationBench | 48.2% | Cavi GLM-5.3 |
| CyberGym | 84.5% | Cavi GLM-5.3 |
| GDPval-AA v2 | 1,769 | Cavi GLM-5.3 |
| Terminal-Bench 3.0 | 28.3% | GPT-5.6 Risoluzione: 34,61 TP3T |
| DeepSWE v1.1 | 66.9% | GPT-5.6 Risultato: 72,71 TP3T |
| ExploitBench | 54.4% | Claude Fable 5: 78.0% |
| Code Bench (50.000 token) | 31.4% | — |
Da parte sua, Artificial Analysis lo stima a 59,5 nell'Indice di intelligenza — statisticamente indistinguibile dal valore di 59,7 ottenuto da Kimi K3, ricavato da un modello con circa un quarto dei parametri.
Un aspetto da tenere presente, emerso dalla stessa valutazione, ma che nessuno nel reparto marketing menziona: GLM-5.3 ha generato 170 milioni di token generati dalla suite Artificial Analysis, a fronte di una mediana di 72 milioni per la sua categoria. È troppo prolisso. sforzo_di_ragionamento il valore predefinito è max e il pensiero non può essere disattivato, quindi si finisce per pagare per un lungo processo di riflessione, indipendentemente dal fatto che l’attività lo giustifichi o meno. Ciononostante, il costo totale per eseguire la valutazione completa è ammontato a $0,68 per attività, contro $0,84 per Kimi K3 e $2,34 per Claude Opus 5 — quindi la maggiore complessità non annulla il vantaggio in termini di prezzo, ma lo riduce soltanto.
Che cos’è in realtà Claude Fable 5.1
Anthropic ha rilasciato Claude Fable 5.1 il 1° settembre 2026, insieme a Claude Mythos 5.1 — lo stesso modello con una configurazione di sicurezza diversa, disponibile esclusivamente tramite programmi di accesso fidato destinati alla sicurezza informatica e alle scienze della vita.
La definizione è precisa e, a mio avviso, corretta: si tratta di “un modello creato per svolgere un lavoro che non si esaurisce in un unico comando”.” Anthropic non sostiene che si tratti di un salto di qualità in termini di intelligenza generale. Afferma piuttosto che le sessioni di interazione con gli agenti, prolungate e della durata di diverse ore, danno risultati migliori.
I numeri
| Punto di riferimento | Fable 5 | Fable 5.1 | Mythos 5.1 |
|---|---|---|---|
| Terminal-Bench 4.0 | 42.0% | 55.8% | 60.9% |
| Terminal-Bench-Science 0.1 | 24.7% | 52.6% | — |
| AutomationBench | 17.1% | 31.4% | — |
| CursorBench 3.2.0 | 70.5% | 73.4% | — |
| OSWorld 2.0 (parziale) | 72.9% | 77.9% | — |
| L'ultimo esame dell'umanità (con strumenti) | 63.8% | 65.0% | — |
| GDPval-AA v2 | — | 1,853 | — |
Per ulteriori informazioni sul contesto esterno relative allo stesso sovrano: GPT-5.6 Sol ottiene un punteggio di 52,31 TP3T su Terminal-Bench 4.0, quindi il 55,81 TP3T di Fable 5.1 si porta in testa e il 60,91 TP3T di Mythos 5.1 ne amplia il vantaggio.
Il salto registrato nel benchmark “Terminal-Bench-Science” — da 24,71 TP3T a 52,61 TP3T, un vero e proprio raddoppio — è quello a cui attribuirei maggiore importanza se il vostro lavoro riguardasse il calcolo scientifico o le pipeline di dati. E nel benchmark di Browserbase relativo all’utilizzo più intensivo del computer, Fable 5.1 ha completato 82% di attività contro le 74% di Claude Opus 5.
Anthropic riferisce inoltre che circa un 60%: riduzione dei falsi positivi nella sicurezza informatica — il modello che rifiuta un’operazione di sicurezza legittima perché la sua corrispondenza con un pattern indica qualcosa di pericoloso. Se vi è mai capitato che un assistente di programmazione si rifiutasse di aiutarvi a scrivere un limitatore di velocità, capirete perché questo è importante.
La novità vera e propria è la modifica dei prezzi
I tassi di riferimento sono rimasti invariati: $10 per milione di token in ingresso, $50 per milione di token in uscita. Ciò che è cambiato è il moltiplicatore di lettura della cache, e il cambiamento è stato notevole.
| Fable 5 | Fable 5.1 | |
|---|---|---|
| Input | $10.00 | $10.00 |
| Risultato | $50.00 | $50.00 |
| Scrittura nella cache (5 min) | $12.50 | $12.50 |
| Scrittura nella cache (1 ora) | $20.00 | $20.00 |
| Lettura dalla cache | $1.00 | $0.25 |
Si tratta di una riduzione denominata 75%, ottenuta portando il moltiplicatore di lettura della cache dal valore standard di 0,1× dell’input di base a 0,025×. Anthropic riporta approssimativamente 25% in meno per carichi di lavoro tipici e fino a 45% in meno per attività altamente agenti — e quella seconda cifra è reale, perché i cicli degli agenti consistono prevalentemente in letture dalla cache. Una sessione di un agente di 60 turni rilegge lo stesso contesto accumulato decine di volte. L’API Batch dimezza inoltre gli input e gli output, arrivando a $5/$25.
Si tratta di una riduzione di prezzo intelligente e mirata. Rende l'aspetto in cui Fable 5.1 eccelle — le sessioni di gioco prolungate — significativamente più conveniente senza sminuire il valore complessivo del modello.
Tre modifiche che comportano incompatibilità, una regressione
Prima di aggiornare un'integrazione in produzione, leggi attentamente queste istruzioni.
1. L'uso obbligatorio degli strumenti è stato eliminato. Richieste che utilizzano tool_choice: "qualsiasi" o scelta_strumento: "strumento" ora restituisci un Errore 400. Il processo di elaborazione è sempre attivo e non può essere bypassato; forzando l'esecuzione di uno strumento, lo si salterebbe. È necessario utilizzare tool_choice: "auto".
2. I blocchi di pensiero sono unidirezionali. Fable 5.1 è in grado di leggere i ’thinking block" generati dai modelli precedenti di Claude, ma i modelli precedenti non sono in grado di leggere quelli di Fable 5.1. Solo Mythos 5.1 è in grado di farlo. Se la vostra architettura cambia modello nel bel mezzo di una conversazione per risparmiare, tale soluzione di ripiego ora impone una ripianificazione.
3. La cronologia delle modifiche rende obsoleti i blocchi mentali. Per gli account creati dopo il 31 agosto 2026, la modifica dei turni precedenti, delle richieste del sistema o della matrice degli strumenti invalida tutti i blocchi di ragionamento successivi — sia che si tratti di errori che di omissioni silenziose. Si tratta di una misura esplicita contro la distillazione che interrompe un modello comune in cui i framework degli agenti riscrivono il contesto tra un turno e l’altro.
E un dato interessante da conoscere: A quanto pare, nella versione 5.1 le chiamate parallele alle funzioni sono più variabili: a volte viene emessa una sola chiamata per turno, mentre nella versione 5 ne venivano raggruppate diverse. In un ciclo lungo dell’agente, ciò comporta un dispendio di tempo reale.
C'è anche un dettaglio relativo al tokenizer che spesso sfugge quando si confrontano i costi: Fable 5.1 utilizza lo stesso tokenizer di Claude Opus 4.7, che produce all'incirca 30% altri gettoni rispetto al tokenizer dei modelli Claude precedenti per lo stesso testo. Se state confrontando i costi con un valore di riferimento del 2025, tenetene conto prima di trarre qualsiasi conclusione.
Il problema del benchmark: tre modelli, tre diversi criteri di valutazione
Ecco il punto su cui la maggior parte degli articoli comparativi sbaglia, e vale la pena dirlo senza mezzi termini.
Non è possibile allineare questi tre modelli su Terminal-Bench. Guarda:
- Kimi K3: 88.3 su Terminal-Bench 2.1
- GLM-5.3: 28.3 su Terminal-Bench 3.0
- Fable 5.1: 55.8 su Terminal-Bench 4.0
Si tratta di tre benchmark diversi che condividono lo stesso nome. Ogni versione è risultata notevolmente più difficile della precedente: è proprio questo lo scopo del rilascio di una nuova versione. Interpretare quei tre numeri come una classifica porterebbe a concludere che Kimi K3 sia tre volte migliore di GLM-5.3 nelle operazioni terminali, il che è una sciocchezza.
Lo stesso problema si riscontra in AutomationBench, dove i numeri di versione vengono riportati in modo incoerente, e nella famiglia SWE-bench, dove le varianti Verified, Pro e Marathon misurano aspetti effettivamente diversi. I laboratori non lo fanno con l’intento di ingannare nessuno: eseguono la valutazione che era in vigore al momento della loro formazione, e le valutazioni cambiano ormai ogni pochi mesi. Tuttavia, l’effetto su un lettore non esperto è lo stesso di un inganno, quindi vale la pena segnalarlo con forza.
Ciò che è possibile confrontare in modo legittimo, all'interno di una stessa versione:
Su Terminal-Bench 2.1: Kimi K3 (88,3), GPT-5.6 Sol (88,8), Claude Fable 5 (88,0) e GLM-5.3-Flash (84,3). Quattro modelli, un unico metro di valutazione. Questo è il confronto che conta davvero e gli ho dedicato una tabella a parte nella sezione successiva.
Su Terminal-Bench 4.0: Fable 5.1 (55,8) contro GPT-5.6 Sol (52,3). Vince Fable.
Oltre a ciò, diffidate di chiunque vi mostri un unico grafico a barre che riporti tutti e tre i modelli citati nel titolo di questo articolo su un asse “Terminal-Bench”. Non è possibile farlo in modo onesto.
Dove si incontrano realmente: i dati comparabili tra loro
Se si escludono le discrepanze tra le versioni dei benchmark, rimangono tre confronti.
1. Indice di intelligenza artificiale applicata all'analisi (settembre 2026)
Gestione indipendente, stessa metodologia applicata a tutti i modelli, nessun coinvolgimento dei fornitori nella selezione.
| Classifica | Modello | Punteggio | Pesi liberi |
|---|---|---|---|
| 1 | Claude Opus 5 | 63.0 | No |
| 2 | Claude Fable 5 | 62.1 | No |
| 3 | Grok 4.6 | 60.9 | No |
| 4 | Kimi K3 | 59.7 | Sì |
| 5 | GLM-5.3 | 59.5 | Sì |
| 6 | GPT-5.6 Sol | 58.9 | No |
| 7 | Anteprima di Qwen 3.8 Max | 58.1 | Pesi da definire |
| 8 | GLM-5.3-Flash | 57.5 | Sì (MIT) |
| 31 | MiniMax M3 | 45.4 | Sì |
Al momento della stesura di questo articolo, Fable 5.1 non era ancora stato valutato da fonti indipendenti, essendo stato lanciato il giorno prima. Considerati i miglioramenti registrati da Fable 5.1 rispetto a Fable 5 nei benchmark, è prevedibile che, una volta valutato, il punteggio si attesti a 62,1 o oltre.
Una nota sulla precisione: i diversi istantanei di questo indice presentano variazioni diverse — in un’analisi di settembre, Kimi K3 e GLM-5.3 risultano entrambi a 60, mentre GPT-5.6 Sol è a 61. L’ordine nella parte alta della classifica è stabile; i distacchi tra i singoli punti non lo sono. Non basare le tue argomentazioni su mezzo punto.
Leggi attentamente quella tabella. Il miglior modello a ponderazione aperta registra un ritardo di 3,3 punti rispetto al miglior modello a ponderazione chiusa sull'indice neutro più ampio disponibile. Un anno fa quel divario era a due cifre.
2. Terminal-Bench 2.1 — l'unica versione in cui i tre si incontrano
Questa è la tabella più utile di tutto l’articolo, e per poco non me la lasciavo sfuggire. Kimi K3, Claude Fable 5 e GPT-5.6 Sol sono stati tutti valutati in base a la stessa versione di Terminal-Bench:
| Modello | Terminal-Bench 2.1 | Pesi liberi |
|---|---|---|
| GPT-5.6 Sol | 88.8 | No |
| Kimi K3 | 88.3 | Sì |
| Claude Fable 5 | 88.0 | No |
| GLM-5.3-Flash | 84.3 | Sì (MIT) |
Mezzo punto separa un modello scaricabile dal miglior modello chiuso nello stesso benchmark, nella stessa versione, nell’ambito del lavoro agentico nativo del terminale.
Questo è il numero da tenere a mente. Non l’aggregato dell’indice, né il grafico del fornitore: proprio questo. Per quanto riguarda la forma specifica del compito che la codifica agentica rappresenta di fatto, il divario è nascosto nel rumore.
Un avvertimento, sottolineato dalla fonte e che vale la pena ribadire: tutti i tempi pubblicati dalla Kimi K3 si riferiscono a giri percorsi al massimo, a sforzo_di_ragionamento massimo e temperatura 1,0. Laboratori diversi utilizzano sistemi di valutazione diversi e impostazioni di sforzo diverse, quindi anche i confronti tra versioni identiche comportano un margine di incertezza maggiore rispetto a quanto suggerirebbe una tabella semplice.
3. Costo di gestione della stessa suite di strumenti di valutazione
Se la tabella sopra riportata rappresenta il miglior confronto in termini di prestazioni, questo è il miglior confronto in termini di rapporto qualità-prezzo — poiché è l’unico dato che riunisce in un unico valore le prestazioni, la verbosità, il sovraccarico di ragionamento e il prezzo:
| Modello | Costo per attività dell’Intelligence Index |
|---|---|
| GLM-5.3 | $0.68 |
| Kimi K3 | $0.84 |
| Claude Opus 5 | $2.34 |
Stesse attività, stesso punteggio, fatture reali. GLM-5.3 raggiunge un punteggio indice di 94% secondo Opus 5 a un costo di 29%.
4. GDPval-AA v2 (lavoro intellettuale)
| Modello | Punteggio |
|---|---|
| Claude Fable 5.1 | 1,853 |
| Claude Fable 5 Max | 1,815 |
| GLM-5.3 | 1,769 |
| GPT-5.6 Sol Max | 1,747.8 |
| Kimi K3 | 1,687 |
È opportuno fare una precisazione: questo dato è stato ricavato da grafici pubblicati dai fornitori piuttosto che da un'unica analisi indipendente, e i suffissi “Max” indicano diverse impostazioni di sforzo. Tuttavia, l’ordine dei risultati è sostanzialmente coerente tra le diverse fonti e rivela un dato di fatto: per quanto riguarda il lavoro intellettuale in generale, a differenza della programmazione, i modelli chiusi mantengono ancora un vantaggio più netto rispetto a quanto suggeriscano i benchmark relativi alla programmazione.
Questo è lo schema che si ripete in tutti e tre i confronti. I modelli aperti hanno sostanzialmente recuperato il ritardo per quanto riguarda la programmazione e le attività di tipo agentico. Sono ancora un po’ indietro per quanto riguarda il lavoro intellettuale in senso lato. Se il vostro carico di lavoro rientra nella prima categoria, le ragioni a favore del pagamento dei prezzi previsti dal modello chiuso sono deboli. Se invece rientra nella seconda, la scelta è comunque giustificabile.
Mostra immagine Tre punti e mezzo separano il miglior modello chiuso da quello scaricabile. Dodici mesi fa quel divario era a due cifre.
Il calcolo dei costi, con numeri reali
I parametri di riferimento sono astratti. Le fatture no.
Lo scenario
Un elemento di media entità, realizzato in modo autonomo: approssimativamente 60 turni di chiamata degli strumenti, con una media di 45.000 token di input per turno man mano che il contesto si arricchisce, e 2.000 token di output per turno.
- Totale immesso: 2,7 milioni di token
- Produzione totale: 120.000 token
- Tasso di successo nella cache dei prompt ipotizzato: 80% (2,16 milioni in cache, 540.000 aggiornati)
Claude Fable 5.1
| Componente | Gettoni | Tasso | Costo |
|---|---|---|---|
| Nuovi spunti | 540.000 | $10,00/M | $5.40 |
| Dati in cache | 2,16 milioni | $0,25/M | $0.54 |
| Risultato | 120.000 | $50,00/M | $6.00 |
| Totale | ≈ $11,94 |
Kimi K3
| Componente | Gettoni | Tasso | Costo |
|---|---|---|---|
| Nuovi spunti | 540.000 | $3,00/M | $1.62 |
| Dati in cache | 2,16 milioni | $0,30/M | $0.65 |
| Risultato | 120.000 | $15,00/M | $1.80 |
| Totale | ≈ $4,07 |
GLM-5.3
| Componente | Gettoni | Tasso | Costo |
|---|---|---|---|
| Nuovi spunti | 540.000 | $1,40/M | $0.76 |
| Dati in cache | 2,16 milioni | $0.26/M | $0.56 |
| Risultato | 120.000 | $4.40/M | $0.53 |
| Totale | ≈ $1.85 |
And for reference — GLM-5.3-Flash
| Componente | Gettoni | Tasso | Costo |
|---|---|---|---|
| Nuovi spunti | 540.000 | $0,15/M | $0.08 |
| Dati in cache | 2,16 milioni | $0.03/M | $0.06 |
| Risultato | 120.000 | $0,50/M | $0.06 |
| Totale | ≈ $0,21 |
The ratios
- GLM-5.3 is 6.5× cheaper than Fable 5.1 for identical work
- Kimi K3 is 2.9× cheaper than Fable 5.1
- GLM-5.3 is 2.2× cheaper than Kimi K3
- GLM-5.3-Flash is 58× cheaper than Fable 5.1, at 57.5 on the intelligence index against Fable 5’s 62.1
Scaling to a working month
Quattro compiti di questo tipo al giorno, venti giorni lavorativi — 80 cicli di azione:
| Modello | Monthly token cost |
|---|---|
| Claude Fable 5.1 | ≈ $955 |
| Kimi K3 | ≈ $326 |
| GLM-5.3 | ≈ $148 |
| GLM-5.3-Flash | ≈ $17 |
Three honest caveats on these figures
Cache writes are excluded. Fable 5.1 charges $12.50/M for 5-minute cache writes, and a long session writes cache repeatedly. Including writes moves Fable’s real number up, not down. The 80% hit rate is also optimistic for short sessions.
Reasoning tokens are billed output tokens. All three models have always-on thinking. GLM-5.3 in particular is documented as verbose — 170M output tokens against a 72M class median across the AA suite — so its real-world output volume runs above what a naive turn count suggests. The $0.68-per-task figure already accounts for this, which is why I trust it more than my own model above.
Fable 5.1’s tokenizer produces ~30% more tokens than older Claude models for the same text. If you are comparing against a historical Claude bill, adjust.
Even after all three corrections, the ordering does not change and the magnitude barely does. Closed-frontier work costs roughly six times what the best open-weight model costs, per completed task.
Mostra immagine Lo stesso compito agentico da 60 turni al prezzo di listino di ciascun fornitore. L’asse verticale rappresenta l’argomento completo per i pesi aperti.
Licences: What “Open” Actually Buys You In 2026
This is the section I would most like people to read, because the vocabulary has drifted badly and it is going to cost somebody a lot of money.
Neither Kimi K3 nor GLM-5.3 is open source. Both publish downloadable weights under bespoke licences with revenue-triggered conditions. That is a meaningfully different thing from MIT or Apache 2.0, and the difference lands squarely on the kind of business that would most benefit from self-hosting.
The Kimi K3 License
You will see “Modified MIT” repeated all over the internet for K3. It is wrong — that described K2. K3 ships under a custom document with two distinct gates:
The Model-as-a-Service gate. If you provide third parties with access to model inference or fine-tuning — where those third parties control inputs, parameters or training data — and “the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars” over any consecutive 12 months, you must negotiate a separate commercial agreement with Moonshot. Note aggregate revenue, across all affiliates, not revenue attributable to K3. A €25M-turnover consultancy that resells inference is over the line even if the AI business is a rounding error.
The attribution gate. Products exceeding 100 million monthly active users or $20 million monthly revenue must display “Kimi K3” prominently in the product interface.
The exemption that matters most: purely internal use is unrestricted. Running K3 to support your own developers, researchers, legal team or general employee productivity does not trigger either gate. For the overwhelming majority of businesses — including essentially all of mine — that is the relevant clause, and the answer is that you are fine.
The clause that is new relative to K2 is the $20M MaaS gate. K2 only required attribution above its thresholds. K3 adds a much lower revenue gate aimed specifically at commercial inference resellers. Moonshot is not trying to stop you using the model; it is trying to stop cloud providers building a business on it for free.
The glm-5.3 License
Z.ai’s flagship went a different direction, and it is a bigger break with precedent than most coverage acknowledged.
GLM-5.2 shipped under MIT. GLM-5.3-Flash shipped under MIT. GLM-5.3 did not.
The custom licence permits individuals and ordinary businesses to run, deploy and fine-tune the model with no additional restrictions. The single gate is aimed at hyperscalers: an entity whose “aggregate revenue of the Licensee and its affiliates exceeds 10 billion US dollars (or the equivalent in other currencies) in total over any consecutive 12 months” mosto pass Z.AI’s security review before using the model or its derivatives for any commercial purpose.
That is a $10 billion threshold. It excludes almost everyone. But two details in the fine print deserve attention:
First, the licence explicitly carves out app developers embedding the model in features, and resellers merely routing requests to other hosts. The target is narrow and clearly labelled: companies that want to host GLM-5.3 commercially at scale.
Second — and this is the genuinely awkward part — the licence publishes no criteria, timeline or appeal process for that security review. It states only that its “scope and method… shall be reasonably determined by Z.AI.” If you are a large enterprise, that is an unbounded dependency on a foreign vendor’s discretion, and your procurement team will say so.
Also worth noting: despite the two-week cyber-capability hold that preceded release, the licence contains no acceptable-use restrictions, no cyber carve-outs and no output-ownership claims. The safety concern shaped the release timing, not the release terms.
So what is actually still MIT?
GLM-5.3-Flash. 320B parameters, 18B active, natively multimodal, MIT-licensed, 57.5 on the intelligence index, and $0.15/$0.50 per million tokens — currently half that under a promotional discount running to 9 September 2026. If your requirement is genuine unrestricted permissive licensing rather than merely downloadable weights, that is where the frontier currently sits, and it is remarkably high.
The practical summary
| Kimi K3 | GLM-5.3 | GLM-5.3-Flash | Fable 5.1 | |
|---|---|---|---|---|
| Weights downloadable | Sì | Sì | Sì | No |
| OSI-style open source | No | No | Sì (MIT) | No |
| Restriction trigger | $20M revenue (MaaS) | $10B revenue (MaaS) | Nessuno | N/D |
| Attribution required | Above 100M MAU / $20M-month | Copyright notice only | Copyright notice only | N/D |
| Internal use restricted | No | No | No | N/D |
| Fine-tuning permitted | Sì | Sì | Sì | No |
| Acceptable-use policy | Standard | None in licence text | Nessuno | Anthropic AUP applies |
For a typical German Mittelstand business, an agency, or a consultancy running models internally or embedding them in a product: all three open options are usable without negotiating anything. The gates exist to catch AWS, not you. But “read the licence” has gone from pedantic advice to actual advice, and if you are above about €20M turnover and reselling inference in any form, get it in front of counsel before you deploy.
The Real Case For Open Models
I want to make this argument properly rather than as a slogan, because “open models are great” is not an argument and the actual reasons are more interesting than the cheerleading.
1. The price collapse is not a subsidy, it is structural
The instinctive objection to cheap Chinese models is that the pricing is loss-leading and will normalise upward. Some of it certainly is. But the structural point stands independently: when weights are public, inference becomes a commodity market. Twenty hosting providers compete to serve the same model, and none of them can charge a capability premium because none of them has a capability moat. That is why GLM-5.3 runs at $1.40/$4.40 and Fable 5.1 runs at $10/$50 for scores three points apart.
Closed-model pricing includes the R&D amortisation, the margin, and the fact that there is exactly one seller. Open-weight pricing includes electricity, hardware amortisation and a thin margin. Those are different businesses, and the gap will not close by open models getting more expensive.
The evidence that this is real rather than temporary: DeepSeek captured roughly 17% of token usage on Vercel by May 2026 while holding about 1% of revenue share. That is not a pricing anomaly. That is what commoditisation looks like on a chart.
2. Exit rights change your negotiating position even if you never exercise them
Most teams reading this will not self-host. The hardware section below explains why — Kimi K3 needs a rack. But the option has value regardless.
If Anthropic raises prices, deprecates the model you built on, changes its acceptable-use policy in a way that breaks your product, or has a bad quarter, your recourse with a closed model is a migration project. With an open-weight model your recourse is downloading a file you could have downloaded any time. You may never do it. The fact that you could is what makes the vendor relationship a commercial one rather than a dependency.
This is not theoretical. Fable 5.1 shipped with three breaking changes on 1 September, one of which invalidates thinking blocks when you edit conversation history for accounts created after 31 August. If that pattern is load-bearing in your architecture, you are re-engineering on Anthropic’s schedule, not yours.
3. Data residency stops being a project
For anyone operating under GDPR, this is the argument that actually closes deals.
With a closed API, every prompt leaves your infrastructure. You need a processor agreement, a transfer impact assessment if the processor is outside the EEA, a data-flow map, and an answer for your DPO about what happens to the data at rest. All of that is doable. It is also weeks of work per vendor, repeated whenever the vendor changes anything.
With open weights on infrastructure you control — your own hardware, or a German or EU GPU host — the cross-border transfer question does not arise, because there is no transfer. The model runs where your data already lives. Legal review goes from a transfer assessment to a licence read.
For German and EU clients this is frequently the deciding factor, and it is worth being precise about the caveat: using GLM-5.3 through Z.ai’s API does not give you this. You get it from running the weights yourself, or from a provider hosting them in your jurisdiction. The licence makes that legal; it does not make it automatic.
4. Inspectability is real, if underused
Open weights mean you can examine the architecture, run your own evaluations on the actual model rather than an endpoint that might be silently updated, quantise it to fit your hardware, fine-tune it on your domain, and — importantly — pin a version forever.
That last one is underrated. Closed APIs get updated behind stable model IDs. Behaviour drifts. Prompts that worked stop working, and you cannot diff the change because you cannot see it. A local checkpoint is byte-identical in a year’s time.
5. The competitive pressure benefits everyone, including closed-model users
Look at what happened this summer. Kimi K3 lands in July at $3/$15 with frontier-adjacent scores. GLM-5.3 lands in August at $1.40/$4.40 within half a point of it. And on 1 September, Anthropic cut cache reads by 75%.
I am not claiming direct causation — Anthropic does not publish its pricing rationale and cache-read cuts are a natural optimisation. But a market with credible cheap substitutes prices differently from one without them, and the substitutes got credible this year. If you use closed models exclusively, the open-weight ecosystem is still quietly making your bills smaller.
6. Post-training is where the gains are, and it is cheap
GLM-5.3’s headline result — a six-fold Terminal-Bench improvement from the same base weights — is the most strategically important number in this entire article, and it is not about GLM.
It says the expensive part of building a frontier model (pretraining) is increasingly a solved, commoditised input, and the differentiating part (post-training, RL environments, long-horizon task design) is comparatively affordable. That is why four Chinese labs with a combined valuation of about $159 billion are keeping pace with American labs valued at multiples of that. Lambert’s assessment is that Chinese labs are running with “orders of magnitude less capital.”
For anyone downstream, the implication is straightforward: the number of credible model vendors is going up, not down. Plan your architecture accordingly.
7. The ecosystem effects compound
Qwen passed one billion cumulative downloads on Hugging Face, overtaking Llama, and now anchors over 200,000 tagged models and 113,000+ derivatives — roughly 40% of all new LLM derivatives on the platform are Qwen-based. That is a tooling, quantisation, fine-tuning and deployment ecosystem that exists only because the weights are public.
You benefit from that ecosystem whether or not you contribute to it. GGUF conversions, vLLM kernels, LoRA adapters, quantisation recipes, evaluation harnesses — all of it exists because thousands of people could get their hands on the actual weights.
8. It is where the developers already went
Chinese-origin models at ~61% of OpenRouter tokens. US model share down from ~70% to ~30% in a year. Four of the five most-used models on the router are Chinese-origin. Xiaomi’s MiMo models alone account for roughly 21% of routed tokens and about 22% of all coding traffic.
Developer behaviour is a leading indicator. It was a leading indicator for Docker, for Postgres, for Linux. Betting against it has a poor historical record.

…And The Honest Case Against
If I only wrote the section above, this would be an advert.
Open weights are not open source, and the vocabulary drift is a real problem. Two of the three models here ship under bespoke revenue-gated licences with no OSI approval. GLM-5.3’s security-review clause has no published criteria or appeal process. If your compliance framework requires OSI-approved licensing, the honest answer is that your options are GLM-5.3-Flash, DeepSeek’s MIT-licensed flagships, and a shrinking list of others.
Vendor benchmarks remain vendor benchmarks. Every headline number Moonshot and Z.ai published is self-reported and self-selected. Where neutral aggregators have measured, the numbers broadly hold up — which is genuinely to both labs’ credit — but “broadly hold up” is doing work in that sentence.
Self-hosting is a fantasy for most teams. Kimi K3 is 1.5 TB and Moonshot recommends 64+ accelerators. GLM-5.3 needs 8×H200 at FP8 as a bare minimum. The exit right is real; the exercise cost is a data centre.
Support is what you make it. When Fable 5.1 breaks, there is a company with an SLA. When your self-hosted GLM-5.3 deployment produces garbage at 300K context on a Friday night, there is a GitHub issue and your own competence.
Geopolitics is a real procurement input. All three open-weight options discussed here come from Chinese labs. For some clients — public sector, defence-adjacent, certain regulated industries — that is a hard blocker regardless of licence terms or where the weights run. It is not my job to tell you whether that concern is well-founded. It is my job to tell you it will come up in the meeting.
The closed models still lead on general knowledge work. GDPval-AA v2 has Fable 5.1 at 1,853 against GLM-5.3’s 1,769 and Kimi K3’s 1,687. On coding the gap is gone. On broad professional knowledge work it is not.
Mostra immagine Pesi scaricabili, tre diversi set di corde. Solo GLM-5.3-Flash è ancora sotto licenza MIT.
Ollama In 2026: The Pricing Change That Actually Matters
Ollama has quietly become the most important piece of infrastructure in this conversation, and on 31 August 2026 it changed how it charges in a way that is worth understanding properly.
What changed
Ollama moved its Pro, Max and Team plans from GPU-time billing to industry-standard per-token pricing, with a pool of usage credits included in every plan. The company’s stated reason is refreshingly concrete: models like Kimi K3 made GPU-time metrics impossible to predict. When one model activates 104B parameters and another activates 3B, “an hour of GPU” stops meaning anything to the person paying.
The new plans
| Plan | Prezzo | Included monthly usage | Concurrency |
|---|---|---|---|
| Gratuito | $0 | Small monthly credit, starter models | 1 request |
| Pro | $20/mo (or $200/yr) | $60 of usage | 3 requests |
| Max | $100/mo | $300 of usage | 10 requests |
| Team | $500/mo | $1,000 shared, unlimited users | 10 requests |
| Impresa | Custom | Custom | Custom |
Read the Pro row again. $20 buys $60 of tokens. That is a 3× multiplier on included usage, and when the pool runs out you continue at exactly the same published per-token rate — no penalty tier, no service fee.
What was removed
This is the part developers actually noticed: no 5-hour resets and no weekly caps. Anyone who has hit a rolling usage window mid-refactor will understand why that mattered more than the price. The monthly pool refreshes on your subscription date and, importantly, does not roll over — so size your plan to your normal month, not your busiest one.
The terms that matter for EU work
Three commitments in the announcement are directly relevant if you are handling client data:
- Zero data retention. Prompts and responses are never logged and never trained on.
- Hosting in the US and Europe, plus Singapore for a limited set of Qwen models.
- Per-request cost visibility in your account — you can see exactly what each call cost.
For a German consultancy that is a materially better compliance story than most gateway providers offer, though “hosted in Europe” is a routing statement rather than a contractual data-residency guarantee. If residency is a hard requirement rather than a preference, ask Ollama for it in writing before you assume it.
The actual per-token rates
This is where it gets interesting. Ollama publishes rates per model, and they track first-party pricing closely:
| Model on Ollama | Ingresso / 1M | Cached / 1M | Uscita / 1M |
|---|---|---|---|
kimi-k3 | $3.00 | $0.30 | $15.00 |
glm-5.3 | $1.40 | $0.26 | $4.40 |
mistral-large-3 | $0.50 | $0.50 | $1.50 |
gemma4 | $0.14 | $0.05 | $0.40 |
nemotron-3-super | $0.015 | $0.015 | $0.60 |
Run the monthly maths from earlier through this. Eighty agentic runs a month on GLM-5.3 costs about $148 in tokens. A Max plan at $100/month covers $300 of usage — so the same workload that would cost roughly $955/month on Fable 5.1 fits comfortably inside a $100 Ollama subscription with headroom to spare.
That is the entire open-model economic argument compressed into one line on an invoice.
Mostra immagine Eighty agentic runs a month on GLM-5.3 fit inside a $100 Ollama Max plan with change. The same work on Fable 5.1 is a $955 invoice.
What else Ollama shipped in 2026
The pricing change did not happen in isolation. The last few months have been busy:
- 25 August — Claude Desktop support. Ollama now works as a third-party gateway provider for Claude Desktop, so you can drive open models through Anthropic’s own client.
- 26 August — v0.33.0. The Claude Desktop gateway integration landed, prefill recovery was restored to the cache, and Ollama disabled Claude Code’s token-countdown system message, which was invalidating the KV cache on every turn. That last one is a quiet but significant performance fix.
- 26 August — v0.33.1. MLX support for Qwen3.8 Flash Next, structured output, and a fix for GPU timeouts when loading models from slower storage.
- 28 August — v0.33.2. Dark mode restored, macOS handoff fixed, and Claude Desktop proxy requests kept alive during model catalogue updates.
- 20 August — v0.32.15. New desktop onboarding, and metadata caching between requests that cut time-to-first-token by roughly half.
- 11 August — NVIDIA Nemotron 3.5 Lightning, a 30B model tuned for agentic workflows with tool calling on personal hardware.
- 10 August — Meta’s Muse Glimmer, a 30B multimodal model under Apache 2.0, accelerated by Ollama’s MLX engine. Worth noting given Llama’s collapse in routed usage — Meta is still shipping genuinely permissive weights.
- 9 July — $88M funding round, with Ollama reporting 8.9 million developers.
- 29 June — Gemma 4 on MLX up to 90% faster via multi-token prediction, which disproportionately benefits coding agents on Apple Silicon.
- 5 June — v0.30 added GGUF compatibility through llama.cpp, broadening hardware support well beyond Apple Silicon.
The through-line is that Ollama has stopped being “the easy way to run a small model on your laptop” and become a general-purpose gateway that happens to also run models locally. The cloud catalogue now includes kimi-k3:cloud, glm-5.3:cloud, glm-5.3-flash, deepseek-v4-pro, deepseek-v4-flash, minimax-m3, qwen3.5 across seven sizes, gpt-oss at 20b and 120b, gemma4, the Nemotron 3 family and mistral-large-3.
Setting It All Up: Copy-Paste Configs
Ollama + Claude Code (the fastest path)
Ollama ships a one-command launcher that handles the environment wiring for you:
bash
ollama launch claude
If you would rather do it manually — which you should if you are scripting it — install Claude Code first:
bash
# macOS / Linux
curl -fsSL https://claude.ai/install.sh | bash
# Windows (PowerShell)
irm https://claude.ai/install.ps1 | iex
Then point it at Ollama:
bash
export ANTHROPIC_AUTH_TOKEN=ollama
export ANTHROPIC_API_KEY=""
export ANTHROPIC_BASE_URL=http://localhost:11434
And run against whichever model you want:
bash
# A local model
claude --model qwen3.5
# A cloud model — note the :cloud suffix
claude --model kimi-k3:cloud
Or inline, without exporting anything globally:
bash
ANTHROPIC_AUTH_TOKEN=ollama \
ANTHROPIC_API_KEY="" \
ANTHROPIC_BASE_URL=http://localhost:11434 \
claude --model glm-5.3:cloud
Two things that will bite you.
Innanzitutto, do not export ANTHROPIC_BASE_URL permanently in your shell profile. It is a global variable and it will silently redirect every Anthropic-speaking tool on your machine to Ollama, including the one you wanted talking to Anthropic. Scope it per-project or per-invocation.
In secondo luogo, set your context length to 64k or higher for anything working on a real repository. Ollama’s default is smaller, and a coding agent that quietly runs out of window will just start forgetting things rather than erroring.
For CI, Docker or scripted runs, --yes skips the interactive prompts:
bash
ollama launch claude --model glm-5.3:cloud --yes -- -p "how does this repository work?"
GLM-5.3 direct from Z.ai
If you want first-party routing rather than going through Ollama, Z.ai exposes three protocol endpoints:
| Protocollo | URL di base |
|---|---|
| Messaggi antropici | https://api.z.ai/api/anthropic |
| Completamenti della chat di OpenAI | https://api.z.ai/api/coding/paas/v4 |
| Risposte di OpenAI | https://api.z.ai/api/v1 |
An OpenCode provider block for GLM-5.3:
json
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"zai": {
"npm": "@ai-sdk/openai-compatible",
"name": "Z.AI",
"options": {
"baseURL": "https://api.z.ai/api/coding/paas/v4",
"apiKey": "{env:ZAI_API_KEY}"
},
"models": {
"glm-5.3": {
"name": "GLM-5.3",
"limit": { "context": 1000000, "output": 128000 },
"options": { "reasoning_effort": "high", "temperature": 1 }
}
}
}
},
"model": "zai/glm-5.3"
}
Important upgrade note: GLM-5.3 requires reasoning to be enabled. If your existing configuration sets thinking.type: "disabled", that will now fail. Change it to "enabled" and set reasoning_effort: "low" if you want the old latency profile back.
sforzo_di_ragionamento is your main quality-versus-cost dial: basso for latency-sensitive calls, alto for substantial work, max (the default) for problems that have genuinely defeated alto. Given the documented verbosity, I would run alto as an everyday setting rather than leaving max on and wondering where the output tokens went.
The routing setup I would actually recommend
None of this is a “pick one” decision. The correct answer for most teams is a three-tier routing table, and every serious agent harness supports per-agent models.
json
{
"$schema": "https://opencode.ai/config.json",
"model": "zai/glm-5.3",
"small_model": "ollama/gemma4",
"agent": {
"plan": {
"description": "Architecture, root-cause analysis, anything expensive to get wrong",
"model": "anthropic/claude-fable-5-1"
},
"build": {
"description": "The bulk of implementation work",
"model": "zai/glm-5.3",
"options": { "reasoning_effort": "high" }
},
"research": {
"description": "Long-context reading, web research, multi-source synthesis",
"model": "ollama/kimi-k3:cloud"
},
"grunt": {
"description": "Renames, docstrings, test scaffolding, lint fixes",
"model": "ollama/glm-5.3-flash"
}
}
}
The reasoning behind each line:
- Planning gets Fable 5.1. Planning is where a bad decision costs you the most downstream tokens, and long-horizon coherence is precisely what 5.1 was built for. It is also where you spend the fewest tokens, so the price premium applies to the smallest slice of your bill.
- Building gets GLM-5.3. Best capability-per-euro on the market, and the bulk of your token spend lands here.
- Research gets Kimi K3. BrowseComp 91.2 and a genuine 1M context make it the best available option for reading a lot of things and synthesising them.
- Mechanical work gets GLM-5.3-Flash. At $0.15/$0.50 it is effectively free, and renaming a symbol across 40 files does not need frontier reasoning.
modello_piccologets something tiny for the harness’s own housekeeping — title generation, summarisation, internal utility calls.
You are not choosing a winner. You are building a gearbox. And write the model IDs so that swapping one is a one-line change, because in this market you will be making that change again within the quarter.
What You Can Actually Run Locally (The Hardware Reality)
Let me kill an assumption before it costs somebody money.
You cannot run Kimi K3 or GLM-5.3 on a workstation. Not with a 5090. Not with two.
Kimi K3
- ~1.5 TB of weights in native MXFP4 — and remember, that è the quantised checkpoint, not a starting point for further compression
- Moonshot recommends a minimum of 64 accelerators for competitive serving
- Realistic self-hosting starts at multi-node clusters; reference deployments use GB300 NVL72 racks
- There is no official Ollama library entry for local Kimi K3 and no consumer GGUF conversion worth pointing you at
- Reports of it running on clusters of consumer RTX 5090s exist, but “a cluster of 5090s” is not a laptop and the throughput is not comparable
GLM-5.3
- ~1.5 TB in BF16, roughly 750 GB at FP8
- Minimum viable single node: 8× H200 (1,128 GB of GPU memory) at FP8, leaving around 375 GB for KV cache
- BF16 needs two 8×H200 nodes, or a single 8×B300 node (2,304 GB)
- A single 8×H200 node in BF16 is too small for the weights plus cache
Mostra immagine The frontier open models need a rack. The 24 GB tier is where “runs on my machine” actually lives — and it has got very good.
So what does “local AI” actually mean in September 2026?
It means a different tier of model, and that tier has got genuinely good:
| Modello | VRAM | What it is for |
|---|---|---|
| Qwen3.8-27B | 24 GB at Q4 | Best all-rounder on consumer hardware — 61.7% SWE-bench |
| gpt-oss:20b | 16 GB | Best small model, adjustable reasoning effort |
| Gemma 4 E4B | ~6 GB | Vision plus tool calling on a modern laptop |
| Mistral 7B | 8 GB | Fastest general-purpose option, 40–60 tok/sec |
| DeepSeek-R1 7B | 5 GB | Chain-of-thought reasoning on a laptop GPU |
| Llama 4 Scout | ~55 GB at Q4 | 10M context, multimodal — workstation territory |
A Qwen3.8-27B scoring 61.7% on SWE-bench, running entirely on a 24 GB consumer GPU with no network connection, is a remarkable thing that would have sounded like science fiction eighteen months ago. It is not GLM-5.3 and it is not pretending to be.
The honest framing: open weights at the frontier buy you sovereignty and price, not local execution. Open weights in the 7B–30B range buy you genuine local execution, at a real but acceptable capability cost, and that is the tier where “runs on my machine, sees no network” is an achievable requirement.
The architecture that works for most of my clients is exactly that split: a small local model for anything touching genuinely sensitive data, and a hosted open-weight frontier model for everything else — with the weights available as insurance rather than as a deployment plan.
Dove si rompe effettivamente ogni modello
No hype. Here are the honest weaknesses.
Kimi K3 weaknesses
It is expensive for an open model. $3/$15 is 2.1× GLM-5.3’s input rate and 3.4× its output rate for essentially the same intelligence index score. If you are choosing K3 over GLM-5.3, be clear about what you are buying with that premium — usually it is the vision stack or the browsing performance.
The licence has the lower gate. A $20M aggregate-revenue MaaS threshold catches far more organisations than GLM-5.3’s $10B. If inference resale is anywhere in your business model, K3 is the more constrained of the two.
Self-hosting is out of reach for almost everybody. 1.5 TB and 64+ accelerators is a serious infrastructure commitment. The exit right is more theoretical here than with any other model in this comparison.
Broad knowledge work trails. GDPval-AA v2 at 1,687 puts it behind GLM-5.3, both Claude Fables and GPT-5.6 Sol Max. It is a coding and agentic specialist that happens to be enormous.
GLM-5.3 weaknesses
Text only. No image input, no video. If your workflow includes screenshot debugging, design-to-code or document vision, GLM-5.3 simply cannot do it and you need GLM-5.3-Flash, Kimi K3 or Fable 5.1 instead. This is the single most common configuration mistake I expect people to make, because the model IDs look related and are not.
It is verbose, and verbosity is billed. 170M output tokens against a 72M class median across the AA suite. sforzo_di_ragionamento il valore predefinito è max, and thinking cannot be turned off. Budget for more output tokens than your turn count implies.
The licence is no longer MIT, and the security-review clause is unbounded. For most readers the $10B threshold makes this academic. For anyone near it, “scope and method shall be reasonably determined by Z.AI” is not a clause your legal team will enjoy.
Terminal-Bench 3.0 at 28.3 trails GPT-5.6 Sol’s 34.6. On the specific benchmark closest to terminal-native agentic coding, it is behind the closed competition on the same ruler.
It is new to open weights. Released 28 August. The community has had days, not months. Long-tail deployment bugs have not surfaced yet.
Claude Fable 5.1 weaknesses
The price. Roughly 6.5× GLM-5.3 per completed agentic task, even after the 75% cache-read cut. For most work that gap is not defensible on capability grounds any more.
Three breaking changes. Forced tool use returns 400. Thinking blocks do not travel backwards to older models. Editing conversation history invalidates thinking blocks on accounts created after 31 August 2026. Any of these can break a working integration on upgrade.
Parallel tool calling regressed. Reports of one call per turn where Fable 5 batched several. On a long agent loop that is wall-clock time you are paying for twice.
Zero exit optionality. No weights, no self-hosting, no version pinning beyond what Anthropic offers, no inspection. When it changes, you adapt.
The tokenizer inflates comparisons. ~30% more tokens than older Claude models for the same text, which makes historical cost comparisons misleading in Anthropic’s favour if you are not careful.
The Decision Framework: Six Scenarios
Mostra immagine Six scenarios, six answers. Find the row that sounds like your week.
1. “I want one model. Set it, forget it, keep the bill sane.”
GLM-5.3. Within half a point of Kimi K3 and roughly three points of the best closed model on the neutral index, at $1.40/$4.40. Run it through Ollama on a Pro or Max plan, set sforzo_di_ragionamento: elevato, and get on with your work. The only thing that should push you off this answer is needing image input.
2. “My work is visual — UI, design-to-code, screenshot debugging, documents.”
Kimi K3 or GLM-5.3-Flash, not GLM-5.3. K3’s MoonViT-V2 handles text, images and video natively and scores 81.6/83.4 on MMMU-Pro. GLM-5.3-Flash is the budget option with native multimodality and an MIT licence. GLM-5.3 is text-only and will simply refuse the input.
3. “Long autonomous sessions where being wrong is expensive.”
Claude Fable 5.1. This is what it was built for and the benchmarks back it: Terminal-Bench-Science doubled, 82% on Browserbase’s hardest computer-use tasks against Opus 5’s 74%, and the largest published gains on multi-hour agentic work. The 75% cache-read cut makes exactly this workload up to 45% cheaper than it was. Pay the premium where a mistake costs more than the tokens.
4. “EU data residency is a hard requirement.”
GLM-5.3-Flash if you need MIT, GLM-5.3 if you need capability. Self-host on your own hardware or an EU GPU provider and the cross-border transfer question stops existing. Budget 8×H200 for GLM-5.3 at FP8; Flash is far more tractable at 320B/18B. If self-hosting is out of budget, Ollama’s Europe hosting with zero data retention is the pragmatic middle ground — but get the residency commitment in writing rather than inferring it from a marketing page.
5. “Small team, tight budget, coding all day.”
GLM-5.3-Flash as default, GLM-5.3 for hard problems, Ollama Pro at $20. Flash costs $0.15/$0.50 — currently half that until 9 September — and scores 57.5 on the intelligence index. Twenty dollars buys sixty dollars of tokens with no weekly caps. For a two-to-four person team this is close to unbeatable. The GLM Coding Plan at $18/month (Lite) is the alternative if you prefer a fixed quota to a credit pool.
6. “I need to justify this to a procurement or legal team.”
GLM-5.3-Flash. It is the only model in this comparison under a standard OSI-approved licence (MIT), it is natively multimodal, it scores 57.5 on the neutral index, and the weights are on Hugging Face with no revenue gates, no attribution mandates and no security-review clause. When the question is “what can we defend in a contract review,” permissive licensing beats three points of benchmark every time.
Cosa terrò d’occhio nel prossimo trimestre
Whether Fable 5.1 lands above 62.1 on the Artificial Analysis index. It launched the day before this article and has not been independently scored. Its benchmark deltas over Fable 5 suggest it should, but “should” is not “did,” and the gap to Kimi K3’s 59.7 is the number the whole open-versus-closed argument turns on.
Whether the licence drift continues. In eight weeks we went from GLM-5.2 under MIT to GLM-5.3 under a bespoke licence with a discretionary security review, and from Kimi K2’s modified-MIT to K3’s revenue-gated terms. If GLM-6 and K4 tighten further, “open weights” becomes a marketing term rather than a meaningful category. GLM-5.3-Flash staying MIT is the counter-signal worth tracking.
Whether other labs adopt Z.ai’s staged-release pattern. A two-week hold with a published safety rationale is a new norm. If it holds, it is a good one. If it becomes a reason weights ship later and later, it is a soft path to not shipping them at all.
Independent replication of the cyber capability claims. 2,436 vulnerabilities across 269 projects is an extraordinary number and it is entirely self-reported. Somebody neutral needs to check it, because if it is accurate it reframes the entire open-weights safety conversation, and if it is not, it was effective marketing.
Ollama’s per-token rates six months from now. $20 for $60 of usage is an aggressive introductory posture from a company that raised $88M in July. Whether those multipliers survive contact with real unit economics is the single biggest variable in the “cheap open models” thesis for small teams.
Whether local models close on the 30B tier. Qwen3.8-27B at 61.7% SWE-bench on 24 GB is the most under-discussed result of the year. The frontier gets the headlines; the 24 GB tier is what changes what an ordinary business can do without an API key.
Domande frequenti
Is Kimi K3 really open source? No. Kimi K3’s weights are freely downloadable from Hugging Face, but under a custom “Kimi K3 License,” not an OSI-approved open-source licence. Model-as-a-Service operators whose aggregate revenue exceeds $20 million over any consecutive 12 months must negotiate a separate commercial agreement with Moonshot, and products above 100 million monthly active users or $20 million monthly revenue must display “Kimi K3” in their interface. Purely internal use is unrestricted.
Is GLM-5.3 better than Kimi K3? They are effectively tied on the neutral aggregate index — 59.5 against 59.7 on Artificial Analysis. GLM-5.3 is 2.2× cheaper per agentic task and scores higher on knowledge work (GDPval-AA v2: 1,769 vs 1,687). Kimi K3 is natively multimodal, leads on agentic browsing (BrowseComp 91.2) and took first place on Arena.AI’s frontend code arena. Choose GLM-5.3 for cost and text-based coding; Kimi K3 for vision, browsing and long-horizon research.
How much cheaper are open models than Claude Fable 5.1? On a modelled 60-turn agentic task, GLM-5.3 costs about $1.85 against Fable 5.1’s $11.94 — roughly 6.5× cheaper. Kimi K3 costs about $4.07, roughly 2.9× cheaper. GLM-5.3-Flash costs about $0.21, roughly 58× cheaper. Over 80 runs a month that is $148 versus $955.
Can I run Kimi K3 or GLM-5.3 locally? Not on consumer hardware. Kimi K3 is roughly 1.5 TB even in its native MXFP4 format and Moonshot recommends 64+ accelerators. GLM-5.3 needs a minimum of 8×H200 (1,128 GB) at FP8. For genuine local execution, look at Qwen3.8-27B (24 GB at Q4), gpt-oss:20b (16 GB) or Gemma 4 E4B (~6 GB).
What changed in Claude Fable 5.1? Released 1 September 2026. Cache reads dropped 75% from $1.00 to $0.25 per million tokens, making typical workloads about 25% cheaper and agentic workloads up to 45% cheaper; input and output stayed at $10/$50. Terminal-Bench 4.0 rose from 42.0% to 55.8%, Terminal-Bench-Science from 24.7% to 52.6%, AutomationBench from 17.1% to 31.4%. Three breaking changes affect forced tool use, thinking-block portability and history editing.
What is Ollama’s new pricing? From 31 August 2026, Pro, Max and Team plans use per-token pricing with included credits: Pro $20/month for $60 of usage, Max $100 for $300, Team $500 for $1,000 shared across unlimited users. The 5-hour and weekly caps were removed entirely, there are no service fees, credits do not roll over, and all plans carry zero data retention with hosting in the US and Europe.
Which of these models is genuinely MIT-licensed? Only GLM-5.3-Flash — a separate 320B/18B natively multimodal model released 26 August 2026, scoring 57.5 on the Artificial Analysis index at $0.15/$0.50 per million tokens. GLM-5.3 and Kimi K3 both use bespoke revenue-gated licences; Claude Fable 5.1 is fully proprietary.
Why can’t I compare these models on Terminal-Bench? Because they were each evaluated on a different version. Kimi K3’s 88.3 is on Terminal-Bench 2.1, GLM-5.3’s 28.3 is on 3.0, and Fable 5.1’s 55.8 is on 4.0. Each version is substantially harder than the last, so the numbers are not on the same scale. Compare within a version only. On Terminal-Bench 2.1, for example, Kimi K3 scores 88.3 against Claude Fable 5’s 88.0, GPT-5.6 Sol’s 88.8 and GLM-5.3-Flash’s 84.3 — that comparison is valid, and it is the one worth quoting.
Should I use Ollama or go direct to the vendor? Ollama if you want one billing relationship, one API surface, easy model switching and EU/US hosting with zero data retention — its per-token rates track first-party pricing closely. Direct if you need first-party features like Z.ai’s sforzo_di_ragionamento controls at full fidelity, vendor SLAs, or subscription plans such as the GLM Coding Plan. Many teams run both and route by workload.
Does GLM-5.3 support images? No. GLM-5.3 is text-only. GLM-5.3-Flash — a completely different model despite the similar name — is natively multimodal and handles text, image, video and file input. Sending images to the wrong model ID is the most common early mistake with the GLM family.
In sintesi
Twelve months ago the open-weight question was whether these models were usable. Six months ago it was whether they were competitive. In September 2026 it is genuinely: what are you still paying a closed-model premium for?
There is a real answer to that question, and it is narrower than it used to be. Claude Fable 5.1 leads on the hardest sustained agentic work, on general knowledge work, and on the class of debugging where a model needs to hold a messy problem in its head for hours without drifting. Anthropic’s 75% cache-read cut targets exactly that workload. If your failure cost exceeds your token cost, that premium is rational.
For everything else, the maths has moved. GLM-5.3 delivers 94% of Claude Opus 5’s index score at 29% of the cost. Kimi K3 scores 88.3 on Terminal-Bench 2.1 against Claude Fable 5’s 88.0 on the same version, and took first place on LMArena’s Frontend Code Arena ahead of both flagship closed models. Ollama will sell you $60 of tokens for $20 and remove the usage caps while doing it. That combination did not exist in the spring.
But do not let the enthusiasm skip the fine print, because there are two of them and both matter.
The first is licensing. “Open weights” and “open source” have quietly stopped meaning the same thing. Kimi K3 and GLM-5.3 both ship under bespoke, revenue-gated licences. The gates are high enough that most readers are unaffected — but “most readers are unaffected” is not the same as “unrestricted,” and GLM-5.3’s undefined security-review clause is a genuine procurement risk for large enterprises. GLM-5.3-Flash under MIT is the last fully permissive frontier-adjacent option, which is precisely why it deserves more attention than it gets.
The second is that self-hosting is mostly aspirational. 1.5 TB of weights and 64 accelerators is not an exit plan for a mid-sized business. What open weights buy you at this tier is a competitive inference market, price transparency, version pinning, jurisdiction choice, and a negotiating position. Those are worth a great deal. They are not the same as running the thing in your basement.
The setup I would actually build: GLM-5.3 as the default, Fable 5.1 on planning and the genuinely hard problems, Kimi K3 for vision and long-horizon research, GLM-5.3-Flash for the grunt work, and a 27B local model for anything that must never leave the building. Route by workload, not by loyalty. Write your configuration so the model IDs are a one-line change.
Because the only prediction I am confident about is that this article will need updating before Christmas.
Are you running any of these three in production? I am particularly interested in whether GLM-5.3’s verbosity shows up as a real cost problem at scale, and whether anyone has actually put the Kimi K3 licence in front of counsel and got a clear read on the MaaS definition. Get in touch — corrections and counter-evidence welcome, and this article gets updated when the picture changes.
Last updated: 2 September 2026.



