Hace seis semanas habría escrito este artículo de otra manera.
En julio, Moonshot AI publicó en Hugging Face un modelo de vanguardia con 2,8 billones de parámetros y permitió que cualquiera pudiera descargarlo. El 28 de agosto, Z.ai publicó por fin los pesos del GLM-5.3 que llevaba dos semanas reteniendo, tras descubrir que su propio modelo podía encadenar vulnerabilidades de formas que nadie le había pedido. Y el 1 de septiembre, Anthropic lanzó Claude Fable 5.1 y redujo el coste del trabajo de los agentes hasta en 451 TP3T sin modificar el precio de referencia.
Tres modelos. Tres respuestas totalmente diferentes a la pregunta de qué le debe un laboratorio de IA a la sociedad. Y, por primera vez desde que empezó todo esto, las opciones de «ponderación abierta» no son la solución de compromiso.
Esa es la historia real de finales de 2026, y da para unas ocho mil palabras, porque es en los detalles donde residen todo el dinero y todos los errores.
He recopilado los documentos de arquitectura, las fichas de modelos, los textos de las licencias —con todo detalle, línea por línea, porque dos de los tres no son lo que la gente suele llamar así—, las pruebas de rendimiento de los proveedores, las puntuaciones de los agregadores independientes, las tasas reales por token y la flamante hoja de precios de Ollama. A continuación, he calculado los costes basándome en un mes laboral normal, en lugar de en un escenario de día de lanzamiento.
Esto es lo que ha salido.
En resumen — El veredicto en 60 segundos
GLM-5.3 es, con diferencia, la mejor opción en inteligencia artificial en este momento. Según Artificial Analysis, se sitúa a menos de medio punto de Kimi K3 y a unos tres puntos y medio por detrás de Claude Opus 5, con $1,40 de entrada y $4,40 de salida por millón de tokens. En la misma tarea de agente, cuesta aproximadamente una sexta parte de lo que cuesta Fable 5.1. Los pesos se pueden descargar. Si quieres un modelo al que orientar a tu agente y no diriges un negocio con $10 mil millones de ingresos, este es el adecuado.
Kimi K3 es lo mejor que puedes descargar legalmente. 2,8 billones de parámetros, 104 000 millones activos, un contexto auténtico de 1 millón, visión nativa, la mejor puntuación en navegación agencial que se haya publicado hasta la fecha (BrowseComp 91,2) y —la cifra que debería poner fin al argumento de que “los modelos abiertos van por detrás”— 88,3 en Terminal-Bench 2.1, frente a los 88,0 de Claude Fable 5 y los 88,8 de GPT-5.6 Sol en la misma versión. Además, cuesta $3/$15 —más del doble que el GLM-5.3— y sus 1,5 TB de pesos requieren un rack, no una estación de trabajo. Es el modelo que se utiliza cuando la tarea es difícil y larga, y el modelo al que se recurre cuando alguien te dice que los pesos abiertos están una generación por detrás.
El Claude Fable 5.1 sigue siendo el modelo al que se recurre cuando cometer un error sale caro. Lidera el sector en lo que respecta al trabajo agente sostenido durante varias horas y a la laboriosa depuración que nunca queda reflejada de forma clara en una tabla de clasificación. Anthropic redujo las lecturas de caché en 75% el 1 de septiembre, lo que hace que las sesiones largas de los agentes sean mucho más económicas de lo que eran. Además, cuesta seis veces más que GLM-5.3 por tarea, es de acceso cerrado y ahora incluye restricciones contra la destilación que impiden el funcionamiento de algunas integraciones existentes.
La frase concisa y sincera: Los modelos de peso abierto han reducido la brecha de capacidad a algo así como entre tres y cinco meses, y han eliminado por completo la diferencia de precio. Lo que realmente estás comprando a Anthropic en septiembre de 2026 es fiabilidad en las tareas más difíciles de 10%, y el derecho a no tener que preocuparte por nada de esto.
Y el giro que nadie se atreve a decir en voz alta: Ni Kimi K3 ni GLM-5.3 se distribuyen ya bajo una licencia de código abierto. Ambos han pasado este verano a condiciones personalizadas, sujetas a ingresos. Peso abierto, sí. Código abierto, no. Lee la sección sobre licencias antes de que lo haga tu equipo jurídico.

Tabla de comparación rápida
| Kimi K3 | GLM-5.3 | Claude Fable 5.1 | |
|---|---|---|---|
| Fabricante | Moonshot AI | Z.ai (Zhipu AI) | Antropico |
| Publicado | 16 de julio de 2026 (API) · 27 de julio (pesos) | 14 de agosto de 2026 (API) · 28 de agosto (pesos) | 1 de septiembre de 2026 |
| Peso disponible | Sí | Sí | No |
| Licencia | Licencia Kimi K3 (personalizada, con restricciones de ingresos) | Licencia glm-5.3 (personalizada, con restricción en función de los ingresos) | Propietario |
| Parámetros totales | 2,8 T | 753B | No revelado |
| Activo por token | 104B (16 de 896 expertos) | ~40 mil millones (estimación) | No revelado |
| Arquitectura | KDA + residuos de atención, LatentMoE estable, 69 capas de KDA + 24 capas de MLA con compuerta | El modelo MoE, con la misma base que el GLM-5.2, solo obtiene mejoras tras el entrenamiento. | No revelado |
| Ventana de contexto | 1,048,576 | 1,000,000 | 1,000,000 |
| Potencia máxima | 131 072 impagos (hasta 1 048 576) | 128,000 | 128,000 |
| Entrada multimodal | Texto, imagen, vídeo (MoonViT-V2) | Solo texto | Texto, imagen |
| Reflexión | Siempre activo | Siempre activo, esfuerzo_de_razonamiento mín./máx./máx. | Siempre activo, adaptativo, esfuerzo por mensaje (beta) |
| Precio de entrada / 1M | $3.00 | $1.40 | $10.00 |
| Entrada almacenada en caché / 1M | $0.30 | $0.26 | $0.25 |
| Precio de venta / 1 millón | $15.00 | $4.40 | $50.00 |
| Índice de Inteligencia de AA | 59.7 | 59.5 | Aún no se ha puntuado (Fable 5: 62,1) |
| Coste por tarea de evaluación de AA | $0.84 | $0.68 | — (Opus 5: $2.34) |
| Huella de autoalojamiento | ~1,5 TB (MXFP4 nativo) | ~1,5 TB BF16 / ~750 GB FP8 | N/A |
| Sobre Ollama | kimi-k3:cloud | glm-5.3:nube | A través de Claude Code / Claude Desktop como cliente |
Todos los precios son precios de catálogo del fabricante. Las cifras de Artificial Analysis proceden de la instantánea del índice de septiembre de 2026; Fable 5.1 se lanzó el 1 de septiembre y, en el momento de redactar este artículo, aún no había sido evaluado de forma independiente.
Por qué esta comparación a tres bandas es la que realmente importa
Durante dos años, el debate entre «abierto» y «cerrado» se mantuvo en un equilibrio cómodo: los modelos cerrados eran mejores, los abiertos eran más baratos, y cada uno elegía su posición en esa disyuntiva. Todo el mundo sabía cuál era su postura.
Esa forma se rompió este verano.
Nathan Lambert, que realiza un seguimiento de este tema con más detenimiento que casi nadie, sitúa la brecha actual de capacidad entre los modelos de peso libre de vanguardia y los de peso cerrado de vanguardia en de tres a cinco meses — lo que supone un descenso respecto a los seis a nueve meses que se barajaban hace un año. En el momento en que escribió el artículo, Kimi K3 ocupaba el puesto #2 en el índice Vals AI y se situaba cerca de la cima del índice de Artificial Analysis, solo superado por Claude Fable y GPT-5.6 Sol Max. La instantánea del índice de septiembre que utilizo más adelante en este artículo lo sitúa en cuarto lugar, por detrás de Grok 4.6. Sea como sea: eso no es “bueno para un modelo abierto”. Es un puesto en el podio.
Mientras tanto, los datos de uso han dado un giro realmente sorprendente. Los modelos de origen chino han captado aproximadamente 61% de todos los tokens enrutados a través de OpenRouter para mayo de 2026. La cuota estadounidense de ese tráfico se desplomó de unos 70% a unos 30% en doce meses. Llama, de Meta —el modelo que inició la ola de los modelos de peso abierto—, cayó por debajo de los 1% de volumen enrutado. La cuota de Google pasó de unos 371 TP3T a 131 TP3T. El propio estudio de 100 billones de tokens realizado por OpenRouter en colaboración con a16z reveló que los modelos de peso abierto representaban aproximadamente un tercio de todo el volumen de tokens de la plataforma.
Se puede discutir sobre qué representa el tráfico de OpenRouter. Sobrerrepresenta a los desarrolladores, a los aficionados y a las cargas de trabajo en las que el coste es un factor determinante, y subestima los contratos empresariales que nunca pasan por un router. De acuerdo. Pero es el mayor conjunto de datos público del que disponemos sobre lo que la gente elige realmente cuando puede elegir libremente, y la tendencia es inequívoca.
Así pues, la pregunta en septiembre de 2026 ya no será “¿son ya lo suficientemente buenos los modelos abiertos?”. Será mucho más concreta:
Si hay tres modelos que cumplen los requisitos, ¿cuál le recomendarías a tu agente? ¿Y cuánto te cuesta realmente equivocarte?
Esa es una cuestión que tiene que ver con los puntos de referencia, los precios, las licencias y la infraestructura, más o menos en ese orden según la importancia que la gente les da, y exactamente en el orden inverso según la importancia que realmente tienen.
Qué es realmente el Kimi K3
Moonshot AI anunció Kimi K3 el 16 de julio de 2026 y publicó todos los pesos el 27 de julio. Se trata, por número de parámetros, del modelo de peso abierto más grande que se haya publicado jamás: 2,8 billones de parámetros en total, aproximadamente 75% más que DeepSeek V4 Pro.
Esa cifra es menos impresionante de lo que parece y más impresionante de lo que parece, en ese orden.
La arquitectura, en un lenguaje sencillo
Kimi K3 es un modelo de mezcla de expertos que activa 104 000 millones de parámetros por token, seleccionando a 16 expertos de un conjunto de 896. Así pues, aunque el total es de 2,8 T, cada paso hacia adelante afecta a menos de 41 TP3T de los pesos. De ahí radica el interés económico: un enorme volumen de conocimiento almacenado y un coste de inferencia moderado.
Lo más interesante desde el punto de vista de la ingeniería está en la pila de atención. K3 utiliza 93 capas formadas por 69 capas KDA (Kimi Delta Attention) y 24 capas Gated MLA — un diseño híbrido que combina atención lineal y atención completa con una relación de intercalación de aproximadamente 3:1. KDA es una variante de la atención lineal que se ejecuta en tiempo lineal en relación con la longitud de la secuencia, en lugar de en tiempo cuadrático, lo que hace que el contexto de 1 millón sea económicamente viable y no solo una promesa publicitaria. Moonshot informa de hasta un 75%: reducción de la caché KV y hasta Rendimiento de decodificación 6 veces mayor con un contexto de 1 M como consecuencia de ello.
Encima hay una capa de Atención a los residuos (AttnRes), que Moonshot describe como un sustituto directo de las conexiones residuales estándar, y un MoE latente estable marco para el enrutamiento. Se afirma que esta combinación ofrece aproximadamente un Mejora de 2,5 veces en la eficiencia global de escalado frente a Kimi K2.
La visión proviene de MoonViT-V2, un codificador de 401M parámetros, y es nativo, no un complemento añadido: K3 gestiona texto, imágenes y vídeo en un solo modelo.
Otro detalle importante para cualquiera que esté pensando en el autoalojamiento: K3 se entrenó con Entrenamiento nativo de MXFP4 con consideración de la cuantificación, con expertos en MXFP4 y activaciones en MXFP8. No se trata de una cuantificación a posteriori. El punto de control publicado es el modelo cuantificado, razón por la cual los 2,8 billones de parámetros se sitúan aproximadamente en 1,5 TB en el disco en lugar de los ~5,6 TB que necesitaría un punto de control de BF16 de ese tamaño.
La historia de referencia
Las cifras publicadas por Moonshot son sólidas en todos los ámbitos y, cosa poco habitual, varias de ellas han resistido el escrutinio de evaluaciones independientes:
- Terminal-Bench 2.1: 88,3 — la puntuación más alta publicada en esa versión de la prueba de rendimiento
- FrontierSWE: 81,2
- DeepSWE: 67,5
- BrowseComp: 91,2 — Estado actual de la investigación sobre la web agentiva
- GPQA Diamond: 93,5
- MMMU-Pro: 81,6 / 83,4 (visión)
- Maratón de SWE: 91,0
- OfficeQA Pro: 81,3
- AutomationBench: 30,8
- LMArena Frontend Code Arena: 1.679 Elo — primer puesto, por delante de Claude Fable 5 (1.631) y GPT-5.6 Sol (1.618), y se ha impuesto en seis de los siete ámbitos de frontend
Esto último merece una reflexión. Un modelo de peso abierto que se puede descargar gratis se alzó con el primer puesto en una competición de programación de interfaces basadas en las preferencias humanas, por delante de los modelos cerrados insignia de los dos laboratorios mejor financiados del mundo. Independientemente de lo que se piense sobre las pruebas comparativas de este tipo como metodología, se trata de una noticia que habría sido impensable en 2025.
En los índices agregados, se sitúa ligeramente por debajo: 59,7 en el Índice de Inteligencia Artificial Analítica, tercero en la clasificación general, por detrás de Claude Opus 5 (63,0) y Claude Fable 5 (62,1). Las propias cifras de Moonshot — GDPval-AA v2 en 1.687 y Maletín AA a 1.527 — lo sitúan en tercer y segundo lugar, respectivamente, en esas evaluaciones del trabajo del conocimiento. En el Índice Vals v2 ocupa el 57.8%, frente a los 67,21 TP3T de Claude Opus 5, los 66,01 TP3T de Claude Fable 5 y los 63,71 TP3T de GPT-5.6 Sol —la mayor diferencia entre ’abierto’ y «cerrado» que he encontrado en ningún índice, y que conviene tener en cuenta a la hora de analizar las cifras de codificación—.

Las demostraciones y cómo interpretarlas
Moonshot publicó dos demostraciones de autonomía a largo plazo que suscitaron gran interés: una Funcionamiento autónomo durante 48 horas completando todo el proceso de diseño de un chip (un diseño de 4 mm con una frecuencia de cierre de 100 MHz, que simula más de 8.700 tokens por segundo) y una reproducción de la astrofísica Relación «I-Love-Q» en unas dos horas, frente a las una o dos semanas que normalmente necesitaría un investigador sénior.
Son realmente impresionantes, pero yo no basaría una decisión de compra en ellas. Las demostraciones organizadas, preparadas y seleccionadas por el proveedor te muestran lo que un modelo es capaz de hacer en un día perfecto y bajo la supervisión de un experto, no lo que hace en tu repositorio un martes cualquiera. Tómalas como una prueba de viabilidad de una capacidad a largo plazo, no como unas especificaciones técnicas.
Mostrar imagen Dos arquitecturas publicadas y una caja negra. Los modelos abiertos explican exactamente cómo funcionan; el modelo cerrado indica qué puntuación obtiene.
Qué es realmente el GLM-5.3
La versión GLM-5.3 tiene la historia más extraña de las tres, y es la que merece la pena contar como es debido, ya que es la primera vez que un laboratorio importante ha dudado abiertamente a la hora de publicar los pesos por motivos que no son de carácter comercial.
Z.ai anunció GLM-5.3 el 14 de agosto de 2026 con el eslogan “Diseñado según las normas. Listo para la ciberdefensa”. El modelo es un Mezcla de expertos con 753.000 millones de parámetros, con unos 40 000 millones de tokens activos según las estimaciones de la comunidad derivadas de la configuración GLM-5.2. Cuenta con una ventana de contexto de 1 millón de tokens y una salida máxima de 128 000, y —a diferencia de GLM-5.3-Flash, que es un modelo completamente distinto— es solo texto.
“Lo único que hicimos fue ampliar la escala tras la formación”
El dato técnico más interesante sobre el GLM-5.3 es que no hay un nuevo modelo base. Reutiliza la misma base preentrenada que el GLM-5.2. Todas las mejoras se deben a un entrenamiento posterior considerablemente más extenso, que Z.ai resumió de la siguiente manera: “Lo único que hicimos con el GLM-5.3 fue ajustar la escala tras el entrenamiento”.”
Las diferencias no son pequeñas:
| Referencia | GLM-5.2 | GLM-5.3 |
|---|---|---|
| Terminal-Bench 3.0 | 4.6 | 28.3 |
| DeepSWE v1.1 | 46.2 | 66.9 |
| CyberGym | 77.2% | 84.5% |
Una mejora de seis veces en Terminal-Bench y un aumento de veinte puntos en DeepSWE, partiendo de los mismos pesos iniciales, es un resultado verdaderamente importante. Esto indica que la vanguardia aún tiene mucho margen de mejora en la fase posterior al entrenamiento, que es, con diferencia, la mitad más económica del coste total del entrenamiento. Esa es precisamente la razón por la que los laboratorios chinos, que operan con, como dice Lambert, “órdenes de magnitud menos de capital” que los estadounidenses, están manteniendo el ritmo.
Las dos semanas que Z.ai pasó sin realizar envíos
Z.ai afirmó en el momento del lanzamiento que los pesos se publicarían “por fases, tras rigurosas evaluaciones de seguridad”, con un plazo previsto de unas dos semanas. En su repositorio de Hugging Face figuraba la fecha del 28 de agosto. Esa fecha llegó, pasó en cuestión de horas y, a continuación, aparecieron los pesos junto con un informe técnico en el que se explicaba el retraso.
El motivo fue la capacidad cibernética. Durante la evaluación, Z.ai informó de que GLM-5.3 identificó 2.436 vulnerabilidades en 269 proyectos de código abierto, de los cuales 1.097 se clasificaron como de gravedad crítica o alta. Más concretamente, el modelo demostró razonamiento sobre cadenas de exploits de varias etapas — encadenando de forma autónoma los pasos de explotación — lo que Z.ai calificó como un comportamiento que “no era del todo intencionado” y que “se fue agravando a medida que se ampliaba el entrenamiento”.”
Quiero ser prudente en este punto, porque esto se puede interpretar fácilmente como una estrategia de marketing o como alarmismo, y probablemente no sea ninguna de las dos cosas. Un modelo que es eficaz a la hora de detectar vulnerabilidades es eficaz a la hora de detectar vulnerabilidades; los usos defensivos y ofensivos son la misma capacidad orientada en direcciones diferentes. La distribución de resultados de las pruebas comparativas de Z.ai refleja esto con sinceridad: CyberGym 84.5% (descubrimiento de vulnerabilidades, adónde conduce) frente a ExploitBench 54.4% (explotación, donde Claude Fable 5 obtiene una puntuación de 78,01 TP3T). El equipo ha optimizado deliberadamente la mitad defensiva y eso se nota.
Lo que realmente llama la atención es que un laboratorio haya retrasado dos semanas el lanzamiento de un producto estrella y haya hecho público su razonamiento. Independientemente de si ese razonamiento te parece convincente o no, se trata de un precedente sin precedentes.
Dónde aterriza realmente el GLM-5.3
La tabla comparativa publicada por Z.ai no pretende dar a entender que haya arrasado, lo cual es de agradecer:
| Referencia | GLM-5.3 | El mejor comparador |
|---|---|---|
| AutomationBench | 48.2% | Cables GLM-5.3 |
| CyberGym | 84.5% | Cables GLM-5.3 |
| GDPval-AA v2 | 1,769 | Cables GLM-5.3 |
| Terminal-Bench 3.0 | 28.3% | GPT-5.6 Sol: 34,61 TP3T |
| DeepSWE v1.1 | 66.9% | GPT-5.6 Sol: 72,71 TP3T |
| ExploitBench | 54.4% | Claude Fable 5: 78,01 TP3T |
| Code Bench (50 000 tokens) | 31.4% | — |
Por su parte, Artificial Analysis lo sitúa en 59,5 en el índice de inteligencia — estadísticamente indistinguible del 59,7 de Kimi K3, obtenido a partir de un modelo con aproximadamente una cuarta parte de los parámetros.
Una salvedad de esa misma evaluación que nadie del departamento de marketing menciona: GLM-5.3 generó 170 millones de tokens generados en el conjunto de herramientas de análisis artificial, frente a una mediana de 72 millones en su categoría. Es demasiado prolijo. esfuerzo_de_razonamiento por defecto es máx. y no se puede desactivar el pensamiento, por lo que estás pagando por un largo proceso de deliberación, independientemente de si la tarea lo merece o no. Aun así, el coste total de llevar a cabo la evaluación completa ascendió a $0,68 por tarea, frente a $0,84 para el Kimi K3 y $2,34 para el Claude Opus 5; por lo tanto, la mayor complejidad no anula la ventaja en el precio, sino que simplemente la reduce.
Qué es realmente «Claude Fable 5.1»
Anthropic lanzó Claude Fable 5.1 el 1 de septiembre de 2026, junto con el Claude Mythos 5.1 —el mismo modelo con una configuración de medidas de seguridad diferente, disponible únicamente a través de programas de acceso de confianza para trabajos en ciberseguridad y ciencias de la vida—.
La definición es concreta y, en mi opinión, acertada: se trata de “Un modelo diseñado para tareas que no se completan con una sola instrucción”.” Anthropic no afirma que se haya producido un avance en la inteligencia general. Lo que afirma es que las sesiones con agentes, de varias horas de duración, dan mejores resultados.
Las cifras
| Referencia | Fábula 5 | Fable 5.1 | Mythos 5.1 |
|---|---|---|---|
| Terminal-Bench 4.0 | 42.0% | 55.8% | 60.9% |
| Terminal-Bench-Science 0.1 | 24.7% | 52.6% | — |
| AutomationBench | 17.1% | 31.4% | — |
| CursorBench 3.2.0 | 70.5% | 73.4% | — |
| OSWorld 2.0 (parcial) | 72.9% | 77.9% | — |
| El último examen de la humanidad (con herramientas) | 63.8% | 65.0% | — |
| GDPval-AA v2 | — | 1,853 | — |
Para obtener información adicional sobre este mismo gobernante: GPT-5.6 Sol obtiene una puntuación de 52,31 TP3T en Terminal-Bench 4.0, por lo que el 55,81 TP3T de Fable 5.1 se sitúa en cabeza y el 60,91 TP3T de Mythos 5.1 amplía su ventaja.
El salto en la prueba «Terminal-Bench-Science» —de 24,71 TP3T a 52,61 TP3T, lo que supone duplicar la puntuación— es al que yo daría más importancia si tu trabajo tiene que ver con la informática científica o los flujos de datos. Y en la prueba de rendimiento más exigente de Browserbase en cuanto a uso del ordenador, Fable 5.1 completó 82% de tareas frente a las 74% de Claude Opus 5.
Anthropic también informa de que, aproximadamente, un Reducción de los falsos positivos en materia de ciberseguridad según el modelo 60% — el modelo rechaza una tarea de seguridad legítima porque su patrón coincide con algo peligroso. Si alguna vez un asistente de programación se ha negado a ayudarte a escribir un limitador de velocidad, entenderás por qué esto es importante.
La noticia de verdad es el cambio en los precios
Los tipos de interés de referencia no variaron: $10 por cada millón de tokens de entrada, $50 por cada millón de tokens de salida. Lo que ha cambiado es el multiplicador de lectura de la caché, y ha cambiado mucho.
| Fábula 5 | Fable 5.1 | |
|---|---|---|
| Entrada | $10.00 | $10.00 |
| Resultado | $50.00 | $50.00 |
| Escritura en caché (5 min) | $12.50 | $12.50 |
| Escritura en caché (1 hora) | $20.00 | $20.00 |
| Lectura de caché | $1.00 | $0.25 |
Se trata de una reducción de 75%, conseguida al cambiar el multiplicador de lectura de la caché del valor estándar de 0,1× de la entrada base a 0,025×. Anthropic cita, aproximadamente, 25% más económico para cargas de trabajo típicas y hasta 45% más económico para tareas que requieren un alto nivel de agilidad — y esa segunda cifra es real, porque los bucles de los agentes consisten, en su gran mayoría, en lecturas de caché. Una sesión de agente de 60 iteraciones vuelve a leer el mismo contexto acumulado docenas de veces. La API por lotes reduce a la mitad las entradas y salidas, además, lo que da como resultado $5/$25.
Se trata de una rebaja de precio inteligente y bien orientada. Hace que aquello en lo que el Fable 5.1 destaca especialmente —las sesiones prolongadas— resulte significativamente más barato sin que ello suponga una devaluación general del modelo.
Tres cambios que afectan a la compatibilidad, una regresión
Antes de actualizar una integración en producción, léete esto con atención.
1. Ya no es obligatorio utilizar herramientas. Solicitudes que utilizan tool_choice: "cualquiera" o tool_choice: "herramienta" ahora devuelve un Error 400. El proceso de reflexión está siempre activo y no se puede eludir; si se forzara la llamada a una herramienta, se saltaría este paso. Debes utilizar tool_choice: "auto".
2. Los bloqueos mentales son unidireccionales. Fable 5.1 puede leer los ’thinking blocks» generados por modelos anteriores de Claude, pero los modelos anteriores no pueden leer los de Fable 5.1. Solo Mythos 5.1 puede hacerlo. Si tu arquitectura cambia de modelo en mitad de una conversación para ahorrar dinero, esa solución alternativa obliga ahora a replantear la estrategia.
3. El historial de ediciones invalida los bloqueos mentales. En el caso de las cuentas creadas después del 31 de agosto de 2026, al editar turnos anteriores, el mensaje del sistema o la gama de herramientas invalida todos los bloques de razonamiento posteriores —ya sean errores o omisiones silenciosas—. Se trata de una medida explícita contra la destilación que rompe con un patrón habitual en el que los marcos de agentes reescriben el contexto entre turnos.
Y una regresión que conviene conocer: Según se informa, la llamada paralela a herramientas es más variable en la versión 5.1, ya que a veces se realiza una sola llamada por turno, mientras que en la versión 5 se agrupaban varias. En un bucle de agente prolongado, esto supone una pérdida de tiempo real.
Hay además un detalle sobre el tokenizador que suele pasar desapercibido a la hora de comparar precios: Fable 5.1 utiliza el mismo tokenizador que Claude Opus 4.7, que genera aproximadamente 30% más fichas que el tokenizador de los modelos Claude más antiguos para el mismo texto. Si estás comparando el coste con una referencia de 2025, tenlo en cuenta antes de sacar conclusiones.
El problema de la referencia: tres modelos, tres reglas diferentes
Este es el punto en el que la mayoría de los artículos comparativos se equivocan, y vale la pena ser franco al respecto.
No es posible alinear estos tres modelos en Terminal-Bench. Mira:
- Kimi K3: 88.3 en Terminal-Bench 2.1
- GLM-5.3: 28.3 en Terminal-Bench 3.0
- Fable 5.1: 55.8 en Terminal-Bench 4.0
Se trata de tres pruebas de rendimiento diferentes que comparten nombre. Cada versión se ha vuelto considerablemente más difícil que la anterior; ese es precisamente el objetivo de lanzar una nueva versión. Si interpretaras esos tres números como una clasificación, pensarías que Kimi K3 es tres veces mejor que GLM-5.3 en tareas de terminal, lo cual es una tontería.
El mismo problema se da en AutomationBench, donde los números de versión se indican de forma inconsistente, y en la familia SWE-bench, donde las variantes Verified, Pro y Marathon miden aspectos realmente diferentes. Los laboratorios no hacen esto con la intención de engañar a nadie: realizan la evaluación que estaba vigente cuando recibieron su formación, y las evaluaciones cambian cada pocos meses en la actualidad. Pero el efecto que esto tiene en un lector ocasional es el mismo que el de un engaño, por lo que merece la pena señalarlo con claridad.
Lo que sí se puede comparar legítimamente, dentro de una misma versión:
En Terminal-Bench 2.1: Kimi K3 (88,3), GPT-5.6 Sol (88,8), Claude Fable 5 (88,0) y GLM-5.3-Flash (84,3). Cuatro modelos, un único criterio de evaluación. Esta es la comparación que realmente importa y le he dedicado una tabla propia en la siguiente sección.
En Terminal-Bench 4.0: Fable 5.1 (55,8) frente a GPT-5.6 Sol (52,3). Gana Fable.
Además, desconfía de cualquiera que te muestre un único gráfico de barras con los tres modelos principales de este artículo en un eje «Terminal-Bench». No se puede hacer de forma honesta.
Dónde coinciden realmente: las cifras comparables entre sí
Si eliminamos las discrepancias entre las versiones de las pruebas de rendimiento, quedan tres comparaciones.
1. Índice de Inteligencia Artificial en el Análisis (septiembre de 2026)
Se lleva a cabo de forma independiente, se aplica la misma metodología en todos los modelos y no interviene ningún proveedor en la selección.
| Clasificación | Modelo | Puntuación | Peso libre |
|---|---|---|---|
| 1 | Claude, Op. 5 | 63.0 | No |
| 2 | Claude Fable 5 | 62.1 | No |
| 3 | Grok 4.6 | 60.9 | No |
| 4 | Kimi K3 | 59.7 | Sí |
| 5 | GLM-5.3 | 59.5 | Sí |
| 6 | GPT-5.6 Sol | 58.9 | No |
| 7 | Qwen 3.8: vista previa de Max | 58.1 | Peso por determinar |
| 8 | GLM-5.3-Flash | 57.5 | Sí (MIT) |
| 31 | MiniMax M3 | 45.4 | Sí |
En el momento de redactar este artículo, Fable 5.1 aún no había sido evaluado de forma independiente, ya que se había lanzado el día anterior. Teniendo en cuenta las diferencias en las pruebas de rendimiento de Fable 5.1 con respecto a Fable 5, cabe esperar que alcance una puntuación igual o superior a 62,1 cuando se evalúe.
Una nota sobre la precisión: las distintas instantáneas de este índice presentan variaciones —en una de septiembre, el Kimi K3 y el GLM-5.3 aparecen ambos en 60, mientras que el GPT-5.6 Sol está en 61—. El orden en los primeros puestos es estable; las diferencias en las subpuntuaciones, no. No bases un argumento en medio punto.
Lee esa tabla con atención. El mejor modelo de peso abierto se sitúa 3,3 puntos por detrás del mejor modelo cerrado en el índice neutral más amplio disponible. Hace un año, esa diferencia era de dos dígitos.
2. Terminal-Bench 2.1: la única versión en la que los tres coinciden
Esta es la tabla más útil de todo el artículo, y por poco se me pasa por alto. Kimi K3, Claude Fable 5 y GPT-5.6 Sol fueron evaluados en la misma versión de Terminal-Bench:
| Modelo | Terminal-Bench 2.1 | Peso libre |
|---|---|---|
| GPT-5.6 Sol | 88.8 | No |
| Kimi K3 | 88.3 | Sí |
| Claude Fable 5 | 88.0 | No |
| GLM-5.3-Flash | 84.3 | Sí (MIT) |
Medio punto separa un modelo que se puede descargar del mejor modelo cerrado en la misma prueba de rendimiento, con la misma versión, en el trabajo agénico nativo del terminal.
Esa es la cifra que hay que recordar. Ni el índice agregado, ni el gráfico del proveedor… esta. En cuanto a la forma específica que adopta la codificación agencial, la diferencia se pierde entre el ruido.
Una salvedad, señalada por la fuente y que merece la pena repetir: Todas las cifras publicadas de Kimi K3 corresponden a vueltas a máxima potencia, a esfuerzo_de_razonamiento máximo y temperatura 1,0. Los distintos laboratorios utilizan diferentes sistemas de evaluación y diferentes ajustes de esfuerzo, por lo que incluso las comparaciones entre versiones idénticas conllevan una mayor incertidumbre de lo que sugiere una tabla sin matices.
3. Coste de ejecutar el mismo conjunto de pruebas de evaluación
Si la tabla anterior es la mejor comparación en cuanto a prestaciones, esta es la mejor comparación en cuanto a relación calidad-precio, ya que es la única cifra que agrupa las prestaciones, la verbosidad, la sobrecarga de razonamiento y el precio en un único valor:
| Modelo | Coste por tarea del Índice de Inteligencia |
|---|---|
| GLM-5.3 | $0.68 |
| Kimi K3 | $0.84 |
| Claude, Op. 5 | $2.34 |
Las mismas tareas, la misma puntuación, facturas reales. GLM-5.3 ofrece una puntuación de 94% en el índice de Opus 5 por un coste de 29%.
4. GDPval-AA v2 (trabajo intelectual)
| Modelo | Puntuación |
|---|---|
| Claude Fable 5.1 | 1,853 |
| Claude Fable 5 Max | 1,815 |
| GLM-5.3 | 1,769 |
| GPT-5.6 Sol Max | 1,747.8 |
| Kimi K3 | 1,687 |
Hay que hacer una salvedad al respecto: se ha elaborado a partir de gráficos publicados por los propios proveedores, en lugar de un único análisis independiente, y los sufijos “Max” indican diferentes configuraciones de esfuerzo. No obstante, el orden es, en líneas generales, coherente en todas las fuentes, y revela algo significativo: en lo que respecta al trabajo intelectual en general, a diferencia de la programación, los modelos cerrados siguen manteniendo una ventaja más clara de lo que sugieren las comparativas de programación.
Ese es el patrón que se observa en las tres comparaciones. Los modelos abiertos prácticamente han alcanzado el nivel de los modelos cerrados en lo que respecta a la programación y las tareas de tipo «agente». Sin embargo, siguen un paso por detrás en lo que se refiere al trabajo de conocimiento en general. Si tu volumen de trabajo se ajusta a la primera situación, los argumentos a favor de pagar los precios de los modelos cerrados son débiles. Si se ajusta a la segunda, sigue siendo defendible.
Mostrar imagen Tres puntos y pico separan al mejor modelo cerrado del mejor modelo descargable. Hace doce meses, esa diferencia era de dos dígitos.
El cálculo de los costes, con cifras reales
Los índices de referencia son abstractos. Las facturas no lo son.
El escenario
Una característica de nivel medio, realizada de forma activa: aproximadamente 60 turnos de selección de herramientas, con una media de 45 000 fichas de entrada por turno a medida que se va acumulando contexto, y 2.000 fichas de salida por turno.
- Total de entradas: 2,7 millones de tokens
- Producción total: 120 000 tokens
- Tasa de aciertos estimada en la caché de comandos: 80% (2,16 millones en caché, 540 000 recientes)
Claude Fable 5.1
| Componente | Fichas | Tarifa | Coste |
|---|---|---|---|
| Novedades | 540 000 | $10,00/M | $5.40 |
| Entrada almacenada en caché | 2,16 millones | $0,25/M | $0.54 |
| Resultado | 120 000 | $50,00/M | $6.00 |
| Total | ≈ $11,94 |
Kimi K3
| Componente | Fichas | Tarifa | Coste |
|---|---|---|---|
| Novedades | 540 000 | $3,00/M | $1.62 |
| Entrada almacenada en caché | 2,16 millones | $0,30/M | $0.65 |
| Resultado | 120 000 | $15,00/M | $1.80 |
| Total | ≈ $4,07 |
GLM-5.3
| Componente | Fichas | Tarifa | Coste |
|---|---|---|---|
| Novedades | 540 000 | $1,40/M | $0.76 |
| Entrada almacenada en caché | 2,16 millones | $0.26/M | $0.56 |
| Resultado | 120 000 | $4.40/M | $0.53 |
| Total | ≈ $1.85 |
And for reference — GLM-5.3-Flash
| Componente | Fichas | Tarifa | Coste |
|---|---|---|---|
| Novedades | 540 000 | $0,15/M | $0.08 |
| Entrada almacenada en caché | 2,16 millones | $0.03/M | $0.06 |
| Resultado | 120 000 | $0,50/M | $0.06 |
| Total | ≈ $0,21 |
The ratios
- GLM-5.3 is 6.5× cheaper than Fable 5.1 for identical work
- Kimi K3 is 2.9× cheaper than Fable 5.1
- GLM-5.3 is 2.2× cheaper than Kimi K3
- GLM-5.3-Flash is 58× cheaper than Fable 5.1, at 57.5 on the intelligence index against Fable 5’s 62.1
Scaling to a working month
Cuatro tareas de este tipo al día, veinte días laborables: 80 ejecuciones de agente:
| Modelo | Monthly token cost |
|---|---|
| Claude Fable 5.1 | ≈ $955 |
| Kimi K3 | ≈ $326 |
| GLM-5.3 | ≈ $148 |
| GLM-5.3-Flash | ≈ $17 |
Three honest caveats on these figures
Cache writes are excluded. Fable 5.1 charges $12.50/M for 5-minute cache writes, and a long session writes cache repeatedly. Including writes moves Fable’s real number up, not down. The 80% hit rate is also optimistic for short sessions.
Reasoning tokens are billed output tokens. All three models have always-on thinking. GLM-5.3 in particular is documented as verbose — 170M output tokens against a 72M class median across the AA suite — so its real-world output volume runs above what a naive turn count suggests. The $0.68-per-task figure already accounts for this, which is why I trust it more than my own model above.
Fable 5.1’s tokenizer produces ~30% more tokens than older Claude models for the same text. If you are comparing against a historical Claude bill, adjust.
Even after all three corrections, the ordering does not change and the magnitude barely does. Closed-frontier work costs roughly six times what the best open-weight model costs, per completed task.
Mostrar imagen La misma tarea agencial de 60 turnos al precio de catálogo de cada proveedor. El eje vertical representa el argumento completo para los pesos abiertos.
Licences: What “Open” Actually Buys You In 2026
This is the section I would most like people to read, because the vocabulary has drifted badly and it is going to cost somebody a lot of money.
Neither Kimi K3 nor GLM-5.3 is open source. Both publish downloadable weights under bespoke licences with revenue-triggered conditions. That is a meaningfully different thing from MIT or Apache 2.0, and the difference lands squarely on the kind of business that would most benefit from self-hosting.
The Kimi K3 License
You will see “Modified MIT” repeated all over the internet for K3. It is wrong — that described K2. K3 ships under a custom document with two distinct gates:
The Model-as-a-Service gate. If you provide third parties with access to model inference or fine-tuning — where those third parties control inputs, parameters or training data — and “the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars” over any consecutive 12 months, you must negotiate a separate commercial agreement with Moonshot. Note aggregate revenue, across all affiliates, not revenue attributable to K3. A €25M-turnover consultancy that resells inference is over the line even if the AI business is a rounding error.
The attribution gate. Products exceeding 100 million monthly active users or $20 million monthly revenue must display “Kimi K3” prominently in the product interface.
The exemption that matters most: purely internal use is unrestricted. Running K3 to support your own developers, researchers, legal team or general employee productivity does not trigger either gate. For the overwhelming majority of businesses — including essentially all of mine — that is the relevant clause, and the answer is that you are fine.
The clause that is new relative to K2 is the $20M MaaS gate. K2 only required attribution above its thresholds. K3 adds a much lower revenue gate aimed specifically at commercial inference resellers. Moonshot is not trying to stop you using the model; it is trying to stop cloud providers building a business on it for free.
The glm-5.3 License
Z.ai’s flagship went a different direction, and it is a bigger break with precedent than most coverage acknowledged.
GLM-5.2 shipped under MIT. GLM-5.3-Flash shipped under MIT. GLM-5.3 did not.
The custom licence permits individuals and ordinary businesses to run, deploy and fine-tune the model with no additional restrictions. The single gate is aimed at hyperscalers: an entity whose “aggregate revenue of the Licensee and its affiliates exceeds 10 billion US dollars (or the equivalent in other currencies) in total over any consecutive 12 months” debe pass Z.AI’s security review before using the model or its derivatives for any commercial purpose.
That is a $10 billion threshold. It excludes almost everyone. But two details in the fine print deserve attention:
First, the licence explicitly carves out app developers embedding the model in features, and resellers merely routing requests to other hosts. The target is narrow and clearly labelled: companies that want to host GLM-5.3 commercially at scale.
Second — and this is the genuinely awkward part — the licence publishes no criteria, timeline or appeal process for that security review. It states only that its “scope and method… shall be reasonably determined by Z.AI.” If you are a large enterprise, that is an unbounded dependency on a foreign vendor’s discretion, and your procurement team will say so.
Also worth noting: despite the two-week cyber-capability hold that preceded release, the licence contains no acceptable-use restrictions, no cyber carve-outs and no output-ownership claims. The safety concern shaped the release timing, not the release terms.
So what is actually still MIT?
GLM-5.3-Flash. 320B parameters, 18B active, natively multimodal, MIT-licensed, 57.5 on the intelligence index, and $0.15/$0.50 per million tokens — currently half that under a promotional discount running to 9 September 2026. If your requirement is genuine unrestricted permissive licensing rather than merely downloadable weights, that is where the frontier currently sits, and it is remarkably high.
The practical summary
| Kimi K3 | GLM-5.3 | GLM-5.3-Flash | Fable 5.1 | |
|---|---|---|---|---|
| Weights downloadable | Sí | Sí | Sí | No |
| OSI-style open source | No | No | Sí (MIT) | No |
| Restriction trigger | $20M revenue (MaaS) | $10B revenue (MaaS) | Ninguno | N/A |
| Attribution required | Above 100M MAU / $20M-month | Copyright notice only | Copyright notice only | N/A |
| Internal use restricted | No | No | No | N/A |
| Fine-tuning permitted | Sí | Sí | Sí | No |
| Acceptable-use policy | Estándar | None in licence text | Ninguno | Anthropic AUP applies |
For a typical German Mittelstand business, an agency, or a consultancy running models internally or embedding them in a product: all three open options are usable without negotiating anything. The gates exist to catch AWS, not you. But “read the licence” has gone from pedantic advice to actual advice, and if you are above about €20M turnover and reselling inference in any form, get it in front of counsel before you deploy.
The Real Case For Open Models
I want to make this argument properly rather than as a slogan, because “open models are great” is not an argument and the actual reasons are more interesting than the cheerleading.
1. The price collapse is not a subsidy, it is structural
The instinctive objection to cheap Chinese models is that the pricing is loss-leading and will normalise upward. Some of it certainly is. But the structural point stands independently: when weights are public, inference becomes a commodity market. Twenty hosting providers compete to serve the same model, and none of them can charge a capability premium because none of them has a capability moat. That is why GLM-5.3 runs at $1.40/$4.40 and Fable 5.1 runs at $10/$50 for scores three points apart.
Closed-model pricing includes the R&D amortisation, the margin, and the fact that there is exactly one seller. Open-weight pricing includes electricity, hardware amortisation and a thin margin. Those are different businesses, and the gap will not close by open models getting more expensive.
The evidence that this is real rather than temporary: DeepSeek captured roughly 17% of token usage on Vercel by May 2026 while holding about 1% of revenue share. That is not a pricing anomaly. That is what commoditisation looks like on a chart.
2. Exit rights change your negotiating position even if you never exercise them
Most teams reading this will not self-host. The hardware section below explains why — Kimi K3 needs a rack. But the option has value regardless.
If Anthropic raises prices, deprecates the model you built on, changes its acceptable-use policy in a way that breaks your product, or has a bad quarter, your recourse with a closed model is a migration project. With an open-weight model your recourse is downloading a file you could have downloaded any time. You may never do it. The fact that you could is what makes the vendor relationship a commercial one rather than a dependency.
This is not theoretical. Fable 5.1 shipped with three breaking changes on 1 September, one of which invalidates thinking blocks when you edit conversation history for accounts created after 31 August. If that pattern is load-bearing in your architecture, you are re-engineering on Anthropic’s schedule, not yours.
3. Data residency stops being a project
For anyone operating under GDPR, this is the argument that actually closes deals.
With a closed API, every prompt leaves your infrastructure. You need a processor agreement, a transfer impact assessment if the processor is outside the EEA, a data-flow map, and an answer for your DPO about what happens to the data at rest. All of that is doable. It is also weeks of work per vendor, repeated whenever the vendor changes anything.
With open weights on infrastructure you control — your own hardware, or a German or EU GPU host — the cross-border transfer question does not arise, because there is no transfer. The model runs where your data already lives. Legal review goes from a transfer assessment to a licence read.
For German and EU clients this is frequently the deciding factor, and it is worth being precise about the caveat: using GLM-5.3 through Z.ai’s API does not give you this. You get it from running the weights yourself, or from a provider hosting them in your jurisdiction. The licence makes that legal; it does not make it automatic.
4. Inspectability is real, if underused
Open weights mean you can examine the architecture, run your own evaluations on the actual model rather than an endpoint that might be silently updated, quantise it to fit your hardware, fine-tune it on your domain, and — importantly — pin a version forever.
That last one is underrated. Closed APIs get updated behind stable model IDs. Behaviour drifts. Prompts that worked stop working, and you cannot diff the change because you cannot see it. A local checkpoint is byte-identical in a year’s time.
5. The competitive pressure benefits everyone, including closed-model users
Look at what happened this summer. Kimi K3 lands in July at $3/$15 with frontier-adjacent scores. GLM-5.3 lands in August at $1.40/$4.40 within half a point of it. And on 1 September, Anthropic cut cache reads by 75%.
I am not claiming direct causation — Anthropic does not publish its pricing rationale and cache-read cuts are a natural optimisation. But a market with credible cheap substitutes prices differently from one without them, and the substitutes got credible this year. If you use closed models exclusively, the open-weight ecosystem is still quietly making your bills smaller.
6. Post-training is where the gains are, and it is cheap
GLM-5.3’s headline result — a six-fold Terminal-Bench improvement from the same base weights — is the most strategically important number in this entire article, and it is not about GLM.
It says the expensive part of building a frontier model (pretraining) is increasingly a solved, commoditised input, and the differentiating part (post-training, RL environments, long-horizon task design) is comparatively affordable. That is why four Chinese labs with a combined valuation of about $159 billion are keeping pace with American labs valued at multiples of that. Lambert’s assessment is that Chinese labs are running with “orders of magnitude less capital.”
For anyone downstream, the implication is straightforward: the number of credible model vendors is going up, not down. Plan your architecture accordingly.
7. The ecosystem effects compound
Qwen passed one billion cumulative downloads on Hugging Face, overtaking Llama, and now anchors over 200,000 tagged models and 113,000+ derivatives — roughly 40% of all new LLM derivatives on the platform are Qwen-based. That is a tooling, quantisation, fine-tuning and deployment ecosystem that exists only because the weights are public.
You benefit from that ecosystem whether or not you contribute to it. GGUF conversions, vLLM kernels, LoRA adapters, quantisation recipes, evaluation harnesses — all of it exists because thousands of people could get their hands on the actual weights.
8. It is where the developers already went
Chinese-origin models at ~61% of OpenRouter tokens. US model share down from ~70% to ~30% in a year. Four of the five most-used models on the router are Chinese-origin. Xiaomi’s MiMo models alone account for roughly 21% of routed tokens and about 22% of all coding traffic.
Developer behaviour is a leading indicator. It was a leading indicator for Docker, for Postgres, for Linux. Betting against it has a poor historical record.

…And The Honest Case Against
If I only wrote the section above, this would be an advert.
Open weights are not open source, and the vocabulary drift is a real problem. Two of the three models here ship under bespoke revenue-gated licences with no OSI approval. GLM-5.3’s security-review clause has no published criteria or appeal process. If your compliance framework requires OSI-approved licensing, the honest answer is that your options are GLM-5.3-Flash, DeepSeek’s MIT-licensed flagships, and a shrinking list of others.
Vendor benchmarks remain vendor benchmarks. Every headline number Moonshot and Z.ai published is self-reported and self-selected. Where neutral aggregators have measured, the numbers broadly hold up — which is genuinely to both labs’ credit — but “broadly hold up” is doing work in that sentence.
Self-hosting is a fantasy for most teams. Kimi K3 is 1.5 TB and Moonshot recommends 64+ accelerators. GLM-5.3 needs 8×H200 at FP8 as a bare minimum. The exit right is real; the exercise cost is a data centre.
Support is what you make it. When Fable 5.1 breaks, there is a company with an SLA. When your self-hosted GLM-5.3 deployment produces garbage at 300K context on a Friday night, there is a GitHub issue and your own competence.
Geopolitics is a real procurement input. All three open-weight options discussed here come from Chinese labs. For some clients — public sector, defence-adjacent, certain regulated industries — that is a hard blocker regardless of licence terms or where the weights run. It is not my job to tell you whether that concern is well-founded. It is my job to tell you it will come up in the meeting.
The closed models still lead on general knowledge work. GDPval-AA v2 has Fable 5.1 at 1,853 against GLM-5.3’s 1,769 and Kimi K3’s 1,687. On coding the gap is gone. On broad professional knowledge work it is not.
Mostrar imagen Peso descargable, tres juegos diferentes de cuerdas. Solo GLM-5.3-Flash sigue siendo del MIT.
Ollama In 2026: The Pricing Change That Actually Matters
Ollama has quietly become the most important piece of infrastructure in this conversation, and on 31 August 2026 it changed how it charges in a way that is worth understanding properly.
What changed
Ollama moved its Pro, Max and Team plans from GPU-time billing to industry-standard per-token pricing, with a pool of usage credits included in every plan. The company’s stated reason is refreshingly concrete: models like Kimi K3 made GPU-time metrics impossible to predict. When one model activates 104B parameters and another activates 3B, “an hour of GPU” stops meaning anything to the person paying.
The new plans
| Plan | Precio | Included monthly usage | Concurrency |
|---|---|---|---|
| Gratis | $0 | Small monthly credit, starter models | 1 request |
| Pro | $20/mo (or $200/yr) | $60 of usage | 3 requests |
| Max | $100/mo | $300 of usage | 10 requests |
| Team | $500/mo | $1,000 shared, unlimited users | 10 requests |
| Empresa | Custom | Custom | Custom |
Read the Pro row again. $20 buys $60 of tokens. That is a 3× multiplier on included usage, and when the pool runs out you continue at exactly the same published per-token rate — no penalty tier, no service fee.
What was removed
This is the part developers actually noticed: no 5-hour resets and no weekly caps. Anyone who has hit a rolling usage window mid-refactor will understand why that mattered more than the price. The monthly pool refreshes on your subscription date and, importantly, does not roll over — so size your plan to your normal month, not your busiest one.
The terms that matter for EU work
Three commitments in the announcement are directly relevant if you are handling client data:
- Zero data retention. Prompts and responses are never logged and never trained on.
- Hosting in the US and Europe, plus Singapore for a limited set of Qwen models.
- Per-request cost visibility in your account — you can see exactly what each call cost.
For a German consultancy that is a materially better compliance story than most gateway providers offer, though “hosted in Europe” is a routing statement rather than a contractual data-residency guarantee. If residency is a hard requirement rather than a preference, ask Ollama for it in writing before you assume it.
The actual per-token rates
This is where it gets interesting. Ollama publishes rates per model, and they track first-party pricing closely:
| Model on Ollama | Entrada / 1M | Cached / 1M | Salida / 1M |
|---|---|---|---|
kimi-k3 | $3.00 | $0.30 | $15.00 |
glm-5.3 | $1.40 | $0.26 | $4.40 |
mistral-large-3 | $0.50 | $0.50 | $1.50 |
gemma4 | $0.14 | $0.05 | $0.40 |
nemotron-3-super | $0.015 | $0.015 | $0.60 |
Run the monthly maths from earlier through this. Eighty agentic runs a month on GLM-5.3 costs about $148 in tokens. A Max plan at $100/month covers $300 of usage — so the same workload that would cost roughly $955/month on Fable 5.1 fits comfortably inside a $100 Ollama subscription with headroom to spare.
That is the entire open-model economic argument compressed into one line on an invoice.
Mostrar imagen Eighty agentic runs a month on GLM-5.3 fit inside a $100 Ollama Max plan with change. The same work on Fable 5.1 is a $955 invoice.
What else Ollama shipped in 2026
The pricing change did not happen in isolation. The last few months have been busy:
- 25 August — Claude Desktop support. Ollama now works as a third-party gateway provider for Claude Desktop, so you can drive open models through Anthropic’s own client.
- 26 August — v0.33.0. The Claude Desktop gateway integration landed, prefill recovery was restored to the cache, and Ollama disabled Claude Code’s token-countdown system message, which was invalidating the KV cache on every turn. That last one is a quiet but significant performance fix.
- 26 August — v0.33.1. MLX support for Qwen3.8 Flash Next, structured output, and a fix for GPU timeouts when loading models from slower storage.
- 28 August — v0.33.2. Dark mode restored, macOS handoff fixed, and Claude Desktop proxy requests kept alive during model catalogue updates.
- 20 August — v0.32.15. New desktop onboarding, and metadata caching between requests that cut time-to-first-token by roughly half.
- 11 August — NVIDIA Nemotron 3.5 Lightning, a 30B model tuned for agentic workflows with tool calling on personal hardware.
- 10 August — Meta’s Muse Glimmer, a 30B multimodal model under Apache 2.0, accelerated by Ollama’s MLX engine. Worth noting given Llama’s collapse in routed usage — Meta is still shipping genuinely permissive weights.
- 9 July — $88M funding round, with Ollama reporting 8.9 million developers.
- 29 June — Gemma 4 on MLX up to 90% faster via multi-token prediction, which disproportionately benefits coding agents on Apple Silicon.
- 5 June — v0.30 added GGUF compatibility through llama.cpp, broadening hardware support well beyond Apple Silicon.
The through-line is that Ollama has stopped being “the easy way to run a small model on your laptop” and become a general-purpose gateway that happens to also run models locally. The cloud catalogue now includes kimi-k3:cloud, glm-5.3:nube, glm-5.3-flash, deepseek-v4-pro, deepseek-v4-flash, minimax-m3, qwen3.5 across seven sizes, gpt-oss at 20b and 120b, gemma4, the Nemotron 3 family and mistral-large-3.
Setting It All Up: Copy-Paste Configs
Ollama + Claude Code (the fastest path)
Ollama ships a one-command launcher that handles the environment wiring for you:
bash
ollama launch claude
If you would rather do it manually — which you should if you are scripting it — install Claude Code first:
bash
# macOS / Linux
curl -fsSL https://claude.ai/install.sh | bash
# Windows (PowerShell)
irm https://claude.ai/install.ps1 | iex
Then point it at Ollama:
bash
export ANTHROPIC_AUTH_TOKEN=ollama
export ANTHROPIC_API_KEY=""
export ANTHROPIC_BASE_URL=http://localhost:11434
And run against whichever model you want:
bash
# A local model
claude --model qwen3.5
# A cloud model — note the :cloud suffix
claude --model kimi-k3:cloud
Or inline, without exporting anything globally:
bash
ANTHROPIC_AUTH_TOKEN=ollama \
ANTHROPIC_API_KEY="" \
ANTHROPIC_BASE_URL=http://localhost:11434 \
claude --model glm-5.3:cloud
Two things that will bite you.
En primer lugar, do not export ANTHROPIC_BASE_URL permanently in your shell profile. It is a global variable and it will silently redirect every Anthropic-speaking tool on your machine to Ollama, including the one you wanted talking to Anthropic. Scope it per-project or per-invocation.
En segundo lugar, set your context length to 64k or higher for anything working on a real repository. Ollama’s default is smaller, and a coding agent that quietly runs out of window will just start forgetting things rather than erroring.
For CI, Docker or scripted runs, --yes skips the interactive prompts:
bash
ollama launch claude --model glm-5.3:cloud --yes -- -p "how does this repository work?"
GLM-5.3 direct from Z.ai
If you want first-party routing rather than going through Ollama, Z.ai exposes three protocol endpoints:
| Protocolo | URL base |
|---|---|
| Mensajes antropicos | https://api.z.ai/api/anthropic |
| Completados de OpenAI Chat | https://api.z.ai/api/coding/paas/v4 |
| Respuestas de OpenAI | https://api.z.ai/api/v1 |
An OpenCode provider block for GLM-5.3:
json
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"zai": {
"npm": "@ai-sdk/openai-compatible",
"name": "Z.AI",
"options": {
"baseURL": "https://api.z.ai/api/coding/paas/v4",
"apiKey": "{env:ZAI_API_KEY}"
},
"models": {
"glm-5.3": {
"name": "GLM-5.3",
"limit": { "context": 1000000, "output": 128000 },
"options": { "reasoning_effort": "high", "temperature": 1 }
}
}
}
},
"model": "zai/glm-5.3"
}
Important upgrade note: GLM-5.3 requires reasoning to be enabled. If your existing configuration sets thinking.type: "disabled", that will now fail. Change it to "enabled" and set reasoning_effort: "low" if you want the old latency profile back.
esfuerzo_de_razonamiento is your main quality-versus-cost dial: bajo for latency-sensitive calls, alto for substantial work, máx. (the default) for problems that have genuinely defeated alto. Given the documented verbosity, I would run alto as an everyday setting rather than leaving máx. on and wondering where the output tokens went.
The routing setup I would actually recommend
None of this is a “pick one” decision. The correct answer for most teams is a three-tier routing table, and every serious agent harness supports per-agent models.
json
{
"$schema": "https://opencode.ai/config.json",
"model": "zai/glm-5.3",
"small_model": "ollama/gemma4",
"agent": {
"plan": {
"description": "Architecture, root-cause analysis, anything expensive to get wrong",
"model": "anthropic/claude-fable-5-1"
},
"build": {
"description": "The bulk of implementation work",
"model": "zai/glm-5.3",
"options": { "reasoning_effort": "high" }
},
"research": {
"description": "Long-context reading, web research, multi-source synthesis",
"model": "ollama/kimi-k3:cloud"
},
"grunt": {
"description": "Renames, docstrings, test scaffolding, lint fixes",
"model": "ollama/glm-5.3-flash"
}
}
}
The reasoning behind each line:
- Planning gets Fable 5.1. Planning is where a bad decision costs you the most downstream tokens, and long-horizon coherence is precisely what 5.1 was built for. It is also where you spend the fewest tokens, so the price premium applies to the smallest slice of your bill.
- Building gets GLM-5.3. Best capability-per-euro on the market, and the bulk of your token spend lands here.
- Research gets Kimi K3. BrowseComp 91.2 and a genuine 1M context make it the best available option for reading a lot of things and synthesising them.
- Mechanical work gets GLM-5.3-Flash. At $0.15/$0.50 it is effectively free, and renaming a symbol across 40 files does not need frontier reasoning.
modelo_pequeñogets something tiny for the harness’s own housekeeping — title generation, summarisation, internal utility calls.
You are not choosing a winner. You are building a gearbox. And write the model IDs so that swapping one is a one-line change, because in this market you will be making that change again within the quarter.
What You Can Actually Run Locally (The Hardware Reality)
Let me kill an assumption before it costs somebody money.
You cannot run Kimi K3 or GLM-5.3 on a workstation. Not with a 5090. Not with two.
Kimi K3
- ~1.5 TB of weights in native MXFP4 — and remember, that es the quantised checkpoint, not a starting point for further compression
- Moonshot recommends a minimum of 64 accelerators for competitive serving
- Realistic self-hosting starts at multi-node clusters; reference deployments use GB300 NVL72 racks
- There is no official Ollama library entry for local Kimi K3 and no consumer GGUF conversion worth pointing you at
- Reports of it running on clusters of consumer RTX 5090s exist, but “a cluster of 5090s” is not a laptop and the throughput is not comparable
GLM-5.3
- ~1.5 TB in BF16, roughly 750 GB at FP8
- Minimum viable single node: 8× H200 (1,128 GB of GPU memory) at FP8, leaving around 375 GB for KV cache
- BF16 needs two 8×H200 nodes, or a single 8×B300 node (2,304 GB)
- A single 8×H200 node in BF16 is too small for the weights plus cache
Mostrar imagen The frontier open models need a rack. The 24 GB tier is where “runs on my machine” actually lives — and it has got very good.
So what does “local AI” actually mean in September 2026?
It means a different tier of model, and that tier has got genuinely good:
| Modelo | VRAM | What it is for |
|---|---|---|
| Qwen3.8-27B | 24 GB at Q4 | Best all-rounder on consumer hardware — 61.7% SWE-bench |
| gpt-oss:20b | 16 GB | Best small model, adjustable reasoning effort |
| Gemma 4 E4B | ~6 GB | Vision plus tool calling on a modern laptop |
| Mistral 7B | 8 GB | Fastest general-purpose option, 40–60 tok/sec |
| DeepSeek-R1 7B | 5 GB | Chain-of-thought reasoning on a laptop GPU |
| Llama 4 Scout | ~55 GB at Q4 | 10M context, multimodal — workstation territory |
A Qwen3.8-27B scoring 61.7% on SWE-bench, running entirely on a 24 GB consumer GPU with no network connection, is a remarkable thing that would have sounded like science fiction eighteen months ago. It is not GLM-5.3 and it is not pretending to be.
The honest framing: open weights at the frontier buy you sovereignty and price, not local execution. Open weights in the 7B–30B range buy you genuine local execution, at a real but acceptable capability cost, and that is the tier where “runs on my machine, sees no network” is an achievable requirement.
The architecture that works for most of my clients is exactly that split: a small local model for anything touching genuinely sensitive data, and a hosted open-weight frontier model for everything else — with the weights available as insurance rather than as a deployment plan.
Dónde falla realmente cada modelo
No hype. Here are the honest weaknesses.
Kimi K3 weaknesses
It is expensive for an open model. $3/$15 is 2.1× GLM-5.3’s input rate and 3.4× its output rate for essentially the same intelligence index score. If you are choosing K3 over GLM-5.3, be clear about what you are buying with that premium — usually it is the vision stack or the browsing performance.
The licence has the lower gate. A $20M aggregate-revenue MaaS threshold catches far more organisations than GLM-5.3’s $10B. If inference resale is anywhere in your business model, K3 is the more constrained of the two.
Self-hosting is out of reach for almost everybody. 1.5 TB and 64+ accelerators is a serious infrastructure commitment. The exit right is more theoretical here than with any other model in this comparison.
Broad knowledge work trails. GDPval-AA v2 at 1,687 puts it behind GLM-5.3, both Claude Fables and GPT-5.6 Sol Max. It is a coding and agentic specialist that happens to be enormous.
GLM-5.3 weaknesses
Text only. No image input, no video. If your workflow includes screenshot debugging, design-to-code or document vision, GLM-5.3 simply cannot do it and you need GLM-5.3-Flash, Kimi K3 or Fable 5.1 instead. This is the single most common configuration mistake I expect people to make, because the model IDs look related and are not.
It is verbose, and verbosity is billed. 170M output tokens against a 72M class median across the AA suite. esfuerzo_de_razonamiento por defecto es máx., and thinking cannot be turned off. Budget for more output tokens than your turn count implies.
The licence is no longer MIT, and the security-review clause is unbounded. For most readers the $10B threshold makes this academic. For anyone near it, “scope and method shall be reasonably determined by Z.AI” is not a clause your legal team will enjoy.
Terminal-Bench 3.0 at 28.3 trails GPT-5.6 Sol’s 34.6. On the specific benchmark closest to terminal-native agentic coding, it is behind the closed competition on the same ruler.
It is new to open weights. Released 28 August. The community has had days, not months. Long-tail deployment bugs have not surfaced yet.
Claude Fable 5.1 weaknesses
The price. Roughly 6.5× GLM-5.3 per completed agentic task, even after the 75% cache-read cut. For most work that gap is not defensible on capability grounds any more.
Three breaking changes. Forced tool use returns 400. Thinking blocks do not travel backwards to older models. Editing conversation history invalidates thinking blocks on accounts created after 31 August 2026. Any of these can break a working integration on upgrade.
Parallel tool calling regressed. Reports of one call per turn where Fable 5 batched several. On a long agent loop that is wall-clock time you are paying for twice.
Zero exit optionality. No weights, no self-hosting, no version pinning beyond what Anthropic offers, no inspection. When it changes, you adapt.
The tokenizer inflates comparisons. ~30% more tokens than older Claude models for the same text, which makes historical cost comparisons misleading in Anthropic’s favour if you are not careful.
The Decision Framework: Six Scenarios
Mostrar imagen Six scenarios, six answers. Find the row that sounds like your week.
1. “I want one model. Set it, forget it, keep the bill sane.”
GLM-5.3. Within half a point of Kimi K3 and roughly three points of the best closed model on the neutral index, at $1.40/$4.40. Run it through Ollama on a Pro or Max plan, set esfuerzo_de_razonamiento: alto, and get on with your work. The only thing that should push you off this answer is needing image input.
2. “My work is visual — UI, design-to-code, screenshot debugging, documents.”
Kimi K3 or GLM-5.3-Flash, not GLM-5.3. K3’s MoonViT-V2 handles text, images and video natively and scores 81.6/83.4 on MMMU-Pro. GLM-5.3-Flash is the budget option with native multimodality and an MIT licence. GLM-5.3 is text-only and will simply refuse the input.
3. “Long autonomous sessions where being wrong is expensive.”
Claude Fable 5.1. This is what it was built for and the benchmarks back it: Terminal-Bench-Science doubled, 82% on Browserbase’s hardest computer-use tasks against Opus 5’s 74%, and the largest published gains on multi-hour agentic work. The 75% cache-read cut makes exactly this workload up to 45% cheaper than it was. Pay the premium where a mistake costs more than the tokens.
4. “EU data residency is a hard requirement.”
GLM-5.3-Flash if you need MIT, GLM-5.3 if you need capability. Self-host on your own hardware or an EU GPU provider and the cross-border transfer question stops existing. Budget 8×H200 for GLM-5.3 at FP8; Flash is far more tractable at 320B/18B. If self-hosting is out of budget, Ollama’s Europe hosting with zero data retention is the pragmatic middle ground — but get the residency commitment in writing rather than inferring it from a marketing page.
5. “Small team, tight budget, coding all day.”
GLM-5.3-Flash as default, GLM-5.3 for hard problems, Ollama Pro at $20. Flash costs $0.15/$0.50 — currently half that until 9 September — and scores 57.5 on the intelligence index. Twenty dollars buys sixty dollars of tokens with no weekly caps. For a two-to-four person team this is close to unbeatable. The GLM Coding Plan at $18/month (Lite) is the alternative if you prefer a fixed quota to a credit pool.
6. “I need to justify this to a procurement or legal team.”
GLM-5.3-Flash. It is the only model in this comparison under a standard OSI-approved licence (MIT), it is natively multimodal, it scores 57.5 on the neutral index, and the weights are on Hugging Face with no revenue gates, no attribution mandates and no security-review clause. When the question is “what can we defend in a contract review,” permissive licensing beats three points of benchmark every time.
Lo que seguiría de cerca durante el próximo trimestre
Whether Fable 5.1 lands above 62.1 on the Artificial Analysis index. It launched the day before this article and has not been independently scored. Its benchmark deltas over Fable 5 suggest it should, but “should” is not “did,” and the gap to Kimi K3’s 59.7 is the number the whole open-versus-closed argument turns on.
Whether the licence drift continues. In eight weeks we went from GLM-5.2 under MIT to GLM-5.3 under a bespoke licence with a discretionary security review, and from Kimi K2’s modified-MIT to K3’s revenue-gated terms. If GLM-6 and K4 tighten further, “open weights” becomes a marketing term rather than a meaningful category. GLM-5.3-Flash staying MIT is the counter-signal worth tracking.
Whether other labs adopt Z.ai’s staged-release pattern. A two-week hold with a published safety rationale is a new norm. If it holds, it is a good one. If it becomes a reason weights ship later and later, it is a soft path to not shipping them at all.
Independent replication of the cyber capability claims. 2,436 vulnerabilities across 269 projects is an extraordinary number and it is entirely self-reported. Somebody neutral needs to check it, because if it is accurate it reframes the entire open-weights safety conversation, and if it is not, it was effective marketing.
Ollama’s per-token rates six months from now. $20 for $60 of usage is an aggressive introductory posture from a company that raised $88M in July. Whether those multipliers survive contact with real unit economics is the single biggest variable in the “cheap open models” thesis for small teams.
Whether local models close on the 30B tier. Qwen3.8-27B at 61.7% SWE-bench on 24 GB is the most under-discussed result of the year. The frontier gets the headlines; the 24 GB tier is what changes what an ordinary business can do without an API key.
Preguntas frecuentes
Is Kimi K3 really open source? No. Kimi K3’s weights are freely downloadable from Hugging Face, but under a custom “Kimi K3 License,” not an OSI-approved open-source licence. Model-as-a-Service operators whose aggregate revenue exceeds $20 million over any consecutive 12 months must negotiate a separate commercial agreement with Moonshot, and products above 100 million monthly active users or $20 million monthly revenue must display “Kimi K3” in their interface. Purely internal use is unrestricted.
Is GLM-5.3 better than Kimi K3? They are effectively tied on the neutral aggregate index — 59.5 against 59.7 on Artificial Analysis. GLM-5.3 is 2.2× cheaper per agentic task and scores higher on knowledge work (GDPval-AA v2: 1,769 vs 1,687). Kimi K3 is natively multimodal, leads on agentic browsing (BrowseComp 91.2) and took first place on Arena.AI’s frontend code arena. Choose GLM-5.3 for cost and text-based coding; Kimi K3 for vision, browsing and long-horizon research.
How much cheaper are open models than Claude Fable 5.1? On a modelled 60-turn agentic task, GLM-5.3 costs about $1.85 against Fable 5.1’s $11.94 — roughly 6.5× cheaper. Kimi K3 costs about $4.07, roughly 2.9× cheaper. GLM-5.3-Flash costs about $0.21, roughly 58× cheaper. Over 80 runs a month that is $148 versus $955.
Can I run Kimi K3 or GLM-5.3 locally? Not on consumer hardware. Kimi K3 is roughly 1.5 TB even in its native MXFP4 format and Moonshot recommends 64+ accelerators. GLM-5.3 needs a minimum of 8×H200 (1,128 GB) at FP8. For genuine local execution, look at Qwen3.8-27B (24 GB at Q4), gpt-oss:20b (16 GB) or Gemma 4 E4B (~6 GB).
What changed in Claude Fable 5.1? Released 1 September 2026. Cache reads dropped 75% from $1.00 to $0.25 per million tokens, making typical workloads about 25% cheaper and agentic workloads up to 45% cheaper; input and output stayed at $10/$50. Terminal-Bench 4.0 rose from 42.0% to 55.8%, Terminal-Bench-Science from 24.7% to 52.6%, AutomationBench from 17.1% to 31.4%. Three breaking changes affect forced tool use, thinking-block portability and history editing.
What is Ollama’s new pricing? From 31 August 2026, Pro, Max and Team plans use per-token pricing with included credits: Pro $20/month for $60 of usage, Max $100 for $300, Team $500 for $1,000 shared across unlimited users. The 5-hour and weekly caps were removed entirely, there are no service fees, credits do not roll over, and all plans carry zero data retention with hosting in the US and Europe.
Which of these models is genuinely MIT-licensed? Only GLM-5.3-Flash — a separate 320B/18B natively multimodal model released 26 August 2026, scoring 57.5 on the Artificial Analysis index at $0.15/$0.50 per million tokens. GLM-5.3 and Kimi K3 both use bespoke revenue-gated licences; Claude Fable 5.1 is fully proprietary.
Why can’t I compare these models on Terminal-Bench? Because they were each evaluated on a different version. Kimi K3’s 88.3 is on Terminal-Bench 2.1, GLM-5.3’s 28.3 is on 3.0, and Fable 5.1’s 55.8 is on 4.0. Each version is substantially harder than the last, so the numbers are not on the same scale. Compare within a version only. On Terminal-Bench 2.1, for example, Kimi K3 scores 88.3 against Claude Fable 5’s 88.0, GPT-5.6 Sol’s 88.8 and GLM-5.3-Flash’s 84.3 — that comparison is valid, and it is the one worth quoting.
Should I use Ollama or go direct to the vendor? Ollama if you want one billing relationship, one API surface, easy model switching and EU/US hosting with zero data retention — its per-token rates track first-party pricing closely. Direct if you need first-party features like Z.ai’s esfuerzo_de_razonamiento controls at full fidelity, vendor SLAs, or subscription plans such as the GLM Coding Plan. Many teams run both and route by workload.
Does GLM-5.3 support images? No. GLM-5.3 is text-only. GLM-5.3-Flash — a completely different model despite the similar name — is natively multimodal and handles text, image, video and file input. Sending images to the wrong model ID is the most common early mistake with the GLM family.
Conclusión
Twelve months ago the open-weight question was whether these models were usable. Six months ago it was whether they were competitive. In September 2026 it is genuinely: what are you still paying a closed-model premium for?
There is a real answer to that question, and it is narrower than it used to be. Claude Fable 5.1 leads on the hardest sustained agentic work, on general knowledge work, and on the class of debugging where a model needs to hold a messy problem in its head for hours without drifting. Anthropic’s 75% cache-read cut targets exactly that workload. If your failure cost exceeds your token cost, that premium is rational.
For everything else, the maths has moved. GLM-5.3 delivers 94% of Claude Opus 5’s index score at 29% of the cost. Kimi K3 scores 88.3 on Terminal-Bench 2.1 against Claude Fable 5’s 88.0 on the same version, and took first place on LMArena’s Frontend Code Arena ahead of both flagship closed models. Ollama will sell you $60 of tokens for $20 and remove the usage caps while doing it. That combination did not exist in the spring.
But do not let the enthusiasm skip the fine print, because there are two of them and both matter.
The first is licensing. “Open weights” and “open source” have quietly stopped meaning the same thing. Kimi K3 and GLM-5.3 both ship under bespoke, revenue-gated licences. The gates are high enough that most readers are unaffected — but “most readers are unaffected” is not the same as “unrestricted,” and GLM-5.3’s undefined security-review clause is a genuine procurement risk for large enterprises. GLM-5.3-Flash under MIT is the last fully permissive frontier-adjacent option, which is precisely why it deserves more attention than it gets.
The second is that self-hosting is mostly aspirational. 1.5 TB of weights and 64 accelerators is not an exit plan for a mid-sized business. What open weights buy you at this tier is a competitive inference market, price transparency, version pinning, jurisdiction choice, and a negotiating position. Those are worth a great deal. They are not the same as running the thing in your basement.
The setup I would actually build: GLM-5.3 as the default, Fable 5.1 on planning and the genuinely hard problems, Kimi K3 for vision and long-horizon research, GLM-5.3-Flash for the grunt work, and a 27B local model for anything that must never leave the building. Route by workload, not by loyalty. Write your configuration so the model IDs are a one-line change.
Because the only prediction I am confident about is that this article will need updating before Christmas.
Are you running any of these three in production? I am particularly interested in whether GLM-5.3’s verbosity shows up as a real cost problem at scale, and whether anyone has actually put the Kimi K3 licence in front of counsel and got a clear read on the MaaS definition. Get in touch — corrections and counter-evidence welcome, and this article gets updated when the picture changes.
Last updated: 2 September 2026.



