
Claude Haiku 5.5 vs Qwen3.8 27B: una barrida total en la API, y por qué los pesos abiertos todavía existen
- openaiNUEVOOpenAI: GPT-6.1 Sol2026-09-2952Inteligencia
- anthropicNUEVOAnthropic: Claude Sonnet 5.52026-09-2856Inteligencia
- typesafeNUEVOTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 por 1M de tokens · 127 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Inteligencia
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Inteligencia
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Inteligencia
- xAIGrok 4.72026-09-2146Inteligencia
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 por 1M de tokens · 58 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 por 1M de tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Inteligencia
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Inteligencia77Código
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Inteligencia76Código
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Inteligencia76Código
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Inteligencia82Código
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 por 1M de tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 por 1M de tokens · 350 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Inteligencia72Código
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 por 1M de tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Inteligencia75Código
- obsidianQwen3.8 27B2026-08-1534Inteligencia68Código
Most of the comparisons in this series come down to a trade. This one does not, and the absence of a trade is itself the finding. Claude Haiku 5.5, shipped on October 7, 2026, beats Qwen3.8 27B, the open-weight 27B dense model from August 14, 2026, on the composite intelligence score, on cost per token, on cost per finished task, on context window, on response ceiling, on agentic coding, and on the science-reasoning rows — and it does so from a single proprietary API that requires no hardware. The one row Qwen3.8 27B wins outright is video input, and the other is a licence.
Eso no es una razón para descartar el modelo, porque una licencia Apache-2.0 y una modalidad de vídeo no son premios de consolación. Es una razón para ser precisos sobre por qué alguien elegiría el lado de los pesos abiertos, y para dejar de fingir que la discusión gira en torno a los benchmarks.
Los números sobre los que realmente se ejecutaron ambos modelos
Per million tokens in US dollars. The vendor's rates are from its own pricing documentation; the shared measurements are Artificial Analysis Intelligence Index v4.3.2, read 2026-10-08, both in their highest published effort configuration.

• Entrada — Claude Haiku 5.5 $0.10 hasta 100.000 tokens y $0.50 por encima; Qwen3.8 27B $0.50 según la tarifa de terceros registrada por el consejo.
• Salida — Claude Haiku 5.5 $0.50 hasta 100,000 tokens y $2.50 por encima; Qwen3.8 27B $3.00.
• Entrada en caché — Claude Haiku 5.5 $0.01 y $0.05, un 90 % de descuento; Qwen3.8 27B $0.10, un 80 % de descuento.
• Puntuación independiente — 43,40 para Claude Haiku 5.5 (Max) frente a 33,70 para Qwen3.8 27B (Xhigh). Una diferencia de 9,70 puntos, la más amplia de este lote.
• Coste por tarea finalizada — $0,21 frente a $1,01, en la misma ejecución de evaluación. Una diferencia de 4,7x, también la más amplia que aparece aquí.
• Contexto y salida — 1M de tokens y un límite de respuesta de 128,000 tokens para Claude Haiku 5.5; un contexto de 256K tokens para Qwen3.8 27B.
• Modalidad — ambas aceptan texto e imágenes y devuelven texto. Qwen3.8 27B añade entrada de video; Claude Haiku 5.5 no.
• Pesos — Qwen3.8 27B son 27B de parámetros densos bajo Apache 2.0, descargable. Claude Haiku 5.5 es propietario y solo mediante API.

Por qué la brecha por tarea es cuatro veces la brecha por token
The rate card already shows a 5x input gap and a 6x output gap in the vendor's favour, so the direction is not in question. What is worth reading is that the per-task figure is smaller than the per-token figure, which is the opposite of what happened in this batch's other matchups, and the reason is instructive.
Claude Haiku 5.5 emits 162,164 output tokens per Intelligence Index task — 129,047 reasoning and 33,118 answer. Qwen3.8 27B emits 66,797, split 47,711 reasoning and 19,087 answer. The vendor's model spends 2.4 times as many tokens per task, so its 6x output-rate advantage compresses to a 4.7x per-task advantage. It is still enormous. It is just not the number the rate card advertises.
The blend math is a cleaner way to see the shape of it. On an input-heavy 7:2:1 blend — the weight most production traffic actually has — Claude Haiku 5.5 comes to $0.077 per million tokens against $0.470 for Qwen3.8 27B. On a balanced 1:1 blend it is $0.30 against $1.75. The vendor's 100,000-token cliff is the only thing that ever closes this: past it, Claude Haiku 5.5's meters become $0.50 and $2.50 on the whole request, and the two models meet at roughly the same price — at which point you are paying a 9.7-point score premium for nothing.
Donde la tarjeta del proveedor y el panel independiente no coinciden
Este es el conjunto de filas en el que vale la pena detenerse, porque la propia ficha del modelo Qwen3.8 27B se lee bastante más fuerte de lo que indican las mediciones independientes, y esa diferencia es la razón de que una afirmación como «un modelo de 27B que rivaliza con otros mucho más grandes» necesite una segunda fuente.
The vendor's own published figures, as recorded on our catalogue page for the model, include 90.3 on LiveCodeBench v6, 89.2 on GPQA Diamond, 84.3 on OSWorld-Verified, and 79.0 on QwenSWEBench. Those are vendor-reported and unreproduced, and they describe a model that is competitive well above its weight class.
En la ejecución independiente de Artificial Analysis, Qwen3.8 27B obtiene 0.0556 en Terminal-Bench 4.0, 0.4664 en SciCode, 0.3392 en Humanity's Last Exam y 1423.21 en GDPval-AA v2.1. Si se coloca el 0.0556 junto al 84.3 que la ficha del proveedor reporta en OSWorld, la forma del desacuerdo queda clara: en el trabajo agéntico de computadora y terminal, el panel independiente sitúa a este modelo más o menos donde corresponde a un modelo denso de 27B, y la ficha del proveedor no.
Dos filas van en sentido contrario y vale la pena mencionarlas por la misma razón. Qwen3.8 27B obtiene 0,4824 en AutomationBench frente al 0,3541 de Claude Haiku 5.5 — una victoria independiente de 12,8 puntos en lo más parecido que cualquiera de las dos tablas tiene a un benchmark de automatización empresarial. Ese es un dato real de la misma ejecución, en la fila más próxima a aquello para lo que la gente realmente despliega modelos pequeños.
La lectura honesta no es «el proveedor lo infló todo». Es que una ficha de modelo mide las tareas que sus autores eligieron, el panel independiente mide un conjunto fijo, y las fortalezas de Qwen3.8 27B están en la lista del proveedor y no en la del panel.
La economía cambia cuando eres dueño del hardware
Cada tarifa citada anteriormente es un precio de API, y para Qwen3.8 27B ese es el marco equivocado.
Este es un modelo denso de 27B bajo Apache 2.0. Cabe en un único acelerador moderno con la cuantización que la mayoría de los equipos aceptaría, lo que significa que el costo marginal de un token deja de ser el medidor de un proveedor y se convierte en tu propia utilización. Por eso el precio de llamarlo varía según el host de una manera que ninguno de los modelos propietarios de esta comparación podría jamás: el tablero registra una tarifa de terceros de $0.50 de entrada y $3.00 de salida, y el catálogo de OrcaRouter lo ofrece a $0.33 de entrada y $2.40 de salida sin ninguna línea de lectura de caché, porque ejecutamos los pesos en nuestra propia infraestructura en lugar de revender el endpoint de alguien más.
For a team already paying for GPUs, the comparison is not $0.21 versus $1.01 per task. It is the vendor's billed rate against the incremental cost of a request on hardware you own, and those are different numbers. That is the argument for the open-weights side, and it is the only one that survives contact with the benchmark table.
Hay una segunda consecuencia, menos obvia, de que la licencia sea Apache 2.0: el modelo puede distribuirse dentro de un producto, afinarse con tu propio tráfico, fijarse a una revisión específica y auditarse hasta los pesos. Nada de eso está disponible a ningún precio en un modelo propietario que solo se ofrece por API, y para cargas de trabajo reguladas suele ser la restricción decisiva más que una preferencia.

Llamando a cualquiera de ellos
Qwen3.8 27B is on the OrcaRouter catalogue as qwen/qwen3.8-27b at $0.33 and $2.40 per million tokens, self-hosted on our own infrastructure, with the OpenAI-compatible endpoint at https://api.orcarouter.ai/v1. Run against our own traffic over the last week it returns a median 1,837 ms to first token at 319 output tokens per second — a fast model, measured on our playground rather than a standardised harness, so treat that as a serving figure and not a benchmark. Video, image and text input, 262K-token context, native tool calling, structured outputs and a reasoning mode are all supported on the same key as everything else on the catalogue.
Claude Haiku 5.5 is not on the catalogue. It is reachable through the vendor's API directly and through AWS, Azure and the major cloud platforms, and that is the accurate statement. What the catalogue does carry from the vendor's line-up is Claude Sonnet 5.5 and Claude Opus 5.5, which makes the escalation path out of a self-hosted small model and into a hosted mid-tier agentic one a routing change rather than a second integration — automatic failover sits underneath it either way.
El veredicto, que es inusualmente desequilibrado
As an API call on text or images, Claude Haiku 5.5 wins this comparison on every row that is not about licensing. Nine point seven index points, 4.7x cheaper per finished task, four times the context window, twice the response ceiling, and a five-to-six-times lower sticker inside the short-prompt regime. There is no workload-shape argument that rescues Qwen3.8 27B on price, because the vendor's disadvantage — heavy output token use — is worth far less than its rate advantage.
Lo que lo rescata es todo aquello que el precio de la API no captura: una licencia que permite la redistribución, pesos que se pueden conservar y fijar, entrada de vídeo y una vía autoalojada en la que el coste por token es tu propio hardware en lugar de la tarjeta de un proveedor. Si buscas la forma más barata de clasificar, extraer, resumir y enrutar texto, Claude Haiku 5.5 es la respuesta, y la diferencia no es ni mucho menos estrecha. Si lo que buscas es un modelo que puedas poner dentro de tu propio producto, en tu propio hardware, bajo tu propia auditoría, entonces la brecha en los benchmarks dejó de ser la cuestión hace ya varios párrafos.
Comparados en este artículo1
Detectado en este artículo · Benchmarks: Artificial Analysis · actualizado a diario
