Una tarjeta destacada generada titulada 'Claude Haiku 5.5 vs Qwen3.8 27B' con el antetítulo 'PESOS ABIERTOS VS LA API', que muestra Claude Haiku 5.5 con Entrada $0.10, Salida $0.50 e Índice 43.40 frente a Qwen3.8 27B con Entrada $0.50, Salida $3.00 e Índice 33.70, sobre una barra que dice 'Costo por tarea finalizada $0.21 vs $1.01'.
Engineering & Research

Claude Haiku 5.5 vs Qwen3.8 27B: una barrida total en la API, y por qué los pesos abiertos todavía existen

Autor

Gideon Frost

Fecha de publicación

Últimos modelos · 20Ver todos los modelos →
Benchmarks: Artificial Analysis · actualizado a diario
Volver a todas las publicaciones

Most of the comparisons in this series come down to a trade. This one does not, and the absence of a trade is itself the finding. Claude Haiku 5.5, shipped on October 7, 2026, beats Qwen3.8 27B, the open-weight 27B dense model from August 14, 2026, on the composite intelligence score, on cost per token, on cost per finished task, on context window, on response ceiling, on agentic coding, and on the science-reasoning rows — and it does so from a single proprietary API that requires no hardware. The one row Qwen3.8 27B wins outright is video input, and the other is a licence.

Eso no es una razón para descartar el modelo, porque una licencia Apache-2.0 y una modalidad de vídeo no son premios de consolación. Es una razón para ser precisos sobre por qué alguien elegiría el lado de los pesos abiertos, y para dejar de fingir que la discusión gira en torno a los benchmarks.

Los números sobre los que realmente se ejecutaron ambos modelos

Per million tokens in US dollars. The vendor's rates are from its own pricing documentation; the shared measurements are Artificial Analysis Intelligence Index v4.3.2, read 2026-10-08, both in their highest published effort configuration.

A generated two-column scoreboard titled 'Claude Haiku 5.5 vs Qwen3.8 27B — the scoreboard' with rows Index v4.3.2 43.40 against 33.70, cost per task $0.21 against $1.01, input price $0.10 against $0.50, context window 1M tokens against 256K tokens, AutomationBench 0.3541 against 0.4824 and weights 'proprietary' against 'open, Apache 2.0', footed 'All figures per Artificial Analysis Intelligence Index v4.3.2; Qwen vendor-card scores are unreproduced by the board.'

• Entrada — Claude Haiku 5.5 $0.10 hasta 100.000 tokens y $0.50 por encima; Qwen3.8 27B $0.50 según la tarifa de terceros registrada por el consejo.

• Salida — Claude Haiku 5.5 $0.50 hasta 100,000 tokens y $2.50 por encima; Qwen3.8 27B $3.00.

• Entrada en caché — Claude Haiku 5.5 $0.01 y $0.05, un 90 % de descuento; Qwen3.8 27B $0.10, un 80 % de descuento.

• Puntuación independiente — 43,40 para Claude Haiku 5.5 (Max) frente a 33,70 para Qwen3.8 27B (Xhigh). Una diferencia de 9,70 puntos, la más amplia de este lote.

• Coste por tarea finalizada — $0,21 frente a $1,01, en la misma ejecución de evaluación. Una diferencia de 4,7x, también la más amplia que aparece aquí.

• Contexto y salida — 1M de tokens y un límite de respuesta de 128,000 tokens para Claude Haiku 5.5; un contexto de 256K tokens para Qwen3.8 27B.

• Modalidad — ambas aceptan texto e imágenes y devuelven texto. Qwen3.8 27B añade entrada de video; Claude Haiku 5.5 no.

• Pesos — Qwen3.8 27B son 27B de parámetros densos bajo Apache 2.0, descargable. Claude Haiku 5.5 es propietario y solo mediante API.

A capture of Anthropic's Claude Platform models documentation showing the Claude model comparison table — Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5.5 and Claude Haiku 5.5 'This model', the latter reading 1M context, 128K max output, 'From $0.10/$0.50', latency 'Fastest', default effort medium — above the Claude Haiku 5.5 specifications block with model ids claude-haiku-5-5, the two-tier pricing lines and 'Released October 7, 2026'.

Por qué la brecha por tarea es cuatro veces la brecha por token

The rate card already shows a 5x input gap and a 6x output gap in the vendor's favour, so the direction is not in question. What is worth reading is that the per-task figure is smaller than the per-token figure, which is the opposite of what happened in this batch's other matchups, and the reason is instructive.

Claude Haiku 5.5 emits 162,164 output tokens per Intelligence Index task — 129,047 reasoning and 33,118 answer. Qwen3.8 27B emits 66,797, split 47,711 reasoning and 19,087 answer. The vendor's model spends 2.4 times as many tokens per task, so its 6x output-rate advantage compresses to a 4.7x per-task advantage. It is still enormous. It is just not the number the rate card advertises.

The blend math is a cleaner way to see the shape of it. On an input-heavy 7:2:1 blend — the weight most production traffic actually has — Claude Haiku 5.5 comes to $0.077 per million tokens against $0.470 for Qwen3.8 27B. On a balanced 1:1 blend it is $0.30 against $1.75. The vendor's 100,000-token cliff is the only thing that ever closes this: past it, Claude Haiku 5.5's meters become $0.50 and $2.50 on the whole request, and the two models meet at roughly the same price — at which point you are paying a 9.7-point score premium for nothing.

Donde la tarjeta del proveedor y el panel independiente no coinciden

Este es el conjunto de filas en el que vale la pena detenerse, porque la propia ficha del modelo Qwen3.8 27B se lee bastante más fuerte de lo que indican las mediciones independientes, y esa diferencia es la razón de que una afirmación como «un modelo de 27B que rivaliza con otros mucho más grandes» necesite una segunda fuente.

The vendor's own published figures, as recorded on our catalogue page for the model, include 90.3 on LiveCodeBench v6, 89.2 on GPQA Diamond, 84.3 on OSWorld-Verified, and 79.0 on QwenSWEBench. Those are vendor-reported and unreproduced, and they describe a model that is competitive well above its weight class.

En la ejecución independiente de Artificial Analysis, Qwen3.8 27B obtiene 0.0556 en Terminal-Bench 4.0, 0.4664 en SciCode, 0.3392 en Humanity's Last Exam y 1423.21 en GDPval-AA v2.1. Si se coloca el 0.0556 junto al 84.3 que la ficha del proveedor reporta en OSWorld, la forma del desacuerdo queda clara: en el trabajo agéntico de computadora y terminal, el panel independiente sitúa a este modelo más o menos donde corresponde a un modelo denso de 27B, y la ficha del proveedor no.

Dos filas van en sentido contrario y vale la pena mencionarlas por la misma razón. Qwen3.8 27B obtiene 0,4824 en AutomationBench frente al 0,3541 de Claude Haiku 5.5 — una victoria independiente de 12,8 puntos en lo más parecido que cualquiera de las dos tablas tiene a un benchmark de automatización empresarial. Ese es un dato real de la misma ejecución, en la fila más próxima a aquello para lo que la gente realmente despliega modelos pequeños.

La lectura honesta no es «el proveedor lo infló todo». Es que una ficha de modelo mide las tareas que sus autores eligieron, el panel independiente mide un conjunto fijo, y las fortalezas de Qwen3.8 27B están en la lista del proveedor y no en la del panel.

La economía cambia cuando eres dueño del hardware

Cada tarifa citada anteriormente es un precio de API, y para Qwen3.8 27B ese es el marco equivocado.

Este es un modelo denso de 27B bajo Apache 2.0. Cabe en un único acelerador moderno con la cuantización que la mayoría de los equipos aceptaría, lo que significa que el costo marginal de un token deja de ser el medidor de un proveedor y se convierte en tu propia utilización. Por eso el precio de llamarlo varía según el host de una manera que ninguno de los modelos propietarios de esta comparación podría jamás: el tablero registra una tarifa de terceros de $0.50 de entrada y $3.00 de salida, y el catálogo de OrcaRouter lo ofrece a $0.33 de entrada y $2.40 de salida sin ninguna línea de lectura de caché, porque ejecutamos los pesos en nuestra propia infraestructura en lugar de revender el endpoint de alguien más.

For a team already paying for GPUs, the comparison is not $0.21 versus $1.01 per task. It is the vendor's billed rate against the incremental cost of a request on hardware you own, and those are different numbers. That is the argument for the open-weights side, and it is the only one that survives contact with the benchmark table.

Hay una segunda consecuencia, menos obvia, de que la licencia sea Apache 2.0: el modelo puede distribuirse dentro de un producto, afinarse con tu propio tráfico, fijarse a una revisión específica y auditarse hasta los pesos. Nada de eso está disponible a ningún precio en un modelo propietario que solo se ofrece por API, y para cargas de trabajo reguladas suele ser la restricción decisiva más que una preferencia.

A capture of the OrcaRouter model page for Qwen3.8 27B under qwen/qwen3.8-27b, showing a 262K-token context chip, text, image and video input, text output, a p50 time to first token of 1.84 seconds, public benchmarks by Qwen dated 2026-08-15, and the page's own description of the model as Alibaba's open-weight 27B dense multimodal model released under Apache-2.0 and self-hosted on OrcaRouter's own infrastructure.

Llamando a cualquiera de ellos

Qwen3.8 27B is on the OrcaRouter catalogue as qwen/qwen3.8-27b at $0.33 and $2.40 per million tokens, self-hosted on our own infrastructure, with the OpenAI-compatible endpoint at https://api.orcarouter.ai/v1. Run against our own traffic over the last week it returns a median 1,837 ms to first token at 319 output tokens per second — a fast model, measured on our playground rather than a standardised harness, so treat that as a serving figure and not a benchmark. Video, image and text input, 262K-token context, native tool calling, structured outputs and a reasoning mode are all supported on the same key as everything else on the catalogue.

Claude Haiku 5.5 is not on the catalogue. It is reachable through the vendor's API directly and through AWS, Azure and the major cloud platforms, and that is the accurate statement. What the catalogue does carry from the vendor's line-up is Claude Sonnet 5.5 and Claude Opus 5.5, which makes the escalation path out of a self-hosted small model and into a hosted mid-tier agentic one a routing change rather than a second integration — automatic failover sits underneath it either way.

El veredicto, que es inusualmente desequilibrado

As an API call on text or images, Claude Haiku 5.5 wins this comparison on every row that is not about licensing. Nine point seven index points, 4.7x cheaper per finished task, four times the context window, twice the response ceiling, and a five-to-six-times lower sticker inside the short-prompt regime. There is no workload-shape argument that rescues Qwen3.8 27B on price, because the vendor's disadvantage — heavy output token use — is worth far less than its rate advantage.

Lo que lo rescata es todo aquello que el precio de la API no captura: una licencia que permite la redistribución, pesos que se pueden conservar y fijar, entrada de vídeo y una vía autoalojada en la que el coste por token es tu propio hardware en lugar de la tarjeta de un proveedor. Si buscas la forma más barata de clasificar, extraer, resumir y enrutar texto, Claude Haiku 5.5 es la respuesta, y la diferencia no es ni mucho menos estrecha. Si lo que buscas es un modelo que puedas poner dentro de tu propio producto, en tu propio hardware, bajo tu propia auditoría, entonces la brecha en los benchmarks dejó de ser la cuestión hace ya varios párrafos.

Comparados en este artículo1

Detectado en este artículo · Benchmarks: Artificial Analysis · actualizado a diario