
Claude Haiku 5.5 vs PPLX 27B: $0.00 a Token Is Not the Same as Free
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 128 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 56 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 347 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
PPLX 27B costs nothing per token. It runs locally, on hardware you own, and once the hardware is bought the marginal cost of the next token really is zero. Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output — which is a number small enough that it is worth doing the subtraction properly rather than accepting the obvious conclusion. A machine capable of running PPLX 27B well is a roughly $4,500 DGX Spark or a desktop with an RTX card of 24GB or more, and $4,500 of Claude Haiku 5.5 output tokens is around nine billion of them. So the question this matchup really asks is not which is cheaper. It is how many tokens you expect to burn, and whether the answer to that changes because the tokens sit on your desk.
Claude Haiku 5.5 is the vendor's small model, released October 7, 2026. PPLX 27B was announced by Perplexity on August 25, 2026 alongside Portable Computer, built with NVIDIA, and it is a post-trained version of the open-weight Qwen 3.8 27B tuned for Perplexity's own harness. Neither is a drop-in for the other, and neither is on OrcaRouter's catalogue.
The two models, and why one of them lives on your desk
Claude Haiku 5.5 — ID claude-haiku-5-5, text and image in, text out, 1M-token context, 128,000-token maximum output with a 300,000-token beta path on the Batch API, June 2026 training cutoff and a commitment not to retire it before October 7, 2027. Effort runs Low through Max with Medium as the default, adaptive thinking is on by default, cache reads cost $0.01 per million tokens and Batch halves the rate. It is a hosted product: you send tokens, you get tokens, you never see a GPU.
PPLX 27B — a 27-billion-parameter dense model derived from Qwen 3.8 27B and post-trained for Perplexity's own agent harness. It runs inside Portable Computer on a DGX Spark or on a Linux machine with an NVIDIA RTX card of 24GB or more, and it does not leave that machine. A hybrid mode added on September 1, 2026 pairs it with a cloud model for heavier work behind an on-device Privacy Gate; Linux support for RTX arrived September 3; Windows support for RTX cards of 24GB and up arrived September 14, together with local MCP and scheduled tasks. The post-trained weights are proprietary and sit behind a Perplexity Pro or Max subscription.

• The dimensions, one line each
• Marginal cost per token — Claude Haiku 5.5 $0.10 per million in and $0.50 out to 100K tokens of prompt, then $0.50 and $2.50; PPLX 27B nothing, once the hardware is paid for.
• Cost to start — nothing beyond an API key for Claude Haiku 5.5; roughly $4,500 for a DGX Spark, or a desktop with an NVIDIA RTX card of 24GB or more, for PPLX 27B.
• Context window — 1,000,000 tokens for Claude Haiku 5.5 against about 262,144 inherited from Qwen 3.8 27B.
• Maximum output — 128,000 tokens for Claude Haiku 5.5, with a 300,000-token beta header on Batch; not published for PPLX 27B.
• Where the data goes — Anthropic's cloud against your own machine, with the hybrid mode's Privacy Gate deciding what leaves it.
• Independent score — Artificial Analysis Intelligence Index 43 for Claude Haiku 5.5, Max configuration 43.40, second of 182; nothing published for PPLX 27B, which Artificial Analysis does not track.
• Output speed — 243.4 tokens per second for Claude Haiku 5.5; dependent on your GPU for PPLX 27B, and not comparable without a published figure.
• Input modalities — text and images against text, inherited from a multimodal base.
• Weights — proprietary in both cases, but for different reasons: no weights released versus a restricted post-train that requires a subscription to obtain.
The hardware is the price, and the price is the honesty test
"Free" is the wrong word for local inference and the vendors know it, which is why the number quoted is always marginal cost. The correct comparison is a break-even: a $4,500 machine at Claude Haiku 5.5's $0.50 per million output tokens equals nine billion output tokens before the hardware pays for itself. A team generating a hundred million output tokens a month reaches that point in about seven and a half years, by which time the GPU is obsolete. A team generating five billion a month reaches it in under two months.
Both answers are correct, and they are correct for different companies. That is the whole reason this comparison cannot be resolved on a spec sheet, and it is also why the arithmetic above ignores everything that makes local inference expensive in practice: the electricity, the rack space, the person who owns the driver updates, and the fact that a machine on a desk is a single point of failure unless you bought two. Add those and the break-even moves further out.
What local inference buys for that money is not cost at all. It is the guarantee that a prompt never leaves the building, which is worth more than any per-token rate to a legal, medical or defence-adjacent team, and no discount Anthropic offers substitutes for it. Conversely, Claude Haiku 5.5's 1M-token context and 128,000-token output ceiling are things you cannot buy at any price on a 24GB card, and a $4,500 machine does not change that.
What the 85.4% actually measures
Perplexity's published number for PPLX 27B is 85.4% on its own 53-task Local Knowledge Work Bench, against 82.6% for the same harness running on the untuned Qwen 3.8 27B base, 77.6% for Pi and 74.0% for Hermes. Three things about that deserve saying plainly, because the launch coverage generally did not say them.
It is vendor-reported and has not been independently reproduced. It is a comparison on the vendor's own harness, which is the fairest possible test for a model post-trained on that harness — the gain over the base model is partly a gain in fitting the test. And it measures local knowledge work specifically: retrieval, summarisation, document handling, the tasks Portable Computer is built for. It is not a general capability score and should not be read as one.
The base model is a different matter. Qwen 3.8 27B is dense, Apache-2.0 licensed, multimodal and released on August 14, 2026 with a 262,144-token native context window, and Artificial Analysis scores its reasoning configuration at Intelligence Index 44 — a point above Claude Haiku 5.5's 43.40, at a different price. PPLX 27B is that model plus Perplexity's post-training, so the honest reading is that the underlying weights are in Claude Haiku 5.5's capability band and the tuning moved them toward one vendor's workload. Which of those two things you are buying is the question the 85.4% is designed not to answer.
Where Claude Haiku 5.5 wins without needing a benchmark
Long context, first: 1M tokens against about 262K, and 128,000 tokens of maximum output against an unpublished figure. Image input, second, which the Qwen 3.8 lineage carries but which Perplexity's harness does not advertise. Effort control, third, and it is the underrated one — the Low setting exists specifically so you can stop paying for reasoning tokens on calls that do not need them, which is a cost lever no local model offers because there is no bill to lower.
And availability, fourth. Claude Haiku 5.5 is a model string on an API, and the independent numbers exist: Intelligence Index 43 overall and 43.40 at Max effort, second of 182, at $0.21 per index task and 243.4 output tokens per second with about 0.3 seconds to first token. When you are deciding whether a workload is ready to move, a published score you did not compute yourself is worth more than a faster number you did.

Where PPLX 27B wins without needing a benchmark
Privacy is the first and most important: the tokens never leave the machine, which no hosted model can match. Cost at volume is the second, and it is real past a few billion output tokens a month. Offline operation is the third — a laptop in a basement with no connection still answers, and Claude Haiku 5.5 does not. And the hybrid mode added on September 1 is the most interesting design in the product: a local model doing the routine work with a cloud model behind an on-device gate for the hard calls is precisely how you get both privacy and capability, and it is the architecture most teams should be copying regardless of which vendor they buy it from.
What PPLX 27B does not offer is a portable answer. It is tied to Portable Computer, it requires a subscription to obtain the weights, and the weights are a derivative of someone else's model — so the licence you actually hold is Perplexity's, not Apache-2.0. If the goal is a model you can fork and ship, the base Qwen 3.8 27B is the one with the permissive licence, but that is a different model from the one this article is about.

Running either of them from one key, and where that stops
PPLX 27B does not come through an API at all: it runs inside Portable Computer on your own hardware, and the hybrid mode reaches a cloud model of Perplexity's choosing on the other side of the Privacy Gate. Claude Haiku 5.5 is available through Anthropic's own API and, from there, wherever Anthropic's models are resold. Neither is on OrcaRouter's catalogue, and we will not suggest otherwise — what we route from these two families is anthropic/claude-haiku-4.5, anthropic/claude-sonnet-5.5, anthropic/claude-opus-5.5 and anthropic/claude-fable-5.1, and separately the Qwen 3.8 27B base that PPLX 27B is post-trained from, which is on the catalogue at qwen/qwen3.8-27b and self-hosted on our own infrastructure.
That last point is the practical one. If the reason you were looking at PPLX 27B was the Qwen 3.8 27B weights underneath it rather than the local execution, you can have the base model through one API for 200-plus models, one key and one bill, with provider list price passed through at 0% markup and automatic failover across providers underneath. If what you actually need is that the prompt never leaves the machine, no gateway is the answer and Portable Computer is.
The decision, without pretending the numbers settle it
If your prompts cannot leave your infrastructure, PPLX 27B is the only one of these two that exists for you, and the $4,500 is not a cost comparison — it is the price of a constraint you cannot negotiate away. If you are burning more than a few billion output tokens a month, the local route pays for itself inside a quarter and then keeps paying.
If neither of those is true, Claude Haiku 5.5 is the better buy at almost any volume: no hardware, no ops, a 1M-token window, a 128,000-token output ceiling, image input, an effort dial that turns cost down on demand, a published independent score of 43, and a price of $0.10 and $0.50 per million that makes the break-even against a workstation about nine billion tokens out. The $0.00-per-token figure is true. It is just not the whole subtraction.
If the reason you were looking at PPLX 27B was the Qwen 3.8 27B weights underneath it rather than the local execution, you can have the base model through one API for 200-plus models, one key and one bill , with provider list price passed through at 0% markup and automatic failover across providers underneath.
