
PPLX 27B Launches in Perplexity's Portable Computer: A Local Agent That Beats the Open Harnesses
- OrcaNEWOrca: OrcaCyber Zero 1.52026-10-10$3.00 / $7.50 per 1M tokens · 87 tok/s
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 115 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 47 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 777 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 60 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 450 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
Perplexity's PPLX 27B is not a cloud model, and that is the entire story. Announced August 25, 2026, alongside Portable Computer — the company's local-first agent, built with NVIDIA — PPLX 27B is a post-trained version of Alibaba's open-weight Qwen 3.8 27B, tuned by Perplexity for its own agent harness and shipped to run entirely on hardware you own. On Perplexity's 53-task Local Knowledge Work Bench, the harness running PPLX 27B scored 85.4%, ahead of the same harness on Qwen 3.8 27B (82.6%) and of two open-source agent harnesses, Pi (77.6%) and Hermes (74.0%). Those are the vendor's own numbers from the launch post and the companion research paper, "A Local-First Agent for Private and Cost-Effective Knowledge Work," not independently reproduced — but they are the first public evidence that a 27B model on a desktop can carry real knowledge work at a fraction of the cloud's price. One week later, on September 1, the same post-trained model turned up in a second deployment: Hybrid Compute, the new Mac mode of Perplexity's Computer agent, which splits a single task between cloud frontier models and PPLX 27B running on Apple silicon — behind a new on-device Privacy Gate that scans for personal data before anything leaves the machine. Two days after the Mac announcement, on September 3, Perplexity said Portable Computer is also now available on Linux for NVIDIA RTX GPUs with 24GB of VRAM or more. On September 14, the company said it is now available on Windows PCs with NVIDIA RTX GPUs — the widest reach the local-first pitch has had, per the company, and the step that makes running the agent on your own hardware a real deployment option on the most common desktop platform. Both claims are vendor-stated.
What actually shipped
Portable Computer is the on-device build of Perplexity's "Computer" agent platform. Where the earlier Computer product ran its orchestrator in the cloud, Portable Computer packages the whole stack into one system that runs on your machine: the agent harness, orchestrator, planner, tool router, scheduler, a durable task queue, a local search index, the inference engine, and the app connectors for Google Drive, Gmail, Slack, GitHub, and Outlook.
Two properties separate it from a DIY local setup. First, code and tool execution happen inside an OS-enforced sandbox, and if the sandbox is unavailable, tool execution is disabled rather than run unprotected. Second, it is local-first by design: every task starts on-device, and if a step needs live web data or frontier reasoning, the orchestrator stops and asks for permission before escalating that single step to one of 15+ cloud models. An on-device PII classifier — the technology Perplexity has since productized as the Privacy Gate — runs over the outgoing context first and shows you exactly what would leave the machine.
The two 27B models
At launch you choose between two 27-billion-parameter models. Qwen 3.8 27B is Alibaba's Apache-2.0 open-weight model — a dense, multimodal 27B with a 262K-token native window, released two weeks earlier and already a strong self-host option. PPLX 27B is that same base post-trained by Perplexity: not a new architecture, but additional training tuned specifically for this harness, with the company claiming more precise and more efficient behavior on its own agent workloads. NVIDIA's Nemotron 3.5 Lightning (30B) is listed as coming soon, and bring-your-own model and inference server are both supported, so the harness is not married to Perplexity's weights.
The benchmark, and what it is not
The headline figure comes from the 53-task Local Knowledge Work Bench, a suite spanning the sort of work a knowledge worker would actually delegate — deep research, financial analysis, document creation. Perplexity says it plans to open-source the benchmark, which will eventually make the comparison repeatable rather than trusted on faith. Until then, these results are vendor-run. On that suite, per the vendor:
• Portable Computer with PPLX 27B — 85.4%
• Portable Computer with Qwen 3.8 27B — 82.6%
• Pi (open-source harness) — 77.6%
• Hermes (open-source harness) — 74.0%

Two more vendor-reported results round out the picture. On BrowseComp, the web-research suite, Computer scored 66.7% against 50.2% for Pi and 43.9% for Hermes, while using 51% less wall time and 70% fewer tokens than Pi. On ParseBench-100, a document-extraction test, it scored 65.1% against 34.6% for Hermes and 13.9% for Pi. The widest gaps land on the tasks that reward a harness's own infrastructure — search, routing, tool use — which is where a post-trained model plus a purpose-built plumbing stack earns its keep.
The economics flip
The more interesting number is the price. Work completed locally carries zero per-token cost and consumes no Perplexity credits — the counter sits at zero for a fully local task. Escalation is the only thing that costs money, and it is opt-in per step. Perplexity published the hybrid math on Terminal Bench 2.1: fully local scored 59.6% at roughly zero cost; adding a frontier model as an advisor lifted it to 73.0% at about $0.415 per rollout; the frontier model alone reached 82.4% at about $0.65 per rollout. That last contrast is the thesis — local inference flattens the cost curve for always-on agents, and cloud escalation becomes a surgical expense where it pays for itself.
None of which is free to start. The box matters. Portable Computer launched Linux-first on August 25 on the NVIDIA DGX Spark, the desktop supercomputer with 128GB of unified memory (list price around $4,500), or on Linux PCs with RTX GPUs of at least 24GB VRAM — roughly an RTX 3090 or newer — with 32GB recommended. On September 3, Perplexity said Portable Computer is now available on Linux for NVIDIA RTX GPUs with 24GB of VRAM or more — per the company — and NVIDIA's IFA 2026 rollout that same week added one-click installers and simplified local-AI setup aimed squarely at RTX GPUs with 24GB or more of VRAM, naming Portable Computer alongside the Hermes agent and OpenClaw. Windows support arrived as scheduled: on September 14, Perplexity said Portable Computer is now available on Windows PCs with NVIDIA RTX GPUs — per the company — in the Perplexity Windows app, with on-device inference requiring an RTX GPU with 24GB of VRAM or higher, matching the Linux bar, and with two additions the Linux route lacks: local MCP for connecting your own tools and app integrations, and scheduled tasks that let Computer run recurring work on your PC while you are away. macOS arrived in a different form a week after the Linux launch — Hybrid Compute, covered next. And the software itself requires a paid Perplexity plan: Pro, Max, Enterprise Pro, or Enterprise Max. This is a subscription-plus-hardware product aimed at people who already pay Perplexity, not a general API.
The Mac turn: Hybrid Compute and the Privacy Gate
On September 1, Perplexity Computer — the agent platform Portable Computer is built from — gained Hybrid Compute in the Mac desktop app. Here the split runs the opposite way from the Linux build: a task starts in the cloud, where frontier models handle the hardest reasoning, web research, and long-horizon planning, and the orchestrator hands sensitive sub-steps to a local subagent running on Apple silicon. Files, local data, and device actions stay on the Mac, context survives the handoff, and none of the local tokens ever reach the cloud. The local engine is the model this piece is about — PPLX 27B, Perplexity's post-trained Qwen 3.8 27B, listed in the company's Mac materials as PPLX Qwen 3.8 27B and installed with a one-click download in the app, no Ollama or terminal setup, per Perplexity. Launch coverage differs on the rest of the lineup: most outlets add Google's Gemma 4 E4B and Alibaba's Qwen3.6 35B-A3B as other on-device options, and one names a Perplexity-post-trained Qwen 3.6 35B instead of the 27B — a detail the press did not land on consistently on day one.
The feature Perplexity is leading with is the Privacy Gate, an on-device classifier trained with its Secure Intelligence Institute. It reads each task — prompt, tool output, memory, logs — before anything is transmitted, looking for personally identifiable information: names, addresses, account numbers, credentials, payment card numbers, and government IDs. Detected values are swapped for stand-ins before the request leaves the Mac and restored when the answer returns; when the gate flags something, the user decides whether that step runs locally, is masked, or is shared — the gate asks, it does not decide. Perplexity also says it has open-sourced the classifier and released PII-TRACE, a benchmark built from 13,148 synthetic conversations across 13 languages for evaluating PII detectors. Those are company claims, not independently audited.
Hybrid Compute runs on any Apple silicon Mac with macOS 15 or later and 24GB of unified memory minimum — 32GB recommended, and Perplexity says 8GB and 16GB machines are out. It is available to Pro, Max, and Enterprise subscribers through the desktop app, with enterprise accounts opting in and admins able to set an organization-wide sensitivity policy and see a record of what leaves each device. Local work uses no cloud credits and locally generated tokens are not charged — only the cloud orchestration and delegation steps cost anything, and no API key is needed. The caveats match the Linux launch: the gate is itself a model, so false negatives are possible, and reviewers were quick to note the protection is only as good as the choices a user makes when prompted. AppleInsider, in particular, advised testing the feature before trusting it with sensitive data.

Where the router fits
PPLX 27B is not something OrcaRouter routes — it runs on hardware you own, behind a Perplexity subscription, and we won't pretend otherwise. What we do route is the comparison set this launch invites. Qwen 3.8 27B, the exact base weights PPLX 27B post-trains from, is live on OrcaRouter self-hosted; so are DeepSeek V4 Flash, Gemini 3.5 Flash Lite, and GPT-5.6 Luna — all behind one API key at provider list price, passed through with zero markup, so a vendor price cut is live on our side the same day it is announced.
That matters here for one specific reason. Portable Computer's pitch is the local-versus-cloud cost contrast, and the cloud half of that contrast is only as honest as the prices you compare against. On OrcaRouter those are the provider's actual list prices, which turns the local-or-cloud decision into a calculation you can trust rather than a marketing comparison. And if you want to prototype the same agent both ways — local harness on one box, cloud models behind one endpoint — automatic failover lets the cloud side stand in when the local model is out of depth, without a second contract or a second code path. Hybrid Compute repeats the same bargain — local tokens free, cloud orchestration the only spend — so the calculation applies whether the local half runs on a DGX Spark, a Linux RTX workstation, a Mac, or — since September 14 — a Windows PC with an RTX GPU.

What to watch
Four things over the next month decide whether this is a niche product or the start of a category. Whether Perplexity actually open-sources the Local Knowledge Work Bench, which would turn the headline numbers from claims into something the community can reproduce. Whether Nemotron 3.5 Lightning lands and beats PPLX 27B on the same harness — the "post-train for the harness" thesis is only as good as the base model beneath it. Whether the Privacy Gate survives adversarial scrutiny now that the classifier is open-sourced — a probabilistic PII detector is the linchpin of the hybrid pitch only if it actually keeps names, account numbers, and credentials off the wire. And whether the reach the Linux launch never had finally arrives through the September widening — the RTX-24GB Linux route confirmed on September 3, Windows support confirmed on September 14 in the Perplexity Windows app, and Hybrid Compute on Mac — putting the same local-agent economics on the widest consumer hardware Perplexity has shipped yet. If the benchmark is real and the harness transfers, "run the agent on your own hardware" stops being a hobbyist pattern and becomes a deployment option with real teeth — and the cloud models it gets compared against will have to earn their per-token prices.
