
GPT-6.1 Sol's 76x Browser-Agent Result: What Asana Actually Measured
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 118 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 53 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 347 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 59 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 361 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 230 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Asana published a browser-agent cost study on October 8, 2026, and GPT-6.1 Sol’s developer wrote it up the next day under the headline "Asana cuts model costs 76x in browser tests with GPT-6.1 Sol." The 76x is real in the sense that someone measured it, but it is not a fact about GPT-6.1 Sol's price. It is a fact about what happens when you fix a broken prompt cache on a browser agent and then swap the model that sits behind it. The same optimization, run on the model Asana already had in production — a model it calls Model B — cut cost 29x on its own. GPT-6.1 Sol supplied the remaining 2.6x. The experiments were run by GPT-6 Astra working in Codex, and the model behind the baseline is a competitor's model Asana calls Model B, so two tiers from the same vendor and one unnamed rival all appear in the story. What follows separates the part of the result you can copy on Monday from the part that belongs to Asana's particular stack.
The comparison, stated exactly
Asana's stack for this is StackAI, the workflow-automation platform it acquired, running a browser agent that navigates websites, fills forms and gathers information without code. The test task was narrow and concrete: collect six fields for each of 32 books from a public demo catalog. That is representative of what some customers run, and it is also small enough that a 144-run study fits in a week.
The design was six caching and screenshot policies at two history budgets, three runs per condition, across four models — 144 runs, plus a 12-run follow-up. Costs were calculated from each provider's own token counters, and every answer was scored against an independently prepared reference. The four models were three unnamed frontier models (Models A, B and C) and GPT-6.1 Sol. Model A is a smaller, cheaper model from another lab, released Fall 2025, priced at half of GPT-6.1 Sol. Model B is the model that was in production, same lab as A, released Summer 2026, priced the same as GPT-6.1 Sol. Model C is a newer version of Model B, released Fall 2026, also priced the same as GPT-6.1 Sol. The three are unnamed in both write-ups, so the comparison is not reproducible by a reader — worth knowing before you treat 29x as a number about somebody else's model. These are Asana's and OpenAI's published figures, not independently audited ones.
The ladder of results, all Asana's own numbers:
• Baseline production on Model B — at least $36.21 per run, at least 22.5 minutes per run. Some baseline runs hit the step limit before finishing, so the mean is a floor rather than a true average.
• Model B, optimized agent — $1.24 per run, 4x faster than baseline, a 29x cost cut.
• GPT-6.1 Sol, same optimized agent — $0.47 per run, about four minutes, a 76x cost cut and 5x faster.
• GPT-6.1 Sol, before vs after the fix — $1.97 to $0.47 per run, a 4x cut from cache and pruning changes alone.
Because the baseline is a lower bound, 76x is itself a floor. The honest reading is "at least 76x," not "76x."

What actually changed, and why it is not a model feature
The mechanism is prompt-cache mechanics, and it is worth understanding because it transfers to any agent you run on any model. A browser agent resends its tools, its system prompt and its growing history of page text and screenshots on every model call. Prompt caching discounts the repeated part, but only the longest unchanged prefix — the moment anything in the middle of the request changes, reuse breaks from that point on.
Asana's production agent had two faults that compounded. It cached its fixed instructions and tool definitions but not its browsing history. And it edited that history on nearly every step: it dropped the previous screenshot each time and trimmed older text to fit a history budget. Every edit invalidated the prefix, so caching would have been near-useless even if it had been switched on. Asana's write-up notes that on the models tested, cache reads cost 0.05x to 0.1x the standard input price — so the prize was large and the agent was systematically refusing it.
The fix has two halves. First, cache the history too, with a cache marker on the latest tool result. Second, stop editing it every call: keep screenshots and prune them in batches at a ratio of 20 to 1, so the agent holds up to 20 and then cuts back to the most recent one. About 19 consecutive calls then reuse an unchanged prefix. Raise the history budget from 120,000 to 480,000 characters so old text stops being trimmed, and the arithmetic lands: on GPT-6.1 Sol, each call cost roughly 3x less because 89% of input came from cache.
The finding that matters most here is a negative one. Caching the history without batch pruning, at the larger budget, cost more than not caching at all on three of the four models — the cache was continually rewritten and rarely read. Cache infrastructure that is switched on without an append-only history discipline is a way to pay a write premium for nothing. That failure mode is model-independent, and it is the reason the same fix moved Model B by 29x.
Where the model choice actually paid
Strip out the workflow fix and compare like with like: at the same optimized agent, GPT-6.1 Sol ran 2.6x cheaper than Model B, at the same listed price. The extra factor is cache-hit behavior, not the rate card. Sol read 89% of its input from cache; Asana does not publish the equivalent share for Model B, so the 2.6x is a measured outcome without a published decomposition. Treat it as "this model, on this workload, used its cache better," not as a general 2.6x advantage over a model we cannot name.
The runtime side is unambiguous: 5x faster than baseline, with the optimized Sol workflow at roughly four minutes against a baseline of at least 22.5 minutes. Speed matters for cost in agents that bill by token, because a slow model that loops pays for its loops.
For anyone costing this out: GPT-6.1 Sol's standard API rates are $2.00 per million input tokens, $0.10 per million cached input tokens and $10.00 per million output tokens, and its cache-read discount is 0.05x the input rate — the deepest on OpenAI's current card. GPT-6 Astra, the model that did the engineering work in Codex, runs $10.00 input and $50.00 output, which is the five-times gap OpenAI's launch framing leaned on. The reason the study produced $0.47 runs rather than $2 runs is not the rate card; it is that 89% of a very repetitive request was billed at a twentieth of the input rate. On a caching-heavy agent the discount line does more work than the headline price, and on our own catalogue provider list prices pass through with 0% markup, so a vendor moving a cached-input meter moves it on your bill the same day.

The other number in the study: the history budget decided whether the agent answered at all
Cost per run is the number everyone quotes, but the study's more useful result is about reliability, and it is the one an operator should read first.
• At the 120,000-character budget, Model C answered in none of its 18 runs and GPT-6.1 Sol answered in 3 of 18 — most runs hit the step limit without producing an answer.
• At 480,000 characters, every run on both models answered, each with the correct answer.
• The newer models burned the smaller budget faster: Model C first trimmed its history at call 10, Model A at call 64.
• In the optimized workflow, every run completed the task and returned the correct answer, and every best-condition run on every model encountered all 192 facts it was supposed to collect.
That is a different argument from "cheaper." A too-small history budget on a capable model produces an agent that fails by running out of room, and it fails by hitting the step cap, which is the most expensive way to fail — you pay for the whole run and get nothing. Raising the budget raised cost per call and lowered cost per answer, which is the only number a production owner should be tracking. If you are evaluating an unproven model for an agent like this, the low-risk pattern is to keep your production route on the model you trust and put the new one behind a failover or a split route, so a step-limit failure shows up as a routing fact rather than an incident. Every model in this comparison is reachable through one API for 200+ models with provider list pricing passed through unchanged, which also makes the cache-read discount comparable across vendors on the same invoice instead of five dashboards.

What to take from this, in order
If you run a browser or computer-use agent, three levers in Asana's study are worth pulling before you look at a rate card:
• Make the request prefix append-only. Any per-call edit to the middle of the history zeros out cache reuse from that point forward.
• Prune in batches. Asana's follow-up found that keeping every screenshot cost 1.2x less per call than the best pruning condition on Model B and GPT-6.1 Sol, and about 5% less on Model C. Pruning still matters for long tasks, small context windows and expensive cache reads — but the case for a 20:1 batch ratio is about cache stability, not about the screenshots themselves.
• Set the history budget per model, and check it against your step cap. A budget that fits one model may starve the next one.
Two guardrails from the study are easy to skip and shouldn't be. No run reached the 480,000-character budget, so the budget never constrained these runs — but a drifting agent grows toward its context limit, and if the cache breaks, every call pays full price. Caps on steps, tokens and cost per run are what bound a bad run. Separately, the study used three or four runs per condition with variable call counts, which Asana itself says is enough to show broad patterns and not enough to separate conditions a few percent apart. Do not read a 5% delta out of this and rebuild your pipeline around it.
The part that is genuinely new, and the part that isn't
The caveats on the setup are worth stating plainly: four models, three of them unnamed, one narrow 32-book task, Asana's own instrumented counters, and a vendor's own write-up hosting the result. Nothing here has been reproduced outside Asana. But the mechanism is fully specified — append-only history, cache marker on the last tool result, batch pruning, larger budget, measure cache reads with the provider's counters — and it is the kind of finding that survives being attributed to the wrong lab. The 76x is an Asana number on an Asana workload. The reason it is worth reading is that it demonstrates a rule you can test in an afternoon: an agent that edits its own history on every step is paying full price for a conversation it has already had.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
