
GPT-6.7 and the "Bob" Question: Will OpenAI's Unreleased Flagship Really Run 2x Slower?
- DeepSeekNEWDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.15 / $0.29 per 1M tokens
- z-aiNEWZ.ai: GLM 5.32026-08-1860Intelligence75Coding
- obsidianNEWQwen3.8 27B2026-08-1552Intelligence68Coding
- qwenNEWQwen: Qwen3.8 27B (free)2026-08-13qwen/qwen3.8-27b-free
- deepseekNEWDeepSeek: DeepSeek V4 Pro 08132026-08-1253Intelligence69Coding
- grokNEWSpaceXAI: Grok 4.62026-08-1261Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0557Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0358Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3152Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2463Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2152Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2137Intelligence49Coding
- metaMeta: Muse Spark 1.12026-07-1653Intelligence71Coding
- kimiMoonshotAI: Kimi K32026-07-1560Intelligence76Coding
- openaiOpenAI: GPT-5.6 Luna2026-07-0952Intelligence71Coding
- openaiOpenAI: GPT-5.6 Terra2026-07-0957Intelligence77Coding
- openaiOpenAI: GPT-5.6 Sol2026-07-0961Intelligence77Coding
On August 15, an X post from the account @Yuchenj_UW compressed a rumor into one sentence: "Bob, the king of GPU kernels, left the company? GPT-6.7 will run 2x slower…." Two claims ride on that question mark — that the company's most secretive engineer has left, and that his departure halves the serving speed of GPT-6.7, the company's unreleased flagship (renamed GPT-6-7 last October). Neither claim is verified. But the post lands at a moment when the company's next model is already a question of when, not if: the Astra family name arrived August 1, a safety review paused it August 7, and GPT-5.6 Sol — the current flagship, live on OrcaRouter — is the only concrete reference point for what the next generation has to beat. Here is what is actually knowable as of today.
One thing up front: nobody outside OpenAI has run GPT-6.7. Any claim about how fast it will serve tokens is extrapolation — including the "2x slower" figure. What follows separates the verified facts (the rename, the release timeline, the departures that are on the record) from the legend (Bob, the kernel, and the number). All of the rumor material is labeled as such.
Who "Bob" is, and why a single engineer became a legend
"Bob" is the nickname OpenAI employees use for an engineer whose name the company has never confirmed. Press reports from September 2025 — including QbitAI, The Paper and itbear, all citing former and current OpenAI staff — describe a secretive researcher who single-handedly writes the CUDA kernels used for inference, the low-level GPU code that turns transformer math into tokens. His attention kernel, called the "Bob kernel" inside OpenAI, is reportedly executed trillions of times a day across hundreds of thousands of GPUs, with the precision to match: a bug would force a rollback of checkpoints and a retraining cycle.
The legend is consistent across accounts. Former employees say Bob fixed in minutes a performance problem colleagues had wrestled with for a week. OpenAI's internal Slack has a "Bob magic" emoji. Industry estimates cited in the same reporting put the number of people alive who can write high-performance training CUDA kernels at under a hundred — which is why a single-person dependency of this kind is treated as genuinely risky rather than a nice-to-have.
On identity, OpenAI has never said who Bob is. The leading hypothesis, repeated across the same press coverage, is Scott Gray — a physics-and-computer-science graduate who joined OpenAI in 2016, first-authored OpenAI's 2017 blog post "Block-sparse GPU kernels," and later co-authored the GPT-4 technical report and the scaling-laws paper, with a bibliography of 51 papers and more than 80,000 citations. The fit is circumstantial but widely repeated; it is an inference, not a confirmation.
Did Bob actually leave?
The tweet is a question, not a claim. But the departure it points at has a paper trail. A leadership-churn roundup published August 15, 2026 lists Scott Gray — described as a "longtime GPU systems engineer" who joined in 2016 and helped build the infrastructure behind sparse transformers, scaling laws, and GPT-3 — among at least twelve senior people to leave OpenAI in 2026. The same list names chief revenue officer Denise Dresser, former COO Brad Lightcap, applications CEO Fidji Simo, Sora head Bill Peebles, and heads of safety, ethics, marketing, and hardware. That is a secondary report, not an official confirmation, and it gives no departure date for Gray.
The broader context makes the rumor less surprising. Reporting around the same week frames the churn as a pre-IPO shakeup: co-founder Greg Brockman is pulling operating authority toward himself ahead of a planned public listing, and CNBC reported the CRO's exit on August 13. And "Bob" has been a known recruitment target for a year — reports from September 2025 said Meta's Mark Zuckerberg had made "who is Bob?" a standing item at hiring meetings. The talent war around this specific skill set is real, which is part of why the leak reads as plausible.
Still, the claim chain has two unverified links: that Scott Gray is Bob, and that the roundup's listing is the same departure the tweet means. It is entirely possible the tweet is repeating a rumor that predates any actual exit. This section is the honest summary: on the record, a senior GPU engineer named Scott Gray appears on a 2026 departure list; off the record, that person may or may not be the Bob the tweet means.
Why "2x slower" would matter — and why it's probably not that simple
Kernels determine the two numbers that decide a model's economics: tokens per second and cost per token. If GPT-6.7 shipped on serving kernels half as efficient as the hypothetical "Bob-optimized" version, the same hardware would serve half the tokens and the same request would take twice as long — at roughly twice the marginal cost. For a developer routing production traffic, that is a 2x latency and 2x cost hit on the flagship, the exact kind of number that changes a model choice. That is why the rumor has teeth.
Now the reasons to distrust it:
• The "2x" has no stated basis. It is a number in a post — a thought experiment, not a measurement, and the linked image does not carry one either.
• Kernels are code, and code outlives authors. The serving stack Bob (or anyone) wrote for earlier models does not disappear when a person leaves; the next model inherits most of it.
• OpenAI's infrastructure organization is large, and the "one person writes everything" legend is almost certainly a simplification. Key-person risk is real but rarely binary.
• Even a genuine efficiency hit would not change capability. It changes cost and latency — painful, but a vendor can partially absorb it or price around it, and OpenAI has already shipped a speed feature (Ultrafast, up to 14x faster on the current flagship) aimed at exactly this kind of serving-cost pressure.
The strongest version of the claim is the under-a-hundred statistic. If the person who wrote the flagship's serving kernels is gone, the replacement is not a headcount hire — it is a multi-year ramp. That is a legitimate risk to OpenAI's serving economics. But it is a risk about the team, not a measured fact about a model that has not shipped.
What we actually know about GPT-6.7
The model the rumor is about has a longer public trail than the tweet suggests, and most of it is recent:
• Oct 31, 2025 — Sam Altman announced on X that "GPT-6" would be renamed "GPT-6-7," widely read as a nod to the Gen-Alpha slang "6-7." Whether "GPT-6.7" and "GPT-6-7" are strictly the same thing is not official; the tweet treats them as interchangeable, and so does most coverage.
• April 2026 — Reports that GPT-6 pre-training was complete, with a rumored release around April 14. That date passed without a launch.
• August 2026 — Leaks said the launch was pulled forward to early August, with OpenAI skipping GPT-5.7 through GPT-5.9. Unverified.
• Aug 1, 2026 — OpenAI announced "Astra" as the name of its next model family, via a research report claiming ten open math problems solved. Whether Astra ships under the GPT-6 name is unconfirmed.
• Aug 7, 2026 — OpenAI paused Astra development and delayed rollout after an internal safety review flagged "critical" cybersecurity risk — the first model to trigger that tier. Altman said broad release "needs a little bit longer."
• Prediction markets — As of mid-August, traders put a GPT-6 release by Sept 30 at roughly 62% and by Dec 31 at about 90%.
• Rumored specs — around a 2M-token context, roughly 40% better performance than the current generation, personalized memory, and native multimodal architecture under the codename "Symphony." All unverified.
So the model behind "2x slower" has not shipped, has no confirmed release date, and is the subject of an active safety review. A serving-speed number attached to it is speculative twice over — once for the date, once for the speed.

The only scoreboard that's honest today
Because GPT-6.7 does not exist yet, a scoreboard can only contrast the rumor with what is actually live:
• GPT-6.7 (unreleased) — status: not shipped; release: rumored autumn 2026; context: ~2M tokens (rumored); speed: "2x slower?" unverified; price: n/a; routable: not yet.
• GPT-5.6 Sol (live) — status: shipping; context: 1.05M tokens, 128K max output; speed: ~290 tok/s observed on OrcaRouter's 7-day stats; price: $5.00 in / $30.00 out per million tokens at the standard input tier; routable: yes, at provider list price.
The interesting gap is the price column. OpenAI's current flagship costs $5/$30 per million tokens through OrcaRouter, and the pass-through rate is the vendor's own list price with zero markup — so whatever OpenAI prices GPT-6.7 at, the number on our side moves the same day it ships, with no renegotiation. That is the useful part of the scoreboard for anyone planning around the next model.

How to treat a leak like this
The pattern is a familiar one: an unreleased model, a dramatic number, a named person. Each part is plausible and none is verified. The right response is neither to ignore the rumor nor to bet on it — it is to structure the decision so you do not have to choose.
That is the case for routing. When GPT-6.7 ships, the way to evaluate a 2x-slower claim without betting a production path on it is to put it behind the same endpoint you already use: one API key, automatic failover to GPT-5.6 Sol the moment the new model underperforms, and real traffic deciding whether the rumor was true. The claim then becomes a measurement instead of a tweet. If OpenAI's price list changes with the new release, the pass-through rate updates the same day — OrcaRouter passes provider list prices through at zero markup, so a vendor price cut (or hike) is live here immediately, no renegotiation.
Meanwhile the current flagship is the concrete thing. GPT-5.6 Sol is live on OrcaRouter today at $5/$30 per million tokens, with the 1.05M-token context and ~290 tok/s the scoreboard shows. The Ultrafast mode announced August 13 — up to 14x faster, roughly 750 tokens per second on Cerebras infrastructure — applies to that model, not to GPT-6.7, and it is a reminder that OpenAI can attack serving cost with infrastructure as well as with kernels.

What we're watching now
• Whether Bob's departure gets confirmed. OpenAI does not comment on who Bob is; the closest on-record event is Scott Gray's name on the 2026 departure list. If OpenAI or Gray confirms it, the first link of the chain becomes fact.
• Whether the "2x slower" number ever gets a basis. If GPT-6.7 ships and independent speed measurements appear, the figure stops being a rumor — either confirmed or retired. Until then it is a number in a post.
• Whether Astra's safety pause slips the date. Prediction markets already moved GPT-6 out of August; the "critical" designation and the government testing OpenAI says it wants are both natural sources of further delay.
• Whether the kernel bench re-sorts. If OpenAI's serving team absorbs the loss — or the legend overstates one person's share — the whole rumor is moot, and the first sign will be the next model's serving performance relative to its predecessor.
The honest read: "2x slower" is a compelling story and, so far, only that. The kernel-wizard departure is a real data point with a real paper trail; the speed number is a number in a tweet. Until GPT-6.7 ships, the defensible position is to plan for both outcomes — assume the rumor could be partly right, and build the routing so a slower-than-expected flagship costs you nothing to test. That is the entire point of routing a model you have not met.
