
GPT-5.6 Sol Ultrafast: 14배 속도, 750 Tok/s — CS-4는 약 1,300으로 알려짐, 아직 가격 미정
- z-aiNEWZ.ai: GLM 5.32026-08-1860지능75코딩
- obsidianNEWQwen3.8 27B Uncensored (Aggressive)2026-08-1552지능68코딩
- qwenNEWQwen: Qwen3.8 27B (free)2026-08-1341 tok/s
- deepseekNEWDeepSeek: DeepSeek V4 Pro 08132026-08-1253지능69코딩
- grokNEWSpaceXAI: Grok 4.62026-08-1261지능77코딩
- metaNEWMeta: Muse Spark 1.22026-08-0557지능72코딩
- qwenQwen: Qwen3.8 Max2026-08-0358지능72코딩
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3152지능69코딩
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 100만 토큰당 · 237 tok/s
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2463지능78코딩
- googleGoogle: Gemini 3.6 Flash2026-07-2152지능69코딩
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2137지능49코딩
- metaMeta: Muse Spark 1.12026-07-1653지능71코딩
- kimiMoonshotAI: Kimi K32026-07-1560지능76코딩
- openaiOpenAI: GPT-5.6 Luna2026-07-0952지능71코딩
- openaiOpenAI: GPT-5.6 Terra2026-07-0957지능77코딩
- openaiOpenAI: GPT-5.6 Sol2026-07-0961지능77코딩
- grokxAI: Grok 4.52026-07-0856지능72코딩
GPT-5.6 Sol Ultrafast mode is the hardware-accelerated serving tier for the flagship — the same full GPT-5.6 Sol model run on Cerebras wafer-scale chips instead of GPU clusters, delivering up to 14× the speed of standard processing and up to 750 output tokens per second on the tier OpenAI announced in August. On 2026-08-18, Cerebras unveiled the next-generation CS-4 system that this tier grows onto, and an early, unverified report claims GPT-5.6 Sol on CS-4 already serves roughly 1,300 tokens per second. It is not a smaller model and not a distillation: only the hardware and scheduling changed, and the intelligence is preserved. What it does not yet have is a price — as of 2026-08-19, Ultrafast is a limited API preview with no published rate card, so the number you can actually budget against today is the standard tier on the live GPT-5.6 Sol model page: $5 per million input tokens and $30 per million output, passed through at zero markup. This article covers what the mode is, how fast the numbers really are — including the new CS-4 hardware and what is still unverified — what it costs (including the honest "no price yet" answer), and what to do until it reaches general availability.
Ultrafast 모드가 실제로 무엇인지
Ultrafast는 새로운 모델이 아닌 서빙 티어입니다. OpenAI는 2026년 8월 13일, 칩 제조사 Cerebras와 2026년 1월 체결한 컴퓨팅 파트너십의 첫 번째 제품으로 이를 발표했습니다. 이 계약은 3년간 약 100억 달러 규모로 알려졌습니다. 표준 GPT-5.6 Sol 추론이 GPU 클러스터에서 실행되며 메모리와 연산 장치 사이에서 가중치를 이동하는 데 많은 시간을 소비하는 반면, Ultrafast는 모델의 가중치를 Cerebras의 웨이퍼 스케일 칩에 있는 44GB 온칩 SRAM에 로드합니다. Sol 크기의 모델은 여러 웨이퍼에 걸쳐 계층으로 펼쳐지므로 가중치는 연산이 이루어지는 곳에 위치합니다. 그 결과 처리량 상한은 메모리 대역폭이 아니라 실리콘에 의해 결정됩니다.
The hardware story moved on 2026-08-18. At its Supernova event, Cerebras unveiled the CS-4, the next-generation rack-scale system built on its Nexus platform — three Wafer-Scale Engine chips per rack, up to 2× the token-generation speed of the prior CS-3, and, per Cerebras, more than 1,000 tokens per second on models above 10 trillion parameters. Those are Cerebras's claims for a system in early access, not yet generally available. Since the unveiling, a single post on X claims CS-4 already serves GPT-5.6 Sol at roughly 1,300 tokens per second — consistent in direction with Cerebras's own numbers, but confirmed by nobody so far. Treat it as an early signal: plausible, unverified, and worth re-checking before you plan capacity around it.
코드에 중요한 부분은 다음과 같습니다: 모델 ID, 가중치, 추론 동작, 출력이 표준 GPT-5.6 Sol과 동일합니다. OpenAI는 Ultrafast를 품질 등급이 아니라 "초당 더 유용한 작업"으로 규정하며, 프리뷰는 전체 플래그십을 실행합니다. 속도는 더 저렴한 모델로 교체해서 얻는 것이 아닙니다. 변하는 것은 토큰이 나오는 속도뿐입니다.
Ultrafast를 기존 fast mode와 혼동하지 마세요. Fast mode는 표준 하드웨어에서 별도로 구매할 수 있는 등급으로, 토큰당 2배의 추가 요금으로 출력 속도가 약 2.5배까지 빨라지며(1M당 입력 $10 / 출력 $60), 품질 변화는 없습니다. Ultrafast는 다른 것입니다 — Cerebras 하드웨어로, 최대 14배, 그리고 현재는 가격이 전혀 없습니다.
얼마나 빠른가 — 중요한 수치들

헤드라인 수치는 14배이며, 이에 대한 정의가 필요합니다. OpenAI는 Ultrafast가 표준 GPT-5.6 Sol 처리 대비 초당 최대 750개의 출력 토큰을 생성한다고 밝혔습니다. 이는 출력 토큰에 대한 처리량 수치이지, '모든 것이 14배 더 빠르게 실행된다'는 단순한 주장이 아닙니다 — 엔드투엔드 요청 시간에는 입력 처리와 모델 자체의 추론도 포함되므로, 14배는 최대치로 표시된 것입니다.
• 최대 출력 처리량 — OpenAI 발표(2026-08-13)에 따르면 초당 최대 750개 출력 토큰.
• 표준 처리 대비 — 동일한 발표에 따르면 최대 14배 더 빠르며, 실제 향상 폭은 입력 길이와 작업 유형에 따라 다릅니다.
• Humanity's Last Exam, 질문 2,500개 — OpenAI/Cerebras는 GPT-5.6 Sol Ultrafast가 11시간 11분 만에 완료했다고 보고했으며, Claude Fable 5는 78시간 27분이 걸렸고 정확도는 비슷했다. 공급업체 보고 자료이며, 비교가 두 공급업체의 하드웨어 스택에 걸쳐 있으므로 확정적인 결과가 아닌 방향성 참고용으로 읽어야 한다.
• GDP-Val, 종단 간 — Cerebras는 전체적으로 5.6배 더 빠르며 측정 가능한 품질 저하가 없다고 보고합니다. 또한 공급업체 보고 수치입니다.
• 빠른 모드 기준 — 기존에 구매 가능한 속도 등급은 2배의 가격 프리미엄으로 약 2.5배의 출력 속도가 최대치입니다.
• CS-4, the next hardware — Cerebras claims the CS-4 rack delivers more than 1,000 tokens per second on models above 10 trillion parameters, up to 2× CS-3's token generation. An unverified community report puts GPT-5.6 Sol on CS-4 at ~1,300 tokens per second — same direction, not yet confirmed by OpenAI or Cerebras.
One caution before you quote the 14×: independent coverage of the launch cites an ~11× generation-speed comparison against Claude Fable 5, while Cerebras's own figures imply ~7× for total test time on the same run — the difference is which portion of a workload you measure. The same discipline applies to the CS-4 speed signal: the ~1,300-token/s figure started as a single X post, and neither OpenAI nor Cerebras has confirmed it. All of these are vendor or community numbers announced this week, and none are independently verified yet. Treat them as claims, not benchmark verdicts.
비용은 얼마인가 — 그리고 솔직한 격차

Here is the state of the price question on 2026-08-19, stated plainly: Ultrafast has no published price. OpenAI has not announced one, and the preview does not appear on any public rate card. Reports that early-access slots are bundled or free are speculation, not policy — and until OpenAI prices the tier, any "GPT-5.6 Sol Ultrafast cost" number you find elsewhere is a guess.
오늘 가격을 책정할 수 있는 것은 GPT-5.6 Sol 라인의 나머지로, OpenAI가 2026-08-18에 게시한 요율과 대조하여 검증되었습니다:
• 표준 GPT-5.6 Sol — 1M 토큰당 입력 $5.00 / 출력 $30.00. 지금 바로 호출할 수 있는 기준 모델입니다.
• 빠른 모드 — $10.00 / $60.00, 대략 2.5배의 출력 속도를 위한 2배 프리미엄. 실제로 구매할 수 있는 속도 등급에 가장 가까운 것입니다.
• 프롬프트 캐시 읽기 — 1M당 $0.50 (신규 입력 대비 90% 할인); 캐시 쓰기는 1M당 $6.25로 청구됩니다.
• 롱 컨텍스트 티어 — 약 272K 토큰을 초과하는 입력은 $10.00 / $45.00으로 과금됩니다.
• Batch API — $2.50 / $15.00, 약 24시간 처리 시간으로 일괄 50% 할인.
기대할 프리미엄에 대한 맥락: fast mode는 표준 요금의 2배에 2.5배 속도를 제공하는 선례를 세웠습니다. Ultrafast도 비슷한 곡선을 따른다면 표준 $5 / $30보다 훨씬 높은 수준에 위치할 것입니다. 그리고 프리뷰에 대한 논평은 정확히 그렇게 예상합니다. Cerebras의 웨이퍼 용량은 희소하고 비싸기 때문입니다. 그러나 예상되는 프리미엄은 가격이 아닙니다. 요율표가 나올 때까지 표준 Sol 또는 fast mode를 기준으로 계획하세요.
사용 방법 — 미리보기의 실제
Access는 두 번째 진짜 격차입니다. Ultrafast는 OpenAI API를 통해 대기자 명단 방식으로 선별된 고객 그룹에게 제공되며, 2026-08-13에 발표되었습니다. OpenAI는 액세스가 "용량이 늘어남에 따라" 확대될 것이라고 밝혔습니다. 일반 공개 날짜는 없습니다. 발표된 프리뷰 코호트는 코딩, 커머스, 금융 리서치, 지원 및 기타 대화형 워크로드입니다. 이러한 팀들의 에이전트는 요청당 많은 호출을 수행하며 응답 지연 시간이 병목 지점입니다. Jane Street, Rogo, Podium이 명시된 초기 테스터 중에 포함됩니다.
이것이 설계된 사용 패턴은 다음과 같습니다: 장애 대응(장애가 진행 중일 때 로그, 트레이스, diff를 읽는 작업), 실시간 데이터에 대한 금융 조사 및 사기 탐지, 실시간 고객 지원 및 음성(0.5초의 침묵도 끊긴 것으로 읽히는 환경), 전자상거래 결제, 그리고 야간 조사 배치를 인터랙티브 세션으로 압축하는 작업입니다. 워크로드가 단일 턴 채팅이라면 14×의 차이는 40회 호출 에이전트 루프에서보다 훨씬 덜 중요합니다.
오늘 할 수 있는 일 — 실용적인 답변

Ultrafast가 프리뷰 단계인 동안, 정식 GPT-5.6 Sol이 지금 바로 제공됩니다. OrcaRouter에서는 표준 티어가 공급자 요율로 제공되며 토큰당 마크업이 $0입니다. 모델 ID는 openai/gpt-5.6-sol이며, OpenAI 호환 엔드포인트 api.orcarouter.ai/v1을 통해 사용할 수 있습니다. 이미 OpenAI 클라이언트가 있다면 마이그레이션은 기본 URL과 모델 ID를 변경하는 것뿐입니다. 그 외에는 아무것도 필요하지 않습니다. 동일한 키와 엔드포인트로 200개 이상의 모델을 사용할 수 있으므로, Sol은 GPT-5.6 Terra, GPT-5.6 Luna, Claude Opus 5, 그리고 비교 대상이 되는 오픈 가중치 모델들과 함께 제공됩니다. 각 모델을 다시 통합할 필요 없이 난이도에 따라 라우팅할 수 있습니다.
속도 페이지에서 언급할 가치가 있는 OrcaRouter 고유의 두 가지가 있습니다. BYOK: OpenAI 키를 직접 가져와서 사용하면 OpenAI가 자체 요율로 직접 청구하며, OrcaRouter는 토큰당 $0를 추가하고 공급업체의 Rate Limit과 크레딧을 그대로 유지합니다. Guardrails: PII Shield와 콘텐츠 정책은 청구 전에 실행되므로 차단된 요청은 깨끗한 400 응답을 반환하고 절대 청구되지 않으며, Agent Firewall은 도구 및 MCP 호출을 실행 전에 평가합니다. Ultrafast가 라우팅 가능한 가격으로 GA(일반 공급)에 도달하면 — 그리고 도달한다면 — 동일한 엔드포인트에 연결하는 것은 모델 ID 변경일 뿐이며, 오늘 구축하는 설정은 그대로 유지됩니다.
정직한 경계 — Ultrafast가 답이 아닐 때
14× 스토리가 성립하지 않는 네 가지 경우가 있으므로, 확정되지 않은 숫자에 맞춰 아키텍처를 설계하거나 예산을 책정하지 마십시오:
14배는 최대치이지 보장이 아닙니다. 이는 Cerebras 하드웨어의 출력 처리량 상한선입니다. 종단 간 지연 시간에는 여전히 입력 처리, 모델의 추론 시간, 그리고 사용자 자신의 네트워크가 포함됩니다. 추론 비중이 높은 프롬프트는 실제 경과 시간의 대부분을 추론에 사용할 수 있으며, 그 경우 내세워진 속도 향상은 줄어듭니다.
It has no price and no GA date. As of 2026-08-19 you cannot buy Ultrafast — only request preview access, and the CS-4 hardware behind it is in early access rather than generally available. Budget against standard Sol ($5 / $30) or fast mode ($10 / $60), and treat any Ultrafast price you see online as unverified.
The benchmarks are vendor-reported. The 11-hour HLE run and the 5.6× GDP-Val figure are OpenAI/Cerebras numbers from the announcement; independent estimates range from roughly 7× to 11× depending on what's measured. Add the CS-4 speed signal to that list: the ~1,300-token/s figure for GPT-5.6 Sol on CS-4 comes from a single X post and has no independent confirmation. None of these are independently verified yet.
없는 병목은 해결해 주지 못합니다.에이전트가 느린 이유가 자체 오케스트레이션, 도구 지연 시간, 또는 입력이 많은 프롬프트 때문이라면, 더 빠른 출력 경로는 꼬리 부분만 약간 개선할 뿐입니다. 티어 매칭 결정 — GPT-5.6 Sol $5/$30, GPT-5.6 Terra $2/$12, GPT-5.6 Luna $0.20/$1.20 — 은 여전히 어떤 속도 등급보다도 청구 금액에 더 큰 영향을 미칩니다.
Bottom line. GPT-5.6 Sol Ultrafast is the fastest serving tier OpenAI has shipped for its flagship — the same weights on Cerebras hardware, up to 14× throughput and 750 output tokens per second on the announced tier. Cerebras's new CS-4 (unveiled 2026-08-18) reportedly pushes GPT-5.6 Sol toward ~1,300 tokens per second, but that figure is a single unverified X post. As of 2026-08-19, Ultrafast is a waitlist preview with no public price. If you are searching for a number to budget against, the honest answer is that there isn't one yet. Standard GPT-5.6 Sol at $5 / $30 is what you can call today, fast mode at $10 / $60 is the purchasable speed tier, and OrcaRouter passes both through at $0 markup on your own key. Build the routing now — and when Ultrafast gets a price and a date, it is one model-id change away.
Ultrafast의 가격이 책정될 때까지, 전체 GPT-5.6 Sol은 GPT-5.6 Sol on OrcaRouter에서 토큰 100만 개당 $5 / $30, 마크업 없이, 본인의 키로 바로 사용할 수 있습니다.
이 글에서 비교한 모델1
이 글에서 자동 인식 · 벤치마크: Artificial Analysis · 매일 업데이트
