히어로 타이틀 카드에는 "GPT-6 vs DeepSeek V4 Pro"가 적혀 있고, 칩에는 "모두를 위한 ChatGPT의 GPT-6 2026-10-07"이 표시되며, 두 개의 가격 줄에는 "GPT-6 Sol $2 / $10"과 "DeepSeek V4 Pro $0.66 / $1.98 peak"가 적혀 있고, 칩에는 "DeepSeek: MIT 가중치, 최대 출력 384K"가 표시되며, 푸터에는 "공급업체 보고 항목은 본문에 표기됨; 지수 수치는 Artificial Analysis 기준."이 적혀 있습니다. OrcaRouter 로고는 오른쪽 하단 모서리에 합성되어 있습니다.
Guides & Insights

GPT-6 vs DeepSeek V4 Pro: 하나는 누구에게서나 빌릴 수 있고, 하나는 집으로 가져갈 수 있다

작성자

Magnus Corvin

게시일

최신 모델 · 20모든 모델 보기 →
벤치마크: Artificial Analysis · 매일 업데이트
모든 게시물로 돌아가기

There is exactly one thing you can do with DeepSeek V4 Pro that you cannot do with GPT-6 Sol: take it home. DeepSeek V4 Pro is an MIT-licensed mixture-of-experts flagship — 1.6 trillion total parameters, 49 billion active per token — released on 13 August 2026, and the checkpoint is downloadable. GPT-6 Sol is closed weights, shipped by Ope​nAI on 22 September 2026, and as of 7 October 2026 it is also the model sitting under the paid tiers of ChatGPT's global rollout to more than 1.2 billion weekly users.

이 비교에서 나머지 모든 것은 어떤 척도를 읽느냐의 문제다. 요금표에서는 둘이 네 배 차이다. 독립 평가 스위트에서 완성된 답변을 산출하는 데 실제로 든 비용을 기준으로 하면 둘은 1.55배 차이다. Intelligence Index에서는 행들이 16포인트 차이로 보이지만, 평가자 자신의 기준에 따르면 전혀 빼서 비교할 수 없는 값이다. 그 하나가 효과가 있다는 것은 표면보다 더 큰 의미인 경우가 많다. 이 글의 비교는 조용한 경고다: 모든 게 어떤 숫자를 사용하느냐에 달려 있고, 올바른 답은 공급업체를 고르는지 모델을 고르는지에 따라 달라진다.

각 모델이 무엇인지, 각각 한 단락씩

GPT-6 Sol is the middle tier of Ope​nAI's GPT-6 line — below the GPT-6 Astra flagship, above the budget GPT-6 Luna — with a 1,050,000-token context window, a 128,000-token output ceiling, text, image and file input, native tool calling, structured outputs, and reasoning effort selectable across five levels from low to max. It is the tier Ope​nAI routes ChatGPT Plus, Pro, Business and Enterprise to, and it is priced at $2.00 per million input tokens and $10.00 per million output, with a long-context tier that reprices the whole request at $4.00 and $15.00 once the prompt passes 272,000 input tokens.

DeepSeek V4 Pro is a text-only flagship with a 1,048,576-token window and a 384,000-token maximum output — three times GPT-6 Sol's ceiling — and no vision at any price. Dee​pSeek's own API documentation lists it at $0.66 per million input tokens and $1.98 per million output during peak hours, halved to $0.33 and $0.99 off-peak, with cache hits at $0.022 off-peak. Peak hours are 01:00–04:00 and 06:00–10:00 UTC on weekdays; everything else, including weekends and Chinese public holidays, is off-peak.

세 개의 계량기

가장 많이 인용되는 것부터 시작하겠습니다. 왜냐하면 그것은 맥락 없이 가장 자주 인용되는 항목이기 때문입니다. Artificial Analysis는 Intelligence Index v4.3.2에서 DeepSeek V4 Pro에 36.0점, GPT-6 Sol에 47.6점을 부여합니다. 11.5점은 비교를 끝내는 그런 격차입니다.

그래서는 안 되며, 평가자 스스로도 그렇게 말한다. Artificial Analysis는 오픈 웨이트 모델을 같은 크기 등급의 다른 오픈 웨이트 모델과만 비교하고, 경계는 1500억 파라미터에 두며, 독점 모델은 입력 대 출력 3:1의 혼합 비율을 기준으로 가격대 전반에 걸쳐 비교한다. 이들은 서로 다른 두 피어 그룹이며, 서로 다른 두 척도를 산출한다. 같은 표의 인접한 행에 놓인 36과 47.6은 일대일 비교가 아니며, 정직한 선택은 한쪽에서 다른 쪽을 빼고 그것을 능력이라고 부르는 것이 아니라 그렇게 말하는 것이다.

A screenshot of the Artificial Analysis leaderboard showing the rows around DeepSeek V4 Pro, with MiMo-V2.6-Flash at 38 and $0.06 per index task, Claude Haiku 5.5 (high) at 38 and $0.08, and DeepSeek V4 Pro 0813 (max) at 36 and $0.67, each row carrying its creator, index, cost per task, output speed and latency columns.

비교할 수 있는 두 미터는 더 유용한 것을 말해준다:

• 3:1 혼합 가격 — GPT-6 Sol은 100만 개당 $4.00, DeepSeek V4 Pro는 피크 요금제 $0.99 또는 오프피크 $0.50; 실행 시점에 따라 4배에서 8배까지 차이
• 완료된 인덱스 작업당 비용 — GPT-6 Sol $1.04 대 DeepSeek V4 Pro $0.67; 1.55배 차이
• 제품군 전체에서 생성된 출력 토큰 — GPT-6 Sol 76.8M 대 DeepSeek V4 Pro 162.9M, 즉 저렴한 모델이 같은 작업을 끝내기 위해 2.1배 더 많이 작성
• 최대 출력 — DeepSeek V4 Pro 384,000 토큰 대 GPT-6 Sol 128,000; 3배 상한
• 멀티모달 입력 — GPT-6 Sol은 텍스트, 이미지, 파일을 받는 반면 DeepSeek V4 Pro는 텍스트만 가능
• 가중치 — DeepSeek V4 Pro는 MIT 라이선스로 다운로드 가능 대 GPT-6 Sol은 비공개

A two-column comparison scoreboard for GPT-6 Sol and DeepSeek V4 Pro, showing GPT-6 Sol at an Intelligence Index of 47.6, $1.04 per index task and a 128,000-token output ceiling, against DeepSeek V4 Pro at 36.0, $0.67 and 384,000 tokens, with input and output prices of $2.00/$10.00 against $0.66/$1.98 and an MIT-weights row. A footer reads "Index figures per Artificial Analysis v4.3.2; row scored in different peer groups, not subtractable."

네 번째와 다섯 번째 줄은 두 번째 줄을 설명한다. 4×라는 대표 수치 격차가 실제 작업에서는 1.55×로 줄어드는 것은, 더 저렴한 모델이 동시에 더 장황한 모델일 때 벌어지는 일이다: DeepSeek V4 Pro는 동일한 평가 세트 전반에서 GPT-6 Sol이 생성한 출력 토큰의 2.1배를 생성했다. 장황함은 가격 페이지에는 결코 나타나지 않는 비용 승수이며, 이번 맞대결에서 가장 제대로 주목받지 못한 단일 수치다.

오픈 모델이 진정으로 승리하는 곳

두 곳이며, 둘 다 주변적이라기보다 구조적입니다.

출력 상한이 첫 번째다. 384,000 토큰 대 128,000은 세 배 차이이며, 이는 작업의 범주 전체를 결정한다: 한 번의 호출로 대규모 구조화 산출물을 생성하기, 체이닝 없이 긴 문서를 작성하기, 대규모 리팩터를 단일 응답으로 산출하기. GPT-6 Sol은 어떤 가격에서도 그런 작업을 할 수 없다. 그 한도는 요금제 등급이 아니라 모델의 최대치이기 때문이다.

Agentic tool use is the second, on one specific measurement. DeepSeek V4 Pro scores 96.2% on τ²-Bench, the agentic tool-use evaluation, and that is the figure Dee​pSeek leads its own card with. It is also the number the model's OrcaRouter catalogue entry repeats, which is the version of the claim our own readers would meet. GPT-6 Sol's comparable τ²-Bench figure is not published on the same board, so treat the comparison as directional: on the one benchmark Dee​pSeek chose to headline, the open model is at the top of the field, and the closed model has not put a number in the same column.

그리고 소유권 항목은, 이건 벤치마크가 전혀 아닙니다. 다운로드 가능한 MIT 체크포인트란 모델을 자신의 하드웨어에서 실행하고, 요청을 누구의 네트워크에도 보내지 않으며, 미세 조정하고, 버전을 고정해 3년 뒤에도 여전히 남아 있을 것임을 알 수 있다는 뜻입니다. 어떤 공급업체도 이를 지원 중단하거나, 가격을 다시 매기거나, 요율표에서 빼버릴 수 없습니다. 규제 산업, 데이터 상주 제약이 있는 사람, 모델 지원 종료로 피해를 본 사람 같은 구매자층에게 그것은 지수 11포인트보다 더 가치가 있습니다.

GPT-6 Sol이 앞서는 부분, 그리고 그것은 단지 지수만이 아니다

Terminal-Bench 4.0은 DeepSeek V4 Pro의 모델 카드에서 가장 초라한 지표입니다: GPT-6 Sol의 43.9% 대비 14.1%입니다. 이는 반올림 오차가 아니며 실제 프로필을 가리킵니다 — 오픈 모델은 다중 턴 도구 오케스트레이션에는 강하지만, 현대 코딩 에이전트의 근간을 이루는 장시간에 걸친 터미널 작업에는 약합니다. 워크로드가 셸 세션을 20분 동안 유지해야 하는 에이전트라면, 이 단일 평가가 인덱스 종합 점수보다 당신의 경험을 더 잘 예측합니다.

멀티모달리티가 두 번째다. DeepSeek V4 Pro는 텍스트 입력, 텍스트 출력이다. GPT-6 Sol은 이미지와 파일을 받아들이는데, 그것은 기능 체크박스가 아니다 — 문서가 많은 전문 업무란 스크린샷, 스캔한 페이지, 도표, PDF를 뜻하며, 텍스트 전용 모델은 그 범주에 아예 들어설 수 없다.

The third is the thing that happened this week. GPT-6 landed in ChatGPT's Chat tab for Plus, Pro, Business and Enterprise on 7 October, with Free and Go following from 8 October, carrying a capability Ope​nAI calls Intelligent UI — answers that compose text, graphics, charts, forms and tappable controls per question rather than defaulting to prose. That is a surface, not an API feature, and it does not change a line of your integration code. What it does change is the ambient default: the model your team already has open in a browser tab is now GPT-6, and prototype work drifts toward the model it can reach without a key. Vendor-reported and unreproduced, Ope​nAI also claims GPT-6 at Extra High begins answering as fast as GPT-5.6 at Medium, and that GPT-6 Instant starts answering 44% sooner on search-backed questions.

하나 실행 또는 둘 다 실행

The reason this particular pairing is easier to hedge than most is that both models sit on the same catalogue. OrcaRouter routes GPT-6 Sol and DeepSeek V4 Pro — along with 200-plus other models — through one Ope​nAI-compatible endpoint on one key, at 0% markup over the provider's list price. That pass-through is what makes the peak/off-peak structure legible rather than mysterious: Dee​pSeek's off-peak halving is the vendor's, we do not touch it, and when a vendor reprices, the change is live here the same day.

여기서 피크 윈도 결정이 라우팅 결정이 되기도 합니다. DeepSeek V4 Pro의 오프피크 요금은 피크 요금의 절반이며, 피크 윈도는 고정된 UTC 시간대입니다. 옮길 수 있는 배치 작업이라면 그 시간대에 맞춰 스케줄링할 가치가 있고, 지연에 민감한 경로라면 그 시간대를 피하는 것이 좋습니다. 두 모델이 하나의 키 뒤에 있으면 이 분리는 두 번째 벤더 계약이 아니라 설정의 문제입니다. 대량 작업과 장문 출력 작업은 오프피크에 DeepSeek V4 Pro로 라우팅하고, 멀티모달 및 추론 비중이 큰 트래픽은 GPT-6 Sol로 라우팅하며, 어느 쪽에서든 속도 제한이 걸리면 프로덕션에서 마주치기보다 자동 페일오버가 이를 감당하도록 하세요.

A screenshot of the OrcaRouter model page for openai/gpt-6-sol, showing the OpenAI vendor label, a 2026-09-22 release date, a 1.05M-token context window, a 128K-token maximum output, text, image and file input and tool calling, and pricing of $2.00 per million input tokens and $10.00 per million output tokens.

내가 실제로 할 일

이미지, 파일 또는 프런티어 추론 점수가 필요하고 API를 구매하려는 경우, GPT-6 Sol이 답이며 11개 지수 포인트는 대체로 요점에서 벗어나 있습니다 — 멀티모달 관문이 벤치마크가 한 표를 행사하기도 전에 결판을 냅니다. 384,000토큰 출력, 계속 보유할 수 있는 라이선스, 온프레미스 배포 또는 순위표에서 완료된 작업당 최저 비용이 필요하다면 DeepSeek V4 Pro가 이기며, 지수 격차는 다른 누군가의 비교 집단에나 해당합니다.

내가 하지 않을 일은 36과 47.6을 비교한 것을 능력 격차로 해석하고 그것을 근거로 구매하는 것이다. 비교 가능한 지표들, 즉 작업당 $0.67 대 $1.04, 그리고 4×라는 헤드라인 격차 안에 숨어 있는 2.1×의 장황함 차이는 인덱스 행들이 들려주는 것보다 훨씬 더 실제에 가까운 이야기를 들려주며, 청구서에 나타나는 것도 바로 이들이다.

The thing to watch next is whether Dee​pSeek closes the Terminal-Bench gap in a V4 Pro refresh, since that is the one measurement where the open model looks genuinely behind rather than differently scored. GPT-6, for its part, just stopped being a choice and became a default, and defaults are hard to argue with even when the benchmark says otherwise.

이 글에서 비교한 모델2

이 글에서 자동 인식 · 벤치마크: Artificial Analysis · 매일 업데이트