
GPT-6 vs DeepSeek V4 Pro: uno che puoi noleggiare da chiunque, uno che puoi portarti a casa
- openaiNUOVOOpenAI: GPT-6.1 Sol2026-09-2952Intelligenza
- anthropicNUOVOAnthropic: Claude Sonnet 5.52026-09-2856Intelligenza
- typesafeNUOVOTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M di token · 128 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligenza
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligenza
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligenza
- xAIGrok 4.72026-09-2146Intelligenza
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M di token · 67 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M di token · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligenza
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligenza77Codice
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligenza76Codice
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligenza76Codice
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligenza82Codice
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M di token · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M di token · 360 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligenza72Codice
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M di token · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligenza75Codice
- obsidianQwen3.8 27B2026-08-1534Intelligenza68Codice
There is exactly one thing you can do with DeepSeek V4 Pro that you cannot do with GPT-6 Sol: take it home. DeepSeek V4 Pro is an MIT-licensed mixture-of-experts flagship — 1.6 trillion total parameters, 49 billion active per token — released on 13 August 2026, and the checkpoint is downloadable. GPT-6 Sol is closed weights, shipped by OpenAI on 22 September 2026, and as of 7 October 2026 it is also the model sitting under the paid tiers of ChatGPT's global rollout to more than 1.2 billion weekly users.
Tutto il resto in questo confronto dipende da quale metrica si legge. Sulla rate card i due sono separati da un fattore quattro. Su quanto è effettivamente costato produrre una risposta finita sulla suite indipendente, sono separati da un fattore 1,55×. Sull'Intelligence Index le righe sembrano separate da sedici punti e, secondo la metodologia documentata dal valutatore stesso, non sono affatto sottraibili. Districare quale di questi tre numeri dovrebbe orientare una decisione è tutto il lavoro, e la risposta cambia a seconda che si stia scegliendo un fornitore o un file.
Che cos'è ciascun modello, in un paragrafo per ognuno
GPT-6 Sol is the middle tier of OpenAI's GPT-6 line — below the GPT-6 Astra flagship, above the budget GPT-6 Luna — with a 1,050,000-token context window, a 128,000-token output ceiling, text, image and file input, native tool calling, structured outputs, and reasoning effort selectable across five levels from low to max. It is the tier OpenAI routes ChatGPT Plus, Pro, Business and Enterprise to, and it is priced at $2.00 per million input tokens and $10.00 per million output, with a long-context tier that reprices the whole request at $4.00 and $15.00 once the prompt passes 272,000 input tokens.
DeepSeek V4 Pro is a text-only flagship with a 1,048,576-token window and a 384,000-token maximum output — three times GPT-6 Sol's ceiling — and no vision at any price. DeepSeek's own API documentation lists it at $0.66 per million input tokens and $1.98 per million output during peak hours, halved to $0.33 and $0.99 off-peak, with cache hits at $0.022 off-peak. Peak hours are 01:00–04:00 and 06:00–10:00 UTC on weekdays; everything else, including weekends and Chinese public holidays, is off-peak.
I tre metri
Comincia da quello che viene citato più spesso, perché è quello che più spesso viene citato senza il suo contesto. Artificial Analysis assegna a DeepSeek V4 Pro 36,0 e a GPT-6 Sol 47,6 sull'Intelligence Index v4.3.2. Undici punti e mezzo è il tipo di divario che mette fine a un confronto.
Non dovrebbe, e lo dice il valutatore stesso. Artificial Analysis confronta i modelli open-weights solo con altri modelli open-weights della stessa classe di dimensione, con il confine a 150 miliardi di parametri, e i modelli proprietari all'interno di una fascia di prezzo con un rapporto misto input-output di 3:1. Quelli sono due gruppi di pari diversi che producono due scale diverse. Un 36 e un 47,6 che si trovano in righe adiacenti della stessa tabella non sono un confronto testa a testa, e la mossa onesta è dirlo invece di sottrarre l'uno dall'altro e chiamarlo capacità.

I due strumenti di misura comparabili dicono qualcosa di più utile:
• Prezzo misto a 3:1 — GPT-6 Sol 4,00 $ per 1M vs DeepSeek V4 Pro 0,99 $ alle tariffe di punta, o 0,50 $ fuori punta; uno scarto da 4× a 8× a seconda di quando si esegue
• Costo per attività di indicizzazione completata — GPT-6 Sol 1,04 $ vs DeepSeek V4 Pro 0,67 $; uno scarto di 1,55×
• Token di output generati nell'intera suite — GPT-6 Sol 76,8M vs DeepSeek V4 Pro 162,9M, quindi il modello economico scrive 2,1× tanto per completare lo stesso lavoro
• Output massimo — DeepSeek V4 Pro 384.000 token vs GPT-6 Sol 128.000; un tetto di 3×
• Input multimodale — GPT-6 Sol accetta testo, immagini e file vs DeepSeek V4 Pro solo testo
• Pesi — DeepSeek V4 Pro con licenza MIT e scaricabile vs GPT-6 Sol chiuso

La quarta e la quinta riga spiegano la seconda. Un divario dichiarato di 4× che si riduce a 1,55× sul lavoro reale è ciò che accade quando il modello più economico è anche quello più prolisso: DeepSeek V4 Pro ha generato 2,1× i token di output rispetto a GPT-6 Sol sullo stesso set di valutazione. La verbosità è un moltiplicatore di costo che non compare mai su una pagina dei prezzi, ed è il numero più sottovalutato in assoluto in questo confronto.
Dove il modello aperto vince davvero
Due punti, ed entrambi sono strutturali anziché marginali.
Il limite di output è il primo. 384.000 token contro 128.000 sono una differenza di tre volte, e decide intere categorie di lavoro: generare un grande artefatto strutturato in una sola chiamata, scrivere un lungo documento senza concatenazioni, produrre un grosso refactoring come singola risposta. GPT-6 Sol non può fare queste cose a qualsiasi prezzo, perché il limite non è un livello di fatturazione — è il massimo del modello.
Agentic tool use is the second, on one specific measurement. DeepSeek V4 Pro scores 96.2% on τ²-Bench, the agentic tool-use evaluation, and that is the figure DeepSeek leads its own card with. It is also the number the model's OrcaRouter catalogue entry repeats, which is the version of the claim our own readers would meet. GPT-6 Sol's comparable τ²-Bench figure is not published on the same board, so treat the comparison as directional: on the one benchmark DeepSeek chose to headline, the open model is at the top of the field, and the closed model has not put a number in the same column.
E quello della proprietà, che non è affatto un benchmark. Un checkpoint MIT scaricabile significa che puoi eseguire il modello sul tuo hardware, tenere una richiesta al di fuori della rete di chiunque, metterlo a punto, fissare una versione e sapere che sarà ancora lì tra tre anni. Nessun fornitore può deprecarlo, rivederne il prezzo o rimuoverlo da un listino. Per una categoria di acquirenti — settori regolamentati, chiunque abbia un vincolo di residenza dei dati, chiunque sia stato scottato dal ritiro di un modello — questo vale più di undici punti di indice.
Dove GPT-6 Sol vince, e non è solo l'indice
Terminal-Bench 4.0 è la voce meno lusinghiera sulla scheda di DeepSeek V4 Pro: 14,1% contro il 43,9% di GPT-6 Sol. Non è una differenza da arrotondamento e indica un profilo reale: il modello aperto è forte nell'orchestrazione di strumenti multi-turno e debole nel lavoro da terminale a lungo orizzonte di cui sono fatti i moderni agenti di programmazione. Se il tuo carico di lavoro è un agente che deve tenere insieme una sessione shell per venti minuti, questa singola valutazione è più predittiva della tua esperienza di quanto lo sia il composito dell'indice.
La multimodalità è la seconda. DeepSeek V4 Pro è testo in ingresso, testo in uscita. GPT-6 Sol accetta immagini e file, e non è una casella da spuntare tra le funzionalità — il lavoro professionale ad alto contenuto documentale comporta screenshot, pagine scansionate, diagrammi e PDF, e un modello solo testuale non può affatto entrare in quella categoria.
The third is the thing that happened this week. GPT-6 landed in ChatGPT's Chat tab for Plus, Pro, Business and Enterprise on 7 October, with Free and Go following from 8 October, carrying a capability OpenAI calls Intelligent UI — answers that compose text, graphics, charts, forms and tappable controls per question rather than defaulting to prose. That is a surface, not an API feature, and it does not change a line of your integration code. What it does change is the ambient default: the model your team already has open in a browser tab is now GPT-6, and prototype work drifts toward the model it can reach without a key. Vendor-reported and unreproduced, OpenAI also claims GPT-6 at Extra High begins answering as fast as GPT-5.6 at Medium, and that GPT-6 Instant starts answering 44% sooner on search-backed questions.
Eseguirne uno, o eseguirli entrambi
The reason this particular pairing is easier to hedge than most is that both models sit on the same catalogue. OrcaRouter routes GPT-6 Sol and DeepSeek V4 Pro — along with 200-plus other models — through one OpenAI-compatible endpoint on one key, at 0% markup over the provider's list price. That pass-through is what makes the peak/off-peak structure legible rather than mysterious: DeepSeek's off-peak halving is the vendor's, we do not touch it, and when a vendor reprices, the change is live here the same day.
È anche il punto in cui la decisione sulla finestra di picco diventa una decisione di routing. La tariffa off-peak di DeepSeek V4 Pro è la metà della sua tariffa di picco, e le finestre di picco sono ore UTC fisse. Un processo batch che può essere spostato vale la pena di essere pianificato attorno a esse, e un percorso sensibile alla latenza vale la pena di tenerlo fuori da esse. Con entrambi i modelli dietro un'unica chiave, la suddivisione è una configurazione anziché un secondo contratto con un fornitore: indirizza il lavoro in blocco e a output lungo verso DeepSeek V4 Pro off-peak, indirizza il traffico multimodale e a forte carico di ragionamento verso GPT-6 Sol, e lascia che il failover automatico copra un limite di velocità su uno dei due lati invece di scoprirlo in produzione.

Quello che farei davvero
Se ti servono immagini, file o un punteggio di ragionamento di frontiera e stai acquistando un'API, GPT-6 Sol è la risposta e gli undici punti dell'indice sono per lo più irrilevanti — il gate multimodale risolve la questione prima che i benchmark abbiano voce in capitolo. Se ti servono output da 384.000 token, una licenza che puoi conservare, un deployment on-premise o il costo per attività completata più basso sul tabellone, DeepSeek V4 Pro vince e il divario nell'indice riguarda il gruppo di riferimento di qualcun altro.
Ciò che non farei è interpretare 36 rispetto a 47,6 come un divario di capacità e comprare su questa base. Le metriche comparabili — 0,67 $ contro 1,04 $ per attività, e una differenza di verbosità di 2,1× nascosta dentro uno scarto di 4× del dato principale — raccontano una storia molto più vicina alla realtà di quanto facciano le righe dell'indice, e sono quelle che compaiono in fattura.
The thing to watch next is whether DeepSeek closes the Terminal-Bench gap in a V4 Pro refresh, since that is the one measurement where the open model looks genuinely behind rather than differently scored. GPT-6, for its part, just stopped being a choice and became a default, and defaults are hard to argue with even when the benchmark says otherwise.
Confrontati in questo articolo2
Rilevato da questo articolo · Benchmark: Artificial Analysis · aggiornato ogni giorno
