A generated hero card titled 'NV-Reason-CT vs Gemini 3.1 Pro' with the subtitle 'The widest input surface, and still not a CT reader'. Two facing cards are separated by a pair of blue arrows: the left card is labelled 'NV-Reason-CT' with a body-scan icon and 'chest and abdomen only - 13,824 visual tokens per volume - OpenMDW-1.1'; the right card is labelled 'Gemini 3.1 Pro' with a chip icon and 'audio, video, image, file, text in - 1M context - preview tier'. The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

NV-Reason-CT vs Gemini 3.1 Pro: The Widest Input Surface, and Still Not a CT Reader

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Ninety-four percent. That is the error rate our own catalogue currently records for one route serving Gemini 3.1 Pro — 93.93%, alongside a median time to first token that reads as a ten-second cap and a throughput figure of 978 tokens per second. Read that as a statement about the model and you will draw exactly the wrong conclusion. Read it as a statement about a preview tier under load and it becomes the most useful thing on the page, because it is what a clinical pipeline actually experiences when it depends on a route that has not been provisioned for production traffic.

The model on the other side of this comparison has no such problem and no such measurement. NV-Reason-CT is NVIDIA's 4.69-billion-parameter native 3D CT reasoner, downloadable under OpenMDW-1.1, with no published throughput figure at all and no hosting anywhere. So one side of this matchup gives you the broadest input surface in the field with an availability profile you have to design around; the other gives you one narrow capability with a latency you have to measure yourself. Neither has a monitoring dashboard you can trust on day one, and for opposite reasons.

Two models, two definitions of "multimodal"

The word is doing a lot of work in both product descriptions and almost none of it is shared.

• Inputs accepted — Gemini 3.1 Pro takes text, images, audio, video and files; NV-Reason-CT takes a single-channel NIfTI volume in Hounsfield units, and nothing else

• What "3D" means — NV-Reason-CT resamples to 2 mm isotropic over a 384-millimetre cube and cuts it into 8 x 8 x 8 patches for a 24 x 24 x 24 grid; Gemini 3.1 Pro's video input is a sequence of two-dimensional frames with a timeline

• Context — 1,048,576 tokens in and 65,536 out for Gemini 3.1 Pro; one volume is 13,824 visual tokens for NV-Reason-CT, with no downsampling and no merging layer, and there is no chat window around it

• Release — 19 February 2026 for Gemini 3.1 Pro, on a preview slug; an unannounced research drop for NV-Reason-CT, with repository activity dated 2 to 25 September 2026

• Price — $2 and $12 per million tokens up to 200,000 tokens then $4 and $18, with cache reads at $0.20 and cache writes at $0.375; against self-hosted weights with no rate card

• Capabilities — vision, audio, tool use, JSON output and reasoning for Gemini 3.1 Pro; structured findings by anatomical region for NV-Reason-CT, with no tools

• Licence and status — a hosted preview API; against OpenMDW-1.1 weights whose model card says research and education only and states plainly that it is not a medical device

A video input and a volumetric CT input look adjacent from a distance and could not be further apart up close. Video is frames over time. A CT study is a single static volume with three spatial axes, a physical scale in millimetres and a calibrated density unit, where the thing being measured is the density of a structure whose size matters. Gemini 3.1 Pro will watch a surgical recording and describe what happened. It will not measure a nodule. That is not a limitation Google failed to address; it is a different problem that nobody has solved with a general model.

What the two benchmark records cover, and what they leave out

Gemini 3.1 Pro has an independent record and it is a good one on the axes it measures.

• GPQA Diamond — 94.1, the highest figure in this entire set of articles

• Humanity's Last Exam — 47, with a long-context recall score of 82, again the highest on this comparison

• IFBench and SciCode — 77.14 and 58.7

• τ²-Bench — 95.61, with tau_banking at 21.44 and Terminal-Bench Hard at 53.79

• Artificial Analysis Intelligence and Coding — 29.7 and 68.8

• Measured latency — a ten-second median to first token, which appears to be a cap rather than a true median, at 978 tokens per second and a 93.93% error rate

Notice the shape of that: a knowledge and long-context profile at the top of the comparison, an intelligence index in the lower-middle, and an availability measurement that is unusable. All three describe the same model on the same route. A high GPQA Diamond means it knows a great deal. A 29.7 composite means it does not lead on everything. A 93.93% error rate means that, on this particular path, most of your requests will not complete — and that is almost certainly a capacity and provisioning fact about a preview endpoint rather than a property of the weights. We cannot prove which from the outside. What we can say is that none of those three numbers is a typo and any architecture that reads only the first one is going to have a bad week.

NV-Reason-CT's record is narrower and comes with a different caveat. Its 0.614 F1 and 0.871 AUROC on CT-RATE, across eighteen labels with a fixed uniform threshold and a direct yes/no prompt and no classification head or task-specific adaptation, are NVIDIA's own numbers from NVIDIA's own paper. The comparison set behind them — VoxelFM at 0.581, Pillar-0 at 0.544, ClinFusion-8B at 0.442, CT-CLIP at 0.398, Merlin at 0.358, MedGemma 1.5 at 0.303 — contains only imaging models. No general multimodal model appears, which means the question "how would Gemini 3.1 Pro do on CT-RATE" has not been asked, let alone answered.

A generated scoreboard titled 'Breadth on one side, depth on the other' with two columns. The left column, headed 'Gemini 3.1 Pro', lists six rows: inputs, text, image, audio, video, file; context, 1,048,576 in and 65,536 out; GPQA Diamond 94.1; AA Intelligence 29.7; route reliability, 93.93% measured error rate; price, $2 and $12 per million tokens, doubling past 200k. The right column, headed 'NV-Reason-CT', lists six rows: input, one-channel NIfTI in Hounsfield units; context, 13,824 visual tokens per volume, no chat window; CT-RATE F1 0.614 and AUROC 0.871; provenance, the authors paper only; route reliability, not applicable, self-hosted; price, no rate card, no published throughput. A footer reads 'The 93.93% error rate describes a preview route under load, not the quality of the weights.'

The preview-tier problem, stated properly

This is the part of the comparison that has nothing to do with CT and everything to do with running anything clinical.

A 93.93% error rate is not a subtle degradation. It is a route where the overwhelming majority of calls fail. The natural reaction is to declare the model unusable, and that reaction is wrong for two reasons. First, an error rate measured on a preview endpoint reflects capacity, quota and regional routing, not the model's ability. Google's preview tiers are famously provisioned for evaluation rather than production, and a slug named for a preview is a slug that can be throttled without notice. Second, the measurement is a snapshot of one path at one moment. It will change. What will not change is the lesson: a dependency on a preview route is a dependency on a capacity decision someone else makes.

The mitigations are unglamorous and they are the whole reason routing exists as a discipline rather than a convenience.

• Never hard-code a preview slug into a production path — put the capability behind an interface so the model behind it can be swapped without a deploy

• Retry and fall back on a different provider or a different model — for a batch extraction job, a 94% failure rate is survivable if the failures route elsewhere and fatal if they do not

• Measure your own error rate rather than inheriting one — the 93.93% figure is our catalogue's number for one route, and yours will depend on your region, your volume and your workload shape

• Keep a second model warm for anything on a clinical timeline — the failure mode you are protecting against is not a bad answer, it is no answer

The same discipline applies in the opposite direction on the CT side, where there is no route at all. A self-hosted model does not have a provider error rate; it has a hardware error rate, a maintenance burden, and a capacity ceiling set by how many GPUs you bought. The failure mode is a queue, and the mitigation is scheduling rather than failover. Different problem, same rule: know which number you are actually depending on.

Where a long context genuinely helps, and where it does not

Gemini 3.1 Pro's 1,048,576-token window and 82-point long-context recall score are real advantages and worth using deliberately rather than accidentally.

The place they pay off in a CT programme is not the scan. It is everything around it. A single extraction pass over a long operative note and its associated reports, holding the whole document in context so the fields are consistent with each other. A cohort summary that has to hold hundreds of patient narratives at once to find the pattern across them. A conversion job where the schema, a hundred example records and the error log all need to be visible in the same call. Those are jobs where a long window changes the answer, not just the convenience.

The place it does not help is the scan itself, and the reason is worth stating precisely because it is the mistake this matchup invites. A chest CT represented the way NV-Reason-CT represents it costs 13,824 tokens. Gemini 3.1 Pro could hold seventy such studies inside its window with room to spare. The constraint is not room. It is that Gemini 3.1 Pro's vision path treats an image as a picture with a pixel scale, so a study presented that way has lost the inter-slice spacing and the density calibration before the first token is emitted. You can spend a million tokens on a volume and still not have a volume. The paper's own comparison makes the cost of that concrete: MedGemma 1.5 at 0.303 F1 when fed up to 85 axial slices, against 0.614 for the native volumetric model, which is a gap of roughly double.

Where we fit, and the one thing we will not say

Gemini 3.1 Pro is served through OrcaRouter as google/gemini-3.1-pro-preview at the provider's list rate with zero markup, inside a single API covering more than 200 models. For a model currently sitting on a preview tier, that combination is more useful than it sounds. Zero markup means a vendor rate change lands on our side the same day rather than being absorbed by a middle layer holding an old number. Automatic failover means a route that starts failing does not take a batch down with it — which, given a measured error rate in the nineties on one path, is not a hypothetical. And a routing DSL for sending individual requests to the model that should handle them lets you put the long-context work on one model and the cheap high-volume work on another without maintaining two integrations.

NV-Reason-CT is not on our platform. No NVIDIA model is. It is weights you download and run yourself, and the honest framing is that the two halves of this architecture live in entirely different places: one behind an API with a rate card and a preview risk, the other on your own hardware with an OpenMDW-1.1 licence and a throughput figure you have to produce yourself.

A generated scoreboard titled 'What each side gives you, and what it costs' listing five paired rows. Row one: widest input surface, audio, video, image and file against text only, one modality at a time. Row two: longest context, 1,048,576 tokens in against 13,824 tokens per volume. Row three: highest benchmark, GPQA Diamond 94.1 against CT-RATE F1 0.614. Row four: weakest measured figure, a 93.93% route error rate against no published throughput at all. Row five: dependency, a hosted preview tier against your own GPUs. A footer reads 'Both sides carry a number nobody should build on without measuring it themselves.'A screenshot of the OrcaRouter model page for Gemini 3.1 Pro, showing the identifier google/gemini-3.1-pro-preview with tags for Vision, Audio, Video, Tools, JSON and Reasoning, the description of it as a multimodal model accepting text, image, audio, video and file input, and a side panel listing a 1,048,576-token context window with up to 65,536 output tokens. A stat strip reads $2.00 per million input tokens, $12.00 per million output tokens, a median time to first token, and an error rate of 93.93 percent.

The decision that survives contact with production

If your problem is text, audio, video or document understanding at long context, Gemini 3.1 Pro is independently measured at the top of the field on knowledge and recall, and the only thing standing between it and a production pipeline is route reliability — which is a solvable engineering problem, not a model problem. Put it behind failover, do not hard-code the preview slug, and measure your own error rate rather than trusting anyone's, including ours.

If your problem is a chest or abdominal CT volume, there is exactly one open model in this comparison that can address it, it is not the one with the million-token window, and the number that should govern your planning is not its F1 score. It is the per-study latency nobody has published. Measure it before you promise anyone a throughput figure.

The comparison is worth making because it separates two things that get conflated constantly: breadth of input and depth of representation. Gemini 3.1 Pro has the widest input surface anyone has shipped and the best long-context recall in this set, and it still cannot read a CT study as a CT study. NV-Reason-CT can read one and nothing else. A pipeline needs both, needs them on different infrastructure, and needs to know which of the two numbers on each side is the one that will page you at three in the morning.