
Nemotron Sports Tennis vs InternLumina U2: The Comparison Neither Model Can Lose
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-10$0.15 / $0.60 per 1M tokens
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.24 / $0.73 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0340Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2134Intelligence69Coding
Ask which is better — NVIDIA NemotronLabs AI for Media - Sports Tennis or InternLumina U2 — and the honest answer is that the question does not resolve. The two repositories both appeared on Hugging Face in September 2026, both are mixture-of-experts multimodal models built by well-funded labs, and they have almost nothing else in common. NVIDIA's model is a 31B specialist that reads a single tennis point and answers questions about it. InternLumina U2, from Shanghai AI Laboratory's InternLM team and published as internlm/InternLumina-U2, is a 16B unified model that understands images, video and 3D and also generates and edits images. They share no benchmark, no target user and no deployment profile. Comparing them is less like choosing a phone and more like choosing between a stethoscope and a 3D printer.
What makes the pairing worth writing about anyway is a pattern both models share: neither has been independently evaluated, and both are being described by their own vendor as something other than finished. That is the real comparison — not which model wins, but how much of either one you can actually verify today.
Two different jobs wearing the same word
"Multimodal" covers both of these and tells you nothing. Here is the split:
• What it takes in — Sports Tennis: video (mp4 up to 2 minutes), audio (wav/mp3), images, text. InternLumina U2: text, images, video, and 3D. The tennis model is the only one of the two with a material audio path, wired through a Parakeet speech encoder that reads crowd and impact sound alongside the frames.
• What it gives back — Sports Tennis: text only. InternLumina U2: text and generated or edited images, via an 8-codebook fully-discrete visual representation built on AToken.
• What the user asks — Sports Tennis: "what happened in this point, where was the player standing, did the ball land in." InternLumina U2: "describe this chart, and then draw me a new version of it."
• Architecture — Sports Tennis: Mamba2-Transformer hybrid MoE, 31B total with about 3B active per token, inheriting its backbone and encoders from Nemotron 3 Nano Omni. InternLumina U2: a 16B sparse MoE with 1B active parameters, built as a multi-codebook diffusion language model rather than a conventional autoregressive one.
• Context — Sports Tennis publishes up to 256k tokens. InternLumina U2's card does not advertise a context figure.
• License — Sports Tennis ships under OpenMDW-1.1 with the card stating commercial and non-commercial use. InternLumina U2 is Apache 2.0.

The thing that actually makes them incomparable
There is no shared scoreboard, and it is worth being precise about why rather than leaving it as a vague shrug.
Sports Tennis fine-tuned on a proprietary NVIDIA tennis dataset — 1,312,129 Q&A examples across 43,084 point-level clips from 239 matches, labelled along 38 annotation categories, all collected in 2025. Its test set is a held-out partition of 12 matches, 2,689 clips and 80,872 questions. Every one of those artifacts is NVIDIA's own. No outside group has the data, so no outside group can reproduce a score — and as it happens, NVIDIA published no score to reproduce. The card documents the evaluation protocol in detail and then reports no result from it. Nothing about accuracy, nothing per category, nothing.
InternLumina U2 sits on the opposite side of that same wall. It publishes numbers, but they are explicitly partial. The model card says so in one line: "Preliminary, partial results. Full comparison tables will appear in the upcoming technical report." What exists is a spread of vendor-reported figures across very different capabilities — ChartQA 86.52 and CharXiv-DQ 83.65 on documents, MathVision 33.22 and DynaMath 56.37 on visual math, VideoMME 51.26 and MVBench 59.74 on video, 3D MM-Vet 41.8, DPG-Bench 87.10 and GenEval 0.81 on generation, ImgEdit 3.83 on editing. And the project page's own table shows the model trailing at least one named competitor on at least one headline metric: GenEval 0.81 against LLaDA2.0-Uni's 0.89.
So one model has a documented evaluation and no results, and the other has results and an explicitly incomplete evaluation. Neither gives you a number you can hold next to the other. Any page that ranks these two has invented its ranking.

Where they genuinely overlap, and what it shows
Video understanding is the one territory both claim, and even there the overlap is thinner than it looks. InternLumina U2 reports VideoMME 51.26 and MLVU 57.03 and MVBench 59.74 — general video comprehension scores, vendor-reported. Sports Tennis reports nothing, but its card describes a much narrower and much deeper task: point-level reasoning over a full rally with audio, across 38 fine-grained annotation categories including shot mechanics, court positioning, movement and score state.
The reasonable inference is that InternLumina U2 would be the better generalist and Sports Tennis the better tennis reader, but that is an inference from task shape, not from data. It could easily be wrong in either direction. What is not an inference is that a general video benchmark and a proprietary point-level tennis benchmark cannot be placed on the same axis, and that a 16B unified generator and a 31B frozen-encoder specialist were not built to answer the same question.
Running either one is a different kind of project
Deployment is where the gap stops being philosophical.
Sports Tennis states a hard floor of one GPU with roughly 80 GB of memory at BF16, with weights around 62 GB, tested across A100, H100, H200, B200, GB200 NVL72, RTX PRO 6000 SE, L40S, DGX Spark, Jetson Thor and RTX 5090. There is no quantized checkpoint, so the 80 GB requirement is not negotiable downward today. It is also English-only.
InternLumina U2's more interesting problem is not memory, it is silicon. The repository hosts checkpoints split by training hardware: a complete ascend/ directory with seven safetensors shards trained on Huawei Ascend NPUs, and an nvidia/ directory marked "coming soon." If you are running NVIDIA hardware — which, for a model whose sibling is a tennis fine-tune of Nemotron, is most of the people reading this — the checkpoint you want is the one that has not landed. Inference code and setup instructions live in the InternLM GitHub repository rather than on the Hub, and the technical report is still pending.
That inverts the usual expectation. The proprietary, unbenchmarked, appliance-shaped model is the one you can download and run right now on day one. The Apache-2.0 open research model is the one waiting on a hardware-specific build.

Where a router fits — and where it doesn't
Neither of these models is hosted on OrcaRouter, and there is no honest way to write around that. Sports Tennis has no endpoint anywhere; NVIDIA published weights and nothing else. InternLumina U2's own repository notes the NVIDIA checkpoints are still pending, so there is nothing to serve on our side either. This is a self-hosting decision on both counts, not a routing one.
The routing question shows up one layer up. A team that wants to evaluate a specialist like Sports Tennis against a generalist like InternLumina U2 does not do it in isolation — it does it as one candidate in a pipeline that also needs speech transcription, document handling, summarization and the application logic around them. Those components are hosted, and they are the parts where a single API key covering 200-plus models with automatic failover across providers saves a week of integration work. One key for the models you can call, so the ones you have to run yourself are the only thing you are actually maintaining.
The open question
Put side by side, NVIDIA NemotronLabs AI for Media - Sports Tennis and InternLumina U2 are a decent illustration of two ways to ship a multimodal model badly in public. One documented its evaluation and withheld the results; the other published results and labeled its evaluation incomplete. Both are, as of September 2026, unfalsifiable claims.
The interesting thing to watch is not which one turns out stronger — they will never be tested on the same thing. It is which lab closes its gap first. If NVIDIA posts tennis numbers, it will have solved the harder problem, because a proprietary held-out benchmark of 12 unseen matches is the only honest way to score a domain fine-tune and NVIDIA already built the split. If InternLM ships the NVIDIA checkpoints and the technical report, it will have made its Apache-2.0 model genuinely usable on the hardware its users own. Whichever happens, this comparison will be worth re-reading — because right now, the most defensible thing either model's card supports is a shrug.
