A generated title card in the OrcaRouter house style reading “PixelUMM vs UI-Mate-27B”, with four statistic tiles: “77.0” (UI-Mate-27B's reported OSWorld-Verified average), “66.2” (its WindowsAgentArena average), “none” (PixelUMM's action space) and “Apache-2.0” (UI-Mate-27B's licence on code and weights, against PixelUMM's noncommercial checkpoint). A footer line reads “Tencent-reported agent scores; PixelUMM figures from its preprint. No independent rerun.”. The OrcaRouter logo is composited into the padded strip at the bottom right.
Guides & Insights

PixelUMM vs UI-Mate-27B: One Model Draws What It Sees, the Other Clicks It

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Tencen​t UI-Mate-27B and NVIDIA PixelUMM look comparable on a spec sheet — 27 billion parameters against roughly 15.2 billion, both open downloads, both built on Qwe​n backbones — and they are not comparable at all. UI-Mate-27B is a foundation GUI agent: it observes live screenshots, plans from what is currently on screen, and emits structured mouse and keyboard actions for a real desktop. PixelUMM is an encoder-free unified multimodal model that reads images and video in raw pixel space and generates images and video from text — it perceives and it creates, but it has no action space, no cursor and no way to touch a screen. One acts on the world. The other describes and depicts it. Everything below follows from that split, and the licensing gap between them is the second thing you need to know.

Both labs report their own numbers and neither has an independent evaluation. Both released quietly — style of release is where the resemblance actually ends.

Actor and perceiver

UI-Mate-27B executes a task from a natural-language instruction and live screenshots alone. It reasons over the visible state, re-plans as the interface changes, and outputs reasoning, a concise action description and structured computer-use tool calls across mouse, keyboard, scrolling, waiting, user interaction and task completion. Its distinguishing feature is demonstration-guided execution: show it a workflow once and it adapts the recorded demonstration to the task at hand, treating the current screenshot as authoritative when the interface has moved on. It is fine-tuned from Qwen3.6-27B with supervised fine-tuning followed by online reinforcement learning in executable GUI environments, and it is served through vLLM with an OpenAI-compatible interface.

• Reported benchmarks — OSWorld-Verified average 77.0, WindowsAgentArena average 66.2, OSWorkerBench strict success 41.00 and progress 76.86. All Tencent-reported and unreproduced.

• Action space — mouse, keyboard, scrolling, waiting, user interaction, task completion, in a normalised coordinate space with pyautogui-compatible structured calls.

• Hardware reality — 27B dense, served with tensor parallelism across two GPUs in Tencent's own quick-start, under Apache-2.0.

• An important caveat on the card itself — UI-Mate-27B is an agent checkpoint rather than a standalone visual-chat model. Tencent recommends the official prompt, response parser and interaction harness from its repository. Point it at "describe this image" and you are using the wrong tool.

PixelUMM occupies the opposite pole. Its entire architecture is an argument about representation: no VAE, no vision encoder, images as 16×16 pixel patches and videos as 4-frame spatiotemporal tubelets, straight into a decoder-only Transformer through single-layer linear projections, with a Mixture-of-Transformers design pairing shared attention with task-specific parameters. Understanding is autoregressive text; generation is pixel-space flow matching. Four checkpoints ship, S8-F22-R05 as the default covering text-to-image, text-to-video at 96 frames and 24 fps, and both understanding directions.

A headless-browser capture of the Hugging Face model card for the Tencent UI-Mate-27B repository, showing the model title, the Apache-2.0 licence tag, and the repository's download and like counters.

The two release patterns are a study in contrast

UI-Mate-27B appeared on Hugging Face on 14 August 2026 with a model card, an arXiv preprint numbered 2608.15930, a project page, a GitHub repository and a documented serving recipe. Between then and early October it accumulated roughly 780 downloads and 93 likes — modest, but a normal research-agent release with the paperwork in order.

PixelUMM's trail is different in kind. The GitHub repository was created on 4 September 2026; the arXiv preprint is dated 29 September; the Hugging Face repository was created at 21:40 UTC on 1 October 2026, minutes after the repository's final commit. No NVIDIA blog post, press release or launch thread surfaced for it. The project page's own URL path still reads pixelumm-project-page-preview. Its download count was zero when this was written.

Reading intent into that is a mistake — a lab can publish a research artifact and schedule a launch later. What is knowable is narrower and more useful: one model has an official serving path and ten weeks of downloads; the other has a paper, a repository, and no vendor statement at all.

A generated scoreboard card titled “PixelUMM vs UI-Mate-27B — the scoreboard”, with the two model names as column headings and six dimension rows spanning both columns: role, backbone, headline benchmarks, action space, deployment and licence. UI-Mate-27B's column shows a foundation GUI agent, a Qwen3.6-27B fine-tune, OSWorld-Verified 77.0 with WindowsAgentArena 66.2, mouse, keyboard, scrolling, waiting and task-completion calls, 27B dense across roughly two GPUs under vLLM, and Apache-2.0; PixelUMM's column shows an encoder-free perceiver and generator, about 15.2B total on a Qwen3-8B backbone, MMMU 41.67 with GenEval 0.83 and VBench Part 1 quality 84.10, no action space, about 30 GB of distributed-checkpoint shards behind a hidden index, and a noncommercial checkpoint licence. A footer notes the agent scores are Tencent-reported and the generation scores are from the PixelUMM preprint, with no independent rerun of either.

Scoreboards that do not overlap

There is no benchmark both models report, which is itself the point: they are not solving the same problem.

• Agentic computer use — UI-Mate-27B: OSWorld-Verified 77.0, WindowsAgentArena 66.2. PixelUMM: no action space, so no score is possible.

• Image understanding — UI-Mate-27B: consumes screenshots as agent input, not evaluated as a general VLM. PixelUMM: MMMU 41.67, AI2D 80.12, DocVQA 90.42, ChartQA 82.96, OCRBench 78.00.

• Video — UI-Mate-27B: none. PixelUMM: MVBench 70.53, Video-MME 57.33, LongVideoBench 59.61, LVBench 40.41.

• Generation — UI-Mate-27B: none. PixelUMM: GenEval 0.83 with a prompt rewriter, DPG-Bench 85.74, VBench Part 1 quality 84.10.

• Deployment cost — UI-Mate-27B: 27B dense across roughly two GPUs via vLLM, plus the official harness. PixelUMM: about 30 GB of distributed-checkpoint shards, CUDA 13 with FlashAttention built from source, a gated guardrail model for the video path and a second Python environment for it.

• Licence — UI-Mate-27B: Apache-2.0, code and weights. PixelUMM: Apache-2.0 code, NVIDIA One-Way Noncommercial License on the checkpoint.

• Scale — 27B dense versus about 15.2B total, rendered as "8B MoT" in PixelUMM's own paper tables.

The only genuinely comparable statement is about provenance, and it applies to both: every number above is a vendor's own, none has been rerun outside the lab that produced it, and one of the two models explicitly labels its own results as preliminary.

Building with them, and why you might use both

These two compose better than they compete. A desktop automation stack wants UI-Mate-27B for the acting layer and something that can reason over what it saw; a content or inspection pipeline wants PixelUMM's reading and generation and has no use for a cursor. The overlap is at the perception boundary, where UI-Mate-27B's screenshot grounding is task-specific and PixelUMM's is general — the same pixels, two entirely different jobs.

When you do run several models against each other — the agent against a generalist VLM, or a new research checkpoint against the one already in production — the cost of the comparison is normally the integration, not the inference: two API shapes, two auth schemes, two contracts. That is the problem a routing layer exists to remove, and it is where OrcaRouter's design pays off: one key across 200+ models, provider list prices passed through at 0% markup, automatic failover across providers, and a routing DSL plus model fusion for composing several models into a single call. Neither PixelUMM nor UI-Mate-27B is on any hosted endpoint today, and OrcaRouter routes neither — so this is about the workflow when they are, not a claim about availability now.

Which one for which job

If the task ends with a click, a keystroke or a filled form, there is only one candidate here and it is UI-Mate-27B. It is Apache-2.0, it has an arXiv paper and an official harness, and its reported OSWorld and WindowsAgentArena scores are the kind of claim that can be checked by anyone with two GPUs and a Windows or Ubuntu test image — which is exactly what makes it worth checking rather than trusting.

If the task ends with an answer or an image, PixelUMM is the one that can do it, and it is a research artifact first. Its paper is unusually honest about its own limits, conceding that differing training data mean its benchmark comparison "cannot establish which architecture is superior". Its checkpoint is noncommercial. Its strengths are architectural novelty and breadth, not state-of-the-art numbers, and its immediate value is to anyone studying whether encoder-free unified models work at video scale. Download it to read the code and run the demos; do not download it to ship a feature.

The two-line summary is the same one the evidence gives: one model can act and cannot create, the other can create and cannot act, and the licensing and maturity gap between them matters more than any parameter count.