A hero title card for the article titled 'GLM-5.3-Flash VRAM Requirements' with subtitle 'How much memory you need to run a 320B MoE', pill badges '320B total · 18B active', '1M-token context', 'Community MLX builds', and a summary card reading '2bit-lite: ~102 GB weights / 112 GB RAM — the build that fits a 128 GB Mac', with the OrcaRouter logo in the bottom-right corner.
Guides & Insights

GLM-5.3-Flash VRAM Requirements: How Much Memory You Need to Run a 320B MoE

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

GLM-5.3-Flash does not fit on any consumer GPU. The smallest usable build, 2bit-lite, needs ~102 GB of weights and 112 GB of RAM. A 128 GB Mac is the entry point; a single H200 fits 2bit-lite with ~39 GB spare; everyone else should use the hosted API, z-ai/glm-5.3-flash.

That is the whole answer in one breath, and it is worth saying plainly right here: there is no consumer hardware configuration that runs this model. Z.ai's 320B-parameter vision-language mixture-of-experts model, released August 26, 2026, needs server-class memory in every quantization. The "18B active" figure describes per-token compute, not resident memory — all 320B weights stay loaded. The realistic local targets are large unified-memory Macs, multi-GPU servers, and a single H200 at 141 GB. If you do not own one of those, the numbers below are not a shopping list; they are the reason to call the model over an API instead.

One framing note before the figures: the memory numbers in this piece are community findings, not vendor guidance. Z.ai publishes weights and API specs; it does not publish local-inference memory requirements. The per-quantization table comes from the orcarouter/GLM-5.3-Flash-MLX model card on Hugging Face, an MLX quantization of the open weights that a community group maintains, and the deployment numbers come from practitioners who have actually loaded the model. Where a figure is vendor-reported — pricing, and the architecture's KV-cache efficiency claims — we label it as such.

The short answer: what fits on what you own

Start from what you actually have, because the quantization ladder only makes sense against a target machine:

• Consumer GPU (RTX 4090, RTX 5090, and everything below) — no. Not a single consumer card has enough VRAM: the smallest build is ~102 GB of weights alone, roughly five RTX 5090s' worth. System RAM does not change this, because the weights must be resident on the device.

• 128 GB Mac (M4 or M5 Max) — the 2bit-lite build (~102 GB weights / 112 GB min RAM), and only after raising the wired-memory limit. More on that below. This is the one consumer-class machine the model fits on.

• 192 GB and 256 GB Mac Studio — 3-bit (~184 / 200 GB) and 2-bit (~145 / 160 GB) become reachable; 4-bit needs ~224 GB, which only the 256 GB tier approaches.

• Single H200 (141 GB VRAM) — 2bit-lite fits with roughly 39 GB left for the KV cache, which means short contexts only.

• Multi-GPU server — everything, including 4-bit, 6-bit, and the ~328 GB FP8 reference build.

• Everyone else — the hosted API, z-ai/glm-5.3-flash. That is not a consolation prize; it is where the model is actually fast and cheap to use, and it is covered at the bottom of this page.

Two numbers, not one: weights vs total memory

Every quantization on the model card has two columns — weight size and minimum RAM — and the gap between them is the part most write-ups skip. The weights column is the model files: every parameter, in that precision, resident in memory. The minimum-RAM column is what the machine must hold at runtime: the weights plus the KV cache, the activations, and the buffers that grow with context length.

The 18B-active figure is where people get misled. GLM-5.3-Flash is a 320B-total / 18B-active MoE: per token, only about 18B parameters compute. That is a compute saving, not a memory saving. All 320B weights sit in memory regardless of which experts fire, because the router does not know which experts it needs until it sees the token. MoE buys speed, not footprint — a point the model's own card makes by example, with the 320B total on display alongside the 18B active.

So when a build says "~204 GB weights / 224 GB min RAM," the extra ~20 GB is runtime overhead — KV cache, activations, context buffers. Push context length up and that delta grows. The min-RAM column, not the weights column, is the one to size a machine against.

A screenshot of the official Hugging Face model card for zai-org/GLM-5.3-Flash (captured August 29, 2026), showing the MIT license, the image-text-to-text and glm5_next tags, and the card's introduction describing GLM-5.3-Flash as the first natively multimodal model in the GLM-5 series with 320B total parameters and 18B active, introducing a hybrid sparse-and-linear attention architecture.

The model itself — architecture, license, and Z.ai's own framing — is documented on the vendor's official card, shown above. Local memory is not covered there; that is why this page exists. The numbers below come from the community MLX port of the open weights.

The quantization ladder

The table on the orcarouter/GLM-5.3-Flash-MLX card is the one practitioners are actually loading from today. It lists five quantized builds plus the FP8 reference, each with weight size and minimum RAM:

• FP8 reference — ~328 GB weights. The un-quantized reference point the open weights ship at.

• 6-bit — ~296 GB weights / 320 GB min RAM. Near-lossless; the best-quality build.

• 4-bit — ~204 GB / 224 GB. The recommended everyday default.

• 3-bit — ~184 GB / 200 GB. Aggressive but usable.

• 2-bit — ~145 GB / 160 GB. Best-effort.

• 2bit-lite — ~102 GB / 112 GB. The smallest build; the only one that fits a 128 GB Mac or a single H200.

Quality falls as the bits drop, and the card quantifies it. Relative to the FP8 reference, perplexity degrades by +0.24% at 6-bit, +2.96% at 4-bit, +9.96% at 3-bit, +56.9% at 2-bit, and +141% at 2bit-lite (perplexity figures from the same model card). The card's own field guidance: 6-bit for near-lossless, 4-bit as the everyday default, 3-bit and 2-bit when memory-constrained, 2bit-lite only when nothing else fits — and long code generation is not reliable at 2bit-lite. That last warning matters for a 320B reasoning model: 2bit-lite is what buys you the 128 GB Mac, and the quality tax lands exactly where coding work hurts most.

A single-column scoreboard titled 'GLM-5.3-Flash — the memory ladder' listing six rows: 'FP8 reference: 328 GB weights', '6-bit: 296 GB / 320 GB RAM', '4-bit: 204 GB / 224 GB RAM', '3-bit: 184 GB / 200 GB RAM', '2-bit: 145 GB / 160 GB RAM', '2bit-lite: 102 GB / 112 GB RAM', with a footer reading 'Community MLX builds (orcarouter/GLM-5.3-Flash-MLX) — not vendor guidance.' and the OrcaRouter logo in the bottom-right corner.

Two numbers in that ladder deserve a closer look because they decide the whole hardware question.

Why 2bit-lite exists

2bit-lite is not an extra tier of quality — it is a size tier created for one reason: regular 2-bit does not fit. At ~102 GB of weights it is the only build that slides under the ~112 GB of usable memory on a 128 GB Mac, and the only one that fits a single H200's 141 GB with room left for a KV cache. The model card says as much: regular 2-bit does not fit a 128 GB Mac; 2bit-lite does, with a raised wired-memory limit. On an H200 it fits "with ~39 GB left for the KV cache" (the card's words). That 39 GB is the entire working budget for everything the model does after loading — which brings us to context.

What the KV cache costs at long context

GLM-5.3-Flash has a 1M-token context window, and long context is where local memory plans die. The model's hybrid sparse-plus-linear attention — NoPE-MLA with an index-pool mechanism — is genuinely efficient: Z.ai reports it cuts attention compute 3.01× and KV cache size 4.44× versus its own GLM-5.3 (vendor-reported figures). But "4.44× smaller than GLM-5.3" still leaves a cache measured in tens of GiB when you push toward a full 1M context.

The best public number we have is from a community deployment that ran the model across four DGX Spark nodes: 16 GiB of KV cache per rank — about 64 GiB total across the four — to hold a full 1M-token context, sized to support a few concurrent full-context requests. That is more than the entire memory budget of most single machines, before a single weight. On a single H200, the arithmetic is the point of this page: 2bit-lite leaves ~39 GB for everything after the weights — comfortable for a short chat, gone in minutes as context climbs toward 100K tokens.

The practical rule: the min-RAM column assumes a reasonable context. If your workload runs long documents, agent loops, or repository-scale code, budget KV-cache memory on top — and for anything near 1M tokens, stop doing the arithmetic and use the API. These sizing observations are community findings; no vendor guidance exists for KV budgeting.

The macOS wired-memory limit: why a 128 GB Mac still fails to load

The most common failure reported by practitioners is not "not enough RAM." It is a 128 GB Mac with a ~102 GB model that refuses to load. The cause is macOS's wired-memory limit. On Apple Silicon, the GPU does not get to address all of unified memory: Metal exposes a "recommended max working set" of roughly two-thirds of physical RAM, and allocations above it fail even when the machine has free memory.

So a 128 GB Mac's default Metal budget sits somewhere in the 80–90 GB range — under what the 2bit-lite build needs (~112 GB min RAM with a ~102 GB resident model). The load fails on the Metal budget, not on RAM capacity. The fix in every practitioner report we found is the same: raise the wired limit with sudo sysctl iogpu.wired_limit_mb=…, setting a value above the model's total footprint in MB, and expect it to reset on reboot. Some community guides also set the wired limit from Python via mlx.metal.set_wired_limit() so the model's weights are pinned and macOS stops compressing idle Metal pages.

One more wrinkle: runtimes behave differently. MLX enforces the Metal budget and fails hard when it is exceeded, while llama.cpp-based runtimes (the GGUF path) generally do not enforce it and let macOS swap instead. That is why the same model can refuse to load in one runtime and "load" in another — and why a swapped model can be slow to the point of unusable. These are community findings about macOS behavior, not Apple guidance.

The hosted alternative: z-ai/glm-5.3-flash

For everyone the numbers above rule out — which is most people — Z.ai serves the model itself as z-ai/glm-5.3-flash, and this is where the "Flash" in the name actually shows up. Z.ai's list price is $0.15 per million input tokens, $0.03 per million cached-input tokens, and $0.50 per million output tokens; a launch promotion runs to September 9, 2026, at $0.075 / $0.015 / $0.25 (vendor-reported pricing, current as of writing). Cached input at a fifth of fresh input is the single biggest cost lever: any workload with a reusable prefix — system prompts, tool definitions, long shared documents — should be hitting cache-read pricing.

The memory math belongs in the decision. Serving a 320B model yourself means dedicating 112–320 GB of memory to it whether it is idle or saturated. The hosted endpoint moves all of that off your machines, and at launch-promo prices the model costs less per token than many models a tenth its size — which is the entire point of an 18B-active MoE.

This is also the natural place for the routing point. Through OrcaRouter, z-ai/glm-5.3-flash is one of 200+ models behind a single API, priced at the provider's list price with 0% markup — so the launch promo, and any future price cut, are live the same day they are announced. Automatic failover means a three-day-old model with a shaky serving path is not a bet on your production stack: if the provider errors or saturates, the request rolls over to a healthy provider instead of failing. Trying an unproven model safely is exactly what a router is for.

A screenshot of the OrcaRouter model page for z-ai/glm-5.3-flash (captured August 29, 2026), showing the model description 'native multimodal model... 320B total / 18B active parameters, 1M-token context, text + image + video in, text out', release date 2026-08-26, endpoint /v1/chat/completions, and prices of $0.07 per million input tokens and $0.25 per million output.

FAQ

Can I run GLM-5.3-Flash on an RTX 5090?

No. The smallest build is ~102 GB of weights alone, and an RTX 5090 has 32 GB of VRAM. No consumer GPU comes close; the model needs unified-memory Macs, a single H200, or a multi-GPU server.

Does 18B active mean GLM-5.3-Flash runs on consumer hardware?

No. The 18B active figure is per-token compute. All 320B parameters stay resident in memory, because the router cannot know which experts a token needs until it sees the token. MoE saves compute, not memory.

What is the cheapest way to try GLM-5.3-Flash?

The hosted API, z-ai/glm-5.3-flash. At the launch-promo price it costs $0.075 per million input tokens, and cached-input tokens cost $0.015. Self-hosting the smallest build means dedicating ~112 GB of RAM to it, which only makes sense if you already own the hardware.

The honest summary: GLM-5.3-Flash is a server-class model in every build. The 128 GB Mac gets the one build that fits — 2bit-lite, with a wired-memory-limit raise, short contexts, and a documented quality tax on long code. A single H200 gets the same build with ~39 GB of KV headroom. Server hardware gets the real ladder, from 4-bit as the default up to near-lossless 6-bit. And for everyone else — most readers — the hosted z-ai/glm-5.3-flash endpoint is the right answer, and the numbers above are the reason, not a shopping list. For the step-by-step install on a MacBook Pro, our MacBook install walkthrough covers the whole flow end to end; for how the quantization builds compare on quality and when to pick each, the MLX build guide has the detail; and the release coverage carries the launch context and benchmark claims.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily