OnpremBench.aiby Understand.tech

Run Qwen3 235B-A22B locally

Which workstation runs Qwen3 235B-A22B?

A 235B mixture-of-experts model with 22B active per token. At 4-bit the weights alone need about 140 GB: more than any 128 GB desktop machine, so this is the model that separates desktop systems from GB300 workstations.

Parameters235B · 22B active
Weights at 4-bit139 GB
Published formatBF16 (published checkpoint)
LicenceApache 2.0
ReleasedAlibaba Qwen · April 2025

Short answer

8 of 24 catalogued workstations pass the memory screen at 4-bit. The cheapest with a public price is the Exxact Valence DGX Station (≈ $95k). On the Exxact Valence DGX Station, memory bandwidth caps one request at ≤ 499 tok/s. No platform has a documented recipe in the Lab yet.

Every workstation, screened

Estimate · 8K context · 1 request
MachineMemoryNeeded at 4-bitSpeed ceilingHighest precision that fitsRecipeList price
Exxact Valence DGX StationNVIDIA GB300252 GB147 GB Fits≤ 499 tok/s4-bitNone yet≈ $95k
MSI XpertStation WS300NVIDIA GB300252 GB147 GB Fits≤ 499 tok/s4-bitNone yet≈ $109k
Dell Pro Max with GB300NVIDIA GB300252 GB147 GB Fits≤ 499 tok/s4-bitNone yet≈ $175k
NVIDIA DGX StationNVIDIA GB300252 GB147 GB Fits≤ 499 tok/s4-bitNone yetPrice from the manufacturer
ASUS ExpertCenter Pro ET900N G3NVIDIA GB300252 GB147 GB Fits≤ 499 tok/s4-bitNone yetPrice from the manufacturer
HP ZGX FuryNVIDIA GB300252 GB147 GB Fits≤ 499 tok/s4-bitNone yetPrice from the manufacturer
GIGABYTE W775-V10-L01NVIDIA GB300252 GB147 GB Fits≤ 499 tok/s4-bitNone yetPrice from the manufacturer
Supermicro Super AI Station · ARS-511GDNVIDIA GB300252 GB147 GB Fits≤ 499 tok/s4-bitNone yetPrice from the manufacturer
Framework Desktop · Ryzen AI Max+AMD Ryzen AI Max128 GB151 GB Too large–NoneNone yet≈ $3.4k
MINISFORUM MS-S1 MAXAMD Ryzen AI Max128 GB151 GB Too large–NoneNone yet≈ $3.8k
NVIDIA DGX SparkNVIDIA GB10128 GB151 GB Too large–NoneNone yet≈ $4.7k
ASUS Ascent GX10NVIDIA GB10128 GB151 GB Too large–NoneNone yet≈ $5.3k
Acer Veriton GN100NVIDIA GB10128 GB151 GB Too large–NoneNone yet≈ $5.5k
MSI EdgeXpertNVIDIA GB10128 GB151 GB Too large–NoneNone yet≈ $6k
Dell Pro Max with GB10NVIDIA GB10128 GB151 GB Too large–NoneNone yet≈ $8.2k
Apple Mac Studio · M3 UltraApple Silicon96 GB151 GB Too large–NoneNone yetPrice from the manufacturer
HP ZGX Nano G1nNVIDIA GB10128 GB151 GB Too large–NoneNone yetPrice from the manufacturer
Lenovo ThinkStation PGXNVIDIA GB10128 GB151 GB Too large–NoneNone yetPrice from the manufacturer
GIGABYTE AI TOP ATOMNVIDIA GB10128 GB151 GB Too large–NoneNone yetPrice from the manufacturer
AMD Ryzen AI HaloAMD Ryzen AI Max128 GB151 GB Too large–NoneNone yetPrice from the manufacturer
HP Z8 Fury G6iNVIDIA RTX PRO96 GB147 GB Too large–NoneNone yetPrice from the manufacturer
Lenovo ThinkStation P7NVIDIA RTX PRO96 GB147 GB Too large–NoneNone yetPrice from the manufacturer
HP Z2 Mini G1aAMD Ryzen AI Max128 GB151 GB Too large–NoneNone yetPrice from the manufacturer
Apple Mac Studio · M5 UltraApple Silicon96 GB151 GB Too large–NoneNone yetPreorder · price from the manufacturer

Fit: weights at 4-bit plus KV cache for 8K tokens and a runtime reserve, against 90% of one memory tier. Speed ceiling: memory bandwidth divided by the bytes read per generated token. Neither is a measured result. With only 22B active parameters, compute and runtime overhead limit speed well before bandwidth on the fastest machines. Method · Change the context, users or precision in the finder

Questions people ask

How much memory does Qwen3 235B-A22B need?

About 139 GB for the weights at 4-bit, plus the KV cache and runtime reserve: roughly 147 GB in total at 8K context and one request. Longer context and more simultaneous users need more.

What is the cheapest workstation that runs Qwen3 235B-A22B?

In the Lab’s catalogue, the Exxact Valence DGX Station (≈ $95k, Supplier list price · United States · Sept 2026) passes the memory screen at 4-bit.

How fast does Qwen3 235B-A22B run locally?

The Lab has no measured run yet. The memory-bandwidth ceiling for one request is ≤ 499 tok/s on the Exxact Valence DGX Station; real single-stream speed lands below the ceiling.

Ran Qwen3 235B-A22B on one of these machines? Publish your run with the exact model revision, runtime version, context and tokens per second. Measured results replace the estimates on this page.