OnpremBench.aiby Understand.tech

Run gpt-oss-120b locally

Which workstation runs gpt-oss-120b?

OpenAI’s open-weight mixture-of-experts model: 117B parameters in total, 5.1B active per token. It ships in MXFP4, so the weights need about 70 GB and a 128 GB machine can hold it.

Parameters117B · 5.1B active
Weights at 4-bit69 GB
Published formatMXFP4 (published checkpoint)
LicenceApache 2.0
ReleasedOpenAI · August 2025

Short answer

24 of 24 catalogued workstations pass the memory screen at 4-bit. The cheapest with a public price is the Framework Desktop · Ryzen AI Max+ (≈ $3.4k). On the NVIDIA DGX Spark, memory bandwidth caps one request at ≤ 77 tok/s. Documented recipes exist for NVIDIA GB10, NVIDIA RTX PRO.

Every workstation, screened

Estimate · 8K context · 1 request
MachineMemoryNeeded at 4-bitSpeed ceilingHighest precision that fitsRecipeList price
Framework Desktop · Ryzen AI Max+AMD Ryzen AI Max128 GB80 GB FitsBandwidth not in record4-bitNone yet≈ $3.4k
MINISFORUM MS-S1 MAXAMD Ryzen AI Max128 GB80 GB FitsBandwidth not in record4-bitNone yet≈ $3.8k
NVIDIA DGX SparkNVIDIA GB10128 GB80 GB Fits≤ 77 tok/s4-bitVendor-listed validationSGLang · MXFP4≈ $4.7k
ASUS Ascent GX10NVIDIA GB10128 GB80 GB Fits≤ 77 tok/s4-bitVendor-listed validationSGLang · MXFP4≈ $5.3k
Acer Veriton GN100NVIDIA GB10128 GB80 GB Fits≤ 77 tok/s4-bitVendor-listed validationSGLang · MXFP4≈ $5.5k
MSI EdgeXpertNVIDIA GB10128 GB80 GB Fits≤ 77 tok/s4-bitVendor-listed validationSGLang · MXFP4≈ $6k
Dell Pro Max with GB10NVIDIA GB10128 GB80 GB Fits≤ 77 tok/s4-bitVendor-listed validationSGLang · MXFP4≈ $8.2k
Exxact Valence DGX StationNVIDIA GB300252 GB76 GB Fits≤ 2008 tok/s4-bitNone yet≈ $95k
MSI XpertStation WS300NVIDIA GB300252 GB76 GB Fits≤ 2008 tok/s4-bitNone yet≈ $109k
Dell Pro Max with GB300NVIDIA GB300252 GB76 GB Fits≤ 2008 tok/s4-bitNone yet≈ $175k
NVIDIA DGX StationNVIDIA GB300252 GB76 GB Fits≤ 2008 tok/s4-bitNone yetPrice from the manufacturer
ASUS ExpertCenter Pro ET900N G3NVIDIA GB300252 GB76 GB Fits≤ 2008 tok/s4-bitNone yetPrice from the manufacturer
HP ZGX FuryNVIDIA GB300252 GB76 GB Fits≤ 2008 tok/s4-bitNone yetPrice from the manufacturer
GIGABYTE W775-V10-L01NVIDIA GB300252 GB76 GB Fits≤ 2008 tok/s4-bitNone yetPrice from the manufacturer
Supermicro Super AI Station · ARS-511GDNVIDIA GB300252 GB76 GB Fits≤ 2008 tok/s4-bitNone yetPrice from the manufacturer
Apple Mac Studio · M3 UltraApple Silicon96 GB80 GB Fits≤ 232 tok/s4-bitNone yetPrice from the manufacturer
HP ZGX Nano G1nNVIDIA GB10128 GB80 GB Fits≤ 77 tok/s4-bitVendor-listed validationSGLang · MXFP4Price from the manufacturer
Lenovo ThinkStation PGXNVIDIA GB10128 GB80 GB Fits≤ 77 tok/s4-bitVendor-listed validationSGLang · MXFP4Price from the manufacturer
GIGABYTE AI TOP ATOMNVIDIA GB10128 GB80 GB Fits≤ 77 tok/s4-bitVendor-listed validationSGLang · MXFP4Price from the manufacturer
AMD Ryzen AI HaloAMD Ryzen AI Max128 GB80 GB FitsBandwidth not in record4-bitNone yetPrice from the manufacturer
HP Z8 Fury G6iNVIDIA RTX PRO96 GB76 GB FitsBandwidth not in record4-bitVendor support statementOllama / llama.cpp · MXFP4Price from the manufacturer
Lenovo ThinkStation P7NVIDIA RTX PRO96 GB76 GB FitsBandwidth not in record4-bitVendor support statementOllama / llama.cpp · MXFP4Price from the manufacturer
HP Z2 Mini G1aAMD Ryzen AI Max128 GB80 GB FitsBandwidth not in record4-bitNone yetPrice from the manufacturer
Apple Mac Studio · M5 UltraApple Silicon96 GB80 GB FitsBandwidth not in record4-bitNone yetPreorder · price from the manufacturer

Fit: weights at 4-bit plus KV cache for 8K tokens and a runtime reserve, against 90% of one memory tier. Speed ceiling: memory bandwidth divided by the bytes read per generated token. Neither is a measured result. With only 5.1B active parameters, compute and runtime overhead limit speed well before bandwidth on the fastest machines. Method · Change the context, users or precision in the finder

Documented configurations

Sourced
Vendor-listed validation

NVIDIA DGX Spark

SGLang · MXFP4

openai/gpt-oss-120b

Follow the checkpoint requirements in the upstream guide.

NVIDIA source
Vendor support statement

RTX PRO workstation / selected 96 GB GPU

Ollama / llama.cpp · MXFP4

openai/gpt-oss-120b

Confirm the selected runtime’s gpt-oss support and GPU offload. Workstation Edition and Max-Q need separate measurements.

NVIDIA source

Questions people ask

How much memory does gpt-oss-120b need?

About 69 GB for the weights at 4-bit, plus the KV cache and runtime reserve: roughly 80 GB in total at 8K context and one request. Longer context and more simultaneous users need more.

What is the cheapest workstation that runs gpt-oss-120b?

In the Lab’s catalogue, the Framework Desktop · Ryzen AI Max+ (≈ $3.4k, Supplier list price · United States · Sept 2026) passes the memory screen at 4-bit.

How fast does gpt-oss-120b run locally?

The Lab has no measured run yet. The memory-bandwidth ceiling for one request is ≤ 77 tok/s on the NVIDIA DGX Spark; real single-stream speed lands below the ceiling.

Ran gpt-oss-120b on one of these machines? Publish your run with the exact model revision, runtime version, context and tokens per second. Measured results replace the estimates on this page.