OnpremBench.aiby Understand.tech

Run Llama 3.3 70B Instruct locally

Which workstation runs Llama 3.3 70B Instruct?

A dense 70B model: every parameter is read for every generated token, so memory bandwidth sets the speed. It fits a 128 GB machine at 4-bit or 8-bit; BF16 needs about 145 GB of weights.

Parameters70.6B
Weights at 4-bit42 GB
Published formatBF16 (published checkpoint)
LicenceLlama 3.3 Community License
ReleasedMeta · December 2024

Short answer

24 of 24 catalogued workstations pass the memory screen at 4-bit. The cheapest with a public price is the Framework Desktop · Ryzen AI Max+ (≈ $3.4k). On the NVIDIA DGX Spark, memory bandwidth caps one request at ≤ 6.3 tok/s. Documented recipes exist for NVIDIA GB10, NVIDIA GB300.

Every workstation, screened

Estimate · 8K context · 1 request
MachineMemoryNeeded at 4-bitSpeed ceilingHighest precision that fitsRecipeList price
Framework Desktop · Ryzen AI Max+AMD Ryzen AI Max128 GB54 GB FitsBandwidth not in record8-bitNone yet≈ $3.4k
MINISFORUM MS-S1 MAXAMD Ryzen AI Max128 GB54 GB FitsBandwidth not in record8-bitNone yet≈ $3.8k
NVIDIA DGX SparkNVIDIA GB10128 GB54 GB Fits≤ 6.3 tok/s8-bitVendor-listed validationSGLang · NVFP4≈ $4.7k
ASUS Ascent GX10NVIDIA GB10128 GB54 GB Fits≤ 6.3 tok/s8-bitVendor-listed validationSGLang · NVFP4≈ $5.3k
Acer Veriton GN100NVIDIA GB10128 GB54 GB Fits≤ 6.3 tok/s8-bitVendor-listed validationSGLang · NVFP4≈ $5.5k
MSI EdgeXpertNVIDIA GB10128 GB54 GB Fits≤ 6.3 tok/s8-bitVendor-listed validationSGLang · NVFP4≈ $6k
Dell Pro Max with GB10NVIDIA GB10128 GB54 GB Fits≤ 6.3 tok/s8-bitVendor-listed validationSGLang · NVFP4≈ $8.2k
Exxact Valence DGX StationNVIDIA GB300252 GB50 GB Fits≤ 164 tok/sBF16Vendor starting pointSGLang · Checkpoint default · confirm dtype≈ $95k
MSI XpertStation WS300NVIDIA GB300252 GB50 GB Fits≤ 164 tok/sBF16Vendor starting pointSGLang · Checkpoint default · confirm dtype≈ $109k
Dell Pro Max with GB300NVIDIA GB300252 GB50 GB Fits≤ 164 tok/sBF16Vendor starting pointSGLang · Checkpoint default · confirm dtype≈ $175k
NVIDIA DGX StationNVIDIA GB300252 GB50 GB Fits≤ 164 tok/sBF16Vendor starting pointSGLang · Checkpoint default · confirm dtypePrice from the manufacturer
ASUS ExpertCenter Pro ET900N G3NVIDIA GB300252 GB50 GB Fits≤ 164 tok/sBF16Vendor starting pointSGLang · Checkpoint default · confirm dtypePrice from the manufacturer
HP ZGX FuryNVIDIA GB300252 GB50 GB Fits≤ 164 tok/sBF16Vendor starting pointSGLang · Checkpoint default · confirm dtypePrice from the manufacturer
GIGABYTE W775-V10-L01NVIDIA GB300252 GB50 GB Fits≤ 164 tok/sBF16Vendor starting pointSGLang · Checkpoint default · confirm dtypePrice from the manufacturer
Supermicro Super AI Station · ARS-511GDNVIDIA GB300252 GB50 GB Fits≤ 164 tok/sBF16Vendor starting pointSGLang · Checkpoint default · confirm dtypePrice from the manufacturer
Apple Mac Studio · M3 UltraApple Silicon96 GB54 GB Fits≤ 19 tok/s4-bitNone yetPrice from the manufacturer
HP ZGX Nano G1nNVIDIA GB10128 GB54 GB Fits≤ 6.3 tok/s8-bitVendor-listed validationSGLang · NVFP4Price from the manufacturer
Lenovo ThinkStation PGXNVIDIA GB10128 GB54 GB Fits≤ 6.3 tok/s8-bitVendor-listed validationSGLang · NVFP4Price from the manufacturer
GIGABYTE AI TOP ATOMNVIDIA GB10128 GB54 GB Fits≤ 6.3 tok/s8-bitVendor-listed validationSGLang · NVFP4Price from the manufacturer
AMD Ryzen AI HaloAMD Ryzen AI Max128 GB54 GB FitsBandwidth not in record8-bitNone yetPrice from the manufacturer
HP Z8 Fury G6iNVIDIA RTX PRO96 GB50 GB FitsBandwidth not in record8-bitNone yetPrice from the manufacturer
Lenovo ThinkStation P7NVIDIA RTX PRO96 GB50 GB FitsBandwidth not in record8-bitNone yetPrice from the manufacturer
HP Z2 Mini G1aAMD Ryzen AI Max128 GB54 GB FitsBandwidth not in record8-bitNone yetPrice from the manufacturer
Apple Mac Studio · M5 UltraApple Silicon96 GB54 GB FitsBandwidth not in record4-bitNone yetPreorder · price from the manufacturer

Fit: weights at 4-bit plus KV cache for 8K tokens and a runtime reserve, against 90% of one memory tier. Speed ceiling: memory bandwidth divided by the bytes read per generated token. Neither is a measured result. Method · Change the context, users or precision in the finder

Documented configurations

Sourced
Vendor-listed validation

NVIDIA DGX Spark

SGLang · NVFP4

nvidia/Llama-3.3-70B-Instruct-FP4

--quantization modelopt_fp4

NVIDIA source
Vendor starting point

NVIDIA DGX Station

SGLang · Checkpoint default · confirm dtype

meta-llama/Llama-3.3-70B-Instruct

Follow the checkpoint requirements in the upstream guide.

NVIDIA source

Questions people ask

How much memory does Llama 3.3 70B Instruct need?

About 42 GB for the weights at 4-bit, plus the KV cache and runtime reserve: roughly 54 GB in total at 8K context and one request. Longer context and more simultaneous users need more.

What is the cheapest workstation that runs Llama 3.3 70B Instruct?

In the Lab’s catalogue, the Framework Desktop · Ryzen AI Max+ (≈ $3.4k, Supplier list price · United States · Sept 2026) passes the memory screen at 4-bit.

How fast does Llama 3.3 70B Instruct run locally?

The Lab has no measured run yet. The memory-bandwidth ceiling for one request is ≤ 6.3 tok/s on the NVIDIA DGX Spark; real single-stream speed lands below the ceiling.

Ran Llama 3.3 70B Instruct on one of these machines? Publish your run with the exact model revision, runtime version, context and tokens per second. Measured results replace the estimates on this page.