OnpremBench.aiby Understand.tech

Run a model locally

Start from the model you want to run.

Pick a model and see which workstations hold it in memory, the highest precision that fits, the speed ceiling and the price. Estimates are labelled as estimates.

OpenAI

gpt-oss-120b

117B parameters · 5.1B active · 69 GB at 4-bit

24 of 24 workstations fit · from ≈ $3.4k

See the machines
Meta

Llama 3.3 70B Instruct

70.6B parameters · 42 GB at 4-bit

24 of 24 workstations fit · from ≈ $3.4k

See the machines
Alibaba Qwen

Qwen3 32B

32.8B parameters · 19 GB at 4-bit

24 of 24 workstations fit · from ≈ $3.4k

See the machines
Alibaba Qwen

Qwen3 235B-A22B

235B parameters · 22B active · 139 GB at 4-bit

8 of 24 workstations fit · from ≈ $95k

See the machines