Run a model locally
Start from the model you want to run.
Pick a model and see which workstations hold it in memory, the highest precision that fits, the speed ceiling and the price. Estimates are labelled as estimates.
OpenAI
gpt-oss-120b
117B parameters · 5.1B active · 69 GB at 4-bit
See the machines MetaLlama 3.3 70B Instruct
70.6B parameters · 42 GB at 4-bit
See the machines Alibaba QwenQwen3 32B
32.8B parameters · 19 GB at 4-bit
See the machines Alibaba QwenQwen3 235B-A22B
235B parameters · 22B active · 139 GB at 4-bit
See the machines