MEMORY AND SPEED
Model fit finder
You need
- Model (eight presets with published architectures, or your own parameters)
- Weight format and KV cache precision
- Context length and simultaneous requests
- Optional budget in USD, EUR or GBP
You get
- Memory demand split into weights, KV cache and reserve
- Every workstation that fits, with headroom
- Decode speed ceiling per stream and batched, from memory bandwidth
- Exclusions with the reason for each
- Exportable JSON of the screen