OnpremBench.aiby Understand.tech

ONPREMBENCH / PLANNING TOOLS

Plan before you invest.

5 tools that use the Lab’s sourced machine data and published model architectures to answer the questions you have before investing. Every result shows its assumptions; measured numbers only come from a run on an identified machine.

THE PLANNING PATH

  1. 1

    Will it fit, and how fast?

    Weights, KV cache and reserve against each machine’s memory tier, plus a bandwidth speed ceiling.

    Model fit finder
  2. 2

    How many people can it serve?

    Turn provisioned users into simultaneous requests and check the demand against the machine’s ceiling.

    Users & workload planner
  3. 3

    What will a token cost me?

    Hardware, power, software and upkeep over your horizon, beside your API bill, per accepted task and per million tokens.

    Cost of local inference

MEMORY AND SPEED

Model fit finder

You need

  • Model (eight presets with published architectures, or your own parameters)
  • Weight format and KV cache precision
  • Context length and simultaneous requests
  • Optional budget in USD, EUR or GBP

You get

  • Memory demand split into weights, KV cache and reserve
  • Every workstation that fits, with headroom
  • Decode speed ceiling per stream and batched, from memory bandwidth
  • Exclusions with the reason for each
  • Exportable JSON of the screen

Limits. Estimates from model-card architecture figures. No CPU offload. Real speed is lower than the ceiling.

Open model fit finder

CAPACITY

Users & workload planner

You need

  • A machine page and one of its model configurations
  • Provisioned users, active share, requests per active user, mean request duration
  • Output tokens per response

You get

  • Active users, requests per minute and simultaneous requests in the busy period
  • Required decode rate versus the machine’s bandwidth ceiling for that model
  • A verdict: comfortable, tight, or measure before committing
  • Saved to your Lab profile when signed in

Limits. Demand is your assumption; supply is a bandwidth ceiling, not a benchmark of the exact machine.

Open users & workload planner

ECONOMICS

Cost of local inference

You need

  • A machine (its published price and power rating prefill the form) or your own quote
  • Software, support and administration per month; electricity tariff and powered hours
  • Monthly token volumes and a hosted price (dated presets for the same open models, or your own), or a monthly cloud bill
  • Horizon and resale value

You get

  • The monthly token volume above which local costs less
  • Total local and cloud cost over the horizon with a cumulative chart
  • Monthly operating cost including electricity
  • Simple payback in months
  • Cost per accepted task and per million output tokens, local and cloud

Limits. No financing, tax or currency conversion. Comparable quality and availability are your responsibility.

Open cost of local inference

SPECIFICATIONS

Compare machines

You need

  • Up to four machines from any card’s Compare box, or a preset

You get

  • Chip, memory, bandwidth, storage, form factor, connectivity, power and price side by side
  • How much evidence exists on each: model configurations, setups, use cases
  • CSV export

Limits. Sourced, dated specifications. A bigger memory number is not a speed ranking.

Open compare machines

DATA

Machine data API

You need

  • Nothing; public and read-only

You get

  • All 27 machine records with specifications, dated offers and evidence notes as JSON
  • The 19 model configurations attached to each platform
  • One machine with ?machine=<id>

Limits. Sourced catalogue data, not a live price feed. Third-party images and linked sources keep their own rights. Cached five minutes.

Open machine data api

HOW THE NUMBERS ARE MADE

Explicit formulas, published inputs.

Memory

Weights = parameters × effective bits ÷ 8 (+3% for embeddings and buffers). KV cache = bytes per token from the model card (2 × layers × KV heads × head size × 2 bytes) × context × simultaneous requests. Reserve = 6 GB on discrete GPUs, 10 GB on unified memory. A machine fits when demand is at most 90% of its memory tier.

Speed ceiling

Decode is memory-bound: ceiling = bandwidth ÷ bytes read per generated token (active weights plus one request’s KV cache). Batched streams share the weight read. Real throughput is lower; prefill is compute-bound and not modelled.

Cost

Local = hardware + setup + months × (software + administration + electricity) − resale. Electricity = watts ÷ 1000 × hours × tariff. Cloud = your monthly bill, or tokens × your provider’s price per million. Payback = upfront ÷ (cloud monthly − local monthly).

Evidence and editorial policy

Missing a tool? Tell the community what you would measure, or publish a measured run so estimates can be checked against reality.