OnpremBench.aiby Understand.tech

Evidence before enthusiasm

Know what every claim means.

Useful decisions need visible sources, assumptions and limitations.

Specifications are not performance.

Manufacturer specification

A published product configuration, linked to the original source. Exact SKUs, regional options and memory differences matter.

Contributor / founder-reported

An attributed observation supplied by its author. Publishing does not establish independent verification.

Lab-verified · future workflow

Reserved for a documented reproduction with pinned setup, raw evidence and review. No performance result published so far carries this label.

Planning scenario

Editable assumptions used to explore a decision. This includes the cost calculator, software allowances and finance scenarios. Published catalog offers are separate, dated supplier snapshots.

MEMBER RATINGS

Ratings are opinions, labelled as such.

Signed-in members can rate each machine once, from one to five stars, and say whether they run it, evaluated it hands-on or are exploring. That relationship is self-declared and shown with the rating. Stars enter the average immediately; written verdicts are published after a review by the Lab team, which checks for confidential information, supplier self-promotion and abuse, not for accuracy of opinion. A rating never stands in for a measured run.

The first-pass screen

Make the memory assumption explicit.

weights = parameters × effective bits ÷ 8 × 1.03 · KV = KB/token × context × requests · reserve = 6 GB (discrete) or 10 GB (unified) · fits if total ≤ 90% of the memory tier

Effective bits include quantisation scales (4-bit formats count 4.6 bits). KV bytes per token come from each model card: 2 × layers × KV heads × head size × 2 bytes, halved for an 8-bit KV cache. The speed ceiling is bandwidth ÷ bytes read per generated token (active weights plus one request’s KV cache); batched streams share the weight read. These are first-order estimates: runtimes, kernels, prefill and scheduling change the real numbers, so every screen ends in a measured run.

The screen compares this estimate with 75% of listed memory. For GB300 it uses only the 252 GB GPU figure; CPU offloading needs separate validation. HP’s system memory figure is not a promise that the same amount is available to its GPU.

Eligibility also checks requested CUDA. A EUR budget is checked only against a published EUR offer; other currencies and quote-only machines require a separate price check. Results prioritize fewer failed technical constraints, with catalog order as a tie-breaker. They do not rank measured speed.

Try the finder

Who publishes the Lab.

OnPremBench is published by Understand Tech. Understand Tech also sells AI-in-a-Box, a managed deployment on NVIDIA GB10 and GB300 appliances, described on its own page. That offer does not change the records: specifications, prices, evidence labels and planning figures follow the same rules for every manufacturer, the catalogue opens sorted by price, and no machine is placed or ranked for payment. Understand Tech’s own deployments appear as use cases labelled founder-reported, like any other attributed source.

Supplier offers and planning scenarios are separate.

Catalog prices are snapshots taken from linked supplier pages on 7 September 2026. Each offer states its currency, region, selected configuration and any published tax treatment. They are not live inventory feeds or UT offers. A missing, zero or ambiguous store price is shown as a quote request. DIY kits and integrated workstations are identified separately. Software, support and setup allowances remain editable examples, not UT pricing.

The calculator uses an independent, editable scenario in the currency you choose (USD by default). It does not convert foreign prices or use PSU ratings as average power. Enter your actual quote and measured consumption. It includes equipment, setup, electricity, software/support and administration, less end-of-horizon residual value. Compare with a cloud option meeting the same quality, demand and availability. Hardware you already own can start at zero incremental purchase cost; include depreciation or replacement costs if relevant to your decision.

The finance scenario uses a 6% nominal annual interest rate with no fees, deposit or residual. It is not a lender offer. Eligibility and commercial terms remain unconfirmed.

Local inference may avoid a per-token API charge, but the system still has equipment, electricity and operating costs. Token speed is different from financial token consumption.

Reviewed 7 September 2026

Real systems. Traceable sources.

NVIDIA DGX Spark

Memory capacity does not establish interactive speed or simultaneous request capacity. UT has documented a Spark deployment. Confirm exact software versions and support scope.

Manufacturer
NVIDIA DGX Station

748 GB coherent memory is 252 GB HBM3e on the GPU plus 496 GB LPDDR5X on the CPU, joined by NVLink-C2C. The two tiers have very different bandwidths; large models spill from the fast tier into the slow one. NVIDIA does not sell the DGX Station directly: it is ordered through partner manufacturers, each with its own chassis, drives, warranty and price. Partner listings observed in 2026 ran from about $85,000 to $175,000 depending on configuration. NVIDIA specifies 1,600 W total system power and a 20 A circuit. Confirm OEM and regional installation requirements before choosing an office outlet. The GB300 can be paired with an RTX PRO 6000 Blackwell workstation GPU for graphics and additional compute. DGX Station for Windows was announced on 31 May 2026 for Q4 2026 availability from the same partners.

Manufacturer
Dell Pro Max with GB300

748 GB coherent memory comprises 252 GB HBM3e GPU plus 496 GB LPDDR5X CPU memory. These tiers have different bandwidths. Regional snapshots cover specific configurations with RTX PRO 2000, four 4 TB drives and 12-month ProSupport. They are not entry prices or directly comparable landed costs. UT’s founder reports a deployment of this system; comparable reproducible measurements are not yet published in this lab.

Manufacturer
AMD Ryzen AI Halo

The reference configuration is Ryzen AI Max+ 395 with 128 GB unified memory. AMD lists a PRO 495 version supporting 192 GB as coming soon; it is a separate configuration. The Lab has not tested this system or validated the UT stack.

Manufacturer
Aetina MegaEdge AIP-FR68

Supports up to two Cloud AI 100 Ultra cards. Confirm the accelerator population and memory with the supplier. 128 GB is the memory-screen reference for one Cloud AI 100 Ultra card; availability to a compiled model depends on the runtime. Qualcomm model conversion and runtime support require validation. A CUDA result does not establish compatibility. The AIP-FR68-A2 orderable bare system excludes CPU, RAM, SSD, GPU card and power supply. Obtain a complete accelerator bundle quote.

Manufacturer
Arduino VENTUNO Q

Manufacturer specifies 40 dense TOPS and 16 GB shared LPDDR5 memory. TOPS is not an LLM tokens-per-second benchmark. A physical-AI evaluation needs the actual sensor, camera or actuator setup as well as the compute board. The Lab’s generic LLM memory heuristic is conservative and is not a prediction of NPU model support.

Manufacturer
AMD Threadripper Halo Station

AMD describes this as a prototype first shown at IFA 2026, coming in 2027. No verified orderable price. 576 GB is aggregate GPU memory across four 144 GB accelerators. It is not one unified GPU allocation; model sharding is required. The model memory screen uses a single 144 GB accelerator. CPU RAM and other accelerators are excluded from that screen.

Manufacturer
ASUS Ascent GX10

Compare exact storage configuration and regional support coverage. Validate the complete runtime and workload before a production choice.

Manufacturer
Dell Pro Max with GB10

Confirm region, storage, warranty and exact shipping configuration with Dell. A shared chip does not imply identical thermals, service terms or measured performance.

Manufacturer
HP ZGX Nano G1n

Product configuration and availability depend on region and SKU. UT compatibility and managed service coverage require review. HP's US store did not expose a usable price on the review date; request a quote for current pricing.

Manufacturer
Lenovo ThinkStation PGX

Specifications vary by configured SKU and region. Managed deployment through UT needs a compatibility and support review.

Manufacturer
Acer Veriton GN100

The listed US offer is the VGN100-UD13 configuration with 128 GB unified memory and a 2 TB NVMe drive; other regional SKUs may differ in storage and bundled software. Acer positions the GN100 on NVIDIA's GB10 reference platform. Vendor and community GB10 guidance applies to the platform; no Lab measurement exists for this exact chassis. Memory capacity alone does not establish speed, runtime compatibility or production capacity.

Manufacturer
GIGABYTE AI TOP ATOM

GIGABYTE publishes three ATAGB10 variants differing in storage (1 TB Gen4, 4 TB Gen4, 4 TB Gen5). Confirm which variant a regional seller offers; no usable public price was verified on the review date. The ATOM shares NVIDIA's GB10 platform. Platform guidance applies; no Lab measurement exists for this exact chassis. Memory capacity alone does not establish speed, runtime compatibility or production capacity.

Manufacturer
MSI EdgeXpert

The US store offer is the EdgeXpert-11SUS single unit with 128 GB unified memory and a 4 TB Gen4 drive. MSI also markets paired-unit configurations; confirm networking accessories for a two-node setup. The EdgeXpert shares NVIDIA's GB10 platform. Platform guidance applies; no Lab measurement exists for this exact chassis. Memory capacity alone does not establish speed, runtime compatibility or production capacity.

Manufacturer
ASUS ExpertCenter Pro ET900N G3

252 GB GPU and 496 GB CPU memory are coherent but do not have equal bandwidth. The manufacturer specifies two pre-installed 2 TB OS drives. Confirm usable RAID capacity and optional data drives.

Manufacturer
MSI XpertStation WS300

The EUR figure is a German reseller preorder for the WS300T60L with two 2 TB SSDs, quoted before final configuration. It is not directly comparable with the Dell GB300 US snapshot, which includes an RTX PRO 2000 card, four 4 TB drives and 12-month ProSupport. The WS300 shares NVIDIA's GB300 platform. Platform guidance applies; no Lab measurement exists for this exact chassis. MSI and the reseller list differing PSU efficiency and chassis dimensions. This page uses MSI's platform specifications; confirm the exact WS300T60L revision in the quote.

Manufacturer
HP ZGX Fury

HP states 748 GB coherent memory. Confirm the exact local configuration and delivery schedule with HP. Manufacturer maximum-model claims depend on quantization and workload; they are not lab benchmarks.

Manufacturer
GIGABYTE W775-V10-L01

This vendor reference is a barebone workstation platform. A supplier must specify SSDs, OS, display GPU and the final ready-to-use configuration. 252 GB GPU plus 496 GB CPU memory; power and circuit requirements need checking before installation.

Manufacturer
Supermicro Super AI Station · ARS-511GD

Optional rackmount kit; confirm rails, rack depth, circuit and acoustic suitability with the integrator. 748 GB combines 252 GB GPU and 496 GB CPU memory. An optional display GPU has its own separate memory.

Manufacturer
Exxact Valence DGX Station

One of the few GB300 workstations with a public starting price. The starting configuration excludes drives, the optional display GPU and support options; the configurator sets the final price. Same GB300 superchip, memory and networking as every DGX Station; the chassis, cooling and warranty are Exxact’s. Supports one double-wide graphics card for display alongside the GB300.

Manufacturer
HP Z8 Fury G6i

Up to 2 TB DDR5 ECC system RAM is separate from GPU memory. The memory screen uses one 96 GB GPU. Multiple GPUs do not automatically form one memory pool. Framework support, sharding and interconnect overhead must be evaluated.

Manufacturer
Lenovo ThinkStation P7

September 2026 PSREF supports up to three RTX PRO 6000 Blackwell Max-Q GPUs or one 600 W Workstation Edition GPU, with orderability still to confirm. Memory screening uses a single 96 GB GPU, not an assumed pool across three cards.

Manufacturer
Framework Desktop · Ryzen AI Max+

The 128 GB memory is soldered. GPU-available memory depends on OS and configuration. The listed $3,449 is the 128 GB system kit selection, not a ready-to-run total. SSD, OS, fan, cable and tiles are additional selections.

Manufacturer
MINISFORUM MS-S1 MAX

The selected 128 GB offer lists estimated mid-September shipping. Confirm delivery before ordering. PCIe x16 physical slot is wired PCIe 4.0 x4; do not assume x16 bandwidth. Shared memory allocation depends on runtime and OS.

Manufacturer
HP Z2 Mini G1a

Reference SKU E07RFPA. GPU-accessible memory depends on configuration; do not assume all system memory is available to the model. Runtime and full UT software support require validation.

Manufacturer
Apple Mac Studio · M5 Ultra

Apple store lists this configuration for pre-order, available from 22 September. No usable price was exposed in the reviewed store page. This is a new configuration. Runtime compatibility, actual GPU memory allowance and performance have not been validated by this lab.

Manufacturer
Apple Mac Studio · M3 Ultra

This is a specific M3 Ultra configuration, not a claim that it is the latest Mac Studio. CUDA workloads need a different platform. UT’s complete managed stack is not validated here.

Manufacturer

Images identify reference hardware and come from manufacturers or product partners. Inclusion does not imply a partnership, certification or permission to display partner badges.

Original image manifest Expanded catalog image manifest

What works here, and what comes next.

The explorer, comparison, workload screening, cost scenarios, scripted walkthrough and downloadable briefs are interactive. Machine comments, replies, helpful votes, concern reports, community notes, linked reproduction reports, saved package progress and pilot requests are saved to the lab database. Authors can edit or remove their machine comments. Roles and affiliations are self-declared; they are not verified badges. Media contributions use external HTTPS links. Deployment packages are versioned evaluation templates, not tested installation bundles. My lab tracks progress in-app; email alerts and automatic evidence capture are not connected. Authors can remove their own notes; pilot records are shown only to their requester.

Approved contributions are public. Comparison preferences stay in this browser; deployment briefs download locally and are not sent to UT.

Live hardware sessions, reservations, payments, live hardware sessions and supplier inventory are not connected. Contributions are moderated by the Lab team before publication; concern reports are recorded and reviewed, but reporting does not automatically hide a comment. Contribution terms are set out in the legal notice. Before any live trial, isolation, data handling and cleanup are agreed in writing.

Explore the adoption roadmap