OnpremBench.aiby Understand.tech

Help improve this record. Suggest a correction

Field notebook

Deployment experience / NVIDIA GB10

Moving the whole stack from cloud GPUs to two Sparks.

UT’s reported migration connects the hardware decision to applications, operating effort and the assumptions behind the savings.

UT field noteUnderstand TechUnderstand Tech · founder-reported experienceReviewed 9 September 2026

CONTEXT

Why this comes up

In the LinkedIn post supplied to the Lab, UT’s founder describes moving its AI application workload from two AWS g6e.8xlarge instances, costing about $6,600 per month, to two NVIDIA DGX Spark systems in its office rack. The described stack includes local models, retrieval, agents and coding tools.

THE USEFUL LESSON

The founder reports an 88% reduction in year-one cost and roughly six-week payback. Separately, UT reports encountering both software and hardware integration challenges while putting its stack on GB10 and GB300. The next useful publication is the exact environment, operating lessons and cost evidence behind those claims.

ON YOUR MACHINE

What to record and test

  1. 01

    Record the before-and-after application stack, model versions and serving configurations.

  2. 02

    Publish the traffic volume, quality requirements, context lengths and concurrency used in the comparison.

  3. 03

    Document the hardware and software problems encountered, the fixes and the versions they apply to.

  4. 04

    Reconcile the cost comparison with equipment, setup, electricity, software, support and backup assumptions.

  5. 05

    Document sustained operation and recovery, including what remains unresolved.

Source: founder-supplied LinkedIn screenshots and deployment discussion. Financial claims have not been independently audited. Exact model configurations, measured capacity and detailed fixes are not published here. These results are not a savings forecast for another organization.