OnpremBench.aiby Understand.tech

THE FIELD NOTEBOOK

The lessons between unboxing and useful AI.

The operating knowledge behind complete AI machines: compatibility, shared workloads, data privacy, secure operation and token costs. Every useful lesson belongs to an exact configuration.

Share your experience

Every hard-won fix can help the next deployment.

UT is learning by putting its stack on GB10 and GB300. This notebook connects UT’s runtime journey, startup investigations and engineering applications with attributed technical sources and evaluation checklists. Detailed UT engineering reports and community reproductions can grow alongside them.

An open notebook in the making
Infrastructure

Ollama, vLLM, NIM: nine months of deployment learning.

UT’s reported journey across serving stacks, hardware integration and customer feedback. A starting record for the decisions and evidence still to publish.

NVIDIA GB10NVIDIA GB300
Startup & recovery

A NIM container that took about 30 minutes to start.

An operator-reported startup observation from UT’s GB10/GB300 work. The next step is to identify the exact configuration and separate the startup phases.

NVIDIA GB10NVIDIA GB300
Engineering applications

Local AI beside a Wi-Fi module test bench.

An execution agent and confidential engineering test assets show how on-premises AI extends beyond office applications.

All platforms
Deployment experience

Moving the whole stack from cloud GPUs to two Sparks.

UT’s reported migration connects the hardware decision to applications, operating effort and the assumptions behind the savings.

NVIDIA GB10
Memory & runtime

The model fits on paper. Why does loading still fail?

On Spark, unified memory and application behavior need to be considered together. A memory estimate is only the first check.

NVIDIA GB10
Reproduction

A second test changed the diagnosis.

A reported multi-GPU limitation turned out to depend on the test environment. Community feedback led to a correction.

AMD
Offline operation

Does the whole application work offline?

Local model inference is one part of the path. Test retrieval, documents, identity, tools and restart behavior as well.

All platforms
Shared inference

One box. Several applications. Where is the limit?

Move beyond a single chat window: measure shared capacity under realistic application traffic.

NVIDIA GB10NVIDIA GB300RTX PRO
Economics

What does an accepted result actually cost?

Put quality, retries, utilization and operating effort beside token throughput.

All platforms

THE NEXT PAGE COULD BE YOURS

What took longer than it should have?

A driver problem. A container that wouldn’t start. A model that slowed down under load. A dependency you discovered after disconnecting the network. Give someone the starting point you wish you had.

Open a field-note template