THE FIELD NOTEBOOK
The lessons between unboxing and useful AI.
The operating knowledge behind complete AI machines: compatibility, shared workloads, data privacy, secure operation and token costs. Every useful lesson belongs to an exact configuration.
Every hard-won fix can help the next deployment.
UT is learning by putting its stack on GB10 and GB300. This notebook connects UT’s runtime journey, startup investigations and engineering applications with attributed technical sources and evaluation checklists. Detailed UT engineering reports and community reproductions can grow alongside them.
Ollama, vLLM, NIM: nine months of deployment learning.
UT’s reported journey across serving stacks, hardware integration and customer feedback. A starting record for the decisions and evidence still to publish.
A NIM container that took about 30 minutes to start.
An operator-reported startup observation from UT’s GB10/GB300 work. The next step is to identify the exact configuration and separate the startup phases.
Local AI beside a Wi-Fi module test bench.
An execution agent and confidential engineering test assets show how on-premises AI extends beyond office applications.
Moving the whole stack from cloud GPUs to two Sparks.
UT’s reported migration connects the hardware decision to applications, operating effort and the assumptions behind the savings.
The model fits on paper. Why does loading still fail?
On Spark, unified memory and application behavior need to be considered together. A memory estimate is only the first check.
A second test changed the diagnosis.
A reported multi-GPU limitation turned out to depend on the test environment. Community feedback led to a correction.
Does the whole application work offline?
Local model inference is one part of the path. Test retrieval, documents, identity, tools and restart behavior as well.
One box. Several applications. Where is the limit?
Move beyond a single chat window: measure shared capacity under realistic application traffic.
What does an accepted result actually cost?
Put quality, retries, utilization and operating effort beside token throughput.
THE NEXT PAGE COULD BE YOURS
What took longer than it should have?
A driver problem. A container that wouldn’t start. A model that slowed down under load. A dependency you discovered after disconnecting the network. Give someone the starting point you wish you had.
Open a field-note template