An AI application stack on two office Sparks
UT’s move from rented cloud GPUs to machines in its own office, and the engineering record behind it.
INFRASTRUCTURE / ACROSS PLATFORMS
Model servers, runtime configurations and deployment recipes. The machine, software, evidence and open questions in one place.
UT’s move from rented cloud GPUs to machines in its own office, and the engineering record behind it.
NVIDIA’s Open WebUI and Ollama guide, connected to the hardware and operating questions it raises.
A professional-services firm’s on-premises deployment: applications, confidentiality and the expertise to operate the system. The firm is not named at its request.
Connect applications to a model server on a GB10 or GB300 reference platform using vLLM.
Explore SGLang serving, structured output and the conditions to check on a GB300 workstation.
A source-based starting point for running llama.cpp on Spark and documenting your own configuration.
AMD’s gpt-oss-20b example connects a GGUF model to llama.cpp with ROCm acceleration.
Use MLX-LM and converted model weights as a starting point on Apple Silicon.
Follow Qualcomm’s Efficient Transformers path from model weights to accelerator execution.
NVIDIA’s local gpt-oss support, connected to the details needed for an OEM workstation evaluation.
Follow a recipe, adapt it to your machine, and share the outcome. Include unsuccessful attempts.