Help improve this record. Suggest a correction
Use casesOFFICE / UNDERSTAND TECH
Two office Sparks serving a company’s AI stack.
Understand Tech moved its applications and shared inference from two rented cloud GPU instances to two DGX Spark units on the office floor: retrieval, agents and coding tools for the team.
- Story
- Stack
- Load
- Measurements
“Two months ago it was a slide. Today it’s on our floor.”Naama Bak · Co-founder, Understand Tech
THE SETTING
Who uses it, and where.
Software company · its own office. Environment: office.
Users. The Understand Tech team. Concurrent users: not yet published.
THE STACK
What runs on the machine.
- Understand Tech platform: retrieval, agents and coding tools.
- Serving stacks tried along the way: Ollama, vLLM, NIM.
- Models and runtime versions: not yet published.
Network posture. Machines in the office; the stack replaced two cloud GPU instances.
THE WORKLOADS
What it does, day to day.
- The company’s own applications and shared inference.
- Coding tools for the team.
- Retrieval over internal documents.
OUTCOME
What changed.
Reported · not yet measuredFounder-reported saving compared with the previous cloud GPU bill; the basis of the figure and a comparable throughput run are not attached.
EVIDENCE & LIMITS
How to read this record.
Founder-supplied screenshots and deployment discussion. Not independently audited.
LESSONS & FAILURES
What the team learned.
- Getting the model to run was only the beginning: nine months across Ollama, vLLM and NIM.
- A NIM container that took about 30 minutes to start: separate download, preparation, loading and readiness.
- Check that every container and native dependency supports arm64.
RELATED RECORDS