Help improve this record. Suggest a correction
An inference endpoint for your local applications
Connect applications to a model server on a GB10 or GB300 reference platform using vLLM.
THE SETUP
An upstream starting point
NVIDIA documents container-based serving and an OpenAI-compatible API for Spark and Station. Its multi-node instructions are scoped to Spark. The linked model catalog supplies launch configurations.
Open NVIDIA’s vLLM playbookWhat supports this record
Upstream platform guidance reviewed 9 September 2026. Listed support is not an exact-OEM benchmark or a UT reproduction.
THE KNOWLEDGE AROUND IT