CONTEXT
Why this comes up
A workstation can be evaluated as a shared inference service for multiple applications. Capacity needs to be established for the actual model, runtime and request mix.
THE USEFUL LESSON
Measure acceptable work completed while application traffic competes for the same machine. Employee count and simultaneous inference requests are different quantities.
ON YOUR MACHINE
What to record and test
- 01
Pin the models, precision, context lengths and serving configuration. Define quality and latency requirements.
- 02
Establish each application’s baseline independently, including preprocessing and retrieval.
- 03
Mix interactive traffic with background work at increasing request concurrency.
- 04
Record p50 and p95 latency, aggregate generation, per-request speed, errors, queueing and memory.
- 05
Test a service restart and document the effect on active work. Publish the tested capacity and its boundaries.