Throughput is proven before procurement is signed
We size the node against the workload you actually have — sustained inference, not a benchmark burst — and demonstrate the figure before a capital request goes in. That exercise is where the real bottleneck appears. On ICCS, 80 configured cameras run on a single GPU node with roughly 15 to 20 concurrently active 1440p streams, and the ceiling turns out to be CPU rather than GPU; the fused eight-model detection stack needs about 28 GB of VRAM. Both are facts you want from a bench, not from a board paper.
