The compute budget is the specification
Most AI briefs describe the output and leave the hardware to the end. We invert that. The node is specified against sustained inference — the load a system carries all day, not a benchmark burst — and throughput is demonstrated before procurement is signed. Quantisation, batching and model count are sizing decisions, not optimisations bolted on afterwards. ICCS runs an eight-model fusion stack in roughly 28 GB of VRAM on a single GPU node, and at load it is CPU-bound rather than GPU-bound. That result came from sizing first.
