The sizing is proven before procurement is signed
A sovereign deployment goes wrong at the purchase order rather than at the model. A node specified from published benchmark figures meets a real workload — concurrent readers, long documents, sustained rather than bursty demand — and falls short of it. We size against that workload and demonstrate throughput before hardware is bought. Quantisation is part of the sizing, not a later economy: the open-weight model is quantised to GGUF and fitted to the hardware that will be on site, not the hardware a model card assumes.
