Verification, not benchmarks
Accuracy on a public benchmark says nothing about a fence line at three in the morning. Detection was evaluated inside the perimeter, on the same GPU hardware the system runs on — not a laboratory rig with a different memory profile — because an air-gapped site leaves nowhere else to run it. The false-positive figure was treated as the harder number, because the cost of a wrong alert is an operator who stops trusting the screen. 0.18% is the measured rate at 40+ concurrent feeds.
