How the accuracy was verified
94.3% is measured at −5 dB signal-to-noise, not on clean laboratory audio: at −5 dB the ambient noise is louder than the sound being classified. Evaluation ran per category rather than as one aggregate, because an average across fourteen classes hides the single class that matters on the night it matters. Latency was measured end to end on the deployment hardware under sustained load, not as a benchmark burst on a larger machine. Sub-50 ms is what the node does in service, not what the model does in isolation.
