How the figure is stated, and what it is not
91.4% is a word-error-rate improvement over the base model, not an absolute accuracy score, and we state it that way deliberately — an absolute number on transcription means little without the audio it was measured on. Evaluation ran language by language on held-out recordings drawn from the same conditions as the work, including poor field audio rather than only clean capture. Twelve languages is twelve evaluation problems: an average across them would hide the weakest, which is the only one worth arguing about.
