Methodology · 27 Jul, 05:14 HKT
三十條片之下,零點幾個百分比就係雜訊;將雜訊排名,就係排行榜變成表演嘅開始。
A scored run draws thirty clips. On thirty clips, the difference between 6.0% and 6.1% character error rate is not a difference: resample the same clips and the order can swap. Publishing a 1 and a 2 in that situation invents a fact.
Every number we publish carries a bootstrap confidence interval, and entries whose intervals overlap share a rank band. They appear as tied, not ordered. On the current board that means bands of three and one rather than four separate ranks, and the report says so in those words.
This is not a hedge. A benchmark that will not publish a ranking over noise is more useful than one that will, because the rankings it does publish mean something. The cost is that we sometimes have to tell an entrant we cannot say they beat someone. That is the correct answer.
Cantonese Voice Benchmark · Published 27 Jul, 05:14 HKT.