Every claim below carries how strongly the data supports it. A benchmark that refuses to publish a ranking over noise cannot then publish prose at uniform confidence, so it does not.
呢一頁暫時只有英文版。翻譯緊,會逐頁推出——首頁、排行榜同評測方法已經有中文。
Established
The data supports this directly and the gap exceeds the noise.
Provisional
The direction is clear. The sample is thin.
Open
We looked and cannot yet say.
Open
The speech pillar is the least populated, which is itself the finding.
No entrant has completed the speech pillar. 0 entered it.
What would change this Enough entrants completing human evaluation to compare synthesis rather than describe its absence.
The field, with its uncertainty
Character error rate with 95% bootstrap intervals, resampled over clips. Entrants whose intervals overlap share a band and are published as tied — the overlap is the reason, and it is visible here rather than asserted. Lower is better.
Pillar coverage
Who entered each pillar, and who finished it. The gap between the two columns is usually the more interesting number.
Pillar
Entered
Scored
How to read this
These findings are generated from the score table, not written by hand, so they change as the edition does. Anything graded open is an absence of evidence rather than evidence of absence — most often it means not enough entrants finished a pillar for a comparison to mean anything.
The methods behind every number, including the exact text normalization applied before scoring, are on the methodology page. The per-entrant results are on the leaderboard.