Every claim below carries how strongly the data supports it. A benchmark that refuses to publish a ranking over noise cannot then publish prose at uniform confidence, so it does not.
Established
The data supports this directly and the gap exceeds the noise.
Provisional
The direction is clear. The sample is thin.
Open
We looked and cannot yet say.
Established
The field separates into 3 bands, not 8 ranks.
Band sizes are 6 and 1 and 1. Entrants within a band cannot be ordered by this data. Character error rates run 13.3% to 52.3%.
What would change this A larger clip set would narrow the intervals and may split bands that currently overlap.
Established
Cantonese recognition quality varies materially between production systems.
A 39.0 point spread in character error rate between the strongest and weakest entrant, over 8 systems.
What would change this More entrants, particularly from outside the invited set.
Provisional
Cantonese speech synthesis remains the weakest link in the stack.
8 of 13 entrants who entered the speech pillar have a published score.
What would change this Enough entrants completing human evaluation to compare synthesis rather than describe its absence.
Established
Endpoint reliability, not model quality, decided at least one result.
7 runs produced no score because too few items completed. Those are reported as insufficient data rather than as a poor score.
What would change this Nothing. This is a measurement about operations, and it is worth reporting as one.
The field, with its uncertainty
Character error rate with 95% bootstrap intervals, resampled over clips. Entrants whose intervals overlap share a band and are published as tied — the overlap is the reason, and it is visible here rather than asserted. Lower is better.
Where the difficulty actually bites
easynormalhard
The same runs, split by clip difficulty. A headline figure is dominated by the easy third of the corpus; the hard tier carries the background noise, overlapping speakers and distressed callers that a contact centre actually produces. Two systems with the same headline can differ sharply here, which is the reason the split is published rather than the average alone.
Pillar coverage
Who entered each pillar, and who finished it. The gap between the two columns is usually the more interesting number.
Pillar
Entered
Scored
ASR
15
8
LLM
14
9
TTS
13
8
E2E
8
0
Recognition accuracy
Best character error rate
13.3%
Median
19.7%
Weakest
52.3%
Taking this away. Print the page and choose “Save as PDF”. The stylesheet drops the interface, keeps figures whole, and spells out where every link goes. We do not generate the PDF ourselves on purpose: one we produced could not embed a Chinese font without shipping megabytes with it, and every character in the 繁中 edition would come out as a box. Your browser already has the fonts.
How to read this
These findings are generated from the score table, not written by hand, so they change as the edition does. Anything graded open is an absence of evidence rather than evidence of absence — most often it means not enough entrants finished a pillar for a comparison to mean anything.
The methods behind every number, including the exact text normalization applied before scoring, are on the methodology page. The per-entrant results are on the leaderboard.