Cantonese Voice Benchmark粵語語音基準

Cantonese Voice Benchmark 2026.1 · 8 entrants

What this edition found.

Every claim below carries how strongly the data supports it. A benchmark that refuses to publish a ranking over noise cannot then publish prose at uniform confidence, so it does not.

Established

The data supports this directly and the gap exceeds the noise.

Provisional

The direction is clear. The sample is thin.

Open

We looked and cannot yet say.

Established

The field separates into 3 bands, not 8 ranks.

Band sizes are 6 and 1 and 1. Entrants within a band cannot be ordered by this data. Character error rates run 13.3% to 52.3%.

What would change this A larger clip set would narrow the intervals and may split bands that currently overlap.

Established

Cantonese recognition quality varies materially between production systems.

A 39.0 point spread in character error rate between the strongest and weakest entrant, over 8 systems.

What would change this More entrants, particularly from outside the invited set.

Provisional

Cantonese speech synthesis remains the weakest link in the stack.

8 of 13 entrants who entered the speech pillar have a published score.

What would change this Enough entrants completing human evaluation to compare synthesis rather than describe its absence.

Established

Endpoint reliability, not model quality, decided at least one result.

7 runs produced no score because too few items completed. Those are reported as insufficient data rather than as a poor score.

What would change this Nothing. This is a measurement about operations, and it is worth reporting as one.

The field, with its uncertainty

Character error rate with 95% confidence intervalsOsprey7: 13.3 percent, interval 10.2 to 17.2. Undertow9: 15.8 percent, interval 12.3 to 19.8. Osprey8: 19.4 percent, interval 16.4 to 22.8. Quadrant10: 19.5 percent, interval 16.4 to 23.3. Zephyr9: 19.7 percent, interval 16.3 to 25.2. Sampan10: 19.8 percent, interval 15.6 to 24.4. Vanguard8: 26.7 percent, interval 23.2 to 30.7. Petrel8: 52.3 percent, interval 48.4 to 56.5. Open-source pipeline: 11.1 percent. Proprietary voice LLM: 8.6 percent20.0%40.0%60.0%Osprey713.3Undertow915.8Osprey819.4Quadrant1019.5Zephyr919.7Sampan1019.8Vanguard826.7Petrel852.3Open-source pipeline11.1Proprietary voice LLM8.6tied
Character error rate with 95% bootstrap intervals, resampled over clips. Entrants whose intervals overlap share a band and are published as tied — the overlap is the reason, and it is visible here rather than asserted. Lower is better.

Where the difficulty actually bites

Character error rate by clip difficultyOsprey7: easy 27.5 percent, normal 10.4 percent, hard 13.5 percent. Undertow9: easy 27.7 percent, normal 13.6 percent, hard 15.4 percent. Zephyr9: easy 41.6 percent, normal 17.0 percent, hard 18.6 percent. Osprey8: easy 30.3 percent, normal 17.2 percent, hard 19.6 percent. Quadrant10: easy 32.2 percent, normal 16.3 percent, hard 19.9 percent. Sampan10: easy 31.3 percent, normal 15.5 percent, hard 21.2 percent. Vanguard8: easy 37.6 percent, normal 24.6 percent, hard 26.8 percent. Petrel8: easy 69.1 percent, normal 55.9 percent, hard 60.8 percent0%40%80%27.510.413.5Osprey727.713.615.4Undertow941.617.018.6Zephyr930.317.219.6Osprey832.216.319.9Quadrant1031.315.521.2Sampan1037.624.626.8Vanguard869.155.960.8Petrel8
The same runs, split by clip difficulty. A headline figure is dominated by the easy third of the corpus; the hard tier carries the background noise, overlapping speakers and distressed callers that a contact centre actually produces. Two systems with the same headline can differ sharply here, which is the reason the split is published rather than the average alone.

Pillar coverage

Who entered each pillar, and who finished it. The gap between the two columns is usually the more interesting number.

PillarEnteredScored
ASR158
LLM149
TTS138
E2E80

Recognition accuracy

Best character error rate13.3%
Median19.7%
Weakest52.3%

Taking this away. Print the page and choose “Save as PDF”. The stylesheet drops the interface, keeps figures whole, and spells out where every link goes. We do not generate the PDF ourselves on purpose: one we produced could not embed a Chinese font without shipping megabytes with it, and every character in the 繁中 edition would come out as a box. Your browser already has the fonts.

How to read this

These findings are generated from the score table, not written by hand, so they change as the edition does. Anything graded open is an absence of evidence rather than evidence of absence — most often it means not enough entrants finished a pillar for a comparison to mean anything.

The methods behind every number, including the exact text normalization applied before scoring, are on the methodology page. The per-entrant results are on the leaderboard.