Cantonese Voice Benchmark粵語語音基準

Cantonese Voice Benchmark 2026.1 · 0 entrants

What this edition found.

Every claim below carries how strongly the data supports it. A benchmark that refuses to publish a ranking over noise cannot then publish prose at uniform confidence, so it does not.

Established

The data supports this directly and the gap exceeds the noise.

Provisional

The direction is clear. The sample is thin.

Open

We looked and cannot yet say.

Open

The speech pillar is the least populated, which is itself the finding.

No entrant has completed the speech pillar. 0 entered it.

What would change this Enough entrants completing human evaluation to compare synthesis rather than describe its absence.

The field, with its uncertainty

Character error rate with 95% confidence intervalsOpen-source pipeline: 11.1 percent. Proprietary voice LLM: 8.6 percent9.0%10.0%11.0%Open-source pipeline11.1Proprietary voice LLM8.6
Character error rate with 95% bootstrap intervals, resampled over clips. Entrants whose intervals overlap share a band and are published as tied — the overlap is the reason, and it is visible here rather than asserted. Lower is better.

Pillar coverage

Who entered each pillar, and who finished it. The gap between the two columns is usually the more interesting number.

PillarEnteredScored

How to read this

These findings are generated from the score table, not written by hand, so they change as the edition does. Anything graded open is an absence of evidence rather than evidence of absence — most often it means not enough entrants finished a pillar for a comparison to mean anything.

The methods behind every number, including the exact text normalization applied before scoring, are on the methodology page. The per-entrant results are on the leaderboard.