Cantonese Voice Benchmark 2026.1 · Qualification Round
Production voice systems evaluated on Hong Kong insurance contact-centre audio across four pillars — hearing, understanding, speaking, and conversation.
Six tones, six words — and every one of them is said in the calls this benchmark is scored on. Play one: the line is its pitch, and getting that shape wrong is what swaps one word for another.
| Rank | Team | Core | ASRCER | LLMJudge | TTSMOS | E2ENatural | TTFB |
|---|---|---|---|---|---|---|---|
| 1tied | MediConCenliveNamed | 95.6 | 13.3%10.2–17.2 | 4.90 / 54.77–5.00 | 5.00 / 5measured, not rated | —Organisers run this | 4.64sp95 6.30 |
| 1 | Sampan10live | 94.7 | 19.8%15.6–24.4 | 5.00 / 55.00–5.00 | 5.00 / 5measured, not rated | —Not entered | 0.18sp95 0.76 |
| 1 | Ener-GliveNamed | 94.1 | 15.8%12.3–19.8 | 4.80 / 54.61–4.95 | 5.00 / 5measured, not rated | —Not entered | 1.31sp95 2.71 |
| 4tied | Vibran-CliveNamed | 91.6 | 19.5%16.4–23.3 | 5.00 / 55.00–5.00 | 4.68 / 5measured, not rated | —Not entered | 0.41sp95 1.65 |
| 4 | Osprey8live | 90.9 | 19.4%16.4–22.8 | 4.53 / 54.14–4.89 | 5.00 / 5measured, not rated5.00–5.00 | —Not entered | 5.17sp95 6.67 |
| Reference baselinesIndicative figures for scale. Not measured on this corpus — no entrant is ranked against them. | |||||||
| A | Open-source pipeline | 76.2 | 11.1% | 3.81 / 5 | 3.44 / 5 | 3.29 / 5 | 2.11sp95 3.88 |
| B | Proprietary voice LLM | 82.7 | 8.6% | 4.14 / 5 | 3.62 / 5 | 4.01 / 5 | 0.88sp95 1.52 |
What the customer says
我想查下我張保單嘅供款
They stop speaking. The clock starts here.
In the gap: hear it, understand it, say it back
What they hear back
你嘅月供係一千二百蚊
Why this is hard
Cantonese carries six lexical tones, writes no spaces between words, and switches into English mid-sentence without marking the boundary. A model tuned on Mandarin does not degrade gracefully here — it fails in ways that read as fluent.
buy · sell
Low rising against low level. In a sales call this inverts who is doing what.
insure · report
High rising against mid level. 保險 is a policy; 報險 is filing a claim against one. Both are said constantly on the same call.
medical · two
High level against low level. A claims call is dense with both hospital words and read-aloud digits, so this one lands in the amount as easily as the diagnosis.
Hear all six tones on the recognition pillar — the contours are plotted and playable.
What the number is
Every score on this page comes from one operation repeated across the draw: line up what was said against what the system sent back, and count the places they differ.
“I want to buy this critical-illness policy, sum insured five hundred thousand.” A line from a claims call, carrying both tone confusions from the pairs above.
Both sides are normalised before they are compared: punctuation dropped, Chinese numerals written as digits. 五十萬 becomes 500000.
Three of the four change what the call is about. 買 buy came back as 賣 sell. 保險 the policy came back as 報險 reporting a claim against it. And the missing character is a missing zero — five hundred thousand of cover, read back as fifty. The rate counts all four the same, which is why the recognition pillar is not the whole score.
No files are uploaded and nothing is scored on request. We connect to a server the entrant runs, stream audio at it, and measure what comes back on our own clock.
A WebSocket server you host. We are the client, so nothing of yours has to be handed over and nothing of ours has to be trusted.
8 clips with published ground truth, downloadable once you have an account. Unlimited runs, full per-clip diffs, no effect on your score.
30 clips drawn at random from a pool of 56, stratified across difficulty tiers, seeded and recorded so the draw can be reconstructed.
Each clip resolves to a verdict and an owner: yours, ours, the network, or nobody. A run we break costs you nothing and does not consume quota.
A bootstrap interval accompanies every number, and entrants whose intervals overlap are published as tied rather than ranked against noise.
A benchmark is only worth reading if it is willing to say less than it could. These are the four cases where we withhold a number rather than produce one.
The widest interval so far is ±4.47 points. Gaps smaller than that are reported as a shared band.
A run that completed too few clips is published as insufficient data. Turning a broken integration into a bad score would misattribute an operations problem as model quality.
Pillars are opt-in. A team that did not enter speech synthesis is shown as not entered, never as scoring nothing.
The language-model judge runs 3 times across 2 Amazon Nova models and is correlated against human raters before its pillar counts at all.
Organised in Hong Kong, on Hong Kong audio, judged by Cantonese speakers. The dataset is real contact-centre recordings, consented and masked, never scraped.
2026-1105 clips held: 56 in rotation, 30 for qualification, 11 held back for final verification, 8 to practise on.
The first real recording handed to us was GSM 6.10, which no browser will decode.
The rules promised it. Until now it required emailing someone, which is a different promise.
One to two pages on what was measured. The Full Score is withheld until it is in.
Taking part
No fee, no exclusivity, and you keep everything. The core is recognition, understanding and speech synthesis; conversation is optional. You do not have to enter all three — the Core Score is rescaled across what you did enter, never counting a pillar you skipped as a zero. Requests are reviewed within one business day.