Cantonese Voice Benchmark粵語語音基準

Cantonese Voice Benchmark 2026.1 · Qualification Round

Cantonese voice, measured on real calls.

Production voice systems evaluated on Hong Kong insurance contact-centre audio across four pillars — hearing, understanding, speaking, and conversation.

Six tones · the words of a claim

Six tones, six words — and every one of them is said in the calls this benchmark is scored on. Play one: the line is its pitch, and getting that shape wrong is what swaps one word for another.

Window
open
Clips drawn
30
Entrants
10
Updated
24 Sept, 00:24 HKT
Core Score (ASR + LLM + TTS). Entries whose 95% confidence intervals overlap share a rank band and are reported as tied.
 Rank Team CoreASRCERLLMJudgeTTSMOSE2ENatural TTFB
1tied
MediConCenliveNamed
95.6
13.3%10.2–17.2
4.90 / 54.77–5.00
5.00 / 5measured, not rated
Organisers run this
4.64sp95 6.30
1
Anonymous by default. Being named is the entrant’s choice, reversible until results publish; after that a named entry can still be un-named on request.Sampan10live
94.7
19.8%15.6–24.4
5.00 / 55.00–5.00
5.00 / 5measured, not rated
Not entered
0.18sp95 0.76
1
Ener-GliveNamed
94.1
15.8%12.3–19.8
4.80 / 54.61–4.95
5.00 / 5measured, not rated
Not entered
1.31sp95 2.71
4tied
Vibran-CliveNamed
91.6
19.5%16.4–23.3
5.00 / 55.00–5.00
4.68 / 5measured, not rated
Not entered
0.41sp95 1.65
4
Anonymous by default. Being named is the entrant’s choice, reversible until results publish; after that a named entry can still be un-named on request.Osprey8live
90.9
19.4%16.4–22.8
4.53 / 54.14–4.89
5.00 / 5measured, not rated5.00–5.00
Not entered
5.17sp95 6.67
Reference baselinesIndicative figures for scale. Not measured on this corpus — no entrant is ranked against them.
A
Open-source pipeline
76.2
11.1%
3.81 / 5
3.44 / 5
3.29 / 5
2.11sp95 3.88
B
Proprietary voice LLM
82.7
8.6%
4.14 / 5
3.62 / 5
4.01 / 5
0.88sp95 1.52
1MediConCenlive95.6
ASR
13.3%
10.2–17.2
LLM
4.90
4.77–5.00
TTS
5.00
measured, not rated
E2E
Organisers run this
TTFB 4.64s p50 · 6.30s p95
1Anonymous by default. Being named is the entrant’s choice, reversible until results publish; after that a named entry can still be un-named on request.Sampan10live94.7
ASR
19.8%
15.6–24.4
LLM
5.00
5.00–5.00
TTS
5.00
measured, not rated
E2E
Not entered
TTFB 0.18s p50 · 0.76s p95
1Ener-Glive94.1
ASR
15.8%
12.3–19.8
LLM
4.80
4.61–4.95
TTS
5.00
measured, not rated
E2E
Not entered
TTFB 1.31s p50 · 2.71s p95
4Vibran-Clive91.6
ASR
19.5%
16.4–23.3
LLM
5.00
5.00–5.00
TTS
4.68
measured, not rated
E2E
Not entered
TTFB 0.41s p50 · 1.65s p95
4Anonymous by default. Being named is the entrant’s choice, reversible until results publish; after that a named entry can still be un-named on request.Osprey8live90.9
ASR
19.4%
16.4–22.8
LLM
4.53
4.14–4.89
TTS
5.00
measured, not rated
5.00–5.00
E2E
Not entered
TTFB 5.17s p50 · 6.67s p95
Reference baselinesIndicative figures for scale. Not measured on this corpus — no entrant is ranked against them.
AOpen-source pipeline76.2
BProprietary voice LLM82.7
Tied
Entries sharing a rule have overlapping 95% confidence intervals. Reported as tied, not ordered.
Not entered
The team did not compete on that pillar. Never shown as a zero.
Pending
Awaiting human evaluation. The Full Score is withheld until every entered pillar is final.
Batch
Results uploaded rather than measured live. Recognition only.
Full leaderboardHow we measureRequest accessRulesFAQ

What the customer says

我想查下我張保單嘅供款

They stop speaking. The clock starts here.

In the gap: hear it, understand it, say it back

Proprietary voice LLM0.88s · p95 1.52s
Open-source pipeline2.11s · p95 3.88s

What they hear back

你嘅月供係一千二百蚊

Median time to first sound after the caller stops, for the two reference systems published on the leaderboard. The bars run for the seconds they show.
ASR
Can you hear?
Character error rate on tonal, code-switched, insurance-domain speech.
LLM
Can you understand?
Policy-grounded reasoning, hallucination avoidance, information extraction.
TTS
Can you speak?
The least-solved pillar. Six tones, domain terms, natural code-switching.
E2E
Can you converse?
Voice in, voice out. Latency and turn-taking under real conditions.

Why this is hard

Six tones. Change the pitch, change the word.

Cantonese carries six lexical tones, writes no spaces between words, and switches into English mid-sentence without marking the boundary. A model tuned on Mandarin does not degrade gracefully here — it fails in ways that read as fluent.

maai5vsmaai6

buy · sell

Low rising against low level. In a sales call this inverts who is doing what.

bou2vsbou3

insure · report

High rising against mid level. 保險 is a policy; 報險 is filing a claim against one. Both are said constantly on the same call.

ji1vsji6

medical · two

High level against low level. A claims call is dense with both hospital words and read-aloud digits, so this one lands in the amount as easily as the diagnosis.

Hear all six tones on the recognition pillar — the contours are plotted and playable.

What the number is

Four errors in seventeen characters.

Every score on this page comes from one operation repeated across the draw: line up what was said against what the system sent back, and count the places they differ.

Said
我想買呢份危疾保險,保額五十萬
Heard
我我想賣呢份危疾報險,保額五萬

“I want to buy this critical-illness policy, sum insured five hundred thousand.” A line from a claims call, carrying both tone confusions from the pairs above.

Both sides are normalised before they are compared: punctuation dropped, Chinese numerals written as digits. 五十萬 becomes 500000.

S
a different character came back
D
a character was not returned
I
a character nobody said
2 substitutions + 1 deletion + 1 insertion17 reference characters
=23.5%CER

Three of the four change what the call is about. 買 buy came back as 賣 sell. 保險 the policy came back as 報險 reporting a claim against it. And the missing character is a missing zero — five hundred thousand of cover, read back as fifty. The rate counts all four the same, which is why the recognition pillar is not the whole score.

What a scored run is.

No files are uploaded and nothing is scored on request. We connect to a server the entrant runs, stream audio at it, and measure what comes back on our own clock.

  1. 01Register an endpoint

    A WebSocket server you host. We are the client, so nothing of yours has to be handed over and nothing of ours has to be trusted.

  2. 02Practise against the sample set

    8 clips with published ground truth, downloadable once you have an account. Unlimited runs, full per-clip diffs, no effect on your score.

  3. 03Book a scored window

    30 clips drawn at random from a pool of 56, stratified across difficulty tiers, seeded and recorded so the draw can be reconstructed.

  4. 04We measure, and attribute every failure

    Each clip resolves to a verdict and an owner: yours, ours, the network, or nobody. A run we break costs you nothing and does not consume quota.

  5. 05Scores publish with their uncertainty

    A bootstrap interval accompanies every number, and entrants whose intervals overlap are published as tied rather than ranked against noise.

What we will not publish.

A benchmark is only worth reading if it is willing to say less than it could. These are the four cases where we withhold a number rather than produce one.

An order inside the noise

The widest interval so far is ±4.47 points. Gaps smaller than that are reported as a shared band.

A score built on too little

A run that completed too few clips is published as insufficient data. Turning a broken integration into a bad score would misattribute an operations problem as model quality.

A zero for not competing

Pillars are opt-in. A team that did not enter speech synthesis is shown as not entered, never as scoring nothing.

A judge we have not checked

The language-model judge runs 3 times across 2 Amazon Nova models and is correlated against human raters before its pillar counts at all.

Who runs this.

Organised in Hong Kong, on Hong Kong audio, judged by Cantonese speakers. The dataset is real contact-centre recordings, consented and masked, never scraped.

Programme
CyberportConvenor and host of the challenge programme.
Data partner
Prudential Hong KongSource of the contact-centre audio and its domain review.
Technical organiser
PubrioBuilds and operates the evaluation platform, publishes the results.

2026-1105 clips held: 56 in rotation, 30 for qualification, 11 held back for final verification, 8 to practise on.

  1. Note27 Jul, 05:14 HKTThe pool now accepts audio in the formats phone systems actually produce

    The first real recording handed to us was GSM 6.10, which no browser will decode.

  2. Note27 Jul, 05:14 HKTWithdrawing is now something you do yourself

    The rules promised it. Until now it required emailing someone, which is a different promise.

  3. Announcement27 Jul, 05:14 HKTA system description is now required before final scoring

    One to two pages on what was measured. The Full Score is withheld until it is in.

Taking part

Have it measured.

No fee, no exclusivity, and you keep everything. The core is recognition, understanding and speech synthesis; conversation is optional. You do not have to enter all three — the Core Score is rescaled across what you did enter, never counting a pillar you skipped as a zero. Requests are reviewed within one business day.