Cantonese Voice Benchmark粵語語音基準

Cantonese Voice Benchmark 2026.1 · Window 2

Cantonese voice, measured on real calls.

Production voice systems evaluated on Hong Kong insurance contact-centre audio across four pillars — hearing, understanding, speaking, and conversation.

Window
open
Clips drawn
30
Entrants
0
Updated

No results published yet

The first scored window has not closed. Reference baselines are shown below so the scale is meaningful the moment real entries land.

A
Open-source pipeline
76.2
11.1%
3.81
3.44
3.29
2.11sp95 3.88
B
Proprietary voice LLM
82.7
8.6%
4.14
3.62
4.01
0.88sp95 1.52
Reference baselines
AOpen-source pipeline76.2
BProprietary voice LLM82.7
Full leaderboardHow we measureRequest accessRulesFAQ
ASR
Can you hear?
Character error rate on tonal, code-switched, insurance-domain speech.
LLM
Can you understand?
Policy-grounded reasoning, hallucination avoidance, information extraction.
TTS
Can you speak?
The least-solved pillar. Six tones, domain terms, natural code-switching.
E2E
Can you converse?
Voice in, voice out. Latency and turn-taking under real conditions.
TiedEntries sharing a rule have overlapping 95% confidence intervals. Reported as tied, not ordered.
Not enteredThe team did not compete on that pillar. Never shown as a zero.
PendingAwaiting human evaluation. The Full Score is withheld until every entered pillar is final.
BatchResults uploaded rather than measured live. Core Score only.

One syllable can be six words.

Cantonese carries six lexical tones, writes no spaces between words, and switches into English mid-sentence without marking the boundary. A model tuned on Mandarin does not degrade gracefully here — it fails in ways that read as fluent.

maai5vsmaai6

buy · sell

Low rising against low level. In a sales call this inverts who is doing what.

bou2vsbou3

insure · report

High rising against mid level. 保險 is a policy; 報險 is filing a claim against one. Both are said constantly on the same call.

ji1vsji6

medical · two

High level against low level. A claims call is dense with both hospital words and read-aloud digits, so this one lands in the amount as easily as the diagnosis.

Hear all six tones on the recognition pillar — the contours are plotted and playable.

What a scored run is.

No files are uploaded and nothing is scored on request. We connect to a server the entrant runs, stream audio at it, and measure what comes back on our own clock.

  1. 01Register an endpoint

    A WebSocket server you host. We are the client, so nothing of yours has to be handed over and nothing of ours has to be trusted.

  2. 02Practise against the sample set

    12 clips with published ground truth, downloadable once you have an account. Unlimited runs, full per-clip diffs, no effect on your score.

  3. 03Book a scored window

    30 clips drawn at random from a pool of 55, stratified across difficulty tiers, seeded and recorded so the draw can be reconstructed.

  4. 04We measure, and attribute every failure

    Each clip resolves to a verdict and an owner: yours, ours, the network, or nobody. A run we break costs you nothing and does not consume quota.

  5. 05Scores publish with their uncertainty

    A bootstrap interval accompanies every number, and entrants whose intervals overlap are published as tied rather than ranked against noise.

What we will not publish.

A benchmark is only worth reading if it is willing to say less than it could. These are the four cases where we withhold a number rather than produce one.

An order inside the noise

Entrants whose confidence intervals overlap share a rank band and are reported as tied.

A score built on too little

A run that completed too few clips is published as insufficient data. Turning a broken integration into a bad score would misattribute an operations problem as model quality.

A zero for not competing

Pillars are opt-in. A team that did not enter speech synthesis is shown as not entered, never as scoring nothing.

A judge we have not checked

The language-model judge runs 3 times across 2 model families and is correlated against human raters before its pillar counts at all.

Who runs this.

Organised in Hong Kong, on Hong Kong audio, judged by Cantonese speakers. The dataset is real contact-centre recordings, consented and masked, never scraped.

Programme
CyberportConvenor and host of the challenge programme.
Data partner
Prudential Hong KongSource of the contact-centre audio and its domain review.
Technical organiser
PubrioBuilds and operates the evaluation platform, publishes the results.
Edition
2026-179 clips held, 55 in rotation, 12 held back for final verification.