Cantonese Voice Benchmark粵語語音基準
EN繁中

Data card

What the audio is.

Real Hong Kong insurance calls, not read speech. This page describes the dataset fully; the audio itself sits behind an account, and the reason is on this page too.

呢一頁暫時只有英文版。翻譯緊,會逐頁推出——首頁、排行榜同評測方法已經有中文。

Composition

SplitClipsUsed forAccess
Sample12Integration and practiceSigned in, ground truth shown
Rotation pool55Scored windows, drawn per runStreamed during a window only
Holdout12Final verificationNever seen until the last run

Difficulty is tiered easy, normal and hard, and every draw holds the same proportional mix. A uniform random draw would hand one run fifteen hard clips and the next run five, and those two scores would not be comparable to each other or to anyone else’s.

What is in a clip

Thirty seconds to three minutes of Hong Kong Cantonese across customer service, sales and internal meetings. Dense English code-switching, insurance and medical vocabulary, policy numbers and claim amounts read aloud, background noise, overlapping speakers, and callers who are frequently distressed. Ground truth is professionally transcribed and reviewed by domain experts, with the key facts of each call recorded separately so information extraction can be scored.

Why it is behind a sign-in

Removing names and policy numbers from a recording does not de-identify it. A voice is a biometric identifier, and under Hong Kong’s Personal Data Privacy Ordinance a voiceprint is personal data about the speaker regardless of what words were said. An anonymous, permanent, public download of real customers’ voices is a materially different act from a controlled evaluation.

So every real clip goes to an identified recipient who has accepted terms. Requests are reviewed within one business day, and approval is a formality for anyone with a genuine system. If we publish a permanently open artifact for citation, it will be recorded by actors rather than drawn from customer calls.

Processing

Masking
Names, policy and ID numbers, and phone numbers are removed from both audio and transcript before anything enters the evaluation platform.
Verification
Native Cantonese reviewers confirm masking is complete and that audio quality survived it. Domain experts spot-check a sample.
Simulation
Where real recordings do not cover an edge case, native speakers record additional calls from scripted scenarios. These are labelled in the dataset and scored identically.
Retention
Raw recordings are deleted once masking is verified. Evaluation copies are deleted thirty days after the edition closes, and a deletion certificate is issued.

Known limitations

Stating these is part of the data card, not a caveat buried at the end. The corpus is Hong Kong Cantonese only and does not represent Guangdong or overseas varieties. It is drawn from one insurer, so vocabulary skews to that book of business. Simulated clips are acted, and actors under-produce the disfluency of a genuinely distressed caller. Speaker demographics follow whoever called, which is not a balanced sample.

The pool is also smaller than it will be. At 55 clips in rotation and a draw of 30, an entrant using ten scored runs has seen 100% of it. Growing the pool between windows is what keeps a score a measurement rather than a memory test.

Request access · How it is scored · Rules