Data card
Real Hong Kong insurance calls, not read speech. This page describes the dataset fully; the audio itself sits behind an account, and the reason is on this page too.
| Split | Clips | Used for | Access |
|---|---|---|---|
| Sample | 8 | Integration and practice | Signed in, ground truth shown |
| Rotation pool | 56 | Scored windows, drawn per run | Streamed during a window only |
| Holdout | 11 | Final verification | Never seen until the last run |
Difficulty is tiered easy, normal and hard, and every draw holds the same proportional mix. A uniform random draw would hand one run fifteen hard clips and the next run five, and those two scores would not be comparable to each other or to anyone else’s.
Every clip is a recorded call, not read speech: 105 customer service.
Thirty seconds to three minutes of Hong Kong Cantonese. Dense English code-switching, insurance and medical vocabulary, policy numbers and claim amounts read aloud, background noise, overlapping speakers, and callers who are frequently distressed. Reference transcripts are delivered with the audio and screened before any of them scores anybody. Every reference is folded to Hong Kong Traditional before scoring, so our own orthography never costs you — your transcript is not folded, so a system that emits Simplified is still charged for it and the run page shows how much that cost. Every clip is checked against its own recording: calls that turn out to be Putonghua leave every draw, a reference whose turns do not run in the order the audio does is withheld until somebody has listened to it, and so is any clip the checks cannot judge either way. The key facts of each call are recorded separately so information extraction can be scored.
Removing names and policy numbers from a recording does not de-identify it. A voice is a biometric identifier, and under Hong Kong’s Personal Data Privacy Ordinance a voiceprint is personal data about the speaker regardless of what words were said. An anonymous, permanent, public download of real customers’ voices is a materially different act from a controlled evaluation.
So every real clip goes to an identified recipient who has accepted terms. Requests are reviewed within one business day, and approval is a formality for anyone with a genuine system. If we publish a permanently open artifact for citation, it will be recorded by actors rather than drawn from customer calls.
Stating these is part of the data card, not a caveat buried at the end. The corpus is Hong Kong Cantonese only and does not represent Guangdong or overseas varieties. It is drawn from one insurer, so vocabulary skews to that book of business. Simulated clips are acted, and actors under-produce the disfluency of a genuinely distressed caller. Speaker demographics follow whoever called, which is not a balanced sample.
The pool is also smaller than it will be. At 56 clips in rotation and a draw of 30, 100% of it. Growing the pool between windows is what keeps a score a measurement rather than a memory test.
Every one of the 105 clips a run can draw carries a reference written by a person transcribing the call, and confirmed here against its own audio. The checks that produced that confirmation also found and corrected specific faults along the way — Simplified characters, and calls held in Putonghua rather than Cantonese.