Version 1.0 · in force for edition Cantonese Voice Benchmark 2026.1
What you agree to by entering, what we agree to in return, and what happens when something goes wrong. Written to be read once, not skimmed forever.
Any company or research group with a working Cantonese speech system. There is no fee and no exclusivity. A rank needs the three core pillars — recognition, understanding and speech (ASR + LLM + TTS) — which is what the Core Score is. Conversation is optional, it sits outside the Core Score, and the organisers run it rather than the entrant. Enter fewer pillars and each one is still published on its own, outside the ranking. Accounts are issued by the organisers; with one, you enter by sending an application from the portal that names your pillars, and it is approved the moment you send it.
You are identified on the leaderboard by a codename by default. Being named is opt-in, you choose it yourself, and you can change your mind up until results are published. We will never name you in the report without written consent. This is stated here rather than buried in terms because it is the first thing most entrants want to know.
Every published number carries a bootstrap confidence interval, and entries whose intervals overlap are reported as tied rather than ordered. The normalization applied before scoring is published as running code, not prose, so you can compute the same number we do. Full methodology.
A pillar you did not enter shows as “not entered”, never as a zero. A run where our platform failed does not count against you and does not consume your quota.
A scored run is a draft until you submit it. Nobody but you can see it, it is not ranked, and it is never submitted for you — press Submit on the run’s own page, before the round closes. The board publishes the latest run you submitted for each pillar, so a worse run you leave unsubmitted costs you nothing.
A programme runs in rounds. A qualification round decides who continues; a challenge round decides the published standing. Which rounds run, how long each lasts and how many scored runs it allows are set by the organisers and shown on your portal while the round is open.
The cut-off is decided by machine, on the same score the round publishes, over the pillars that round marks as required. Nobody is advanced or eliminated by opinion, and the round does not rank you on a pillar it did not ask you for.
A tie is not split. The board refuses to order entries whose confidence intervals overlap, and the cut-off honours that: if the line falls inside a tied band, the whole band goes through — so "the top four advance" sometimes advances five. Eliminating a team on a difference this benchmark has said it cannot measure would be the worse answer. A team that ran but could not be scored is not treated as a team that came last; an organiser is told about those separately and decides them by hand.
We do not claim the audio cannot be captured. It is streamed to your process, so of course it can. We claim something narrower and true: the response window precludes manual transcription, every dataset pack is issued to a named recipient and recorded, and the final set is fresh. We do not watermark the audio and do not claim to: attribution rests on the issue record, which names a person and is therefore enforceable.
The evaluation audio is real Hong Kong contact-centre material with personal information removed. You may use it to produce your submissions during the edition. You may not redistribute it, publish it, or train on it. Access is logged. All copies you hold are to be deleted within thirty days of the edition closing.
Scores are append-only: a correction is published as a new value with a note, never as a silent edit. If you believe a number is wrong, open a dispute and we will show you the per-item record behind it, including which failures we attributed to ourselves. Results enter a review window before publication so disagreements are settled privately first.
You can withdraw at any time before publication — from your portal, without asking anyone. Your scores are voided and you disappear from the leaderboard immediately. The run records are kept, so a dispute already raised about a run can still be answered. After publication, an anonymous entry stays anonymous, and a named entry can be un-named on request.