Cantonese Voice Benchmark 2026.1
Everything a team meets, from the public board to a published result, in the order you meet it. Every number on this page is read from the platform when you open it, so the dates, weights and quotas here are the ones in force.
One file: what is measured, both round dates, how a score is made, and the whole protocol with a server you can run. Opens anywhere, prints to PDF.
Fifteen screens of the live platform, photographed on 2 September 2026, for presenting rather than reading. A screen that has changed since is described above, not in the file.
Everything a stranger can see without signing in, which is the whole benchmark except the audio itself.
Production Cantonese voice systems, scored on real Hong Kong insurance calls rather than read speech. Four pillars, one score, and every number published with the uncertainty around it. There is no public sign-up — accounts are issued by the organisers, because an account is the only thing between a stranger and real customers’ recorded voices.
Look forThe four pillar cards. Each links to exactly how that pillar is scored.
Every published figure carries a confidence interval, and entries whose intervals overlap are reported as tied rather than ordered. At 30 clips a run, a fraction of a percent is noise, and a board that ranks noise is a board that misleads.
Look forCore against Full, and the difficulty filter: the same teams, three honest views.
Four questions, weighted. A system that transcribes a call perfectly and loses the policy number has not understood it; a flat contour turns “buy” into “sell”, which in an insurance call is not cosmetic; and conversation is the one pillar where a pause is the product, so latency is rewarded there and nowhere else. Enter as few or as many as you like.
| Pillar | Weight | Asks | Who runs it |
|---|---|---|---|
| Recognition ASR | 20% | Can you hear? | Every entrant |
| Understanding LLM | 25% | Can you answer? | Every entrant |
| Speech TTS | 30% | Can you speak? | Every entrant |
| Conversation E2E | 25% | Can you converse? | Optional |
Look forThe weight, and whether the pillar gates the result, stated before anything else.
The whole method is public: how clips are drawn, how the interval is computed, what is withheld and why. There is a scorer on the page you can type into, so you can produce the same figure we would. A benchmark nobody can check is a press release.
Look forThe worked example — this is what turns one run into an interval.
Ask for access and the organisers issue the account. That request is the review — there is no second one — and it is answered in one business day. The letter that arrives carries a link that sets your own password, once.
Look forOne account per person. A shared login makes the audit log useless.
From the first sign-in to a system we can reach. None of it costs a scored run.
One screen answering two questions: what do I do next, and where do I stand. The next step moves as you progress and is never blank — a team that has to work out its own next step is a team that files a support ticket.
Look forYour next step, at the top of the Now tab.
You run a WebSocket server; we connect to it. Register the URL, then run the connection test — three real clips, about ten seconds, and a readable diagnosis if anything fails. The starter kit is a working server in Python and Node, so the first green tick takes minutes rather than an afternoon.
Look forThe four endpoints and their state. One socket may answer all four.
8 recordings with their correct transcripts, laid out one line per speaking turn. Click a timestamp to jump there, paste your own attempt, and see the difference character by character. The recordings are real customer calls, so they never leave the platform — and none of this costs a scored run.
Look forPractice is unlimited. Only a scored run spends quota.
Invite colleagues and set what each can do: who can start a scored run, and who can only look. One account per person rather than a shared login, so the audit log names who did what.
Look forYour published identity. A codename is the default until you choose otherwise.
The one step between an account and a scored run.
Say which pillars you are entering and what you have built. It is approved the moment you send it — the organisers issued your account, and that was the review — so scored runs open on the pillars you named as soon as the form is in. A pillar you did not enter is not scored, and the board says “not entered” rather than showing a zero.
Look forThe pillars you enter. You can add one later; one with a scored run stays in.
One to two pages, answering six questions, written in the portal. Your Full Score is withheld until it is in; your Core Score is not affected. A leaderboard of anonymous numbers teaches nobody anything, and every hosted component has to be declared — a result nobody can attribute is not a result.
Look forThe six questions. They are the ones other entrants will read.
What a scored run costs, what it tells you, and when it becomes public.
3 scored runs per pillar in a round, against 30 clips drawn for that round. Practice stays unlimited beside it. A run our platform broke does not count against you and does not spend a run; a run you stop after the first clip has gone out does.
Look forHow many runs are left, beside the pillar you are about to run.
A finished run reports the error rate with the interval around it, the response times, and what the error was actually made of — mishearing, writing Simplified and answering in Mandarin are three different problems with three different fixes.
Look forThe breakdown under the headline number. That is the part you can act on.
Nobody can see the number but you until you submit it. Submitting adds these clips to your standing rather than replacing it — every submitted clip is scored together, so more runs narrow the interval instead of moving the number around. One-way: there is nothing to gain by cherry-picking a lucky run.
Look forWhat your standing would become, shown before you decide.
Your Core Score and each pillar, exactly as the public board has them. The figures move only when you submit. The qualification round closes 30 September 2026; the challenge round runs 12 October 2026 to 26 October 2026.
Look forThese are the published figures, and your results download as CSV.
Every way out of a problem, and how long each one takes. Use the fastest, not the politest.
Access requests answered in one business day, support the same, security reports the same day if exploitable, and disputes answered with the per-item record behind the number rather than an assurance.
Look forA dispute goes to a queue with a stated turnaround, not an inbox.
Quotas, deadlines, what ends an entry, and how to appeal — written before the round opened rather than after somebody complained. Hosted APIs are allowed and must be declared. Undisclosed human involvement in a scored run is the one thing that ends an entry.
Look forThe cut-off rule: a tie at the line advances the whole tied band.
An error message or an instruction that did not make sense to you is a bug in our writing rather than a question you should have known the answer to. Tell us and we will fix the page.