Cantonese Voice Benchmark粵語語音基準

LLM · 25% of the score · every entrant

Can you understand?

Given the call, does the system grasp what the caller needs, ground its answer in the actual policy, and refuse to invent coverage that does not exist?

Both of these are fluent.

One of them costs the insurer a claim. Hallucinated coverage does not look like an error in a transcript — it looks like good service, which is exactly why it has to be scored separately from whether the words were transcribed correctly.

The caller asks

我做physiotherapy,一個療程十次,一次八百蚊,我份plan cover唔cover到?

I need physiotherapy — ten sessions at $800 each. Does my plan cover it?

Grounded

你張保單有物理治療嘅保障,不過我喺度睇唔到每次嘅上限同埋年度限額,要幫你查返份保單條款先可以講到準確數字。要唔要我而家幫你開個查詢,跟進返俾你?

Says what the policy shows, names what it cannot see, and offers the next step. It answers less than the caller wanted, which is the correct amount.

Invented

有嘅,你份plan完全cover到物理治療,每次八百蚊都喺保障範圍之內,一年最多可以做二十次,你直接去做就得,唔使預先approve。

Every marked clause is a fact the policy documents do not contain: the coverage, the per-session limit, the annual cap, and the claim that no pre-approval is needed. The caller books ten sessions on the strength of it.

What the judge scores, per answer
Instruction following
Did it address what the caller actually asked, and complete the task?
Cantonese localisation
Natural Hong Kong Cantonese, or written Chinese read aloud?
Hallucination
Did it invent coverage, amounts, or procedures that do not exist?
Policy accuracy
Where it cited the policy, was the citation correct?

What to send back

One reply to the caller, in Cantonese — not a transcript, not a summary of the call, not a script for somebody else to read. The same sentence is in every hello frame as `task`, so your server can read it rather than remember it:

Answer the caller. In Cantonese, say what they need and what was agreed, grounded only in what the call actually contains. Do not repeat the call back — a transcript is not an answer — and do not invent coverage, amounts or policy terms that were not said.

What is measured

How it is judged

Structured rubrics over instruction following, Cantonese localisation, hallucination and policy accuracy, each scored 1 to 5 against published anchors. Two Amazon Nova models — Pro and Lite — score every answer three times each, at temperature 0. The median is taken, the spread between runs is kept, and an item where the two models differ by more than one rubric point is marked contested rather than averaged into agreement. Before any model is called, a deterministic check floors a reply that reproduces the call: that is a string question with a certain answer, and on short calls a model gets it wrong. Both models come from one vendor, so this is independence from one model’s quirks and not from one vendor’s; two providers is still the goal and this page will say so until it is true. Choosing them was a measurement, not a preference: across five random production clips a genuine answer ranked above a copy-paste, an echo and a non-answer, and a frontier pair cost twenty times more for the same work. A deployment without model credentials scores no reasoning runs at all — it never falls back to the word-overlap rubric this replaced, which ranked reproducing the call above answering it.

Information extraction

Did the system capture the facts a downstream process needs: policy numbers, claim amounts, dates, medical terms, and what was actually decided.

Why it is hard here

Hallucinated coverage
The single most expensive failure. A confident wrong answer about what is covered is worse than no answer.
Cantonese register
Written Chinese and spoken Cantonese diverge sharply. A reply that reads correctly can still sound wrong.
Grounding
The answer has to come from the policy documents, not from what the model expects an insurer to say.