LLM · 25% of the score · every entrant
Given the call, does the system grasp what the caller needs, ground its answer in the actual policy, and refuse to invent coverage that does not exist?
One of them costs the insurer a claim. Hallucinated coverage does not look like an error in a transcript — it looks like good service, which is exactly why it has to be scored separately from whether the words were transcribed correctly.
我做physiotherapy,一個療程十次,一次八百蚊,我份plan cover唔cover到?
I need physiotherapy — ten sessions at $800 each. Does my plan cover it?
你張保單有物理治療嘅保障,不過我喺度睇唔到每次嘅上限同埋年度限額,要幫你查返份保單條款先可以講到準確數字。要唔要我而家幫你開個查詢,跟進返俾你?
Says what the policy shows, names what it cannot see, and offers the next step. It answers less than the caller wanted, which is the correct amount.
有嘅,你份plan完全cover到物理治療,每次八百蚊都喺保障範圍之內,一年最多可以做二十次,你直接去做就得,唔使預先approve。
Every marked clause is a fact the policy documents do not contain: the coverage, the per-session limit, the annual cap, and the claim that no pre-approval is needed. The caller books ten sessions on the strength of it.
One reply to the caller, in Cantonese — not a transcript, not a summary of the call, not a script for somebody else to read. The same sentence is in every hello frame as `task`, so your server can read it rather than remember it:
Answer the caller. In Cantonese, say what they need and what was agreed, grounded only in what the call actually contains. Do not repeat the call back — a transcript is not an answer — and do not invent coverage, amounts or policy terms that were not said.
Structured rubrics over instruction following, Cantonese localisation, hallucination and policy accuracy, each scored 1 to 5 against published anchors. Two Amazon Nova models — Pro and Lite — score every answer three times each, at temperature 0. The median is taken, the spread between runs is kept, and an item where the two models differ by more than one rubric point is marked contested rather than averaged into agreement. Before any model is called, a deterministic check floors a reply that reproduces the call: that is a string question with a certain answer, and on short calls a model gets it wrong. Both models come from one vendor, so this is independence from one model’s quirks and not from one vendor’s; two providers is still the goal and this page will say so until it is true. Choosing them was a measurement, not a preference: across five random production clips a genuine answer ranked above a copy-paste, an echo and a non-answer, and a frontier pair cost twenty times more for the same work. A deployment without model credentials scores no reasoning runs at all — it never falls back to the word-overlap rubric this replaced, which ranked reproducing the call above answering it.
Did the system capture the facts a downstream process needs: policy numbers, claim amounts, dates, medical terms, and what was actually decided.