LLM · 25% of the score · every entrant
Given the call, does the system grasp what the caller needs, ground its answer in the actual policy, and refuse to invent coverage that does not exist?
One of them costs the insurer a claim. Hallucinated coverage does not look like an error in a transcript — it looks like good service, which is exactly why it has to be scored separately from whether the words were transcribed correctly.
我做physiotherapy,一個療程十次,一次八百蚊,我份plan cover唔cover到?
I need physiotherapy — ten sessions at $800 each. Does my plan cover it?
你張保單有物理治療嘅保障,不過我喺度睇唔到每次嘅上限同埋年度限額,要幫你查返份保單條款先可以講到準確數字。要唔要我而家幫你開個查詢,跟進返俾你?
Says what the policy shows, names what it cannot see, and offers the next step. It answers less than the caller wanted, which is the correct amount.
有嘅,你份plan完全cover到物理治療,每次八百蚊都喺保障範圍之內,一年最多可以做二十次,你直接去做就得,唔使預先approve。
Every marked clause is a fact the policy documents do not contain: the coverage, the per-session limit, the annual cap, and the claim that no pre-approval is needed. The caller books ten sessions on the strength of it.
Structured rubrics over instruction following, Cantonese localisation, hallucination, and policy accuracy. Three runs, median taken. Judged by two model families from different providers, because a judge scoring its own vendor is a conflict.
Did the system capture the facts a downstream process needs: policy numbers, claim amounts, dates, medical terms, and what was actually decided.