Cantonese Voice Benchmark粵語語音基準

TTS · 30% of the score · optional

Can you speak?

The least-solved piece of the stack. No dominant Cantonese voice exists, and the reasons are structural rather than a matter of more training data.

Six contours, produced.

Recognition has to tell 6 tones apart. Synthesis has to make them distinguishable in the first place, and the ways that fails do not show up as a bad word — they show up as a voice a Hong Kong listener does not trust. These are the four we hear most.

  1. Tone flattening

    市 si5 → 是 si6

    Low rising collapses into low level under fast speech or at the end of a phrase.

    The listener hears a different word and repairs it from context — until the word is a policy number, where there is no context to repair from.

  2. Reading written Chinese aloud

    沒有 → 冇

    Output is grammatical Standard Written Chinese pronounced in Cantonese, rather than spoken Cantonese.

    Every sentence is correct and nobody talks like that. It is the single most common way a synthesised Hong Kong voice announces itself as one.

  3. Code-switch prosody

    你個 claim 已經 approve 咗

    The English words are produced with English stress and timing, so the sentence audibly changes speaker mid-phrase.

    Hong Kong speakers switch constantly and smoothly. A seam here is more noticeable than a wrong word.

  4. Digit strings

    HK 七三二一四五

    Numbers read with uniform timing and no grouping, or with the wrong tone on 一 and 七.

    A policy number that has to be said twice is a call that takes twice as long, and the caller writes it down wrong the first time.

None of this is captured by asking “does it sound natural?” alone, which is why naturalness and intelligibility are rated separately and why a tonal check runs alongside both. A warm voice that flattens tone five is worse on a claims line than a robotic one that does not.

What is measured

MOS

Naturalness and intelligibility, rated separately by native listeners. They diverge: a flat voice can be perfectly clear, and a warm one can still mangle a policy number.

Consistency

Does the voice hold its identity across a whole call, or drift between turns.

Loop-back CER

The generated audio is fed back through a reference recogniser. If a machine cannot read it back, a person is guessing.

Domain pronunciation

Binary, expert-checked: policy numbers, medical terms, and English words embedded in Cantonese sentences.

Why it is hard here

Tone production
Getting six tones right, every time, is harder to generate than to recognise.
Code-switch seams
Most systems audibly change accent at the boundary into English and back.
Conversational markers
Real Hong Kong service speech is full of 係喂, 嗯, 唔該. Without them the voice reads as a recording.