Cantonese Voice Benchmark粵語語音基準
EN繁中

Personal Data (Privacy) Ordinance · Hong Kong

What we hold, and why.

This benchmark handles two kinds of personal data: the voices of people who called an insurance line, and the details of the people who compete and rate. They carry different risks and different rules, so they are described separately.

呢一頁暫時只有英文版。翻譯緊,會逐頁推出——首頁、排行榜同評測方法已經有中文。

Whose data this is

Pubrio Limited operates this platform as data processor for the challenge. Prudential Hong Kong is the data user for the source recordings and obtained the consent under which they were collected. Cyberport convenes the programme and receives results, not recordings.

Questions about the recordings go to the data user. Questions about this platform, your account, or a request to see what we hold go to [email protected], and we answer within 40 days as the Ordinance requires.

The recordings

A voice is a biometric identifier. Under the Ordinance a voiceprint is personal data about the speaker whatever words were said, so removing names from a recording does not de-identify it — which is the entire reason the dataset sits behind an account rather than a public download link.

Before any clip reaches this platform, names, policy and identity numbers and phone numbers are removed from both audio and transcript, and native Cantonese reviewers confirm the masking is complete. Raw recordings are deleted once that check passes. We never receive them.

Clips are streamed to an entrant’s system during an evaluation and are logged against the account that received them. Entrants agree in the rules not to retain, redistribute, or train on the audio, and the sample pack names its recipient inside the archive so a leaked copy is traceable.

Account and rater data

For entrants: the name, work email and organisation given on the access request, the endpoint address you register, and a log of your runs. Passwords are stored as a hash and are never recoverable — an account has no password at all until you set one.

For raters: the same identity fields, plus your ratings, listening durations, and calibration results. Those exist to detect a rater clicking through without listening, which is a quality control on the benchmark, not a performance review of a person. They are not shared with entrants and never leave aggregate form in anything published.

Who sees what

The public
Scores, confidence intervals and pillar coverage, under a codename unless you have opted into being named. Never audio, never per-clip output, never rater identities.
Entrants
Their own runs in full, including per-clip diagnostics on practice sets. Nothing about any other entrant beyond what is published.
Raters
Audio to rate, with system identity hidden. A rater is never told whose speech they are scoring.
Organisers
Everything, with every privileged action written to an audit log that records who did what and when.
Sub-processors
Render, for hosting in Singapore. No analytics, no advertising, no third-party trackers on any page of this site.

How long we keep it

Evaluation copies of the audio are deleted thirty days after an edition closes and a deletion certificate is issued to the data user. Scores, intervals and the audit log are kept indefinitely — a published benchmark that cannot show its own working later is not a record of anything.

Account records are kept while the account is active. Disabling an account keeps the row so past runs and ratings stay attributable; deletion of the identifying fields is available on request once the edition it belongs to has closed.

Your rights

You may ask what personal data we hold about you and ask us to correct it. Write to [email protected]. We do not charge for a first request and we answer within 40 days.

If you are a caller whose recording may be in the corpus, the data user holds the consent record and the ability to identify a clip as yours — we deliberately cannot. Contact Prudential Hong Kong’s data protection officer, and we will act on any instruction they pass to us.

Terms · Data card · Rules