CantoNotes
Blog

Cantonese transcription4 min read

Mixed Cantonese–English transcription: names, brands, and 中英夾雜

Hong Kong speech is not Mandarin with loanwords. A usable transcript keeps the English terms, the Cantonese grammar, and the names your CRM already uses.

Mixed Cantonese and English transcript with names and brands kept accurate

Listen to a real Hong Kong meeting for thirty seconds. You will hear a Cantonese frame, an English noun, a particle, a brand, a nickname, then a number said in two systems. This is not a bug in the speakers. It is how the city talks at work. A transcription that “fixes” it into one language is already editorializing.

Teams then waste the next hour correcting the same class of errors: a colleague’s English name spelled as a similar-sounding Chinese character, Slack written as a random transliteration, KPI expanded or dropped, 報價 turned into a generic “quote” in a sentence that still needed the Cantonese verb. The transcript becomes something you cannot search, because nobody will type the model’s guesses.

Keep the mix. Decide later whether minutes should be formal.

Transcript and minutes have different jobs. The transcript should follow the mouth. If someone said “我哋下個 sprint 先 lock scope,” the line should still contain sprint and lock and scope. The minutes can say “Scope to be locked next sprint” if you are writing for a US parent company. Those are two documents. Collapsing them in the recognizer is how you lose both the quote and the searchable term.

The same rule applies to 口語. Particles tell you whether a “yes” was reluctant. Do not strip them from the transcript so the page looks like written Chinese. You can tone down the minutes. You cannot reconstruct hesitation from a sanitized line.

Names should match the CRM, not the homophone

Hong Kong names live in two scripts at once. The person in the seat might be 陳嘉儀 on the ID and Karen Chan in email. The client might be a legal Chinese name and an English trading name. If the transcript picks 陳 when every Slack mention is Chan, later search and later AI answers will miss the thread.

  • Prefer the form your team already files: email display name, CRM, or the Zoom participant list.
  • Do not “helpfully” convert an English given name into Chinese characters unless that is how they write it.
  • Keep titles and romanizations stable across the meeting. One spelling per person per workspace is the goal.

Brands and shop terms stay in the language they were said

Product names, law firm names, mall names, and internal code names are where generic speech-to-text panics. It will pick a common character that sounds close. Your invoice will not. Lock a small list of terms you care about — client names, product lines, the word you always say in English (onboarding, SLA, retainer) — and judge a transcript by those first, not by overall “accuracy %.”

中英夾雜 is consistent inside a company even when it looks messy from outside. If your team always says “meeting” and never 會議, the transcript that writes 會議 is not more professional. It is less searchable. Agree the rule per workspace: keep the English token, or always use the Chinese token. Then stick to it.

The first two minutes of QA that actually matter

  1. Scan every proper noun in the first page. If a name is wrong once, it is wrong everywhere.
  2. Find one mixed sentence you remember from the call. If the English island survived, the rest of the file is probably usable.
  3. Check a number: a price, a date, a headcount. Speech models love confident wrong digits.
  4. Only then skim for meaning. Fixing tone before names is backwards.
Accuracy you cannot search is decoration. Accuracy on names and terms is the product.

When the same term is always wrong

Some errors are not random. A client name, an internal product, a mall, a law firm — if it is wrong in meeting one, it will be wrong in meeting twelve. Correct it once in context and keep a short list of house terms. Judging a vendor by a single demo file that does not contain your names is how you buy a toy that fails on Monday.

Do not wait for a mythical 99% score. Wait for the names your invoices already use to survive a noisy, mixed-language standup.

Where CantoNotes is honest about the problem

CantoNotes uses CantoSub speech AI because that stack was already built for Cantonese, English, and Mandarin in the same recording — including the ugly, normal 中英夾雜 of a Hong Kong office. The meeting workspace then keeps the transcript next to notes and to-dos, so you are not copying a “cleaned” paragraph into a doc that nobody can grep.

It will still mishear a rare brand the first time. That is speech. The difference is the default: keep the mix, label the speaker, let you correct a name in context instead of shipping a Mandarin-shaped summary and calling it localization. If you need subtitles for a cut, that is a different export (SRT on paid plans). For the meeting record, readable mixed-language text is the feature.

Related

Stop writing minutes at 11 p.m.

Sign up