Tutorials

Claude Opus 5.5: build a pronunciation sheet for product voiceovers

Use Claude Opus 5.5 to separate the written brand name, the approved spoken form, and the voice engine's supported pronunciation controls.

A steaming ivory coffee cup and saucer on a walnut table, with a microphone and headphones for recording product sounds.

Use Claude Opus 5.5 to separate the written brand name, the approved spoken form, and the voice engine's supported pronunciation controls. Keep the product label intact. Changing every occurrence of a name to a phonetic spelling can repair the audio while damaging captions, product search, and campaign copy.

A multilingual TTS discussion describes inconsistent brand pronunciation across languages. It is a user's report, not a test of Opus. The workflow below is an editorial proposal; no generated voice or model performance is claimed.

Decide what kind of mistake you heard

SymptomNeeded evidenceNext step
A model code becomes a numberApproved reading of each tokenSpecify letters, digits, and separators
Brand stress is wrongOwner-approved reference recordingAdd a locale-specific pronunciation entry
Only one voice failsVoice ID, locale, engine versionTest that exact voice separately
Caption spelling changesOriginal display textKeep display and speech fields separate

Google's SSML reference documents pronunciation markup. That does not establish support in every voice or service. Check the selected engine and voice before producing markup; plain phonetic text is also a candidate to audition, not a guaranteed fix.

Build one traceable entry

Imagine a fictional product called `Navo Q7`. Its owner has approved “NAH-voh cue seven” for one English voice. This is an invented naming decision, not the pronunciation of a real brand. Store the display string `Navo Q7`, locale `en-US`, approved audio reference, intended reading, engine, voice ID, and review status in separate fields. Do not reuse the English reading automatically for Japanese or Arabic.

Test three sentences: the name alone, the name at the beginning of the actual ad, and the name near the end of the actual ad. Context can change delivery. If the sample set contains 4 product names, 6 locales, and 3 contexts, it requires 72 clips for one voice per locale. Two voices per locale require 144 clips. These are planning counts, not measured failure rates.

Copyable review prompt

Use the official Opus guide with this task brief:

Create a pronunciation review sheet from the supplied product list, approved reference recordings or transcripts, locale list, and target voice documentation.
Preserve every display name and SKU exactly. Separate written text, intended spoken form, and engine-specific input. For each locale, identify ambiguous letters, digits, abbreviations, stress, and pauses. Never invent an approved pronunciation when no reference exists; mark it for owner review.
Generate three short test sentences per entry: isolated name, opening context, closing context. Propose SSML only where the exact voice documentation supports it; otherwise propose a plain-text audition candidate. Do not change captions or on-screen brand spelling.
Return a table with entry ID, locale, display text, reference, proposed input, voice ID, and pending/pass/fail review status. Leave listening results pending until actual clips are supplied. Flag unmatched references and repeated entries.

Listen before replacing the campaign audio

Render the small test set in the actual voice tool. Listen to pronunciation, stress, pauses, and whether the surrounding sentence still sounds natural. Have an appropriate language reviewer compare the result with the approved reference. Record the accepted clip ID and voice settings, not just “sounds good.” If a voice changes, reopen the affected entries.

Then render the full ad. Check the final product name, the caption spelling, and the cut timing together. A longer pronunciation can push a line past its visual moment; edit the script or timing deliberately instead of speeding up the whole narration.

Practical questions

**Can Opus certify pronunciation from spelling?** Spelling alone cannot establish the owner's intended reading.

**Should phonetic spellings appear on screen?** Keep the approved display name unless the campaign explicitly calls for a pronunciation explanation.

**Does passing one sentence finish the review?** No. Audition the final sentence and final voice settings.

Use the approved display names and visual references in Panelly Studio when preparing the accompanying campaign images. If narration length affects on-screen text, the subtitle timing workflow helps review that separate delivery step.

Sources and further reading

Four panels. One ad.

Your next ad starts here.

Describe your product, its benefits and its audience. Create a four-panel ad with Panelly.

Start creating ↗View credit packs