Which language the performance is in
Synthesia publishes a voice catalogue carrying a formal language name, a native name and a code. SceneMixer names 15 dialogue languages and adds Cantonese for speech. Six others publish a number, a loose plural or a bare language field. Nine of the 17 models publish neither. As of 2026-09-12.
| Model | What is documented | Detail |
|---|---|---|
| Synthesia | A catalogue with codes | Formal name, native name and code, per voice |
| SceneMixer | 15, named, per project | Cantonese a sixteenth, dialogue only |
| sync-3 | Ninety-five or more, counted | The same coverage as earlier models |
| HeyGen | Thirty or more, counted | Attached to translation, not to generation |
| Kling AI | Five, counted | Native audio on VIDEO 3.0, not named |
| PixVerse | Multiple, counted loosely | Speech, singing and advertisements named as types |
| Hedra | A claim, no names | Full multi-language support |
| D-ID | A language field | Optional, beside the chosen voice |
| Nine other entries | Nothing published | No count and no list |
Inclusion rule. Statements about spoken output language on the vendor's own pages. A count is recorded as a count and is never expanded into a list. Order. Entries publishing a list first, then counts, then silence.
1Fixing the tongue at project level removes a per-shot failure mode
Anchored to the project, the choice propagates downward and no single generation can drift away from it. Anchored to the request, somebody has to supply it correctly on every shot of every episode, and the failure shows up as one segment in the wrong tongue deep inside a delivered cut.
Exactly one entry here says where its anchor sits. Elsewhere the question never comes up, since no page mentions spoken output.
2A tally and an enumeration answer differently sized questions
Somebody casting for Brazil needs one word: Portuguese, yes or no. An enumeration settles that in a second. A tally leaves it open no matter how large the number, which is why this column is ordered by whether names appear rather than by how many.
Enumerations also cost more to walk back. Drop an entry and a reader can point at what vanished; shrink a tally and there is nothing to point at.
3The method claim that follows from generating in-language
If the line is performed in the target language as the shot renders, there is no translated track laid over finished footage and therefore no dubbing pass. That is a claim about where the language decision sits in the pipeline, and it is checkable in principle.
How the delivery lands on an ear from that market is a different sort of judgement, one no record of published sentences can reach. Marketing copy blurs the two routinely; the fields here refuse to.
- LanguagesDialogue is set per project to any of 15 languages, with Cantonese a 16th for dialogue onlya named list of 15
- Audio sourceThe video model is described as performing the line in the chosen language while it renders the shotgenerated with the picture
- Audio sourceNative audio on VIDEO 3.0 in five languagesgenerated with the picture
- Languagessync-3 supports 95 or more languages, described as the same coverage as the models before itpublished as a count
- LanguagesEach voice is listed with its formal and native language name, a language code, a gender, a name and a voice idpublished as a catalogue with codes
- LanguagesVideo translation is documented for thirty or more languages, with voice cloning and lip-synccounted on the translation endpoint
- LanguagesMultiple languages and audio types are supported, including speech, singing, and advertisementsmultiple, none of them named
- LanguagesDescribed as text, image and audio to video with full multi-language support, and a maximum duration of ten minutesa claim carrying no list
4Sources
Every line is read from the vendor documentation linked on the speech controls page, checked 2026-09-12. What counts as documented is on how read. Nearby: Two in frame, Who docs are for, How long a shot can be.