What Synthesia documents about speech
Synthesia (docs.synthesia.io) publishes a voice catalogue in which every entry carries a language name, a language code, a gender and a voice id, and states that lip sync and facial expressions are generated from the spoken content. As of 2026-09-22.
| Field | What the vendor documents |
|---|---|
| Audio source | Synthesised from the script, or an uploaded recording |
| Languages | A catalogue, each voice carrying a language name and code |
| Voice source | The catalogue, or a cloned voice with a language list of its own |
| Per-character binding | Not documented |
| Lip-sync | Generated from the spoken content, with a framing condition attached |
Inclusion rule. Fields are the same five for every model here; a field the vendor does not document is recorded as not documented rather than inferred from generated output. Order. Fixed field order, identical on every model page.
1A language list a schedule can be built on
Instead of a number in a paragraph, this vendor publishes rows: the formal language name, the native name, a language code, a gender, a voice name and an id. A planner can compare that list between quarters, pick a market off it, and pass the id straight into a request.
The register records it as a catalogue rather than a count, because the difference is operational. A count tells a planner the capability exists somewhere; a list with codes tells them whether the market they were asked about is inside it.
2Lip sync is tied to the script, and comes with a condition
Lip sync and facial expressions are generated from the spoken content, which places this entry with the models that derive the mouth from the words rather than repairing it afterwards. Then the documentation adds something almost unheard of here: lip sync performs best when the avatar is framed closer in the scene rather than positioned far from the camera.
That sentence is a shot-list instruction wearing the clothes of a help note. It says the weak case is the wide shot, which is precisely the shot a director reaches for when two people are in a scene together.
3A cloned voice has its own reach
Voice cloning carries a language list of its own, running from Afrikaans to Zulu and described as the same list used for standalone cloning. So the question is not only which languages the platform speaks but which of them a particular voice can speak, and the two lists are kept separately.
What is not published is how a voice stays attached to a character across a run of videos. An id can be reused; nothing states that the pairing is held anywhere except in a production's own notes.
4This entry, one column at a time
Each of these stays inside a single field: what Synthesia puts there, what the wording settles, and what it leaves for a take to answer.
- Audio source — from the script, or uploaded.
- Voice source — a catalogue voice, or a cloned one.
- Per-character binding — not documented by the vendor.
- Languages — a catalogue with codes.
- Lip-sync — the spoken content, framed close.
5Read against another entry
Each of these puts Synthesia beside one other entry on a column where the two land at opposite grades of answer.
- Synthesia and sync-3 — on languages.
- Synthesia and D-ID — on lip-sync.
- HeyGen and Synthesia — on languages.
- Kling AI and Synthesia — on lip-sync.
- SceneMixer and Synthesia — on languages.
- Hedra and Synthesia — on lip-sync.
6Where this model sits against the rest
| Model | Audio source |
|---|---|
| D-ID | A text script read aloud, or an audio url supplied |
| Hedra | Supplied; audio is a required input |
| HeyGen | A text to speech endpoint, apart from the picture |
| Kling AI | Generated with the picture, VIDEO 3.0 |
| LTX Studio | Generated jointly with the picture, plus audio-to-video |
| Luma Ray | Not documented; the video reference does not mention sound |
| MiniMax | Native speech, generated with the picture |
| PixVerse | Supplied to a speech endpoint, or read from text there |
| Runway | Not documented where voices are defined |
| SceneMixer | Generated with the picture, in the language set for the project |
| Sora 2 | Generated with the picture; video and audio are both listed as output |
| sync-3 | Supplied, or read from text on the call |
| Synthesia | Synthesised from the script, or an uploaded recording |
| Veo | Generated with the video rather than added afterwards |
| Vidu | Q3 generates speech, effects and music natively, on by default |
| Wan 3.0 | On by default; false returns a file with no audio track |
| Wan2.2-S2V | Supplied; the card describes the model as audio-driven |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Languages |
|---|---|
| D-ID | A language field beside the voice; no list on the reference |
| Hedra | Full multi-language support claimed, none named |
| HeyGen | Thirty or more, counted on the translation endpoint |
| Kling AI | Five, not named |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Not documented by the vendor |
| PixVerse | Multiple, with speech, singing and advertisements named as types |
| Runway | Not documented |
| SceneMixer | 15, named, with Cantonese a 16th for dialogue only |
| Sora 2 | Not documented |
| sync-3 | Ninety-five or more, counted rather than named |
| Synthesia | A catalogue, each voice carrying a language name and code |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Not documented |
| Wan2.2-S2V | Not documented |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Voice source |
|---|---|
| D-ID | A voice id from one of five named speech providers |
| Hedra | An uploaded track, or a voice id from the voices endpoint |
| HeyGen | A clone, instant from one recording or professional from 20 minutes or more |
| Kling AI | Not documented by the vendor |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Reference audio, 15 s total across 3 clips |
| PixVerse | Built-in voices, or a custom voice from a supplied sample |
| Runway | An audio sample of 10 seconds to 5 minutes, or a description in words |
| SceneMixer | Preset voice library, or the user's own recordings |
| Sora 2 | Not documented |
| sync-3 | Not documented on the model page |
| Synthesia | The catalogue, or a cloned voice with a language list of its own |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Reference audio, 15 seconds in total, WAV or MP3 up to 15 MB |
| Wan2.2-S2V | Whatever track is handed over; no catalogue |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
7Sources
Taken from the voices reference and avatar documentation at docs.synthesia.io on 2026-09-22. Every entry against the same five fields sits on the speech controls page; the rule behind each field is on how read. The entries either side of this one: Veo, Vidu.
- LanguagesEach voice is listed with its formal and native language name, a language code, a gender, a name and a voice idpublished as a catalogue with codes
- Lip-syncLip sync and facial expressions are generated from the spoken contentdriven by the script
- A framing condition attached to lip-syncLip sync performs best when the avatar is framed closer in the scene rather than positioned far from the camera
- Voice sourceVoice cloning carries a language list of its own, running from Afrikaans to Zulu and described as the same list used for standalone cloninga cloned voice with its own reach