Synthesia: the script is the source of the sound
Lip sync and facial expressions are generated from the spoken content, which makes the script the origin of the sound and of the face. A recording can be uploaded instead, and each voice in the catalogue carries a language name and a code. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| Lip sync and facial expressions are generated from the spoken content | How much of a performance the text can carry without direction |
| Lip sync performs better with the avatar framed closer rather than far from the camera | Where the framing threshold is, since only the direction is given |
| Each voice is listed with a language name, a code, a gender, a name and an id | Which of those voices suits a dramatic read rather than a presentation |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1Driving the face from the text is a different bargain
If the expression comes from the spoken content, then rewriting a line rewrites the performance. That is convenient for corporate work, where the script is the deliverable, and awkward for drama, where the same words can be said several ways and the choice between them is the craft.
It also removes a lever. There is no place in this arrangement to ask for the line to be delivered colder, faster or through clenched teeth, because the model is reading the words rather than taking notes about them.
2A published framing condition is rarer than a published capability
Most vendors that mention lip-sync name it and move on. This one attaches a condition: closer framing performs better than a figure far from the camera. That is a sentence a shot list can act on, and it quietly tells a reader which shot is the weak case.
Wide two-shots are the staple of dialogue coverage, so a condition favouring close framing is a constraint on coverage rather than a tip. Recording it in this field keeps that consequence attached to the claim it comes from.
3The others that synthesise before they render
Three more entries build speech in a stage of its own. This is the one whose documentation also ties the face to the same script.
- D-ID — from a script, or a supplied url.
- HeyGen — from a speech endpoint, then rendered.
- PixVerse — supplied, or read from text.
- Lip-syncLip sync and facial expressions are generated from the spoken contentdriven by the script
- A framing condition attached to lip-syncLip sync performs best when the avatar is framed closer in the scene rather than positioned far from the camera
- LanguagesEach voice is listed with its formal and native language name, a language code, a gender, a name and a voice idpublished as a catalogue with codes
4Sources
Read from the voices reference and avatar documentation at docs.synthesia.io on 2026-09-22. The same column across every entry is on audio source; everything this vendor publishes about speech is on Synthesia. What counts as documented is on how read.