Synthesia: generated from the spoken content, framed close
Lip sync and facial expressions are generated from the spoken content, which names the driver, and a second sentence attaches a condition: performance is better with the avatar framed closer rather than far from the camera. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| Lip sync and facial expressions are generated from the spoken content | How much of a performance a script can carry unaided |
| Lip sync performs better with the avatar framed closer rather than far from the camera | Where the threshold sits, since only a direction is given |
| Each voice is listed with a language name, a code, a gender, a name and an id | Whether alignment differs between catalogue voices and clones |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1A driver and a condition, which is the fullest answer here
Two sentences do more than most entries manage with one. The first says what the movement comes from; the second says when it works less well. Together they let a production plan coverage rather than guess at it, which is the practical test this column applies.
The condition is also the more useful half, because it names the weak case. Wide framing is the staple of dialogue coverage, and a vendor saying alignment prefers closer framing has constrained a shot list rather than offered a tip.
2Driving from the script has a consequence nobody states
If the face comes from the spoken content, then the same words will always produce roughly the same performance, and a rewrite is the only lever. There is no documented way to ask for the line colder, faster or through clenched teeth while keeping the words.
That is a real limitation for drama and a non-issue for the presentations this platform is mostly used for. It is recorded here because this register is read by people making the first kind of thing.
3The others that name what the mouth is following
Four more entries name a driver. All four of those follow a supplied or generated waveform; this is the only one driving the face from the script itself.
- Hedra — the supplied audio.
- MiniMax — whoever is on screen.
- sync-3 — the audio it is given.
- Wan2.2-S2V — the audio input.
- Lip-syncLip sync and facial expressions are generated from the spoken contentdriven by the script
- A framing condition attached to lip-syncLip sync performs best when the avatar is framed closer in the scene rather than positioned far from the camera
- LanguagesEach voice is listed with its formal and native language name, a language code, a gender, a name and a voice idpublished as a catalogue with codes
4Sources
Read from the voices reference and avatar documentation at docs.synthesia.io on 2026-09-22. The same column across every entry is on lip-sync; everything this vendor publishes about speech is on Synthesia. What counts as documented is on how read.