Sentioscope

Speech and voice controls, as each vendor documents them

Synthesia: the script is the source of the sound

Lip sync and facial expressions are generated from the spoken content, which makes the script the origin of the sound and of the face. A recording can be uploaded instead, and each voice in the catalogue carries a language name and a code. As of 2026-09-22.

One script driving the sound and the face togetherLip sync and facial expressions are generated from the spoken content, so rewriting a line rewrites the performance. A published condition follows: closer framing performs better than a figure far from camera.What the text decides once it is writtenScriptThe spoken content is the single input that decides everythingbelow it.VoiceA catalogue entry with a language name and code, or a clonedvoice.FaceLip sync and expression are generated from the same spokencontent.FramingPerformance is stated to improve with the avatar closer tocamera.A rewrite is a new performance, not a new read
Fig. 1 There is no place in this arrangement to ask for the same words colder or faster, because the model is reading them rather than taking notes.
Synthesia on audio source, statement by statement. Read from the vendor's voices reference and avatar documentation on 2026-09-22.
What the documentation settlesWhat it leaves to a take
Lip sync and facial expressions are generated from the spoken contentHow much of a performance the text can carry without direction
Lip sync performs better with the avatar framed closer rather than far from the cameraWhere the framing threshold is, since only the direction is given
Each voice is listed with a language name, a code, a gender, a name and an idWhich of those voices suits a dramatic read rather than a presentation

Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.

1Driving the face from the text is a different bargain

If the expression comes from the spoken content, then rewriting a line rewrites the performance. That is convenient for corporate work, where the script is the deliverable, and awkward for drama, where the same words can be said several ways and the choice between them is the craft.

It also removes a lever. There is no place in this arrangement to ask for the line to be delivered colder, faster or through clenched teeth, because the model is reading the words rather than taking notes about them.

2A published framing condition is rarer than a published capability

Most vendors that mention lip-sync name it and move on. This one attaches a condition: closer framing performs better than a figure far from the camera. That is a sentence a shot list can act on, and it quietly tells a reader which shot is the weak case.

Wide two-shots are the staple of dialogue coverage, so a condition favouring close framing is a constraint on coverage rather than a tip. Recording it in this field keeps that consequence attached to the claim it comes from.

3The others that synthesise before they render

Three more entries build speech in a stage of its own. This is the one whose documentation also ties the face to the same script.

  • D-ID — from a script, or a supplied url.
  • HeyGen — from a speech endpoint, then rendered.
  • PixVerse — supplied, or read from text.
  • Lip-sync
    Lip sync and facial expressions are generated from the spoken contentdriven by the scriptSynthesia, create an avatar / recorded 2026-09-22
  • A framing condition attached to lip-sync
    Lip sync performs best when the avatar is framed closer in the scene rather than positioned far from the cameraSynthesia, create an avatar / recorded 2026-09-22
  • Languages
    Each voice is listed with its formal and native language name, a language code, a gender, a name and a voice idpublished as a catalogue with codesSynthesia, list of supported voices / recorded 2026-09-22

4Sources

Read from the voices reference and avatar documentation at docs.synthesia.io on 2026-09-22. The same column across every entry is on audio source; everything this vendor publishes about speech is on Synthesia. What counts as documented is on how read.