Synthesia: a voice id, an avatar, and no stated pairing
Both halves of a character exist here as documented objects. Voices carry ids in a published catalogue and avatars are created and reused. What is missing is the sentence that joins them, which is the odd gap in an otherwise structured platform. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| Each voice is listed with a language name, a code, a gender, a name and an id | Which id belongs to which avatar by default, if any |
| Lip sync and facial expressions are generated from the spoken content | Whether the avatar constrains which voices suit it |
| Voice cloning carries a language list of its own | Whether a cloned voice can be made the avatar's own |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1Two well-documented objects and a missing join
Most blanks in this column belong to products that document almost nothing about voices. This one documents them thoroughly and still leaves the pairing unstated, which is the harder kind of gap to read: the pieces are all present and their relationship is assumed.
In practice a production will pass a voice id with an avatar on every request and treat the combination as the character. That works, and it means the character lives in the calling code rather than on the platform.
2The framing condition hints that pairing matters
Lip sync is documented as performing better with the avatar framed closer to camera, which says the relationship between a face and a spoken line has conditions on it. If framing affects it, a voice's pitch and pace plausibly do too, and nothing addresses whether some voices suit some avatars.
That is left as a question rather than inferred. The register logs that both objects are published with ids and that no document connects a given pair.
3Others whose documentation never addresses persistence
Seven more entries reach this column empty. This is the only one where both a voice and a character are published as durable objects and the link between them is not.
- Hedra — not documented by the vendor.
- Luma Ray — not documented by the vendor.
- SceneMixer — not documented by the vendor.
- sync-3 — not documented by the vendor.
- Veo — not documented by the vendor.
- Vidu — not documented by the vendor.
- Wan2.2-S2V — not documented by the vendor.
- LanguagesEach voice is listed with its formal and native language name, a language code, a gender, a name and a voice idpublished as a catalogue with codes
- A framing condition attached to lip-syncLip sync performs best when the avatar is framed closer in the scene rather than positioned far from the camera
- Voice sourceVoice cloning carries a language list of its own, running from Afrikaans to Zulu and described as the same list used for standalone cloninga cloned voice with its own reach
4Sources
Read from the voices reference and avatar documentation at docs.synthesia.io on 2026-09-22. The same column across every entry is on per-character binding; everything this vendor publishes about speech is on Synthesia. What counts as documented is on how read.