Hedra and Synthesia: a waveform, or the script itself
Both name what the mouth follows, which only five entries here manage. One follows the audio a production supplies; the other generates lip sync and expression from the script, and adds a condition about framing. As of 2026-09-22.
| Field | Hedra | Synthesia | Where they part |
|---|---|---|---|
| Audio source | Supplied; audio is required | From the script, or uploaded | Different answers |
| Voice source | A track, or a voice id from the endpoint | A catalogue voice, or a cloned one | Different answers |
| Per-character binding | Not documented by the vendor | Not documented by the vendor | Same answer |
| Languages | A claim with no names | A catalogue with codes | Different answers |
| Lip-sync | The supplied audio | The spoken content, framed close | Different answers |
Inclusion rule. Two entries are given a page together when at least one column puts them at opposite grades of answer. Pairs that agree on every column, or that are both blank throughout, do not get a page. Order. Fixed field order, identical on every side-by-side page.
1Following a waveform against following words
Where the driver is a recording, a performer's timing reaches the screen: pauses, emphasis and breath are all in the file. Where the driver is the script, the model decides those things, and the same words will always produce roughly the same reading.
For drama the first preserves a performance and the second removes one. For a presentation the second is cheaper and entirely adequate, which is why these two products have different customers despite looking alike.
2Only one of them publishes a weak case
Saying performance improves with the avatar framed closer names the shot where alignment struggles, which is a constraint on coverage rather than a tip. The audio-driven entry names its driver and says nothing about when it holds less well.
Both sentences are worth having and the second kind is rarer. A vendor willing to publish where something fails has given a production more than one willing only to publish that it works.
3Length comes from opposite places
One says the audio normally determines the video length, so duration lives in the recording and a slot is fitted by re-recording. The other drives everything from the script, so a rewrite is the lever and a length change is a text change.
Neither publishes a language list. One claims full multi-language support; the other publishes a code on every voice, which is the widest gap between these two rows and the one that decides which markets are reachable.
4Each of them on its own
The column this pair was chosen for is lip-sync, and each entry has a page of its own on it. The full row for either, all five columns with the wording behind each cell, is on its model note.
- Hedra on lip-sync — the supplied audio.
- Synthesia on lip-sync — the spoken content, framed close.
- Hedra, all five fields — read from the Character 3 model page and avatar video guide.
- Synthesia, all five fields — read from the voices reference and avatar documentation.
- Audio sourceAvatar videos are driven by audio, and the character in the image will lip-sync and move to the audio providedhanded to the model rather than made by it
- Shot length set by the takeThe audio normally determines the video length, and omitting the duration follows the source audio
- LanguagesDescribed as text, image and audio to video with full multi-language support, and a maximum duration of ten minutesa claim carrying no list
- Lip-syncLip sync and facial expressions are generated from the spoken contentdriven by the script
- A framing condition attached to lip-syncLip sync performs best when the avatar is framed closer in the scene rather than positioned far from the camera
5Sources
Every cell above is read from the documentation each vendor publishes, on the dates carried by the two model notes. The whole register on one column is on the field note; entries grouped by the shape of their answer are on the routes. Other pairs: Wan 3.0 and Wan2.2-S2V, Veo and Hedra.