Veo and Hedra: sound out of one, into the other
Veo documents audio generated with the video and stops. Hedra requires audio as an input and says the character lip-syncs and moves to it. Two sparse entries, sparse in opposite places. As of 2026-09-22.
| Field | Veo | Hedra | Where they part |
|---|---|---|---|
| Audio source | With the picture | Supplied; audio is required | Different answers |
| Voice source | Not documented by the vendor | A track, or a voice id from the endpoint | Only Hedra answers |
| Per-character binding | Not documented by the vendor | Not documented by the vendor | Same answer |
| Languages | Not documented by the vendor | A claim with no names | Only Hedra answers |
| Lip-sync | Not documented by the vendor | The supplied audio | Only Hedra answers |
Inclusion rule. Two entries are given a page together when at least one column puts them at opposite grades of answer. Pairs that agree on every column, or that are both blank throughout, do not get a page. Order. Fixed field order, identical on every side-by-side page.
1Two thin pages, and the thinness lands differently
Neither of these publishes much. The difference is which question each one answers. A sentence saying audio arrives with the video tells a production a dubbing stage is unnecessary and nothing about casting. A sentence saying audio is required tells a production to book a session and nothing about what comes back.
For planning, the second is the more actionable. A required input is a task on a schedule; a produced output is a hope with a date on it.
2Only one of them addresses the mouth
The audio-driven entry names its driver because the driver is the input, which costs the vendor nothing to say. The native-audio entry never mentions mouths, which is the common position among entries that make sound and picture together.
That pattern holds across this register and is worth stating plainly: a production that needs a published statement about lip movement is mostly choosing between products that are handed a recording.
3Both leave the cast entirely open
Neither publishes a language list, and neither describes a voice attaching to a character. One of them at least lets a production supply the voice, which makes consistency an archival problem rather than an impossible one.
For a serial that is the deciding difference. A voice a production owns can be repeated in episode nine; a voice the model invents on each call cannot be, and no sentence on either page says otherwise.
4Each of them on its own
The column this pair was chosen for is audio source, and each entry has a page of its own on it. The full row for either, all five columns with the wording behind each cell, is on its model note.
- Veo on audio source — with the picture.
- Hedra on audio source — supplied; audio is required.
- Veo, all five fields — read from the API documentation.
- Hedra, all five fields — read from the Character 3 model page and avatar video guide.
- Audio sourceAudio is generated with the video rather than added afterwardsgenerated with the picture
- Audio sourceAvatar videos are driven by audio, and the character in the image will lip-sync and move to the audio providedhanded to the model rather than made by it
- Shot length set by the takeThe audio normally determines the video length, and omitting the duration follows the source audio
- LanguagesDescribed as text, image and audio to video with full multi-language support, and a maximum duration of ten minutesa claim carrying no list
5Sources
Every cell above is read from the documentation each vendor publishes, on the dates carried by the two model notes. The whole register on one column is on the field note; entries grouped by the shape of their answer are on the routes. Other pairs: Vidu and sync-3, Synthesia and sync-3.