What Veo documents about speech
Veo (ai.google.dev) documents audio generated with the video rather than added afterwards. The other four fields this register keeps are not addressed in the API documentation. As of 2026-09-12.
| Field | What the vendor documents |
|---|---|
| Audio source | Generated with the video rather than added afterwards |
| Languages | Not documented |
| Voice source | Not documented |
| Per-character binding | Not documented |
| Lip-sync | Not documented |
Inclusion rule. Fields are the same five for every model here; a field the vendor does not document is recorded as not documented rather than inferred from generated output. Order. Fixed field order, identical on every model page.
1One field filled, and it is the field everyone fills
Audio generated with the picture is the common position among the entries read first here, so the one documented field on this entry is also the least distinguishing. What separates the entries in this register is what comes after that sentence, and for this model nothing does.
The emptiness is not evidence of a weaker model. It is evidence about the documentation, which is written for developers calling an API and describes parameters rather than production behaviour. A field this site keeps is a field a series has to plan around, and those two audiences do not want the same page.

2What a drama production cannot plan from this page
Whether a character can be given a voice that persists, which languages the model will speak, whether a supplied recording is accepted, and what happens to mouths when two people talk. Four questions, none answered, and each of them decides a workflow rather than a preference.
The honest reading is that the answer may exist and may be good; it is not published where a developer or a producer would find it before committing. Recording that as not documented keeps the comparison on the same footing as every other row, which is the only way a register of this kind stays worth reading.
3Why the entry stays in the register
Dropping the sparse rows would leave a table that flatters the vendors who write the most. Seventeen entries have now been read the same way against the same five fields, and the shape of the result, one common answer and four blanks, is itself the finding for this entry.
If the documentation grows a paragraph about voices, this page changes and its date moves. Until then a reader can take one thing from it with confidence and take nothing else, which is more useful than a page that guesses at the other four.
4Read against another entry
Each of these puts Veo beside one other entry on a column where the two land at opposite grades of answer.
- Veo and Hedra — on audio source.
5Where this model sits against the rest
| Model | Audio source |
|---|---|
| D-ID | A text script read aloud, or an audio url supplied |
| Hedra | Supplied; audio is a required input |
| HeyGen | A text to speech endpoint, apart from the picture |
| Kling AI | Generated with the picture, VIDEO 3.0 |
| LTX Studio | Generated jointly with the picture, plus audio-to-video |
| Luma Ray | Not documented; the video reference does not mention sound |
| MiniMax | Native speech, generated with the picture |
| PixVerse | Supplied to a speech endpoint, or read from text there |
| Runway | Not documented where voices are defined |
| SceneMixer | Generated with the picture, in the language set for the project |
| Sora 2 | Generated with the picture; video and audio are both listed as output |
| sync-3 | Supplied, or read from text on the call |
| Synthesia | Synthesised from the script, or an uploaded recording |
| Veo | Generated with the video rather than added afterwards |
| Vidu | Q3 generates speech, effects and music natively, on by default |
| Wan 3.0 | On by default; false returns a file with no audio track |
| Wan2.2-S2V | Supplied; the card describes the model as audio-driven |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Languages |
|---|---|
| D-ID | A language field beside the voice; no list on the reference |
| Hedra | Full multi-language support claimed, none named |
| HeyGen | Thirty or more, counted on the translation endpoint |
| Kling AI | Five, not named |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Not documented by the vendor |
| PixVerse | Multiple, with speech, singing and advertisements named as types |
| Runway | Not documented |
| SceneMixer | 15, named, with Cantonese a 16th for dialogue only |
| Sora 2 | Not documented |
| sync-3 | Ninety-five or more, counted rather than named |
| Synthesia | A catalogue, each voice carrying a language name and code |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Not documented |
| Wan2.2-S2V | Not documented |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Voice source |
|---|---|
| D-ID | A voice id from one of five named speech providers |
| Hedra | An uploaded track, or a voice id from the voices endpoint |
| HeyGen | A clone, instant from one recording or professional from 20 minutes or more |
| Kling AI | Not documented by the vendor |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Reference audio, 15 s total across 3 clips |
| PixVerse | Built-in voices, or a custom voice from a supplied sample |
| Runway | An audio sample of 10 seconds to 5 minutes, or a description in words |
| SceneMixer | Preset voice library, or the user's own recordings |
| Sora 2 | Not documented |
| sync-3 | Not documented on the model page |
| Synthesia | The catalogue, or a cloned voice with a language list of its own |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Reference audio, 15 seconds in total, WAV or MP3 up to 15 MB |
| Wan2.2-S2V | Whatever track is handed over; no catalogue |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
6Sources
Taken from the API documentation at ai.google.dev on 2026-09-12. Every entry against the same five fields sits on the speech controls page; the rule behind each field is on how read. The entries either side of this one: Vidu, Wan 3.0.
- Audio sourceAudio is generated with the video rather than added afterwardsgenerated with the picture