What SceneMixer documents about speech
SceneMixer (scenemixer.com), an AI short-drama generator, publishes a named list of 15 dialogue languages and states that the model performs the line in the chosen language as it renders the shot. Voices come from a preset library or the user's own recordings. As of 2026-09-12.
| Field | What the vendor documents |
|---|---|
| Audio source | Generated with the picture, in the language set for the project |
| Languages | 15, named, with Cantonese a 16th for dialogue only |
| Voice source | Preset voice library, or the user's own recordings |
| Per-character binding | Not documented |
| Lip-sync | Stated as unnecessary: no dubbing pass, so no patch |
Inclusion rule. Fields are the same five for every model here; a field the vendor does not document is recorded as not documented rather than inferred from generated output. Order. Fixed field order, identical on every model page.
1A voice statement written as a rule, not as a feature
The sentence that fills this entry's voice field is in a checklist about likeness and publishing risk rather than in product documentation. It names two routes, a preset library and the user's own recordings, and rules out a third by treating cloning of an identifiable voice as a likeness issue, citing US digital-replica statutes.
That framing is unusual in this register and it is worth reading precisely. It confirms a preset library exists and that supplied recordings are accepted. It does not say how many voices the library holds, which languages they cover, or how a voice is attached to a character.

2The language answer is the distinctive one
Every other entry in this register treats language as a property of the voice library. This one treats it as a property of the performance: the project fixes one of 15 named languages, and the model is described as speaking the line in that language while it renders the shot. Cantonese is listed as a sixteenth option for dialogue only.
The consequence the vendor draws from that is worth recording precisely, because it is a claim about method rather than about quality: if the performance is generated in the target language there is no translated track to lay over finished footage, so no dubbing pass and no lip-sync patch. Whether the result convinces a native speaker is not something a published sentence settles, and this site does not test.
3What is still unstated
How a voice is attached to a particular character across episodes is not described anywhere public. Two entries in this register bind a voice to a reusable character object and one takes a reference clip per call; this one names a library without saying how a selection persists.
The compliance guide adds one boundary from the other direction: cloning an identifiable voice is treated as a likeness issue, with US digital-replica statutes named. That rules a method out rather than describing the one in use.
4This entry, one column at a time
Each of these stays inside a single field: what SceneMixer puts there, what the wording settles, and what it leaves for a take to answer.
- Audio source — with the picture, in the language set for the project.
- Voice source — a preset library, or the user's own recordings.
- Per-character binding — not documented by the vendor.
- Languages — a named list, set per project.
- Lip-sync — argued out of existence.
5Read against another entry
Each of these puts SceneMixer beside one other entry on a column where the two land at opposite grades of answer.
- SceneMixer and Kling AI — on languages.
- SceneMixer and Synthesia — on languages.
6Where this model sits against the rest
| Model | Audio source |
|---|---|
| D-ID | A text script read aloud, or an audio url supplied |
| Hedra | Supplied; audio is a required input |
| HeyGen | A text to speech endpoint, apart from the picture |
| Kling AI | Generated with the picture, VIDEO 3.0 |
| LTX Studio | Generated jointly with the picture, plus audio-to-video |
| Luma Ray | Not documented; the video reference does not mention sound |
| MiniMax | Native speech, generated with the picture |
| PixVerse | Supplied to a speech endpoint, or read from text there |
| Runway | Not documented where voices are defined |
| SceneMixer | Generated with the picture, in the language set for the project |
| Sora 2 | Generated with the picture; video and audio are both listed as output |
| sync-3 | Supplied, or read from text on the call |
| Synthesia | Synthesised from the script, or an uploaded recording |
| Veo | Generated with the video rather than added afterwards |
| Vidu | Q3 generates speech, effects and music natively, on by default |
| Wan 3.0 | On by default; false returns a file with no audio track |
| Wan2.2-S2V | Supplied; the card describes the model as audio-driven |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Languages |
|---|---|
| D-ID | A language field beside the voice; no list on the reference |
| Hedra | Full multi-language support claimed, none named |
| HeyGen | Thirty or more, counted on the translation endpoint |
| Kling AI | Five, not named |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Not documented by the vendor |
| PixVerse | Multiple, with speech, singing and advertisements named as types |
| Runway | Not documented |
| SceneMixer | 15, named, with Cantonese a 16th for dialogue only |
| Sora 2 | Not documented |
| sync-3 | Ninety-five or more, counted rather than named |
| Synthesia | A catalogue, each voice carrying a language name and code |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Not documented |
| Wan2.2-S2V | Not documented |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
| Model | Voice source |
|---|---|
| D-ID | A voice id from one of five named speech providers |
| Hedra | An uploaded track, or a voice id from the voices endpoint |
| HeyGen | A clone, instant from one recording or professional from 20 minutes or more |
| Kling AI | Not documented by the vendor |
| LTX Studio | Not documented |
| Luma Ray | Not documented |
| MiniMax | Reference audio, 15 s total across 3 clips |
| PixVerse | Built-in voices, or a custom voice from a supplied sample |
| Runway | An audio sample of 10 seconds to 5 minutes, or a description in words |
| SceneMixer | Preset voice library, or the user's own recordings |
| Sora 2 | Not documented |
| sync-3 | Not documented on the model page |
| Synthesia | The catalogue, or a cloned voice with a language list of its own |
| Veo | Not documented |
| Vidu | Not documented |
| Wan 3.0 | Reference audio, 15 seconds in total, WAV or MP3 up to 15 MB |
| Wan2.2-S2V | Whatever track is handed over; no catalogue |
Inclusion rule. Models whose vendor documents something about generated speech. A model with nothing published on the point is still listed, with the cell marked hollow. Order. Alphabetical by model name.
7Sources
Taken from the languages guide at scenemixer.com on 2026-09-12. Every entry against the same five fields sits on the speech controls page; the rule behind each field is on how read. The entries either side of this one: Sora 2, sync-3.
- Audio sourceThe video model is described as performing the line in the chosen language while it renders the shotgenerated with the picture
- LanguagesDialogue is set per project to any of 15 languages, with Cantonese a 16th for dialogue onlya named list of 15
- Lip-syncThe vendor states there is no separate dubbing step and no lip-sync patch afterwardsstated as unnecessary rather than as a feature
- Voice sourceThe guide tells users to use a preset voice library or their own recordingstwo routes named
- What the vendor rules outCloning an identifiable voice is described as a likeness issue, with US digital-replica statutes named
- Per-character bindingNot documented by the vendor (as of 2026-09-12)no public description