SceneMixer: the line is performed as the shot renders
SceneMixer describes the video model as performing the line in the language chosen for the project while it renders the shot. The vendor draws a consequence from that: no separate dubbing step, and no lip-sync patch afterwards. As of 2026-09-12.
| What the documentation settles | What it leaves to a take |
|---|---|
| The model performs the line in the chosen language while it renders the shot | Whether the delivery convinces a listener who speaks that language at home |
| Dialogue is set per project to any of fifteen languages, with Cantonese a sixteenth | Whether register and idiom hold up across all sixteen |
| There is no separate dubbing step and no lip-sync patch afterwards | What happens when a line has to change after the shot is approved |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1A claim about method rather than about quality
Most entries treat language as a property of a voice library. This one treats it as a property of the performance: the project fixes the language, and the model speaks the line in it as the picture is made. That is a mechanism statement, and it is falsifiable in a way a capability claim is not.
The consequence the vendor draws follows logically. If nothing is translated after the fact, there is no track to lay over finished footage and therefore nothing to re-synchronise. Whether the result sounds native is a separate question that no page settles and this register does not test.
2Where the setting lives decides how often it can change
Setting dialogue language per project rather than per shot makes the choice durable and makes it early. A season is committed to a language before the first episode is generated, which is the right order for consistency and the wrong order for experiments.
It also means a market added later is a new project rather than a re-render, since the performance was made in the old language rather than laid over silent footage. That is the trade the method implies, and the documentation does not discuss it.
3The others that make sound while they make the picture
Seven more entries generate audio in the same pass. This is the only one whose vendor names the languages that audio is spoken in.
- Kling AI — with the picture, on video 3.0.
- LTX Studio — with the picture, and audio to video.
- MiniMax — with the picture.
- Sora 2 — with the picture.
- Veo — with the picture.
- Vidu — with the picture, speech available on its own.
- Wan 3.0 — with the picture, unless switched off.
- Audio sourceThe video model is described as performing the line in the chosen language while it renders the shotgenerated with the picture
- LanguagesDialogue is set per project to any of 15 languages, with Cantonese a 16th for dialogue onlya named list of 15
- Lip-syncThe vendor states there is no separate dubbing step and no lip-sync patch afterwardsstated as unnecessary rather than as a feature
4Sources
Read from the languages guide at scenemixer.com on 2026-09-12. The same column across every entry is on audio source; everything this vendor publishes about speech is on SceneMixer. What counts as documented is on how read.