Two entries read against each other, column by column
A pair earns a page when one column puts the two at opposite grades: a list against a total, a named driver against a silence, a stored voice against a clip supplied on every call. As of 2026-09-12.
The pairing rule is deliberately narrow. Two entries that agree everywhere teach a reader nothing that their own notes do not already say in fewer words, and two entries that are blank everywhere produce a third silence. What is worth a page is a disagreement, because a disagreement names the choice a production is actually making.
Every page here carries the same five-column table, because the disagreement that justified the pairing is rarely the only one. Entries chosen for the language column routinely turn out to answer the mouth question at opposite grades as well, and that second gap is often the one that decides a schedule.
None of these pages scores anything. The gap column says where the two part, and the prose says what each position costs, which is as far as documentation can take a comparison. How either one sounds is settled by a take.
1The pairs
- Wan 3.0 and Wan2.2-S2V — Two entries from the same model family sit at opposite ends of this register: one makes the sound with the picture, the other is driven by a recording.
- Veo and Hedra — One entry generates audio with the video and documents nothing else; the other cannot run without a recording and names what the mouth follows.
- Vidu and sync-3 — One entry generates speech, effects and music in one pass; the other takes a waveform and moves a mouth to it, and publishes a driver for doing so.
- Synthesia and sync-3 — One entry publishes a language code on every voice; the other publishes ninety-five or more as a total and never mentions a voice at all.
- SceneMixer and Kling AI — Both generate audio with the picture.
- Runway and MiniMax — One entry builds a voice once, previews it and reuses it by id; the other rebuilds it from fifteen seconds of audio on every generation.
- HeyGen and LTX Studio — Both let a voice outlast a call.
- MiniMax and Wan 3.0 — Two entries publish an identical fifteen-second total for reference audio and arrange everything around it differently, from clip counts to pricing.
- Sora 2 and Luma Ray — One entry publishes where a line goes, how to label speakers and how much speech fits a shot; the other has a reference that never mentions sound.
- Synthesia and D-ID — Two talking-head products.
- HeyGen and Synthesia — Two avatar platforms answering the language question at different grades: thirty or more counted on translation, against a code published on every voice.
- D-ID and Hedra — Two avatar products taking opposite inputs: one reads a script of up to forty thousand characters, the other requires a recording and follows its length.
- Kling AI and Synthesia — Both name lip-sync.
- SceneMixer and Synthesia — Both entries that name languages here do it in a different shape: a per-project setting of fifteen, against a code published on every voice.
- PixVerse and D-ID — Both pass a voice identifier on each request.
- Sora 2 and Wan 3.0 — Both put dialogue into the prompt and disagree on how.
- Runway and Luma Ray — Both entries reach most fields empty.
- LTX Studio and Vidu — Both generate audio with the picture.
- MiniMax and Sora 2 — Two native-audio entries that say something about a shot with two people in it, one as a behaviour and one as an instruction to the writer.
- Hedra and Synthesia — Two avatar products naming a driver for the mouth.
A pair is added when the two entries can be shown to disagree from published wording alone; a pairing that needed a listening test to justify it is not made here.
Other notes: Speech controls, Models, Fields, Routes, Questions, Terms, Learn, Data. What counts as a documented control is set out on the reading page.