Three ways to get a voice, and what each costs
Three routes exist to give a character a voice. A preset library is instant and finite; cloning gives a unique voice and asks a consent question first; speech generated with the picture arrives aligned and cannot be retaken on its own. As of 2026-09-12.
| Route | Billed by | What it gives | What it costs |
|---|---|---|---|
| Cloning a voice | Per voice, one-off or by subscription | A unique voice, reusable for the run | A consent question before a price question |
| Native audio with the picture | Inside the per-second video rate | Alignment by construction | A retake redoes the picture too |
| Preset library | Per character of speech | Instant, auditionable, stable across episodes | A finite catalogue |
Inclusion rule. Routes documented by vendors for putting speech on a generated character. Recording a human performer is a fourth option and is outside what these tools publish. Order. Alphabetical by route.
1A preset library is cheaper than it looks and narrower
Billing by character of speech makes a season's dialogue cost very little, and the catalogue is the constraint rather than the price. Two characters that should sound different may have to share a register because nothing else in the catalogue fits.
The compensating advantage is stability: a preset sounds the same in episode forty as in episode one, which is the property serial work needs most.
Billing by character also means the cost is set by the script rather than by the running time, so a dialogue-heavy episode and a quiet one differ in audio cost far more than they differ in picture cost.
2Cloning is a rights question before it is a price question
Whose voice is being cloned, and on what consent, is settled before any of the pricing matters. Vendors publish minimum sample lengths and plan tiers; none of that addresses permission.
Where the answer is clean, a cloned voice gives a character a voice no other production has, and it stays available for the whole run.
Where the answer is not clean, no pricing tier makes it cleaner, and a cloned voice is the item in a finished series hardest to replace afterwards. That argues for settling it before the first episode rather than at the point of delivery.
3Native audio trades control for alignment
Speech generated with the picture is aligned by construction, which removes the hardest problem. It also welds the two together: a line that needs re-reading means regenerating the shot, and an approved picture is lost with it.
For dialogue-heavy serial work that trade is usually worth making early and regretted late, which is why productions often mix routes rather than choosing one.
A common split is native audio for shots where alignment carries the scene and a preset or cloned voice for everything else. That keeps the expensive coupling to the shots that need it.
4The three real causes of broken lip alignment
Audio laid over picture that was generated silent. A dubbed line longer than the original because the target language expanded. A regenerated shot matched against an older audio take.
None of the three are model defects, and all three are avoidable by deciding the route before the shot list exists. What each model documents is on the control table.
Alignment is judged on the delivered cut rather than on the clip, so it is worth checking after assembly rather than at generation. A shot that looked right on its own can drift once it sits between two others.
5Where the published detail is logged
What is described above is how the techniques work in general. Which models document which of them, in whose words, is kept in the speech controls table with a date on every field.
- MiniMax pay-as-you-go pricing — publishes pay-as-you-go rates per generation
A mechanism note rather than a documented field. Nothing here is attributed to a model, and nothing here is a claim about one. The sourced material is on the speech controls table. Related: Reading a blank cell, Counts and lists.