Sentioscope

Speech and voice controls, as each vendor documents them

HeyGen and LTX Studio: a voice object, or a character

Two entries here with a durable voice, and they store it in different places. One enrols a clone and reuses it by id. The other attaches a voice to a character Element and never says where the voice came from. As of 2026-09-22.

A voice object and a character object, weighedAn id can be passed wrongly, so a production keeps a record of which id is which character. A character object cannot be passed wrongly, and cannot be moved, shared or reconstructed either.A clone with an idA voice on a characterCan be passed wronglyYes, so keep a recordNo, there is nothing to passCan be reused elsewhereYes, across series if wantedNo, it belongs to that characterHow it was madeOne recording, or twenty minutesNot stated anywhereIf it is lostEnrol again from the same audioNo published way to reconstruct itPersistence and selection come apart here
Fig. 1 Both solve keeping a voice; only one says what the voice will be before a season is committed to it.
Controls documented, model by modelFilled where the model documents that control, hollow where nothing is published about it.Controls documented, model by modelHeyGenLTX StudioWhere they partAudio sourceAudio source — HeyGen: From a speech endpoint, then renderedAudio source — LTX Studio: With the picture, and audio to videoAudio source — Where they part: Different answersVoice sourceVoice source — HeyGen: A recording to clone fromVoice source — LTX Studio: Nothing documented as an inputVoice source — Where they part: Different answersPer-character bindingPer-character binding — HeyGen: On a stored clonePer-character binding — LTX Studio: On the character ElementPer-character binding — Where they part: Different answersLanguagesLanguages — HeyGen: A count, attached to translationLanguages — LTX Studio: Not documented by the vendorLanguages — Where they part: Only HeyGen answersLip-syncLip-sync — HeyGen: Named, alongside translationLip-sync — LTX Studio: Not documented by the vendorLip-sync — Where they part: Only HeyGen answers
Fig. 2 Filled where the model documents that control, hollow where nothing is published about it.
Models documenting each controlHow many models document each control. A hollow column is a statement about documentation, not capability.Models documenting each controlHeyGen4 of 5LTX Studio3 of 5Where they part5 of 5
Fig. 3 How many models document each control. A hollow column is a statement about documentation, not capability.
HeyGen and LTX Studio on the five fields this register keeps, with the gap on each. Recorded 2026-09-22.
FieldHeyGenLTX StudioWhere they part
Audio sourceFrom a speech endpoint, then renderedWith the picture, and audio to videoDifferent answers
Voice sourceA recording to clone fromNothing documented as an inputDifferent answers
Per-character bindingOn a stored cloneOn the character ElementDifferent answers
LanguagesA count, attached to translationNot documented by the vendorOnly HeyGen answers
Lip-syncNamed, alongside translationNot documented by the vendorOnly HeyGen answers

Inclusion rule. Two entries are given a page together when at least one column puts them at opposite grades of answer. Pairs that agree on every column, or that are both blank throughout, do not get a page. Order. Fixed field order, identical on every side-by-side page.

1A voice object and a character object fail differently

An id can be passed wrongly, so a production has to keep a record of which id belongs to which character. A character object cannot be passed wrongly, because there is nothing to pass, which removes the commonest continuity error rather than guarding against it.

The trade is flexibility. A voice with an id can be reused across series, shared deliberately between characters, or retired without touching a character. A voice welded to a character cannot be moved, and nothing published says how to reconstruct it if the character is lost.

2One publishes how the voice was made, the other does not

A clone instant from one recording or professional from twenty minutes or more is a figure a production can plan a session around. An attached voice with no stated origin means the first generation of each character is a casting session with one candidate.

So persistence and selection come apart here. Both entries solve the harder problem of keeping a voice; only one tells a production what the voice will be before it commits a season to it.

3Their other columns barely overlap at all

One publishes a language count attached to translation and names lip-sync in the same sentence. The other publishes no language, no voice source and nothing about mouths, and documents an audio-to-video route nobody else here has.

Read together they are a reminder that a shared answer in one column predicts very little about the rest of a row, which is why this register keeps five columns rather than a score.

4Each of them on its own

The column this pair was chosen for is per-character binding, and each entry has a page of its own on it. The full row for either, all five columns with the wording behind each cell, is on its model note.

  • Voice source
    A voice clone is instant from one recording or professional from twenty minutes or more, and is then passed as a voice idtwo enrolment routes with a thresholdHeyGen, API quick start / recorded 2026-09-22
  • Languages
    Video translation is documented for thirty or more languages, with voice cloning and lip-synccounted on the translation endpointHeyGen, API quick start / recorded 2026-09-22
  • Per-character binding
    Voices are attached to character Elementstied to the element systemLTX Studio / recorded 2026-09-12
  • Audio source
    Joint audio and video generation, plus audio-to-video, on LTX-2.5two directionsLTX Studio / recorded 2026-09-12

5Sources

Every cell above is read from the documentation each vendor publishes, on the dates carried by the two model notes. The whole register on one column is on the field note; entries grouped by the shape of their answer are on the routes. Other pairs: MiniMax and Wan 3.0, Sora 2 and Luma Ray.