Hedra: audio or an id travels with each call, and that is all
Each call carries either a track or a voice id, and neither is described as attaching to a character. The consistency of a cast therefore rests on whatever a production keeps in its own notes, not on anything the platform holds. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| Speech can be generated inline by naming a voice id from the voices endpoint | Whether that id is guaranteed to stay available |
| Audio videos are driven by audio, and the character lip-syncs to the audio provided | How a supplied performer is kept consistent across episodes |
| The audio normally determines the video length | Whether two calls for the same character can be made to match in pace |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1When the audio is supplied, persistence is a filing problem
An uploaded recording carries its own consistency: the same performer, recorded in the same room, sounds the same in episode nine as in episode one. The platform is not being asked to hold anything, which is why the blank here is less alarming than the identical blank on a model that invents the voice.
What replaces a binding mechanism is a production discipline. The performer, the room, the microphone and the processing chain all have to be recorded and repeated, and none of that is visible in a returned file.
2The inline route is where the gap actually bites
Naming a voice id instead of uploading moves the voice inside the platform, and now nothing published says whether that id is durable, whether the catalogue is versioned, or what happens if a voice is withdrawn. A character built on an inline voice is built on an identifier with no stated lifetime.
The register records the absence rather than assuming durability. Voice catalogues behind endpoints do change, and a series is a long time to depend on one that nobody promised to keep.
3Others whose documentation never addresses persistence
Seven more entries reach this column empty. Several of those make the sound themselves, which makes the gap sharper than it is here, where a production can simply keep the recording.
- Luma Ray — not documented by the vendor.
- SceneMixer — not documented by the vendor.
- sync-3 — not documented by the vendor.
- Synthesia — not documented by the vendor.
- Veo — not documented by the vendor.
- Vidu — not documented by the vendor.
- Wan2.2-S2V — not documented by the vendor.
- Voice sourceSpeech can be generated inline instead of uploaded, by naming a voice id drawn from the voices endpointa catalogue behind an endpoint
- Audio sourceAvatar videos are driven by audio, and the character in the image will lip-sync and move to the audio providedhanded to the model rather than made by it
- Shot length set by the takeThe audio normally determines the video length, and omitting the duration follows the source audio
4Sources
Read from the Character 3 model page and avatar video guide at hedra.com on 2026-09-22. The same column across every entry is on per-character binding; everything this vendor publishes about speech is on Hedra. What counts as documented is on how read.