Hedra: upload a track, or name a voice from the list
Two routes are documented. Audio can be uploaded, which makes the performance a production decision, or speech can be generated inline by naming a voice id drawn from the voices endpoint. The track is required either way. As of 2026-09-22.
| What the documentation settles | What it leaves to a take |
|---|---|
| Speech can be generated inline by naming a voice id from the voices endpoint | How many voices that endpoint holds, and in which languages |
| Audio videos are driven by audio, and the character lip-syncs to the audio provided | Whether an uploaded take outperforms an inline voice on the same line |
| Omitting the duration makes the result follow the source audio | Whether an inline voice produces a predictable length |
Inclusion rule. Only statements found on the vendor page named in the caption, and only ones bearing on this one field. Nothing here is filled in from generated output. Order. In the order the statements arrive on the vendor's page.
1Both routes end in the same required input
What makes this row different from a catalogue-only entry is that the catalogue is an alternative to a recording rather than the only option. The model needs a waveform; the vendor will make one if a production does not bring one. That is a convenience with a consequence, because the two routes give a director very different amounts of control.
Uploading means the read was auditioned by a human before any money was spent on frames. Naming a voice id means the read arrives with the render, and a bad one costs a generation to discover.
2An endpoint is a catalogue without a page
Putting the voices behind an endpoint rather than in the documentation means the inventory is current and unreadable at the same time. Current because it cannot go stale against the code; unreadable because a producer deciding whether to use the platform will not call it.
For this register the consequence is that the cell records a route rather than a range. How many voices, in which languages, and how they were made are all outside what the avatar guide publishes.
3The others that hand a production a catalogue
Four more entries answer with something to choose from. This is the only one where choosing is optional because the model will accept a finished recording instead.
- D-ID — a voice id, or a recording by url.
- PixVerse — a sample, or a built-in voice.
- SceneMixer — a preset library, or the user's own recordings.
- Synthesia — a catalogue voice, or a cloned one.
- Voice sourceSpeech can be generated inline instead of uploaded, by naming a voice id drawn from the voices endpointa catalogue behind an endpoint
- Audio sourceAvatar videos are driven by audio, and the character in the image will lip-sync and move to the audio providedhanded to the model rather than made by it
- Shot length set by the takeThe audio normally determines the video length, and omitting the duration follows the source audio
4Sources
Read from the Character 3 model page and avatar video guide at hedra.com on 2026-09-22. The same column across every entry is on voice source; everything this vendor publishes about speech is on Hedra. What counts as documented is on how read.