PixVerse and D-ID: one id, two supply chains
Two entries put an identifier on every generation. One passes a speaker id covering both built-in voices and custom ones built from a sample. The other passes a voice id drawn from five named speech providers. As of 2026-09-22.
| Field | PixVerse | D-ID | Where they part |
|---|---|---|---|
| Audio source | Supplied, or read from text | From a script, or a supplied url | Different answers |
| Voice source | A sample, or a built-in voice | A voice id, or a recording by url | Different answers |
| Per-character binding | On the request | On the request | Same answer |
| Languages | Multiple, none of them named | A language field, no list | Different answers |
| Lip-sync | Named as the endpoint's purpose | Not documented by the vendor | Only PixVerse answers |
Inclusion rule. Two entries are given a page together when at least one column puts them at opposite grades of answer. Pairs that agree on every column, or that are both blank throughout, do not get a page. Order. Fixed field order, identical on every side-by-side page.
1A stable id is most of what persistence needs
Neither platform holds a pairing between a character and a voice, and neither needs to for the voice itself to stay put. An id cannot drift, so a production that writes it down gets consistency without storing audio anywhere.
What neither offers is a check. Passing the wrong id for a character produces a perfectly valid call and a different person, and nothing in a returned file records which id spoke. The shot list is the authoritative record either way.
2Where the ids come from changes the dependency
One platform's ids are its own, covering a catalogue it controls and clones a user built. The other's belong to five outside services, each with its own deprecation schedule, licence terms and language codes.
Five suppliers is a wide net and a wide surface. A voice withdrawn upstream takes a character with it, and the talk reference is not the document that would announce that.
3One id also decides whose mouth moves
One of these describes its identifier as the lip-sync speaker on the request, which quietly says the same handle governs the voice and the face. That is half of the two-speaker problem answered by a routing detail.
The other never mentions a mouth. Two products passing what looks like the same field, and only one of them has said what the field does to the picture.
4Each of them on its own
The column this pair was chosen for is per-character binding, and each entry has a page of its own on it. The full row for either, all five columns with the wording behind each cell, is on its model note.
- PixVerse on per-character binding — on the request.
- D-ID on per-character binding — on the request.
- PixVerse, all five fields — read from the speech and lip sync guide.
- D-ID, all five fields — read from the create a talk reference.
- Per-character bindingThe chosen speaker id is passed as the lip-sync speaker on the generation requesta speaker id per call
- Voice sourceText to speech supports both built-in voices and custom voices created from user-provided sample audiocatalogue or supplied sample
- Length per callAudio and video are each capped at sixty seconds and one hundred megabytes
- Per-character bindingThe voice is a voice id selected from the list of available voices, with an optional language field beside ita voice id on every request
- Voice sourceFive speech providers are named for the voice: Microsoft, ElevenLabs, Amazon, Google and Azure OpenAIoutside catalogues named on the page
5Sources
Every cell above is read from the documentation each vendor publishes, on the dates carried by the two model notes. The whole register on one column is on the field note; entries grouped by the shape of their answer are on the routes. Other pairs: Sora 2 and Wan 3.0, Runway and Luma Ray.