D-ID generates talking video from a still image plus a script or audio, and the documentation is clear about how to do that. The API surface is small, the integration path is short, and someone competent can have working output inside an afternoon. The material is also more honest than I expected about quality, meaning it acknowledges the uncanny valley, explains which source images produce better results and does not pretend the output is indistinguishable from a filmed person. That candour is worth something in a category that usually oversells.
The interactive and streaming material is decent. Building an avatar that responds in something close to real time is a more demanding engineering problem than generating a clip, and the documentation covers the latency considerations and the integration pattern properly. If your use case is a presenter for internal training, a multilingual product explainer or an interactive kiosk, the material gets you there. The ethics are handled the way this category always handles them, which is with a reference to terms of service and a note that you should have rights to the likeness you use.
That is a legal position rather than guidance. Generating video of a person saying words they never said is the exact capability that makes this category consequential, and the questions that follow are not hard to anticipate. What consent looks like when someone agrees to a likeness once and the model outlives the agreement. Whether an employee who appears in company training material can withdraw.
How you disclose to a viewer that what they are watching was synthesised. What you do when someone asks for a likeness of a person who cannot consent. None of this is addressed and it is not unreasonable to expect a page on it. Disclosure in particular deserves attention because it costs the vendor nothing and protects everyone.
A short section on watermarking, on stating clearly in the video that the presenter is synthetic and on the emerging regulatory expectations around that would be genuinely useful, and its absence suggests nobody wanted to raise the question with customers. Output quality dependence on the source image is underexplained relative to how much it matters. Lighting, angle, resolution and expression in the original photograph drive the result more than any setting you can adjust afterwards, and the guidance on choosing a source is thinner than it should be. People conclude the model is poor when the actual problem is a bad input.
Pricing is per minute of generated video and the arithmetic is unkind once retakes enter the picture. Script changes, pronunciation problems and quality rejects all cost the same as usable output, so a finished five minute video may have consumed fifteen. Budget accordingly. Three point two.
Competent technical documentation for a capability that works reasonably well within its limits, marked down for engineering honesty paired with ethical silence on exactly the questions this technology raises.