Hedra generates video of a character speaking from a still image and an audio track, and it does it well enough that the results are frequently convincing. The documentation reflects a company that has decided its centre of gravity is developers. Of roughly a hundred and thirty pages, the great majority sit in the developer platform, the API reference and the model catalogue, with a much smaller set covering getting started and content creation for people who just want to use the product. As API documentation it is good.
Endpoints are properly specified, the job and webhook model is explained, the MCP integration is documented, and somebody building an application on this has what they need. Credit mechanics are also laid out clearly, which I appreciate more than it might sound. Generative video burns credits fast, iteration is where the cost lives, and vendors who explain the meter before you hit it are being straight with you. The free plan exists, so you can assess quality before paying, and pricing runs through Basic, Creator and Professional tiers in the usual shape.
The company raised a thirty two million dollar Series A led by a16z in May 2025, so the platform is reasonably well resourced. Now the part that I cannot let pass. This product animates a face from a photograph and makes it speak words supplied by whoever is operating it. That is, in plain terms, a tool for producing footage of a person saying something they did not say.
There are entirely legitimate uses, including your own likeness, licensed actors, fictional characters and consenting participants, and the documentation covers none of that distinction in any depth. There is no meaningful guidance on obtaining permission from the person whose face you are using, no discussion of what that permission should cover, nothing about jurisdictions where likeness and biometric data are regulated, and nothing about disclosure to the eventual audience. A hundred and thirty pages of documentation with essentially nothing on the single most consequential decision a user of this product makes is a choice, and it is the wrong one. The other gap is craft.
The developer weighting means creators get comparatively little, and there is real skill in making this output convincing. Source image quality, framing, how the audio is recorded, where the uncanny artefacts appear and how to avoid them. That knowledge exists inside the company and has not been written down for the people making things. What you get instead is thorough instruction on calling the endpoint and rather less on producing something worth generating.
Three. Solid, well maintained developer documentation for a genuinely capable model, marked down heavily for treating consent as somebody else's problem and for leaving creators with far less than developers.