Luma is two companies wearing one coat. The original work was neural capture, meaning turning photographs and video of a real scene into a navigable 3D reconstruction, and it was genuinely novel. The thing everyone knows now is Dream Machine, which generates video from text and images. Both are documented, and the documentation behaves as though each is unaware of the other, which is a shame because the interesting future is obviously where they meet.
The keyframe workflow is the strongest practical material. Rather than describing a clip and accepting whatever arrives, you supply a start image and an end image and let the system work out the motion between them. That is a fundamentally more controllable way to work and it is closer to how anyone with an animation background thinks. It also sidesteps the biggest weakness of text to video, which is that language is a terrible instrument for specifying visual motion.
The documentation explains this well and it is the single technique I would want a newcomer to learn first. Camera motion guidance is decent without reaching the specificity of the better competitors, and motion coherence is where the material gets vague. Every video model has failure modes, things that morph, hands that reorganise themselves, objects that pass through each other, and the honest way to document that is to show it. Luma describes limitations in careful general language and does not show a single failure, which leaves users to discover the boundaries by spending credits.
A page of here is what breaks and why would be worth more than another prompting tip. The capture documentation is the part I keep returning to. Reconstructing a real space from a phone walkthrough has obvious uses in property, in previsualisation, in preserving a location, in giving an artist a real environment to work in, and the material is technically solid on how to shoot for good results, which is where most people fail. It is also almost invisible next to the video product, and I suspect a lot of people never find it.
That is a marketing decision rather than a documentation failure, but the effect is that the more distinctive capability gets less attention. What is missing across both is anything about after. You have a four second clip. Now what.
Sequencing, cutting, matching shots, sound, the fact that a video is made of many clips that have to agree with each other. The material stops at generation, and generation is the easy part now. Anyone producing something longer than a demo runs into assembly immediately and gets no help. The free tier is small enough that you cannot build intuition without paying, which for a tool where intuition is the whole skill is a real barrier to fair evaluation.
Three point four. Technically credible material with the most controllable workflow in AI video, let down by a reluctance to show its own failure modes and by treating its two halves as separate businesses when the combination is the most interesting thing here.