Cutting AI footage like film
Generative models hand you twenty seconds at a time. Everything that makes those seconds read as a film happens afterwards, in the edit.
The honest constraint of generative video in 2026 is duration. A model will give you a beautiful twenty-second clip and then stop. It will not remember the grade it just used, the lens it implied, or where the light was coming from. Ask for a continuation and you get a cousin, not a next shot.
Most work made this way betrays that constraint immediately. You can hear it: a cut, a beat of black, a new clip that starts from a standstill. Twenty seconds, breath, twenty seconds. It reads as a folder of outputs rather than a film, and no amount of grading fixes a structural problem.
Continuity is a decision made before generation
When we assembled THE ROOT — six clips, two minutes, one continuous idea — the sequencing existed before the first prompt. We wrote the film as six beats of an argument: ground broken, structure raised, system taking hold, growth outpacing the plan, the city breathing, the root beneath it all. Each beat had a stated end state, and the following beat opened on that state.
That means the last frame of one clip and the first frame of the next are describing the same moment from the same distance. When they are, a cross-dissolve does not read as a transition. It reads as the camera continuing to run.
What we actually do to the clips
- Upscale every clip to a single resolution before anything else, so no shot arrives softer than its neighbour.
- Trim the first and last beats of each clip — models tend to settle in the middle and drift at the edges.
- Overlap the joins rather than butting them, so the eye is given motion to follow across the cut.
- Grade the whole assembly once, at the end, as a single piece. Grading clip by clip guarantees six different films.
- Cut to a duration the story earns, not to the duration the model produced.
A generated clip is not a shot. It becomes a shot when it is cut against another one.
Why the black gaps were the real problem
Early versions of DUST and RUST had pauses between their clips. It was the most technically correct thing to do and the most damaging. A pause tells the viewer the piece is a compilation, and once they are counting clips they have stopped watching the film. Removing the gaps changed nothing about the images and everything about how the work is read.
The lesson generalises past AI footage: the material is rarely the bottleneck. The structure around the material is. This is the same reason we do not sell generation as a service. Anyone can produce twenty seconds. Producing two minutes that hold together is direction, and direction is the part that is hard to buy.
How this shows up in client work
On commercial engagements the practical consequence is scheduling. We do not treat generation as production and the edit as post. They are one phase, run by one person, with the assembly re-cut every day as new material lands. It costs less than the traditional split and it is the only way we know to keep a piece coherent when the material is arriving in twenty-second pieces.
