Comparison
Cutroom vs Runway: one makes the shots, one makes the sequence
Runway generates footage that never existed, better than most people could shoot it. Cutroom generates none. It takes a read of yours that already exists and arranges it into a finished ad. This page covers what a shot is worth, what a sequence is worth, and where a generated clip drops into ours. By the end you will know which half you are missing.
- the finished thing Cutroom returns
- 9:16 MP4
- transitions: cut, whip, punch, glitch, sweep
- 5
- upload ceiling
- 3 min
Runway makes the shot nobody could film. It has no opinion about the order
Making a shot that does not exist and cannot be filmed. A product floating in a space you cannot reach. A texture. An impossible camera move. That is the job and it does it.
Add rotoscoping, inpainting and clean background replacement. That set of tools replaced a category of post work which used to cost serious money and take a week.
For brand films and music video it is in a different league to everything else on this page. Where the image is the point, we do not compete.
An ad is a different object. It is a sequence of claims in an order, with a person behind them. Deciding that order is the job Cutroom does. No image model has a view on it.
An ad is a sequence of claims, and no image model has an opinion about the order
A direct-response ad is not one beautiful image. It is a sequence of claims in a specific order. The unit is the sentence, not the frame.
Generating a great shot leaves that problem untouched. You still decide which sentence opens and which one dies. You still decide where the cut lands and how long the caption holds.
Those are timing decisions about speech. No image model has a view on any of them. It cannot tell you the middle sags at fourteen seconds, because it never heard the middle.
So a generative tool bought to make ads leaves you with impressive clips and no ad. The clips are inputs. Something has to build the argument around them. Until now that was an evening on a timeline.
This is the whole editor
Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.
Cutroom is the assembly half, and you direct the sequence by reading it
You film one take of up to three minutes. Cutroom transcribes it and drafts a full cut. B-roll goes over the phrases that need showing. A headline goes on the opening. Captions get styled. A 9:16 MP4 comes out.
The interface after that is the transcript and nothing else. Highlight a phrase and footage lands over exactly those words. Delete a line and the cut rebuilds. Change the pace and the whole video re-cuts.
Swap the clip in any spot, or search millions of free ones. Drop a generated clip in as the b-roll for a named phrase. It behaves like any other footage, under the same coverage cap.
No timeline and no keyframes. Those exist to specify a shot. Specifying shots is not the job here. The compiler owns every boundary in time.
So a pace change re-times the video around the same clips. You never re-place a single one. The shot list stays valid after you change your mind about the opening.
- Coverage caps by pace: 40 percent chill, 45 normal, 52 fast
- Eleven editor actions, all of them on the words
- 100 credits a batch, 20 credits per output minute
- Music module: describe a style, get a track for 20 credits
Where a generated clip enters the pipeline
A generated shot arrives at step four, as the b-roll for one named phrase. Nothing upstream of it changes.
Film the read
3 min ceiling
Transcribe
Word-level timing
Read the shot list
3 or 4 phrases
Swap in your clip
Under the coverage cap
Export
9:16 MP4, 20 CR a minute
Film first and you generate four shots instead of forty
Order of operations decides the budget here. Film the read first, because the read decides everything else.
Once it is transcribed you can see which phrases are asking for evidence. A number. A before and after. A process. Those phrases are your shot list. It is short.
Most ads need three or four moments covered, not thirty. Now you know what to generate. You know it before you have spent anything per attempt.
Coverage stays under 40, 45 or 52 percent depending on pace. The visuals support the person making the argument instead of replacing them.
Generating first and cutting second gives you good shots in search of a video. Cutting first gives you a short shot list and an ad that already works without any of it.
How many shots a forty-second ad actually uses
The coverage caps decide this, not taste. At 52 percent on fast, there is not room for thirty clips in forty seconds.
- Clips generated before a script existsabout 30
- Phrases in the read asking for evidence4
Generated shots inside a finished ad, and the facts about the rest
The best version of both tools is sequential. Cut first, generate second. Let the shot list come out of the transcript. Three or four clips, dropped onto named phrases, inside an ad that already worked without them.
Now the facts a buyer needs. Cutroom generates no footage. There is no effects browser, no keyframing, no compositing, no colour grading. Five clip transitions and six caption packs is the whole visual vocabulary.
It exports 9:16 and nothing else. It refuses a source over three minutes at upload. It expects one person talking to a camera.
The avatar module is the way around the camera. It costs 320 credits a minute against 20 for cutting a take you filmed. That gap is a signal rather than an oversight.
Runway is the buy for imagery that does not exist, and for visual effects work. Cutroom is the buy when you have a person, a phone and something worth saying, and the missing piece is the edit.
On cost, a batch is 100 credits and export is 20 per output minute. Read Runway's own page for theirs.
Runway and Cutroom, row by row
Two rows go to Runway, and they are the imagery ones. The other six are what turns clips into an ad.
| Runway | Cutroom | |
|---|---|---|
| Shots that were never filmed | Generates footage from nothing | You supply the take |
| Visual effects and compositing | Rotoscope, inpaint, replace | No effects browser at all |
| Coverage capped so the speaker stays | Nothing to cap, no speaker | 40, 45 or 52 percent by pace |
| Takes your generated clip as b-roll | Output, not a host | Swap it onto a named phrase |
| A finished ad, not a shot | Clips, not a sequence | One take in, MP4 out |
| Decides the order of claims | No opinion about sentences | You direct on the transcript |
| Captions burned into the export | Add them somewhere else | Six packs, fully restyleable |
| Cost of the next cut | Generate again from scratch | 20 credits per output minute |
Questions people ask
- Can Cutroom generate video from a text prompt?
- No. It cuts a take you filmed. The only generation in the product is the avatar module, the music module, and the caption and headline drafting.
- Can I use generated clips as b-roll in Cutroom?
- Yes. You can swap the clip in any spot, so a generated shot sits over a specific phrase. It obeys the same coverage caps as any other footage.
- Which one is better for social ads?
- For talking-head direct response, Cutroom, because the hard part is timing speech. For visually driven brand work, Runway, because the hard part is the image.
- Is it right for every ad you run?
- It is built for ads made of a person talking. There is no effects browser and no keyframing. If motion design is what your ads are made of, a generator is closer to the job.
Runway makes the shot. Cutroom is what turns your read plus that shot into a 9:16 file you can run tomorrow.