Comparison

Cutroom vs Runway: one makes the shots, one makes the sequence

Runway generates footage that never existed, better than most people could shoot it. Cutroom generates none. It takes a read of yours that already exists and arranges it into a finished ad. This page covers what a shot is worth, what a sequence is worth, and where a generated clip drops into ours. By the end you will know which half you are missing.

the finished thing Cutroom returns
9:16 MP4
transitions: cut, whip, punch, glitch, sweep
5
upload ceiling
3 min

Runway makes the shot nobody could film. It has no opinion about the order

Making a shot that does not exist and cannot be filmed. A product floating in a space you cannot reach. A texture. An impossible camera move. That is the job and it does it.

Add rotoscoping, inpainting and clean background replacement. That set of tools replaced a category of post work which used to cost serious money and take a week.

For brand films and music video it is in a different league to everything else on this page. Where the image is the point, we do not compete.

An ad is a different object. It is a sequence of claims in an order, with a person behind them. Deciding that order is the job Cutroom does. No image model has a view on it.

An ad is a sequence of claims, and no image model has an opinion about the order

A direct-response ad is not one beautiful image. It is a sequence of claims in a specific order. The unit is the sentence, not the frame.

Generating a great shot leaves that problem untouched. You still decide which sentence opens and which one dies. You still decide where the cut lands and how long the caption holds.

Those are timing decisions about speech. No image model has a view on any of them. It cannot tell you the middle sags at fourteen seconds, because it never heard the middle.

So a generative tool bought to make ads leaves you with impressive clips and no ad. The clips are inputs. Something has to build the argument around them. Until now that was an evening on a timeline.

This is the whole editor

Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.

CLIPS · 5I have thisexact conversationeverysingle week. Somebody sits down and says,oh yeah, I takecinnamonevery day.And honestly, doc, I have no idea if it works.So let me tell you what isin that capsule.a clip lands on these wordscut from the editTAKING CINNAMONEVERY DAY?is it doing anythingHeadlineMusicCaptionsTHIS VIDEOLength25.0sClips5Words removed18Export video

Cutroom is the assembly half, and you direct the sequence by reading it

You film one take of up to three minutes. Cutroom transcribes it and drafts a full cut. B-roll goes over the phrases that need showing. A headline goes on the opening. Captions get styled. A 9:16 MP4 comes out.

The interface after that is the transcript and nothing else. Highlight a phrase and footage lands over exactly those words. Delete a line and the cut rebuilds. Change the pace and the whole video re-cuts.

Swap the clip in any spot, or search millions of free ones. Drop a generated clip in as the b-roll for a named phrase. It behaves like any other footage, under the same coverage cap.

No timeline and no keyframes. Those exist to specify a shot. Specifying shots is not the job here. The compiler owns every boundary in time.

So a pace change re-times the video around the same clips. You never re-place a single one. The shot list stays valid after you change your mind about the opening.

  • Coverage caps by pace: 40 percent chill, 45 normal, 52 fast
  • Eleven editor actions, all of them on the words
  • 100 credits a batch, 20 credits per output minute
  • Music module: describe a style, get a track for 20 credits

Where a generated clip enters the pipeline

A generated shot arrives at step four, as the b-roll for one named phrase. Nothing upstream of it changes.

  1. Film the read

    3 min ceiling

  2. Transcribe

    Word-level timing

  3. Read the shot list

    3 or 4 phrases

  4. Swap in your clip

    Under the coverage cap

  5. Export

    9:16 MP4, 20 CR a minute

Film first and you generate four shots instead of forty

Order of operations decides the budget here. Film the read first, because the read decides everything else.

Once it is transcribed you can see which phrases are asking for evidence. A number. A before and after. A process. Those phrases are your shot list. It is short.

Most ads need three or four moments covered, not thirty. Now you know what to generate. You know it before you have spent anything per attempt.

Coverage stays under 40, 45 or 52 percent depending on pace. The visuals support the person making the argument instead of replacing them.

Generating first and cutting second gives you good shots in search of a video. Cutting first gives you a short shot list and an ad that already works without any of it.

How many shots a forty-second ad actually uses

The coverage caps decide this, not taste. At 52 percent on fast, there is not room for thirty clips in forty seconds.

  • Clips generated before a script existsabout 30
  • Phrases in the read asking for evidence4

Generated shots inside a finished ad, and the facts about the rest

The best version of both tools is sequential. Cut first, generate second. Let the shot list come out of the transcript. Three or four clips, dropped onto named phrases, inside an ad that already worked without them.

Now the facts a buyer needs. Cutroom generates no footage. There is no effects browser, no keyframing, no compositing, no colour grading. Five clip transitions and six caption packs is the whole visual vocabulary.

It exports 9:16 and nothing else. It refuses a source over three minutes at upload. It expects one person talking to a camera.

The avatar module is the way around the camera. It costs 320 credits a minute against 20 for cutting a take you filmed. That gap is a signal rather than an oversight.

Runway is the buy for imagery that does not exist, and for visual effects work. Cutroom is the buy when you have a person, a phone and something worth saying, and the missing piece is the edit.

On cost, a batch is 100 credits and export is 20 per output minute. Read Runway's own page for theirs.

Runway and Cutroom, row by row

Two rows go to Runway, and they are the imagery ones. The other six are what turns clips into an ad.

RunwayCutroom
Shots that were never filmedGenerates footage from nothingYou supply the take
Visual effects and compositingRotoscope, inpaint, replaceNo effects browser at all
Coverage capped so the speaker staysNothing to cap, no speaker40, 45 or 52 percent by pace
Takes your generated clip as b-rollOutput, not a hostSwap it onto a named phrase
A finished ad, not a shotClips, not a sequenceOne take in, MP4 out
Decides the order of claimsNo opinion about sentencesYou direct on the transcript
Captions burned into the exportAdd them somewhere elseSix packs, fully restyleable
Cost of the next cutGenerate again from scratch20 credits per output minute

Questions people ask

Can Cutroom generate video from a text prompt?
No. It cuts a take you filmed. The only generation in the product is the avatar module, the music module, and the caption and headline drafting.
Can I use generated clips as b-roll in Cutroom?
Yes. You can swap the clip in any spot, so a generated shot sits over a specific phrase. It obeys the same coverage caps as any other footage.
Which one is better for social ads?
For talking-head direct response, Cutroom, because the hard part is timing speech. For visually driven brand work, Runway, because the hard part is the image.
Is it right for every ad you run?
It is built for ads made of a person talking. There is no effects browser and no keyframing. If motion design is what your ads are made of, a generator is closer to the job.

Runway makes the shot. Cutroom is what turns your read plus that shot into a 9:16 file you can run tomorrow.

Start with one take300 free credits · no card · cancel anytime