Comparison

Cutroom vs InVideo: generated supply, or a face a stranger believes

InVideo builds a whole video from a prompt, a stock library and a generated voice. Cutroom cuts a take you filmed of your own face. It hands back a finished vertical ad with the captions already burned in. This page covers what generated supply is worth, what it cannot buy, and what one cut costs here. By the end you will know which of the two your next ad needs.

take of your own face, in
1
credits per batch
100
credits per output minute
20

InVideo fills an empty project in minutes. It cannot put a believable person in it

Cold start is the real problem it solves. No footage, no camera, no script you like. Prompt-to-video hands back a draft in minutes. Reacting to a draft beats staring at an empty project.

It is built for volume in a category where volume is the point. Faceless channels, listicles, explainers, product roundups. Anything where the viewer does not need to know who is talking.

The template library is a real asset rather than a filler feature. Someone with no design instinct picks a look that already works.

What it hands back is supply. A cold feed asks a second question. Does a stranger believe the person saying this. Cutroom starts at that question, from a face that already exists.

Prompt-to-video solves supply. A cold feed is asking about belief

Prompt-to-video turns text you have into video you do not have, at close to zero marginal cost. That is worth money whenever nothing exists yet.

What it cannot convert is credibility. A stock shot of a smiling person is not evidence. A generated voice reading a claim is not a person making a claim.

A viewer works that out faster than any of us would like. For a lot of content it does not matter. The viewer came for the information and does not care who supplies it.

It starts mattering the second the video asks for money. Direct response lives or dies on whether somebody believes the person saying it.

That is the whole distinction. One tool fixes how much video exists. The other fixes whether a stranger trusts forty seconds of it.

This is the whole editor

Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.

CLIPS · 5I have thisexact conversationeverysingle week. Somebody sits down and says,oh yeah, I takecinnamonevery day.And honestly, doc, I have no idea if it works.So let me tell you what isin that capsule.a clip lands on these wordscut from the editTAKING CINNAMONEVERY DAY?is it doing anythingHeadlineMusicCaptionsTHIS VIDEOLength25.0sClips5Words removed18Export video

Cutroom cuts a read you actually gave, and you steer it on the transcript

You film yourself saying the thing, up to three minutes, on a phone. Cutroom transcribes it and drafts the cut. B-roll lands on the phrases that need showing.

Captions get styled. An opening headline gets drafted. A 9:16 MP4 comes back. Direction then happens on the words rather than on a playhead.

Highlight a phrase and a clip lands over exactly those words. Delete a line and the cut rebuilds. Change the pace and the whole piece re-cuts.

Coverage is capped at 40 percent on chill, 45 on normal and 52 on fast. Your face never disappears behind stock.

There is an avatar module for the days you cannot film. A photo plus a voice-over becomes a lip-synced presenter. It then gets cut like any other take. Avatar render is 320 credits a minute and voice-over is 70.

Cutting a take you filmed costs 20 credits a minute. The gap between 20 and 320 is the product pointing at the path it was built for.

  • Your own face and your own read, not a licensed performer
  • One take in, one finished vertical ad out, captions burned in
  • Six caption packs and five clip transitions, cut through to sweep
  • 300 free credits over 7 days, no card, enough for two finished ads

From a phone in a kitchen to a 9:16 MP4

The 100 credit batch covers steps two and three. Step five is 20 credits a finished minute, every time after that.

  1. Film one take

    3 min ceiling, phone is fine

  2. Transcribe

    Word-level timing

  3. Draft the cut

    Headline, b-roll, captions

  4. Direct by reading

    Delete, highlight, re-pace

  5. Export

    9:16 MP4, 20 CR a minute

A new prompt changes four variables. A new cut changes one

Both tools produce a vertical video with captions, so on a feature grid they look adjacent. One fills an empty timeline from text. The other decides what survives from a performance that already happened.

When a generated ad fatigues, the next move is a new prompt. New script, new visuals, new pacing, new voice, all changed at once. When it performs better, nobody can say which change did it.

When a filmed ad fatigues, the next move is a different cut of the same read. Start on sentence three. Move faster. Put the footage on the claim instead of the greeting. One item moved, and you know which.

The read is the control variable and the edit is what you test. Transcription and the director pass are paid once at 100 credits. Every further version costs only the export, at 20 credits a finished minute.

Five edits of one take is an afternoon rather than a budget conversation. Published benchmarks put winning creatives at roughly 5 to 8 percent, per Motion's analysis of 550,000 plus Meta ads. Cheap, readable attempts are the only lever you own.

How many variables move between version one and version two

A test with four moving parts cannot tell you which part won. That is what makes a cheap re-cut worth more than a cheap re-prompt.

  • New prompt: script, visuals, pace, voice4 variables
  • New cut: one of them, on purpose1 variable

Your own face, cut five ways in an afternoon, for 20 credits a version

That is what this side is for. Coaching, supplements, skincare, services. Wherever the face is the mechanism the offer runs on, a filmed read out-earns a generated one and keeps earning.

Now the facts a buyer needs. There is no text-to-video path here and no script input. No horizontal or square. No source over three minutes. No timeline to escape into when the draft gets a beat wrong.

The avatar module is the one exception and it is narrow. It needs a photo you supply plus a voice-over. There is no library of licensed performers to browse.

InVideo is the buy for faceless content, or for many videos from text you already wrote. Cutroom is the buy when somebody will film.

A middle path exists and it is the one we would recommend to a friend. Use generated video to find which angle holds attention in a cold feed. Then film the winner yourself and cut it here. Testing text is cheap. Testing a shoot is not.

Run the numbers on your own volume. A one-minute ad here is about 120 credits. Basic carries 2,500 credits for $39.99 a month. Read InVideo's pricing page at the source, because a table written by a competitor is not evidence.

InVideo and Cutroom, row by row

Two rows belong to InVideo, and they matter if nobody will film. The other six are what a paid feed is judging.

InVideoCutroom
Video without filming anythingPrompt in, finished video outYou film a take first
Many videos from text you haveBuilt for exactly thatNo text input anywhere
Captions burned into the exportStyled over generated speechSix packs, from your own read
B-roll on phrases you nameStock chosen by the promptHighlight words, clip lands
Your own face and readStock and generated narrationYour camera, your voice
Edit driven by the transcriptPrompt again to change itDelete a line, cut rebuilds
Same read, different editsA new prompt changes everything20 credits per output minute
Speaker stays on screenStock can fill the frameCoverage capped at 52 percent

Questions people ask

Can Cutroom write the script for me?
No. It works from what you said. It drafts an opening headline from your read and decides what to cut, but the body of the ad is your own words.
Does Cutroom use stock footage?
Yes, as b-roll over specific phrases. You can swap any clip or search millions of free ones. Stock is the supporting layer, capped at 52 percent, never the whole video.
How long can my upload be?
Three minutes. That ceiling is deliberate. This is built for the read that becomes a short vertical ad, not for a webinar.
Is it right for every ad?
It is built for ads where a person is the reason someone believes the claim. If you have no footage and nobody who will film this quarter, a prompt-to-video tool is closer to your problem today.

You can generate almost everything now. A face a stranger believes is the one thing still worth filming. Cutroom is what turns that read into a finished ad.

Start with one take300 free credits · no card · cancel anytime