Comparison
Cutroom vs InVideo: generated supply, or a face a stranger believes
InVideo builds a whole video from a prompt, a stock library and a generated voice. Cutroom cuts a take you filmed of your own face. It hands back a finished vertical ad with the captions already burned in. This page covers what generated supply is worth, what it cannot buy, and what one cut costs here. By the end you will know which of the two your next ad needs.
- take of your own face, in
- 1
- credits per batch
- 100
- credits per output minute
- 20
InVideo fills an empty project in minutes. It cannot put a believable person in it
Cold start is the real problem it solves. No footage, no camera, no script you like. Prompt-to-video hands back a draft in minutes. Reacting to a draft beats staring at an empty project.
It is built for volume in a category where volume is the point. Faceless channels, listicles, explainers, product roundups. Anything where the viewer does not need to know who is talking.
The template library is a real asset rather than a filler feature. Someone with no design instinct picks a look that already works.
What it hands back is supply. A cold feed asks a second question. Does a stranger believe the person saying this. Cutroom starts at that question, from a face that already exists.
Prompt-to-video solves supply. A cold feed is asking about belief
Prompt-to-video turns text you have into video you do not have, at close to zero marginal cost. That is worth money whenever nothing exists yet.
What it cannot convert is credibility. A stock shot of a smiling person is not evidence. A generated voice reading a claim is not a person making a claim.
A viewer works that out faster than any of us would like. For a lot of content it does not matter. The viewer came for the information and does not care who supplies it.
It starts mattering the second the video asks for money. Direct response lives or dies on whether somebody believes the person saying it.
That is the whole distinction. One tool fixes how much video exists. The other fixes whether a stranger trusts forty seconds of it.
This is the whole editor
Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.
Cutroom cuts a read you actually gave, and you steer it on the transcript
You film yourself saying the thing, up to three minutes, on a phone. Cutroom transcribes it and drafts the cut. B-roll lands on the phrases that need showing.
Captions get styled. An opening headline gets drafted. A 9:16 MP4 comes back. Direction then happens on the words rather than on a playhead.
Highlight a phrase and a clip lands over exactly those words. Delete a line and the cut rebuilds. Change the pace and the whole piece re-cuts.
Coverage is capped at 40 percent on chill, 45 on normal and 52 on fast. Your face never disappears behind stock.
There is an avatar module for the days you cannot film. A photo plus a voice-over becomes a lip-synced presenter. It then gets cut like any other take. Avatar render is 320 credits a minute and voice-over is 70.
Cutting a take you filmed costs 20 credits a minute. The gap between 20 and 320 is the product pointing at the path it was built for.
- Your own face and your own read, not a licensed performer
- One take in, one finished vertical ad out, captions burned in
- Six caption packs and five clip transitions, cut through to sweep
- 300 free credits over 7 days, no card, enough for two finished ads
From a phone in a kitchen to a 9:16 MP4
The 100 credit batch covers steps two and three. Step five is 20 credits a finished minute, every time after that.
Film one take
3 min ceiling, phone is fine
Transcribe
Word-level timing
Draft the cut
Headline, b-roll, captions
Direct by reading
Delete, highlight, re-pace
Export
9:16 MP4, 20 CR a minute
A new prompt changes four variables. A new cut changes one
Both tools produce a vertical video with captions, so on a feature grid they look adjacent. One fills an empty timeline from text. The other decides what survives from a performance that already happened.
When a generated ad fatigues, the next move is a new prompt. New script, new visuals, new pacing, new voice, all changed at once. When it performs better, nobody can say which change did it.
When a filmed ad fatigues, the next move is a different cut of the same read. Start on sentence three. Move faster. Put the footage on the claim instead of the greeting. One item moved, and you know which.
The read is the control variable and the edit is what you test. Transcription and the director pass are paid once at 100 credits. Every further version costs only the export, at 20 credits a finished minute.
Five edits of one take is an afternoon rather than a budget conversation. Published benchmarks put winning creatives at roughly 5 to 8 percent, per Motion's analysis of 550,000 plus Meta ads. Cheap, readable attempts are the only lever you own.
How many variables move between version one and version two
A test with four moving parts cannot tell you which part won. That is what makes a cheap re-cut worth more than a cheap re-prompt.
- New prompt: script, visuals, pace, voice4 variables
- New cut: one of them, on purpose1 variable
Your own face, cut five ways in an afternoon, for 20 credits a version
That is what this side is for. Coaching, supplements, skincare, services. Wherever the face is the mechanism the offer runs on, a filmed read out-earns a generated one and keeps earning.
Now the facts a buyer needs. There is no text-to-video path here and no script input. No horizontal or square. No source over three minutes. No timeline to escape into when the draft gets a beat wrong.
The avatar module is the one exception and it is narrow. It needs a photo you supply plus a voice-over. There is no library of licensed performers to browse.
InVideo is the buy for faceless content, or for many videos from text you already wrote. Cutroom is the buy when somebody will film.
A middle path exists and it is the one we would recommend to a friend. Use generated video to find which angle holds attention in a cold feed. Then film the winner yourself and cut it here. Testing text is cheap. Testing a shoot is not.
Run the numbers on your own volume. A one-minute ad here is about 120 credits. Basic carries 2,500 credits for $39.99 a month. Read InVideo's pricing page at the source, because a table written by a competitor is not evidence.
InVideo and Cutroom, row by row
Two rows belong to InVideo, and they matter if nobody will film. The other six are what a paid feed is judging.
| InVideo | Cutroom | |
|---|---|---|
| Video without filming anything | Prompt in, finished video out | You film a take first |
| Many videos from text you have | Built for exactly that | No text input anywhere |
| Captions burned into the export | Styled over generated speech | Six packs, from your own read |
| B-roll on phrases you name | Stock chosen by the prompt | Highlight words, clip lands |
| Your own face and read | Stock and generated narration | Your camera, your voice |
| Edit driven by the transcript | Prompt again to change it | Delete a line, cut rebuilds |
| Same read, different edits | A new prompt changes everything | 20 credits per output minute |
| Speaker stays on screen | Stock can fill the frame | Coverage capped at 52 percent |
Questions people ask
- Can Cutroom write the script for me?
- No. It works from what you said. It drafts an opening headline from your read and decides what to cut, but the body of the ad is your own words.
- Does Cutroom use stock footage?
- Yes, as b-roll over specific phrases. You can swap any clip or search millions of free ones. Stock is the supporting layer, capped at 52 percent, never the whole video.
- How long can my upload be?
- Three minutes. That ceiling is deliberate. This is built for the read that becomes a short vertical ad, not for a webinar.
- Is it right for every ad?
- It is built for ads where a person is the reason someone believes the claim. If you have no footage and nobody who will film this quarter, a prompt-to-video tool is closer to your problem today.
You can generate almost everything now. A face a stranger believes is the one thing still worth filming. Cutroom is what turns that read into a finished ad.