Comparison
ElevenLabs produces a voice. Cutroom produces the ad around it.
A voice engine turns text into a convincing read, in as many languages as you need. ElevenLabs is the best of them. A read is one ingredient of five in an ad. Here is what the other four cost, what a finished ad costs in credits, and how Cutroom produces all of it.

- per minute of generated voice-over
- 70 CR
- per minute of avatar render
- 320 CR
- one batch: transcript, director pass, b-roll
- 100 CR
For voice quality, languages and an API, a voice specialist is the correct buy
Control over delivery, a large library of voices, cloning, and output good enough that listeners stop noticing it is synthetic.
Many languages from one script, which matters when the same audiobook, course or product tour has to exist in nine markets.
And an API, so voice can sit inside a product rather than inside a person's afternoon.
Consistency is the underrated part. One voice across two hundred assets is a brand decision, and only a voice library holds it.
For an audio deliverable that is the correct buy, and nothing here competes with it.
Here is what a clean read still needs. Pictures, captions, pace, a first line that interrupts, an export sized for a feed.
That is four more jobs, and a voice engine correctly claims none of them.
Cutroom does every one of them from a single 100-credit batch, and generates a read at 70 credits a minute when nobody can film.
What comes back is a 9:16 MP4 with captions burned in and coverage sitting on the words you marked.
A voice file is not an ad, and the gap between them is the whole afternoon
The read arrives clean. Then it needs pictures, captions, pace, a first line that interrupts, and an export sized for a feed.
That second stage is where the hours live, and it is the stage a voice engine correctly does not attempt.
There is a harder problem underneath. A generated read is even by design, and evenness is the tell that makes a claim sound rented.
In direct response the hesitation is doing work. The half-second before a number is what makes the number sound arrived at rather than written.
That is why the strongest input for a paid ad is usually a read you already gave, in your own voice, with the roughness intact.
Cutroom is built around that recording. The voice module exists for the case where no recording is possible.
The half-second before a number is what makes the number sound arrived at rather than written.
Bring that recording here and the roughness survives into the export, which is the point of it.
From a spoken read to an uploadable file
A voice engine produces step one. Cutroom produces steps two to five, from a 100-credit batch.
The read exists
Filmed, or generated at 70 CR a minute
Transcribed
Every word indexed
Cut drafted
Director pass places coverage
You mark the transcript
Delete, highlight, pace
Captioned 9:16 MP4
20 CR a minute
This is the whole editor
Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.
Highlight the phrase and the picture arrives on those exact words
Upload a talking-head take of up to three minutes. It returns transcribed, drafted and captioned as a vertical MP4.
Highlight a phrase and a clip lands over exactly it. Swap the clip, or search millions of free ones.
Delete a line and the cut rebuilds around the gap. Change the pace and the whole piece re-cuts.
Emphasise the word your claim turns on so it pops in the captions. Fix a mis-heard word and the burned-in text follows.
Where nobody can film, the avatar module takes a photo you own plus a generated voice-over and produces a lip-synced presenter.
That combination costs 320 credits a minute for the render and 70 for the voice, on top of the 100-credit batch.
Coverage is capped by pace: 40 percent on chill, 45 on normal, 52 on fast. The presenter stays the spine of the ad.
- Six caption packs, adjustable for colour, weight, size and position
- Coverage capped at 40, 45 or 52 percent by pace
- Five entrances for a clip: cut, whip, punch, glitch or sweep
- Music generated from a described style, 20 credits a track, sitting under the read
Video in, video out: the specification a buyer plans around
The route in is a video take of up to three minutes. The route out is a 9:16 MP4.
There is no audio-only export. This cannot hand you a WAV at any price.
There is no voice cloning and no library of characters to audition. One generated voice per job, priced per minute.
There is no translation route. One take, one language, one finished file.
There is no way to reuse one generated voice across a campaign. Each job generates its own.
There is no timeline underneath and nothing of yours gets stamped on the file.
Those are the boundaries. Inside them sits a finished ad, which is what the next section prices.
The voice is one ingredient of five, and this ships all five
Winners run at 5 to 8 percent of creatives, per Motion's analysis of 550,000+ Meta ads.
Six a month draws under one winner. Thirty draws two, and the difference is what a second attempt costs.
Here that is 20 credits a finished minute and about four minutes of attention, from a take already batched.
One finished minute from a fresh take is about 120 credits. An avatar minute adds 320 on top.
Seven days and 300 credits with no card covers two finished ads before any decision is made.
Keep the voice specialist for anything where sound is the deliverable. It wins that outright.
Bring the read here when the deliverable is a file you upload to a feed.
ElevenLabs and Cutroom, row by row
The top two rows go to ElevenLabs, and an audio deliverable should weigh them heavily. The rest is the other four fifths of an ad.
| ElevenLabs | Cutroom | |
|---|---|---|
| Voice quality, cloning, languages | The core product | One read, 70 CR a minute |
| Audio-only files and an API | Use it anywhere | Video export only |
| Cut drafted before you start | Not the job | Director pass first |
| Captions burned in for a muted feed | Audio only | Six packs, one tap |
| Finished captioned video | Not the job | 9:16 MP4, six caption packs |
| Pictures tied to a phrase | Audio only | Highlight, clip lands |
Questions people ask
- Can I bring my own voice-over file into Cutroom?
- The route into the product is a video take. Generated voice sits inside the avatar module at 70 credits a minute, paired with a photo you own to produce a lip-synced presenter. There is no audio-first project type.
- Is the generated voice as good as a specialist engine?
- No. It is a working read priced at 70 credits a minute, built to feed the avatar module rather than to win a listening test. Where voice quality is the deliverable, buy the specialist.
- Why does Cutroom prefer a real recording?
- Because the unevenness carries information. A pause before a number reads as somebody arriving at it. A generated read is even by construction, and evenness is the first thing a sceptical viewer notices.
- Who should not buy Cutroom?
- Anyone whose deliverable is sound, in many languages, inside software. Buy the voice specialist for that. Bring the read here when the deliverable is a file that runs in a feed.
A voice engine hands you the read. Cutroom hands you the ad around it, captions burned in and coverage on the phrase you marked.