Feature

The transcript is the project. The video is the output

Timelines think in seconds. You think in words, and the gap between the two is where an editing evening goes. Here the take is transcribed with word-level timing and every instruction points at a word. Cover this phrase. Remove this line. Here is why that survives the second edit, what transcription gets wrong, and what one take costs to run.

A close-up shot of hands typing a film script on a laptop, focusing on creativity and screenplay writing.
Photo by Ron Lach on Pexels
Timestamps you ever type
0
A read-through of the transcript
30s
Credits per take through the machine
100

Word-level timing is what lets a shot land on two exact words

When your take is uploaded it is transcribed with word-level timing. Every word carries its own position in the audio, and that index makes the rest of the system possible.

The transcript is not a paragraph sitting next to a video player. It is an addressable list of every word you said and when you said it.

That difference decides what a tool can offer. If you only know where sentences begin, all you can sell is sentence-level trimming.

If you know where every word sits, you can place a shot over two specific words and have it land exactly there. Starting on the first, gone on the last.

Captions are timed from that index. Cuts are computed from it. The footage search reads it to work out what the ad is about.

Why an instruction survives an edit above it

Nothing in this chain stores a timestamp, so deleting a line changes nothing about what your other marks mean.

  1. Word index

    Every word, its position

  2. Your marks

    Cover this phrase

  3. Delete a line

    Words go, marks stay

  4. Recompute

    Whole cut, from scratch

Timestamps drift on the second edit. Word positions never do

Your instructions are expressed in words, never in seconds. Cover this phrase. Remove this line. Emphasise this word. None of them carries a time value.

That matters on the second edit rather than the first, which is why it is easy to miss when comparing tools.

If your instructions were timestamps and you deleted an earlier line, every timestamp after it would be wrong. The tool would then have to guess how to repair them.

That is the drift everyone has lived through. Captions slipping. A shot landing over the wrong sentence. By version five the project needs supervision rather than direction.

Instructions point at words here, so deleting a line changes nothing about what the other instructions mean. The whole edit is recomputed from the take and the current instruction set.

Which is why the count of edits appears nowhere in the interface. There is no version history to manage, because there is no accumulated version to manage it against.

This is the whole editor

Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.

CLIPS · 5I have thisexact conversationeverysingle week. Somebody sits down and says,oh yeah, I takecinnamonevery day.And honestly, doc, I have no idea if it works.So let me tell you what isin that capsule.a clip lands on these wordscut from the editTAKING CINNAMONEVERY DAY?is it doing anythingHeadlineMusicCaptionsTHIS VIDEOLength25.0sClips5Words removed18Export video

It mis-hears the words you care about most, and that is a thirty second fix

It will mis-hear words. Transcription is good and not perfect, and it degrades where you would predict. Proper nouns, product names, industry jargon, numbers said quickly, anything spoken over a fan or a road.

You fix it in place. Correcting a word updates the caption, so the export is right.

It also updates what the machine understands the line to be about, which is what the footage search reads. A line transcribed as our new sweet was about the wrong subject until you corrected it to suite.

So the habit worth forming is to read the transcript once before you direct anything. Thirty seconds on a short take.

The errors cluster in the words that matter most to you. Your brand. Your product line. The number in your offer. Those are the words the search leans on hardest, which makes it the cheapest thirty seconds you will spend.

One example from a real take. A founder said a product name the machine wrote as two ordinary words, and the footage search spent the whole ad illustrating those two words instead of the product.

Why the transcript is a record rather than a script field

Correcting the transcript does not change your audio. You are fixing the caption and the machine's understanding, not the recording.

You cannot type in words you did not speak. The transcript is a record, and editing it into something else would leave the caption disagreeing with the audio.

That rule is what makes the captions worth trusting. Every word on screen was said out loud by the person on screen.

The cut follows the order of the take for the same reason. Spoken sentences joined out of sequence carry mismatched breath and audible joins, and a viewer hears that before they hear the claim.

Removal works by the line, or by changing the pace and letting the read tighten. Extracting a single word from a spoken sentence is the operation that makes an edit sound spliced.

What one take costs, end to end, and what the trial covers

A batch is 100 credits. It covers transcription, the directing pass and the footage search. Export is 20 credits per minute of finished video.

So a thirty second ad, made and downloaded, is about 110 credits. Marking the transcript costs nothing until you export it.

The trial is 7 days and 300 credits with no card, which is two finished ads end to end. Basic is $39.99 a month for 2,500 credits, Premium is $79.99 for 5,000, and a $15 top-up adds 1,000.

The scope is one talking-head take of up to three minutes becoming a 9:16 vertical ad. Everything above was built for that one job.

Assembling an hour of interview footage or chaptering a podcast belongs in a timeline. Saying so here is cheaper for both of us than a refund conversation later.

For the job it was built for, you can change your mind eleven times and pay only for the download.

How far 300 trial credits actually go

Two finished thirty second ads, end to end, with change left over. No card to find out.

  • One thirty second export10 CR
  • One batch100 CR
  • Two finished ads220 CR
  • The 7 day trial300 CR

Questions people ask

How accurate is the transcription?
Good enough that reading it once takes thirty seconds. Errors cluster in proper nouns, product names, jargon and anything said over background noise.
Does correcting a word change what I hear?
No. It corrects the caption and what the machine understands the line to be about, which improves the footage placed on it. Your audio is untouched.
Can I type in something I did not say?
No. The transcript is a record of the take, not a script box. If the caption disagreed with the audio, both would be wrong at once.
What happens to my edits if I delete a line above them?
Nothing. Your instructions point at words rather than times, so the whole cut is recomputed and every other clip, caption and emphasis stays where you put it.
Is it right for every project?
It is built for one spoken take of up to three minutes becoming one vertical ad. Interview assemblies and podcast chapters belong in a timeline. Bring the ads here, where the eleventh version costs 10 credits.

Read the words once. Then mark them, and let the video be the output rather than the workspace.

Start with one take300 free credits · no card · cancel anytime