Feature

Lip sync video generator: the mouth is reading the sound

Bad lip sync gets blamed on the model almost every time. It is usually the recording. Mouth shapes are derived from the waveform you uploaded. A hard room shows up as a soft mouth. Here is what the render takes, the four recordings that break the sync, and how the clip becomes a finished ad.

Close-up of a young man holding a microphone with facial hair, outdoors.
Photo by Samer Daboul on Pexels
Audio length the render accepts
3s to 5min
Per minute if you generate the read
70 CR
Per minute of finished lip-synced video
320 CR

In: one still and one sound file. Out: a vertical MP4.

The portrait goes in as a PNG, JPEG or WebP under 15 MB.

The audio goes in as a file you recorded, or you generate a read here for 70 credits a minute.

Audio shorter than three seconds is refused. Audio longer than five minutes is refused.

Back comes one continuous head-and-shoulders shot in 9:16. The mouth is in time with the sound.

Billing is by finished minute of render at 320 credits. Forty seconds is about 213 before the voice-over.

The trial is 300 credits with no card. That is roughly one minute of render.

Spend twenty seconds of it on a test, for about 107 credits, before you write the rest of the script.

You will hear whether your recording syncs cleanly in the first five seconds of that test.

The four recordings that produce a mushy mouth

Sync is derived from the waveform. Anything that blurs the waveform blurs the mouth.

Reverberant rooms are the worst offender. A kitchen, a bathroom, a bare office with hard walls.

The tail of each word smears into the next and the mouth never fully closes.

Music or ambience mixed under the voice before upload is the second. The model cannot tell which part of the sound is speech.

Give it speech alone. Add music later, at the edit, for 20 credits a track.

Heavy compression or a low bitrate export is the third. Consonants are the highest frequency information in the file. They go first.

Speaking too fast is the fourth. Above roughly 180 words a minute the mouth shapes start overlapping.

A duvet over your head and a phone at arm's length beats a good microphone in a hard room.

That is not a joke. It is what the waveform responds to.

  • Record in a soft room. Carpet, curtains, a wardrobe of clothes behind you.
  • Speech only. No bed, no ambience, no ducking applied before upload.
  • Export at a high bitrate. Consonants live in the part a low bitrate deletes.
  • Aim near 150 words a minute. Above 180 the sync starts reading late.
  • Leave half a second of silence at the top and tail. The mouth needs somewhere to start and stop.

What actually decides whether the sync looks right

Buyers spend their effort on the portrait. The recording is what the model is reading.

  • Room and recording qualityMost of it
  • Pace of the readThe next slice
  • The portraitRules options out

This is the whole editor

Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.

CLIPS · 5I have thisexact conversationeverysingle week. Somebody sits down and says,oh yeah, I takecinnamonevery day.And honestly, doc, I have no idea if it works.So let me tell you what isin that capsule.a clip lands on these wordscut from the editTAKING CINNAMONEVERY DAY?is it doing anythingHeadlineMusicCaptionsTHIS VIDEOLength25.0sClips5Words removed18Export video

A generated read costs 70 credits a minute and has no room in it

A generated voice-over is recorded in no room at all. No reverb, no fridge hum, no bitrate history.

That makes it the most reliable input the sync can have.

For a forty second script it is about 47 credits, which is small next to the 213 the render costs.

It also needs nobody's time. Write, generate, render, finished, without leaving the desk.

The trade is delivery. A generated read is even. Evenness is what makes a long piece feel synthetic.

Under forty seconds it is barely noticeable. Over ninety it is the main thing a viewer hears.

So use the generated read for information. Use your own voice where a line needs a person in it.

Whichever you pick, listen to the audio alone before you spend a credit on the render.

If it sounds unclear to your ear, it will look unclear on the mouth.

The order that saves credits

Every check here is free. The render is the only expensive step.

  1. Write it

    Read it out loud once

  2. Record or generate

    70 CR a minute if generated

  3. Listen alone

    Free, and catches most failures

  4. Render 20 seconds

    About 107 CR as a test

  5. Render the rest

    320 CR a minute

The sync is fine and the video still feels wrong

This is the complaint underneath most lip sync disappointment. It is not a sync problem.

A face holding one expression for forty seconds is the tell. Real people shift and blink irregularly while they talk.

The fix is to stop showing the face for the whole runtime.

Hand the render to the editor as a source take. One batch is 100 credits.

You direct the cut on the transcript. Highlight a phrase and a clip lands over exactly those words.

Delete a line and the cut rebuilds around it. Change the pace and the whole piece re-cuts.

At normal pace up to 45 percent of the ad is covered with footage, so the unvarying face holds barely half of it.

Captions from six style packs carry the words through those covered sections, which is how most people watch anyway.

A perfectly synced forty second stare is still a forty second stare. This is the step that fixes it, and it is the step other lip sync tools do not have.

What this makes, and the job it is not for

It renders a new video from a still. That is the whole mechanism. It explains the strengths and the edges.

It is strongest on scripts where the information is the value and a face is carrying it.

It does not dub existing footage. There is no original performance to keep. The video is being made rather than altered.

If you need a real take re-voiced, film it again or use a dubbing service. Different job, different shape of tool.

The portrait is yours, or belongs to somebody who gave explicit written permission. There is no third case.

The plain facts. Output is 9:16. Audio caps at five minutes. There is no timeline, by design.

If somebody will speak the words to a phone, that take batches for 100 credits and finishes in the same editor.

Both routes end in the same 9:16 file with captions and b-roll already on it.

That is worth knowing before you compare lip sync tools on render quality alone. None of the others own the step after the render.

Questions people ask

Does a better microphone improve the sync?
Less than a better room does. Reverb is what blurs the waveform. A phone recorded in a wardrobe full of clothes syncs better than a condenser microphone in a bare office.
Can I lip sync to a song?
The sync is built for speech. Music mixed with a vocal gives the model two signals. It cannot tell which one is the mouth, so the result drifts. Give it a clean vocal alone if you try it.
What is the longest video I can make?
Five minutes of audio, which is 1,600 credits to render. In practice almost nobody goes past sixty seconds, because an even delivery loses viewers first.
Who should not buy this?
Anybody trying to dub existing footage, and anybody wanting words in a recognisable person's mouth. Dubbing needs a dubbing service. The second is not something this product will help with.

Fix the room, then the photo, then let one batch cover half the runtime. Every lip sync tool renders a mouth. This one hands back an ad.

Start with one take300 free credits · no card · cancel anytime