Feature
One photo, one audio track, and a presenter renders
Two files decide everything here. One front-facing photograph and one audio track return a lip-synced presenter take at 320 credits a rendered minute. That take then runs through the editor exactly like filmed footage and exports as a finished 9:16 ad. Here is what makes a photo work, what the audio decides, and what a campaign costs. By the end you will know which scripts to render and which to film.

- Photo required
- 1
- Credits per avatar minute
- 320
- Credits per voice-over minute
- 70
What you hand it, what it returns, and what the minute costs
The inputs are two. One still photograph of a person, front-facing, with nothing crossing the mouth or jaw. One audio track of the read.
The output is a video take of that person speaking your words. It enters the editor exactly like a filmed one. Transcribed, cut, covered with footage on the words that need showing, captioned, headlined, exported 9:16.
A rendered minute is 320 credits. That makes it the most expensive thing in the product by a wide margin, because every frame is generated rather than found.
A forty second read is about 213 credits of render. The 100 credit batch and the export at 20 credits a minute sit on top of it.
Basic at $39.99 for 2,500 credits a month is a little under eight minutes of presenter footage. Premium at $79.99 for 5,000 is about fifteen. A $15 top-up adds 1,000 credits, roughly three more minutes.
Credits per minute, by what the machine makes
Generating frames costs sixteen times what exporting them does. Write the avatar read short.
- Export a finished minute20 CR
- Generated voice-over, a minute70 CR
- One batch, whole take100 CR
- Avatar render, a minute320 CR
It animates a photograph in time with a waveform, and that predicts everything
The model does one narrow task. It drives a mouth, a jaw and some head motion from an audio waveform. The face it works from is a single still frame.
Hold that sentence and you can predict the result before spending a credit. Anything near the mouth tends to look right, because that is the whole job.
Anything far from the mouth is where the illusion goes thin. Hands, shoulders, the background, the sense that a body is holding itself up. None of it was ever being modelled.
Which is why the practical advice is short reads and frequent cuts to footage. The illusion is strongest in bursts and thins over long holds.
What the model is handed, and what it drives
Nothing outside the mouth and jaw is being modelled, which is exactly where viewers look next.
One still photo
Front-facing, mouth clear
One audio track
Yours, or 70 CR a minute
Mouth, jaw, head
Driven from the waveform
A video take
320 CR per rendered minute
This is the whole editor
Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.
The photo sets the ceiling, the audio decides whether you reach it
Front-facing beats three-quarter. A neutral or slightly open mouth beats a wide grin, because a fixed smile has to be broken apart to speak. Even, soft light beats hard shadow.
Nothing may cross the mouth or jaw. No hand on the chin, no microphone, no heavy shadow at the lip line.
Resolution matters less than people expect. Framing matters far more. A clean head-and-shoulders shot at moderate resolution beats a high-resolution full-body photo where the face is a hundred pixels wide.
The audio carries more of the result than the image does. Clear speech, even pace, no traffic behind it, no music baked in. Muddy audio gives the mouth nothing crisp to follow and the sync goes soft.
Record the read yourself where you can. It usually sounds better and it costs nothing extra.
Test the photo before you write the script. One thirty second render is 160 credits, and it tells you more than any checklist can.
The presenter carries the words. Footage carries what it cannot hold
The render gives you speech and modest head movement. That is the performance in full, and for an information-led script it is enough.
There are no hands. The presenter cannot pick anything up, point at anything, or demonstrate a step.
So you cover those lines with real footage of the product, placed on the exact words that name it. Highlight the phrase and the shot lands there. It does the job a gesture would have done, and does it more clearly.
There is no acting either. No glance to camera at the punchline, no shift in posture when the tone changes. Write for the voice and let the pictures do the pointing.
One photo means one look for the length of the take. A two minute render holds the same expression at the end as at the start, which is why short reads and frequent cuts win.
And the face has to be one you may use. Your own, or somebody who agreed in writing that their likeness can appear in your paid advertising.
Where a render beats a booking, and where a phone still wins
The render earns its price on three things. Nobody has to be available. Ad twenty is framed exactly like ad one. Twenty scripts can be rendered inside a week.
That is a strong case. A spokesperson effect that used to need twenty bookings now needs one photograph and one audio file.
A filmed take is cheaper. One batch is 100 credits against 320 a minute of render. It can hold the product, and audiences believe it more.
So the split is clean. Render the scripts nobody will read on camera. Film the ones where a specific human has to be believed. Both routes use the same editor.
A medical result or a financial outcome belongs in a filmed take. Some viewers clock a render on a long hold, and the claim inherits the doubt.
What is left for the render is large. The ad carried by what is said. A founder who will not be on camera. A presenter you want in twenty ads over six months without booking anybody twice.
A generated presenter against filming it yourself
Two rows go to the camera. Render the five it wins, and film the two it does not.
| Filming a real take | Cutroom | |
|---|---|---|
| Holding or pointing at the product | Hands are in frame, doing the selling | No hands, no props, no gestures |
| Being believed on a trust-led claim | A real person carries the claim | Some viewers clock it on longer holds |
| Ad twenty framed exactly like ad one | Light and wardrobe drift over six months | Same photo, same framing, every render |
| Works when nobody will go on camera | Somebody has to do the read | A photo and an audio track is enough |
| Twenty scripts turned around in a week | Twenty set-ups and twenty diaries | 320 credits per rendered minute |
| Changing the offer the afternoon it changes | Book, light and film it again | New audio, same photo, re-render |
| Captions and footage on the same file | A second tool after the shoot | The render goes straight into the editor |
Questions people ask
- How many photos do I need?
- One. A front-facing head-and-shoulders shot in even light, with nothing crossing the mouth or jaw, beats a whole folder of badly framed images.
- Can I use my own voice?
- Yes, and it usually sounds better. Record the read and supply it as the audio. If you would rather the machine spoke it, generated voice-over costs 70 credits per minute.
- What does a minute of avatar cost?
- 320 credits. Basic at $39.99 for 2,500 credits a month is a little under eight minutes, Premium at $79.99 for 5,000 is about fifteen, and a $15 top-up adds 1,000 credits.
- Can I use a photo of a public figure or a stock model?
- No. Use yourself, a colleague who has agreed in writing, or a licensed likeness where the licence explicitly covers synthetic video. Public figures are never available.
- Is it right for every script?
- It is built for scripts where the information is the value. If the ad has to show the product being handled, film that part on a phone and run it through the same editor for 100 credits, then place it under the presenter's words.
One photo, one audio track, and the same presenter in ad twenty. Then footage on the words a face cannot hold, in the same tool.