Feature

AI spokesperson for ecommerce: one face, forty listings

A catalogue is a scheduling problem before it is a creative one. Forty SKUs means forty scripts. Nobody stands in front of a camera forty times. A rendered spokesperson removes that constraint at 320 credits a finished minute. Here is what one SKU ad costs, how a single render fronts many listings, and where to spend the credits first.

East Asian woman unpacking a package indoors, with natural lighting creating shadows on her white knitted sweater.
Photo by Pavel Danilyuk on Pexels
Per minute of spokesperson render
320 CR
A forty second clip with a generated read
260 CR
Credits a month on Basic at $39.99
2,500

What one SKU costs, before you multiply it

One forty second spokesperson clip with a generated read is about 213 credits of render plus 47 of voice-over.

Call it 260. Batching that clip into a real ad with captions and product footage adds 100. The export adds about 14.

So one finished spokesperson ad for one SKU is around 374 credits.

Basic is $39.99 a month for 2,500 credits. That is six of those. Premium is $79.99 for 5,000, which is thirteen.

A top-up is $15 for 1,000 credits. That buys two and a half more.

The trial is 300 credits with no card. Enough to take one SKU all the way to a finished file before you decide.

Filmed source takes run on the same plan at 100 credits a batch. The same budget stretches further where a camera is possible.

Both numbers matter. No catalogue is written in a single register.

Finished ads per month, same budget

Same spend, two production routes, one account. The mix is the decision.

  • Filmed source, Basic plan21 ads
  • Rendered spokesperson, Basic plan6 ads

One render can front many SKUs, because the footage underneath changes

There are no hands in a rendered frame. No unboxing, no label turned to the lens, no pump on the back of a wrist.

That sounds like a wall. It is actually what makes the render reusable.

Shoot one session of product footage on a phone. An afternoon covers a catalogue.

Then render the spokesperson once for a script shape you use often. An offer. A spec comparison. A restock.

In the editor you highlight the phrase that names the product. A clip lands over exactly those words.

Swap that clip in any spot, or search millions of free ones. The same rendered read then fronts three SKUs with different footage under each.

Change the caption headline and the pack. The two ads read as siblings rather than as copies.

That is the version of this that pays back. Rendering a fresh face per SKU and posting it bare is the version that does not.

None of it needs a second tool. The swap is a highlight in the same editor that made the cut.

One render, many SKUs

The expensive part is the render. The cheap part is swapping what sits under the words.

  1. Shoot product footage once

    Phone, one session, all SKUs

  2. Render the spokesperson

    320 CR a minute

  3. Batch it

    100 CR, transcript and cut

  4. Swap the clips per SKU

    Highlight the phrase, drop the shot

  5. Export each

    20 CR per output minute

This is the whole editor

Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.

CLIPS · 5I have thisexact conversationeverysingle week. Somebody sits down and says,oh yeah, I takecinnamonevery day.And honestly, doc, I have no idea if it works.So let me tell you what isin that capsule.a clip lands on these wordscut from the editTAKING CINNAMONEVERY DAY?is it doing anythingHeadlineMusicCaptionsTHIS VIDEOLength25.0sClips5Words removed18Export video

Consistency is the reason ecommerce teams want this

The same presenter fronts January and June. Nobody is rebooked, restyled or unavailable.

That matters more for a catalogue than for a single product. A recognisable face across forty listings does work no individual ad does.

It also survives staff turnover. A filmed spokesperson who leaves takes every ad they appear in with them.

There is a quieter benefit. A fixed presenter means the only variable between two ads is the script.

So your test results tell you which script won. With filmed creators, every ad varies in lighting, energy and room.

The winner you find there may be a person rather than a message. That is an expensive thing to learn late.

The permission rule keeps this safe. The portrait is yours, or belongs to somebody who gave explicit written permission. That permission should name how long it runs.

A stock or image-model presenter sidesteps the question, at the cost of a face nobody recognises.

Decide it once at the start. Changing presenter across a catalogue costs more than choosing carefully did.

Which sentences to render and which to film

There is a rule teams land on after a month of running both. It fits in one line.

If the sentence contains a number or a policy, render it. If it contains an adjective about how something feels, film it.

Render: free shipping over forty euros, ships Tuesday, two hundred servings a tub, restocked this week.

Film: how heavy it is, how the fabric moves, how it smells when the lid comes off, the moment somebody sees the result.

Scale needs a hand next to it. Held next to a hand is a fact. Described as compact is a claim.

First-person testimony about a physical experience belongs on a phone too. A rendered face performing a memory it never had reads wrong.

Long scripts belong to neither. At 320 credits a minute, ninety seconds is 480 credits of render and loses viewers anyway.

Forty seconds is the working ceiling for a rendered read, which is also the length a feed rewards.

That split decides where the 320 credits a minute earns its price, and it takes ten minutes with a pen.

Where to spend the credits first

Buy the render first for the scripts nobody will film. Policy updates, shipping cutoffs, spec comparisons, restock announcements.

Those are the ads that never get made otherwise. They are the ones the render makes cheap.

If you already pay creators for footage, you own faces and hands. Batch those takes at 100 credits each and spend the difference on more attempts.

Published benchmarks put the winner rate at roughly 5 to 8 percent, from Motion's analysis of 550,000+ Meta ads.

Cost per attempt is the number that decides whether you reach a winner. The mix matters more than the tool choice.

The plain facts to plan around. Output is 9:16. A filmed source take caps at three minutes. Rendered audio caps at five.

There is no timeline, by design. Every lever is a decision about the ad.

What you get on both routes is a finished vertical ad rather than a clip to take somewhere else.

For a catalogue that is the part that decides whether forty listings ever get forty ads.

Questions people ask

Can the spokesperson show the product?
Not in the render. Film the product separately on a phone and place that footage under the phrases where the spokesperson describes it. One product shoot serves every script you render afterwards.
How many SKUs can one presenter cover?
As many as you like, because the presenter is not tied to a product. The limit is credits: at roughly 374 credits per finished ad, Basic covers six a month and Premium thirteen.
Is a stock presenter better than using my own face?
It is safer and less memorable. A stock or image-model face removes the permission question and cannot leave the company. Your own face carries more trust if customers already know it.
Who should not buy this?
Anybody whose catalogue sells purely on texture and scale. Put that budget into filmed takes at 100 credits a batch and keep the render for announcements and specs.

One face across a catalogue, different footage under every script, and a finished 9:16 file at the end of each one.

Start with one take300 free credits · no card · cancel anytime