What One Second of AI Video Costs Us, Layer by Layer

September 24, 2026Startup Ops9 min read
What One Second of AI Video Costs Us, Layer by Layer

Between 12:01 and 18:01 UTC on 10 September, our balance at the avatar vendor dropped by 67 cents. That delta is the only current AI video cost per second in our pipeline that anyone has measured with a wallet.

Everything else we could tell you about what a clip costs us comes from a price list, a calculation, or a guess.

On 21 September, @samgrows asked(opens in new tab) the people shipping video tools for two things: a full-length sample, and this: "I just wish everyone was a little bit more transparent about the costs behind things." The post has passed fifteen thousand views. This post is our answer to the second half, for Clipwright(opens in new tab), the talking-actor video API we build.

What it won't give you, up front: a total, or a margin. Most of the stack isn't measured, and the stack isn't finished. We're still adding image generation, and generated footage may come after that. A total built on those gaps would look more certain than it is.

The layers in one clip, and what kind of number each one is

A Clipwright clip passes through five paid layers today, plus a one-off portrait for each new actor, and two more layers are planned. Here is each one, what it costs, and what kind of number that cost is.

Table: Layer, Cost, What kind of number
LayerCostWhat kind of number
Script cleanup (an LLM call)
about $0.01 per clipPrice list
Voice (text-to-speech)
about $0.10 per 1,000 characters, halved on the faster modelsPrice list
Avatar and lip-sync (HeyGen Avatar IV)
$0.0385 per secondMeasured by wallet delta
Actor portrait (an image model)
$0.015 per image, once per actorMeasured once, from a vendor dashboard
Render
about $0.03 per clip, roughly two minutes of CPUEstimate
Storage
about $0.002 per clipPrice list
B-roll images
nothing yetPlanned, no data
Generated footage (image-to-video, prompt-to-video)
nothing yetPlanned, no data

We name one vendor here, the one we measured. Every other row is described by category on purpose. A named vendor's list price goes stale within weeks, and what we want you to take from this table is the right-hand column.

Each row is billed in its own unit: per second, per thousand characters, per image, per clip. We left them that way on purpose. Converting them all to one clip length would amount to printing a total, which we said up front we wouldn't do.

Sorted by evidence instead of by dollars, the same rows look like this:

How we know each cost number How we know each cost number, by layer One dot per layer, in the column for the strongest evidence we have. No dollar amounts. Wallet Dashboard Price list Estimate No data Avatar and lip-sync Actor portrait Voice Script cleanup Render Storage B-roll images Generated footage Strongest evidence on the left. Each layer is billed in its own unit (per second, per character, per image, per clip).

How each of our cost numbers is known, as of 23 September 2026.

One dot sits in the wallet column. One more comes from a single reading of a vendor dashboard. The render dot is our own estimate of CPU time, three dots are what vendors say they charge, and the last two rows are empty circles.

The one layer we measured with a wallet

On 25 August, the avatar vendor announced that from 28 August, Avatar IV would drop from $0.05 to $0.0385 per second. We did not change our cost table that day. We have a project rule that a rate from an announcement doesn't go into the table until a wallet confirms it, since the balance shows what the vendor actually took.

We didn't need a special paid run to check it. A background job reads our vendor balance every six hours and stores the result. On 10 September the probes looked like this:

How the avatar rate was measured on 10 September Vendor balance probes, 10 September 2026 (UTC) Change in balance between probes, in cents 06:01 12:01 18:01 −67¢ no renders control window one render, 17.9 s clip $0.67 ÷ 17.9 s = $0.0374/s At the old $0.05 rate the same clip would have cost $0.90.

Balance deltas only. The first window is the control: with nothing rendering, the balance didn't move, so the second delta belongs to the one run inside it.

The idle window is the easy part to skip. Without it, 67 cents is a number near a render. With it, the wallet is shown to sit still on its own, and the delta can only belong to the run.

That run was the only avatar render in the window, and the clip came out 17.9 seconds long by ffprobe. $0.67 over 17.9 seconds is $0.0374 per second, about 2.8% under the announced rate. At the old $0.05 rate the same clip would have cost $0.90, so the cut is real, by far more than rounding could explain.

Here is what it cost us to learn that. The clip never reached a customer, so the run earned nothing. The 67 cents was our money, spent on a video nobody received, and it is still the best-evidenced cost number we have. (Why a lost clip is ours to eat and not the customer's is its own post: four outcomes, not two.)

Our table still uses the announced $0.0385, not the measured $0.0374. The measurement divides by our duration, from ffprobe. The vendor bills by its duration count.

$0.67 fits 17.9 seconds at $0.0374 just as well as 17.4 seconds at $0.0385, and one wallet delta can't tell those two apart. When the denominator is yours and the bill is theirs, take the higher rate.

An earlier wallet measurement, in August at the old rate, had also come in a little under list, so the direction is consistent. Still, this is one reading at the new rate, and one reading can't show how much the rate varies.

Two more limits, stated because they're real. Six-hour probes can't tell you when inside the window the vendor charged, so this says nothing about whether it bills before or after a clip is ready. And anything spent through the vendor's own dashboard, by hand, wouldn't appear in our run list.

Why the measured layer happens to be the one that matters

The avatar is the only layer priced per second of finished video. Script cleanup and storage are priced per clip, voice per character of script, render by our estimate of CPU time, and the portrait once per actor.

We didn't plan it that way, but it means the one rate set directly in seconds of output is the one we've checked against money. An error there repeats in every second of every clip.

The portrait row shows why "measured" needs a sub-label. We measured it once, on 20 September, and not with the wallet. The image vendor's dashboard splits spending by model endpoint: two images went through, and two endpoint lines moved by a total of three cents.

We used the breakdown because the overall balance lags spending. It showed four cents against the three in the breakdown. The breakdown has its own trap, though: it splits by model, not by run, so a second portrait job running at the same moment would land on the same lines and add to ours without a sign.

So one clean observation gives the order of magnitude and nothing more. The high-quality setting hasn't been measured at all.

Then there are the empty circles. B-roll images and generated footage have no numbers of ours yet, which is a different thing from costing nothing. The public price tables we've seen for generated footage are per second of output (here is one(opens in new tab)), like the avatar, so nothing on this page tells you what it will do to our costs. We'll know when there's a wallet delta to show.

A cost table that prints a blank as $0.00 is the easiest way for one of these tables to mislead, ours included.

A per-second vendor rate is not the cost per second of a video

$0.0385 is what one vendor charges us for one step. It isn't what a second of finished video costs to make.

A tool's price per minute of finished video covers every layer in the table above, plus the runs that fail (who pays for those is a rule of its own), plus whatever the tool keeps. Put $0.0385 next to someone's per-minute price and you're comparing one ingredient with the menu. The two numbers don't share a unit, and the division that makes them look comparable is the error.

That's also the honest reply to the "show me a full ten-minute sample" half of the request. Our rows are per second or per clip, and a ten-minute render is exactly where our table is weakest. The render line for a short clip is an estimate of CPU time. For a long composition we haven't measured it at all.

If you're pricing AI video from the outside, the same unit discipline applies to what vendors sell you. We walked through three pricing pages that never state what a finished video costs in what a credit buys.

How to label your own AI video cost table

The line-by-line AI video cost posts on this topic are useful, and some are careful about method. One breakdown of an AI cartoon pipeline(opens in new tab) says plainly that its prices are the vendors' posted API rates, cross-referenced with its last hundred or so generations, and that its generated-seconds figure matches the invoice.

Another, billed as real numbers from a production pipeline(opens in new tab), is mostly vendor rate tables. Neither is wrong. But neither labels, line by line, which rows were checked against money, and that is the column that decides how much weight each line can carry.

Here is the column we use, from strongest to weakest:

Table: Label, What it means, Example from our table
LabelWhat it meansExample from our table
Measured by wallet
Balance before and after a known run, with an idle control windowAvatar and lip-sync
Measured by breakdown
A vendor dashboard split by model, one clean observationActor portrait
Price list
What the vendor's page or announcement saysVoice, script cleanup, storage
Estimate
Our own calculation (CPU time, a cloud pricing page)Render
No data
The layer exists or is planned, and has no number yetB-roll images, generated footage

Four rules came out of building ours:

  • An announced rate is a price list until a wallet agrees with it.
  • A wallet delta counts only next to a window where nothing ran.
  • When you divide by your own duration and the vendor bills by theirs, keep the more conservative rate.
  • A total inherits the weakest label of anything in it. If you print one, print that label next to it.

The ten-minute version: open your cost sheet and add one column called "how do we know". Fill it in honestly. The rows that come back "price list" are your to-do list, and the biggest of them is the one to put a wallet on first.

FAQ

What does one second of AI avatar video cost to generate?

For the avatar and lip-sync step alone, the AI video cost per second we measured was $0.0374 against our vendor's announced $0.0385 in September 2026, and we budget at $0.0385. That's the vendor's charge for one step. A finished clip also carries script, voice, render and storage costs, and a customer price also has to cover the runs that fail.

Why not publish a total cost per clip?

Because four of the five per-clip layers are price lists or estimates, and two planned layers have no numbers. A total would carry the weakest of those labels while looking like the strongest. We'll publish more rows as they're measured.

How do you measure a vendor cost with a wallet?

Read the balance on a schedule. Find a window with exactly one known run and a neighbouring window with none. If the idle window didn't move, the delta in the busy window belongs to that run. Divide by the output length, and remember the vendor may count that length differently.

Our cost table is still being measured. Your price shouldn't be.

You should not have to budget around our unknowns. Clipwright bills the finished seconds of the clip you get back, from credits you buy once that never expire.

About the Author

Dzmitry Vladyka
Dzmitry Vladyka

Dimantika

Founder of Dimantika. Co-founded and exited a SaaS at $1.2M ARR. Now building AI tools for founders who want autonomous growth without blind trust in agents.

View all posts