The cheapest model usually wins the spreadsheet. Its price per image is lower, so the monthly API total looks better. Then production starts.
Take a routine packshot edit. The first result changes the product. The next keeps the packshot but breaks the type. A third reaches review and comes back with the same fidelity problem. The provider records three inexpensive generations; the team pays for three attempts and another round of work.
This is the gap an AI model cost comparison needs to capture. Providers bill the generation event. Creative teams keep paying until an asset can leave production.
The rate card stops too early
AI model pricing measures an image, a video second, a token, a credit, or a completed task. It does not include the option nobody uses, the prompt rewrite, the manual repair, or the second review. Those costs sit with the production team.
The FinOps Foundationdistinguishes resource units such as cost per token from business units such as cost per transaction or case resolved. For creative production, the useful business unit is often an approved asset: work that meets the required format, brand, product, claim, market, and usage conditions.
Approval does not promise campaign performance. It means the asset can leave production without another repair or decision. That is also different from review-ready output: the point at which a person has enough context to make the decision.
01 / Public evidence
The cheaper image model cost more per successful edit
The January 2026 HYPE-EDIT-1 benchmark makes the hidden part of AI image generation cost visible. It tested 100 marketing and design edits, ran every task ten times per model, and had five blinded raters mark each result pass or fail.
The benchmark allowed up to four attempts and added 20 seconds of human inspection at $50 an hour: $0.278 of review for every candidate. Its combined results make the rate-card trap visible.
| Model | Candidate | First-pass | Attempts | Per success |
|---|---|---|---|---|
| Gemini 3 Pro Preview | $0.134 | 63.8% | 1.85 | $0.95 |
| GPT Image 1.5 | $0.170 | 61.2% | 2.04 | $1.30 |
| Qwen Image Edit 2511 | $0.030 | 45.4% | 2.48 | $1.33 |
| Seedream 4.0 | $0.030 | 35.6% | 2.64 | $1.42 |
Seedream 4.0 was about 4.5 times cheaper per candidate than Gemini 3 Pro Preview. Once retries and review were included, it cost about 49% more per successful edit: $1.42 versus $0.95.
02 / AI video evidence
Similar video ratings appeared at five times the price
AI video generation cost shows why the decision must stay task-specific. In a 9 September 2026 snapshot, Artificial Analysisgave statistically similar ratings to models priced at $2.40, $6, and $12 per generated minute.
The three aggregate ratings were statistically close, while listed prices differed by as much as five times. That does not make the $2.40 model the automatic winner. The benchmark’s blind votescannot tell you which model will preserve your packshot or clear your reviewer. They show why price alone is a poor shortcut for predicting the approved result.
03 / Worked example
Count the work between generation and approval
Here is the same mechanism without model names. Two fictional routes receive the same input and must pass the same review.
$2 route · two attempts
- Generation
- $4.00
- Review
- $0.56
- One approval
- $4.56
$3 route · one attempt
- Generation
- $3.00
- Review
- $0.28
- One approval
- $3.28
The review line uses the same $0.278-per-candidate assumption as HYPE-EDIT-1. Even before review, two $2 attempts cost 33% more than one $3 attempt. With review, the $4.56 route costs about 39% more than the $3.28 route.
If both models pass on the first attempt, the $2 model wins. If the cheaper one needs more retries or repair, the saving can disappear. The rate card cannot tell you which situation you have.
With these assumptions, the $2 route needs roughly a 70% first-attempt pass rate to break even against a $3 route that passes first time.
Use one number the team can audit
Approval-adjusted model cost includes calls, failed runs, abandoned options, repair, assembly, and review. Use the same loaded labor rate for every model. The point is not to price creative judgment; it is to stop treating human attention as free.
Keep time to approval and mistakes found after approval beside the cost figure. A product error, unsupported claim, likeness problem, or rights issue can rule out a route regardless of its average.
04 / Route by task
Cheap wins—when it clears the bar
Replacing every cheap model with a premium one repeats the same mistake. A low-priced model that handles the job cleanly is the right choice. A premium model that still needs retries is not.
Put model choice inside a complete AI ad workflow, where the brief, production steps, review, and output history stay together.
Feed the result back into the next run
Runway’s Media Routerchooses among eligible models using cost, quality, latency, price caps, and allow or deny lists. That answers which model should receive a request.
The useful evidence arrives later: whether the product stayed faithful, how many attempts were abandoned, what a person repaired, and whether the asset passed review. Return that evidence to the next AI model routing decision, or the router keeps optimizing the call instead of the outcome.
05 / Qualified Yield Test
Test one job before changing your stack
“Video generation” is too broad. “Extend this approved product shot into an eight-second 9:16 clip without changing the packshot, claim, or logo” is a job you can test.
Define the job
Use the same kind of source, instructions, output specification, and reviewer for every model.
Write the pass criteria
Decide what cannot break—product, copy, claim, likeness, rights, market, or format—before seeing results.
Run the same batch
Record calls, abandoned outputs, repair time, review time, and approved results.
Set a retry limit
After that limit, move the job to another model or a named person.
Test again after a change
A new model version, source type, prompt structure, or pass criterion can change the result.
Different jobs need different reviewers. A background extension and a synthetic spokesperson should not follow the same path. The AI content approval workflow guide shows how to route review by what changed.
Keep the scorecard small
- approval-adjusted model cost;
- first-attempt approval rate;
- human minutes and model calls per approved creative;
- approved outputs that needed no repair;
- time to approval and problems found afterwards.
Drop a model if it repeatedly misses a must-have requirement. From the rest, choose the one with the lowest observed total cost and an acceptable review record.
06 / Production evidence
Keep the result next to the workflow
A fair test falls apart when prompts live in one tool, outputs in another, costs in a provider dashboard, and revisions in chat.
Pawook’s multi-model production layer can connect separate models to defined production steps and keep runs, intermediate outputs, revisions, execution status, and usage costs with the workflow. Its collaboration and governance layer keeps that production record visible to the people responsible for the process.
Pawook does not decide whether an asset is good or legally safe. The team still needs a reviewer and, when approval happens elsewhere, a matching outcome record.
Start with one weekly job. Compare the current model with one alternative for two weeks, using the same inputs and reviewer. Then look at approved work, total dollars, and human minutes—not the number of files generated.
Sometimes the cheap model wins. Sometimes it sends the same job around the loop three times. The rate card cannot tell you which one happened.
Sources
