I would budget an AI video integration around cost per accepted clip, not cost per generated second. The rate card is useful for building a shortlist. It does not tell me how many retries, review minutes, or edits a workload will need.
The directly comparable first-party rates here run from $0.03 to $0.70 per generated second. That is $0.30–$7.00 per normalized 10-second equivalent, not necessarily per supported request: some endpoints only generate specific durations.
Two things affect the shortlist immediately:
- Veo 3.1 Lite has the lowest listed 720p video-only rate at $0.03/second, but it is Preview.
- Sora 2 and Sora 2 Pro are deprecated, with the Videos API scheduled for removal on September 24, 2026. I would treat them as migration benchmarks, not new long-term dependencies.
Start with the invoice denominator
A cheap generation is not cheap if most outputs are rejected.
For planning, I use:
Average attempts per accepted clip = 1 / acceptance rate
Expected cost per usable clip =
base generation cost × average attempts per accepted clip
+ input, audio, editing, storage, and review costs
Here is the source pricing scenario translated into an accepted-output budget:
| Item | Assumption | Cost or calculation |
|---|---|---|
| Generation baseline | Normalized 10-second output | $1.20 |
| Acceptance rate | 60% | 1 / 0.60 ≈ 1.67 attempts |
| Generation per usable clip | Baseline divided by acceptance rate | $2.00 |
| Human review | 2 minutes at $30/hour | $1.00 |
| Storage and transfer | Planning assumption | $0.02 |
| Total per usable clip | Generation + review + storage/transfer | $3.02 |
Acceptance rate alone changes that $1.20 generation baseline substantially:
| Acceptance rate | Attempts per accepted clip | Generation cost per usable clip |
|---|---|---|
| 80% | 1.25 | $1.50 |
| 60% | 1.67 | $2.00 |
| 40% | 2.5 | $3.00 |
I would calculate those numbers separately for product shots, people, dialogue, visible text, camera motion, and multi-shot scenes. A blended acceptance rate can hide the prompt classes consuming most of the budget.
The production winner is the lowest-total-cost route that passes the workload’s acceptance criteria. It is not automatically the cheapest row below.
Keep pricing surfaces separate
Before comparing models, I separate three different products:
| Surface | What the price represents | How I would use it |
|---|---|---|
| First-party API | Provider’s published developer rate | Neutral cross-provider baseline |
| API gateway | Price for a specific routed model and parameter set | Actual integration budget, checked against the live catalog |
| Creator subscription | Web-app credits, limits, and UI features | Not API economics unless developer calls are explicitly included |
A gateway alias does not guarantee identical capabilities, billing units, or parameters to the first-party endpoint.
For a multi-provider evaluation, a unified API such as CometAPI can reduce separate authentication, billing, and client work. I would still use first-party rates for comparison and the gateway’s actual usage records for budgeting.
All rates below exclude taxes, storage, transfer, editing, rejected outputs, and human review. Before deployment, verify the exact model ID, platform, region, resolution, audio mode, duration, billing policy, and lifecycle date.
Fixed-rate routes: the useful shortlist
This table uses first-party prices or official credit conversions. Kling values retain the precise per-second equivalents rather than rounding away differences.
| Route | Configuration | Per second | Normalized 10 seconds |
|---|---|---|---|
| Veo 3.1 Lite | 720p, video only | $0.03 | $0.30 |
| Runway Gen-4 Turbo | Standard API generation | $0.05 | $0.50 |
| Veo 3.1 Fast | 720p, video only | $0.08 | $0.80 |
| Kling 3.0 | 720p, no native audio | $0.084 | $0.84 |
| Sora 2 | 720p | $0.10 | $1.00 |
| Kling 3.0 Turbo | 720p, native audio | $0.112 | $1.12 |
| Runway Gen-4.5 | Standard API generation | $0.12 | $1.20 |
| Kling 3.0 | 720p, native audio, no voice control | $0.126 | $1.26 |
| Veo 3.1 | 720p or 1080p, video only | $0.20 | $2.00 |
| Sora 2 Pro | 720p | $0.30 | $3.00 |
| Sora 2 Pro | 1080p | $0.70 | $7.00 |
For Veo 3.1 Lite, an allowed 8-second 720p silent output costs $0.24 before other expenses. Its fixed-quota access and regional availability vary.
At 1080p, Lite is $0.05/second, Kling 3.0 without native audio is $0.112/second, and Sora 2 Pro reaches $0.70/second. Those are useful shortlist numbers, not evidence of equivalent output quality.
Seedance is deliberately absent from this table: its configuration-dependent, token-metered billing does not reduce to one universal per-second rate.
Provider details that change the budget
Veo: distinguish the SKU from the callable endpoint
Google’s generative AI pricing page lists these output-second rates:
| Model | Configuration | 720p | 1080p | 4K |
|---|---|---|---|---|
| Veo 3.1 Lite | Video only | $0.03/s | $0.05/s | Not listed |
| Veo 3.1 Lite | Video + audio | $0.05/s | $0.08/s | Not supported |
| Veo 3.1 Fast | Video only | $0.08/s | $0.10/s | $0.25/s |
| Veo 3.1 Fast | Video + audio | $0.10/s | $0.12/s | $0.30/s |
| Veo 3.1 | Video only | $0.20/s | $0.20/s | $0.40/s |
| Veo 3.1 | Video + audio | $0.40/s | $0.40/s | $0.60/s |
Google documents 4-, 6-, and 8-second outputs. Standard and Fast -001 Agent Platform endpoints are GA, with retirement dates of November 17, 2026 or later. Lite is Preview.
The audio distinction matters: pricing pages list video-with-audio SKUs, but the current Agent Platform documentation marks sound generation as unsupported on the standard and Fast -001 endpoints while supporting it on Lite.
I would verify the exact callable route before budgeting for audio. A listed SKU is not enough.
Kling: log the full configuration
Kling’s developer pricing gives these per-second equivalents:
| Route | Configuration | 720p | 1080p |
|---|---|---|---|
| Kling 3.0 | No native audio | $0.084/s | $0.112/s |
| Kling 3.0 | Native audio, no voice control | $0.126/s | $0.168/s |
| Kling 3.0 Turbo | Native audio | $0.112/s | $0.14/s |
Kling 3.0 supports 3–15-second outputs. Its model guide documents native audio, multi-shot generation, and multilingual support.
At 720p, native audio without voice control raises the full model’s rate by 50%, from $0.084 to $0.126/second. Resolution, voice control, and Turbo versus full-model routing also affect pricing.
“Used Kling” is not enough information for a billing log. I would retain the route and submitted parameters with every job.
Runway: credits have a straightforward conversion
Runway developer credits cost $0.01 each. Its API pricing documentation lists:
| Model | Credits per second | USD per second | Normalized 10 seconds |
|---|---|---|---|
gen4_turbo |
5 | $0.05 | $0.50 |
gen4.5 |
12 | $0.12 | $1.20 |
For Gen-4.5, I would want a measurable quality or acceptance-rate gain to justify the higher generation price.
Runway also routes third-party models, including Veo and Seedance. Those belong in the gateway category: record the selected model and realized credit cost from response metadata rather than treating all Runway jobs as the same pricing surface.
Seedance: budget from usage, not a fabricated fixed rate
BytePlus’s ModelArk pricing page uses configuration-dependent token metering for Seedance 2.0. These are per-video ranges for video-input workloads:
| Model | 480p | 720p | 1080p | 4K |
|---|---|---|---|---|
| Seedance 2.0 Mini | $0.19–$0.42 | $0.41–$0.91 | Not supported | Not supported |
| Seedance 2.0 Fast | $0.30–$0.66 | $0.64–$1.43 | Not supported | Not supported |
| Seedance 2.0 | $0.39–$0.86 | $0.84–$1.86 | $2.06–$4.57 | $4.20–$9.33 |
Shorter inputs correspond to the lower end. Longer inputs and higher resolutions increase the charge. I would use the live calculator or provider-reported usage for a production estimate, not divide these ranges into a supposedly universal second rate.
ByteDance has also announced Seedance 2.5. Compared with Seedance 2.0, it expands single-generation output from up to 15 seconds to up to 30 seconds. ByteDance says one task can accept up to 30 images, 10 video clips, and 10 audio clips, supporting larger reference sets for continuity, storytelling, and editing.
Until the production route, supported parameters, and live billing are confirmed, I would leave Seedance 2.5 out of fixed-price comparisons.
Sora: useful migration baseline, poor new dependency
OpenAI’s official pricing lists:
| Model | Resolution | Standard | Batch | Normalized 10 seconds, standard |
|---|---|---|---|---|
sora-2 |
720p | $0.10/s | $0.05/s | $1.00 |
sora-2-pro |
720p | $0.30/s | $0.15/s | $3.00 |
sora-2-pro |
1024p | $0.50/s | $0.25/s | $5.00 |
sora-2-pro |
1080p | $0.70/s | $0.35/s | $7.00 |
The video documentation covers asynchronous jobs, synchronized audio, image guidance, editing, extensions, and outputs of up to 20 seconds.
Sora 2 Pro at 1080p costs about 2.33× its 720p rate. More importantly, OpenAI’s deprecation schedule says the Videos API, both models, and listed snapshots will be removed on September 24, 2026. The table lists no direct replacement.
For an existing integration, I would use August and early September as the migration window. Keep the same prompts and reference assets, then test at least one low-cost route, one native-audio route, and one higher-quality candidate across Veo, Kling, Seedance, or Runway.
Pick candidates by workload, not model reputation
These are starting points based on documented pricing and capabilities, not a quality ranking.
| Workload | Routes I would test first | What decides the result |
|---|---|---|
| Silent drafts | Veo 3.1 Lite; Runway Gen-4 Turbo | Access, prompt adherence, acceptance rate |
| Fast iteration | Veo 3.1 Fast; Runway Gen-4 Turbo; Kling 3.0 Turbo | Queue time, retries, visual consistency |
| Native-audio clips | Supported Veo audio routes; Kling 3.0; Seedance 2.0 | Lip sync, languages, endpoint support, audio billing |
| Reference-driven ads | Kling 3.0; Seedance 2.0; supported Veo routes | Subject fidelity, input charges, moderation, callback reliability |
| Higher-resolution delivery | Veo 3.1; Seedance 2.0; Kling 3.0 | Delivered resolution, compression, editing effort, usable cost |
| Sora migration | Existing Sora route plus at least two replacements | Controlled comparison and completion before September 24, 2026 |
Input assets deserve their own budget check. I would not reuse a text-to-video estimate for image-to-video or video-reference jobs unless the provider confirms identical billing rules.
The same goes for unsuccessful jobs: record terminal state and reported charge for failed, moderated, cancelled, and timed-out requests. Submission count is not an invoice.
My evaluation plan: 30 jobs, then repeat
A practical first pass is 30 jobs:
- Ten text-to-video prompts covering people, products, camera motion, visible text, and multi-subject scenes.
- Ten image-to-video jobs using the same licensed reference assets.
- Ten workload-specific jobs, such as audio ads, product demonstrations, loops, or multi-shot sequences.
Where supported, keep duration, resolution, aspect ratio, references, audio settings, seed behavior, and reviewer rubric constant.
I would compare supported durations directly and normalize cost separately. Stitching outputs merely to force every provider into a 10-second test changes the workload.
Log enough to reconstruct both the job and the bill
| Field | Reason |
|---|---|
| Provider, route, model ID, version date | Avoid comparing different releases or aliases unknowingly |
| Submitted parameters and input assets | Reconstruct resolution, audio, and reference costs |
| Task ID and terminal state | Distinguish completion, failure, cancellation, and moderation |
| Provider-reported usage and charge | Measure billed usage rather than infer it from submissions |
| Queue and generation time | Separate interactive suitability from batch suitability |
| Reviewer verdict: accepted, fixable, rejected | Establish the usable-output denominator |
| Retry and fallback reason | Explain why a low-rate route becomes expensive |
| Review and editing minutes | Expose labor costs |
| Output-copy timestamp | Track preservation of outputs behind temporary URLs |
For each prompt class, report:
- Acceptance rate and cost per accepted clip
- p50 and p95 completion time
- Technical failure, moderation, and fallback rates
- Reviewer minutes
Run at least two passes because generation is stochastic.
Before increasing volume, add idempotency, webhook retries, timeout handling, per-job cost ceilings, version monitoring, and automatic copying of temporary outputs into durable storage. Also verify callback behavior, failure billing, output retention, and lifecycle status for the exact route.
That is the comparison I would trust: not which model advertises the cheapest second, but which integration delivers acceptable clips at a predictable total cost.
Originally published at cometapi.com
Top comments (0)