Omni is a workflow upgrade, not a quality one. If a project needs the same face, the same bottle, and the same studio across several shots, it earns its keep. If you want one good clip, the standard model is cheaper and often ranks higher.
The order that works: validate a single short shot, approve the references, add audio, then expand to a storyboard. Jumping straight to multi-shot sequencing is where most of my wasted credits went.
What Omni supports
The usual pitch is “video with sound.” That is misleading. Standard Kling VIDEO 3.0 also does native audio, multi-shot generation, and element binding. The real distinction is how references get organized and reused.
| Spec (official generation guide) | Omni capability |
|---|---|
| Input options | Text, images, video references, reusable elements, voice references |
| Model-level output duration | 3–15 seconds |
| Video resolution | 720p and 1080p, native 4K in supported workflows |
| Audio | Native audiovisual generation with supported character–voice binding |
| Shot control | Single-shot generation and custom multi-shot sequencing |
| Multi-image element creation | 2–4 images |
| Video-based character creation | 3–8 second character clip |
| Recommended voice sample | 5–30 seconds of clear, single-speaker speech |
| Combined image and element budget | Up to 7 without video input, up to 4 with video input |
Two traps live in that table.
Creating an element is not the same as attaching it to one request. Several views can belong to a single reusable element, and they do not count as separate element slots.
A model-level capability is not a provider-level one. A 15-second workflow in Kling’s interface and a third-party API request can expose different parameters and different duration values.
Pick by reference layout, not by name
Stop framing this as “basic vs pro.” Ask where the essential information lives.
| Dimension | Kling VIDEO 3.0 | Kling VIDEO 3.0 Omni |
|---|---|---|
| Reference organization | Element binding around an approved start frame or start/end frames | Broader composition with images, elements, and video references |
| Element relationship | Best when the important subjects already exist in the defining frame | Best when separate assets must be combined or reused |
| Character voice | Can use voice-bound elements | Visual and voice references in one workflow |
| Multi-shot generation | Supported | Supported with custom storyboard controls |
| Recommended starting point | One approved keyframe that already defines the scene | A collection of assets that must stay recognizable across shots |
Start with the standard workflow when one approved frame already carries the essential information. Reach for Omni when reference composition, asset reuse, or cross-shot continuity is the job.
The benchmarks do not favor Omni
These are point estimates for the 1080p Pro, audio-enabled entries in Artificial Analysis, checked September 9, 2026. Values after ± are published 95% confidence intervals.
| Evaluation (text-to-video / image-to-video) | Standard | Omni |
|---|---|---|
| Text-to-video Elo | 1108 ± 5 | 1090 ± 6 |
| Image-to-video Elo | 1072 ± 6 | 1060 ± 7 |
These are preference-based ratings, not generation-success percentages. The standard model holds the higher point estimate in both rows, and the image-to-video intervals overlap. Scores shift as new evaluations arrive.
So Omni’s case rests on reference handling and production continuity. It is not a universal quality ranking, and the name does not win a benchmark.
Build one reference-led shot first
Illustrative setup: a vertical ad with a presenter, a reusable bottle, and one studio setting. Not a reported generation test.
Prepare a small asset set
| Asset | Preparation | Acceptance criterion |
|---|---|---|
| Presenter | Clear views with consistent hair, clothing, and lighting | Recognizable face, unchanged wardrobe |
| Bottle | Front, side, and three-quarter views | Stable body shape, cap, color, label |
| Setting | One approved studio reference | Consistent background and light direction |
| Voice | Clean speech from an authorized speaker | Recognizable voice, no competing speech |
Fix conflicting references before you try to compensate with a longer prompt. Use only faces, recordings, and brand assets you are authorized to use.
Create and name elements
In Kling’s Element Library, create reusable assets instead of re-describing the same subject on every shot. The official workflow supports multi-image and video-based character creation.
Names like Maya, BlueBottle, and Studio are organizational aids. Selecting the saved element is what attaches the reference. For uploaded voices, use clear, single-speaker audio without music or overlapping voices.
In the prompts below, @maya, @BlueBottle, and @studio stand for elements selected in the interface. Typing the names without attaching the assets creates no reference.
Generate a single shot
Open the Omni generation workspace, pick the model, attach the assets. Start at five seconds, one camera move, no dialogue.
Create a single continuous product shot using @BlueBottle in @Studio.
The bottle stands upright on a clean tabletop. Keep its body shape,
cap, surface color, and label placement consistent with the references.
Camera: a slow, straight push-in from a medium close-up.
Lighting: soft studio light from camera left.
Action: the bottle remains stationary.
Composition: leave clear space above the bottle for text added later.
No cuts, no extra products, and no generated captions.
Inspect the beginning, middle, and end. Check the cap, silhouette, label placement, and tabletop contact. If the asset drifts, simplify the shot or improve the references before adding demands.
Add the presenter and native audio
Once the bottle shot passes, add the presenter and one short line. Kling documents speech generation in several languages, but quality tracks reference quality, timing, and prompt clarity.
Create a five-second continuous shot using @Maya, @BlueBottle,
and @Studio.
Maya stands beside the bottle, looks toward the camera, and says:
“Ready when you are.”
Use Maya’s bound voice. Keep the bottle unchanged. Use one static
camera angle. Do not add music, captions, cuts, or additional speech.
Review audio apart from picture. Confirm the right person speaks, the line is complete, and the timing matches the visible performance.
Go multi-shot only after that
Reuse the approved references and plan the sequence. Give each shot a duration and a distinct visual purpose. Custom storyboard controls support shot-level duration, framing, camera behavior, and narrative progression.
| Time | Purpose | Visual instruction | Audio instruction |
|---|---|---|---|
| 0–5 seconds | Establish the product | Close-up of the bottle, slow push-in | Quiet room tone |
| 5–10 seconds | Introduce the presenter | Medium shot of Maya beside the bottle, static camera | “Ready when you are.” |
| 10–15 seconds | Finish on the product | Clean hero shot, hold the final composition | Room tone |
Create a three-shot product advertisement using @Maya, @BlueBottle,
and @Studio.
Continuity rules:
Use the same presenter, outfit, bottle, and studio throughout.
Keep the bottle’s shape, cap, color, and label unchanged.
Maintain the same lighting direction.
Use Maya’s bound voice only.
Do not add people, products, captions, or extra dialogue.
Follow the separate shot descriptions and durations.
End on a stable product composition.
Check transitions as closely as the shots. Compare the last clear product view before each cut with the first clear view after it. Add legal copy, exact prices, subtitles, and brand typography in an editor so wording stays under direct control.
Video references: identity vs editing
A clip that creates a reusable character is not source footage for an edit. The first establishes identity. The second supplies existing motion and composition to transform or preserve.
For character creation, use a 3–8-second clip of one character. For video editing, Kling’s official Omni guide permits one 3–10-second source video, up to 200 MB and 2K resolution. Check the selected API route’s current limits and test a short clip before scaling.
Use the uploaded video as the source.
Change only the background to a softly lit studio.
Preserve the presenter’s identity, clothing, actions, and timing.
Preserve the bottle and the original camera movement.
Do not add cuts, gestures, objects, or dialogue.
A controlled edit request is not a guarantee of frame-perfect preservation. Decide separately whether to keep the source audio or generate new sound, then compare against the original clip.
The API path
For programmatic access, the documented identifier is kling-v3-omni. The examples follow the published schema, not a live generation test. The route documents duration values of “5” and “10” for the relevant modes. Do not assume a 15-second interface workflow maps to the same request.
Submit a task
export COMETAPI_KEY="YOUR_COMETAPI_KEY"
curl --fail --silent --show-error \
--request POST \
"https://api.cometapi.com/kling/v1/videos/omni-video" \
--header "Authorization: Bearer ${COMETAPI_KEY}" \
--header "Content-Type: application/json" \
--data-raw '{
"model_name": "kling-v3-omni",
"prompt": "A blue reusable bottle stands on a clean studio tabletop. One continuous shot with a slow straight push-in. Soft light from camera left. No people, no cuts, no captions.",
"mode": "std",
"aspect_ratio": "9:16",
"duration": "5",
"sound": "off"
}'
For a first-frame reference, merge this fragment into the request and swap the sample URL for a reachable image:
{
"image_list": [
{
"image_url": "https://your-public-host.example/bottle.jpg",
"type": "first_frame"
}
],
"prompt": "Starting from <<>>, create one continuous slow push-in. Preserve the bottle and studio composition. No cuts or captions."
}
The API’s <<>> syntax is not the same as selected @Element references in the interface. Documented sound values are on and off.
Retrieve the result
Store the returned data.task_id, then query the status route.
TASK_ID="REPLACE_WITH_RETURNED_TASK_ID"
curl --fail --silent --show-error \
"https://api.cometapi.com/kling/v1/videos/omni-video/${TASK_ID}" \
--header "Authorization: Bearer ${COMETAPI_KEY}"
Documented states: submitted, processing, succeed, failed. A creation response with code: 0 means the request was accepted, not that the video is ready. On success, read the result from data.task_result.videos[0].url.
Use bounded polling with backoff, keep the task ID next to the prompt and settings, and read the failure message before retrying. A client timeout may still have created a billable task.
Cost
Price by configuration, not a generic per-video figure. Figures checked September 9, 2026, and they can change.
Kling’s official Omni guide lists credit rates. The dollar figures below use different billing units, and credits do not convert to dollars without Kling’s applicable credit purchase rate.
| Omni configuration | Credits/second | Credits/5s | Credits/10s |
|---|---|---|---|
| 720p, no video input, native audio off | 6 | 30 | 60 |
| 1080p, no video input, native audio off | 8 | 40 | 80 |
| 1080p, no video input, native audio on | 12 | 60 | 120 |
| 1080p, video input, native audio off | 16 | 80 | 160 |
No 4K credit rate is listed, and native audio with video input is marked unsupported in that pricing table. Confirm current in-product rates before ordering.
| Omni API configuration | 5-second output | 10-second output |
|---|---|---|
| 720p, no video input | $0.336 | $0.672 |
| 1080p, no video input | $0.448 | $0.896 |
| 1080p, native audio | $0.560 | $1.120 |
| 1080p, video input | $0.672 | $1.344 |
| 4K, no video input | $1.680 | $3.360 |
These are separate API configurations. Do not add rows together, and do not infer that every combination of resolution, audio, and video input exists. A 4K price also does not establish which parameter enables 4K.
Kling added native 4K after the original launch. Approve content and motion before paying for higher-resolution iterations.
Budget against accepted clips
Cost per accepted clip = total generation spend ÷ accepted clips.
Three five-second attempts at the 1080p native-audio rate cost 3 × $0.560 = $1.680. If only one passes, that clip’s effective generation cost is $1.680. This is budgeting math, not a measured retry rate.
Keep Kling application credits separate from dollar billing. Kling’s credit-cost guide treats duration, resolution, audio, and workflow as independent planning factors.
Troubleshooting
| Problem | First check | Adjustment |
|---|---|---|
| Character changes between shots | Conflicting references or wardrobe descriptions | Reuse one approved element, simplify the shot |
| Product shape or label changes | Reference quality and demanding motion | Add a clearer product view, test a stationary shot |
| Wrong voice or extra speech | Voice binding and conflicting instructions | Keep one speaker and one short line |
| Unwanted cuts | Storyboard settings and scene changes | Test one continuous shot without narrative jumps |
| API rejects the request | Route-specific schema, types, and duration | Return to the documented minimal request |
| Task exists but has no video URL | Generation may still be pending | Query the existing task instead of creating another |
| Higher resolution still looks wrong | The motion or identity problem remains | Revise the scene before increasing output quality |
Change one variable at a time. Record the assets, prompt, settings, output, and reason for rejection so improvements trace back to a specific decision.
FAQ
Is Omni always better than the standard model? No. Preference benchmarks can favor the standard model. Omni pays off when its broader reference workflow improves reuse, continuity, or editing control.
Can every workflow hit 15 seconds? No. Fifteen seconds is a model-level capability in supported workflows. Individual interfaces and routes may expose shorter options.
Do @Element names work in an API request? Not automatically. Selected elements in the interface and an API’s image-reference syntax are different mechanisms. Follow the route’s schema.
How do I estimate a production budget? Track total generation spend, divide by accepted clips, and include rejected attempts, higher-resolution reruns, storage, and post-production in the working number.
Originally published at cometapi.com
Top comments (0)