The last time I integrated a video model, the job table had one row type. A prompt went in, a clip came out, and every user revision was a brand-new job with no relationship to the previous one. That worked because the model only had one operation. The schema matched the tool.
The published capabilities of Kling 4.0, announced September 28, 2026, break that assumption in a useful way. There are three operations now, and they have parent-child relationships. This post is about what that does to the backend. No code, just the design decisions.
Three job types, one lineage
generate takes a prompt plus optional references and produces a shot of 3 to 30 seconds. extend takes an existing output and continues it forward or backward, preserving character and background. edit takes up to five input videos, one of them the main clip, and modifies only the named attributes: expression, motion, camera, style or background.
The obvious model is a jobs table with a kind column and a nullable parent_output_id. Extend and edit always have a parent; generate never does. That gives you lineage for free, which matters because users will want to see the revision tree and roll back.
Reference assets need role slots
Omni Reference accepts up to fifteen inputs per request, but not fifteen of anything. The composition limit is up to ten images, up to five videos totaling 30 seconds or less, and up to seven elements that may include voice. Inside the prompt each file is addressed by handle: @image1, @video1, @audio1.
So the reference set is not a flat list. Store it as a structured object with typed slots and validate the per-type counts and the aggregate video duration before you submit. Reject early on the client, because a failed generation still costs a round trip. Keep the handle-to-asset mapping on your side so a prompt template can say "@image1 is the character" and you inject the right file at submit time.
Image-to-video separately accepts up to ten keyframe images. Treat keyframes as an ordered list, not a set; order is semantic.
Prompt budget changes the UI
Prompts can run to 8,000 tokens. If your product exposes a single text box, you are wasting the budget. The vendor's own guide suggests structuring by subject, action, background, camera, lighting and sound, with references called out explicitly. A sectioned form that compiles into one prompt string will produce better results than free text, and it gives you fields you can validate.
Concurrency is plan-dependent
Starter allows one concurrent job, Standard four, Premium eight. Read the plan at queue time and cap in-flight jobs per account accordingly. Submitting four jobs on a Starter plan does not fail loudly; it just serializes on the vendor side, and your users will see inconsistent latency with no explanation.
Audio is in the artifact
Stereo audio is generated with the picture, with lip-synced dialogue in nine languages. If your pipeline currently has a TTS step and a mux step, both become optional. Keep them behind a flag for models that do not emit audio, but do not run them by default.
Output handling
Resolutions are 720p, 1080p and 4K, with 10-bit HDR coming to the top two. Store the resolution and bit depth as output metadata, because a 4K HDR file needs different transcoding presets than a 720p SDR one. Every plan ships without a watermark, so there is no per-tier post-processing branch.
Pricing for capacity planning
Annual billing: $21/month for 180 credits (about 11 videos), $56 for 580 (about 36), $90 for 1,300 (about 81). Roughly 16 credits per video, though extend and edit may price differently; the public page does not say. There is a pay-as-you-go option, which is the right default while your usage curve is still unknown.
What is still undocumented
Rate limits, webhook or polling semantics, and per-operation credit costs are not on the public page. The full model ships in October; the Flash build available now is 720p and tops out at 20 seconds, so benchmark against Flash with the expectation that numbers will move. One more note: kling4.kr is an independent service, not the official Kling AI site, and says so. Its demo cards list inputs, prompt and mode per example, which is the fastest way to see what the three job types actually produce.
Top comments (0)