The landing page for Wan 3.0 promises "up to 20 reference assets", but the underlying API does not accept a single list of twenty items. It splits those inputs across three separate arrays, each with its own count, file type, and duration cap.
What one wan3.0-video request actually takes
The API uses a single model identifier, wan3.0-video, across every generation mode. The fields you include in the payload determine which mode runs:
-
prompt: Text string up to 20,000 characters describing the shot. -
resolution: Output size, accepting480P,720P, or1080P. The provider default is1080P. -
duration: Video length from 2 to 30 seconds, or-1to let the model select the duration. Default is 5. -
audio: Boolean controlling sound generation, defaulting totrue. Sound generates in the same pass as the video frames rather than as a post-process. -
generation_type: Processing mode flag, set toframefor frame-driven generation orreferencefor asset-conditioned generation. -
image_urls: Array of image URLs for single-frame generation or reference images. -
image_with_roles: Array of image objects containing URLs and explicit roles (first_frame,last_frame, orreference_image). -
video_urls: Array of reference video URLs. -
audio_urls: Array of reference audio URLs. -
file_url: URL for one reference document, up to 100 MB and 50 pages. -
link_url: URL for one public web page reference. -
size: Output dimensions and aspect ratio. -
seed: Integer seed for output repeatability. -
nsfw_check: Content filtering toggle. -
watermark: Watermark inclusion toggle.
Payload construction determines the generation mode. Passing two frames into image_with_roles triggers first-and-last frame interpolation. Passing a single image into image_urls with generation_type: frame runs standard image-to-video. Passing image URLs with generation_type: reference runs reference-conditioned generation.
The twenty is three arrays, not one
Alibaba groups reference handling under the name Omni-Creation. The advertised twenty-asset capacity is divided into three fixed buckets:
| Group | How many | Other limits |
|---|---|---|
| Reference images | up to 10 | 20 MB each; JPG, JPEG, PNG, BMP or WebP |
| Reference video | up to 5 clips | 1-15 seconds each, 15 seconds in total; MP4 or MOV |
| Reference audio | up to 5 clips | 1-15 seconds each, 15 seconds in total; WAV or MP3 |
The total of 10 images, 5 video clips, and 5 audio clips makes up the advertised twenty. Because each category is isolated, you cannot submit 20 reference images or a single 20-second reference video.
References bind to prompt text using explicit tags: @Image1, @Video1, and @Audio1. The numeric index maps to the position of the asset inside its respective request array. The API also accepts standalone reference audio in audio_urls without any attached visual reference.
One code block that does the arithmetic
Pricing scales along resolution and duration tiers without offering audio discounts:
Resolution cost scaling (fixed duration):
480p = 1x (baseline)
720p = 2x
1080p = 4x
Duration cost scaling (fixed resolution):
5s = 1x (baseline)
30s = 6x
Audio setting:
audio: true = 1x
audio: false = 1x (disabling audio does not reduce cost)
Scaling a batch from 480p at 5 seconds up to 1080p at 30 seconds increases resource consumption by a factor of 24. You can calculate specific requirements using the Wan 3.0 cost calculator.
What we shipped through it, and what the clips showed
We tested the pipeline by sending four clips through the API at 480p, 16:9 aspect ratio, 5 seconds per clip, with audio enabled. Prompts were transmitted exactly as written.
Clip 1 structured the prompt using a five-element order: shot type, subject, action, lighting, and sound. The render delivered the ceramicist lifting the bowl from the wheel, low window light entering from the left, background studio elements in shadow, and a visible camera push-in.
Generated with Wan 3.0, 480p, 5 seconds, audio on, from the prompt above, unedited — 2026-09-05.
Clip 2 formatted scene ambience as active clauses rather than descriptor lists: rain striking the glass, sodium streetlights crossing the subject's face sequentially, interior lighting restricted to an overhead strip, and a locked-off camera position.
Clip 3 supplied two reference stills from Nano Banana 2 tagged as @Image1 and @Image2. The yellow jacket, dark hair, and bag strap mapped from the first image, while the magenta neon, fire escape, and wet ground mapped from the second. Because text-to-image calls retain no state between separate generations, the open-shutter reference still had to be created by editing the closed-shutter still directly.
Generated with Wan 3.0, 480p, 5 seconds, audio on, with two stills attached as @Image1 and @image2 — 2026-09-05.
A single reference test confirms that label binding functioned in this instance, but it does not isolate whether the model resolved the assets via the @Image tokens or via matching text descriptions like "the courier" and "the alley".
Clip 4 tested frame-to-video with defined start and end points. The initial frame showed a closed shutter, which raised across the runtime before settling on the final framing. An additional run on 2026-08-16 evaluated a full 30-second continuous shot at 1080p, executing an unbroken crane pull-back from a rooftop herb garden to a dawn skyline. Full prompt text and media outputs are documented in the Wan 3.0 prompt guide.
What our own front end does not expose
The interface on the Seadanse model page provides access to text-to-video and frame-to-video generation, durations from 4 to 30 seconds, three resolution profiles, six aspect ratio presets plus Auto, a 10,000-character prompt field, and native audio generation.
Certain API capabilities are omitted from the front end:
- Duration minimum: The composer enforces a 4-second minimum rather than the API's 2-second limit due to a shared schema constraint across our interface layer.
-
Document parsing: The API accepts PDF and text files via
file_url, which is not exposed in the web composer. -
Web page parsing: The
link_urlinput is not wired to the UI. - In-place video editing: Direct video transformation via API flags is omitted.
- Upload constraints: While the upload panel lists accepted file extensions (JPG, PNG, WebP, MP4, MOV), it does not display the API's per-group quantity and duration limits.
A checklist for the next model API page you read
When evaluating a video model API from its documentation, check these five constraints before writing integration code:
- Array segmentation: Verify whether multi-asset reference claims allow arbitrary file combinations or mandate fixed allocations across image, video, and audio arrays.
- Per-array duration caps: Check if video and audio inputs have aggregate duration ceilings in addition to individual file limits.
- Default resolutions: Check the default API resolution value, as provider defaults like 1080p will run at higher cost multipliers than baseline 480p settings.
- Duration floors: Compare the documented API minimum length against front-end limits to identify where interface validation schemas diverge from underlying endpoints.
-
Weight availability: Verify if model weights are downloadable. For Wan 3.0, no public weights exist on Hugging Face; open checkpoints under Alibaba's organisation stop at Wan 2.2 releases such as
Wan2.2-TI2V-5B-Diffusers.
The author works on Seadanse, which runs Wan 3.0.


Top comments (0)