Sign InOpen Brain
VercelEngineering PostOfficial Source

Wan 3.0 now available on AI Gateway

Wan 3.0 gives AI Gateway one video model ID for text, image, frame, and reference workflows, with async renders up to 30 seconds at 1080p and synchronized audio.

Vercel · Aug 25, 2026
Open Source Open MarkdownOpen JSON
Source Summary

Alibaba’s **Wan 3.0** is available as alibaba/wan-v3.0-video. One model handles text-to-video, image-to-video, frame conditioning, and references, producing up to **30 seconds at 30 fps** in resolutions through **1080p** with synchronized audio.

Practical Implication

Consolidate Wan integrations around the new model ID and use asynchronous generation plus a verified webhook instead of holding an HTTP request open. inputReferences accepts images, video, or audio, with media-specific source rules.

Agent-Ready Context
Alibaba’s **Wan 3.0** is available as alibaba/wan-v3.0-video. One model handles text-to-video, image-to-video, frame conditioning, and references, producing up to **30 seconds at 30 fps** in resolutions through **1080p** with synchronized audio.

Consolidate Wan integrations around the new model ID and use asynchronous generation plus a verified webhook instead of holding an HTTP request open. inputReferences accepts images, video, or audio, with media-specific source rules.

First- and last-frame conditioning cannot be combined with other references. Video and audio inputs require hosted URLs, and the material gives no pricing, latency, or output-quality measurements.
Connected Context · Feed7 Judgment

Wan 3.0 consolidates several video-generation and conditioning workflows behind one model ID, including synchronized audio and clips up to 30 seconds. That simplifies integration breadth but not request design: applications still need asynchronous jobs, verified webhooks, media-specific hosting, and explicit branching because frame conditioning cannot coexist with other references. Availability alone does not establish a preferred route.

Context Map
modelvideoaudio#generative-media#model-selection
Uncertainty
First- and last-frame conditioning cannot be combined with other references. Video and audio inputs require hosted URLs, and the material gives no pricing, latency, or output-quality measurements.