Alibaba Wan3.0 AI Video Model Unveiled, Generates Up to 30-Second Videos from Multimodal Inputs
Alibaba Wan3.0 expands AI video generation with support for clips lasting up to 30 seconds and multimodal reference inputs including text, images, videos, audio, web pages, PDFs, and PowerPoint presentations.
- Longer AI Videos: Wan3.0 can natively generate videos lasting up to 30 seconds, twice the 15-second maximum clip duration supported by the previous Wan2.7-Video model, giving creators more room for continuous scenes, complex camera movements, and storytelling.
- Multimodal Inputs: The new model can work with text, images, video, and audio simultaneously while also supporting web pages, PDFs, and PowerPoint presentations, potentially making it easier to transform static and information-heavy materials into engaging video content.
- Greater Consistency: Alibaba says Wan3.0 delivers improved visual continuity, reference accuracy, realistic faces and expressions, multilingual voice generation, stable user interfaces and motion graphics, and better consistency for characters, products, layouts, and audio.
Alibaba has officially unveiled Wan3.0, the latest version of its advanced AI video generation model, bringing longer video outputs of up to 30 seconds along with expanded multimodal input support for text, images, videos, audio, web pages, PDFs, and PowerPoint presentations.
The rapid development of generative AI is opening new possibilities not only for creating images and written content but also for producing increasingly sophisticated videos. With Wan3.0, Alibaba is looking to give creators and businesses a more capable AI video generation platform that can handle longer sequences while accepting a much wider variety of source materials.
Currently available in public beta, Wan3.0 is designed to combine high-fidelity video generation, visual consistency, and multimodal reference capabilities within a single model. Interested users can apply to test the technology through Alibaba Cloud's AI development platform Model Studio and the AI-native Qwen Cloud platform.
Up to 30 Seconds of AI-Generated Video
One of the most notable improvements introduced with Wan3.0 is its ability to natively generate videos lasting up to 30 seconds.
Mainstream AI video generators have traditionally focused on relatively short clips lasting anywhere from just a few seconds to around 15 seconds. Alibaba's own preceding Wan2.7-Video model, for example, supports a maximum clip duration of 15 seconds.
By doubling that limit to 30 seconds, Wan3.0 gives creators more room to develop scenes, perform more complex camera movements, and produce continuous shots without having to immediately combine multiple independently generated clips.
The longer duration could be particularly useful for social media creators, filmmakers, advertisers, educators, and businesses that need more time to communicate an idea or tell a short story within a single generated sequence.
Wan3.0 also comes with an intelligent duration feature that can recommend an optimal video length depending on the user's prompt. Additionally, video extension capabilities can help creators expand existing sequences and develop longer narrative timelines.
Multimodal Inputs from Text to PowerPoint Presentations
Beyond longer video generation, another major strength of Wan3.0 is the variety of reference materials that it can process.
The AI model supports text, images, video, and audio inputs simultaneously. It can also work with information sourced from web pages and documents such as PDFs and PowerPoint presentations.
This expanded multimodal support could prove particularly valuable for professionals and businesses with large amounts of existing static content.
Instead of manually converting presentations, reports, marketing materials, and other information-heavy documents into videos, users could potentially provide these materials directly to Wan3.0 and use the AI model to transform them into more dynamic visual content.
For content creators, meanwhile, the ability to combine several types of reference materials could provide greater control over the appearance, sound, pacing, and overall direction of generated videos.
Improved Visual Continuity and Realism
Consistency remains one of the biggest challenges in AI-generated video. Characters can unexpectedly change their appearance between frames, objects can become distorted, and backgrounds or other visual elements can drift as a generated sequence progresses.
Alibaba says Wan3.0 addresses these issues through high-precision visual continuity.
The model is designed to render realistic human faces while maintaining synchronized micro-expressions, helping characters appear more natural throughout a sequence. It can also generate natural multilingual voice outputs, expanding its potential usefulness for creators targeting audiences in different markets.
Wan3.0 is likewise designed to accurately reproduce stable software user interfaces and motion graphics.
This capability could be especially useful when creating technology-related videos, product demonstrations, instructional materials, presentations, and advertisements where interface elements and graphical details need to remain recognizable throughout the clip.
Better Control Over Characters, Products, and Audio
Wan3.0 also focuses on accurately reproducing fine details from the reference materials provided by users.
Rather than simply generating an approximation of the original reference, Alibaba says the model can maintain characters, props, product details, spatial relationships, audio characteristics, layouts, and visual styles with greater consistency.
This could be particularly important for brands using generative AI for advertising and marketing content. Keeping a product's design, proportions, placement, and other recognizable details consistent throughout a video can be crucial when presenting it to potential customers.
Voice consistency is another important capability, helping maintain the same audio identity throughout a generated sequence.
Combined with more natural movement and emotional expressions, these improvements are designed to give creators greater control while making generated clips feel more cohesive and immersive.
Applications for Creators, Businesses, and Developers
Alibaba sees potential applications for Wan3.0 across a wide variety of industries.
For filmmakers, the technology could help streamline certain parts of the creative process, including conceptualization and pre-production, while also making it easier to experiment with different visual ideas before moving into traditional production.
Creators could use the platform for short dramas and social media content, while businesses could transform existing text, images, documents, and presentations into marketing and educational videos.
The model could also have applications beyond entertainment and content creation.
According to Alibaba, Wan3.0 can potentially generate realistic simulation videos that can be used to help train autonomous driving and robotics systems, giving developers another way to produce visual data and scenarios for emerging technologies.
Alibaba Continues to Advance Its Wan AI Platform
Alibaba first introduced its Wan series of visual generation models in July 2023 and has continued to improve the technology with each succeeding generation.
The company's upgrades have focused on making AI-generated images and videos more realistic, easier to control, and more practical for creators and professional workflows.
Wan3.0 represents another significant step in that direction. Its native support for videos lasting up to 30 seconds immediately doubles the maximum duration offered by Wan2.7-Video, while its ability to accept everything from images and audio to PDFs and PowerPoint presentations greatly expands the types of materials that creators can bring into the generation process.
As Wan3.0 enters public beta testing, its combination of longer video generation, multimodal inputs, improved visual continuity, and tighter reference control could make it an interesting new option for creators, businesses, educators, and developers looking to explore the rapidly expanding possibilities of generative AI video.


