Wan 3.0 debuts at #1 on the Artificial Analysis Video Editing Leaderboard, and is a close #2 in Text to Video with Audio

Wan 3.0 is Alibaba’s new all-in-one video generation and editing model, positioned as a single system for turning multimodal creative direction into video. It generates up to 30 seconds at 1080p with native audio and accepts text, images, video, audio, documents, and web pages as creative references. The same model supports Text to Video, Image to Video, reference-based generation, and instruction-led editing, including changes to visuals, plot, dialogue, and sound.

In the Artificial Analysis Video Arena, Wan 3.0 ranks #1 in Video Editing with Audio, #2 in Text to Video with Audio, and #5 in Image to Video with Audio.

Wan 3.0 marks a large generational improvement: against the most recent Wan 2.7 version on each leaderboard, it rises from #5 to #1 in Video Editing with Audio, #6 to #2 in Text to Video with Audio, and #12 to #5 in Image to Video with Audio.

Wan 3.0 is available now in public preview through Alibaba Cloud Model Studio. Pricing starts at $0.05 per second for 480p, increasing to $0.10 for 720p and $0.20 for 1080p.