Wan 3.0 in Genosai — 30 seconds of AI video with sound by Alibaba
Wan 3.0 is Alibaba's new flagship video model, available in Genosai. It generates a coherent clip from 2 to 30 seconds in a single pass, creates the soundtrack and speech together with the visuals, accepts up to 10 reference images, and keeps your character and product recognizable across the whole scene. In Genosai the model runs online — no developer account, no API setup, from 8 credits per second.
Updated: August 31, 2026
- Up to 30 seconds in one pass — A coherent 2-to-30-second scene with no stitching — twice as long as the previous Wan 2.7.
- Character and product references — Up to 10 images per generation: your character, product, style and set stay consistent through the entire clip.
- Audio and speech built in — Ambience, music and multilingual speech are generated in the same pass as the video.
- First and last frame — Set a start image and optionally an end image — the model builds a smooth transition between them.
- Cheap 480p drafts — 8 credits per second at 480p — iterate on drafts, then rebuild the final in 1080p.
Contents
- What is Wan 3.0
- Capabilities
- Examples and prompts
- How to use it in Genosai
- Pricing
- Comparison
- Limitations and tips
- FAQ
What is Wan 3.0
Wan 3.0 is Alibaba's flagship video model, built by the same Tongyi team behind the open Wan family and the Qwen language models. The public beta opened on August 6, 2026 on Alibaba Cloud Model Studio, and the official launch followed on August 24 — one day after Alibaba raised 10.2 billion dollars to fund its AI push. During the beta the model was already used in short-drama production, advertising, tourism promos and music videos.
The headline feature is scene length and coherence. Wan 3.0 outputs up to 30 seconds of video in a single pass — twice the previous Wan 2.7 — while keeping the character, product and environment recognizable from the first frame to the last. Audio is generated together with the visuals rather than layered on afterwards: ambience, music and multilingual dialogue come out of one pass. Face detail, on-screen interface stability and motion consistency place it among the strongest releases of summer 2026.
In Genosai the model runs online: no Alibaba Cloud developer account, no API keys. Open the video studio in your browser, describe the scene — and download a finished clip with sound. Billing is per second, from 8 credits.
Capabilities
Here is what Wan 3.0 can do in Genosai in practice.
A long coherent scene
The duration slider goes from 2 to 30 seconds in one-second steps. Most models cap out at 8–15 seconds per pass: Veo 3.1 at 8, Kling 3.0 at 15. A 30-second unstitched scene is rare — besides Wan 3.0, only Seedance 2.5 offers it in Genosai. At the other end of the slider, a two-second 480p draft costs just 16 credits — perfect for testing an idea.
Character and product references
In reference mode the model accepts up to 10 images per generation. Alibaba calls the technology Omni-Reference: a character, product or set defined by your images stays itself in every shot, so a sequence reads as one piece of footage rather than random takes. For brand content this is the key feature — the logo and product do not drift from scene to scene.
Audio and speech in one pass
The soundtrack is generated with the video: city noise, a crackling oven, background music, spoken lines. Speech is multilingual, so a short talking scene comes out complete in a single generation — no separate voice-over or editing. If you do not need sound, switch it off with a toggle for a more predictable silent clip.
First and last frame
The classic keyframe mode: upload a start image, optionally an end image, and the model builds a smooth transition between them. Great for animating stills, seamless loops and precise composition control. Leave the aspect ratio on adaptive in this mode — the image dictates the proportions.
Audio references and formats
Beyond images, the model takes up to 5 audio files with a total length of 15 seconds — a voice or a music sample the soundtrack should follow. Aspect ratios: 16:9, 4:3, 1:1, 3:4, 9:16 and adaptive, where the model matches the proportions to the scene.
Examples and prompts
Below are real generations: the prompt and the clip Wan 3.0 produced from it in Genosai on the first try, with no cherry-picking. All three examples are cheap 480p drafts.
A cinematic city scene with rain ambience, 16:9, 8 seconds:
A rainy night megacity: neon signs reflect off wet asphalt, a yellow taxi drives past the camera, slow-motion spray from under the wheels, cinematic look, smooth forward dolly. Audio: rain noise, distant city hum.
Night city in the rain, 16:9, 8 seconds — generated in Genosai.io with Wan 3.0
Speech generated together with the video, 16:9, 8 seconds:
A cozy kitchen in the morning, an elderly baker in an apron takes a tray of croissants out of the oven, looks at the camera, smiles and says in Russian: "Fresh — straight from the oven!". Warm light, light steam over the pastries. Audio: the baker's voice, oven crackle.
A baker with a spoken line, 16:9, 8 seconds — generated in Genosai.io with Wan 3.0
A vertical product clip for social feeds, 9:16, 6 seconds:
Vertical social media clip: a drop of honey slowly runs down a stack of fluffy pancakes, macro shot, soft morning window light, shallow depth of field, appetizing breakfast ad. Audio: light background music.
Honey and pancakes, 9:16, 6 seconds — generated in Genosai.io with Wan 3.0
A few rules that noticeably improve results. Be concrete: subject, action, light, camera move — and describe the audio you want separately. For clips longer than 10 seconds, script the action over time ranges, or the model may loop the middle. Put spoken lines in quotes and name the language — lip articulation matches the speech much better.
How to use it in Genosai
No setup required — everything happens in the browser.
- Sign in to Genosai and open the video studio.
- Pick Wan 3.0 from the model list.
- Choose the mode: text to video, first and last frame, or references.
- Describe the scene and the sound; upload images and audio if needed.
- Set the duration slider, resolution and aspect ratio, then start the generation.
- Download the clip, or refine the prompt and regenerate.
Pricing
Billing is per second and depends on the resolution: 8 credits per second at 480p, 16 at 720p, 32 at 1080p. A five-second 480p draft is 40 credits, eight seconds — 64, and a full 30-second 1080p clip — 960 credits. For reference, Alibaba's official API charges 5, 10 and 20 cents per second — Genosai's rates are 20 percent lower. The exact cost is shown on the button before every run, and current credit packages are listed on the Pricing page. Free starter credits let you try the model after signing up.
A practical strategy: iterate on composition and pacing with 480p drafts, then rebuild the final version in 1080p — same prompt, four times the price, but only once.
Comparison
Versus Seedance 2.5: both deliver 30 seconds in one pass with audio. Seedance wins on reference count (up to 50 files, including video) and editing features; Wan 3.0 wins on draft price — 8 credits per second at 480p versus 22. For a cheap long draft or a product clip with a couple of references, start with Wan 3.0; for projects built on dozens of references and timecode edits, take Seedance.
Versus Kling 3.0: Kling caps at 15 seconds but excels at motion dynamics and multi-shot sequences. Versus Veo 3.1 Fast: Veo makes striking 8-second clips with sound, but 8 seconds is the ceiling. When a scene needs room to breathe — storytelling, product walkthroughs, travel promos — Wan 3.0 wins on length and price. And for multimodal scenes with a video input, look at Gemini Omni 1.1 Flash.
Limitations and tips
Alibaba's model also accepts video, web pages, PDFs and slide decks as scene sources, but these inputs are not yet exposed in Genosai: you get text, first and last frame, up to 10 images and up to 5 audio files. Video references were left out deliberately — the provider bills them by the combined length of input and output, which cannot yet be shown honestly on the price button. The intelligent-duration mode, where the model decides the clip length itself, is also not exposed — you set the duration explicitly with the slider.
Practical tips: remember that first/last frame and references are mutually exclusive provider scenes — the studio will guide you. An audio reference is not accepted alone: add at least one image. Write spoken lines in quotes right inside the prompt. And start with short 480p drafts — they are four times cheaper than the final render and almost always enough to see where the scene is going. The release is fresh, limits may still be tuned — trust the actual result in the studio.
FAQ
What is Wan 3.0 and who developed it?
Wan 3.0 is Alibaba's flagship video model, released to public beta on August 6, 2026 and officially launched on August 24. Its highlights are 30-second single-pass clips, multimodal references and audio generated together with the visuals. In Genosai it works online with no API setup.
How is Wan 3.0 different from Wan 2.7?
Double the single-pass duration (30 seconds versus 15), consistent characters and products across shots thanks to the reference mode, smarter duration handling, and more stable rendering of faces, interfaces and textures.
How long can a Wan 3.0 video be?
From 2 to 30 seconds in a single pass, set with a one-second-step slider. It is one of only two models in Genosai that produce a coherent 30-second scene without stitching — the other is Seedance 2.5.
Does Wan 3.0 generate audio and speech?
Yes. The soundtrack — ambience, music and dialogue — is created in the same pass as the video, and speech works in multiple languages. You can switch audio off with a toggle if you only need the visuals.
How much does Wan 3.0 cost in Genosai?
8 credits per second at 480p, 16 at 720p and 32 at 1080p. A five-second 480p draft costs 40 credits, and a full 30-second 1080p clip costs 960. That is 20 percent below Alibaba's official API pricing.
Which aspect ratios are supported?
16:9, 4:3, 1:1, 3:4, 9:16 and an adaptive mode where the model picks the proportions based on the scene and references. That covers YouTube as well as vertical Reels and Shorts.
Can I use videos or documents as references?
Alibaba's model itself accepts video, PDF and slide inputs, but these modes are not yet exposed in Genosai: you get text, first and last frame, up to 10 images and up to 5 audio files. The input set will expand.