
Wan is Alibaba's AI video model, and it split itself in two. Its open-weight 2.1 and 2.2 models run free on your own GPU under an Apache 2.0 license, the best local video AI you can self-host. Have a capable GPU? Use them today. If not, the paid, API-only 2.5 and 2.6 add audio and native 1080p, but only match the top closed models like Kling and Veo, they do not beat them.
Wan, sometimes written as WAN or Wan Video, is the text-to-video and image-to-video model family from Alibaba's Tongyi lab.
It first drew attention in early 2025 because Alibaba did something most video labs refused to do: it released the actual model weights under a permissive open-source license.
That single decision made Wan the default choice for anyone who wanted to generate video locally, fine-tune it, or build it into a product without paying per clip.
Since then the family has grown fast, and it has also split. Some Wan models are open and free to run yourself. The most capable ones are now closed and only reachable through a paid API.
Understanding that split is the whole point of this review, because which Wan you mean changes the answer to almost every question about cost, quality, and control.
This is the most important thing to understand before you commit to Wan.
The open-weight track is Wan 2.1 (February 2025) and Wan 2.2 (July 2025). Both ship under an Apache 2.0 license, which means the weights are public, you can run them on your own hardware, and you can use the output commercially with no per-clip fee.
The API-only track is Wan 2.5 (September 2025) and Wan 2.6 (December 2025). These weights were never published. You reach them through Alibaba Cloud's platform or resellers, and you pay by the second.
The community consensus is blunt: the open-source era for Wan ended with 2.2. Everything genuinely new after that, native audio, 1080p, longer clips, multi-shot, arrived behind the paywall.
So when a tool or a tutorial says it uses Wan, ask which one. Free and local means 2.1 or 2.2. Paid and cloud with sound means 2.5 or 2.6.
These are the reason Wan matters, and they are still the best free video models you can run yourself in 2026.
Wan 2.1 shipped in two sizes, a small 1.3B model and a full 14B model, both using a diffusion transformer. The small model is the headline act: it generates 480p to 720p clips and runs on a consumer GPU with as little as about 8GB of VRAM, so a mid-range gaming card can make video at home.
Wan 2.2 pushed the architecture to a mixture-of-experts design that splits the work between a high-noise and a low-noise expert, which gets more quality out of less memory. Its hybrid 5B model runs comfortably on a 24GB card like an RTX 4090 and does both text-to-video and image-to-video at 720p and 24fps.
Neither open model generates sound. Wan 2.2 has a separate speech-to-video companion model for lip-sync, but out of the box these are silent generators, and you add music, voiceover, and effects in post.
Both slot cleanly into ComfyUI and the wider open-source tooling, which is why they show up inside so many other apps and pipelines. If you want unlimited iteration with zero marginal cost and full data privacy, this is the only serious game in town.
The closed models are where Wan started chasing the commercial leaders, and where it started charging.
Wan 2.5, launched as a preview in September 2025, was the turning point. It added native, synchronized audio, dialogue, sound effects, and music generated in the same pass as the picture, plus lip-sync, native 1080p, and clips up to about 10 seconds. It is trained jointly on audio and visual data rather than bolting sound on afterward.
Wan 2.6 followed in December 2025 and pushed further into production territory. It keeps the single-pass audio and 1080p, and adds reference-to-video for character consistency, multi-shot storytelling, and a wider set of aspect ratios covering 16:9, 9:16, 1:1, 4:3, and 3:4.
The catch is that neither runs locally. They live on Alibaba Cloud and on resellers like Replicate and fal.ai, and every second costs money.
In other words, the moment you want the features that make a clip feel finished, sound and full HD, you leave the free world and enter the same paid arena as Kling and Hailuo.
Wan's quality reputation is strong for its class, with a couple of honest caveats.
Prompt following and motion are consistent highlights. The models track what you asked for and produce smooth, coherent movement, and Wan has long punched above its size on physical realism for the amount of compute it uses.
The open 1.3B model is the standout value story in all of AI video: usable 720p motion from a card most gamers already own is something no closed rival offers.
The weaknesses track the version split. The open models cap at 720p, run about five seconds, and are silent, so they are prototyping and iteration engines, not one-click final-cut tools.
The usual failure modes apply too. Like every current model, Wan can still mangle hands and fingers, garble any on-screen text, and drift on fast or complex motion, and the open models show these cracks sooner than the paid ones because they run at lower resolution and shorter length.
In practice Wan is strongest on slow, physical shots, a dolly across a landscape, a product rotating on a table, liquid pouring in close-up, and weakest on multi-person dialogue and busy crowd scenes. Prompt camera moves explicitly too: "slow dolly forward" lands far better than a vague "cinematic," and Wan repays the specificity with steadier motion.
Even the paid 2.5 and 2.6 models, while competitive, do not clearly beat the very top of the market. For raw single-subject realism many creators still rank Hailuo and Google Veo ahead, and for cinematic camera work Kling remains the reference point. Wan wins on openness and price-to-capability, not on a realism leaderboard.
Wan has the widest price range in AI video, because half the family is free.
The open-weight 2.1 and 2.2 models cost nothing to use beyond your own hardware and electricity. Download the weights, run them locally, and generate as much as you want with full commercial rights. That is the cheapest serious video generation available anywhere.
The API models are billed by the second and by resolution. On Alibaba's own platform, Wan 2.5 runs roughly $0.05 per second at 480p, $0.10 at 720p, and $0.20 at 1080p, so a 10-second full-HD clip lands around $2. Third-party hosts like Replicate and fal.ai tend to charge more, often in the $0.30 to $0.40 per second range, in exchange for an easier setup.
The practical takeaway: prototype for free on the open models, then spend on the API only for the finished, audio-carrying clips a client actually sees.
Against the paid leaders, Wan competes on a different axis.
Kling, Hailuo, Google Veo, and Runway are all closed, cloud-only, pay-per-clip services. On pure output quality at the top end, the best of them still edge out Wan's API models, and their apps are more polished.
What none of them offer is a free, open, self-hostable model with commercial rights. That is Wan's moat. If you need to run video on your own infrastructure, keep data private, fine-tune on your own footage, or embed generation in a product without a per-clip bill, Wan is effectively the only option at this quality level.
So the comparison is not really "which clip looks best." It is "do you value ownership and zero marginal cost, or do you value the last 10 percent of polish and native audio out of the box." For a fuller picture of the closed leaders, see our best AI video generators guide and our Google Veo review.
Wan fits a specific set of teams unusually well.
If you are a developer or a technical founder with a GPU, the open 2.1 and 2.2 models are a gift: free, unlimited, private iteration you can wire straight into a product or pipeline.
If you are a startup that needs finished, sound-carrying clips for ads or launches but wants to keep costs sane, the workflow is to draft ideas for free on the open models, then render the keepers through the 2.5 or 2.6 API.
If you are a non-technical creator who just wants the best possible clip with the least friction, and you do not care about local control, a polished closed app like Kling, Hailuo, or Veo will serve you better than fighting with weights and ComfyUI.
The starting point depends on which track you want.
For the open models, the weights live on Hugging Face under the Wan-AI organization. The easiest local setup is to install ComfyUI, drop the Wan 2.2 checkpoint into your models folder, and load a Wan text-to-video workflow. Begin with the 5B hybrid model if you have a 24GB card, or the 1.3B model if you are on 8GB, keep clips to five seconds at 720p, and iterate freely because nothing you generate costs a cent.
For the paid models, you do not need a GPU at all. Reach Wan 2.5 and 2.6 through Alibaba Cloud Model Studio, or through resellers like Replicate and fal.ai if you prefer a simpler API and do not mind paying a premium per second.
The handoff between the two is the real workflow. Nail the prompt, framing, and camera move on the free open model until a five-second draft looks right, then feed that same prompt, and often the last frame as an image input, into the 2.5 or 2.6 API to render the final 1080p clip with sound. You pay only for the shots that already earned it.
| Version | Released | Open? | Resolution | Audio | Runs locally |
|---|---|---|---|---|---|
| Wan 2.1 | Feb 2025 | Yes (Apache 2.0) | Up to 720p | No | Yes (1.3B from ~8GB VRAM) |
| Wan 2.2 | Jul 2025 | Yes (Apache 2.0) | 720p 24fps | No (separate S2V model) | Yes (5B on 24GB) |
| Wan 2.5 | Sep 2025 | No (API only) | Up to 1080p | Native, synced | No |
| Wan 2.6 | Dec 2025 | No (API only) | 1080p | Native, single pass | No |
Alibaba has also begun releasing select Wan 2.7 models as open-weight in 2026 and has pre-announced Wan 3.0, so the open track is not dead, it just trails the paid one.
Wan is the most important open project in AI video, and in 2026 it is best understood as two tools wearing one name.
The open-weight 2.1 and 2.2 models are the clear winners of the free-and-local category. Nothing else lets you generate this quality of video on your own hardware, at no marginal cost, with full commercial rights. For developers, tinkerers, and cost-sensitive teams, that is genuinely category-defining.
The paid 2.5 and 2.6 models are good, and their native audio and 1080p make them viable for finished work, but they do not dethrone the best closed leaders on pure quality. Reach for them when you need Wan's ecosystem plus sound, not because they beat Kling or Veo.
Bottom line: if you can run a GPU, Wan should be in your stack today. If you only want the single best-looking clip with zero setup, pay for a closed leader instead.
Partly. Wan 2.1 and 2.2 are open-weight under an Apache 2.0 license, so they are free to run on your own GPU with full commercial rights. Wan 2.5 and 2.6 are API-only and billed by the second.
The older models are. Wan 2.1 (Feb 2025) and 2.2 (Jul 2025) shipped their weights under Apache 2.0. The newer 2.5 and 2.6 were never open-sourced, so the community view is that the open era ended with 2.2, though select Wan 2.7 models are open again in 2026.
Yes, for the open models. Wan 2.1's small 1.3B model runs on a consumer GPU with about 8GB of VRAM, and Wan 2.2's 5B hybrid model runs on a 24GB card like an RTX 4090. The API-only 2.5 and 2.6 cannot run locally.
Only the paid models. Wan 2.5 and 2.6 generate synced audio, including dialogue, effects and music, in the same pass as the video. The open 2.1 and 2.2 models are silent and need audio added in post.
Not on raw quality. The best closed models like Kling, Hailuo and Google Veo still edge out Wan's paid models on realism and polish. Wan wins on being free, open and self-hostable, which none of them offer.