Checked 21 August 2026. Two things most Wan 2.5 articles get wrong: it is not open weights — the model shipped as an API-only preview checkpoint, while Wan 2.1 and 2.2 are the versions actually released under Apache 2.0 — and it has been superseded twice, by Wan 2.6 and Wan 2.7. Specs below are from Alibaba Cloud Model Studio's text-to-video API reference; per-second rates are from API aggregators and should be confirmed against your own billing.
Wan 2.5 is Alibaba's text-to-video and image-to-video model, notable for being the first in the Wan line to generate synchronised audio — dialogue with lip-sync, ambient sound and music — in the same pass as the video. It produces 5–10 second clips at up to 1080p and 24fps, costs roughly $0.07 per second of output through the API, and is accessed as wan2.5-t2v-preview or wan2.5-i2v-preview. It is not downloadable, and two newer Wan versions now exist.
Wan 2.5 Specs
| Property | Wan 2.5 |
|---|---|
| Model IDs | wan2.5-t2v-preview, wan2.5-i2v-preview |
| Clip length | 5–10 seconds |
| Resolutions | 480p, 720p, 1080p |
| Frame rate | 24fps |
| Audio | Generated by default — dialogue, lip-sync, ambience, music |
| Aspect ratios | 16:9, 9:16, 1:1, 4:3, 3:4 |
| Access | Alibaba Cloud Model Studio / DashScope API, plus third-party gateways |
| Open weights | No |
The 9:16 support and native audio are the two properties that matter for short-form ad creative. Most models of this generation give you silent video and leave you to add a voiceover afterwards; Wan 2.5 producing lip-synced dialogue in the same generation removes a step, at the cost of giving you less control over the voice.
Is Wan 2.5 Open Source? No — and This Trips People Up
This is the most repeated error about Wan 2.5. The Wan family does have open-weights releases: Wan 2.1 and Wan 2.2 were published under Apache 2.0 on github.com/Wan-Video, with inference code and weights you can download and run locally. Wan 2.5 was not. It shipped as a commercial API preview checkpoint, with no weights on Hugging Face or GitHub.
So "Wan is open source" and "Wan 2.5 is open source" are different claims, and only the first one is true. If your reason for looking at Wan is that you want to self-host and pay nothing per clip, the version you actually want is Wan 2.2, not 2.5 — you trade the synced audio and the longer clips for zero marginal cost and full control.
Real Cost per Clip
Wan 2.5 bills on output duration, capped at ten seconds — the rule is min(10, rounded duration), so an eleven-second request is still billed as ten.
| Model | Rate | 5s clip | 10s clip |
|---|---|---|---|
| Wan 2.5 | $0.0708/s | $0.35 | $0.71 |
| Wan 2.6 | $0.0708/s | $0.35 | $0.71 |
| Wan 2.6 Flash | $0.021–$0.069/s | $0.11–$0.35 | $0.21–$0.69 |
| Wan 2.7 (720p) | $0.086/s | $0.43 | $0.86 |
| Wan 2.7 (1080p) | $0.144/s | $0.72 | $1.44 |
OpenAI sora-2 (720p) | $0.10/s | $0.50 | $1.00 |
OpenAI sora-2-pro (1080p) | $0.70/s | $3.50 | $7.00 |
Wan rates as published by API gateways tracking Alibaba's endpoints, August 2026. Alibaba's own list price varies by region and by whether you buy through Model Studio directly; confirm before you budget.
The headline: a ten-second 1080p clip with audio from Wan 2.5 costs about 71 cents. The same length from #f4f4f5] px-1.5 py-0.5 text-sm font-mono font-medium text-[#7d6cea]">sora-2-pro at 1080p was ten times that — and sora-2 is being [removed on 24 September 2026 anyway. On cost per finished second with audio, Wan is currently the cheapest credible option in this class.
Should You Use Wan 2.6 or 2.7 Instead?
Probably, and this is the part the version-specific articles skip. Alibaba shipped two newer models since:
- Wan 2.6 — same $0.0708/s as 2.5, but the clip range widens to 2–15 seconds, and audio remains on by default. Strictly better value than 2.5 at an identical rate. There is also a Flash tier as low as $0.021/s for faster, slightly lower-quality output, which is the right choice for high-volume creative testing where you are throwing most variants away.
- Wan 2.7 — 2–15 seconds, 720p and 1080p, audio either auto-generated or supplied as your own file. That last point matters for ads: you can bring a real voiceover instead of accepting a generated one. It costs more, at $0.086/s for 720p and $0.144/s for 1080p.
Practical read: use 2.6 Flash for volume testing, 2.7 for the variants you are actually going to spend on, and 2.2 if you want to self-host and pay nothing per clip. There is no scenario in August 2026 where 2.5 is the right pick over 2.6 at the same price — the only reason to run 2.5 specifically is an existing integration that pins the model ID.
Is Wan 2.5 Good for UGC Ads?
Partly. Where it holds up: vertical 9:16 output is native rather than cropped, ten seconds is enough for a hook plus one beat, and the synced dialogue means a talking-to-camera clip does not need a separate voice pass. Where it does not: the audio is generated, so you get a voice rather than your voice, ten seconds is short for a full ad structure, and — as with every raw model — subject consistency across multiple clips is something you manage yourself with reference images.
That last constraint is the real one. A UGC ad is usually three to five shots of the same person. Raw models generate shots; keeping one face consistent across all of them, then cutting them together with captions and a CTA, is a separate job. The AI video generator buyer's guide covers which tools handle which half, and Sora 2 vs Veo 3.1 covers the same trade-off among the frontier models.
Wan 2.5 vs the Rest, in One Line Each
| Model | Best at | Weak at |
|---|---|---|
| Wan 2.5 / 2.6 | Cheapest synced-audio video, native 9:16 | Short clips, generated-only voice |
| Wan 2.7 | Bring-your-own audio, up to 15s | Twice the price of 2.6 |
| Wan 2.2 | Free to run, full control | Self-hosted GPU, no audio, fixed 5s |
| Veo 3.1 | Prompt fidelity, motion quality | Cost, access limits |
sora-2 | Was strong on physics | Removed 24 Sept 2026 |
| Higgsfield | Camera control, presets | Credit system, no public REST API |
Clips Are Not Ads
Every model above hands you a clip. Turning clips into an ad that performs means a hook that lands in three seconds, a script written for muted autoplay, captions burned in, and a cut that holds attention to the CTA. Creetr does that end to end from a product link — AI actor, script, hook, captions, auto-edit — on a free plan with no watermark on export. If you want to work the other way round and nail the hook before you generate anything, the video hook generator is free and needs no signup, and the script length calculator will tell you whether your script actually fits in ten seconds before you pay for the render.
