The 30-Second Era: Seedance 2.5 vs Wan 3.0 vs FLUX 3 Video
In one week of August 2026, three labs pushed AI video past the 15-second ceiling and made native audio standard. An honest comparison of what each model does and which you can actually call today.
For most of 2026, AI video meant clips of five to fifteen seconds with silent output you scored yourself. In the first week of August that changed three times over. Black Forest Labs shipped FLUX 3 Video on August 4, Alibaba opened Wan 3.0's public beta on August 6, and Seedance 2.5 brought 30-second generation with native audio into general availability.
Two things happened at once across all three: clip length broke past 15 seconds, and audio stopped being a separate step. This comparison covers what each actually does, and — the part most write-ups skip — which of them you can call today.
The Specs
| Seedance 2.5 | Wan 3.0 | FLUX 3 Video | |
|---|---|---|---|
| Lab | ByteDance | Alibaba | Black Forest Labs |
| Released | Available now | August 6, 2026 (beta) | August 4, 2026 |
| Max single-pass duration | 30s | 30s | 20s |
| Resolutions | 480p, 720p | Not published | 720p, 1080p |
| Native audio | Yes, included | Yes, synchronized | Yes, lip-synced in 12+ languages |
| Reference inputs | 9 images | Text, image, audio, video, plus documents | Images and keyframes |
| API access today | Yes | Rolling out, no date | BFL API and select partners |
Why Single-Pass Duration Is The Real Story
Thirty seconds sounds like an incremental number until you consider how people previously got there: generating three ten-second clips and joining them. Every join is a place where the character's face shifts, the light temperature jumps, or the wardrobe changes. A single pass eliminates those seams, which is what makes a continuous camera move or an unbroken one-take shot possible at all.
This is why Alibaba described Wan 3.0's duration jump as its clearest change, and why the same capability in Seedance 2.5 matters more than the raw seconds suggest. It is not "longer clips" so much as "shots that do not need to be assembled."
FLUX 3 Video: The Strongest Audio Story
FLUX 3 Video is the video half of FLUX 3, the multimodal model Black Forest Labs announced in July with a single set of weights trained across images, video, audio, and robot actions. Video generation became generally available on August 4.
- Up to 20 seconds, output at 720p with 1080p available via upscaling.
- Native audio including spoken dialogue with lip-sync across 12+ languages — the most ambitious audio claim of the three.
- Keyframes. You can set a start image, an end frame, or multiple keyframes and have the model connect them in sequence.
- Video continuation and multi-shot sequences, plus a Draft Mode that renders cheap previews before you commit to a full-quality render.
BFL's own evaluation reports human raters preferring it for both text-to-video and image-to-video. Access is through the BFL API and selected partner platforms; pricing was not disclosed at launch.
Wan 3.0: The Most Unusual Input Surface
Wan 3.0's distinguishing feature is not duration but what it will accept as a reference. Its Omni-Reference capability parses PDFs, plain text, Microsoft Office files, Apple Keynote and Pages documents, and webpages, and uses them as source material for a video.
That makes document-to-video a genuine new capability rather than a spec bump — turning a deck or a spec sheet directly into a clip is something none of the others do. Wan 3.0 also folds generation, editing, visual replication, and character driving into one model, where Wan 2.7 split those across separate ones, and adds an intelligent duration feature that suggests an optimal length from your prompt.
The caveats are real, though. It is a public beta, reachable through Alibaba's own surfaces — Model Studio, Qwen Cloud, the Wan site — with full API access described as coming "soon" and no date attached. There are no published weights. The honest position on Wan 3.0 today is interest, not migration.
Seedance 2.5: The One You Can Actually Call
Seedance 2.5 matches Wan 3.0's 30-second ceiling and, unlike the other two, is available through a normal API and a normal interface right now.
- Up to 30 seconds at 480p or 720p, across 5s, 10s, 15s, and 30s.
- Native audio generated in the same pass, included at the same credit cost — there is no audio surcharge.
- Up to 9 reference images for subject consistency, plus start and end frame control.
- Transparent per-render pricing: 80 credits for 480p/5s up to 1,055 for 720p/30s.
Its clear weakness is resolution. Seedance 2.5 stops at 720p while FLUX 3 Video reaches 1080p. If you need full HD today, Kling 3.0 delivers 1080p at 5s and 10s and can be paired with Seedance 2.5 for the long shots.
The Honest Scoreboard
| If you need… | Use | Why |
|---|---|---|
| A 30s take you can render today | Seedance 2.5 | Only 30s model with open access and published prices |
| Spoken dialogue with lip-sync | FLUX 3 Video | 12+ language lip-synced audio, if you have API access |
| Video from a document or deck | Wan 3.0 | Only model that parses PDFs and Office files — once the API opens |
| 1080p output right now | Kling 3.0 | Seedance 2.5 caps at 720p |
| Character consistency across shots | Seedance 2.5 | 9 reference images, available today |
What This Means Practically
Announcement quality and availability have drifted apart. FLUX 3 Video has the most impressive audio and no public price. Wan 3.0 has the most novel inputs and an API that has not opened. If your work is scheduled rather than speculative, the model you can call matters more than the model that demoed best — and that argues for building on what is available now while keeping the pipeline loose enough to swap in the others when they open up.
That lesson has a sharper edge this quarter: the Sora 2 API shuts down on September 24, 2026 with no replacement model, which is a reminder of what single-vendor dependency costs.
Try The 30-Second Format
Seedance 2.5 is live in Text to Video, Image to Video, and Reference to Video. Start with a 480p 5-second test at 80 credits to tune the prompt, then commit to the long render. The Seedance 2.5 guide covers prompt structure for 30-second shots.