> This article is written based on publicly verifiable information. Sources include Alibaba Cloud official announcements, arXiv technical papers, PixelDojo early access pages, and third-party platform test data. As of early August 2026, the latest officially available version from Alibaba is Wan 2.7. Information about "WAN 3.0" in this article primarily comes from early platform testing leaks and is cited with clear sources.
I. Introduction: AI Video Generation Enters the Era of "30-Second Complete Narratives"

On July 31, 2026, AI tool platform PixelDojo posted a message on its official X account that sent shockwaves through the industry: they had obtained early access to the WAN 3.0 video model and demonstrated multiple single-generation complete 30-second video demos.[^1]
Among these demos were a complete scene of two chefs exchanging four lines of dialogue in a kitchen, a 30-second single-take tracking shot of a letter traveling from a mailbox to a doorstep, and adaptive aspect ratio demonstrations that automatically fitted vertical ads and square product loops. Most striking of all—every video's dialogue, sound effects, ambient audio, and music were generated synchronously with the visuals, not layered on in post-production.
What does this mean? AI video generation is evolving from "generating a few seconds of visual footage" to "single-output complete short films with sound." For content creators, advertising professionals, and technical developers, this is a qualitative leap.
However, an important context must be clarified first: As of August 3, 2026, Alibaba Cloud's official platforms (wan.video and Alibaba Cloud Model Studio) still have Wan 2.7 as the latest publicly available version, supporting up to approximately 15 seconds of 1080P video generation.[^2] The name "WAN 3.0" has not yet appeared in any official Alibaba documentation or API.[^3] The 3.0 capabilities cited in this article primarily come from early testing demonstrations on platforms like PixelDojo; its official name, specifications, and release date remain to be confirmed by Alibaba.
II. The Full History of the WAN Video Model Series

To understand the significance of WAN 3.0, one must first understand the evolution of the entire Wan series. It is one of the most important open-source video model families in China, and its development trajectory clearly demonstrates a strategy of "open-source first to build foundations → rapid iteration to enhance capabilities → move toward production-grade."
2.1 Origins: From Image Generation to Video (2023-2024)
Tongyi Wanxiang was originally the visual generation brand of Alibaba's Tongyi Lab. When it launched in July 2023, it focused primarily on AI image generation, integrated into office scenarios like DingTalk. At the Apsara Conference in September 2024, Tongyi Wanxiang added video generation capabilities for the first time, supporting scene-based applications like illustration, doodling, and short clips—marking the formal transition from static images to dynamic video.
2.2 The Breakthrough Moment: Wan 2.1 Fully Open-Source (February 2025)
On February 26, 2025, Alibaba Cloud announced the open-sourcing of four models in the Wan 2.1 series, including T2V-14B, T2V-1.3B, I2V-14B-720P, and I2V-14B-480P.[^4] This move created huge waves in the industry, with many practitioners calling it "the DeepSeek moment for video generation."
Wan 2.1's core technology is based on the Diffusion Transformer (DiT) architecture, combined with Alibaba's proprietary Spatio-Temporal Variational Autoencoder (Wan-VAE) and flow matching training framework.[^5] Its technical paper, "Wan: Open and Advanced Large-Scale Video Generative Models" (arXiv:2503.20314), detailed the architecture: Wan-VAE achieves a 4x temporal compression ratio and 8x8 spatial compression ratio, efficiently encoding video into latent space; the text encoder uses the Qwen LLM in place of traditional CLIP or T5, supporting complex prompts of up to 512 tokens in both Chinese and English.[^5]
Key breakthroughs of Wan 2.1:
- Dual-version strategy: The 14B version targets professional production environments, while the 1.3B version requires only 8.19 GB of VRAM to run on consumer-grade GPUs (such as RTX 4090)[^5]
- Benchmark supremacy: Ranked first overall in VBench evaluations with a score of 86.22%, surpassing multiple closed-source models at the time[^4]
- Visual text generation: The first model to support bilingual Chinese-English text rendering within generated videos[^4]
- Fully open-source: Apache 2.0 license, with source code, model weights, and training code all freely available[^4]
By April 2025, the Wan 2.1 series had accumulated over 2.2 million downloads on Hugging Face and ModelScope.[^6] The ecosystem around ComfyUI, LoRA fine-tuning, and other community tools rapidly flourished, significantly lowering the barrier to high-quality video generation.
2.3 Rapid Iteration: From 2.2 to 2.7 (Mid-2025 to Early 2026)
The Wan series iterates faster than virtually any other large model in China. Here are the key version milestones:
Wan 2.2 (July 29, 2025) — The industry's first open-source MoE video model[^7]
Alibaba Cloud released the Wan 2.2 series, with the biggest highlight being the introduction of the Mixture-of-Experts (MoE) architecture. The series includes three models: Wan2.2-T2V-A14B (text-to-video), Wan2.2-I2V-A14B (image-to-video), and Wan2.2-TI2V-5B (hybrid text+image-to-video).
The MoE architecture is elegantly designed: two expert models handle the high-noise and low-noise stages respectively—the high-noise expert is responsible for overall scene layout and motion planning, while the low-noise expert focuses on detail texture and consistency refinement. Although the total parameter count reaches 27 billion, only 14 billion parameters are activated per inference step, reducing computational consumption by approximately 50%.[^7]
In terms of data, Wan 2.2's training dataset was substantially expanded compared to 2.1—image data increased by 65.6% and video data by 83.2%.[^7] Motion performance improved significantly in challenging scenarios like complex facial expressions, dynamic hand gestures, and sports movements.
Wan 2.5 (approximately September 2025) — First native audio-visual synchronization
Cross-verified from multiple sources, Wan 2.5 was the first version in the series to introduce native audio generation capabilities, supporting synchronous generation of human voice, sound effects, and background music, with duration extending to the 10-second range at 1080P. This capability was further refined in version 2.7.
Wan 2.6 (approximately December 2025) — Enhanced Reference-to-Video (R2V)
The Wan 2.6 series further strengthened reference-to-video capabilities, supporting characters with voice performing in new scenes and multi-shot narratives. It performed excellently on evaluation platforms such as Artificial Analysis Video Arena.
HappyHorse 1.0 (April 2026) — Anonymous Arena-topping performance
On April 7, 2026, an anonymous model named "HappyHorse 1.0" appeared on Video Arena, achieving a text-to-video Elo score of 1384 and an image-to-video Elo of 1416, both setting historical Arena records.[^8] The model was subsequently confirmed to belong to the Wan series (Wan 2.7), demonstrating Alibaba's technical prowess in video generation.
Wan 2.7 (March-April 2026 to present) — Current highest officially available version
Wan 2.7 is the latest version currently available on Alibaba Cloud Model Studio and the wan.video official platform.[^2] It unifies image and video generation, supporting text-to-video, image-to-video (first frame / first-and-last frames / video continuation), reference-to-video, and video editing across multiple modes. Key capabilities include:
- Up to approximately 15 seconds of 1080P video generation
- Native audio synchronization (dialogue + sound effects + ambient audio)
- Multi-shot narratives (natural language shot descriptions supported)
- First-and-last frame control and video continuation
- Multilingual visual text generation
The Alibaba Cloud DashScope API currently provides three endpoints: wan2.7-t2v-2026-06-12 (text-to-video), wan2.7-i2v-2026-04-25 (image-to-video), and wan2.7-videoedit (video editing).[^2][^9]
2.4 Core Evolution Logic
Looking across the entire Wan series, three clear technical trajectories emerge:
| Dimension | 2.1 (Feb 2025) | 2.2 (Jul 2025) | 2.7 (Apr 2026) | 3.0 (Jul 2026, Testing) |
|---|---|---|---|---|
| Duration | 5 seconds | 5 seconds | ~15 seconds | 2-30 seconds |
| Architecture | Dense DiT | MoE DiT | Unified multi-task | TBD |
| Audio | None | None | Native sync | Native sync |
| Reference Input | Limited | Limited | Image + first/last frame | Image + video + audio hybrid |
| Resolution | 720P | 720P | 1080P | 1080P (demonstrated) |
Strategically, Alibaba adopted a hybrid approach of "aggressively open-source early to build the ecosystem, then use APIs and closed-source testing to push toward production viability." Wan 2.1/2.2's complete open-sourcing (Apache 2.0) brought enormous community ecosystem dividends, while some capabilities from 2.5 onward began to be provided through closed-source APIs. This strategy bears similarities to Meta's LLaMA series—using openness to gain ecosystem, using closed-source to maintain commercial competitiveness.
III. WAN 3.0 Core Capabilities in Detail (Based on Early Testing)
The following information primarily comes from PixelDojo's early access preview pages and public demo demonstrations between July 31 and August 2, 2026.[^1] PixelDojo explicitly states that the model is "still in testing and not in the tool list yet." Specifications claimed by third-party marketing sites (wan30.co, wan3pro.com, etc.), such as native 4K, 60fps, 60 seconds, and Neural Physics Engine, do not match official documentation or actual test platforms. This article uses information from reliable test platforms like PixelDojo as the baseline.
3.1 Duration Breakthrough: Single-Generation 30-Second Complete Videos

This is WAN 3.0's biggest breakthrough. PixelDojo demonstrated the ability to generate complete videos of 2-30 seconds in a single pass, without stitching or post-processing.[^1] Even more noteworthy is the "intelligent duration" feature—when duration is set to -1, the model automatically decides the video length based on the prompt. In PixelDojo's test, describing "a paper boat washing down a gutter" resulted in the model autonomously generating a 20-second video, with the length entirely determined by the model.
This means video generation is shifting from "user precisely controls every frame" to "user describes intent, model orchestrates autonomously."
3.2 Native Audio Integration

WAN 3.0's audio is not layered on in post-production but generated synchronously during the diffusion process.[^1] Dialogue, music, room tone, and foley are all directly embedded in the output MP4 file. PixelDojo specifically notes that "turning sound off costs exactly what leaving it on costs," indicating that audio generation is architecturally native, not an optional post-processing step.
In the demos, the 30-second chef dialogue scene showcased the synchronous generation of four natural lines of dialogue, with sound matching lip movements; the rain-on-canvas scene featured thunder, rain, and distant fire effects precisely aligned with the visuals.
3.3 Multi-Modal Reference System

WAN 3.0 supports inputting up to approximately 10 images, 5 video clips, and 5 audio clips as references in a single request.[^1] Users can extract a character's appearance from one image, a scene environment from another, voice characteristics from an audio file, and describe in text how they should combine.
PixelDojo's demo showcased this capability: a single still photo of a person was input, and the output was a dynamic video of that person in a specific environment, with the character's appearance maintained consistently.
3.4 Pixel-Level Consistency
Across multiple demos, WAN 3.0 demonstrated character and scene consistency throughout the entire video.[^1] In a 30-second rescue scene, two characters' clothing, facial features, and spatial relationships remained stable throughout, with no noticeable drift or deformation. This is critical for narrative content—viewers won't be pulled out of the story by a character suddenly "changing faces."
3.5 Adaptive Aspect Ratio

WAN 3.0 supports setting the aspect ratio to an "adaptive" mode where the model automatically selects the most appropriate frame shape based on the prompt content.[^1] In PixelDojo's tests, prompting "vertical social ad" returned 720x1280, and prompting "square product loop" returned 960x960—neither request specified any dimensions.
3.6 First-and-Last Frame Control
Users can specify the first and last frames of a video (both as still images), and the model automatically generates the motion transition between them.[^1] This feature is highly practical for scenarios requiring precise control over start and end states, such as product showcases or action transitions.
IV. Head-to-Head Comparison: WAN 3.0 vs. Leading Video Models

The AI video generation landscape in mid-2026 is fiercely competitive, with each leading player having distinct strengths. Below is a comparison based on verifiable information.
4.1 Comparison Subjects
- Seedance 2.5 (ByteDance): API officially launched in July 2026, currently the most feature-aggressive commercial model[^10]
- Kling 3.0 (Kuaishou): Officially released in February 2026, the most mature commercial closed-source model[^11]
- Gemini Omni Flash (Google DeepMind): Released in May 2026, with unique conversational editing capabilities[^8]
4.2 Core Specifications Comparison
| Dimension | WAN 3.0 (Early Testing) | Seedance 2.5 | Kling 3.0 | Advantage |
|---|---|---|---|---|
| Max single-pass duration | 30 seconds | 30 seconds | 15 seconds | WAN/Seedance tied |
| Resolution | 1080P (demonstrated) | Native 4K (10-bit color) | Native 4K / 1080P / 60fps | Seedance/Kling |
| Native audio | Yes (dialogue+SFX+ambient) | Yes (joint latent space generation) | Yes (multilingual lip-sync+dialects) | All three |
| Multi-modal references | ~10 images+5 videos+5 audio | Up to 50 reference inputs | Multi-image+element reference | Seedance clearly leads |
| Multi-shot narratives | Supports coherent scenes | Multi-shot+region editing | Up to 6 shots (AI director) | Kling most refined |
| Character consistency | Pixel-level (strong in testing) | Excellent (abundant references) | Excellent (Element Reference) | All three comparable |
| Availability | Early testing, official TBD | API open, expanding | Fully open (website+API) | Kling most mature |
| Open source | Series traditionally open | Closed-source | Closed-source | WAN potential advantage |
4.3 Detailed Analysis
Duration and narrative completeness: Both WAN 3.0 and Seedance 2.5 achieve single-pass 30-second generation, which is hugely significant for advertising and short drama scenarios—a complete narrative arc can be achieved in a single generation. Kling 3.0 is capped at 15 seconds but maximizes information density through up to 6 intelligent shots, making it suitable for faster-paced social media content.[^10][^11]
Reference control: Seedance 2.5 is far ahead in this regard—up to 50 multi-modal references per generation (images + video + audio + 3D blockouts), plus support for localized region editing (bounding-box inpainting) and 3D camera pre-visualization.[^10] WAN 3.0's reference capabilities are already strong in early testing but still fall short of Seedance's 50-reference capacity. Kling 3.0's advantage lies in the refinement of character element binding and multi-shot customization.
Audio and multilingual support: All three support native audio-visual synchronization. Kling 3.0 has the most comprehensive multilingual support (Chinese, English, Japanese, Korean, Spanish + dialects/accents, multi-character mixed speech).[^11] Seedance 2.5 emphasizes the higher synchronization precision achieved through joint latent-space generation. WAN 3.0's audio naturalness performs well in demos, but multilingual details have not yet been disclosed.
Visual quality: Seedance 2.5 and Kling 3.0's native 4K capabilities have been widely verified. WAN 3.0's current demonstrations remain at 1080P, with physics and detail performance (such as honey pouring, rain canvas, interface text rendering) looking good, but no large-scale 4K evidence has been seen yet.
Cost-efficiency: Kling 3.0 has the lowest API pricing (standard mode at approximately $0.075/second), making it the top choice for high-frequency daily use.[^8] Seedance 2.5's 720P standard tier costs approximately $0.30/second, and the 4K tier approximately $0.68/second—significantly more expensive. WAN 3.0's official pricing has not yet been announced.
4.4 Use Case Recommendations
- Longest complete narrative (30-second single take) → WAN 3.0 or Seedance 2.5
- Maximum reference materials + fine-grained control (professional production) → Seedance 2.5
- Mature and stable, multilingual dialogue, refined shot control, immediate output → Kling 3.0
- Budget-sensitive / high cost-efficiency for daily use → Kling 3.0 or future WAN official release
- Local deployment / self-hosting / model fine-tuning → WAN series (open-source DNA)
V. Technical Speculation and Analysis
Based on the Wan series' publicly available technical papers[^5] and the architectural evolution of existing versions, combined with early testing performance, we can make some reasonable speculations about WAN 3.0's technical direction.
5.1 The Architectural Challenge of 30-Second Generation
Going from 15 to 30 seconds is not simply "doubling"—the core challenge of long-form video generation is the exponential decay of temporal coherence. Early DiT models tend to exhibit object drift, style discontinuities, and spatial relationship breakdowns beyond 10 seconds.
WAN 3.0 likely achieved breakthroughs in one or more of the following areas:
- Spatio-temporal attention optimization: May have introduced multi-scale temporal attention, processing at different temporal resolutions in parallel to balance long-range consistency and short-range motion detail
- VAE compression ratio improvement: The Wan-VAE's temporal compression ratio may have increased from 4x to higher, reducing computational load for long sequences
- Flow matching mechanism improvements: May have drawn from the "world model + event stream" decomposition approach described in the Wan-Streamer paper, splitting long videos into semantic event sequences
5.2 Audio-Visual Joint Generation
WAN 3.0's native audio is not added in post-production but produced synchronously during generation.[^1] It likely employs a joint audio-video diffusion architecture that simultaneously models visual and audio signals in the same latent space. This is similar to Seedance 2.5's technical approach, though the specific implementation details may differ.
The advantage of this architecture lies in: higher audio-visual synchronization precision (no post-alignment needed), and the ability to learn deeper "sound-action" associations (such as the natural matching between footsteps and walking rhythm).
5.3 Multi-Reference and Identity Locking
The leap from Wan 2.7's limited reference inputs to WAN 3.0's "10 images + 5 videos + 5 audio" hybrid references represents a significant capability jump. This likely introduces stronger cross-attention mechanisms or dedicated identity locking modules that fuse features from multiple reference sources into the generation process at the cross-attention layers.
5.4 On the Credibility of the "Neural Physics Engine"
Some third-party sites claim WAN 3.0 features a "Neural Physics Engine" supporting precise physical simulation of liquids, cloth, and hair.[^12] From PixelDojo's demos, physics realism has indeed improved (honey pouring, paper boat drifting, etc.), but this is more likely a side effect of improved training data quality and temporal modeling rather than an independent physics engine module. In current diffusion model architectures, "true physics engine integration" remains an unsolved research problem.
VI. Application Scenarios and Significance for Creators


6.1 A Paradigm Shift in Content Creation
What does 30 seconds + native audio mean? It means AI video generation has transformed from a "footage supplier" to a "short film producer." Previously, creators needed to generate 5-10 second clips with AI, then manually stitch, dub, and edit. Now, a complete advertising scene, a dialogue-driven short drama, or a product showcase can go from prompt to finished product in a single step.
Specific application scenarios:
- Short-form / social media content: 30 seconds perfectly covers the single-video duration for TikTok, Instagram Reels, and similar platforms
- Ad creative: Complete 15-30 second commercials can be produced as first drafts through a single generation
- Short dramas / narrative content: Complete story arcs with beginning, development, and conclusion
- Product demos: Product close-ups with sound effects, 360-degree showcases
- Education / training: Narrated operation demonstrations, concept explanations
6.2 Multi-Model Workflows Become the Norm
In actual creative workflows, a single model rarely meets all needs. Many creators are already mixing multiple models:
- Using Seedance 2.5 for high-reference-control brand visual content (requiring abundant reference materials + precise editing)
- Using Kling 3.0 for multilingual dialogue scenes and social media clips (requiring mature multi-shot + lip-sync)
- Waiting for the WAN 3.0 official release to fill the "30-second long narrative" gap (requiring single-pass complete generation + potential low-cost / open-source solution)
This "toolbox" mindset will be the norm for future AI content creation—no single model does it all, but multiple complementary models can cover virtually every scenario.
6.3 Deeper Industry Impact
If WAN 3.0 maintains the Wan series' open-source tradition upon official release, it will have profound impacts on the industry:
- Lowering barriers to entry: Small and medium creators and independent developers can access near-commercial-grade video generation capabilities
- Promoting localized deployment: For scenarios with privacy requirements or cost sensitivity, self-hosted solutions are essential
- Accelerating ecosystem innovation: ComfyUI workflows, LoRA style fine-tuning, customized inference pipelines, and other community innovations will continue to flourish
VII. Future Outlook and Risks
7.1 Release Timeline Speculation
Based on the Wan series' historical pace (2.1 to 2.7 iterated through 6 major versions in approximately one year) and the fact that multiple platforms have already received early access, the speculated official release timeline for WAN 3.0 may be:
- August-September 2026: Alibaba Cloud Model Studio API opens 3.0 version (closed-source API)
- Q3-Q4 2026: Lightweight versions (e.g., 1.3B) open-sourced
- Q4 2026 or later: Full version open-sourced (if the tradition continues)
However, it should be noted that Wan 2.5/2.6 experienced situations where "open-source promises were not fully delivered," so the community should maintain cautious optimism about 3.0's open-source strategy.
7.2 Computational Cost Challenges
Generating 30-second videos will cost significantly more than 15-second ones. Based on Kling 3.0 and Seedance 2.5 pricing, a 30-second 1080P video API call could cost in the $2-10 range. For high-frequency creators, this is a factor that needs consideration.
7.3 Competitive Landscape Outlook
The AI video race in the second half of 2026 will intensify further:
- ByteDance (Seedance): Most feature-aggressive, but closed-source and expensive
- Kuaishou (Kling): Most mature and stable, best cost-efficiency, continuously iterating
- Alibaba (WAN): Open-source DNA + long-duration advantage, needs to catch up on image quality and 4K
- Google (Gemini Omni): Unique conversational editing paradigm, broadest ecosystem
- OpenAI (Sora): Technologically cutting-edge but commercially conservative
WAN series' greatest differentiated competitive advantage lies in its open-source DNA + Chinese-language scenario optimization + sustained rapid iteration. If 3.0 continues this tradition, it will become the most cost-effective choice for developers and small-to-medium creators.
7.4 Risks to Watch
- Marketing exaggeration: Specifications claimed by third-party sites (wan30.co, wan3pro.com, etc.), such as "native 4K, 60fps, 60 seconds, Neural Physics Engine," do not match official documentation or actual test platforms and may be exaggerated.[^3] It is recommended to use actually testable platforms like PixelDojo and Alibaba's official documentation as the baseline
- Copyright and ethics: Issues around the boundaries of reference material usage, commercial copyright attribution of generated content, and other matters still require industry consensus
- Technical limitations: Complex high-speed motion, extreme physical interactions, quality degradation in ultra-long videos (beyond 30 seconds), and other issues remain to be verified
VIII. Conclusion
> Prefer learning by doing? Try text-to-video and image-to-video in the browser studio at wan3video.art (credit-based; live generation caps follow the connected production API).
WAN 3.0's early testing information indicates that it is a pragmatism-driven major upgrade—focused on solving the problem of "generating complete, usable short films" (duration + audio + consistency), rather than simply stacking resolution or parameter counts. From PixelDojo's actual demos, the quality is already commercially viable.
However, we must also remain rational: information about WAN 3.0 is currently highly fragmented, ranging from reliable early testing data to exaggerated third-site marketing claims. Until Alibaba makes an official announcement, it is recommended to judge based on actually testable demo results.
For developers and technical teams, the WAN series' open-source tradition means lower trial-and-error costs and greater customization space. For content creators, 30-second + native audio capabilities will significantly shorten the distance from concept to finished product. Regardless of which group you belong to, staying tuned to the PixelDojo preview page, Alibaba Cloud Model Studio official documentation, and Wan-AI's GitHub/Hugging Face pages is the best way to get the latest updates.
In this era of multi-model workflows, there is no single "best model"—only the one best suited to your specific scenario. WAN 3.0 is filling in the crucial "long narrative" puzzle piece, and the entire industry will accelerate its evolution toward "single-generation complete narrated storytelling" as a result.
References
[^1]: PixelDojo, "WAN 3.0 makes 30 second videos," July 31, 2026. https://pixeldojo.ai/wan-3-0-preview
[^2]: Alibaba Cloud Model Studio, "Wan2.7 - text-to-video API reference," updated July 22, 2026. https://help.aliyun.com/en/model-studio/text-to-video-api-reference
[^3]: Imagera.ai, "Is Wan 3.0 Real? What Alibaba Has Actually Shipped in 2026," July 29, 2026. https://imagera.ai/blog/is-wan-3-0-real-what-alibaba-shipped-2026
[^4]: Alibaba Group, "Alibaba Cloud Open Sources its AI Models for Video Generation," February 26, 2025. https://www.alibabagroup.com/en-US/document-1831486012178563072
[^5]: Wan Team, Alibaba Group, "Wan: Open and Advanced Large-Scale Video Generative Models," arXiv:2503.20314, March 2025. https://arxiv.org/abs/2503.20314
[^6]: Alibaba Group, "Alibaba Unveils its Latest Open-Source Video Generation Model," April 18, 2025. https://www.alibabagroup.com/en-US/document-1851424828087599104
[^7]: Alibaba Cloud, "Alibaba Releases Wan2.2 to Uplift Cinematic Video Production," July 29, 2025. https://www.alibabacloud.com/en/press-room/alibaba-releases-wan-2-2-to-uplift-cinematic
[^8]: CSDN, "HappyHorse, Seedance, Kling, Gemini Omni: Deep Cross-Review of 2026 Video Large Models," July 30, 2026. https://blog.csdn.net/aidoudoulong/article/details/163328673
[^9]: Alibaba Cloud Model Studio, "Wan 2.7 - image-to-video API," updated July 22, 2026. https://help.aliyun.com/en/model-studio/image-to-video-general-api-reference
[^10]: OrcaRouter, "Seedance 2.5 vs Kling 3.0: Best AI Video Model in 2026," June 30, 2026. https://www.orcarouter.ai/blog/seedance-2-5-vs-kling-3-0
[^11]: Dreamina (Seedance Official), "Seedance 2.5 Vs Kling 3.0: Which AI Video Model Fits Modern Content Production?" June 29, 2026. https://dreamina.capcut.com/seedance/seedance-2-5-vs-kling-3-0
[^12]: wan30.co, "Wan 3.0 Review," 2026. https://wan30.co/review (Note: Third-party marketing site, not Alibaba official; information credibility should be treated with caution)