The following tables list model releases. For model deprecation rules and lists, see Model decommissioning policy.
China (Beijing)
|
Model type |
Date |
Model ID |
Description |
|
Speech recognition |
2026-07-30 |
|
Added the Qwen-Audio-3.0-ASR-Flash-Streaming (real-time), Qwen-Audio-3.0-ASR-Flash-Filetrans (non-real-time), and Qwen-Audio-3.0-ASR-Flash (non-real-time) models: Dialect support: Supports the seven major Chinese dialect groups (Mandarin, Wu, Xiang, Gan, Hakka, Min, and Yue) and more than 20 regional accents; Classical poetry optimization: Improves recognition accuracy for classical Chinese poetry, making it suitable for education, culture, and audiobook scenarios; Text optimization: Enhances punctuation prediction and text normalization, automatically converting numbers, dates, and monetary amounts to standard formats; Multilingual expansion: Supports 30 languages, including Chinese, English, Japanese, and Korean; Hotwords and context: Supports hotwords (precompiled and on-the-fly) and context input to improve recognition accuracy for domain-specific terms. |
|
Text generation, Reasoning, Visual understanding |
2026-07-21 |
|
The Qwen3.7 native vision-language series Flash models comprehensively enhance multimodal understanding and Agent execution capabilities compared to 3.6-Flash. Key improvements include strengthened foundational multimodal abilities, enhanced object recognition, improved real-world perception and spatial intelligence. Multimodal Agent scenarios such as Search Agent and CI Agent have seen significant upgrades, with more stable end-to-end task execution. Multimodal coding capabilities are optimized, delivering a smoother vibe coding experience. |
|
Image generation |
2026-07-20 |
|
Rich content: Supports input of up to 4.5k tokens and dense information layout with images-within-images, enabling complex layouts like newspapers, storyboards, menus, and exam papers to be generated in a single pass. Authentic detail: Supports precise rendering of text as small as 10px, and vividly reproduces fine details such as micro-expressions, pores, and individual strands of hair—approaching the quality of real photography. Deep knowledge: Supports native rendering of 12 languages and 20+ fonts, realistic simulation of mainstream interfaces such as web pages, games, and live streams, fully incorporating external knowledge. Qwen-Image-3.0-Pro isn't just pursuing "good looks"—it's pursuing "usefulness", making image generation a truly deployable productivity tool. |
|
Text generation, Reasoning, Visual understanding |
2026-07-17 |
|
Kimi K3 is Kimi's most powerful flagship model with 2.8 trillion parameters. Built on KDA hybrid linear attention (Kimi Delta Attention) and Attention Residuals technologies, it natively supports vision understanding and features a 1 million token context window. It is the world's first open-source 3-trillion-parameter model designed for long-context programming, knowledge work, and advanced reasoning. |
|
Text embedding |
2026-07-15 |
|
Qwen3-7-Text-Embedding is a multilingual text vector model based on Qwen3.7, offering significant improvements in text retrieval, clustering, and classification over text-embedding-v4. It achieves 20% better performance in MTEB multilingual, Chinese-English, and code retrieval tasks and supports customizable vector dimensions (256-2560). |
|
Realtime speech synthesis |
2026-07-14 |
|
Qwen-Audio-3.0-TTS-Plus is a high-performance speech synthesis model, designed for high-quality speech generation scenarios. Compared with the previous version, it supports more low-resource languages and Chinese dialects, significantly improves dialect authenticity, and enhances free-style instruction following and fine-grained tag control for more accurate control over emotion, tone, character, speaking rate, volume, and synthesis style. It is also more robust under noisy and reverberant acoustic conditions, with further improvements in audio quality, clarity, resolution, and overall expressiveness. The Plus version focuses more on synthesis quality and detailed expressiveness, making it suitable for professional scenarios with higher requirements for audio quality, naturalness, and expressiveness, such as content creation, audiobooks, film and video dubbing, brand voice design, and premium speech services. |
|
Realtime speech synthesis |
2026-07-14 |
|
qwen-audio-3.0-tts-flash is a high-performance speech synthesis model, optimized for real-time interactive scenarios. Compared with the previous version, it supports more low-resource languages and Chinese dialects, improves dialect authenticity, and enhances free-style instruction following and fine-grained tag control for more flexible control over emotion, tone, character, speaking rate, volume, and expressive style. It is also more robust under noisy and reverberant acoustic conditions, with improved audio quality, clarity, and overall expressiveness. The Flash version focuses on real-time synthesis, making it suitable for voice assistants, real-time dialogue, intelligent customer service, and other low-latency interactive applications. |
|
Realtime chat |
2026-07-14 |
|
Qwen 3.0 Realtime Speech Model Standard Edition - next-generation duplex speech model ranked #1 globally in Artificial Analysis Speech-to-Speech benchmark. Balances model intelligence with duplex dialogue rhythm for natural interaction and enhanced response quality. |
|
Realtime chat |
2026-07-14 |
|
Qwen 3.0 Realtime Speech Dialogue Model Flash Edition - a next-generation duplex speech model ranked #1 globally in Artificial Analysis Speech-to-Speech benchmark. Combines high intelligence with optimized duplex rhythm for low-latency (parallel inference, full-streaming optimization) 'fast and smart' dialogue experience focusing on response speed. |
|
Video generation |
2026-07-14 |
|
Upscale enhances videos of any resolution to 4K, improving clarity and detail for high-quality presentation and redistribution. |
|
Video generation |
2026-07-14 |
|
Motion Control extracts actions from reference videos and transfers them to target character images, generating new videos of characters performing the same actions. |
|
Video generation |
2026-07-14 |
|
Lip-sync capability aligns characters' mouth movements in videos with input audio/TTS, enhancing natural speech and narrative expressiveness. |
|
Video generation |
2026-07-09 |
|
Accepts images and text prompts for video generation. ViduQ3-Pro-fast offers faster generation speed and higher cost-effectiveness. Compared to ViduQ2-Pro-fast, it extends generation duration from 10s to 16s, enabling more complex shot transitions and narrative logic. |
|
Image generation |
2026-07-09 |
|
Accepts 0-14 reference images or text prompts for reference-to-image generation, text-to-image generation, and image editing. Prioritizes high speed, quality, and low cost, with costs approximately 50% lower than Pro. |
|
Video generation |
2026-07-09 |
|
ViduQ3-Drama is a premium drama/AI manga production model, upgraded across 'consistency, motion effects, detail aesthetics, and cost-effectiveness'. Delivers stable visuals, authentic emotions, and enhanced dynamics. |
|
Video generation |
2026-07-09 |
|
ViduQ3-Ad is an advertising-specific model featuring 'marketing-grade scene transitions, intelligent camera movement, and direct audio output'. Upload a product image to generate a 16-second ad video, lowering production barriers and costs. |
|
Image generation |
2026-07-09 |
|
Generates images from 0-14 references or text, excelling in complex logic with strong context consistency and industrial-grade stability. Suitable for professional design and web drama production. |
|
Image generation |
2026-07-09 |
|
Generates images from 0-14 references or text, with enhanced semantic understanding and support for diverse styles. |
|
Image generation |
2026-07-09 |
|
Generates images from 0-14 reference images or text, supporting reference-based generation, text-to-image, and editing. Ensures pixel-level accuracy for Chinese/English text and UI/chart details, ideal for posters and infographics. |
|
Text generation |
2026-07-09 |
|
GLM-5.2-Fast-Preview is the high-speed variant of Zhipu AI's GLM-5.2, with 1M context and capabilities on par with the standard version. Inference-optimized to deliver 1.5–2× the output TPS, it fits latency-sensitive use cases such as real-time chat, multi-turn agents, and streaming code generation. |
|
Video generation |
2026-07-01 |
|
Wan2.7 text to video, performance fully reimagined. Delivering nuanced and organic emotional depth in narrative arcs and visceral, bone-crunching impact in action sequences. Enhanced by rhythmic cinematic cuts for unparalleled storytelling power.This version is a snapshot as of June 12, 2026. |
|
Video generation |
2026-07-01 |
|
Wan2.7 reference to video, enhanced consistency & performance. Delivering superior stability for characters, props, and scenes. Supports hybrid referencing of up to 5 mixed image/video inputs and audio timbre cloning. Together with core engine upgrades, it achieves unprecedented cinematic expressive power. This version is a snapshot as of June 12, 2026. |
|
Image generation |
2026-06-25 |
|
The Qwen-Image-2.0 series full-fledged model integrates image generation and editing; it boasts more professional text rendering capabilities with 1k token command support, more delicate and realistic textures, meticulous depiction of realistic scenes, and stronger semantic adherence. The full-fledged version possesses the strongest text rendering capabilities and realistic textures in the 2.0 series. |
|
Text generation, Reasoning, Visual understanding |
2026-06-17 |
|
K2.7 Code High-Speed Edition shares the same model as the standard version but delivers 5-6x faster output speed, achieving ~180 tokens/s in median programming scenarios and ~260 tokens/s in short-context scenarios, enhancing programming efficiency. |
|
Speech recognition |
2026-06-17 |
|
The Bailing ASR version, updated in June 2026, fully supports the seven major dialect systems of traditional Chinese (Mandarin/Wu/Xiang/Gan/Hakka/Min/Yue) and is compatible with over 20 regional Mandarin accents. It features specific optimizations for the rhythm, meter, and literary expression characteristics of classical Chinese poetry, improving the accuracy of poetry recognition and making it suitable for cultural heritage, educational explanations, and audiobooks. Improved punctuation prediction and text normalization capabilities make the output text more consistent with written expression habits; numbers, dates, and amounts are automatically converted to standard formats, enhancing readability and professionalism. The language support has been expanded to include English, Japanese, Korean, Vietnamese, Thai, Indonesian, Malay, Filipino, Hindi, Arabic, French, German, Spanish, Portuguese, Russian, Italian, Dutch, Swedish, Danish, Finnish, Norwegian, Greek, Polish, Czech, Hungarian, Romanian, Bulgarian, Croatian, and Slovak, totaling 30 languages. It supports contextualization capabilities and can transcribe audio up to 5 minutes in length. |
|
Visual understanding |
2026-06-16 |
|
Qwen3.5-OCR is an upgraded OCR model with enhanced document parsing, text localization, and key information extraction. It significantly improves extraction performance for real-world documents (e.g., ID cards, driver's licenses). |
|
Video generation |
2026-06-16 |
|
HappyHorse-1.1-T2V supports Text-to-Video generation with improved semantic understanding, cinematic shot control, and dynamic motion rendering. It more accurately captures creative intent, producing high-quality videos with smoother motion, richer details, stronger visual consistency, and more natural character actions, scene atmosphere, and physical dynamics. |
|
Video generation |
2026-06-16 |
|
HappyHorse-1.1-R2V supports Reference-to-Video generation with significantly improved stability in subject, scene style, and visual consistency. With support for up to 9 reference images, it can more accurately understand and preserve creative intent, delivering stronger controllability and expressiveness across characters, scenes, styles, and cinematic motion. |
|
Video generation |
2026-06-16 |
|
HappyHorse-1.1-I2V supports Image-to-Video generation with improved visual quality, dynamic performance, and cross-clip consistency. It more accurately understands the input image and preserves creative intent, delivering significant improvements in skin texture realism, character ID consistency across clips, motion smoothness, text rendering stability, and audio-visual synchronization, producing high-quality videos with greater realism, richer details, and stronger overall consistency. |
|
Text generation, Reasoning |
2026-06-16 |
|
GLM-5.2 is the latest flagship model from Zhipu AI, designed for long-horizon tasks with support for an ultra-long 1M context window. It features powerful logical reasoning, long-text comprehension, and code generation capabilities, balancing performance with inference efficiency. It excels across multi-task benchmarks and is well-suited for intelligent interaction, enterprise applications, and development assistance scenarios. |
|
Text generation, Reasoning, Visual understanding |
2026-06-15 |
|
kimi-k2.7-code is Kimi's most intelligent coding model to date. It follows instructions more reliably over long contexts and completes programming tasks with higher success rates. It supports text, image, and video inputs, along with thinking mode, conversation, and agent tasks. |
|
Text generation, Reasoning, Visual understanding |
2026-06-09 |
|
The Max model, the largest and most capable in the Qwen3.7 series, has added visual‑modal understanding compared to the May 20 snapshot, enabling it to perceive real‑world scenes and supporting multimodal interactive hybrid agent capabilities. This version is based on a snapshot taken on June 8, 2026. |
|
Video generation |
2026-06-08 |
|
V6 is a new model launched by PixVerse in late March 2026. The r2v (multi-subject reference video generation) model ranks second globally, accepting 2-7 input images to intelligently fuse different subjects. It excels in complex mid-shot and long-shot video scenarios while retaining t2v's prompt control capabilities and it2v's consistency preservation. It offers enhanced emotional expression and smoother high-speed motion visuals, supports 15-second long videos, direct music-and-video output, and multiple languages. It enables one-click generation for scenarios like e-commerce product close-ups, advertisement trailers, and simulating C4D modeling to showcase product structures. |
|
Text generation, Reasoning, Visual understanding |
2026-06-01 |
|
Among the Qwen3.7 series, the cost-effective Plus model builds on its robust text capabilities while delivering a comprehensive upgrade to its vision‑language abilities, all while preserving its full‑stack agent‑level intelligence for coding, tool use, and productivity workflows. Its key distinguishing feature is multi‑modal interactive hybrid agent capabilities, enabling it to perceive real‑world scenes, read screens and interact with GUIs, generate code based on visual references, and perform end‑to‑end navigation within mobile apps. |
|
Speech synthesis |
2026-06-01 |
|
The Fun Music large model supports song creation based on open-ended lyrics or composition requirements, generating full Chinese/English songs with male/female vocals. The songs are accessible and emotionally progressive, representing a perfect fusion of human inspiration and LLM capabilities. |
|
Text generation, Reasoning, Visual understanding |
2026-06-01 |
|
MiniMax M3 excels in enterprise-grade long-document understanding, high-quality content generation, code writing, bug fixing, and native application building, powered by industry-leading coding and agentic capabilities, a 1M ultra-long context window, and native multimodal features. Its agentic capabilities enable end-to-end workflow integration, while native multimodal support delivers seamless text-image interaction. |
|
Text generation |
2026-05-29 |
|
A flagship MoE large model with 1.6 trillion parameters and 49 billion activated parameters, natively supporting context lengths of up to one million tokens. Trained on a vast corpus of high-quality data, it excels in advanced mathematical reasoning, complex logical inference, specialized coding, and deep analysis of long-form text, making it well-suited for demanding applications such as cutting-edge research, sophisticated office workflows, and advanced AI agents. |
|
Text generation, Visual understanding |
2026-05-25 |
|
Step 3.7 Flash is a production-grade Agent-efficient Flash model newly launched by Stepfun for building high-performance Agents. Optimized for speed, cost efficiency, execution reliability, and complex task completion, it features multimodal perception, visual search, tool enhancement, high-reliability tool orchestration, and Agent ecosystem compatibility. |
|
Text generation, Reasoning |
2026-05-21 |
|
The Max model, the largest and most capable in the Qwen3.7 series, currently offers a pure‑text‑only interface for public experimentation. Qwen3.7 is a next‑generation flagship model designed for the agent‑centric era, with its core strengths lying in the breadth and depth of its agent‑level capabilities: it excels at programming, office and productivity tasks, and long‑term autonomous execution. |
|
Text generation, Reasoning |
2026-05-20 |
|
The Max model, the largest and most capable variant in the Qwen3.7 series, is available as a preview and supports only thinking mode, offering a pure text‑only interface for experimentation. It is primarily optimized for general‑purpose conversational use cases, such as knowledge‑based question answering, instruction following, and creative writing. |
|
Realtime speech translation |
2026-05-19 |
|
The real-time version of Qwen3.5-LiveTranslate-Flash, which is a high-precision, highly responsive, and robust multilingual simultaneous audio and video interpretation model. Leveraging Qwen3.5-Omni's powerful infrastructure, massive multimodal data, cross-language and cross-modal alignment, and visual enhancement technologies, Qwen3.5-LiveTranslate-Flash offers both offline and real-time audio and video translation capabilities. It can understand 60 languages and speak 29 languages. |
|
Text generation |
2026-05-18 |
|
MiMo-V2.5-Pro is Xiaomi's latest flagship model. Compared to predecessors, it significantly improves general agent capabilities, complex software engineering, and long-horizon tasks, leading benchmarks like ClawEval, GDPVal, and SWE-bench Pro. It autonomously completes expert-level tasks requiring thousands of tool calls over days/weeks, with a 1 million token context length ideal for agent frameworks. |
|
Text generation |
2026-05-18 |
|
GLM-5.1 is a model developed by Zhipu AI, specifically designed for long-horizon tasks. It has 744 billion parameters, supports an ultra-long context of 200k tokens, and can generate up to 128k tokens in a single response. GLM-5.1 excels in logical reasoning, long-text understanding, and code generation, while balancing performance with inference efficiency. It delivers outstanding results across multiple multi-task benchmarks and is well-suited for applications such as intelligent human-computer interaction, enterprise solutions, and developer assistance. |
|
Speech synthesis |
2026-05-06 |
|
A music generation model creating complete Chinese/English songs (male/female vocals) from lyrics or creative requirements, blending human inspiration with AI capabilities. |
|
Video generation |
2026-04-28 |
|
Accepts 1-7 reference images and text prompts for video generation. ViduQ3-mix delivers strong visual quality and balanced performance. Designed for episodic content creation, it integrates 6 visual effects (particle/fluid/dynamics/camera movement/transition/lighting), 5 audio effects (environment/dynamic/atmosphere/sound design/emotional), and 4 scenarios (short films/animated series/TV dramas/advertisements), enabling versatile creation of short films, animations, and ads. |
|
Video generation |
2026-04-27 |
|
Accepts 1-7 reference images and text prompts for video generation. ViduQ3 supports intelligent scene transitions with enhanced multi-camera consistency. Designed for episodic content creation, it integrates 6 visual effects (particle/fluid/dynamics/camera movement/transition/lighting), 5 audio effects (environment/dynamic/atmosphere/sound design/emotional), and 4 scenarios (short films/animated series/TV dramas/advertisements), enabling versatile creation of short films, animations, and ads. |
|
Video generation |
2026-04-27 |
|
Accepts 1-7 reference images and text prompts for video generation. ViduQ3-Turbo offers fast generation speed and high cost-effectiveness. Designed for episodic content creation, it integrates 6 visual effects (particle/fluid/dynamics/camera movement/transition/lighting), 5 audio effects (environment/dynamic/atmosphere/sound design/emotional), and 4 scenarios (short films/animated series/TV dramas/advertisements), enabling versatile creation of short films, animations, and ads. |
|
Video generation |
2026-04-27 |
|
Generates videos from images and text. ViduQ2-Pro-fast offers cost-effective, stable performance with 2-3x faster speed than turbo. As the world's first 'Everything as Reference' video model, it supports reference in six dimensions (effects, expressions, textures, actions, characters, scenes) for precise editing, tailored for web dramas, short films, and production. |
|
3D generation |
2026-04-27 |
|
Tripo P1.0 is a real-time 3D generation model for developers and creators needing clean topology and engine-ready meshes. It generates professional-grade 3D assets in ~2 seconds, optimized for gaming, Web3D, and interactive scenarios. It prioritizes speed and out-of-the-box usability for UGC pipelines, accelerating integration into real-time engines. |
|
3D generation |
2026-04-27 |
|
Tripo H3.1 is a high-precision 3D generation model designed for creators requiring extreme visual quality and detail. With 20B-level parameters, it supports billion-voxel resolution and up to 2M-polygon generation. It enhances geometric fidelity and texture alignment for complex structures like character forms, facial details, and text, suitable for high-quality visual production and 3D printing. |
|
Video generation |
2026-04-26 |
|
Wan2.7 text to video, performance fully reimagined. Delivering nuanced and organic emotional depth in narrative arcs and visceral, bone-crunching impact in action sequences. Enhanced by rhythmic cinematic cuts for unparalleled storytelling power.This version is a snapshot as of April 25, 2026. |
|
Video generation |
2026-04-26 |
|
Wan2.7 image to video, performance fully reimagined. Delivering nuanced and organic emotional depth in narrative arcs and visceral, bone-crunching impact in action sequences. Enhanced by rhythmic cinematic cuts for unparalleled storytelling power.This version is a snapshot as of April 25, 2026. |
|
Text generation, Reasoning, Visual understanding |
2026-04-26 |
|
Kimi-k2.6 is the latest intelligent model in Kimi series, featuring enhanced long-context code generation, improved instruction following and self-correction capabilities, supporting text/image/video inputs and multiple operational modes. |
|
Video generation |
2026-04-26 |
|
HappyHorse-1.0-V2V supports advanced video editing through natural language instructions. It allows for local or global editing of video elements using up to 5 reference images, precisely preserving original motion dynamics to achieve superior expressiveness. |
|
Video generation |
2026-04-26 |
|
HappyHorse-1.0-R2V supports Reference-to-Video generation, offering enhanced stability in subject and scene referencing. Capable of processing up to 9 reference images, it precisely preserves creative intent to deliver superior performance. |
|
Text generation, Reasoning |
2026-04-24 |
|
A flagship MoE large model with 1.6 trillion parameters and 49 billion activated parameters, natively supporting context lengths of up to one million tokens. Trained on a vast corpus of high-quality data, it excels in advanced mathematical reasoning, complex logical inference, specialized coding, and deep analysis of long-form text, making it well-suited for demanding applications such as cutting-edge research, sophisticated office workflows, and advanced AI agents. |
|
Text generation, Reasoning |
2026-04-24 |
|
A highly efficient, lightweight MoE model with 284 billion parameters in total and 13 billion activated parameters, natively supporting context windows of up to one million tokens. It offers fast inference speed, low latency, and cost-effective invocation, delivering well-balanced overall performance. Designed for high-concurrency, lightweight workloads, it is ideally suited for common, essential use cases such as everyday dialogue, content creation, basic RAG applications, and batch text processing. |
|
Text generation, Reasoning, Visual understanding |
2026-04-23 |
|
The Qwen3.5 native vision-language series Plus model has seen a substantial improvement in agentic coding capabilities compared to the February 15th snapshot. Inference speed has also been significantly enhanced, while its knowledge retention, reasoning ability, and long-context processing remain at a high level, making it well-suited for complex agent-based tasks. It is ideal for applications such as coding agents, production workflows, and high-throughput scenarios. This version is based on a snapshot taken on April 20, 2026. |
|
Image generation |
2026-04-23 |
|
The full-featured Qwen-Image-2.0 series models integrate image generation and image editing, offering enhanced text rendering with support for 1,000-token prompts, more refined realistic textures, detailed depiction of photorealistic scenes, and stronger semantic adherence. The full-featured version delivers the strongest text rendering and most lifelike textures in the 2.0 series. |
|
Text generation, Reasoning, Visual understanding |
2026-04-22 |
|
The Qwen3.6 27B native vision-language dense model builds upon the 3.5-27B version, with key improvements in agentic coding capabilities and enhanced STEM reasoning and inference skills. In the vision modality, it demonstrates significant advances in spatial intelligence, object localization, and detection, while video understanding, document OCR, and visual agent capabilities continue to improve steadily. |
|
Video generation |
2026-04-22 |
|
HappyHorse-1.0-I2V enables image-to-video generation, featuring highly realistic dynamic rendering. It accurately comprehends both text and image semantics to produce high-quality videos that are fluid, natural, and rich in detail. |
|
Text generation, Reasoning, Visual understanding |
2026-04-21 |
|
Kimi-k2.6 is the latest intelligent model in Kimi series, featuring enhanced long-context code generation, improved instruction following and self-correction capabilities, supporting text/image/video inputs and multiple operational modes. |
|
Video generation |
2026-04-21 |
|
HappyHorse-1.0-T2V supports text-to-video generation, featuring highly realistic dynamic rendering. It accurately comprehends text semantics to produce high-quality videos that are fluid, natural, and rich in detail. |
|
Text generation, Reasoning |
2026-04-20 |
|
The Max model, the largest and most capable variant in the Qwen3.6 series, is now available in a preview version. At present, only its plain-text capabilities are open for experimentation. Compared with the previously released Qwen3-Max and Qwen3.6-Plus, this model features enhanced vibe coding abilities, more efficient coding agent execution, and significantly improved front-end development skills. Additionally, its long-tail knowledge retention has been further upgraded. |
|
Text generation, Reasoning, Visual understanding |
2026-04-16 |
|
The Qwen3.6 native vision-language Flash model series delivers a significant performance boost over the 3.5-Flash version. This model particularly excels in agentic coding capabilities, substantially outperforming its predecessor on multiple code-agent benchmarks, as well as in mathematical and code reasoning. In terms of vision, it features markedly improved spatial intelligence, with especially notable enhancements in object localization and object detection. |
|
Text generation, Reasoning, Visual understanding |
2026-04-16 |
|
The Qwen3.6 35B-A3B native vision-language model is built on a hybrid architecture that integrates linear attention mechanisms with a sparse mixture-of-experts framework, achieving higher inference efficiency. Compared with the 3.5-35B-A3B, this model demonstrates significantly improved agentic coding capabilities, mathematical and code reasoning abilities, spatial intelligence, as well as object localization and object detection performance. |
|
Text generation, Reasoning |
2026-04-14 |
|
GLM-5.1 is a model developed by Zhipu AI, specifically designed for long-horizon tasks. It has 744 billion parameters, supports an ultra-long context of 200k tokens, and can generate up to 128k tokens in a single response. GLM-5.1 excels in logical reasoning, long-text understanding, and code generation, while balancing performance with inference efficiency. It delivers outstanding results across multiple multi-task benchmarks and is well-suited for applications such as intelligent human-computer interaction, enterprise solutions, and developer assistance. |
|
Video generation |
2026-04-09 |
|
C1, PixVerse's film-industry large model launched in March 2026, uses t2v (text-to-video) to precisely control visuals via prompts, accurately reproducing cinematic techniques (push-ins, pull-outs, pans, tracking shots). Supports 15s videos, direct music/video output, and multilingual text. |
|
Video generation |
2026-04-09 |
|
C1, PixVerse's film-industry large model launched in March 2026, employs r2v (multi-reference-to-video) to input 2-7 images, intelligently fusing multiple subjects. Combines t2v prompt control with it2v consistency and cinematic combat/action effects. Supports 15s videos, direct music/video output, and multilingual text. Ideal for multi-character scenes, dialogues, and interactions with medium/long shots. If a multi-panel storyboard (up to 3x3 grid) is provided, generates continuous storyboard videos with one click. |
|
Video generation |
2026-04-09 |
|
C1, PixVerse's film-industry large model launched in March 2026, uses kf2v (keyframe-to-video) to seamlessly connect any two images with natural transitions. Supports 15s videos, direct music/video output, and multilingual text. |
|
Video generation |
2026-04-09 |
|
C1, PixVerse's film-industry large model launched in March 2026, extends t2v prompt control with it2v (image-to-video) capabilities to precisely replicate reference images' colors, saturation, scenes, and character features. Compared to V6, it enhances prompt expressiveness, imagination, cinematic combat actions, and visual effects. Supports 15s videos, direct music/video output, and multilingual text. Suitable for close-ups, monologues, stop-motion/slow-motion, and empty-shot transitions. |
|
Text generation, Reasoning |
2026-04-07 |
|
DeepSeek-V3.2 harmonizes high computational efficiency with superior reasoning and agent performance. Built on DeepSeek-V3, it incorporates key innovations like DeepSeek Sparse Attention (DSA), scalable reinforcement learning frameworks, and large-scale agent task synthesis pipelines, advancing the frontier of open-source LLMs. |
|
Text generation, Reasoning |
2026-04-07 |
|
DeepSeek-V3.1-Terminus is an updated version of DeepSeek-V3.1, maintaining the model's core capabilities while addressing user feedback through fixes and optimizations. It retains the same architecture as DeepSeek-V3 and achieves significant enhancements in specific domains. |
|
Text generation |
2026-04-07 |
|
A self-developed Mixture-of-Experts (MoE) model with 671B parameters (activating 37B), pre-trained on 14.8T tokens. Demonstrates excellent capabilities in long-text processing, coding, mathematics, encyclopedic knowledge, and Chinese language tasks. |
|
Text generation, Reasoning |
2026-04-07 |
|
A large language model enhanced through extensive reinforcement learning during post-training, achieving strong reasoning capabilities with minimal labeled data. Performs well in mathematics, coding, and natural language reasoning tasks. |
|
Text generation, Visual understanding |
2026-04-07 |
|
DeepSeek-OCR, centered on 'exploring vision-text compression boundaries', redefines the functional positioning of vision encoders from a large language model (LLM) perspective, providing a new solution balancing accuracy and efficiency for document recognition, image-to-text conversion, and other high-frequency scenarios. |
|
Video generation |
2026-04-03 |
|
Wan2.7 video edit, supports both localized and global editing with prompt. Seamlessly replace elements using image references and replicate complex dynamic processes, including motion, special effects, and camera movements. |
|
Video generation |
2026-04-03 |
|
Wan2.7 text to video, performance fully reimagined. Delivering nuanced and organic emotional depth in narrative arcs and visceral, bone-crunching impact in action sequences. Enhanced by rhythmic cinematic cuts for unparalleled storytelling power. |
|
Video generation |
2026-04-03 |
|
Wan2.7 reference to video, enhanced consistency & performance. Delivering superior stability for characters, props, and scenes. Supports hybrid referencing of up to 5 mixed image/video inputs and audio timbre cloning. Together with core engine upgrades, it achieves unprecedented cinematic expressive power. |
|
Video generation |
2026-04-03 |
|
Wan2.7 image to video, performance fully reimagined. Delivering nuanced and organic emotional depth in narrative arcs and visceral, bone-crunching impact in action sequences. Enhanced by rhythmic cinematic cuts for unparalleled storytelling power. |
|
Image generation |
2026-04-01 |
|
Wan2.7–image-pro, supports text to image, text/image to sequential images, image editing, multi-image reference generation, and interactive editing. Delivers enhanced performance in text rendering, subject consistency, and complex instruction following. |
|
Image generation |
2026-04-01 |
|
Wan2.7 – image generation and editing, supports text to image, text/image to sequential images, image editing, multi-image reference generation, and interactive editing. Delivers enhanced performance in text rendering, subject consistency, and complex instruction following. |
|
Text generation, Reasoning, Visual understanding |
2026-04-01 |
|
The Qwen3.6 native vision-language Plus series models demonstrate exceptional performance on par with the current state-of-the-art models, with a significant improvement in overall results compared to the 3.5 series. The models have been markedly enhanced in code-related capabilities such as agentic coding, front-end programming, and Vibe coding, as well as in multi-modal general object recognition, OCR, and object localization. |
|
Video generation |
2026-04-01 |
|
V6 is a new model launched by PixVerse in late March 2026. The t2v (text-to-video) model enables precise video control through prompts, accurately reproducing various cinematic techniques including push-in, pull-out, panning, tracking, and follow shots with natural transitions and precise perspective control. It supports 15-second long videos, direct music-and-video output, and multiple languages. |
|
Video generation |
2026-04-01 |
|
V6 is a new model launched by PixVerse in late March 2026. The kf2v (keyframe-to-video) model seamlessly connects any two images, achieving smooth and natural video transitions. It supports 15-second long videos, direct music-and-video output, and multiple languages. |
|
Video generation |
2026-04-01 |
|
V6 is a new model launched by PixVerse in late March 2026. The it2v (image-to-video) model ranks second globally, offering not only the prompt control capabilities of t2v (text-to-video) but also high-fidelity reproduction of reference images' colors, saturation, scenes, and character features. It provides enhanced emotional expression and high-speed motion performance, supports 15-second long videos, direct music-and-video output, and multiple languages. It enables one-click generation for scenarios like e-commerce product close-ups, advertisement trailers, and simulating C4D modeling to showcase product structures. |
|
Realtime omni-modal |
2026-03-30 |
|
Qwen 3.5-Omni is the latest generation of Qwen's multimodal large model, supporting text, image, audio, and audio-visual understanding and interaction. As a fully evolved version of Qwen3-Omni, it supports audio input in 60+ languages, voice output in 30+ languages, and controllable voice dialogue, WebSearch and complex FunctionCall invocation, and has intelligent semantic interruption interaction capabilities. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interactive experience. |
|
Omni-modal |
2026-03-30 |
|
Qwen 3.5-Omni is the latest generation of Qwen's multimodal large model, supporting text, image, audio, and audio-visual understanding and interaction. As a comprehensive evolution of Qwen 3-Omni, it supports over 10 hours of audio understanding and over 400 seconds of 720P (1 FPS) audio-visual understanding and dialogue. It further expands the language range, supporting audio input in 60+ languages and speech output in 30+ languages. It also possesses powerful structured audio-visual understanding capabilities and is widely used in text creation, voice assistants, multimedia analysis, and other scenarios, providing a natural and fluent multimodal understanding and interactive experience. |
|
Realtime omni-modal |
2026-03-30 |
|
Qwen 3.5-Omni is the latest generation of Qwen's multimodal large model, supporting text, image, audio, and audio-visual understanding and interaction. As a fully evolved version of Qwen3-Omni, it supports audio input in 60+ languages, voice output in 30+ languages, and controllable voice dialogue, WebSearch and complex FunctionCall invocation, and has intelligent semantic interruption interaction capabilities. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interactive experience. |
|
Omni-modal |
2026-03-30 |
|
Qwen 3.5-Omni is the latest generation of Qwen's multimodal large model, supporting text, image, audio, and audio-visual understanding and interaction. As a comprehensive evolution of Qwen 3-Omni, it supports over 10 hours of audio understanding and over 400 seconds of 720P (1 FPS) audio-visual understanding and dialogue. It further expands the language range, supporting audio input in 60+ languages and speech output in 30+ languages. It also possesses powerful structured audio-visual understanding capabilities and is widely used in text creation, voice assistants, multimedia analysis, and other scenarios, providing a natural and fluent multimodal understanding and interactive experience. |
|
Video generation |
2026-03-26 |
|
Accepts text prompts for video generation. ViduQ3-Turbo is a high-performance accelerated model with extreme generation efficiency, combining superior image quality and dynamic performance. Excels in combat scenes, emotional rendering, and semantic understanding. Outstanding cost-effectiveness for image-based social platforms, AI companionship, and VFX material creation. |
|
Video generation |
2026-03-26 |
|
Accepts start/end frames and text prompts for video generation. ViduQ3-Turbo is a high-performance accelerated model with extreme generation efficiency, combining superior image quality and dynamic performance. Excels in combat scenes, emotional rendering, and semantic understanding. Outstanding cost-effectiveness for image-based social platforms, AI companionship, and VFX material creation. |
|
Video generation |
2026-03-26 |
|
Accepts images and text prompts for video generation. ViduQ3-Turbo is a high-performance accelerated model with extreme generation efficiency, combining superior image quality and dynamic performance. Excels in combat scenes, emotional rendering, and semantic understanding. Outstanding cost-effectiveness for image-based social platforms, AI companionship, and VFX material creation. |
|
Video generation |
2026-03-26 |
|
Accepts text prompts for video generation. ViduQ3-Pro is a flagship audiovisual-native model supporting up to 16s of synchronized audio-visual generation. It enables free multi-shot transitions with precise control over pacing, emotions, and narrative continuity. Leading parameter scale ensures cinematic-grade quality with superior image fidelity, character consistency, and emotional expression, suitable for professional production in advertising (e-commerce, TVC, performance ads), animated series, live-action dramas, and gaming. |
|
Video generation |
2026-03-26 |
|
Accepts start/end frames and text prompts for video generation. ViduQ3-Pro is a flagship audiovisual-native model supporting up to 16s of synchronized audio-visual generation. It enables free multi-shot transitions with precise control over pacing, emotions, and narrative continuity. Leading parameter scale ensures cinematic-grade quality with superior image fidelity, character consistency, and emotional expression, suitable for professional production in advertising (e-commerce, TVC, performance ads), animated series, live-action dramas, and gaming. |
|
Video generation |
2026-03-26 |
|
Accepts images and text prompts for video generation. ViduQ3-Pro is a flagship audiovisual-native model supporting up to 16s of synchronized audio-visual generation. It enables free multi-shot transitions with precise control over pacing, emotions, and narrative continuity. Leading parameter scale ensures cinematic-grade quality with superior image fidelity, character consistency, and emotional expression, suitable for professional production in advertising (e-commerce, TVC, performance ads), animated series, live-action dramas, and gaming. |
|
Video generation |
2026-03-26 |
|
Generates videos from text. ViduQ2-text2video ensures precise instruction following and emotional nuance capture, with strong narrative control, micro-expression rendering, and dynamic cinematography. Widely used in film, advertising, web dramas, and tourism. |
|
Video generation |
2026-03-26 |
|
Generates videos from reference images and text. ViduQ2-reference2video ensures precise instruction following and emotional nuance capture, with strong narrative control, micro-expression rendering, and dynamic cinematography. Widely used in film, advertising, web dramas, and tourism. |
|
Video generation |
2026-03-26 |
|
Generates videos from start/end frames and text. ViduQ2-Turbo is an ultra-fast engine: 720P 5s video in 19s, 1080P in 27s. Delivers natural human actions/expressions and high-quality motion effects for dynamic scenes like combat. |
|
Video generation |
2026-03-26 |
|
Generates videos from images and text. ViduQ2-Turbo is an ultra-fast engine: 720P 5s video in 19s, 1080P in 27s. Delivers natural human actions/expressions and high-quality motion effects for dynamic scenes like combat. |
|
Video generation |
2026-03-26 |
|
Generates videos from start/end frames and text. ViduQ2-Pro is the world's first 'Everything as Reference' video model, supporting six-dimensional reference for precise editing. Designed for web dramas, short films, and production. |
|
Video generation |
2026-03-26 |
|
Generates videos from reference videos, images, and text. ViduQ2-Pro-reference2video is the world's first 'Everything as Reference' video model, supporting six-dimensional reference for advanced editing. Enables controlled modifications for precise video creation, tailored for web dramas, short films, and production. |
|
Video generation |
2026-03-26 |
|
Generates videos from images and text. ViduQ2-Pro is the world's first 'Everything as Reference' video model, supporting six-dimensional reference for comprehensive editing. Enables controlled modifications for precise video creation, optimized for web dramas, short films, and production. |
|
Video generation |
2026-03-26 |
|
Intelligent scene segmentation interprets script transitions, automatically adjusting camera angles and shot types. Native multimodal framework ensures audio-visual coherence. Breaks duration limits for flexible multi-shot storytelling. |
|
Video generation |
2026-03-26 |
|
Introduces 'All-in-One Reference' supporting 3-8s videos or multi-image anchoring for character elements. Synchronizes voice dubbing and lip movements for authentic character portrayal. Enhances video consistency and dynamic expression with audio-visual synchronization and intelligent scene segmentation. |
|
Image generation |
2026-03-26 |
|
Unlocks cinematic narrative visuals with new series generation and 2K/4K direct output. Deeply analyzes audio-visual elements in prompts, accurately executing creative instructions. Supports multi-reference images and comprehensive effect upgrades, ideal for storyboarding, plot concept art, and scene design. |
|
Image generation |
2026-03-26 |
|
Supports up to 10 reference images, locking subjects, elements, and tones for style consistency. Integrates style transfer, portrait/character reference, multi-image fusion, and localized redrawing with flexible operations. Delivers realistic portrait details, refined visuals, and cinematic color atmospheres. |
|
Text generation, Reasoning |
2026-03-26 |
|
GLM-5 targets coding and agent scenarios, achieving open-source SOTA in complex systems engineering with capabilities approaching Claude Opus, built on a 744B foundation with asynchronous reinforcement learning and sparse attention. |
|
Text generation |
2026-03-23 |
|
Qwen-Deep-Research is an advanced agent system for complex research tasks, equipped with multi-step reasoning and global planning capabilities. It leverages tools like internet search to perform detailed task decomposition, reasoning, and analysis, ultimately generating traceable, logically rigorous research reports. |
|
Multimodal embedding |
2026-03-20 |
|
Embedding-Vision is a vision-centric multimodal embedding model powered by an LLM, featuring outstanding domain-specific performance and high cost-effectiveness in various domains (e.g., e-commerce, photo galleries, security, autonomous driving). With support for text, image, and video, it is applicable to downstream retrieval tasks, including text-to-image, image-to-image, text-to-video and video-to-video. |
|
Multimodal embedding |
2026-03-20 |
|
Embedding-Vision is a vision-centric multimodal embedding model powered by an LLM, featuring outstanding domain-specific performance and high cost-effectiveness in various domains (e.g., e-commerce, photo galleries, security, autonomous driving). With support for text, image, and video, it is applicable to downstream retrieval tasks, including text-to-image, image-to-image, text-to-video and video-to-video. |
|
Text generation, Reasoning |
2026-03-20 |
|
M2.7 autonomously constructs complex agent frameworks, leveraging capabilities like Agent Teams, complex Skills, and Tool Search to execute highly sophisticated productivity tasks. |
|
Video generation |
2026-03-19 |
|
Input text descriptions to generate high-quality videos with precise semantic matching in seconds, supporting multiple styles. PixVerse V5.6 is a proprietary video generation large model developed by A-Soul Technology, featuring comprehensive upgrades in text-to-video and image-to-video capabilities. The model significantly improves image clarity, complex motion stability, and audio-visual synchronization, achieving more accurate lip-syncing and natural emotional expression in multi-character dialogue scenarios. It also optimizes composition, lighting, and texture consistency for further enhanced generation quality. PixVerse V5.6 ranks in the global top tier of the Artificial Analysis text-to-video and image-to-video leaderboards. |
|
Video generation |
2026-03-19 |
|
Input 2-7 images to intelligently fuse different subjects while maintaining style consistency and motion coordination, enabling easy construction of rich narrative scenes with enhanced content controllability and creative flexibility. PixVerse V5.6 is a proprietary video generation large model developed by A-Soul Technology, featuring comprehensive upgrades in text-to-video and image-to-video capabilities. The model significantly improves image clarity, complex motion stability, and audio-visual synchronization, achieving more accurate lip-syncing and natural emotional expression in multi-character dialogue scenarios. It also optimizes composition, lighting, and texture consistency for further enhanced generation quality. PixVerse V5.6 ranks in the global top tier of the Artificial Analysis text-to-video and image-to-video leaderboards. |
|
Video generation |
2026-03-19 |
|
Achieves seamless transitions between any two images for smooth, visually striking scene shifts. PixVerse V5.6, developed by Artifactory, upgrades text-to-video and image-to-video capabilities. The model significantly improves visual clarity, complex motion stability, and audio-visual coordination, with accurate lip-sync and natural emotion in multi-character dialogues. It optimizes composition, lighting, and texture consistency for enhanced overall quality. PixVerse V5.6 ranks in the global top tier of Artificial Analysis text-to-video and image-to-video benchmarks. |
|
Video generation |
2026-03-19 |
|
Upload any image and generate dynamic, coherent videos with customizable plots, pacing, and styles. PixVerse V5.6, developed by Artifactory, upgrades text-to-video and image-to-video capabilities. The model significantly improves visual clarity, complex motion stability, and audio-visual coordination, with accurate lip-sync and natural emotion in multi-character dialogues. It optimizes composition, lighting, and texture consistency for enhanced overall quality. PixVerse V5.6 ranks in the global top tier of Artificial Analysis text-to-video and image-to-video benchmarks. |
|
Speech synthesis |
2026-03-19 |
|
MiniMax speech large model intelligently predicts text emotions and intonation based on context to generate ultra-natural, high-fidelity, and personalized speech. Demonstrates strong capabilities in social interaction, podcasting, audiobooks, news, education, digital humans, and other scenarios. |
|
Speech synthesis |
2026-03-19 |
|
MiniMax speech large model intelligently predicts text emotions and intonation based on context to generate ultra-natural, high-fidelity, and personalized speech. Demonstrates strong capabilities in social interaction, podcasting, audiobooks, news, education, digital humans, and other scenarios. |
|
Speech synthesis |
2026-03-19 |
|
MiniMax speech large model intelligently predicts text emotions and intonation based on context to generate ultra-natural, high-fidelity, and personalized speech. Demonstrates strong capabilities in social interaction, podcasting, audiobooks, news, education, digital humans, and other scenarios. |
|
Speech synthesis |
2026-03-19 |
|
MiniMax speech large model intelligently predicts text emotions and intonation based on context to generate ultra-natural, high-fidelity, and personalized speech. Demonstrates strong capabilities in social interaction, podcasting, audiobooks, news, education, digital humans, and other scenarios. |
|
Reasoning, Visual understanding |
2026-03-18 |
|
GUI-series foundation model for cross-platform (mobile/desktop) interface understanding and interaction, supporting complex multi-step tasks and multi-agent collaboration. |
|
Speech recognition |
2026-03-05 |
|
This is the real-time version of Tongyi Lab's next-generation end-to-end speech recognition model, based on leading proprietary speech technology, and boasts exceptional contextual awareness and high-precision speech transcription capabilities. Based on an end-to-end architecture, Fun-ASR integrates innovative RAG technology, supporting multi-dimensional features such as large-scale hotword customization, automatic filtering of sensitive and modal particles, ITN normalization, and punctuation prediction, significantly improving overall recognition accuracy and contextual relevance. Furthermore, Fun-ASR supports flexible switching between Chinese and English, covers multiple regional dialects, and boasts enhanced noise robustness, adapting to diverse and complex environments. |
|
Speech recognition |
2026-03-03 |
|
Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves highly accurate speech recognition, automatically determining the language and accurately identifying speech in multiple languages, while ensuring precise transcription even in complex audio environments.This version is a snapshot dated February 10, 2026. |
|
Image generation |
2026-03-03 |
|
The full-featured Qwen-Image-2.0 series models integrate image generation and image editing, offering enhanced text rendering with support for 1,000-token prompts, more refined realistic textures, detailed depiction of photorealistic scenes, and stronger semantic adherence. The full-featured version delivers the strongest text rendering and most lifelike textures in the 2.0 series.This version is a snapshot as of March 3, 2026. |
|
Image generation |
2026-03-03 |
|
The Qwen-Image-2.0 series of accelerated models integrates image generation and image editing, offering enhanced text-rendering capabilities with support for 1,000-token prompts, more realistic textures, finely detailed photorealistic scenes, and improved semantic consistency. The accelerated version effectively strikes an optimal balance between model performance and quality. |
|
Text generation |
2026-02-28 |
|
The Qwen Role-Playing Model Series is specifically optimized for muti-language anthropomorphic interaction scenarios. It demonstrates advanced capabilities in character consistency maintenance, context-aware dialogue progression, and empathetic engagement, enabling precise personalized character embodiment. This version significantly enhances Japanese linguistic localization (including dialects and honorifics), human-like role-playing authenticity, narrative coherence control, and scenario-based cognitive intelligence. |
|
Speech synthesis |
2026-02-27 |
|
A high-expressiveness speech synthesis model in the CosyVoice series with enhanced voice cloning and design capabilities. Supports free-style instruction control while maintaining speaker similarity, offering rich synthesis styles. Reduces first-syllable latency, improves pronunciation accuracy, and enhances prosody and sound quality. Enables ultra-natural multilingual (Chinese, English, German, French, Russian, Japanese, Korean, Portuguese, Thai, Indonesian, Vietnamese) real-time speech synthesis. |
|
Speech synthesis |
2026-02-27 |
|
A high-performance speech synthesis model in the CosyVoice series with enhanced voice cloning and design capabilities. Supports free-style instruction control while maintaining speaker similarity, offering rich synthesis styles. Reduces first-syllable latency, improves pronunciation accuracy, and enhances prosody and sound quality. Enables ultra-natural multilingual (Chinese, English, German, French, Russian, Japanese, Korean, Portuguese, Thai, Indonesian, Vietnamese) real-time speech synthesis. |
|
Text generation, Reasoning |
2026-02-24 |
|
MiniMax-M2.5 is MiniMax's flagship open-source large model, trained on hundreds of thousands of real-world complex scenarios through large-scale reinforcement learning, achieving or surpassing industry SOTA in productivity scenarios like programming, tool calls, search, and office tasks. |
|
Text generation, Reasoning, Visual understanding |
2026-02-23 |
|
The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the 3 series, these models deliver a leap forward in performance for both pure text and multimodal tasks, offering fast response times while balancing inference speed and overall performance. |
|
Text generation, Reasoning, Visual understanding |
2026-02-23 |
|
The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall performance is comparable to that of the Qwen3.5-27B. |
|
Text generation, Reasoning, Visual understanding |
2026-02-23 |
|
The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of the Qwen3.5-122B-A10B. |
|
Text generation, Reasoning, Visual understanding |
2026-02-23 |
|
The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of overall performance, this model is second only to Qwen3.5-397B-A17B. Its text capabilities significantly outperform those of Qwen3-235B-2507, and its visual capabilities surpass those of Qwen3-VL-235B. |
|
Text generation |
2026-02-19 |
|
The new-generation code generation model in the Qwen3 series delivers performance close to that of Qwen3-Coder-Plus while offering even better capabilities. The model has been optimized with a focus on repository-level understanding, supports multi-turn tool interactions, and enhances its compatibility with agentic coding tools. |
|
Text generation, Reasoning |
2026-02-18 |
|
GLM-5 targets coding and agent scenarios, achieving open-source SOTA in complex systems engineering with capabilities approaching Claude Opus, built on a 744B foundation with asynchronous reinforcement learning and sparse attention. |
|
Text generation, Reasoning, Visual understanding |
2026-02-15 |
|
The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of task evaluations, the 3.5 series consistently demonstrates performance on par with state-of-the-art leading models. Compared to the 3 series, these models show a leap forward in both pure-text and multimodal capabilities. |
|
Text generation, Reasoning, Visual understanding |
2026-02-15 |
|
The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers state-of-the-art performance comparable to leading-edge models across a wide range of tasks, including language understanding, logical reasoning, code generation, agent-based tasks, image understanding, video understanding, and graphical user interface (GUI) interactions. With its robust code-generation and agent capabilities, the model exhibits strong generalization across diverse agent. |
|
Speech recognition |
2026-02-13 |
|
The real-time version of Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves highly accurate speech recognition, automatically determining the language and accurately identifying speech in multiple languages, while ensuring precise transcription even in complex audio environments.This version is a snapshot dated February 10, 2026. |
|
Text generation, Reasoning, Visual understanding |
2026-02-13 |
|
Kimi-k2.5 is Moonshot AI's most versatile model with native multimodal architecture, supporting visual/text inputs, thinking/non-thinking modes, and both dialogue and agent tasks. |
|
Text generation, Reasoning |
2026-02-13 |
|
MiniMax-M2.5 is MiniMax's flagship open-source large model, trained on hundreds of thousands of real-world complex scenarios through large-scale reinforcement learning, achieving or surpassing industry SOTA in productivity scenarios like programming, tool calls, search, and office tasks. |
|
Text generation, Reasoning |
2026-02-13 |
|
MiniMax-M2.1 is MiniMax's flagship open-source large model, focusing on real-world complex tasks with core strengths in multilingual programming and long-chain agent capabilities. |
|
Realtime speech recognition |
2026-02-12 |
|
A lightweight real-time Mandarin speech recognition model optimized for Chinese call-center scenarios, supporting multi-dialect accents and achieving low-latency, high-accuracy transcription in low-sample-rate/low-SNR environments. |
|
Speech synthesis |
2026-02-10 |
|
Qwen3-TTS-VD model is Tongyi's latest real-time speech synthesis model. It can perform high-fidelity real-time speech synthesis on the voices designed by the qwen3-voice-design service, and supports speech output in 11 languages using the same voice. This model has been trained on massive amounts of data, and the synthesized audio can adaptively adjust its tone according to the text, demonstrating good processing capabilities for complex text synthesis. This model is a snapshot version from January 26, 2026. |
|
Speech synthesis |
2026-02-10 |
|
Qwen3-TTS-Flash model is Tongyi's latest real-time speech synthesis model. It can perform high-fidelity real-time speech synthesis on voices replicated from the qwen-voice-enrollment service, and supports speech output in 11 languages using the same voice timbre. This model has been trained on massive amounts of data, and the synthesized audio can adaptively adjust tone according to the text, demonstrating good processing capabilities for complex text synthesis. This model is a snapshot version from January 22, 2026. |
|
Speech synthesis |
2026-02-10 |
|
Qwen3-TTS-Flash model is Tongyi's latest real-time speech synthesis model. The Instruct model processes the synthesis effect through natural language, ensuring highly appropriate emotional and expressive speech in different contexts. Currently, it supports 25 timbres for both Chinese and English Instruct adjustments. |
|
Text generation, Reasoning, Visual understanding |
2026-01-30 |
|
Kimi-k2.5 is Moonshot AI's most versatile model with native multimodal architecture, supporting visual/text inputs, thinking/non-thinking modes, and both dialogue and agent tasks. |
|
Video generation |
2026-01-29 |
|
Wan2.6 reference to video flash, faster and more cost-effective generation. Supports using a specified person or any object as a reference, precisely maintaining consistency of appearance and voice, and allows multi‑character reference for joint performances. |
|
Text embedding |
2026-01-29 |
|
Qwen3-VL-Rerank is a reranking model that deeply understands rich multimodal information (text, images, video). After initial retrieval, it applies advanced cross-modal correlation to intelligently re-rank candidates, prioritizing the most relevant results. It improves cross-modal search accuracy, optimizes image/video retrieval precision, enhances image clustering quality, and enables efficient complex multimodal retrieval and tagging. |
|
Text generation, Reasoning |
2026-01-28 |
|
DeepSeek-V3.2 is the official release of a model that incorporates DeepSeek Sparse Attention—a sparse attention mechanism. It’s also the first model launched by DeepSeek that integrates reasoning into tool usage, supporting both reasoning-enabled and non-reasoning tool calls. |
|
Text generation, Reasoning |
2026-01-28 |
|
DeepSeek-V3.1-Terminus is an updated version of DeepSeek-V3.1, maintaining the model's core capabilities while addressing user feedback through fixes and optimizations. It retains the same architecture as DeepSeek-V3 and achieves significant enhancements in specific domains. |
|
Text generation |
2026-01-28 |
|
A self-developed Mixture-of-Experts (MoE) model with 671B parameters (activating 37B), pre-trained on 14.8T tokens. Demonstrates excellent capabilities in long-text processing, coding, mathematics, encyclopedic knowledge, and Chinese language tasks. |
|
Text generation, Reasoning |
2026-01-28 |
|
A large language model enhanced through extensive reinforcement learning during post-training, achieving strong reasoning capabilities with minimal labeled data. Performs well in mathematics, coding, and natural language reasoning tasks. |
|
Reasoning, Visual understanding |
2026-01-26 |
|
The Qwen3 series VL models effectively integrates thinking and non-thinking modes, achieving world-leading performance in visual agent capabilities on public benchmark datasets such as OS World. This version features comprehensive upgrades in areas like visual coding, spatial perception, and multimodal reasoning, significantly enhancing visual perception and recognition abilities, and supporting the understanding of ultra-long videos. |
|
Text generation, Reasoning |
2026-01-23 |
|
Compared with the snapshot as of September 23, 2025, the Qwen-3 series Max model in this release achieves an effective integration of thinking and non-thinking modes, resulting in a comprehensive and substantial improvement in the model's overall performance. In thinking mode, the model simultaneously supports web search, web information extraction, and a code interpreter tool, enabling it to tackle more complex and challenging problems with greater accuracy by leveraging external tools while engaging in slow, deliberative reasoning. This version is based on a snapshot taken on January 23, 2026. |
|
Text generation, Reasoning |
2026-01-23 |
|
MiniMax-M2.1 is MiniMax's flagship open-source large model, focusing on real-world complex tasks with core strengths in multilingual programming and long-chain agent capabilities. |
|
Visual understanding |
2026-01-22 |
|
The Qwen3 series of small-sized visual understanding models effectively integrates thinking and non-thinking modes. Compared with the snapshot taken on October 15, 2025, the overall performance of the model has improved significantly: it delivers enhanced capabilities in general visual recognition and reasoning, and shows marked improvements in recognition accuracy across various business scenarios such as security, in-store inspections, equipment monitoring, and photo-based problem solving. This version is a snapshot as of January 22, 2026. |
|
Multimodal embedding |
2026-01-21 |
|
Qwen3-VL-Embedding is a unified multimodal vector model based on Qwen3-VL, supporting text, image, video (single or mixed modality) input. It outputs unified representation vectors, suitable for cross-modal retrieval, image/video search, clustering, complex multimodal retrieval, and tagging. |
|
Speech synthesis |
2026-01-21 |
|
qwen3-TTS-Flash model is Tongyi's latest real-time speech synthesis model. The Instruct model processes the synthesis effect through natural language, ensuring highly appropriate emotional and expressive speech in different contexts. Currently, it supports 25 timbres for both Chinese and English Instruct adjustments. This model is equivalent to the snapshot version released on January 22, 2026. |
|
Video generation |
2026-01-15 |
|
Wan2.6 image to video flash, faster and more cost-effective generation. Intelligent shot scheduling enables multi‑camera storytelling, supports stable multi‑speaker dialogue with more natural and realistic vocal timbres, and supports generation of clips up to 15 seconds in length. |
|
Image generation |
2026-01-15 |
|
The Max series Qwen's image editing models delivers more stable and versatile editing capabilities: enhanced industrial design and geometric reasoning, improved character consistency, reduced offset issues, and integrated LoRA capabilities for a wider range of image editing functions. |
|
Speech synthesis |
2026-01-14 |
|
Qwen3-TTS-VD model is Tongyi's latest real-time speech synthesis model. It can perform high-fidelity real-time speech synthesis on the voices designed by the qwen3-voice-design service, and supports speech output in 11 languages using the same voice. This model has been trained on massive amounts of data, and the synthesized audio can adaptively adjust its tone according to the text, demonstrating good processing capabilities for complex text synthesis. This model is a snapshot version from January 15, 2026. |
|
Speech synthesis |
2026-01-14 |
|
Qwen 3-TTS-Flash model is Tongyi's latest real-time speech synthesis model. It can perform high-fidelity real-time speech synthesis on voices replicated from the qwen-voice-enrollment service, and supports speech output in 11 languages using the same voice. This model has been trained on massive amounts of data, and the synthesized audio can adaptively adjust tone according to the text, demonstrating good processing capabilities for complex text synthesis. This model is a snapshot version from January 15, 2026. |
|
Text generation |
2026-01-13 |
|
The Qwen Role-Playing Model Series is specifically optimized for muti-language anthropomorphic interaction scenarios. It demonstrates advanced capabilities in character consistency maintenance, context-aware dialogue progression, and empathetic engagement, enabling precise personalized character embodiment. This version significantly enhances Japanese linguistic localization (including dialects and honorifics), human-like role-playing authenticity, narrative coherence control, and scenario-based cognitive intelligence. |
|
Text generation |
2026-01-09 |
|
Dialogue Analysis-Pro model for advanced complex analysis of intricate quality inspection rules with complex business logic, supporting fine-grained analysis standards and featuring enhanced multi-turn context modeling, deep semantic understanding, and reasoning. |
|
Text generation |
2026-01-09 |
|
Dialogue Analysis-Flash model for daily tasks like dialogue information extraction and scenario classification, supporting custom analysis standards with enhanced semantic understanding for low-latency offline/online analysis. |
|
Image generation |
2026-01-09 |
|
The Qwen series of image-generation models boasts exceptional text-rendering capabilities and excels in complex text rendering as well as a wide range of generation and editing tasks. This version, a snapshot taken on January 9, 2026, is a distilled and accelerated variant of Qwen-Image-Max, enabling faster generation of high-quality images. |
|
Image generation |
2025-12-30 |
|
The Max series of qwen's image generation model excels across a wide range of generation tasks. Compared with the Plus series, it significantly reduces the "AI-like" feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering. |
|
Text generation, Reasoning |
2025-12-25 |
|
Zhipu's latest flagship with enhanced coding and multi-step reasoning capabilities, supporting long-term task planning, tool collaboration, and immersive writing/role-playing. |
|
Image generation |
2025-12-18 |
|
Z-Image-Turbo is a highly efficient image-generation model that has topped the Artificial Analysis benchmark as the world's No. 1 open-source text-to-image model. With just 6 billion parameters and an 8-step inference process, it generates photo-realistic images comparable to those produced by large-scale commercial models, while excelling in bilingual Chinese–English text rendering, complex semantic understanding, and diverse thematic generation. |
|
Reasoning, Visual understanding |
2025-12-18 |
|
The Qwen3 series of visual understanding models effectively integrates thinking and non-thinking modes. Compared to the snapshot released on September 23, this version delivers superior performance in reasoning and analysis tasks as well as style control, while also offering lower latency and faster response speeds. This version is based on a snapshot taken on December 19, 2025. |
|
Video generation |
2025-12-16 |
|
Wan2.6 reference to video, supports using a specified person or any object as a reference, precisely maintaining consistency of appearance and voice, and allows multi‑character reference for joint performances. |
|
Image generation |
2025-12-15 |
|
Wan2.6 text to image, Upgraded visual quality, aesthetics, and instruction-following deliver precise style control, realistic portraits, long-text understanding, and broad historical/cultural IP coverage, enabling high-quality, highly expressive visual generation. |
|
Image generation |
2025-12-15 |
|
Wan2.6 Image, An all-round image generation model that supports joint text–image reasoning, multi-image creative fusion, commercial-grade consistency, aesthetic style transfer, and precise control of framing and lighting, significantly enhancing consistency, controllability, and expressiveness in image generation. |
|
Image generation |
2025-12-15 |
|
The Qianwen series of Image Editing Plus models features enhanced character consistency, industrial design capabilities, and geometric reasoning abilities compared to the snapshot as of October 30. Additionally, it integrates LoRA capabilities such as lighting effects and effectively mitigates offset issues. This version is based on a snapshot taken on December 15, 2025. |
|
Speech synthesis |
2025-12-12 |
|
Qwen 3-TTS-VD model is Tongyi's latest real-time speech synthesis model. It can perform high-fidelity real-time speech synthesis on the voices designed by the qwen3-voice-design service, and supports speech output in 11 languages using the same voice. This model has been trained on massive amounts of data, and the synthesized audio can adaptively adjust its tone according to the text, demonstrating good processing capabilities for complex text synthesis. This model is a snapshot version from December 16, 2025. |
|
Speech synthesis |
2025-12-12 |
|
Qwen Voice-Design model is a series of voice design models from Qwen Speech Model. It only requires a simple text description to quickly design a suitable voice. When used in conjunction with the qwen3-tts-vd-realtime model, it can design and output speech in 10 languages. Furthermore, the synthesized audio can adaptively adjust its tone based on the text and has good processing capabilities for complex text synthesis. |
|
Realtime omni-modal |
2025-12-04 |
|
The real-time version of the Qwen3-Omni-Flash multimodal large-scale model, based on the Thinker–Talker Mixed Expert (MoE) architecture, supports efficient understanding and speech generation of text, images, audio, and video. It can interact with text in 119 languages and speech in 20 languages, generating human-like speech for precise cross-lingual communication. The model boasts powerful command-following and system prompt customization capabilities, flexibly adapting to conversational styles and character settings. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interaction experience. |
|
Omni-modal |
2025-12-04 |
|
Qwen3-Omni-Flash multimodal large-scale model, based on the Thinker–Talker Mixed Expert (MoE) architecture, supports efficient understanding and speech generation of text, images, audio, and video. It can interact with text in 119 languages and speech in 20 languages, generating human-like speech for precise cross-lingual communication. The model boasts powerful command-following and system prompt customization capabilities, flexibly adapting to conversational styles and character settings. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interaction experience. |
|
Realtime speech recognition |
2025-12-04 |
|
Qwen3-LiveTranslate-Flash is a high-precision, highly responsive, and robust multilingual real-time audio and video interpretation model. Leveraging Qwen3-Omni's powerful infrastructure, massive multimodal data, cross-language and cross-modal alignment, and visual enhancement technologies, Qwen3-LiveTranslate-Flash provides both offline and real-time audio and video translation capabilities. It can understand 19 languages and speak 10 languages, and also supports 8 Chinese dialects. |
|
Video generation |
2025-12-03 |
|
Wan2.6 text to video, Intelligent shot scheduling supports multi-shot storytelling, generating multi-shot narrative videos with consistent subjects, scenes, and atmosphere, with a maximum duration of 15 seconds. It delivers better instruction following, and improved visual fidelity. |
|
Video generation |
2025-12-03 |
|
Wan2.6 image to video, intelligent shot scheduling enables multi‑camera storytelling, delivers higher‑quality voice generation, supports stable multi‑speaker dialogue with more natural and realistic vocal timbres, and supports generation of clips up to 15 seconds in length. |
|
Text generation, Reasoning |
2025-12-02 |
|
DeepSeek-V3.2 is the official release of a model that incorporates DeepSeek Sparse Attention—a sparse attention mechanism. It's also the first model launched by DeepSeek that integrates reasoning into tool usage, supporting both reasoning-enabled and non-reasoning tool calls. |
|
Text generation, Reasoning |
2025-12-01 |
|
This version is a snapshot as of December 1, 2025, and features improved reasoning capabilities compared to the July 28 snapshot. Agent capabilities and multi-turn tool invocation abilities have been further enhanced, and performance on subjective creative tasks has improved significantly. It supports a context length of up to 1 million tokens, with tiered pricing based on context length. |
|
Speech synthesis |
2025-11-27 |
|
Qwen3-TTS-Flash model is Tongyi's latest real-time speech synthesis model. It can perform high-fidelity real-time speech synthesis on voices replicated by the qwen3-voice-enrollment service, and supports speech output in 11 languages with the same voice timbre. This model has been trained on massive amounts of data, and the synthesized audio can adaptively adjust the tone according to the text, and it also has good processing capabilities for complex text synthesis.This model is provided as a snapshot version. |
|
Speech synthesis |
2025-11-27 |
|
The Qwen3-TTS-Flash-Realtime model is Tongyi's latest real-time speech synthesis foundation model, featuring 17 expressive voices while delivering low-latency, high-stability audio synthesis. It supports multilingual and dialect outputs with consistent voice characteristics across languages. Trained on extensive datasets, the system autonomously adjusts vocal tones based on text semantics and demonstrates robust capabilities for complex content synthesis. |
|
Speech synthesis |
2025-11-27 |
|
The Qwen3-TTS-Flash is Tongyi's latest offline text-to-speech foundation model, featuring 17 expressive voices while enabling low-latency, high-stability audio synthesis. It supports multilingual and dialect outputs with consistent voice characteristics across languages. Trained on massive datasets, the system automatically adjusts vocal tones based on text semantics and demonstrates robust capabilities for synthesizing complex content. |
|
Speech synthesis |
2025-11-27 |
|
The Qwen Voice-Enrollment model is a series of voice replication models from the qwen speech model. It can quickly replicate highly similar voices using audio of only 5 seconds or more. When used in conjunction with the qwen3-tts-vc-realtime model, it can replicate a person's voice with high fidelity and output speech in 10 languages. Furthermore, the synthesized audio can adaptively adjust its tone according to the text and has good processing capabilities for complex text synthesis. |
|
Speech recognition |
2025-11-21 |
|
The Fun ASR version, updated in April 2026, fully supports the seven major dialect systems of traditional Chinese (Mandarin/Wu/Xiang/Gan/Hakka/Min/Yue) and is compatible with over 20 regional Mandarin accents. It features specific optimizations for the rhythm, meter, and literary expression characteristics of classical Chinese poetry, improving the accuracy of poetry recognition and making it suitable for cultural heritage, educational explanations, and audiobooks. Improved punctuation prediction and text normalization capabilities make the output text more consistent with written expression habits; numbers, dates, and amounts are automatically converted to standard formats, enhancing readability and professionalism. The language support has been expanded to include English, Japanese, Korean, Vietnamese, Thai, Indonesian, Malay, Filipino, Hindi, Arabic, French, German, Spanish, Portuguese, Russian, Italian, Dutch, Swedish, Danish, Finnish, Norwegian, Greek, Polish, Czech, Hungarian, Romanian, Bulgarian, Croatian, and Slovak, totaling 30 languages. This version is equivalent to the snapshot released on November 7, 2025. |
|
Visual understanding |
2025-11-20 |
|
Qwen-VL_OCR is an OCR model trained based on Qwen-VL. It aggregates various image-text recognition, parsing, and processing tasks through a unified model approach, offering powerful image-text recognition capabilities. |
|
Text generation |
2025-11-19 |
|
Qwen-MT-Lite is a large language model of the Qwen model series that specializes in multi-lingual translation. It provides high-quality and rapid translation services across 32 languages at a cost-effective price. It offers features such as terminology intervention, format preservation, and domain-specific translation to cater to the diverse needs of various applications, ensuring both efficiency and performance. |
|
Speech recognition |
2025-11-17 |
|
The large file transcription version of Qwen3-ASR-Flash. Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves high-precision speech recognition. It can automatically determine the language and accurately recognize speech in multiple languages, ensuring precise transcription even in complex audio environments. |
|
Realtime speech recognition |
2025-11-17 |
|
This is the real-time version of Tongyi Lab's next-generation end-to-end speech recognition model, based on leading proprietary speech technology, and boasts exceptional contextual awareness and high-precision speech transcription capabilities. Based on an end-to-end architecture, Fun-ASR integrates innovative RAG technology, supporting multi-dimensional features such as large-scale hotword customization, automatic filtering of sensitive and modal particles, ITN normalization, and punctuation prediction, significantly improving overall recognition accuracy and contextual relevance. Furthermore, Fun-ASR supports flexible switching between Chinese and English, covers multiple regional dialects, and boasts enhanced noise robustness, adapting to diverse and complex environments.This is a snapshot released on November 7, 2025. |
|
Speech recognition |
2025-11-17 |
|
The Fun ASR version, updated in April 2026, fully supports the seven major dialect systems of traditional Chinese (Mandarin/Wu/Xiang/Gan/Hakka/Min/Yue) and is compatible with over 20 regional Mandarin accents. It features specific optimizations for the rhythm, meter, and literary expression characteristics of classical Chinese poetry, improving the accuracy of poetry recognition and making it suitable for cultural heritage, educational explanations, and audiobooks. Improved punctuation prediction and text normalization capabilities make the output text more consistent with written expression habits; numbers, dates, and amounts are automatically converted to standard formats, enhancing readability and professionalism. The language support has been expanded to include English, Japanese, Korean, Vietnamese, Thai, Indonesian, Malay, Filipino, Hindi, Arabic, French, German, Spanish, Portuguese, Russian, Italian, Dutch, Swedish, Danish, Finnish, Norwegian, Greek, Polish, Czech, Hungarian, Romanian, Bulgarian, Croatian, and Slovak, totaling 30 languages. This is a snapshot released on November 7, 2025. |
|
Speech synthesis |
2025-11-17 |
|
Synthesis Capabilities: CosyVoice-v3-Flash is the latest high-performance speech synthesis model in the CosyVoice series from Tongyi Labs, offering improved naturalness, timbre, prosody, and emotional expressiveness compared to previous versions. This model supports real-time streaming text-to-speech synthesis. Cloning Capabilities: CosyVoice-v3-Flash is also the latest speech cloning model in the CosyVoice series from Tongyi Labs. Compared to previous versions, it improves pronunciation accuracy and timbre similarity, and adds support for more less commonly spoken languages (German, Spanish, French, Italian, Russian, Japanese). It can quickly generate highly similar and naturally sounding custom voices from just 5-20 seconds of reference audio. |
|
Visual understanding |
2025-11-12 |
|
GUI-series foundation model for cross-platform (mobile/desktop) interface understanding and interaction, supporting complex multi-step tasks and multi-agent collaboration. |
|
Text generation, Reasoning |
2025-11-10 |
|
Kimi-k2-thinking is a general agentic reasoning model developed by Moonshot AI, specialized in deep reasoning and multi-step tool integration to solve complex problems. |
|
Text generation |
2025-11-06 |
|
Qwen-MT-Flash, a large language model from the Qwen series, has been fully upgraded with the Qwen 3 architecture for significantly enhanced performance and translation quality. It provides rapid, cost-effective translation across 92 languages, while supporting advanced features such as terminology intervention, format preservation, and domain-specific adaptation. It is the ideal choice for applications requiring a powerful balance of speed, quality, and cost. |
|
Image generation |
2025-10-30 |
|
The qwen series of image editing Plus models further optimizes inference performance and system stability based on the initial Edit model, significantly reducing the response time for image generation and editing. It also supports returning multiple images in a single request, greatly enhancing user experience. |
|
Realtime speech recognition |
2025-10-27 |
|
The real-time version of Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves highly accurate speech recognition, automatically determining the language and accurately identifying speech in 11 languages, while ensuring precise transcription even in complex audio environments. |
|
Reasoning, Visual understanding |
2025-10-21 |
|
The largest dense model in the Qwen3-VL series, its reasoning version boasts multimodal reasoning capabilities second only to Qwen3-VL-235B-Thinking. It excels in STEM and math problem-solving, general image and video understanding, and achieves state-of-the-art performance in multimodal agent capabilities, making it ideal for complex multimodal reasoning tasks. |
|
Visual understanding |
2025-10-21 |
|
The largest dense model in the Qwen3-VL series, in its non-inference version, delivers overall performance second only to Qwen3-VL-235B-Instruct. It excels in document recognition and comprehension, demonstrates strong spatial awareness and object identification capabilities, and achieves state-of-the-art performance in 2D visual detection and spatial reasoning. It is well-suited for complex perception tasks across a wide range of general-purpose scenarios. |
|
Text embedding |
2025-10-21 |
|
A text-ranking model trained on the Qwen LLM foundation performs relevance ranking for input queries and candidate documents. It supports over 100 languages and long-text inputs, and is suitable for applications such as text retrieval and RAG. Its performance is aligned with the open-source Qwen3-Rerank series models. |
|
Multimodal embedding |
2025-10-21 |
|
Qwen2-5-VL-Embedding is a unified multimodal vector model based on Qwen2.5-VL, supporting text, image, video (single or mixed modality) input. It outputs unified representation vectors, suitable for cross-modal retrieval, image/video search, clustering, complex multimodal retrieval, and tagging. |
|
Text generation, Reasoning |
2025-10-21 |
|
GLM's new flagship model with comprehensive capability improvements over 4.5, featuring a 200K context window. |
|
Reasoning, Visual understanding |
2025-10-14 |
|
The Qwen3 series of small-scale visual understanding models effectively integrates thinking and non-thinking modes, delivering superior performance compared to the open-source Qwen3-VL-30B-A3B while maintaining fast response speeds. It features a comprehensive upgrade in image/video understanding, supporting ultra-long contexts such as extended videos and documents, spatial awareness, and object recognition across various domains. Equipped with 2D/3D visual localization capabilities, it is well-suited for tackling complex real-world tasks. |
|
Reasoning, Visual understanding |
2025-09-30 |
|
The Thinking version of the 8B Dense model in the Qwen3-VL series consumes less GPU memory and is capable of performing multimodal understanding and reasoning. It supports extremely long contexts such as lengthy videos and documents, 2D/3D visual localization, and features comprehensively upgraded image/video understanding, spatial awareness, and object recognition abilities. |
|
Visual understanding |
2025-09-30 |
|
The Instruct version of the 8B Dense model in the Qwen3-VL series requires less GPU memory and provides comprehensively upgraded image/video understanding, support for extremely long contexts like lengthy videos and documents, spatial awareness, and object recognition capabilities, making it suitable for tackling complex real-world tasks. |
|
Reasoning, Visual understanding |
2025-09-30 |
|
The Thinking version of the second-largest MoE model in the Qwen3-VL series features fast response speeds and enhanced multimodal understanding and reasoning capabilities, visual agents, and support for extremely long contexts such as lengthy videos and documents. It also boasts comprehensively upgraded image/video understanding, spatial awareness, and object recognition abilities, making it well-suited for complex real-world tasks. |
|
Visual understanding |
2025-09-30 |
|
The Instruct version of the second-largest MoE model in the Qwen3-VL series offers rapid response speeds and supports extremely long contexts like lengthy videos and documents. It includes comprehensively upgraded image/video understanding, spatial awareness, and object recognition capabilities, as well as 2D/3D visual localization, enabling it to handle intricate real-world challenges. |
|
Text generation, Reasoning |
2025-09-30 |
|
Experimental version introducing DeepSeek Sparse Attention (a sparse attention mechanism), exploring optimization and validation for training and inference efficiency on long texts. |
|
Speech recognition |
2025-09-25 |
|
Fun's multilingual speech recognition model supports over 31 languages and allows for free language switching, making it the top choice for users expanding overseas, especially to Southeast Asia. Fun-asr is an upgraded version of this model; switching to Fun-asr is recommended. |
|
Image generation |
2025-09-23 |
|
The upgraded Wan2.5 Preview image edit model, newly upgraded model architecture supports rich image editing capabilities via instruction control, with enhanced instruction adherence. It also enables multi-image reference generation with high consistency and demonstrates excellent text generation performance. |
|
Multimodal embedding |
2025-09-23 |
|
Embedding-Vision is a vision-centric multimodal embedding model powered by an LLM, featuring outstanding domain-specific performance and high cost-effectiveness in various domains (e.g., e-commerce, photo galleries, security, autonomous driving). With support for text, image, and video, it is applicable to downstream retrieval tasks, including text-to-image, image-to-image, text-to-video and video-to-video. |
|
Multimodal embedding |
2025-09-23 |
|
Embedding-Vision is a vision-centric multimodal embedding model powered by an LLM, featuring outstanding domain-specific performance and high cost-effectiveness in various domains (e.g., e-commerce, photo galleries, security, autonomous driving). With support for text, image, and video, it is applicable to downstream retrieval tasks, including text-to-image, image-to-image, text-to-video and video-to-video. |
|
Reasoning, Visual understanding |
2025-09-23 |
|
The Qwen3 series VL models effectively integrates thinking and non-thinking modes, achieving world-leading performance in visual agent capabilities on public benchmark datasets such as OS World. This version features comprehensive upgrades in areas like visual coding, spatial perception, and multimodal reasoning, significantly enhancing visual perception and recognition abilities, and supporting the understanding of ultra-long videos.This version is a snapshot as of September 23, 2025 |
|
Reasoning, Visual understanding |
2025-09-23 |
|
Qwen3 series VL models feature significantly enhanced multimodal reasoning capabilities, with a particular focus on optimizing the model for STEM and mathematical reasoning. Visual perception and recognition abilities have been comprehensively improved, and OCR capabilities have undergone a major upgrade. |
|
Visual understanding |
2025-09-23 |
|
The Qwen3 series VL models has been comprehensively upgraded in areas such as visual coding and spatial perception. Its visual perception and recognition capabilities have significantly improved, supporting the understanding of ultra-long videos, and its OCR functionality has undergone a major enhancement. |
|
Text generation, Reasoning |
2025-09-23 |
|
The Qwen 3 series Max model has undergone specialized upgrades in agent programming and tool invocation compared to the preview version. The officially released model this time has achieved state-of-the-art (SOTA) performance in its field and is better suited to meet the demands of agents operating in more complex scenarios. |
|
Realtime speech translation |
2025-09-23 |
|
The real-time version of Qwen3-LiveTranslate-Flash, which is a high-precision, highly responsive, and robust multilingual simultaneous audio and video interpretation model. Leveraging Qwen3-Omni's powerful infrastructure, massive multimodal data, cross-language and cross-modal alignment, and visual enhancement technologies, Qwen3-LiveTranslate-Flash offers both offline and real-time audio and video translation capabilities. It can understand 19 languages and speak 10 languages, including 8 Chinese dialects. |
|
Text generation |
2025-09-23 |
|
Qwen3-based code generation model with strong coding agent power, excels at tool calling and environment interaction, capable of autonomous programming with outstanding code capability while maintaining general ability. This is a snapshot from 23 September, 2025.Compared to the previous version (snapshot from July 22), it demonstrates improved robustness in downstream task performance and tool invocation, along with enhanced code security. |
|
Visual understanding |
2025-09-23 |
|
Qwen-VL_OCR is an OCR model trained based on Qwen-VL. It aggregates various image-text recognition, parsing, and processing tasks through a unified model approach, offering powerful image-text recognition capabilities. |
|
Image generation |
2025-09-23 |
|
The first image generation foundation model in the Qwen series that achieves significant advances in complex text rendering and precise image editing. Experiments show strong general capabilities in both image generation and editing, with exceptional performance in text rendering, especially for Chinese. |
|
Realtime speech recognition |
2025-09-23 |
|
This is the real-time version of Tongyi Lab's next-generation end-to-end speech recognition model, based on leading proprietary speech technology, and boasts exceptional contextual awareness and high-precision speech transcription capabilities. Based on an end-to-end architecture, Fun-ASR integrates innovative RAG technology, supporting multi-dimensional features such as large-scale hotword customization, automatic filtering of sensitive and modal particles, ITN normalization, and punctuation prediction, significantly improving overall recognition accuracy and contextual relevance. Furthermore, Fun-ASR supports flexible switching between Chinese and English, covers multiple regional dialects, and boasts enhanced noise robustness, adapting to diverse and complex environments. |
|
Image generation |
2025-09-22 |
|
The first Qwen image editing model extends Qwen-Image's text rendering to editing tasks. It offers precise bilingual (Chinese/English) text editing, dual visual and semantic editing, and strong cross-benchmark performance. |
|
Video generation |
2025-09-19 |
|
The upgraded Wan2.5 Preview text to video model, newly upgraded model architecture supports synchronized audio generation with visuals, enables 10-second long video generation, and offers enhanced instruction adherence, improved motion capabilities, and superior image quality. |
|
Image generation |
2025-09-19 |
|
The upgraded Wan2.5 Preview text to image model, newly upgraded model architecture significantly enhances visual aesthetics, design sensibility, and realistic texture. It excels in precise instruction adherence, generates text proficiently in English, Chinese, and less common languages, and supports the generation of complex structured long texts, charts, and architectural diagrams. |
|
Video generation |
2025-09-19 |
|
The upgraded Wan2.5 Preview image to video model, newly upgraded model architecture supports synchronized audio generation with visuals, enables 10-second long video generation, and offers enhanced instruction adherence, improved motion capabilities, and superior image quality. |
|
Video generation |
2025-09-19 |
|
wan2.2-animate-move is a character animation generation model. Users simply upload a character photo and a reference performance video, and the model transfers the expressions and actions from the video onto the character in the image, producing a high-fidelity animated video. |
|
Video generation |
2025-09-19 |
|
wan2.2-animate-mix is a character replacement model product. By uploading a character photo and a performance video, users can accurately replace the character in the original video with the character from the photo, while completely preserving environmental details such as the scene, lighting, and color tone of the original video. |
|
Realtime omni-modal |
2025-09-17 |
|
The real-time version of the Qwen3-Omni-Flash multimodal large-scale model, based on the Thinker–Talker Mixed Expert (MoE) architecture, supports efficient understanding and speech generation of text, images, audio, and video. It can interact with text in 119 languages and speech in 20 languages, generating human-like speech for precise cross-lingual communication. The model boasts powerful command-following and system prompt customization capabilities, flexibly adapting to conversational styles and character settings. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interaction experience. |
|
Omni-modal |
2025-09-17 |
|
Qwen3-Omni-Flash multimodal large-scale model, based on the Thinker–Talker Mixed Expert (MoE) architecture, supports efficient understanding and speech generation of text, images, audio, and video. It can interact with text in 119 languages and speech in 20 languages, generating human-like speech for precise cross-lingual communication. The model boasts powerful command-following and system prompt customization capabilities, flexibly adapting to conversational styles and character settings. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interaction experience.This version is a snapshot version from September 15, 2025. |
|
Speech recognition |
2025-09-17 |
|
Qwen3-Omni-30b-a3b-Captioner is a powerful fine-grained audio analysis model designed to generate accurate and comprehensive content descriptions in complex and changing audio scenarios. It can automatically parse and describe various audio content, from complex speech and ambient sounds to music and film and television sound effects, and can maintain stable and reliable output even in multi-source and mixed environments. |
|
Speech synthesis |
2025-09-16 |
|
The Qwen3-TTS-Flash-Realtime-2025-09-18 model is Tongyi's latest real-time speech synthesis foundation model, featuring 17 expressive voices while delivering low-latency, high-stability audio synthesis. It supports multilingual and dialect outputs with consistent voice characteristics across languages. Trained on extensive datasets, the system autonomously adjusts vocal tones based on text semantics and demonstrates robust capabilities for complex content synthesis.This model is provided as a snapshot version. |
|
Speech synthesis |
2025-09-16 |
|
The Qwen 3-TTS-Flash-2025-09-18 is Tongyi's latest offline text-to-speech foundation model, featuring 17 expressive voices while enabling low-latency, high-stability audio synthesis. It supports multilingual and dialect outputs with consistent voice characteristics across languages. Trained on massive datasets, the system automatically adjusts vocal tones based on text semantics and demonstrates robust capabilities for synthesizing complex content. This model is provided as a snapshot version. |
|
Visual understanding |
2025-09-13 |
|
Qwen-VL_OCR is an OCR model trained based on Qwen-VL. It aggregates various image-text recognition, parsing, and processing tasks through a unified model approach, offering powerful image-text recognition capabilities. |
|
Video generation |
2025-09-12 |
|
The all-new Wan2.2 First and Last Frame to video model is here. We've optimized motion stability and success rates, enhanced prompt adherence, and enabled seamless transitions between two images. |
|
Text generation, Reasoning |
2025-09-11 |
|
A new generation of Qwen3-based open-source thinking mode models. This version offers improved instruction following and streamlined summary responses over the previous iteration (Qwen3-235B-A22B-Thinking-2507). |
|
Text generation |
2025-09-11 |
|
A new generation of open-source, non-thinking mode model powered by Qwen3. This version demonstrates superior Chinese text understanding, augmented logical reasoning, and enhanced capabilities in text generation tasks over the previous iteration (Qwen3-235B-A22B-Instruct-2507). |
|
Text generation, Reasoning |
2025-09-11 |
|
As of the September 11, 2025 snapshot, this release features enhanced instruction following and streamlined summary responses in thinking mode. Non-thinking mode provides superior Chinese text understanding and augmented logical reasoning capabilities. The model supports a 1M context length with tiered pricing. |
|
Speech recognition |
2025-09-08 |
|
Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves high-precision speech recognition. It can automatically determine the language and accurately recognize speech in 11 languages, ensuring precise transcription even in complex audio environments. |
|
Text generation, Reasoning |
2025-09-05 |
|
A preview version of the Max model in the Qwen 3 series, achieving an effective integration of thinking and non-thinking modes. In thinking mode, there is a significant enhancement in capabilities such as intelligent agent programming, common-sense reasoning, and reasoning across mathematics, science, and general domains. |
|
Speech synthesis |
2025-09-03 |
|
Cloning capability: CosyVoice-v3-plus is the latest large voice cloning model in the CosyVoice series from Tongyi Lab. It offers superior sound quality and cloning fidelity, ideal for professional scenarios. With just 5-20 seconds of reference audio, it can rapidly generate a highly similar and natural-sounding custom voice. Synthesis capability: CosyVoice-v3-plus is the latest large speech synthesis model in the CosyVoice series from Tongyi Lab. It features enhanced sound quality and expressiveness, ideal for professional scenarios. The model supports real-time, streaming text-to-speech synthesis. |
|
Video generation |
2025-08-25 |
|
Auxiliary model for wan2.2-s2v to validate input portrait images against required specifications, ensuring compatibility with video generation via wan2.2-s2v. |
|
Video generation |
2025-08-25 |
|
High-quality Video Generation model that creates dynamic character videos (speaking/singing/performing) from input character images and voice audio files. |
|
Text generation, Reasoning |
2025-08-23 |
|
A hybrid inference architecture model supporting both thinking mode and non-thinking mode, featuring higher reasoning efficiency and stronger agent capabilities. |
|
Image generation |
2025-08-22 |
|
Qwen-MT-Image specializes in image translation, converting images across 11 languages (Chinese, English, Japanese, etc.) to target languages. It accurately preserves layout and content, supporting custom features like terminology definitions, sensitive word filtering, and product detection for flexible, accurate image localization. |
|
Text generation |
2025-08-22 |
|
Qwen-Deep-Research is an advanced agent system for complex research tasks, equipped with multi-step reasoning and global planning capabilities. It leverages tools like internet search to perform detailed task decomposition, reasoning, and analysis, ultimately generating traceable, logically rigorous research reports. |
|
Speech recognition |
2025-08-22 |
|
This is a new-generation large speech recognition model that focuses on Mandarin Chinese, English, and Japanese. It supports a wide range of regional dialects, offers enhanced noise robustness, and adapts to diverse and complex environments, making it the top recommendation for users in China.This is a snapshot released on August 25, 2025. |
|
Image generation |
2025-08-13 |
|
The first image generation foundation model in the Qwen series that achieves significant advances in complex text rendering and precise image editing. Experiments show strong general capabilities in both image generation and editing, with exceptional performance in text rendering, especially for Chinese. |
|
Video generation |
2025-08-11 |
|
The upgraded Wan 2.2 image to video Flash model, delivers faster speed with optimized stability, more powerful prompt following, improved consistency for text, portraits, and products, and precise shot control. |
|
Text generation |
2025-08-06 |
|
GLM-4.5-Air uses MoE architecture (106B total, 12B active parameters), a compact variant of GLM-4.5 for resource-constrained scenarios. |
|
Text generation |
2025-08-06 |
|
GLM-4.5 adopts a Mixture-of-Experts (MoE) architecture with 355B total parameters and 32B active parameters, excelling in complex reasoning, code generation, and agent interactions. |
|
Text generation |
2025-08-05 |
|
Based on Qwen3, this code generation model inherits the coding agent capabilities of Qwen3-Coder-Plus and supports multi-turn tool interaction. It features focused optimizations on repository-level understanding and enhanced tool-calling stability. |
|
Text generation, Reasoning |
2025-08-05 |
|
The Qwen3 Flash model offers a powerful fusion of thinking and non-thinking modes with dynamic in-conversation switching, excelling in complex reasoning while showing significant gains in instruction following and text comprehension. It supports a 1M context length and is billed on a tiered model corresponding to context usage. |
|
Text generation |
2025-07-31 |
|
Qwen3-based code generation model that inherits the coding agent ability of Qwen3-Coder-480B-A35B-Instruct; code capability reaches SOTA at the same scale. |
|
Text generation, Reasoning |
2025-07-31 |
|
The Qwen series model optimized for balanced performance, offering inference efficiency between Qwen-Max and Qwen-Turbo, is designed to handle moderately complex tasks effectively. This dynamically updated version implements changes without prior notice. |
|
Text generation, Reasoning |
2025-07-30 |
|
Open-source Qwen3 thinking model; compared to the previous version (Qwen3-30B-A3B) excels in complex thinking tasks, including logic, math, science, code, and other challenging scenarios; instruction following, text understanding, and multilingual translation capabilities significantly improved. |
|
Text generation, Reasoning |
2025-07-30 |
|
Qwen3 series Plus model, integrates thinking and non-thinking modes and can switch modes during dialogue. Compared to the prior version, adds dedicated enhancements for Chinese & English capabilities and tool calling. This is a snapshot from 28 July, 2025; first to support 1 M context length, and uses tiered pricing. |
|
Text generation |
2025-07-29 |
|
Open-source Qwen3 non-thinking model; compared to the previous version (Qwen3-30B-A3B) shows major improvements in Chinese, English, and overall multilingual general capabilities. Optimized for subjective open-ended tasks, delivering responses significantly more aligned with user preferences and more helpful. |
|
Video generation |
2025-07-28 |
|
The upgraded Wan 2.2 Plus text to video model, delivers higher quality results with stable sweeping complex movements, cinematic vision control, more powerful prompt following, and realistic world recreation. |
|
Image generation |
2025-07-28 |
|
The upgraded Wan 2.2 Plus text to image model, delivers richer image detail with enhanced creativity, stability, and realism. It also features stronger prompt following and native support for multiple styles. Up to 2 million pixel generation and prompt enhancement are supported as well. |
|
Image generation |
2025-07-28 |
|
The upgraded Wan 2.2 Flash text to image model, delivers faster speed with enhanced creativity, stability, and realism. It also features stronger prompt following and native support for multiple styles. Up to 2 million pixel generation and prompt enhancement are supported as well. |
|
Video generation |
2025-07-28 |
|
The upgraded Wan 2.2 Plus image to video model, delivers higher quality results with optimized stability, more powerful prompt following, improved consistency for text, portraits, and products, and precise shot control. |
|
Text generation, Reasoning |
2025-07-25 |
|
Open-source Qwen3 thinking model; compared to the previous version (Qwen3-235B-A22B) shows major improvements in logical ability, general capabilities, knowledge enhancement, and creativity, suitable for high-difficulty, strong-thinking scenarios. |
|
Text generation |
2025-07-23 |
|
Qwen-Doc-Turbo rapidly extracts precise information from documents, supporting tagging, classification, content moderation, and summarization. |
|
Text generation |
2025-07-22 |
|
Qwen3-based code generation model with strong coding agent power, excels at tool calling and environment interaction, capable of autonomous programming with outstanding code capability while maintaining general ability. |
|
Text generation |
2025-07-22 |
|
Qwen3-based code generation model with strong coding agent power; code capability reaches open-source SOTA. |
|
Text generation |
2025-07-22 |
|
Open-source Qwen3 non-thinking model; compared to the previous version (Qwen3-235B-A22B) shows slight improvements in subjective creativity and model safety. |
|
Text generation |
2025-07-22 |
|
Qwen-MT-Turbo is a large language model within the Qwen model series that specializes in multi-lingual translation. It provides high-quality and rapid translation services across 92 languages at a cost-effective price point. It also offers features such as terminology intervention, format preservation, and domain-specific translation to cater to the diverse needs of various applications, ensuring both efficiency and performance. |
|
Text generation |
2025-07-22 |
|
Qwen-MT-Plus, the flagship translation model from our Qwen series, is now fully upgraded with the Qwen3 architecture. It supports 92 languages and delivers exceptionally accurate and natural-sounding translations. Its advanced capabilities in contextual understanding, terminology control, and format preservation make it a superior choice over traditional models, especially for specialized domains. |
|
Realtime speech synthesis |
2025-07-16 |
|
Qwen-TTS-Realtime is a speech synthesis model in the Qwen series with bidirectional context awareness. It enables low-latency, high-fidelity generation of multi-voice, dialectal, and long-text bidirectional streaming. |
|
Text generation, Reasoning |
2025-07-16 |
|
The Plus model of the Qwen3 Series, achieving effective integration of thinking mode and non-thinking mode, allows switching modes during conversations. This is a snapshot from July 14, 2025. Compared to the previous version, there has been a significant improvement in both Chinese and English capabilities under non-thinking mode, with enhanced tool-calling abilities. |
|
Text generation |
2025-07-16 |
|
Kimi-K2 is Moonshot's first open-source trillion-parameter MoE model in China, featuring 32B activated parameters with exceptional coding and tool-calling capabilities. |
|
Speech synthesis |
2025-06-26 |
|
Qwen-TTS is the first speech synthesis model in the Qwen series, supporting Chinese, English, and mixed inputs. It adaptively adjusts output tone based on text, offers natural voice quality, and supports streaming output. |
|
Text generation, Reasoning |
2025-06-24 |
|
The Turbo model of the Qwen3 series. It effectively integrates thinking mode and non-thinking mode, allowing for mode switching during conversations. Its reasoning capabilities rival those of QwQ-32B with a smaller parameter size, while its general capabilities significantly surpass those of Qwen2.5-Turbo, achieving the SOTA level in the same scale within the industry. |
|
Text generation, Reasoning |
2025-06-24 |
|
Qwen-Plus is an enhanced version of the Qwen ultra-large language model that supports multiple input languages such as Chinese and English. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
|
Reasoning |
2025-06-18 |
|
A large language model enhanced through extensive reinforcement learning during post-training, achieving strong reasoning capabilities with minimal labeled data. Performs well in mathematics, coding, and natural language reasoning tasks. |
|
Visual understanding |
2025-06-13 |
|
Qwen-VL-Plus is the enhanced version of the large visual language model. It significantly improves detail recognition and text recognition capabilities, supporting images with resolutions exceeding one million pixels and any aspect ratio specifications. The model delivers exceptional performance across a wide range of visual tasks. |
|
Text embedding |
2025-06-05 |
|
The General Text Vector V4 version is a multi-language text vector model developed by the Tongyi Lab based on Qwen3. Compared to the V3 version, it significantly improves performance in text retrieval, clustering, and classification tasks. It achieves a 15% to 40% improvement in evaluation tasks such as MTEB multilingual, Chinese-English, and code retrieval. Additionally, it supports user-defined vector dimensions ranging from 64 to 2048. |
|
Reasoning |
2025-06-04 |
|
A large language model enhanced through extensive reinforcement learning during post-training, achieving strong reasoning capabilities with minimal labeled data. Performs well in mathematics, coding, and natural language reasoning tasks. |
|
Reasoning, Visual understanding |
2025-06-03 |
|
Qwen QVQ Visual Reasoning Model Plus version supporting visual input and chain-of-thought output with enhanced capabilities in mathematics, programming, visual analysis, creation, and general tasks. |
|
Speech synthesis |
2025-05-27 |
|
A generative speech synthesis model combining text understanding and speech generation through large-scale pretrained language models, supporting real-time streaming text-to-speech synthesis. |
|
Visual understanding |
2025-05-26 |
|
Qwen-VL-Max is a large-scale visual language model of the Qwen series. Compared to the Plus version, it further enhances visual reasoning capabilities and instruction-following abilities, offering higher levels of visual perception and cognition. It delivers optimal performance on more complex tasks. |
|
Video generation |
2025-05-14 |
|
Wanxiang 2.1 - Video Editing Unified Model - Plus. Supports localized editing, video redrawing, background expansion, duration extension, and reference-based generation. Enables multi-modal control through text, images, or videos. |
|
Realtime omni-modal |
2025-05-08 |
|
The real-time version of Qwen's new large multimodal understanding and generation model, suitable for real-time audio interaction scenarios. It supports the understanding of audio accompanied by text, images, and video mixed inputs, and can simultaneously generate speech and text in stream, providing four natural tones. |
|
Text generation, Reasoning |
2025-04-29 |
|
Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability reaches the SOTA level in the same scale industry, and its general capability significantly surpasses Qwen2.5-7B. |
|
Text generation, Reasoning |
2025-04-29 |
|
Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability significantly surpasses QwQ, and its general capability markedly exceeds Qwen2.5-32B-Instruct, reaching the SOTA level in the same scale industry. |
|
Text generation, Reasoning |
2025-04-29 |
|
Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability rivals QwQ-32B with a smaller parameter size, and its general capability significantly surpasses Qwen2.5-14B, reaching the SOTA level in the same scale industry. |
|
Text generation, Reasoning |
2025-04-29 |
|
Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability significantly surpasses QwQ, and its general capability markedly exceeds Qwen2.5-72B-Instruct, reaching the SOTA level in the same scale industry. |
|
Text generation, Reasoning |
2025-04-29 |
|
The Plus model of the Qwen3 series, effectively integrates thinking mode and non-thinking mode, allowing for mode switching during conversations. Its reasoning capabilities significantly surpass those of QwQ, and its general capabilities notably exceed those of Qwen2.5-Plus, reaching the SOTA level in the same scale within the industry. This model is the snapshot from April 28, 2025. |
|
Text generation, Reasoning |
2025-04-28 |
|
Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability reaches the SOTA level in the same scale industry, and its general capability significantly surpasses Qwen2.5-14B. |
|
Image generation |
2025-04-28 |
|
A high-quality virtual try-on image generation model that produces try-on effects with enhanced image clarity, garment texture details, and logo reconstruction compared to aitryon. Requires longer generation time, suitable for non-time-critical scenarios. |
|
Visual understanding |
2025-04-23 |
|
Qwen-VL_OCR is an OCR model trained based on Qwen-VL. It aggregates various image-text recognition, parsing, and processing tasks through a unified model approach, offering powerful image-text recognition capabilities. |
|
Video generation |
2025-04-21 |
|
Wanxiang 2.1 - Keyframe-to-Video - Plus. Generates smooth transition videos from two input images. Supports complex large-scale motion, physics adherence, diverse artistic styles, and cinematic-grade visual quality. Enhanced instruction-following capability and richer visual details. |
|
Speech synthesis |
2025-04-20 |
|
Qwen-TTS is the first speech synthesis model in the Qwen series, supporting Chinese, English, and mixed inputs. It adaptively adjusts output tone based on text, offers natural voice quality, and supports streaming output. |
|
Omni-modal |
2025-03-27 |
|
The new multimodal understanding and generation model, supporting text, image, speech, and video input as well as mixed input understanding. It can simultaneously stream text and speech, with significantly enhanced speed for multimodal content understanding. Additionally, it offers four natural dialogue tones. This model is the snapshot from March 26, 2025, with significant improvements on visual capabilities over the snapshot of January 19, 2025. |
|
Omni-modal |
2025-03-26 |
|
The new multi-modal understanding and generation model trained based on Qwen2.5. It supports text, image, speech, video, and mixed input understanding and can simultaneously generate streams of text and speech, significantly improves the speed of multi-modal content understanding. It provides four natural tones. |
|
Reasoning, Visual understanding |
2025-03-26 |
|
The Tongyi Qianwen QVQ visual reasoning model supports visual input and chain-of-thought output, demonstrating stronger capabilities in mathematics, programming, visual analysis, creation, and general tasks. |
|
Image generation |
2025-03-25 |
|
Image Editing model supporting preset and command-based tasks, including global/local editing (style transfer, inpainting, expansion, super-resolution) and reference-based generation. |
|
Speech synthesis |
2025-03-20 |
|
Provides text-to-speech service combining SAMBERT+NSFGAN deep neural network algorithms with traditional domain expertise, featuring accurate pronunciation, natural prosody, high voice fidelity, and strong expressiveness. |
|
Text generation |
2025-03-20 |
|
The role-playing model of the Qwen series. This is a dynamically updated version, and notifications will be provided in advance for any model updates. It is suitable for anthropomorphic role-playing and has optimized capabilities in following predefined character instructions, advancing conversations, and demonstrating active listening and empathy. Additionally, it supports the deep restoration of personalized characters. |
|
Text embedding |
2025-03-20 |
|
A multilingual text reranking model for semantic retrieval and RAG, ranking candidate documents by relevance to queries. |
|
Text generation |
2025-03-19 |
|
Qwen-Long is a large language model designed for ultra-long context processing, supporting Chinese, English, and other languages. It handles up to 10 million tokens (approximately 15 million characters or 15,000 document pages) in dialogues. Integrated with document services, it supports parsing and dialogue for text files (TXT, DOCX, PDF, XLSX, EPUB, MOBI, MD, CSV) and image files (BMP, PNG, JPG/JPEG, GIF, PDF scans). Notes: HTTP requests support up to 1M tokens; file submission is recommended for longer content. |
|
Speech synthesis |
2025-03-18 |
|
Provides text-to-speech service combining SAMBERT+NSFGAN deep neural network algorithms with traditional domain expertise, featuring accurate pronunciation, natural prosody, high voice fidelity, and strong expressiveness. |
|
Speech synthesis |
2025-03-18 |
|
A next-generation generative speech synthesis model integrating text understanding and speech generation via large-scale pretrained language models, supporting real-time streaming text-to-speech synthesis. |
|
Omni-modal |
2025-03-17 |
|
The new multimodal understanding and generation model, supporting text, image, speech, and video input as well as mixed input understanding. It can simultaneously stream text and speech, with significantly enhanced speed for multimodal content understanding. Additionally, it offers four natural dialogue tones. |
|
Reasoning |
2025-03-05 |
|
The enhanced version of the Qwen QwQ reasoning model, trained on the Qwen2.5 model, has significantly improved its reasoning capabilities through reinforcement learning. The model's core metrics in mathematics and coding (e.g., AIME 24/25, LiveCodeBench) as well as some general metrics (e.g., IFEval, LiveBench) have reached the level of the full version of DeepSeek-R1. |
|
Realtime speech translation |
2025-03-04 |
|
A multilingual speech-to-text and translation model providing high-accuracy real-time transcription for 10 languages (including Chinese/English/Japanese/Korean) with cross-lingual translation. |
|
Speech recognition |
2025-03-04 |
|
A multilingual speech-to-text and translation model supporting real-time recognition of 10 mixed languages (60s limit) with Chinese/English/Japanese/Korean mutual translation. |
|
Video generation |
2025-02-27 |
|
Image-to-Video Generation-Turbo model that transforms images into dynamic videos with complex movements and cinematic aesthetics, featuring faster generation speed and enhanced instruction compliance. |
|
Omni-modal |
2025-02-14 |
|
The new multimodal understanding and generation model, supporting text, image, speech, and video input as well as mixed input understanding. It can simultaneously stream text and speech, with significantly enhanced speed for multimodal content understanding. Additionally, it offers four natural dialogue tones. |
|
Text generation |
2025-02-05 |
|
A distilled large language model based on Qwen2.5-Math-7B, trained using DeepSeek R1's outputs. |
|
Text generation |
2025-02-05 |
|
A distilled large language model based on Qwen2.5-32B, trained using DeepSeek R1's outputs. |
|
Text generation |
2025-02-05 |
|
A distilled large language model based on Qwen2.5-14B, trained using DeepSeek R1's outputs. |
|
Text generation |
2025-02-05 |
|
A distilled large language model based on Qwen2.5-Math-1.5B, trained using DeepSeek R1's outputs. |
|
Text generation |
2025-02-03 |
|
The Qwen series of models, which are well-balanced in capabilities, offer reasoning performance and speed that fall between Qwen-Max and Qwen-Turbo, making them suitable for moderately complex tasks. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
|
Text generation, Reasoning |
2025-01-27 |
|
A self-developed Mixture-of-Experts (MoE) model with 671B parameters (activating 37B), pre-trained on 14.8T tokens. Demonstrates excellent capabilities in long-text processing, coding, mathematics, encyclopedic knowledge, and Chinese language tasks. |
|
Video generation |
2025-01-20 |
|
Image-to-Video Generation-Plus model that converts images into dynamic videos with complex motions, physical realism, cinematic quality, and improved instruction adherence for higher video fidelity. |
|
Image generation |
2025-01-20 |
|
Enhanced Text-to-Image model specializing in realistic portraits and creative designs, upgraded in aesthetics, realism, and artistic quality with up to 2MP resolution and smart prompt rewriting support. |
|
Video generation |
2025-01-16 |
|
Emoji is a facial animation video generation model that creates facial animation videos using face images and preset dynamic templates. |
|
Video generation |
2025-01-16 |
|
Emoji-Detect is an image detection model assisting emoji generation, used to detect whether character images meet video generation requirements. |
|
Text generation |
2025-01-15 |
|
Qwen-Plus is an enhanced version of the Qwen ultra-large language model that supports multiple input languages such as Chinese and English. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
|
Image generation |
2025-01-15 |
|
An image segmentation model serving as a supporting module for AI try-on OutfitAnyone, capable of segmenting model images and garment images for pre/post-processing of try-on images. |
|
Video generation |
2025-01-09 |
|
Wanxiang 2.1 - Text-to-Video - Turbo. High-speed video generation with support for complex motion, physics simulation, artistic styles, and cinematic quality. Improved instruction-following capability. |
|
Video generation |
2025-01-09 |
|
Wanxiang 2.1 - Text-to-Video - Plus. Generates high-quality videos from text prompts. Supports complex motion, realistic physics simulation, diverse artistic styles, and cinematic visuals. Enhanced instruction-following performance. |
|
Image generation |
2025-01-09 |
|
Wanxiang 2.1 - Text-to-Image - Turbo. Accelerated generation speed with enhanced aesthetics, realism, and artistry. Maintains strong semantic understanding, style diversity, and 2MP resolution support. Includes smart prompt rewriting functionality. |
|
Image generation |
2025-01-09 |
|
Wanxiang 2.1 - Text-to-Image - Plus. Upgraded image generation with enhanced aesthetics, realism, and artistry. Improved semantic understanding, style generalization, and support for up to 2MP resolution. Features smart prompt rewriting and richer visual details. |
|
Realtime speech recognition |
2024-12-31 |
|
Recommended Paraformer real-time speech recognition model supporting multilingual switching for live streaming/meetings. Enables language selection via language_hints parameter for enhanced accuracy in 8kHz customer service scenarios. Supports: Mandarin (including dialects), English, Japanese, Korean. |
|
Text generation |
2024-12-26 |
|
Qwen-Plus is an enhanced version of the Qwen ultra-large language model that supports multiple input languages such as Chinese and English. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
|
Multimodal embedding |
2024-12-23 |
|
Multimodal vector model developed by Tongyi Lab based on pre-trained multimodal foundation models. Generates high-dimensional continuous vectors from text, image, or video inputs for downstream tasks including search, classification, and content moderation. |
|
Speech recognition |
2024-12-19 |
|
Latest Paraformer Mandarin speech recognition model with enhanced architecture, improved recognition accuracy, and 8kHz telephone speech support (Chinese hotword only). |
|
Text generation |
2024-12-12 |
|
Intent Detection and Slot Filling model for dialogue systems, enabling joint prediction of API-based intents and slot parameters in a single output, returning standardized JSON results with multiple API commands and filled slots. |
|
Video generation |
2024-12-10 |
|
Character Video Generation model that synthesizes lip-synced videos based on input character videos and voice audio, matching mouth movements to the audio content. |
|
Video generation |
2024-12-10 |
|
A motion template generation model assisting AnimateAnyone, capable of extracting human motions from videos and creating templates. |
|
Video generation |
2024-12-10 |
|
A video generation model that creates full-body motion videos based on human portraits and motion templates. |
|
Video generation |
2024-12-10 |
|
An image detection model assisting AnimateAnyone, used to verify if human figures in images meet video generation requirements. |
|
Visual understanding |
2024-11-14 |
|
Qwen-VL_OCR is an OCR model trained based on Qwen-VL. It aggregates various image-text recognition, parsing, and processing tasks through a unified model approach, offering powerful image-text recognition capabilities. |
|
Text generation |
2024-11-12 |
|
Qwen-Coder-Plus is a specialized language model for programming and code generation, delivering excellent performance and outstanding results. |
|
Video generation |
2024-11-07 |
|
LivePortrait-detect is an auxiliary image detection model for assessing whether human figures in images meet video generation requirements. |
|
Video generation |
2024-11-07 |
|
LivePortrait is a video generation model that creates lightweight dynamic portrait videos from static images. |
|
Video generation |
2024-11-07 |
|
EMO is a video generation model that generates high-quality dynamic portrait videos based on character images. |
|
Video generation |
2024-11-07 |
|
EMO-Detect is an image detection model assisting EMO, used to detect whether character images meet video generation requirements. |
|
Text generation |
2024-10-15 |
|
Qwen-Max supports a parameter scale of hundreds of billions and multiple input languages such as Chinese and English. Qwen-Max is updated in a rolling manner. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
|
Text generation |
2024-09-19 |
|
Qwen-Math-Turbo is a specialized language model for math problem-solving, featuring high inference speed and low cost. |
|
Text generation |
2024-09-19 |
|
Qwen-Math-Plus is a powerful math problem-solving model excelling in Chinese and English mathematical tasks, including equations, calculations, and proofs. |
|
Text generation |
2024-09-19 |
|
Qwen-Coder-Turbo is a specialized language model for programming and code generation, featuring high inference speed and low cost. |
|
Speech synthesis |
2024-09-13 |
|
A large-model voice replication service used in conjunction with Cosyvoice-v3. Utilizing advanced large-model technology for feature extraction, it can replicate voices without a training process. Only a very short audio clip is required to quickly generate a highly similar and natural-sounding custom voice. |
|
Video generation |
2024-09-13 |
|
Video Style Repainting model that applies artistic styles (e.g., Japanese comic, 3D cartoon) to input video frames while preserving original appearances, generating stylized videos with diverse visual effects. |
|
Speech recognition |
2024-09-13 |
|
Speech-Biasing allows users to define hotwords (specific terms or phrases) with prioritized recognition/translation. This improves accuracy for domain-specific terms in speech recognition and translation tasks. |
|
Text generation |
2024-09-13 |
|
Qwen-Math-Plus is a powerful math problem-solving model excelling in Chinese and English mathematical tasks, including equations, calculations, and proofs. |
|
Speech synthesis |
2024-09-13 |
|
A voice cloning model leveraging advanced large model technology for feature extraction without training. Generates highly similar, natural-sounding custom voices from short audio samples. |
|
Speech recognition |
2024-09-06 |
|
Recommended Paraformer speech recognition model supporting multilingual recognition with language_hints parameter. Supports all sampling rates and hotwords across: Mandarin (including dialects), English, Japanese, Korean. |
|
Realtime speech recognition |
2024-09-06 |
|
Recommended Paraformer real-time speech recognition model supporting multilingual switching with language_hints parameter. Supports all sampling rates and hotwords across: Mandarin (including dialects), English, Japanese, Korean. |
|
Image generation |
2024-08-19 |
|
Person instance segmentation employs detection and segmentation techniques to identify objects in images and generate pixel-level masks for precise object boundary delineation. |
|
Image generation |
2024-08-19 |
|
Image erasure and completion tool removing specified elements (people, objects, text, watermarks) while preserving backgrounds using computer vision and AIGC inpainting techniques. |
|
Text embedding |
2024-07-12 |
|
The general text vectorization model is a multilingual unified text vectorization model developed by Tongyi Lab based on the large language model (LLM) foundation. This model is designed for multiple mainstream languages worldwide and provides advanced vectorization services to help developers convert text data into high-quality vector data. |
|
Image generation |
2024-06-28 |
|
An image refinement module for secondary generation of AI try-on results, producing higher-fidelity try-on images with improved realism. |
|
Image generation |
2024-06-25 |
|
Virtual Model that replaces models and backgrounds in product images with virtual alternatives while maintaining poses, supporting interactive products (e.g., apparel, footwear, accessories). |
|
Image generation |
2024-06-25 |
|
Virtual Model that intelligently replaces models and backgrounds in product images while retaining poses, enabling diverse showcase of interactive products (e.g., apparel, accessories) with virtual alternatives. |
|
Image generation |
2024-06-21 |
|
Creative Poster Generation model that automatically creates poster layouts and typography based on user input, supporting diverse styles (promotion, greeting) for personalized designs without design expertise. |
|
Image generation |
2024-06-21 |
|
ShoeModel-V1 supports AI try-on for footwear by inputting multi-view shoe images and reshaping shoe regions in template images. It generates natural layouts, rich details, and realistic try-on results for applications like model product design, AI shoe try-on, and layout refinement. |
|
Image generation |
2024-06-11 |
|
Doodle-to-Image model that generates artistic doodles from hand-drawn sketches and text descriptions, offering 5 styles (flat illustration, oil painting, etc.) for creative, educational, and design applications. |
|
Realtime speech recognition |
2024-06-06 |
|
Paraformer Mandarin real-time speech recognition model for 16kHz+ live streaming/meeting scenarios. |
|
Realtime speech recognition |
2024-06-06 |
|
Paraformer Mandarin real-time speech recognition model optimized for 8kHz customer service scenarios. |
|
Image generation |
2024-05-28 |
|
Image Local Repainting model based on the Composer framework, generating semantically consistent partial edits with diverse styles using input images, masks, and text prompts, delivering natural layouts with rich details. |
|
Image generation |
2024-05-24 |
|
Image expansion model for free-form image extension with rotation support. Allows expansion via expansion ratio or pixel count parameters for creative design, auxiliary drawing, and film post-production applications. |
|
Image generation |
2024-05-24 |
|
A high-performance virtual try-on image generation model that produces try-on effect images based on garment flat lay photos and human frontal full-body portraits. It generates try-on images quickly, suitable for time-sensitive scenarios. |
|
Text generation |
2024-05-20 |
|
Qwen-Long is a large language model designed for ultra-long context processing, supporting Chinese, English, and other languages. It handles up to 10 million tokens (approximately 15 million characters or 15,000 document pages) in dialogues. Integrated with document services, it supports parsing and dialogue for text files (TXT, DOCX, PDF, XLSX, EPUB, MOBI, MD, CSV) and image files (BMP, PNG, JPG/JPEG, GIF, PDF scans). Notes: HTTP requests support up to 1M tokens; file submission is recommended for longer content. |
|
Text generation |
2024-05-14 |
|
Tongyi FaRui is a legal industry large model trained on legal data and knowledge, integrating fine-tuning, reinforcement learning, RAG retrieval augmentation, and legal Agent technology, capable of answering legal questions, case analysis, legal document generation, and contract review. |
|
Image generation |
2024-04-09 |
|
WordArt Jinshu - Text Texture Generation. Applies material and texture overlays to text content or images based on prompts, achieving 3D effects or scene integration. Produces high-quality artistic text suitable for poster backgrounds. |
|
Image generation |
2024-04-09 |
|
WordArt Jinshu - Text Deformation. Creatively morphs input text outlines based on prompts, enabling versatile typographic designs. Outputs black-background white-mask images with transformed text content. |
|
Text embedding |
2024-04-09 |
|
Universal Text Vector, a multilingual text embedding model developed by Tongyi Lab based on the LLM foundation, providing high-quality vector services for major global languages to convert text data into high-quality vector representations. |
|
Text embedding |
2024-04-09 |
|
Text-Embedding-V1 is a multilingual text vector model based on Qwen's LLM foundation, providing high-quality vector services for major global languages. It helps developers convert text data into high-quality vectors efficiently. |
|
Text embedding |
2024-04-09 |
|
Text-Embedding-Async-V2 is a batch processing interface for general text vectors. Users submit bulk vector computation requests via text, and the platform stores results in downloadable files after processing. |
|
Text embedding |
2024-04-09 |
|
Text-Embedding-Async-V1 is a batch processing interface for general text vectors. Users submit bulk vector computation requests via text, and the platform stores results in downloadable files after processing. |
|
Speech recognition |
2024-04-09 |
|
Paraformer bilingual Mandarin-English speech recognition model for 16kHz+ audio/video processing. |
|
Speech recognition |
2024-04-09 |
|
Paraformer multilingual speech recognition model supporting audio/video with 16kHz+ sampling rates across 14 languages/dialects including Mandarin, Cantonese, Wu, Minnan, and major global languages. |
|
Speech recognition |
2024-04-09 |
|
Paraformer speech recognition file transcription API for converting common audio/video files to text, supporting 8kHz Mandarin telephone speech recognition. |
|
Image generation |
2024-04-09 |
|
Generates portraits of trained character models in preset styles (e.g., ID photos, business portraits). |
|
Image generation |
2024-04-09 |
|
Detects user-uploaded character images to determine if faces meet FaceChain fine-tuning standards, evaluating face count, size, angle, lighting, clarity, etc., supporting batch input with per-image results. |
|
Image generation |
2024-03-22 |
|
Portrait Style Repainting model that transforms input portraits into artistic styles (e.g., watercolor, oil painting) while preserving original facial features. |
|
Image generation |
2024-03-22 |
|
Image Background Generation model that extends foreground images with natural lighting and realistic details, supporting text/image-guided generation and intelligent text overlay. |
|
Image generation |
2024-01-05 |
|
Multilingual Text-to-Image Generation model supporting Chinese/English inputs, covering key styles including watercolor, oil painting, Chinese ink painting, sketch, flat illustration, anime, and 3D cartoon. |
Singapore
|
Model type |
Date |
Service scope |
Model ID |
Description |
|
Speech recognition |
2026-07-30 |
International |
|
Added the Qwen-Audio-3.0-ASR-Flash-Streaming (real-time), Qwen-Audio-3.0-ASR-Flash-Filetrans (non-real-time), and Qwen-Audio-3.0-ASR-Flash (non-real-time) models: Dialect support: Supports the seven major Chinese dialect groups (Mandarin, Wu, Xiang, Gan, Hakka, Min, and Yue) and more than 20 regional accents; Classical poetry optimization: Improves recognition accuracy for classical Chinese poetry, making it suitable for education, culture, and audiobook scenarios; Text optimization: Enhances punctuation prediction and text normalization, automatically converting numbers, dates, and monetary amounts to standard formats; Multilingual expansion: Supports 30 languages, including Chinese, English, Japanese, and Korean; Hotwords and context: Supports hotwords (precompiled and on-the-fly) and context input to improve recognition accuracy for domain-specific terms. |
|
Text generation, Reasoning, Visual understanding |
2026-07-21 |
International |
|
The Qwen3.7 native vision-language series Flash models comprehensively enhance multimodal understanding and Agent execution capabilities compared to 3.6-Flash. Key improvements include strengthened foundational multimodal abilities, enhanced object recognition, improved real-world perception and spatial intelligence. Multimodal Agent scenarios such as Search Agent and CI Agent have seen significant upgrades, with more stable end-to-end task execution. Multimodal coding capabilities are optimized, delivering a smoother vibe coding experience. |
|
Image generation |
2026-07-20 |
International |
|
Rich content: Supports input of up to 4.5k tokens and dense information layout with images-within-images, enabling complex layouts like newspapers, storyboards, menus, and exam papers to be generated in a single pass. Authentic detail: Supports precise rendering of text as small as 10px, and vividly reproduces fine details such as micro-expressions, pores, and individual strands of hair—approaching the quality of real photography. Deep knowledge: Supports native rendering of 12 languages and 20+ fonts, realistic simulation of mainstream interfaces such as web pages, games, and live streams, fully incorporating external knowledge. Qwen-Image-3.0-Pro isn't just pursuing "good looks"—it's pursuing "usefulness", making image generation a truly deployable productivity tool. |
|
Realtime speech synthesis |
2026-07-14 |
International |
|
Qwen-Audio-3.0-TTS-Plus is a high-performance speech synthesis model, designed for high-quality speech generation scenarios. Compared with the previous version, it supports more low-resource languages and Chinese dialects, significantly improves dialect authenticity, and enhances free-style instruction following and fine-grained tag control for more accurate control over emotion, tone, character, speaking rate, volume, and synthesis style. It is also more robust under noisy and reverberant acoustic conditions, with further improvements in audio quality, clarity, resolution, and overall expressiveness. The Plus version focuses more on synthesis quality and detailed expressiveness, making it suitable for professional scenarios with higher requirements for audio quality, naturalness, and expressiveness, such as content creation, audiobooks, film and video dubbing, brand voice design, and premium speech services. |
|
Realtime speech synthesis |
2026-07-14 |
International |
|
qwen-audio-3.0-tts-flash is a high-performance speech synthesis model, optimized for real-time interactive scenarios. Compared with the previous version, it supports more low-resource languages and Chinese dialects, improves dialect authenticity, and enhances free-style instruction following and fine-grained tag control for more flexible control over emotion, tone, character, speaking rate, volume, and expressive style. It is also more robust under noisy and reverberant acoustic conditions, with improved audio quality, clarity, and overall expressiveness. The Flash version focuses on real-time synthesis, making it suitable for voice assistants, real-time dialogue, intelligent customer service, and other low-latency interactive applications. |
|
Text generation |
2026-07-10 |
International |
|
GLM-5.2-Fast-Preview is the high-speed variant of Zhipu AI's GLM-5.2, with 1M context and capabilities on par with the standard version. Inference-optimized to deliver 1.5–2× the output TPS, it fits latency-sensitive use cases such as real-time chat, multi-turn agents, and streaming code generation. |
|
Video generation |
2026-07-01 |
International |
|
Wan2.7 text to video, performance fully reimagined. Delivering nuanced and organic emotional depth in narrative arcs and visceral, bone-crunching impact in action sequences. Enhanced by rhythmic cinematic cuts for unparalleled storytelling power.This version is a snapshot as of June 12, 2026. |
|
Video generation |
2026-07-01 |
International |
|
Wan2.7 reference to video, enhanced consistency & performance. Delivering superior stability for characters, props, and scenes. Supports hybrid referencing of up to 5 mixed image/video inputs and audio timbre cloning. Together with core engine upgrades, it achieves unprecedented cinematic expressive power. This version is a snapshot as of June 12, 2026. |
|
Image generation |
2026-06-25 |
International |
|
The Qwen-Image-2.0 series full-fledged model integrates image generation and editing; it boasts more professional text rendering capabilities with 1k token command support, more delicate and realistic textures, meticulous depiction of realistic scenes, and stronger semantic adherence. The full-fledged version possesses the strongest text rendering capabilities and realistic textures in the 2.0 series. |
|
Text generation, Reasoning, Visual understanding |
2026-06-25 |
International |
|
kimi-k2.7-code is Kimi's most intelligent coding model to date. It follows instructions more reliably over long contexts and completes programming tasks with higher success rates. It supports text, image, and video inputs, along with thinking mode, conversation, and agent tasks. |
|
Text generation, Reasoning |
2026-06-25 |
International |
|
GLM-5.2 is the latest flagship model from Zhipu AI, designed for long-horizon tasks with support for an ultra-long 1M context window. It features powerful logical reasoning, long-text comprehension, and code generation capabilities, balancing performance with inference efficiency. It excels across multi-task benchmarks and is well-suited for intelligent interaction, enterprise applications, and development assistance scenarios. |
|
Speech recognition |
2026-06-17 |
International |
|
The Bailing ASR version, updated in June 2026, fully supports the seven major dialect systems of traditional Chinese (Mandarin/Wu/Xiang/Gan/Hakka/Min/Yue) and is compatible with over 20 regional Mandarin accents. It features specific optimizations for the rhythm, meter, and literary expression characteristics of classical Chinese poetry, improving the accuracy of poetry recognition and making it suitable for cultural heritage, educational explanations, and audiobooks. Improved punctuation prediction and text normalization capabilities make the output text more consistent with written expression habits; numbers, dates, and amounts are automatically converted to standard formats, enhancing readability and professionalism. The language support has been expanded to include English, Japanese, Korean, Vietnamese, Thai, Indonesian, Malay, Filipino, Hindi, Arabic, French, German, Spanish, Portuguese, Russian, Italian, Dutch, Swedish, Danish, Finnish, Norwegian, Greek, Polish, Czech, Hungarian, Romanian, Bulgarian, Croatian, and Slovak, totaling 30 languages. It supports contextualization capabilities and can transcribe audio up to 5 minutes in length. |
|
Video generation |
2026-06-16 |
International |
|
HappyHorse-1.1-T2V supports Text-to-Video generation with improved semantic understanding, cinematic shot control, and dynamic motion rendering. It more accurately captures creative intent, producing high-quality videos with smoother motion, richer details, stronger visual consistency, and more natural character actions, scene atmosphere, and physical dynamics. |
|
Video generation |
2026-06-16 |
International |
|
HappyHorse-1.1-R2V supports Reference-to-Video generation with significantly improved stability in subject, scene style, and visual consistency. With support for up to 9 reference images, it can more accurately understand and preserve creative intent, delivering stronger controllability and expressiveness across characters, scenes, styles, and cinematic motion. |
|
Video generation |
2026-06-16 |
International |
|
HappyHorse-1.1-I2V supports Image-to-Video generation with improved visual quality, dynamic performance, and cross-clip consistency. It more accurately understands the input image and preserves creative intent, delivering significant improvements in skin texture realism, character ID consistency across clips, motion smoothness, text rendering stability, and audio-visual synchronization, producing high-quality videos with greater realism, richer details, and stronger overall consistency. |
|
Text generation, Reasoning, Visual understanding |
2026-06-10 |
International |
|
The Max model, the largest and most capable in the Qwen3.7 series, has added visual‑modal understanding compared to the May 20 snapshot, enabling it to perceive real‑world scenes and supporting multimodal interactive hybrid agent capabilities. This version is based on a snapshot taken on June 8, 2026. |
|
Text generation, Reasoning, Visual understanding |
2026-06-01 |
International |
|
Among the Qwen3.7 series, the cost-effective Plus model builds on its robust text capabilities while delivering a comprehensive upgrade to its vision‑language abilities, all while preserving its full‑stack agent‑level intelligence for coding, tool use, and productivity workflows. Its key distinguishing feature is multi‑modal interactive hybrid agent capabilities, enabling it to perceive real‑world scenes, read screens and interact with GUIs, generate code based on visual references, and perform end‑to‑end navigation within mobile apps. |
|
Text generation, Reasoning |
2026-05-21 |
International |
|
The Max model, the largest and most capable in the Qwen3.7 series, currently offers a pure‑text‑only interface for public experimentation. Qwen3.7 is a next‑generation flagship model designed for the agent‑centric era, with its core strengths lying in the breadth and depth of its agent‑level capabilities: it excels at programming, office and productivity tasks, and long‑term autonomous execution. |
|
Realtime speech translation |
2026-05-19 |
International |
|
The real-time version of Qwen3.5-LiveTranslate-Flash, which is a high-precision, highly responsive, and robust multilingual simultaneous audio and video interpretation model. Leveraging Qwen3.5-Omni's powerful infrastructure, massive multimodal data, cross-language and cross-modal alignment, and visual enhancement technologies, Qwen3.5-LiveTranslate-Flash offers both offline and real-time audio and video translation capabilities. It can understand 60 languages and speak 29 languages. |
|
Text generation, Reasoning |
2026-05-11 |
International |
|
A flagship MoE large model with 1.6 trillion parameters and 49 billion activated parameters, natively supporting context lengths of up to one million tokens. Trained on a vast corpus of high-quality data, it excels in advanced mathematical reasoning, complex logical inference, specialized coding, and deep analysis of long-form text, making it well-suited for demanding applications such as cutting-edge research, sophisticated office workflows, and advanced AI agents. |
|
Text generation, Reasoning |
2026-05-11 |
International |
|
A highly efficient, lightweight MoE model with 284 billion parameters in total and 13 billion activated parameters, natively supporting context windows of up to one million tokens. It offers fast inference speed, low latency, and cost-effective invocation, delivering well-balanced overall performance. Designed for high-concurrency, lightweight workloads, it is ideally suited for common, essential use cases such as everyday dialogue, content creation, basic RAG applications, and batch text processing. |
|
Video generation |
2026-04-26 |
International |
|
Wan2.7 text to video, performance fully reimagined. Delivering nuanced and organic emotional depth in narrative arcs and visceral, bone-crunching impact in action sequences. Enhanced by rhythmic cinematic cuts for unparalleled storytelling power.This version is a snapshot as of April 25, 2026. |
|
Video generation |
2026-04-26 |
International |
|
Wan2.7 image to video, performance fully reimagined. Delivering nuanced and organic emotional depth in narrative arcs and visceral, bone-crunching impact in action sequences. Enhanced by rhythmic cinematic cuts for unparalleled storytelling power.This version is a snapshot as of April 25, 2026. |
|
Video generation |
2026-04-26 |
International |
|
HappyHorse-1.0-V2V supports advanced video editing through natural language instructions. It allows for local or global editing of video elements using up to 5 reference images, precisely preserving original motion dynamics to achieve superior expressiveness. |
|
Video generation |
2026-04-26 |
International |
|
HappyHorse-1.0-R2V supports Reference-to-Video generation, offering enhanced stability in subject and scene referencing. Capable of processing up to 9 reference images, it precisely preserves creative intent to deliver superior performance. |
|
Text generation, Reasoning, Visual understanding |
2026-04-23 |
International |
|
The Qwen3.5 native vision-language series Plus model has seen a substantial improvement in agentic coding capabilities compared to the February 15th snapshot. Inference speed has also been significantly enhanced, while its knowledge retention, reasoning ability, and long-context processing remain at a high level, making it well-suited for complex agent-based tasks. It is ideal for applications such as coding agents, production workflows, and high-throughput scenarios. This version is based on a snapshot taken on April 20, 2026. |
|
Image generation |
2026-04-23 |
International |
|
The full-featured Qwen-Image-2.0 series models integrate image generation and image editing, offering enhanced text rendering with support for 1,000-token prompts, more refined realistic textures, detailed depiction of photorealistic scenes, and stronger semantic adherence. The full-featured version delivers the strongest text rendering and most lifelike textures in the 2.0 series. |
|
Text generation, Reasoning, Visual understanding |
2026-04-22 |
International |
|
The Qwen3.6 27B native vision-language dense model builds upon the 3.5-27B version, with key improvements in agentic coding capabilities and enhanced STEM reasoning and inference skills. In the vision modality, it demonstrates significant advances in spatial intelligence, object localization, and detection, while video understanding, document OCR, and visual agent capabilities continue to improve steadily. |
|
Video generation |
2026-04-22 |
International |
|
HappyHorse-1.0-T2V supports text-to-video generation, featuring highly realistic dynamic rendering. It accurately comprehends text semantics to produce high-quality videos that are fluid, natural, and rich in detail. |
|
Video generation |
2026-04-22 |
International |
|
HappyHorse-1.0-I2V enables image-to-video generation, featuring highly realistic dynamic rendering. It accurately comprehends both text and image semantics to produce high-quality videos that are fluid, natural, and rich in detail. |
|
Text generation, Reasoning, Visual understanding |
2026-04-17 |
International |
|
The Qwen3.6 native vision-language Flash model series delivers a significant performance boost over the 3.5-Flash version. This model particularly excels in agentic coding capabilities, substantially outperforming its predecessor on multiple code-agent benchmarks, as well as in mathematical and code reasoning. In terms of vision, it features markedly improved spatial intelligence, with especially notable enhancements in object localization and object detection. |
|
Text generation, Reasoning, Visual understanding |
2026-04-17 |
International |
|
The Qwen3.6 35B-A3B native vision-language model is built on a hybrid architecture that integrates linear attention mechanisms with a sparse mixture-of-experts framework, achieving higher inference efficiency. Compared with the 3.5-35B-A3B, this model demonstrates significantly improved agentic coding capabilities, mathematical and code reasoning abilities, spatial intelligence, as well as object localization and object detection performance. |
|
Speech synthesis |
2026-04-15 |
International |
|
A large-model voice replication service used in conjunction with Cosyvoice-v3. Utilizing advanced large-model technology for feature extraction, it can replicate voices without a training process. Only a very short audio clip is required to quickly generate a highly similar and natural-sounding custom voice. |
|
Text generation, Reasoning |
2026-04-14 |
International |
|
The Max model, the largest and most capable variant in the Qwen3.6 series, is now available in a preview version. At present, only its plain-text capabilities are open for experimentation. Compared with the previously released Qwen3-Max and Qwen3.6-Plus, this model features enhanced vibe coding abilities, more efficient coding agent execution, and significantly improved front-end development skills. Additionally, its long-tail knowledge retention has been further upgraded. |
|
Text generation, Reasoning |
2026-04-14 |
International |
|
GLM-5.1 is a model developed by Zhipu AI, specifically designed for long-horizon tasks. It has 744 billion parameters, supports an ultra-long context of 200k tokens, and can generate up to 128k tokens in a single response. GLM-5.1 excels in logical reasoning, long-text understanding, and code generation, while balancing performance with inference efficiency. It delivers outstanding results across multiple multi-task benchmarks and is well-suited for applications such as intelligent human-computer interaction, enterprise solutions, and developer assistance. |
|
Video generation |
2026-04-03 |
International |
|
Wan2.7 video edit, supports both localized and global editing with prompt. Seamlessly replace elements using image references and replicate complex dynamic processes, including motion, special effects, and camera movements. |
|
Video generation |
2026-04-03 |
International |
|
Wan2.7 text to video, performance fully reimagined. Delivering nuanced and organic emotional depth in narrative arcs and visceral, bone-crunching impact in action sequences. Enhanced by rhythmic cinematic cuts for unparalleled storytelling power. |
|
Video generation |
2026-04-03 |
International |
|
Wan2.7 reference to video, enhanced consistency & performance. Delivering superior stability for characters, props, and scenes. Supports hybrid referencing of up to 5 mixed image/video inputs and audio timbre cloning. Together with core engine upgrades, it achieves unprecedented cinematic expressive power. |
|
Video generation |
2026-04-03 |
International |
|
Wan2.7 image to video, performance fully reimagined. Delivering nuanced and organic emotional depth in narrative arcs and visceral, bone-crunching impact in action sequences. Enhanced by rhythmic cinematic cuts for unparalleled storytelling power. |
|
Image generation |
2026-04-01 |
International |
|
Wan2.7–image-pro, supports text to image, text/image to sequential images, image editing, multi-image reference generation, and interactive editing. Delivers enhanced performance in text rendering, subject consistency, and complex instruction following. |
|
Image generation |
2026-04-01 |
International |
|
Wan2.7 – image generation and editing, supports text to image, text/image to sequential images, image editing, multi-image reference generation, and interactive editing. Delivers enhanced performance in text rendering, subject consistency, and complex instruction following. |
|
Text generation, Reasoning, Visual understanding |
2026-04-01 |
International |
|
The Qwen3.6 native vision-language Plus series models demonstrate exceptional performance on par with the current state-of-the-art models, with a significant improvement in overall results compared to the 3.5 series. The models have been markedly enhanced in code-related capabilities such as agentic coding, front-end programming, and Vibe coding, as well as in multi-modal general object recognition, OCR, and object localization. |
|
Realtime omni-modal |
2026-03-26 |
International |
|
Qwen 3.5-Omni is the latest generation of Qwen's multimodal large model, supporting text, image, audio, and audio-visual understanding and interaction. As a fully evolved version of Qwen3-Omni, it supports audio input in 60+ languages, voice output in 30+ languages, and controllable voice dialogue, WebSearch and complex FunctionCall invocation, and has intelligent semantic interruption interaction capabilities. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interactive experience. |
|
Omni-modal |
2026-03-26 |
International |
|
Qwen 3.5-Omni is the latest generation of Qwen's multimodal large model, supporting text, image, audio, and audio-visual understanding and interaction. As a comprehensive evolution of Qwen 3-Omni, it supports over 10 hours of audio understanding and over 400 seconds of 720P (1 FPS) audio-visual understanding and dialogue. It further expands the language range, supporting audio input in 60+ languages and speech output in 30+ languages. It also possesses powerful structured audio-visual understanding capabilities and is widely used in text creation, voice assistants, multimedia analysis, and other scenarios, providing a natural and fluent multimodal understanding and interactive experience. |
|
Realtime omni-modal |
2026-03-26 |
International |
|
Qwen 3.5-Omni is the latest generation of Qwen's multimodal large model, supporting text, image, audio, and audio-visual understanding and interaction. As a fully evolved version of Qwen3-Omni, it supports audio input in 60+ languages, voice output in 30+ languages, and controllable voice dialogue, WebSearch and complex FunctionCall invocation, and has intelligent semantic interruption interaction capabilities. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interactive experience. |
|
Omni-modal |
2026-03-26 |
International |
|
Qwen 3.5-Omni is the latest generation of Qwen's multimodal large model, supporting text, image, audio, and audio-visual understanding and interaction. As a comprehensive evolution of Qwen 3-Omni, it supports over 10 hours of audio understanding and over 400 seconds of 720P (1 FPS) audio-visual understanding and dialogue. It further expands the language range, supporting audio input in 60+ languages and speech output in 30+ languages. It also possesses powerful structured audio-visual understanding capabilities and is widely used in text creation, voice assistants, multimedia analysis, and other scenarios, providing a natural and fluent multimodal understanding and interactive experience. |
|
Realtime omni-modal |
2026-03-25 |
International |
|
Qwen 3.5-Omni is the latest generation of Qwen's multimodal large model, supporting text, image, audio, and audio-visual understanding and interaction. As a fully evolved version of Qwen3-Omni, it supports audio input in 60+ languages, voice output in 30+ languages, and controllable voice dialogue, WebSearch and complex FunctionCall invocation, and has intelligent semantic interruption interaction capabilities. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interactive experience.This version is a snapshot from March 15, 2026. |
|
Text generation, Reasoning |
2026-03-20 |
International |
|
DeepSeek-V3.2 is the official release of a model that incorporates DeepSeek Sparse Attention—a sparse attention mechanism. It's also the first model launched by DeepSeek that integrates reasoning into tool usage, supporting both reasoning-enabled and non-reasoning tool calls. |
|
Image generation |
2026-03-03 |
International |
|
The full-featured Qwen-Image-2.0 series models integrate image generation and image editing, offering enhanced text rendering with support for 1,000-token prompts, more refined realistic textures, detailed depiction of photorealistic scenes, and stronger semantic adherence. The full-featured version delivers the strongest text rendering and most lifelike textures in the 2.0 series.This version is a snapshot as of March 3, 2026. |
|
Image generation |
2026-03-03 |
International |
|
The Qwen-Image-2.0 series of accelerated models integrates image generation and image editing, offering enhanced text-rendering capabilities with support for 1,000-token prompts, more realistic textures, finely detailed photorealistic scenes, and improved semantic consistency. The accelerated version effectively strikes an optimal balance between model performance and quality. |
|
Speech recognition |
2026-03-02 |
International |
|
Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves highly accurate speech recognition, automatically determining the language and accurately identifying speech in multiple languages, while ensuring precise transcription even in complex audio environments.This version is a snapshot dated February 10, 2026. |
|
Text generation, Reasoning, Visual understanding |
2026-02-23 |
International |
|
The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the 3 series, these models deliver a leap forward in performance for both pure text and multimodal tasks, offering fast response times while balancing inference speed and overall performance. |
|
Text generation, Reasoning, Visual understanding |
2026-02-23 |
International |
|
The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall performance is comparable to that of the Qwen3.5-27B. |
|
Text generation, Reasoning, Visual understanding |
2026-02-23 |
International |
|
The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of the Qwen3.5-122B-A10B. |
|
Text generation, Reasoning, Visual understanding |
2026-02-23 |
International |
|
The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of overall performance, this model is second only to Qwen3.5-397B-A17B. Its text capabilities significantly outperform those of Qwen3-235B-2507, and its visual capabilities surpass those of Qwen3-VL-235B. |
|
Text generation |
2026-02-20 |
International |
|
The new-generation code generation model in the Qwen3 series delivers performance close to that of Qwen3-Coder-Plus while offering even better capabilities. The model has been optimized with a focus on repository-level understanding, supports multi-turn tool interactions, and enhances its compatibility with agentic coding tools. |
|
Text generation, Reasoning, Visual understanding |
2026-02-15 |
International |
|
The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of task evaluations, the 3.5 series consistently demonstrates performance on par with state-of-the-art leading models. Compared to the 3 series, these models show a leap forward in both pure-text and multimodal capabilities. |
|
Text generation, Reasoning, Visual understanding |
2026-02-15 |
International |
|
The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers state-of-the-art performance comparable to leading-edge models across a wide range of tasks, including language understanding, logical reasoning, code generation, agent-based tasks, image understanding, video understanding, and graphical user interface (GUI) interactions. With its robust code-generation and agent capabilities, the model exhibits strong generalization across diverse agent. |
|
Speech recognition |
2026-02-13 |
International |
|
The real-time version of Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves highly accurate speech recognition, automatically determining the language and accurately identifying speech in multiple languages, while ensuring precise transcription even in complex audio environments.This version is a snapshot dated February 10, 2026. |
|
Speech synthesis |
2026-02-10 |
International |
|
Qwen3-TTS-VD model is Tongyi's latest real-time speech synthesis model. It can perform high-fidelity real-time speech synthesis on the voices designed by the qwen3-voice-design service, and supports speech output in 11 languages using the same voice. This model has been trained on massive amounts of data, and the synthesized audio can adaptively adjust its tone according to the text, demonstrating good processing capabilities for complex text synthesis. This model is a snapshot version from January 26, 2026. |
|
Speech synthesis |
2026-02-10 |
International |
|
Qwen3-TTS-Flash model is Tongyi's latest real-time speech synthesis model. It can perform high-fidelity real-time speech synthesis on voices replicated from the qwen-voice-enrollment service, and supports speech output in 11 languages using the same voice timbre. This model has been trained on massive amounts of data, and the synthesized audio can adaptively adjust tone according to the text, demonstrating good processing capabilities for complex text synthesis. This model is a snapshot version from January 22, 2026. |
|
Speech synthesis |
2026-02-10 |
International |
|
Qwen3-TTS-Flash model is Tongyi's latest real-time speech synthesis model. The Instruct model processes the synthesis effect through natural language, ensuring highly appropriate emotional and expressive speech in different contexts. Currently, it supports 25 timbres for both Chinese and English Instruct adjustments. |
|
Speech synthesis |
2026-02-10 |
International |
|
Cloning capability: CosyVoice-v3-plus is the latest large voice cloning model in the CosyVoice series from Tongyi Lab. It offers superior sound quality and cloning fidelity, ideal for professional scenarios. With just 5-20 seconds of reference audio, it can rapidly generate a highly similar and natural-sounding custom voice. Synthesis capability: CosyVoice-v3-plus is the latest large speech synthesis model in the CosyVoice series from Tongyi Lab. It features enhanced sound quality and expressiveness, ideal for professional scenarios. The model supports real-time, streaming text-to-speech synthesis. |
|
Speech synthesis |
2026-02-09 |
International |
|
Synthesis Capabilities: CosyVoice-v3-Flash is the latest high-performance speech synthesis model in the CosyVoice series from Tongyi Labs, offering improved naturalness, timbre, prosody, and emotional expressiveness compared to previous versions. This model supports real-time streaming text-to-speech synthesis. Cloning Capabilities: CosyVoice-v3-Flash is also the latest speech cloning model in the CosyVoice series from Tongyi Labs. Compared to previous versions, it improves pronunciation accuracy and timbre similarity, and adds support for more less commonly spoken languages (German, Spanish, French, Italian, Russian, Japanese). It can quickly generate highly similar and naturally sounding custom voices from just 5-20 seconds of reference audio. |
|
Text generation |
2026-01-30 |
International |
|
The role-playing model of the Qwen series. This is a dynamically updated version, and notifications will be provided in advance for any model updates. It is suitable for anthropomorphic role-playing and has optimized capabilities in following predefined character instructions, advancing conversations, and demonstrating active listening and empathy. Additionally, it supports the deep restoration of personalized characters. |
|
Video generation |
2026-01-29 |
International |
|
Wan2.6 reference to video flash, faster and more cost-effective generation. Supports using a specified person or any object as a reference, precisely maintaining consistency of appearance and voice, and allows multi‑character reference for joint performances. |
|
Text embedding |
2026-01-27 |
International |
|
A text-ranking model trained on the Qwen LLM foundation performs relevance ranking for input queries and candidate documents. It supports over 100 languages and long-text inputs, and is suitable for applications such as text retrieval and RAG. Its performance is aligned with the open-source Qwen3-Rerank series models. |
|
Text generation, Reasoning |
2026-01-23 |
International |
|
Compared with the snapshot as of September 23, 2025, the Qwen-3 series Max model in this release achieves an effective integration of thinking and non-thinking modes, resulting in a comprehensive and substantial improvement in the model's overall performance. In thinking mode, the model simultaneously supports web search, web information extraction, and a code interpreter tool, enabling it to tackle more complex and challenging problems with greater accuracy by leveraging external tools while engaging in slow, deliberative reasoning. This version is based on a snapshot taken on January 23, 2026. |
|
Visual understanding |
2026-01-22 |
International |
|
The Qwen3 series of small-sized visual understanding models effectively integrates thinking and non-thinking modes. Compared with the snapshot taken on October 15, 2025, the overall performance of the model has improved significantly: it delivers enhanced capabilities in general visual recognition and reasoning, and shows marked improvements in recognition accuracy across various business scenarios such as security, in-store inspections, equipment monitoring, and photo-based problem solving. This version is a snapshot as of January 22, 2026. |
|
Speech synthesis |
2026-01-21 |
International |
|
qwen3-TTS-Flash model is Tongyi's latest real-time speech synthesis model. The Instruct model processes the synthesis effect through natural language, ensuring highly appropriate emotional and expressive speech in different contexts. Currently, it supports 25 timbres for both Chinese and English Instruct adjustments. This model is equivalent to the snapshot version released on January 22, 2026. |
|
Video generation |
2026-01-15 |
International |
|
Wan2.6 image to video flash, faster and more cost-effective generation. Intelligent shot scheduling enables multi‑camera storytelling, supports stable multi‑speaker dialogue with more natural and realistic vocal timbres, and supports generation of clips up to 15 seconds in length. |
|
Image generation |
2026-01-15 |
International |
|
The Max series Qwen's image editing models delivers more stable and versatile editing capabilities: enhanced industrial design and geometric reasoning, improved character consistency, reduced offset issues, and integrated LoRA capabilities for a wider range of image editing functions. |
|
Speech synthesis |
2026-01-14 |
International |
|
Qwen3-TTS-VD model is Tongyi's latest real-time speech synthesis model. It can perform high-fidelity real-time speech synthesis on the voices designed by the qwen3-voice-design service, and supports speech output in 11 languages using the same voice. This model has been trained on massive amounts of data, and the synthesized audio can adaptively adjust its tone according to the text, demonstrating good processing capabilities for complex text synthesis. This model is a snapshot version from January 15, 2026. |
|
Speech synthesis |
2026-01-14 |
International |
|
Qwen 3-TTS-Flash model is Tongyi's latest real-time speech synthesis model. It can perform high-fidelity real-time speech synthesis on voices replicated from the qwen-voice-enrollment service, and supports speech output in 11 languages using the same voice. This model has been trained on massive amounts of data, and the synthesized audio can adaptively adjust tone according to the text, demonstrating good processing capabilities for complex text synthesis. This model is a snapshot version from January 15, 2026. |
|
Text generation |
2026-01-13 |
International |
|
The Qwen Role-Playing Model Series is specifically optimized for muti-language anthropomorphic interaction scenarios. It demonstrates advanced capabilities in character consistency maintenance, context-aware dialogue progression, and empathetic engagement, enabling precise personalized character embodiment. This version significantly enhances Japanese linguistic localization (including dialects and honorifics), human-like role-playing authenticity, narrative coherence control, and scenario-based cognitive intelligence. |
|
Image generation |
2026-01-09 |
International |
|
The Qwen series of image-generation models boasts exceptional text-rendering capabilities and excels in complex text rendering as well as a wide range of generation and editing tasks. This version, a snapshot taken on January 9, 2026, is a distilled and accelerated variant of Qwen-Image-Max, enabling faster generation of high-quality images. |
|
Image generation |
2025-12-30 |
International |
|
The Max series of qwen's image generation model excels across a wide range of generation tasks. Compared with the Plus series, it significantly reduces the "AI-like" feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering. |
|
Image generation |
2025-12-22 |
International |
|
Z-Image-Turbo is a highly efficient image-generation model that has topped the Artificial Analysis benchmark as the world's No. 1 open-source text-to-image model. With just 6 billion parameters and an 8-step inference process, it generates photo-realistic images comparable to those produced by large-scale commercial models, while excelling in bilingual Chinese–English text rendering, complex semantic understanding, and diverse thematic generation. |
|
Reasoning, Visual understanding |
2025-12-18 |
International |
|
The Qwen3 series of visual understanding models effectively integrates thinking and non-thinking modes. Compared to the snapshot released on September 23, this version delivers superior performance in reasoning and analysis tasks as well as style control, while also offering lower latency and faster response speeds. This version is based on a snapshot taken on December 19, 2025. |
|
Video generation |
2025-12-16 |
International |
|
Wan2.6 reference to video, supports using a specified person or any object as a reference, precisely maintaining consistency of appearance and voice, and allows multi‑character reference for joint performances. |
|
Image generation |
2025-12-15 |
International |
|
Wan2.6 text to image, Upgraded visual quality, aesthetics, and instruction-following deliver precise style control, realistic portraits, long-text understanding, and broad historical/cultural IP coverage, enabling high-quality, highly expressive visual generation. |
|
Image generation |
2025-12-15 |
International |
|
Wan2.6 Image, An all-round image generation model that supports joint text–image reasoning, multi-image creative fusion, commercial-grade consistency, aesthetic style transfer, and precise control of framing and lighting, significantly enhancing consistency, controllability, and expressiveness in image generation. |
|
Image generation |
2025-12-15 |
International |
|
The Qianwen series of Image Editing Plus models features enhanced character consistency, industrial design capabilities, and geometric reasoning abilities compared to the snapshot as of October 30. Additionally, it integrates LoRA capabilities such as lighting effects and effectively mitigates offset issues. This version is based on a snapshot taken on December 15, 2025. |
|
Speech synthesis |
2025-12-12 |
International |
|
Qwen 3-TTS-VD model is Tongyi's latest real-time speech synthesis model. It can perform high-fidelity real-time speech synthesis on the voices designed by the qwen3-voice-design service, and supports speech output in 11 languages using the same voice. This model has been trained on massive amounts of data, and the synthesized audio can adaptively adjust its tone according to the text, demonstrating good processing capabilities for complex text synthesis. This model is a snapshot version from December 16, 2025. |
|
Speech synthesis |
2025-12-12 |
International |
|
Qwen Voice-Design model is a series of voice design models from Qwen Speech Model. It only requires a simple text description to quickly design a suitable voice. When used in conjunction with the qwen3-tts-vd-realtime model, it can design and output speech in 10 languages. Furthermore, the synthesized audio can adaptively adjust its tone based on the text and has good processing capabilities for complex text synthesis. |
|
Realtime omni-modal |
2025-12-04 |
International |
|
The real-time version of the Qwen3-Omni-Flash multimodal large-scale model, based on the Thinker–Talker Mixed Expert (MoE) architecture, supports efficient understanding and speech generation of text, images, audio, and video. It can interact with text in 119 languages and speech in 20 languages, generating human-like speech for precise cross-lingual communication. The model boasts powerful command-following and system prompt customization capabilities, flexibly adapting to conversational styles and character settings. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interaction experience. |
|
Omni-modal |
2025-12-04 |
International |
|
Qwen3-Omni-Flash multimodal large-scale model, based on the Thinker–Talker Mixed Expert (MoE) architecture, supports efficient understanding and speech generation of text, images, audio, and video. It can interact with text in 119 languages and speech in 20 languages, generating human-like speech for precise cross-lingual communication. The model boasts powerful command-following and system prompt customization capabilities, flexibly adapting to conversational styles and character settings. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interaction experience. |
|
Realtime speech recognition |
2025-12-04 |
International |
|
Qwen3-LiveTranslate-Flash is a high-precision, highly responsive, and robust multilingual real-time audio and video interpretation model. Leveraging Qwen3-Omni's powerful infrastructure, massive multimodal data, cross-language and cross-modal alignment, and visual enhancement technologies, Qwen3-LiveTranslate-Flash provides both offline and real-time audio and video translation capabilities. It can understand 19 languages and speak 10 languages, and also supports 8 Chinese dialects. |
|
Video generation |
2025-12-03 |
International |
|
Wan2.6 text to video, Intelligent shot scheduling supports multi-shot storytelling, generating multi-shot narrative videos with consistent subjects, scenes, and atmosphere, with a maximum duration of 15 seconds. It delivers better instruction following, and improved visual fidelity. |
|
Video generation |
2025-12-03 |
International |
|
Wan2.6 image to video, intelligent shot scheduling enables multi‑camera storytelling, delivers higher‑quality voice generation, supports stable multi‑speaker dialogue with more natural and realistic vocal timbres, and supports generation of clips up to 15 seconds in length. |
|
— |
2025-12-01 |
International |
|
This version is a snapshot as of December 1, 2025, and features improved reasoning capabilities compared to the July 28 snapshot. Agent capabilities and multi-turn tool invocation abilities have been further enhanced, and performance on subjective creative tasks has improved significantly. It supports a context length of up to 1 million tokens, with tiered pricing based on context length. |
|
Speech synthesis |
2025-11-27 |
International |
|
Qwen3-TTS-Flash model is Tongyi's latest real-time speech synthesis model. It can perform high-fidelity real-time speech synthesis on voices replicated by the qwen3-voice-enrollment service, and supports speech output in 11 languages with the same voice timbre. This model has been trained on massive amounts of data, and the synthesized audio can adaptively adjust the tone according to the text, and it also has good processing capabilities for complex text synthesis.This model is provided as a snapshot version. |
|
Speech synthesis |
2025-11-27 |
International |
|
The Qwen3-TTS-Flash-Realtime model is Tongyi's latest real-time speech synthesis foundation model, featuring 17 expressive voices while delivering low-latency, high-stability audio synthesis. It supports multilingual and dialect outputs with consistent voice characteristics across languages. Trained on extensive datasets, the system autonomously adjusts vocal tones based on text semantics and demonstrates robust capabilities for complex content synthesis. |
|
Speech synthesis |
2025-11-27 |
International |
|
The Qwen3-TTS-Flash is Tongyi's latest offline text-to-speech foundation model, featuring 17 expressive voices while enabling low-latency, high-stability audio synthesis. It supports multilingual and dialect outputs with consistent voice characteristics across languages. Trained on massive datasets, the system automatically adjusts vocal tones based on text semantics and demonstrates robust capabilities for synthesizing complex content. |
|
Speech synthesis |
2025-11-27 |
International |
|
The Qwen Voice-Enrollment model is a series of voice replication models from the qwen speech model. It can quickly replicate highly similar voices using audio of only 5 seconds or more. When used in conjunction with the qwen3-tts-vc-realtime model, it can replicate a person's voice with high fidelity and output speech in 10 languages. Furthermore, the synthesized audio can adaptively adjust its tone according to the text and has good processing capabilities for complex text synthesis. |
|
Visual understanding |
2025-11-21 |
International |
|
This model is a snapshot version from November 20, 2025, and is based on the latest Qwen-VL3 architecture with a comprehensive upgrade. It features significant improvements in document parsing and text localization capabilities, as well as substantial reductions in end-to-end latency and illusions. |
|
Speech recognition |
2025-11-21 |
International |
|
The Fun ASR version, updated in April 2026, fully supports the seven major dialect systems of traditional Chinese (Mandarin/Wu/Xiang/Gan/Hakka/Min/Yue) and is compatible with over 20 regional Mandarin accents. It features specific optimizations for the rhythm, meter, and literary expression characteristics of classical Chinese poetry, improving the accuracy of poetry recognition and making it suitable for cultural heritage, educational explanations, and audiobooks. Improved punctuation prediction and text normalization capabilities make the output text more consistent with written expression habits; numbers, dates, and amounts are automatically converted to standard formats, enhancing readability and professionalism. The language support has been expanded to include English, Japanese, Korean, Vietnamese, Thai, Indonesian, Malay, Filipino, Hindi, Arabic, French, German, Spanish, Portuguese, Russian, Italian, Dutch, Swedish, Danish, Finnish, Norwegian, Greek, Polish, Czech, Hungarian, Romanian, Bulgarian, Croatian, and Slovak, totaling 30 languages. This is a snapshot released on November 7, 2025. |
|
Visual understanding |
2025-11-20 |
International |
|
Qwen-VL_OCR is an OCR model trained based on Qwen-VL. It aggregates various image-text recognition, parsing, and processing tasks through a unified model approach, offering powerful image-text recognition capabilities. |
|
Text generation |
2025-11-19 |
International |
|
Qwen-MT-Lite is a large language model of the Qwen model series that specializes in multi-lingual translation. It provides high-quality and rapid translation services across 32 languages at a cost-effective price. It offers features such as terminology intervention, format preservation, and domain-specific translation to cater to the diverse needs of various applications, ensuring both efficiency and performance. |
|
Speech recognition |
2025-11-18 |
International |
|
The large file transcription version of Qwen3-ASR-Flash. Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves high-precision speech recognition. It can automatically determine the language and accurately recognize speech in multiple languages, ensuring precise transcription even in complex audio environments. |
|
Realtime speech recognition |
2025-11-17 |
International |
|
This is the real-time version of Tongyi Lab's next-generation end-to-end speech recognition model, based on leading proprietary speech technology, and boasts exceptional contextual awareness and high-precision speech transcription capabilities. Based on an end-to-end architecture, Fun-ASR integrates innovative RAG technology, supporting multi-dimensional features such as large-scale hotword customization, automatic filtering of sensitive and modal particles, ITN normalization, and punctuation prediction, significantly improving overall recognition accuracy and contextual relevance. Furthermore, Fun-ASR supports flexible switching between Chinese and English, covers multiple regional dialects, and boasts enhanced noise robustness, adapting to diverse and complex environments.This is a snapshot released on November 7, 2025. |
|
Text generation |
2025-11-11 |
International |
|
Qwen-MT-Flash, a large language model from the Qwen series, has been fully upgraded with the Qwen 3 architecture for significantly enhanced performance and translation quality. It provides rapid, cost-effective translation across 92 languages, while supporting advanced features such as terminology intervention, format preservation, and domain-specific adaptation. It is the ideal choice for applications requiring a powerful balance of speed, quality, and cost. |
|
Video generation |
2025-11-10 |
International |
|
wan2.2-animate-move is a character animation generation model. Users simply upload a character photo and a reference performance video, and the model transfers the expressions and actions from the video onto the character in the image, producing a high-fidelity animated video. |
|
Video generation |
2025-11-10 |
International |
|
wan2.2-animate-mix is a character replacement model product. By uploading a character photo and a performance video, users can accurately replace the character in the original video with the character from the photo, while completely preserving environmental details such as the scene, lighting, and color tone of the original video. |
|
Multimodal embedding |
2025-10-31 |
International |
|
Embedding-Vision is a vision-centric multimodal embedding model powered by an LLM, featuring outstanding domain-specific performance and high cost-effectiveness in various domains (e.g., e-commerce, photo galleries, security, autonomous driving). With support for text, image, and video, it is applicable to downstream retrieval tasks, including text-to-image, image-to-image, text-to-video and video-to-video. |
|
Multimodal embedding |
2025-10-31 |
International |
|
Embedding-Vision is a vision-centric multimodal embedding model powered by an LLM, featuring outstanding domain-specific performance and high cost-effectiveness in various domains (e.g., e-commerce, photo galleries, security, autonomous driving). With support for text, image, and video, it is applicable to downstream retrieval tasks, including text-to-image, image-to-image, text-to-video and video-to-video. |
|
Image generation |
2025-10-31 |
International |
|
The qwen series of image editing Plus models further optimizes inference performance and system stability based on the initial Edit model, significantly reducing the response time for image generation and editing. It also supports returning multiple images in a single request, greatly enhancing user experience. |
|
Realtime speech recognition |
2025-10-29 |
International |
|
The real-time version of Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves highly accurate speech recognition, automatically determining the language and accurately identifying speech in 11 languages, while ensuring precise transcription even in complex audio environments. |
|
Reasoning, Visual understanding |
2025-10-21 |
International |
|
The largest dense model in the Qwen3-VL series, its reasoning version boasts multimodal reasoning capabilities second only to Qwen3-VL-235B-Thinking. It excels in STEM and math problem-solving, general image and video understanding, and achieves state-of-the-art performance in multimodal agent capabilities, making it ideal for complex multimodal reasoning tasks. |
|
Visual understanding |
2025-10-21 |
International |
|
The largest dense model in the Qwen3-VL series, in its non-inference version, delivers overall performance second only to Qwen3-VL-235B-Instruct. It excels in document recognition and comprehension, demonstrates strong spatial awareness and object identification capabilities, and achieves state-of-the-art performance in 2D visual detection and spatial reasoning. It is well-suited for complex perception tasks across a wide range of general-purpose scenarios. |
|
Reasoning, Visual understanding |
2025-10-15 |
International |
|
The Qwen3 series of small-scale visual understanding models effectively integrates thinking and non-thinking modes, delivering superior performance compared to the open-source Qwen3-VL-30B-A3B while maintaining fast response speeds. It features a comprehensive upgrade in image/video understanding, supporting ultra-long contexts such as extended videos and documents, spatial awareness, and object recognition across various domains. Equipped with 2D/3D visual localization capabilities, it is well-suited for tackling complex real-world tasks. |
|
Reasoning, Visual understanding |
2025-10-03 |
International |
|
The Thinking version of the second-largest MoE model in the Qwen3-VL series features fast response speeds and enhanced multimodal understanding and reasoning capabilities, visual agents, and support for extremely long contexts such as lengthy videos and documents. It also boasts comprehensively upgraded image/video understanding, spatial awareness, and object recognition abilities, making it well-suited for complex real-world tasks. |
|
Visual understanding |
2025-10-03 |
International |
|
The Instruct version of the second-largest MoE model in the Qwen3-VL series offers rapid response speeds and supports extremely long contexts like lengthy videos and documents. It includes comprehensively upgraded image/video understanding, spatial awareness, and object recognition capabilities, as well as 2D/3D visual localization, enabling it to handle intricate real-world challenges. |
|
Reasoning, Visual understanding |
2025-09-30 |
International |
|
The Thinking version of the 8B Dense model in the Qwen3-VL series consumes less GPU memory and is capable of performing multimodal understanding and reasoning. It supports extremely long contexts such as lengthy videos and documents, 2D/3D visual localization, and features comprehensively upgraded image/video understanding, spatial awareness, and object recognition abilities. |
|
Visual understanding |
2025-09-30 |
International |
|
The Instruct version of the 8B Dense model in the Qwen3-VL series requires less GPU memory and provides comprehensively upgraded image/video understanding, support for extremely long contexts like lengthy videos and documents, spatial awareness, and object recognition capabilities, making it suitable for tackling complex real-world tasks. |
|
Speech recognition |
2025-09-25 |
International |
|
Fun's multilingual speech recognition model supports over 31 languages and allows for free language switching, making it the top choice for users expanding overseas, especially to Southeast Asia. Fun-asr is an upgraded version of this model; switching to Fun-asr is recommended. |
|
Image generation |
2025-09-24 |
International |
|
The upgraded Wan2.5 Preview text to image model, newly upgraded model architecture significantly enhances visual aesthetics, design sensibility, and realistic texture. It excels in precise instruction adherence, generates text proficiently in English, Chinese, and less common languages, and supports the generation of complex structured long texts, charts, and architectural diagrams. |
|
Image generation |
2025-09-24 |
International |
|
The upgraded Wan2.5 Preview image edit model, newly upgraded model architecture supports rich image editing capabilities via instruction control, with enhanced instruction adherence. It also enables multi-image reference generation with high consistency and demonstrates excellent text generation performance. |
|
Text generation, Reasoning |
2025-09-24 |
International |
|
The Qwen 3 series Max model has undergone specialized upgrades in agent programming and tool invocation compared to the preview version. The officially released model this time has achieved state-of-the-art (SOTA) performance in its field and is better suited to meet the demands of agents operating in more complex scenarios. |
|
Realtime speech translation |
2025-09-24 |
International |
|
The real-time version of Qwen3-LiveTranslate-Flash, which is a high-precision, highly responsive, and robust multilingual simultaneous audio and video interpretation model. Leveraging Qwen3-Omni's powerful infrastructure, massive multimodal data, cross-language and cross-modal alignment, and visual enhancement technologies, Qwen3-LiveTranslate-Flash offers both offline and real-time audio and video translation capabilities. It can understand 19 languages and speak 10 languages, including 8 Chinese dialects. |
|
Speech recognition |
2025-09-24 |
International |
|
The Fun ASR version, updated in April 2026, fully supports the seven major dialect systems of traditional Chinese (Mandarin/Wu/Xiang/Gan/Hakka/Min/Yue) and is compatible with over 20 regional Mandarin accents. It features specific optimizations for the rhythm, meter, and literary expression characteristics of classical Chinese poetry, improving the accuracy of poetry recognition and making it suitable for cultural heritage, educational explanations, and audiobooks. Improved punctuation prediction and text normalization capabilities make the output text more consistent with written expression habits; numbers, dates, and amounts are automatically converted to standard formats, enhancing readability and professionalism. The language support has been expanded to include English, Japanese, Korean, Vietnamese, Thai, Indonesian, Malay, Filipino, Hindi, Arabic, French, German, Spanish, Portuguese, Russian, Italian, Dutch, Swedish, Danish, Finnish, Norwegian, Greek, Polish, Czech, Hungarian, Romanian, Bulgarian, Croatian, and Slovak, totaling 30 languages. This version is equivalent to the snapshot released on November 7, 2025. |
|
Video generation |
2025-09-23 |
International |
|
The upgraded Wan2.5 Preview text to video model, newly upgraded model architecture supports synchronized audio generation with visuals, enables 10-second long video generation, and offers enhanced instruction adherence, improved motion capabilities, and superior image quality. |
|
Video generation |
2025-09-23 |
International |
|
The upgraded Wan2.5 Preview image to video model, newly upgraded model architecture supports synchronized audio generation with visuals, enables 10-second long video generation, and offers enhanced instruction adherence, improved motion capabilities, and superior image quality. |
|
Visual understanding |
2025-09-23 |
International |
|
The Qwen3 series VL models effectively integrates thinking and non-thinking modes, achieving world-leading performance in visual agent capabilities on public benchmark datasets such as OS World. This version features comprehensive upgrades in areas like visual coding, spatial perception, and multimodal reasoning, significantly enhancing visual perception and recognition abilities, and supporting the understanding of ultra-long videos. |
|
Reasoning, Visual understanding |
2025-09-23 |
International |
|
Qwen3 series VL models feature significantly enhanced multimodal reasoning capabilities, with a particular focus on optimizing the model for STEM and mathematical reasoning. Visual perception and recognition abilities have been comprehensively improved, and OCR capabilities have undergone a major upgrade. |
|
Visual understanding |
2025-09-23 |
International |
|
The Qwen3 series VL models has been comprehensively upgraded in areas such as visual coding and spatial perception. Its visual perception and recognition capabilities have significantly improved, supporting the understanding of ultra-long videos, and its OCR functionality has undergone a major enhancement. |
|
Realtime speech translation |
2025-09-23 |
International |
|
The real-time version of Qwen3-LiveTranslate-Flash, which is a high-precision, highly responsive, and robust multilingual simultaneous audio and video interpretation model. Leveraging Qwen3-Omni's powerful infrastructure, massive multimodal data, cross-language and cross-modal alignment, and visual enhancement technologies, Qwen3-LiveTranslate-Flash offers both offline and real-time audio and video translation capabilities. It can understand 19 languages and speak 10 languages, including 8 Chinese dialects.This version is a snapshot version from September 22, 2025. |
|
Text generation |
2025-09-23 |
International |
|
Qwen3-based code generation model with strong coding agent power, excels at tool calling and environment interaction, capable of autonomous programming with outstanding code capability while maintaining general ability. This is a snapshot from 23 September, 2025.Compared to the previous version (snapshot from July 22), it demonstrates improved robustness in downstream task performance and tool invocation, along with enhanced code security. |
|
Image generation |
2025-09-23 |
International |
|
The first image generation foundation model in the Qwen series that achieves significant advances in complex text rendering and precise image editing. Experiments show strong general capabilities in both image generation and editing, with exceptional performance in text rendering, especially for Chinese. |
|
Realtime speech recognition |
2025-09-23 |
International |
|
This is the real-time version of Tongyi Lab's next-generation end-to-end speech recognition model, based on leading proprietary speech technology, and boasts exceptional contextual awareness and high-precision speech transcription capabilities. Based on an end-to-end architecture, Fun-ASR integrates innovative RAG technology, supporting multi-dimensional features such as large-scale hotword customization, automatic filtering of sensitive and modal particles, ITN normalization, and punctuation prediction, significantly improving overall recognition accuracy and contextual relevance. Furthermore, Fun-ASR supports flexible switching between Chinese and English, covers multiple regional dialects, and boasts enhanced noise robustness, adapting to diverse and complex environments. |
|
Speech recognition |
2025-09-19 |
International |
|
Qwen3-Omni-30b-a3b-Captioner is a powerful fine-grained audio analysis model designed to generate accurate and comprehensive content descriptions in complex and changing audio scenarios. It can automatically parse and describe various audio content, from complex speech and ambient sounds to music and film and television sound effects, and can maintain stable and reliable output even in multi-source and mixed environments. |
|
Realtime omni-modal |
2025-09-17 |
International |
|
The real-time version of the Qwen3-Omni-Flash multimodal large-scale model, based on the Thinker–Talker Mixed Expert (MoE) architecture, supports efficient understanding and speech generation of text, images, audio, and video. It can interact with text in 119 languages and speech in 20 languages, generating human-like speech for precise cross-lingual communication. The model boasts powerful command-following and system prompt customization capabilities, flexibly adapting to conversational styles and character settings. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interaction experience. |
|
Omni-modal |
2025-09-17 |
International |
|
Qwen3-Omni-Flash multimodal large-scale model, based on the Thinker–Talker Mixed Expert (MoE) architecture, supports efficient understanding and speech generation of text, images, audio, and video. It can interact with text in 119 languages and speech in 20 languages, generating human-like speech for precise cross-lingual communication. The model boasts powerful command-following and system prompt customization capabilities, flexibly adapting to conversational styles and character settings. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interaction experience.This version is a snapshot version from September 15, 2025. |
|
Speech synthesis |
2025-09-16 |
International |
|
The Qwen3-TTS-Flash-Realtime-2025-09-18 model is Tongyi's latest real-time speech synthesis foundation model, featuring 17 expressive voices while delivering low-latency, high-stability audio synthesis. It supports multilingual and dialect outputs with consistent voice characteristics across languages. Trained on extensive datasets, the system autonomously adjusts vocal tones based on text semantics and demonstrates robust capabilities for complex content synthesis.This model is provided as a snapshot version. |
|
Speech synthesis |
2025-09-16 |
International |
|
The Qwen 3-TTS-Flash-2025-09-18 is Tongyi's latest offline text-to-speech foundation model, featuring 17 expressive voices while enabling low-latency, high-stability audio synthesis. It supports multilingual and dialect outputs with consistent voice characteristics across languages. Trained on massive datasets, the system automatically adjusts vocal tones based on text semantics and demonstrates robust capabilities for synthesizing complex content. This model is provided as a snapshot version. |
|
Speech recognition |
2025-09-16 |
International |
|
Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves high-precision speech recognition. It can automatically determine the language and accurately recognize speech in 11 languages, ensuring precise transcription even in complex audio environments. |
|
Visual understanding |
2025-09-16 |
International |
|
Qwen-VL-Plus is the enhanced version of the large visual language model. It significantly improves detail recognition and text recognition capabilities, supporting images with resolutions exceeding one million pixels and any aspect ratio specifications. The model delivers exceptional performance across a wide range of visual tasks. |
|
Visual understanding |
2025-09-16 |
International |
|
Qwen-VL-Max is a large-scale visual language model of the Qwen series. Compared to the Plus version, it further enhances visual reasoning capabilities and instruction-following abilities, offering higher levels of visual perception and cognition. It delivers optimal performance on more complex tasks. |
|
Text generation, Reasoning |
2025-09-16 |
International |
|
Qwen-Plus is an enhanced version of the Qwen ultra-large language model that supports multiple input languages such as Chinese and English. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
|
Visual understanding |
2025-09-12 |
International |
|
The all-new Wan2.2 First and Last Frame to video model is here. We've optimized motion stability and success rates, enhanced prompt adherence, and enabled seamless transitions between two images. |
|
Text generation, Reasoning |
2025-09-11 |
International |
|
A new generation of Qwen3-based open-source thinking mode models. This version offers improved instruction following and streamlined summary responses over the previous iteration (Qwen3-235B-A22B-Thinking-2507). |
|
Text generation |
2025-09-11 |
International |
|
A new generation of open-source, non-thinking mode model powered by Qwen3. This version demonstrates superior Chinese text understanding, augmented logical reasoning, and enhanced capabilities in text generation tasks over the previous iteration (Qwen3-235B-A22B-Instruct-2507). |
|
Text generation, Reasoning |
2025-09-11 |
International |
|
As of the September 11, 2025 snapshot, this release features enhanced instruction following and streamlined summary responses in thinking mode. Non-thinking mode provides superior Chinese text understanding and augmented logical reasoning capabilities. The model supports a 1M context length with tiered pricing. |
|
Text generation, Reasoning |
2025-09-05 |
International |
|
A preview version of the Max model in the Qwen 3 series, achieving an effective integration of thinking and non-thinking modes. In thinking mode, there is a significant enhancement in capabilities such as intelligent agent programming, common-sense reasoning, and reasoning across mathematics, science, and general domains. |
|
Text embedding |
2025-08-25 |
International |
|
The General Text Vector V4 version is a multi-language text vector model developed by the Tongyi Lab based on Qwen3. Compared to the V3 version, it significantly improves performance in text retrieval, clustering, and classification tasks. It achieves a 15% to 40% improvement in evaluation tasks such as MTEB multilingual, Chinese-English, and code retrieval. Additionally, it supports user-defined vector dimensions ranging from 64 to 2048. |
|
Image generation |
2025-08-18 |
International |
|
The first Qwen image editing model extends Qwen-Image's text rendering to editing tasks. It offers precise bilingual (Chinese/English) text editing, dual visual and semantic editing, and strong cross-benchmark performance. |
|
Video generation |
2025-08-15 |
International |
|
The upgraded Wan 2.2 image to video Flash model, delivers faster speed with optimized stability, more powerful prompt following, improved consistency for text, portraits, and products, and precise shot control. |
|
Realtime omni-modal |
2025-08-14 |
International |
|
The real-time version of Qwen's new large multimodal understanding and generation model, suitable for real-time audio interaction scenarios. It supports the understanding of audio accompanied by text, images, and video mixed inputs, and can simultaneously generate speech and text in stream, providing four natural tones. |
|
Image generation |
2025-08-14 |
International |
|
The first image generation foundation model in the Qwen series that achieves significant advances in complex text rendering and precise image editing. Experiments show strong general capabilities in both image generation and editing, with exceptional performance in text rendering, especially for Chinese. |
|
Text generation, Reasoning |
2025-08-01 |
International |
|
The Qwen3 Flash model offers a powerful fusion of thinking and non-thinking modes with dynamic in-conversation switching, excelling in complex reasoning while showing significant gains in instruction following and text comprehension. It supports a 1M context length and is billed on a tiered model corresponding to context usage. |
|
Text generation |
2025-07-31 |
International |
|
Qwen3-based code generation model that inherits the coding agent ability of Qwen3-Coder-480B-A35B-Instruct; code capability reaches SOTA at the same scale. |
|
Text generation, Reasoning |
2025-07-31 |
International |
|
The Qwen series model optimized for balanced performance, offering inference efficiency between Qwen-Max and Qwen-Turbo, is designed to handle moderately complex tasks effectively. This dynamically updated version implements changes without prior notice. |
|
Text generation, Reasoning |
2025-07-30 |
International |
|
Open-source Qwen3 thinking model; compared to the previous version (Qwen3-30B-A3B) excels in complex thinking tasks, including logic, math, science, code, and other challenging scenarios; instruction following, text understanding, and multilingual translation capabilities significantly improved. |
|
Text generation |
2025-07-29 |
International |
|
Based on Qwen3, this code generation model inherits the coding agent capabilities of Qwen3-Coder-Plus and supports multi-turn tool interaction. It features focused optimizations on repository-level understanding and enhanced tool-calling stability. |
|
Text generation |
2025-07-29 |
International |
|
Open-source Qwen3 non-thinking model; compared to the previous version (Qwen3-30B-A3B) shows major improvements in Chinese, English, and overall multilingual general capabilities. Optimized for subjective open-ended tasks, delivering responses significantly more aligned with user preferences and more helpful. |
|
Text generation, Reasoning |
2025-07-29 |
International |
|
Qwen3 series Plus model, integrates thinking and non-thinking modes and can switch modes during dialogue. Compared to the prior version, adds dedicated enhancements for Chinese & English capabilities and tool calling. This is a snapshot from 28 July, 2025; first to support 1 M context length, and uses tiered pricing. |
|
Video generation |
2025-07-28 |
International |
|
The upgraded Wan 2.2 Plus text to video model, delivers higher quality results with stable sweeping complex movements, cinematic vision control, more powerful prompt following, and realistic world recreation. |
|
Image generation |
2025-07-28 |
International |
|
The upgraded Wan 2.2 Plus text to image model, delivers richer image detail with enhanced creativity, stability, and realism. It also features stronger prompt following and native support for multiple styles. Up to 2 million pixel generation and prompt enhancement are supported as well. |
|
Image generation |
2025-07-28 |
International |
|
The upgraded Wan 2.2 Flash text to image model, delivers faster speed with enhanced creativity, stability, and realism. It also features stronger prompt following and native support for multiple styles. Up to 2 million pixel generation and prompt enhancement are supported as well. |
|
Video generation |
2025-07-28 |
International |
|
The upgraded Wan 2.2 Plus image to video model, delivers higher quality results with optimized stability, more powerful prompt following, improved consistency for text, portraits, and products, and precise shot control. |
|
Text generation, Reasoning |
2025-07-25 |
International |
|
Open-source Qwen3 thinking model; compared to the previous version (Qwen3-235B-A22B) shows major improvements in logical ability, general capabilities, knowledge enhancement, and creativity, suitable for high-difficulty, strong-thinking scenarios. |
|
Text generation |
2025-07-24 |
International |
|
Qwen-MT-Turbo is a large language model within the Qwen model series that specializes in multi-lingual translation. It provides high-quality and rapid translation services across 92 languages at a cost-effective price point. It also offers features such as terminology intervention, format preservation, and domain-specific translation to cater to the diverse needs of various applications, ensuring both efficiency and performance. |
|
Text generation |
2025-07-24 |
International |
|
Qwen-MT-Plus, the flagship translation model from our Qwen series, is now fully upgraded with the Qwen3 architecture. It supports 92 languages and delivers exceptionally accurate and natural-sounding translations. Its advanced capabilities in contextual understanding, terminology control, and format preservation make it a superior choice over traditional models, especially for specialized domains. |
|
Text generation |
2025-07-23 |
International |
|
Qwen3-based code generation model with strong coding agent power, excels at tool calling and environment interaction, capable of autonomous programming with outstanding code capability while maintaining general ability. |
|
Text generation |
2025-07-23 |
International |
|
Qwen3-based code generation model with strong coding agent power; code capability reaches open-source SOTA. |
|
Text generation |
2025-07-23 |
International |
|
Open-source Qwen3 non-thinking model; compared to the previous version (Qwen3-235B-A22B) shows slight improvements in subjective creativity and model safety. |
|
Text generation, Reasoning |
2025-07-23 |
International |
|
The Plus model of the Qwen3 Series, achieving effective integration of thinking mode and non-thinking mode, allows switching modes during conversations. This is a snapshot from July 14, 2025. Compared to the previous version, there has been a significant improvement in both Chinese and English capabilities under non-thinking mode, with enhanced tool-calling abilities. |
|
Text generation |
2025-07-21 |
International |
|
The Qwen Role-Playing Model Series is specifically optimized for Japanese anthropomorphic interaction scenarios. It demonstrates advanced capabilities in character consistency maintenance, context-aware dialogue progression, and empathetic engagement, enabling precise personalized character embodiment. This version significantly enhances Japanese linguistic localization (including dialects and honorifics), human-like role-playing authenticity, narrative coherence control, and scenario-based cognitive intelligence. |
|
Omni-modal |
2025-07-18 |
International |
|
The new multimodal understanding and generation model, supporting text, image, speech, and video input as well as mixed input understanding. It can simultaneously stream text and speech, with significantly enhanced speed for multimodal content understanding. Additionally, it offers four natural dialogue tones. |
|
Image generation |
2025-05-22 |
International |
|
Wan2.1 Text-to-Image Turbo version, faster generation speed. Upgraded in image beauty, realism, and artistry. Stronger semantic understanding ability, rich style generalization ability, supports up to 2 million pixel generation, supports smart prompt rewriting. |
|
Image generation |
2025-05-22 |
International |
|
Wan2.1 Text-to-Image Plus version, Generate more image details. Upgraded in image beauty, realism, and artistry. Stronger semantic understanding ability, rich style generalization ability, supports up to 2 million pixel generation, supports smart prompt rewriting. |
|
Video generation |
2025-05-14 |
International |
|
Wan2.1-VACE-Plus All-in-One Video Creation and Editing model.It supports local editing, video repainting, background outpainting, duration extension, image reference, and other video editing and generation tasks, and supports multimodal conditional control through text, images, and videos. |
|
Text generation, Reasoning |
2025-05-12 |
International |
|
Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability reaches the SOTA level in the same scale industry, and its general capability significantly surpasses Qwen2.5-7B. |
|
Text generation, Reasoning |
2025-05-12 |
International |
|
Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability significantly surpasses QwQ, and its general capability markedly exceeds Qwen2.5-32B-Instruct, reaching the SOTA level in the same scale industry. |
|
Text generation, Reasoning |
2025-05-12 |
International |
|
Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability rivals QwQ-32B with a smaller parameter size, and its general capability significantly surpasses Qwen2.5-14B, reaching the SOTA level in the same scale industry. |
|
Text generation, Reasoning |
2025-05-12 |
International |
|
Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability significantly surpasses QwQ, and its general capability markedly exceeds Qwen2.5-72B-Instruct, reaching the SOTA level in the same scale industry. |
|
Text generation, Reasoning |
2025-05-12 |
International |
|
Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability reaches the SOTA level in the same scale industry, and its general capability significantly surpasses Qwen2.5-14B. |
|
Text generation, Reasoning |
2025-04-29 |
International |
|
The Plus model of the Qwen3 series, effectively integrates thinking mode and non-thinking mode, allowing for mode switching during conversations. Its reasoning capabilities significantly surpass those of QwQ, and its general capabilities notably exceed those of Qwen2.5-Plus, reaching the SOTA level in the same scale within the industry. This model is the snapshot from April 28, 2025. |
|
Video generation |
2025-04-07 |
International |
|
Wan2.1 start and end frames to video Plus version, generate a smooth transition video for two images. Support for large and complex movements, adherence to physical laws, rich artistic styles, and visual quality at the film and television level. The ability to follow instructions is further enhanced, resulting in richer details in the generated video. |
|
Omni-modal |
2025-03-26 |
International |
|
The new multi-modal understanding and generation model trained based on Qwen2.5. It supports text, image, speech, video, and mixed input understanding and can simultaneously generate streams of text and speech, significantly improves the speed of multi-modal content understanding. It provides four natural tones. |
|
Reasoning, Visual understanding |
2025-03-26 |
International |
|
The Tongyi Qianwen QVQ visual reasoning model supports visual input and chain-of-thought output, demonstrating stronger capabilities in mathematics, programming, visual analysis, creation, and general tasks. |
|
Reasoning |
2025-03-05 |
International |
|
The enhanced version of the Qwen QwQ reasoning model, trained on the Qwen2.5 model, has significantly improved its reasoning capabilities through reinforcement learning. The model's core metrics in mathematics and coding (e.g., AIME 24/25, LiveCodeBench) as well as some general metrics (e.g., IFEval, LiveBench) have reached the level of the full version of DeepSeek-R1. |
|
Video generation |
2025-02-27 |
International |
|
Wan2.1 image to video Turbo version, make the static image generated video. Support for large and complex movements, adherence to physical laws, artistic styles, and visual quality of movies. The ability to follow instructions is further improved, and generation more fast. |
|
Text generation, Reasoning |
2025-01-30 |
International |
|
The Turbo model of the Qwen3 series. It effectively integrates thinking mode and non-thinking mode, allowing for mode switching during conversations. Its reasoning capabilities rival those of QwQ-32B with a smaller parameter size, while its general capabilities significantly surpass those of Qwen2.5-Turbo, achieving the SOTA level in the same scale within the industry. |
|
Text generation |
2025-01-30 |
International |
|
The Qwen series of models, which are well-balanced in capabilities, offer reasoning performance and speed that fall between Qwen-Max and Qwen-Turbo, making them suitable for moderately complex tasks. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
|
Text generation |
2025-01-30 |
International |
|
Qwen-Max supports a parameter scale of hundreds of billions and multiple input languages such as Chinese and English. Qwen-Max is updated in a rolling manner. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
|
Video generation |
2025-01-20 |
International |
|
Wan2.1 image to video Plus version, make the static image generated video. Support for large and complex movements, adherence to physical laws, artistic styles, and visual quality of movies. The ability to follow instructions is further improved, and better video quality. |
|
Video generation |
2025-01-09 |
International |
|
Wan2.1 text to video Turbo version, one sentence generated video. Support for large and complex movements, adherence to physical laws, artistic styles, and visual quality of movies. The ability to follow instructions is further improved, and the generation speed is faster. |
|
Video generation |
2025-01-09 |
International |
|
Wan2.1 text to video Plus version, one sentence generated video. Support for large and complex movements, adherence to physical laws, artistic styles, and visual quality of movies. The ability to follow instructions is further improved, and better video quality. |
|
Text embedding |
2024-07-12 |
International |
|
The general text vectorization model is a multilingual unified text vectorization model developed by Tongyi Lab based on the large language model (LLM) foundation. This model is designed for multiple mainstream languages worldwide and provides advanced vectorization services to help developers convert text data into high-quality vector data. |
US (Virginia)
|
Model type |
Date |
Service scope |
Model ID |
Description |
|
Text generation, Reasoning |
2026-07-07 |
US |
|
GLM-5.2 is the latest flagship model from Zhipu AI, designed for long-horizon tasks with support for an ultra-long 1M context window. It features powerful logical reasoning, long-text comprehension, and code generation capabilities, balancing performance with inference efficiency. It excels across multi-task benchmarks and is well-suited for intelligent interaction, enterprise applications, and development assistance scenarios. |
|
Text generation, Reasoning |
2026-07-07 |
Global |
|
GLM-5.2 is the latest flagship model from Zhipu AI, designed for long-horizon tasks with support for an ultra-long 1M context window. It features powerful logical reasoning, long-text comprehension, and code generation capabilities, balancing performance with inference efficiency. It excels across multi-task benchmarks and is well-suited for intelligent interaction, enterprise applications, and development assistance scenarios. |
|
Text generation, Reasoning, Visual understanding |
2026-07-03 |
US |
|
The Qwen3.6 native vision-language Flash model series delivers a significant performance boost over the 3.5-Flash version. This model particularly excels in agentic coding capabilities, substantially outperforming its predecessor on multiple code-agent benchmarks, as well as in mathematical and code reasoning. In terms of vision, it features markedly improved spatial intelligence, with especially notable enhancements in object localization and object detection. |
|
Text generation, Reasoning, Visual understanding |
2026-07-03 |
Global |
|
The Qwen3.6 native vision-language Flash model series delivers a significant performance boost over the 3.5-Flash version. This model particularly excels in agentic coding capabilities, substantially outperforming its predecessor on multiple code-agent benchmarks, as well as in mathematical and code reasoning. In terms of vision, it features markedly improved spatial intelligence, with especially notable enhancements in object localization and object detection. |
|
Text generation |
2026-07-02 |
Global |
|
The role-playing model of the Qwen series. This is a dynamically updated version, and notifications will be provided in advance for any model updates. It is suitable for anthropomorphic role-playing and has optimized capabilities in following predefined character instructions, advancing conversations, and demonstrating active listening and empathy. Additionally, it supports the deep restoration of personalized characters. |
|
Text generation, Reasoning, Visual understanding |
2026-06-26 |
US |
|
Among the Qwen3.7 series, the cost-effective Plus model builds on its robust text capabilities while delivering a comprehensive upgrade to its vision‑language abilities, all while preserving its full‑stack agent‑level intelligence for coding, tool use, and productivity workflows. Its key distinguishing feature is multi‑modal interactive hybrid agent capabilities, enabling it to perceive real‑world scenes, read screens and interact with GUIs, generate code based on visual references, and perform end‑to‑end navigation within mobile apps. |
|
Text generation, Reasoning, Visual understanding |
2026-06-26 |
Global |
|
Among the Qwen3.7 series, the cost-effective Plus model builds on its robust text capabilities while delivering a comprehensive upgrade to its vision‑language abilities, all while preserving its full‑stack agent‑level intelligence for coding, tool use, and productivity workflows. Its key distinguishing feature is multi‑modal interactive hybrid agent capabilities, enabling it to perceive real‑world scenes, read screens and interact with GUIs, generate code based on visual references, and perform end‑to‑end navigation within mobile apps. |
|
Text generation, Reasoning |
2026-06-26 |
US |
|
The Max model, the largest and most capable in the Qwen3.7 series, currently offers a pure‑text‑only interface for public experimentation. Qwen3.7 is a next‑generation flagship model designed for the agent‑centric era, with its core strengths lying in the breadth and depth of its agent‑level capabilities: it excels at programming, office and productivity tasks, and long‑term autonomous execution. |
|
Text generation, Reasoning |
2026-06-26 |
Global |
|
The Max model, the largest and most capable in the Qwen3.7 series, currently offers a pure‑text‑only interface for public experimentation. Qwen3.7 is a next‑generation flagship model designed for the agent‑centric era, with its core strengths lying in the breadth and depth of its agent‑level capabilities: it excels at programming, office and productivity tasks, and long‑term autonomous execution. |
|
Video generation |
2026-06-16 |
Global |
|
HappyHorse-1.1-T2V supports Text-to-Video generation with improved semantic understanding, cinematic shot control, and dynamic motion rendering. It more accurately captures creative intent, producing high-quality videos with smoother motion, richer details, stronger visual consistency, and more natural character actions, scene atmosphere, and physical dynamics. |
|
Video generation |
2026-06-16 |
Global |
|
HappyHorse-1.1-R2V supports Reference-to-Video generation with significantly improved stability in subject, scene style, and visual consistency. With support for up to 9 reference images, it can more accurately understand and preserve creative intent, delivering stronger controllability and expressiveness across characters, scenes, styles, and cinematic motion. |
|
Video generation |
2026-06-16 |
Global |
|
HappyHorse-1.1-I2V supports Image-to-Video generation with improved visual quality, dynamic performance, and cross-clip consistency. It more accurately understands the input image and preserves creative intent, delivering significant improvements in skin texture realism, character ID consistency across clips, motion smoothness, text rendering stability, and audio-visual synchronization, producing high-quality videos with greater realism, richer details, and stronger overall consistency. |
|
Text generation, Reasoning |
2026-06-16 |
Global |
|
GLM-5.2 is the latest flagship model from Zhipu AI, designed for long-horizon tasks with support for an ultra-long 1M context window. It features powerful logical reasoning, long-text comprehension, and code generation capabilities, balancing performance with inference efficiency. It excels across multi-task benchmarks and is well-suited for intelligent interaction, enterprise applications, and development assistance scenarios. |
|
Text generation, Reasoning, Visual understanding |
2026-06-15 |
Global |
|
kimi-k2.7-code is Kimi's most intelligent coding model to date. It follows instructions more reliably over long contexts and completes programming tasks with higher success rates. It supports text, image, and video inputs, along with thinking mode, conversation, and agent tasks. |
|
Text generation, Reasoning, Visual understanding |
2026-06-09 |
Global |
|
The Max model, the largest and most capable in the Qwen3.7 series, has added visual‑modal understanding compared to the May 20 snapshot, enabling it to perceive real‑world scenes and supporting multimodal interactive hybrid agent capabilities. This version is based on a snapshot taken on June 8, 2026. |
|
Text generation, Reasoning, Visual understanding |
2026-06-01 |
Global |
|
Among the Qwen3.7 series, the cost-effective Plus model builds on its robust text capabilities while delivering a comprehensive upgrade to its vision‑language abilities, all while preserving its full‑stack agent‑level intelligence for coding, tool use, and productivity workflows. Its key distinguishing feature is multi‑modal interactive hybrid agent capabilities, enabling it to perceive real‑world scenes, read screens and interact with GUIs, generate code based on visual references, and perform end‑to‑end navigation within mobile apps. |
|
Text generation, Reasoning |
2026-05-20 |
Global |
|
The Max model, the largest and most capable in the Qwen3.7 series, currently offers a pure‑text‑only interface for public experimentation. Qwen3.7 is a next‑generation flagship model designed for the agent‑centric era, with its core strengths lying in the breadth and depth of its agent‑level capabilities: it excels at programming, office and productivity tasks, and long‑term autonomous execution. |
|
Text generation, Reasoning |
2026-05-11 |
US |
|
A flagship MoE large model with 1.6 trillion parameters and 49 billion activated parameters, natively supporting context lengths of up to one million tokens. Trained on a vast corpus of high-quality data, it excels in advanced mathematical reasoning, complex logical inference, specialized coding, and deep analysis of long-form text, making it well-suited for demanding applications such as cutting-edge research, sophisticated office workflows, and advanced AI agents. |
|
Text generation, Reasoning |
2026-05-11 |
Global |
|
A flagship MoE large model with 1.6 trillion parameters and 49 billion activated parameters, natively supporting context lengths of up to one million tokens. Trained on a vast corpus of high-quality data, it excels in advanced mathematical reasoning, complex logical inference, specialized coding, and deep analysis of long-form text, making it well-suited for demanding applications such as cutting-edge research, sophisticated office workflows, and advanced AI agents. |
|
Text generation, Reasoning |
2026-05-11 |
US |
|
A highly efficient, lightweight MoE model with 284 billion parameters in total and 13 billion activated parameters, natively supporting context windows of up to one million tokens. It offers fast inference speed, low latency, and cost-effective invocation, delivering well-balanced overall performance. Designed for high-concurrency, lightweight workloads, it is ideally suited for common, essential use cases such as everyday dialogue, content creation, basic RAG applications, and batch text processing. |
|
Text generation, Reasoning |
2026-05-11 |
Global |
|
A highly efficient, lightweight MoE model with 284 billion parameters in total and 13 billion activated parameters, natively supporting context windows of up to one million tokens. It offers fast inference speed, low latency, and cost-effective invocation, delivering well-balanced overall performance. Designed for high-concurrency, lightweight workloads, it is ideally suited for common, essential use cases such as everyday dialogue, content creation, basic RAG applications, and batch text processing. |
|
Text generation, Reasoning, Visual understanding |
2026-04-29 |
Global |
|
Kimi-k2.5 is Moonshot AI's most versatile model with native multimodal architecture, supporting visual/text inputs, thinking/non-thinking modes, and both dialogue and agent tasks. |
|
Video generation |
2026-04-26 |
Global |
|
HappyHorse-1.0-V2V supports advanced video editing through natural language instructions. It allows for local or global editing of video elements using up to 5 reference images, precisely preserving original motion dynamics to achieve superior expressiveness. |
|
Video generation |
2026-04-26 |
Global |
|
HappyHorse-1.0-R2V supports Reference-to-Video generation, offering enhanced stability in subject and scene referencing. Capable of processing up to 9 reference images, it precisely preserves creative intent to deliver superior performance. |
|
Text generation, Reasoning |
2026-04-24 |
Global |
|
A flagship MoE large model with 1.6 trillion parameters and 49 billion activated parameters, natively supporting context lengths of up to one million tokens. Trained on a vast corpus of high-quality data, it excels in advanced mathematical reasoning, complex logical inference, specialized coding, and deep analysis of long-form text, making it well-suited for demanding applications such as cutting-edge research, sophisticated office workflows, and advanced AI agents. |
|
Text generation, Reasoning |
2026-04-24 |
Global |
|
A highly efficient, lightweight MoE model with 284 billion parameters in total and 13 billion activated parameters, natively supporting context windows of up to one million tokens. It offers fast inference speed, low latency, and cost-effective invocation, delivering well-balanced overall performance. Designed for high-concurrency, lightweight workloads, it is ideally suited for common, essential use cases such as everyday dialogue, content creation, basic RAG applications, and batch text processing. |
|
Video generation |
2026-04-22 |
Global |
|
HappyHorse-1.0-T2V supports text-to-video generation, featuring highly realistic dynamic rendering. It accurately comprehends text semantics to produce high-quality videos that are fluid, natural, and rich in detail. |
|
Video generation |
2026-04-22 |
Global |
|
HappyHorse-1.0-I2V enables image-to-video generation, featuring highly realistic dynamic rendering. It accurately comprehends both text and image semantics to produce high-quality videos that are fluid, natural, and rich in detail. |
|
Text generation, Reasoning, Visual understanding |
2026-04-17 |
Global |
|
The Qwen3.6 native vision-language Flash model series delivers a significant performance boost over the 3.5-Flash version. This model particularly excels in agentic coding capabilities, substantially outperforming its predecessor on multiple code-agent benchmarks, as well as in mathematical and code reasoning. In terms of vision, it features markedly improved spatial intelligence, with especially notable enhancements in object localization and object detection. |
|
Text generation, Reasoning, Visual understanding |
2026-04-17 |
Global |
|
The Qwen3.6 35B-A3B native vision-language model is built on a hybrid architecture that integrates linear attention mechanisms with a sparse mixture-of-experts framework, achieving higher inference efficiency. Compared with the 3.5-35B-A3B, this model demonstrates significantly improved agentic coding capabilities, mathematical and code reasoning abilities, spatial intelligence, as well as object localization and object detection performance. |
|
Text generation, Reasoning |
2026-04-14 |
Global |
|
GLM-5.1 is a model developed by Zhipu AI, specifically designed for long-horizon tasks. It has 744 billion parameters, supports an ultra-long context of 200k tokens, and can generate up to 128k tokens in a single response. GLM-5.1 excels in logical reasoning, long-text understanding, and code generation, while balancing performance with inference efficiency. It delivers outstanding results across multiple multi-task benchmarks and is well-suited for applications such as intelligent human-computer interaction, enterprise solutions, and developer assistance. |
|
Text generation, Reasoning, Visual understanding |
2026-04-01 |
Global |
|
The Qwen3.6 native vision-language Plus series models demonstrate exceptional performance on par with the current state-of-the-art models, with a significant improvement in overall results compared to the 3.5 series. The models have been markedly enhanced in code-related capabilities such as agentic coding, front-end programming, and Vibe coding, as well as in multi-modal general object recognition, OCR, and object localization. |
|
Text generation |
2026-03-30 |
US |
|
Qwen-MT-Lite is a large language model of the Qwen model series that specializes in multi-lingual translation. It provides high-quality and rapid translation services across 32 languages at a cost-effective price. It offers features such as terminology intervention, format preservation, and domain-specific translation to cater to the diverse needs of various applications, ensuring both efficiency and performance. |
|
Text generation |
2026-03-30 |
Global |
|
Qwen-MT-Lite is a large language model of the Qwen model series that specializes in multi-lingual translation. It provides high-quality and rapid translation services across 32 languages at a cost-effective price. It offers features such as terminology intervention, format preservation, and domain-specific translation to cater to the diverse needs of various applications, ensuring both efficiency and performance. |
|
Visual understanding |
2026-03-14 |
US |
|
The Qwen3 series of small-sized visual understanding models effectively integrates thinking and non-thinking modes. Compared with the snapshot taken on October 15, 2025, the overall performance of the model has improved significantly: it delivers enhanced capabilities in general visual recognition and reasoning, and shows marked improvements in recognition accuracy across various business scenarios such as security, in-store inspections, equipment monitoring, and photo-based problem solving. This version is a snapshot as of January 22, 2026. |
|
Visual understanding |
2026-03-14 |
Global |
|
The Qwen3 series of small-sized visual understanding models effectively integrates thinking and non-thinking modes. Compared with the snapshot taken on October 15, 2025, the overall performance of the model has improved significantly: it delivers enhanced capabilities in general visual recognition and reasoning, and shows marked improvements in recognition accuracy across various business scenarios such as security, in-store inspections, equipment monitoring, and photo-based problem solving. This version is a snapshot as of January 22, 2026. |
|
Text generation, Reasoning, Visual understanding |
2026-02-23 |
Global |
|
The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the 3 series, these models deliver a leap forward in performance for both pure text and multimodal tasks, offering fast response times while balancing inference speed and overall performance. |
|
Text generation, Reasoning, Visual understanding |
2026-02-23 |
Global |
|
The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall performance is comparable to that of the Qwen3.5-27B. |
|
Text generation, Reasoning, Visual understanding |
2026-02-23 |
Global |
|
The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of the Qwen3.5-122B-A10B. |
|
Text generation, Reasoning, Visual understanding |
2026-02-23 |
Global |
|
The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of overall performance, this model is second only to Qwen3.5-397B-A17B. Its text capabilities significantly outperform those of Qwen3-235B-2507, and its visual capabilities surpass those of Qwen3-VL-235B. |
|
Text generation, Reasoning, Visual understanding |
2026-02-15 |
Global |
|
The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of task evaluations, the 3.5 series consistently demonstrates performance on par with state-of-the-art leading models. Compared to the 3 series, these models show a leap forward in both pure-text and multimodal capabilities. |
|
Text generation, Reasoning, Visual understanding |
2026-02-15 |
Global |
|
The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers state-of-the-art performance comparable to leading-edge models across a wide range of tasks, including language understanding, logical reasoning, code generation, agent-based tasks, image understanding, video understanding, and graphical user interface (GUI) interactions. With its robust code-generation and agent capabilities, the model exhibits strong generalization across diverse agent. |
|
Video generation |
2025-12-16 |
Global |
|
Wan2.6 reference to video, supports using a specified person or any object as a reference, precisely maintaining consistency of appearance and voice, and allows multi‑character reference for joint performances. |
|
Image generation |
2025-12-15 |
Global |
|
Wan2.6 text to image, Upgraded visual quality, aesthetics, and instruction-following deliver precise style control, realistic portraits, long-text understanding, and broad historical/cultural IP coverage, enabling high-quality, highly expressive visual generation. |
|
Image generation |
2025-12-15 |
Global |
|
Wan2.6 Image, An all-round image generation model that supports joint text–image reasoning, multi-image creative fusion, commercial-grade consistency, aesthetic style transfer, and precise control of framing and lighting, significantly enhancing consistency, controllability, and expressiveness in image generation. |
|
Video generation |
2025-12-03 |
US |
|
Wan2.6 text to video, Intelligent shot scheduling supports multi-shot storytelling, generating multi-shot narrative videos with consistent subjects, scenes, and atmosphere, with a maximum duration of 15 seconds. It delivers better instruction following, and improved visual fidelity. |
|
Video generation |
2025-12-03 |
Global |
|
Wan2.6 text to video, Intelligent shot scheduling supports multi-shot storytelling, generating multi-shot narrative videos with consistent subjects, scenes, and atmosphere, with a maximum duration of 15 seconds. It delivers better instruction following, and improved visual fidelity. |
|
Video generation |
2025-12-03 |
Global |
|
Wan2.6 text to video, Intelligent shot scheduling supports multi-shot storytelling, generating multi-shot narrative videos with consistent subjects, scenes, and atmosphere, with a maximum duration of 15 seconds. It delivers better instruction following, and improved visual fidelity. |
|
Video generation |
2025-12-03 |
US |
|
Wan2.6 image to video, intelligent shot scheduling enables multi‑camera storytelling, delivers higher‑quality voice generation, supports stable multi‑speaker dialogue with more natural and realistic vocal timbres, and supports generation of clips up to 15 seconds in length. |
|
Video generation |
2025-12-03 |
Global |
|
Wan2.6 image to video, intelligent shot scheduling enables multi‑camera storytelling, delivers higher‑quality voice generation, supports stable multi‑speaker dialogue with more natural and realistic vocal timbres, and supports generation of clips up to 15 seconds in length. |
|
Video generation |
2025-12-03 |
Global |
|
Wan2.6 image to video, intelligent shot scheduling enables multi‑camera storytelling, delivers higher‑quality voice generation, supports stable multi‑speaker dialogue with more natural and realistic vocal timbres, and supports generation of clips up to 15 seconds in length. |
|
Text generation, Reasoning |
2025-12-01 |
US |
|
This version is a snapshot as of December 1, 2025, and features improved reasoning capabilities compared to the July 28 snapshot. Agent capabilities and multi-turn tool invocation abilities have been further enhanced, and performance on subjective creative tasks has improved significantly. It supports a context length of up to 1 million tokens, with tiered pricing based on context length. |
|
Text generation, Reasoning |
2025-12-01 |
Global |
|
This version is a snapshot as of December 1, 2025, and features improved reasoning capabilities compared to the July 28 snapshot. Agent capabilities and multi-turn tool invocation abilities have been further enhanced, and performance on subjective creative tasks has improved significantly. It supports a context length of up to 1 million tokens, with tiered pricing based on context length. |
|
Text generation, Reasoning |
2025-12-01 |
Global |
|
This version is a snapshot as of December 1, 2025, and features improved reasoning capabilities compared to the July 28 snapshot. Agent capabilities and multi-turn tool invocation abilities have been further enhanced, and performance on subjective creative tasks has improved significantly. It supports a context length of up to 1 million tokens, with tiered pricing based on context length. |
|
Visual understanding |
2025-11-21 |
Global |
|
This model is a snapshot version from November 20, 2025, and is based on the latest Qwen-VL3 architecture with a comprehensive upgrade. It features significant improvements in document parsing and text localization capabilities, as well as substantial reductions in end-to-end latency and illusions. |
|
Text generation |
2025-11-19 |
Global |
|
Qwen-MT-Lite is a large language model of the Qwen model series that specializes in multi-lingual translation. It provides high-quality and rapid translation services across 32 languages at a cost-effective price. It offers features such as terminology intervention, format preservation, and domain-specific translation to cater to the diverse needs of various applications, ensuring both efficiency and performance. |
|
Text generation |
2025-11-11 |
Global |
|
Qwen-MT-Flash, a large language model from the Qwen series, has been fully upgraded with the Qwen 3 architecture for significantly enhanced performance and translation quality. It provides rapid, cost-effective translation across 92 languages, while supporting advanced features such as terminology intervention, format preservation, and domain-specific adaptation. It is the ideal choice for applications requiring a powerful balance of speed, quality, and cost. |
|
Reasoning, Visual understanding |
2025-10-21 |
Global |
|
The largest dense model in the Qwen3-VL series, its reasoning version boasts multimodal reasoning capabilities second only to Qwen3-VL-235B-Thinking. It excels in STEM and math problem-solving, general image and video understanding, and achieves state-of-the-art performance in multimodal agent capabilities, making it ideal for complex multimodal reasoning tasks. |
|
Visual understanding |
2025-10-21 |
Global |
|
The largest dense model in the Qwen3-VL series, in its non-inference version, delivers overall performance second only to Qwen3-VL-235B-Instruct. It excels in document recognition and comprehension, demonstrates strong spatial awareness and object identification capabilities, and achieves state-of-the-art performance in 2D visual detection and spatial reasoning. It is well-suited for complex perception tasks across a wide range of general-purpose scenarios. |
|
Visual understanding |
2025-10-15 |
US |
|
The Qwen3 series of small-scale visual understanding models effectively integrates thinking and non-thinking modes, delivering superior performance compared to the open-source Qwen3-VL-30B-A3B while maintaining fast response speeds. It features a comprehensive upgrade in image/video understanding, supporting ultra-long contexts such as extended videos and documents, spatial awareness, and object recognition across various domains. Equipped with 2D/3D visual localization capabilities, it is well-suited for tackling complex real-world tasks. |
|
Visual understanding |
2025-10-15 |
Global |
|
The Qwen3 series of small-scale visual understanding models effectively integrates thinking and non-thinking modes, delivering superior performance compared to the open-source Qwen3-VL-30B-A3B while maintaining fast response speeds. It features a comprehensive upgrade in image/video understanding, supporting ultra-long contexts such as extended videos and documents, spatial awareness, and object recognition across various domains. Equipped with 2D/3D visual localization capabilities, it is well-suited for tackling complex real-world tasks. |
|
Reasoning, Visual understanding |
2025-10-15 |
Global |
|
The Qwen3 series of small-scale visual understanding models effectively integrates thinking and non-thinking modes, delivering superior performance compared to the open-source Qwen3-VL-30B-A3B while maintaining fast response speeds. It features a comprehensive upgrade in image/video understanding, supporting ultra-long contexts such as extended videos and documents, spatial awareness, and object recognition across various domains. Equipped with 2D/3D visual localization capabilities, it is well-suited for tackling complex real-world tasks. |
|
Reasoning, Visual understanding |
2025-10-03 |
Global |
|
The Thinking version of the second-largest MoE model in the Qwen3-VL series features fast response speeds and enhanced multimodal understanding and reasoning capabilities, visual agents, and support for extremely long contexts such as lengthy videos and documents. It also boasts comprehensively upgraded image/video understanding, spatial awareness, and object recognition abilities, making it well-suited for complex real-world tasks. |
|
Visual understanding |
2025-10-03 |
Global |
|
The Instruct version of the second-largest MoE model in the Qwen3-VL series offers rapid response speeds and supports extremely long contexts like lengthy videos and documents. It includes comprehensively upgraded image/video understanding, spatial awareness, and object recognition capabilities, as well as 2D/3D visual localization, enabling it to handle intricate real-world challenges. |
|
Reasoning, Visual understanding |
2025-09-30 |
Global |
|
The Thinking version of the 8B Dense model in the Qwen3-VL series consumes less GPU memory and is capable of performing multimodal understanding and reasoning. It supports extremely long contexts such as lengthy videos and documents, 2D/3D visual localization, and features comprehensively upgraded image/video understanding, spatial awareness, and object recognition abilities. |
|
Visual understanding |
2025-09-30 |
Global |
|
The Instruct version of the 8B Dense model in the Qwen3-VL series requires less GPU memory and provides comprehensively upgraded image/video understanding, support for extremely long contexts like lengthy videos and documents, spatial awareness, and object recognition capabilities, making it suitable for tackling complex real-world tasks. |
|
Text generation, Reasoning |
2025-09-24 |
Global |
|
The Qwen 3 series Max model has undergone specialized upgrades in agent programming and tool invocation compared to the preview version. The officially released model this time has achieved state-of-the-art (SOTA) performance in its field and is better suited to meet the demands of agents operating in more complex scenarios. |
|
Visual understanding |
2025-09-23 |
Global |
|
The Qwen3 series VL models effectively integrates thinking and non-thinking modes, achieving world-leading performance in visual agent capabilities on public benchmark datasets such as OS World. This version features comprehensive upgrades in areas like visual coding, spatial perception, and multimodal reasoning, significantly enhancing visual perception and recognition abilities, and supporting the understanding of ultra-long videos. |
|
Reasoning, Visual understanding |
2025-09-23 |
Global |
|
Qwen3 series VL models feature significantly enhanced multimodal reasoning capabilities, with a particular focus on optimizing the model for STEM and mathematical reasoning. Visual perception and recognition abilities have been comprehensively improved, and OCR capabilities have undergone a major upgrade. |
|
Visual understanding |
2025-09-23 |
Global |
|
The Qwen3 series VL models has been comprehensively upgraded in areas such as visual coding and spatial perception. Its visual perception and recognition capabilities have significantly improved, supporting the understanding of ultra-long videos, and its OCR functionality has undergone a major enhancement. |
|
Text generation |
2025-09-23 |
Global |
|
Qwen3-based code generation model with strong coding agent power, excels at tool calling and environment interaction, capable of autonomous programming with outstanding code capability while maintaining general ability. This is a snapshot from 23 September, 2025.Compared to the previous version (snapshot from July 22), it demonstrates improved robustness in downstream task performance and tool invocation, along with enhanced code security. |
|
Speech recognition |
2025-09-16 |
US |
|
Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves high-precision speech recognition. It can automatically determine the language and accurately recognize speech in 11 languages, ensuring precise transcription even in complex audio environments. |
|
Speech recognition |
2025-09-16 |
Global |
|
Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves high-precision speech recognition. It can automatically determine the language and accurately recognize speech in 11 languages, ensuring precise transcription even in complex audio environments. |
|
Text generation, Reasoning |
2025-09-16 |
US |
|
Qwen-Plus is an enhanced version of the Qwen ultra-large language model that supports multiple input languages such as Chinese and English. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
|
Text generation, Reasoning |
2025-09-16 |
Global |
|
Qwen-Plus is an enhanced version of the Qwen ultra-large language model that supports multiple input languages such as Chinese and English. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
|
Text generation, Reasoning |
2025-09-16 |
Global |
|
Qwen-Plus is an enhanced version of the Qwen ultra-large language model that supports multiple input languages such as Chinese and English. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
|
Text generation, Reasoning |
2025-09-11 |
Global |
|
A new generation of Qwen3-based open-source thinking mode models. This version offers improved instruction following and streamlined summary responses over the previous iteration (Qwen3-235B-A22B-Thinking-2507). |
|
Text generation |
2025-09-11 |
Global |
|
A new generation of open-source, non-thinking mode model powered by Qwen3. This version demonstrates superior Chinese text understanding, augmented logical reasoning, and enhanced capabilities in text generation tasks over the previous iteration (Qwen3-235B-A22B-Instruct-2507). |
|
Text generation, Reasoning |
2025-09-11 |
Global |
|
As of the September 11, 2025 snapshot, this release features enhanced instruction following and streamlined summary responses in thinking mode. Non-thinking mode provides superior Chinese text understanding and augmented logical reasoning capabilities. The model supports a 1M context length with tiered pricing. |
|
Text generation, Reasoning |
2025-09-05 |
Global |
|
A preview version of the Max model in the Qwen 3 series, achieving an effective integration of thinking and non-thinking modes. In thinking mode, there is a significant enhancement in capabilities such as intelligent agent programming, common-sense reasoning, and reasoning across mathematics, science, and general domains. |
|
Text generation, Reasoning |
2025-08-01 |
US |
|
The Qwen3 Flash model offers a powerful fusion of thinking and non-thinking modes with dynamic in-conversation switching, excelling in complex reasoning while showing significant gains in instruction following and text comprehension. It supports a 1M context length and is billed on a tiered model corresponding to context usage. |
|
Text generation, Reasoning |
2025-08-01 |
Global |
|
The Qwen3 Flash model offers a powerful fusion of thinking and non-thinking modes with dynamic in-conversation switching, excelling in complex reasoning while showing significant gains in instruction following and text comprehension. It supports a 1M context length and is billed on a tiered model corresponding to context usage. |
|
Text generation, Reasoning |
2025-08-01 |
Global |
|
The Qwen3 Flash model offers a powerful fusion of thinking and non-thinking modes with dynamic in-conversation switching, excelling in complex reasoning while showing significant gains in instruction following and text comprehension. It supports a 1M context length and is billed on a tiered model corresponding to context usage. |
|
Text generation |
2025-07-31 |
Global |
|
Qwen3-based code generation model that inherits the coding agent ability of Qwen3-Coder-480B-A35B-Instruct; code capability reaches SOTA at the same scale. |
|
Text generation, Reasoning |
2025-07-30 |
Global |
|
Open-source Qwen3 thinking model; compared to the previous version (Qwen3-30B-A3B) excels in complex thinking tasks, including logic, math, science, code, and other challenging scenarios; instruction following, text understanding, and multilingual translation capabilities significantly improved. |
|
Text generation |
2025-07-29 |
Global |
|
Based on Qwen3, this code generation model inherits the coding agent capabilities of Qwen3-Coder-Plus and supports multi-turn tool interaction. It features focused optimizations on repository-level understanding and enhanced tool-calling stability. |
|
Text generation |
2025-07-29 |
Global |
|
Open-source Qwen3 non-thinking model; compared to the previous version (Qwen3-30B-A3B) shows major improvements in Chinese, English, and overall multilingual general capabilities. Optimized for subjective open-ended tasks, delivering responses significantly more aligned with user preferences and more helpful. |
|
Text generation, Reasoning |
2025-07-29 |
Global |
|
Qwen3 series Plus model, integrates thinking and non-thinking modes and can switch modes during dialogue. Compared to the prior version, adds dedicated enhancements for Chinese & English capabilities and tool calling. This is a snapshot from 28 July, 2025; first to support 1 M context length, and uses tiered pricing. |
|
Text generation, Reasoning |
2025-07-25 |
Global |
|
Open-source Qwen3 thinking model; compared to the previous version (Qwen3-235B-A22B) shows major improvements in logical ability, general capabilities, knowledge enhancement, and creativity, suitable for high-difficulty, strong-thinking scenarios. |
|
Text generation |
2025-07-24 |
Global |
|
Qwen-MT-Plus, the flagship translation model from our Qwen series, is now fully upgraded with the Qwen3 architecture. It supports 92 languages and delivers exceptionally accurate and natural-sounding translations. Its advanced capabilities in contextual understanding, terminology control, and format preservation make it a superior choice over traditional models, especially for specialized domains. |
|
Text generation |
2025-07-23 |
Global |
|
Qwen3-based code generation model with strong coding agent power, excels at tool calling and environment interaction, capable of autonomous programming with outstanding code capability while maintaining general ability. |
|
Text generation |
2025-07-23 |
Global |
|
Qwen3-based code generation model with strong coding agent power; code capability reaches open-source SOTA. |
|
Text generation |
2025-07-23 |
Global |
|
Open-source Qwen3 non-thinking model; compared to the previous version (Qwen3-235B-A22B) shows slight improvements in subjective creativity and model safety. |
|
Visual understanding |
2025-07-07 |
Global |
|
Qwen-VL_OCR is an OCR model trained based on Qwen-VL. It aggregates various image-text recognition, parsing, and processing tasks through a unified model approach, offering powerful image-text recognition capabilities. |
|
Text generation, Reasoning |
2025-05-12 |
Global |
|
Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability reaches the SOTA level in the same scale industry, and its general capability significantly surpasses Qwen2.5-7B. |
|
Text generation, Reasoning |
2025-05-12 |
Global |
|
Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability significantly surpasses QwQ, and its general capability markedly exceeds Qwen2.5-32B-Instruct, reaching the SOTA level in the same scale industry. |
|
Text generation, Reasoning |
2025-05-12 |
Global |
|
Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability rivals QwQ-32B with a smaller parameter size, and its general capability significantly surpasses Qwen2.5-14B, reaching the SOTA level in the same scale industry. |
|
Text generation, Reasoning |
2025-05-12 |
Global |
|
Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability significantly surpasses QwQ, and its general capability markedly exceeds Qwen2.5-72B-Instruct, reaching the SOTA level in the same scale industry. |
|
Text generation, Reasoning |
2025-05-12 |
Global |
|
Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability reaches the SOTA level in the same scale industry, and its general capability significantly surpasses Qwen2.5-14B. |
Germany (Frankfurt)
|
Model type |
Date |
Service scope |
Model ID |
Description |
|
Text generation |
2026-07-02 |
Global |
|
The role-playing model of the Qwen series. This is a dynamically updated version, and notifications will be provided in advance for any model updates. It is suitable for anthropomorphic role-playing and has optimized capabilities in following predefined character instructions, advancing conversations, and demonstrating active listening and empathy. Additionally, it supports the deep restoration of personalized characters. |
|
Video generation |
2026-06-16 |
Global |
|
HappyHorse-1.1-T2V supports Text-to-Video generation with improved semantic understanding, cinematic shot control, and dynamic motion rendering. It more accurately captures creative intent, producing high-quality videos with smoother motion, richer details, stronger visual consistency, and more natural character actions, scene atmosphere, and physical dynamics. |
|
Video generation |
2026-06-16 |
Global |
|
HappyHorse-1.1-R2V supports Reference-to-Video generation with significantly improved stability in subject, scene style, and visual consistency. With support for up to 9 reference images, it can more accurately understand and preserve creative intent, delivering stronger controllability and expressiveness across characters, scenes, styles, and cinematic motion. |
|
Video generation |
2026-06-16 |
Global |
|
HappyHorse-1.1-I2V supports Image-to-Video generation with improved visual quality, dynamic performance, and cross-clip consistency. It more accurately understands the input image and preserves creative intent, delivering significant improvements in skin texture realism, character ID consistency across clips, motion smoothness, text rendering stability, and audio-visual synchronization, producing high-quality videos with greater realism, richer details, and stronger overall consistency. |
|
Text generation, Reasoning |
2026-06-16 |
Global |
|
GLM-5.2 is the latest flagship model from Zhipu AI, designed for long-horizon tasks with support for an ultra-long 1M context window. It features powerful logical reasoning, long-text comprehension, and code generation capabilities, balancing performance with inference efficiency. It excels across multi-task benchmarks and is well-suited for intelligent interaction, enterprise applications, and development assistance scenarios. |
|
Text generation, Reasoning, Visual understanding |
2026-06-15 |
Global |
|
kimi-k2.7-code is Kimi's most intelligent coding model to date. It follows instructions more reliably over long contexts and completes programming tasks with higher success rates. It supports text, image, and video inputs, along with thinking mode, conversation, and agent tasks. |
|
Text generation, Reasoning, Visual understanding |
2026-06-09 |
Global |
|
The Max model, the largest and most capable in the Qwen3.7 series, has added visual‑modal understanding compared to the May 20 snapshot, enabling it to perceive real‑world scenes and supporting multimodal interactive hybrid agent capabilities. This version is based on a snapshot taken on June 8, 2026. |
|
Text generation, Reasoning, Visual understanding |
2026-06-01 |
Global |
|
Among the Qwen3.7 series, the cost-effective Plus model builds on its robust text capabilities while delivering a comprehensive upgrade to its vision‑language abilities, all while preserving its full‑stack agent‑level intelligence for coding, tool use, and productivity workflows. Its key distinguishing feature is multi‑modal interactive hybrid agent capabilities, enabling it to perceive real‑world scenes, read screens and interact with GUIs, generate code based on visual references, and perform end‑to‑end navigation within mobile apps. |
|
Text generation, Reasoning |
2026-05-20 |
Global |
|
The Max model, the largest and most capable in the Qwen3.7 series, currently offers a pure‑text‑only interface for public experimentation. Qwen3.7 is a next‑generation flagship model designed for the agent‑centric era, with its core strengths lying in the breadth and depth of its agent‑level capabilities: it excels at programming, office and productivity tasks, and long‑term autonomous execution. |
|
Text generation, Reasoning, Visual understanding |
2026-04-29 |
Global |
|
Kimi-k2.5 is Moonshot AI's most versatile model with native multimodal architecture, supporting visual/text inputs, thinking/non-thinking modes, and both dialogue and agent tasks. |
|
Video generation |
2026-04-26 |
Global |
|
HappyHorse-1.0-V2V supports advanced video editing through natural language instructions. It allows for local or global editing of video elements using up to 5 reference images, precisely preserving original motion dynamics to achieve superior expressiveness. |
|
Video generation |
2026-04-26 |
Global |
|
HappyHorse-1.0-R2V supports Reference-to-Video generation, offering enhanced stability in subject and scene referencing. Capable of processing up to 9 reference images, it precisely preserves creative intent to deliver superior performance. |
|
Text generation, Reasoning |
2026-04-24 |
Global |
|
A flagship MoE large model with 1.6 trillion parameters and 49 billion activated parameters, natively supporting context lengths of up to one million tokens. Trained on a vast corpus of high-quality data, it excels in advanced mathematical reasoning, complex logical inference, specialized coding, and deep analysis of long-form text, making it well-suited for demanding applications such as cutting-edge research, sophisticated office workflows, and advanced AI agents. |
|
Text generation, Reasoning |
2026-04-24 |
Global |
|
A highly efficient, lightweight MoE model with 284 billion parameters in total and 13 billion activated parameters, natively supporting context windows of up to one million tokens. It offers fast inference speed, low latency, and cost-effective invocation, delivering well-balanced overall performance. Designed for high-concurrency, lightweight workloads, it is ideally suited for common, essential use cases such as everyday dialogue, content creation, basic RAG applications, and batch text processing. |
|
Video generation |
2026-04-22 |
Global |
|
HappyHorse-1.0-T2V supports text-to-video generation, featuring highly realistic dynamic rendering. It accurately comprehends text semantics to produce high-quality videos that are fluid, natural, and rich in detail. |
|
Video generation |
2026-04-22 |
Global |
|
HappyHorse-1.0-I2V enables image-to-video generation, featuring highly realistic dynamic rendering. It accurately comprehends both text and image semantics to produce high-quality videos that are fluid, natural, and rich in detail. |
|
Text generation, Reasoning, Visual understanding |
2026-04-17 |
Global |
|
The Qwen3.6 native vision-language Flash model series delivers a significant performance boost over the 3.5-Flash version. This model particularly excels in agentic coding capabilities, substantially outperforming its predecessor on multiple code-agent benchmarks, as well as in mathematical and code reasoning. In terms of vision, it features markedly improved spatial intelligence, with especially notable enhancements in object localization and object detection. |
|
Text generation, Reasoning, Visual understanding |
2026-04-17 |
Global |
|
The Qwen3.6 35B-A3B native vision-language model is built on a hybrid architecture that integrates linear attention mechanisms with a sparse mixture-of-experts framework, achieving higher inference efficiency. Compared with the 3.5-35B-A3B, this model demonstrates significantly improved agentic coding capabilities, mathematical and code reasoning abilities, spatial intelligence, as well as object localization and object detection performance. |
|
Text generation, Reasoning |
2026-04-14 |
Global |
|
GLM-5.1 is a model developed by Zhipu AI, specifically designed for long-horizon tasks. It has 744 billion parameters, supports an ultra-long context of 200k tokens, and can generate up to 128k tokens in a single response. GLM-5.1 excels in logical reasoning, long-text understanding, and code generation, while balancing performance with inference efficiency. It delivers outstanding results across multiple multi-task benchmarks and is well-suited for applications such as intelligent human-computer interaction, enterprise solutions, and developer assistance. |
|
Text generation, Reasoning, Visual understanding |
2026-04-01 |
Global |
|
The Qwen3.6 native vision-language Plus series models demonstrate exceptional performance on par with the current state-of-the-art models, with a significant improvement in overall results compared to the 3.5 series. The models have been markedly enhanced in code-related capabilities such as agentic coding, front-end programming, and Vibe coding, as well as in multi-modal general object recognition, OCR, and object localization. |
|
Text generation, Reasoning, Visual understanding |
2026-02-23 |
EU |
|
The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the 3 series, these models deliver a leap forward in performance for both pure text and multimodal tasks, offering fast response times while balancing inference speed and overall performance. |
|
Text generation, Reasoning, Visual understanding |
2026-02-23 |
Global |
|
The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the 3 series, these models deliver a leap forward in performance for both pure text and multimodal tasks, offering fast response times while balancing inference speed and overall performance. |
|
Text generation, Reasoning, Visual understanding |
2026-02-23 |
Global |
|
The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall performance is comparable to that of the Qwen3.5-27B. |
|
Text generation, Reasoning, Visual understanding |
2026-02-23 |
Global |
|
The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of the Qwen3.5-122B-A10B. |
|
Text generation, Reasoning, Visual understanding |
2026-02-23 |
Global |
|
The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of overall performance, this model is second only to Qwen3.5-397B-A17B. Its text capabilities significantly outperform those of Qwen3-235B-2507, and its visual capabilities surpass those of Qwen3-VL-235B. |
|
Text generation |
2026-02-20 |
EU |
|
The new-generation code generation model in the Qwen3 series delivers performance close to that of Qwen3-Coder-Plus while offering even better capabilities. The model has been optimized with a focus on repository-level understanding, supports multi-turn tool interactions, and enhances its compatibility with agentic coding tools. |
|
Text generation, Reasoning, Visual understanding |
2026-02-15 |
Global |
|
The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of task evaluations, the 3.5 series consistently demonstrates performance on par with state-of-the-art leading models. Compared to the 3 series, these models show a leap forward in both pure-text and multimodal capabilities. |
|
Text generation, Reasoning, Visual understanding |
2026-02-15 |
Global |
|
The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers state-of-the-art performance comparable to leading-edge models across a wide range of tasks, including language understanding, logical reasoning, code generation, agent-based tasks, image understanding, video understanding, and graphical user interface (GUI) interactions. With its robust code-generation and agent capabilities, the model exhibits strong generalization across diverse agent. |
|
Text generation, Reasoning |
2026-01-23 |
EU |
|
Compared with the snapshot as of September 23, 2025, the Qwen-3 series Max model in this release achieves an effective integration of thinking and non-thinking modes, resulting in a comprehensive and substantial improvement in the model's overall performance. In thinking mode, the model simultaneously supports web search, web information extraction, and a code interpreter tool, enabling it to tackle more complex and challenging problems with greater accuracy by leveraging external tools while engaging in slow, deliberative reasoning. This version is based on a snapshot taken on January 23, 2026. |
|
Visual understanding |
2026-01-22 |
EU |
|
The Qwen3 series of small-sized visual understanding models effectively integrates thinking and non-thinking modes. Compared with the snapshot taken on October 15, 2025, the overall performance of the model has improved significantly: it delivers enhanced capabilities in general visual recognition and reasoning, and shows marked improvements in recognition accuracy across various business scenarios such as security, in-store inspections, equipment monitoring, and photo-based problem solving. This version is a snapshot as of January 22, 2026. |
|
Video generation |
2025-12-16 |
Global |
|
Wan2.6 reference to video, supports using a specified person or any object as a reference, precisely maintaining consistency of appearance and voice, and allows multi‑character reference for joint performances. |
|
Image generation |
2025-12-15 |
Global |
|
Wan2.6 text to image, Upgraded visual quality, aesthetics, and instruction-following deliver precise style control, realistic portraits, long-text understanding, and broad historical/cultural IP coverage, enabling high-quality, highly expressive visual generation. |
|
Image generation |
2025-12-15 |
Global |
|
Wan2.6 Image, An all-round image generation model that supports joint text–image reasoning, multi-image creative fusion, commercial-grade consistency, aesthetic style transfer, and precise control of framing and lighting, significantly enhancing consistency, controllability, and expressiveness in image generation. |
|
Video generation |
2025-12-03 |
Global |
|
Wan2.6 text to video, Intelligent shot scheduling supports multi-shot storytelling, generating multi-shot narrative videos with consistent subjects, scenes, and atmosphere, with a maximum duration of 15 seconds. It delivers better instruction following, and improved visual fidelity. |
|
Video generation |
2025-12-03 |
Global |
|
Wan2.6 image to video, intelligent shot scheduling enables multi‑camera storytelling, delivers higher‑quality voice generation, supports stable multi‑speaker dialogue with more natural and realistic vocal timbres, and supports generation of clips up to 15 seconds in length. |
|
— |
2025-12-01 |
EU |
|
This version is a snapshot as of December 1, 2025, and features improved reasoning capabilities compared to the July 28 snapshot. Agent capabilities and multi-turn tool invocation abilities have been further enhanced, and performance on subjective creative tasks has improved significantly. It supports a context length of up to 1 million tokens, with tiered pricing based on context length. |
|
— |
2025-12-01 |
Global |
|
This version is a snapshot as of December 1, 2025, and features improved reasoning capabilities compared to the July 28 snapshot. Agent capabilities and multi-turn tool invocation abilities have been further enhanced, and performance on subjective creative tasks has improved significantly. It supports a context length of up to 1 million tokens, with tiered pricing based on context length. |
|
Visual understanding |
2025-11-21 |
Global |
|
This model is a snapshot version from November 20, 2025, and is based on the latest Qwen-VL3 architecture with a comprehensive upgrade. It features significant improvements in document parsing and text localization capabilities, as well as substantial reductions in end-to-end latency and illusions. |
|
Visual understanding |
2025-11-20 |
Global |
|
Qwen-VL_OCR is an OCR model trained based on Qwen-VL. It aggregates various image-text recognition, parsing, and processing tasks through a unified model approach, offering powerful image-text recognition capabilities. |
|
Text generation |
2025-11-19 |
Global |
|
Qwen-MT-Lite is a large language model of the Qwen model series that specializes in multi-lingual translation. It provides high-quality and rapid translation services across 32 languages at a cost-effective price. It offers features such as terminology intervention, format preservation, and domain-specific translation to cater to the diverse needs of various applications, ensuring both efficiency and performance. |
|
Text generation |
2025-11-11 |
Global |
|
Qwen-MT-Flash, a large language model from the Qwen series, has been fully upgraded with the Qwen 3 architecture for significantly enhanced performance and translation quality. It provides rapid, cost-effective translation across 92 languages, while supporting advanced features such as terminology intervention, format preservation, and domain-specific adaptation. It is the ideal choice for applications requiring a powerful balance of speed, quality, and cost. |
|
Reasoning, Visual understanding |
2025-10-21 |
Global |
|
The largest dense model in the Qwen3-VL series, its reasoning version boasts multimodal reasoning capabilities second only to Qwen3-VL-235B-Thinking. It excels in STEM and math problem-solving, general image and video understanding, and achieves state-of-the-art performance in multimodal agent capabilities, making it ideal for complex multimodal reasoning tasks. |
|
Visual understanding |
2025-10-21 |
Global |
|
The largest dense model in the Qwen3-VL series, in its non-inference version, delivers overall performance second only to Qwen3-VL-235B-Instruct. It excels in document recognition and comprehension, demonstrates strong spatial awareness and object identification capabilities, and achieves state-of-the-art performance in 2D visual detection and spatial reasoning. It is well-suited for complex perception tasks across a wide range of general-purpose scenarios. |
|
Reasoning, Visual understanding |
2025-10-15 |
EU |
|
The Qwen3 series of small-scale visual understanding models effectively integrates thinking and non-thinking modes, delivering superior performance compared to the open-source Qwen3-VL-30B-A3B while maintaining fast response speeds. It features a comprehensive upgrade in image/video understanding, supporting ultra-long contexts such as extended videos and documents, spatial awareness, and object recognition across various domains. Equipped with 2D/3D visual localization capabilities, it is well-suited for tackling complex real-world tasks. |
|
Reasoning, Visual understanding |
2025-10-15 |
Global |
|
The Qwen3 series of small-scale visual understanding models effectively integrates thinking and non-thinking modes, delivering superior performance compared to the open-source Qwen3-VL-30B-A3B while maintaining fast response speeds. It features a comprehensive upgrade in image/video understanding, supporting ultra-long contexts such as extended videos and documents, spatial awareness, and object recognition across various domains. Equipped with 2D/3D visual localization capabilities, it is well-suited for tackling complex real-world tasks. |
|
Reasoning, Visual understanding |
2025-10-03 |
Global |
|
The Thinking version of the second-largest MoE model in the Qwen3-VL series features fast response speeds and enhanced multimodal understanding and reasoning capabilities, visual agents, and support for extremely long contexts such as lengthy videos and documents. It also boasts comprehensively upgraded image/video understanding, spatial awareness, and object recognition abilities, making it well-suited for complex real-world tasks. |
|
Visual understanding |
2025-10-03 |
Global |
|
The Instruct version of the second-largest MoE model in the Qwen3-VL series offers rapid response speeds and supports extremely long contexts like lengthy videos and documents. It includes comprehensively upgraded image/video understanding, spatial awareness, and object recognition capabilities, as well as 2D/3D visual localization, enabling it to handle intricate real-world challenges. |
|
Reasoning, Visual understanding |
2025-09-30 |
Global |
|
The Thinking version of the 8B Dense model in the Qwen3-VL series consumes less GPU memory and is capable of performing multimodal understanding and reasoning. It supports extremely long contexts such as lengthy videos and documents, 2D/3D visual localization, and features comprehensively upgraded image/video understanding, spatial awareness, and object recognition abilities. |
|
Visual understanding |
2025-09-30 |
Global |
|
The Instruct version of the 8B Dense model in the Qwen3-VL series requires less GPU memory and provides comprehensively upgraded image/video understanding, support for extremely long contexts like lengthy videos and documents, spatial awareness, and object recognition capabilities, making it suitable for tackling complex real-world tasks. |
|
Text generation, Reasoning |
2025-09-24 |
EU |
|
The Qwen 3 series Max model has undergone specialized upgrades in agent programming and tool invocation compared to the preview version. The officially released model this time has achieved state-of-the-art (SOTA) performance in its field and is better suited to meet the demands of agents operating in more complex scenarios. |
|
Text generation, Reasoning |
2025-09-24 |
Global |
|
The Qwen 3 series Max model has undergone specialized upgrades in agent programming and tool invocation compared to the preview version. The officially released model this time has achieved state-of-the-art (SOTA) performance in its field and is better suited to meet the demands of agents operating in more complex scenarios. |
|
Visual understanding |
2025-09-23 |
EU |
|
The Qwen3 series VL models effectively integrates thinking and non-thinking modes, achieving world-leading performance in visual agent capabilities on public benchmark datasets such as OS World. This version features comprehensive upgrades in areas like visual coding, spatial perception, and multimodal reasoning, significantly enhancing visual perception and recognition abilities, and supporting the understanding of ultra-long videos. |
|
Visual understanding |
2025-09-23 |
Global |
|
The Qwen3 series VL models effectively integrates thinking and non-thinking modes, achieving world-leading performance in visual agent capabilities on public benchmark datasets such as OS World. This version features comprehensive upgrades in areas like visual coding, spatial perception, and multimodal reasoning, significantly enhancing visual perception and recognition abilities, and supporting the understanding of ultra-long videos. |
|
Reasoning, Visual understanding |
2025-09-23 |
Global |
|
Qwen3 series VL models feature significantly enhanced multimodal reasoning capabilities, with a particular focus on optimizing the model for STEM and mathematical reasoning. Visual perception and recognition abilities have been comprehensively improved, and OCR capabilities have undergone a major upgrade. |
|
Visual understanding |
2025-09-23 |
Global |
|
The Qwen3 series VL models has been comprehensively upgraded in areas such as visual coding and spatial perception. Its visual perception and recognition capabilities have significantly improved, supporting the understanding of ultra-long videos, and its OCR functionality has undergone a major enhancement. |
|
Text generation |
2025-09-23 |
Global |
|
Qwen3-based code generation model with strong coding agent power, excels at tool calling and environment interaction, capable of autonomous programming with outstanding code capability while maintaining general ability. This is a snapshot from 23 September, 2025.Compared to the previous version (snapshot from July 22), it demonstrates improved robustness in downstream task performance and tool invocation, along with enhanced code security. |
|
Text generation, Reasoning |
2025-09-16 |
EU |
|
Qwen-Plus is an enhanced version of the Qwen ultra-large language model that supports multiple input languages such as Chinese and English. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
|
Text generation, Reasoning |
2025-09-16 |
Global |
|
Qwen-Plus is an enhanced version of the Qwen ultra-large language model that supports multiple input languages such as Chinese and English. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
|
Text generation, Reasoning |
2025-09-11 |
Global |
|
A new generation of Qwen3-based open-source thinking mode models. This version offers improved instruction following and streamlined summary responses over the previous iteration (Qwen3-235B-A22B-Thinking-2507). |
|
Text generation |
2025-09-11 |
Global |
|
A new generation of open-source, non-thinking mode model powered by Qwen3. This version demonstrates superior Chinese text understanding, augmented logical reasoning, and enhanced capabilities in text generation tasks over the previous iteration (Qwen3-235B-A22B-Instruct-2507). |
|
Text generation, Reasoning |
2025-09-11 |
Global |
|
As of the September 11, 2025 snapshot, this release features enhanced instruction following and streamlined summary responses in thinking mode. Non-thinking mode provides superior Chinese text understanding and augmented logical reasoning capabilities. The model supports a 1M context length with tiered pricing. |
|
Text generation, Reasoning |
2025-09-05 |
Global |
|
A preview version of the Max model in the Qwen 3 series, achieving an effective integration of thinking and non-thinking modes. In thinking mode, there is a significant enhancement in capabilities such as intelligent agent programming, common-sense reasoning, and reasoning across mathematics, science, and general domains. |
|
Text generation, Reasoning |
2025-08-01 |
Global |
|
The Qwen3 Flash model offers a powerful fusion of thinking and non-thinking modes with dynamic in-conversation switching, excelling in complex reasoning while showing significant gains in instruction following and text comprehension. It supports a 1M context length and is billed on a tiered model corresponding to context usage. |
|
Text generation |
2025-07-31 |
Global |
|
Qwen3-based code generation model that inherits the coding agent ability of Qwen3-Coder-480B-A35B-Instruct; code capability reaches SOTA at the same scale. |
|
Text generation, Reasoning |
2025-07-30 |
Global |
|
Open-source Qwen3 thinking model; compared to the previous version (Qwen3-30B-A3B) excels in complex thinking tasks, including logic, math, science, code, and other challenging scenarios; instruction following, text understanding, and multilingual translation capabilities significantly improved. |
|
Text generation |
2025-07-29 |
Global |
|
Based on Qwen3, this code generation model inherits the coding agent capabilities of Qwen3-Coder-Plus and supports multi-turn tool interaction. It features focused optimizations on repository-level understanding and enhanced tool-calling stability. |
|
Text generation |
2025-07-29 |
Global |
|
Open-source Qwen3 non-thinking model; compared to the previous version (Qwen3-30B-A3B) shows major improvements in Chinese, English, and overall multilingual general capabilities. Optimized for subjective open-ended tasks, delivering responses significantly more aligned with user preferences and more helpful. |
|
Text generation, Reasoning |
2025-07-29 |
Global |
|
Qwen3 series Plus model, integrates thinking and non-thinking modes and can switch modes during dialogue. Compared to the prior version, adds dedicated enhancements for Chinese & English capabilities and tool calling. This is a snapshot from 28 July, 2025; first to support 1 M context length, and uses tiered pricing. |
|
Text generation, Reasoning |
2025-07-25 |
Global |
|
Open-source Qwen3 thinking model; compared to the previous version (Qwen3-235B-A22B) shows major improvements in logical ability, general capabilities, knowledge enhancement, and creativity, suitable for high-difficulty, strong-thinking scenarios. |
|
Text generation |
2025-07-24 |
Global |
|
Qwen-MT-Plus, the flagship translation model from our Qwen series, is now fully upgraded with the Qwen3 architecture. It supports 92 languages and delivers exceptionally accurate and natural-sounding translations. Its advanced capabilities in contextual understanding, terminology control, and format preservation make it a superior choice over traditional models, especially for specialized domains. |
|
Text generation |
2025-07-23 |
Global |
|
Qwen3-based code generation model with strong coding agent power, excels at tool calling and environment interaction, capable of autonomous programming with outstanding code capability while maintaining general ability. |
|
Text generation |
2025-07-23 |
Global |
|
Qwen3-based code generation model with strong coding agent power; code capability reaches open-source SOTA. |
|
Text generation |
2025-07-23 |
Global |
|
Open-source Qwen3 non-thinking model; compared to the previous version (Qwen3-235B-A22B) shows slight improvements in subjective creativity and model safety. |
|
Text generation, Reasoning |
2025-05-12 |
Global |
|
Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability reaches the SOTA level in the same scale industry, and its general capability significantly surpasses Qwen2.5-7B. |
|
Text generation, Reasoning |
2025-05-12 |
Global |
|
Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability significantly surpasses QwQ, and its general capability markedly exceeds Qwen2.5-32B-Instruct, reaching the SOTA level in the same scale industry. |
|
Text generation, Reasoning |
2025-05-12 |
Global |
|
Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability rivals QwQ-32B with a smaller parameter size, and its general capability significantly surpasses Qwen2.5-14B, reaching the SOTA level in the same scale industry. |
|
Text generation, Reasoning |
2025-05-12 |
Global |
|
Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability significantly surpasses QwQ, and its general capability markedly exceeds Qwen2.5-72B-Instruct, reaching the SOTA level in the same scale industry. |
Japan (Tokyo)
|
Model type |
Date |
Service scope |
Model ID |
Description |
|
Text generation, Reasoning, Visual understanding |
2026-07-29 |
Global |
|
kimi-k2.7-code is Kimi's most intelligent coding model to date. It follows instructions more reliably over long contexts and completes programming tasks with higher success rates. It supports text, image, and video inputs, along with thinking mode, conversation, and agent tasks. |
|
Text generation |
2026-07-02 |
Global |
|
The role-playing model of the Qwen series. This is a dynamically updated version, and notifications will be provided in advance for any model updates. It is suitable for anthropomorphic role-playing and has optimized capabilities in following predefined character instructions, advancing conversations, and demonstrating active listening and empathy. Additionally, it supports the deep restoration of personalized characters. |
|
Video generation |
2026-06-16 |
Global |
|
HappyHorse-1.1-T2V supports Text-to-Video generation with improved semantic understanding, cinematic shot control, and dynamic motion rendering. It more accurately captures creative intent, producing high-quality videos with smoother motion, richer details, stronger visual consistency, and more natural character actions, scene atmosphere, and physical dynamics. |
|
Video generation |
2026-06-16 |
Global |
|
HappyHorse-1.1-R2V supports Reference-to-Video generation with significantly improved stability in subject, scene style, and visual consistency. With support for up to 9 reference images, it can more accurately understand and preserve creative intent, delivering stronger controllability and expressiveness across characters, scenes, styles, and cinematic motion. |
|
Video generation |
2026-06-16 |
Global |
|
HappyHorse-1.1-I2V supports Image-to-Video generation with improved visual quality, dynamic performance, and cross-clip consistency. It more accurately understands the input image and preserves creative intent, delivering significant improvements in skin texture realism, character ID consistency across clips, motion smoothness, text rendering stability, and audio-visual synchronization, producing high-quality videos with greater realism, richer details, and stronger overall consistency. |
|
Text generation, Reasoning, Visual understanding |
2026-06-01 |
Global |
|
Among the Qwen3.7 series, the cost-effective Plus model builds on its robust text capabilities while delivering a comprehensive upgrade to its vision‑language abilities, all while preserving its full‑stack agent‑level intelligence for coding, tool use, and productivity workflows. Its key distinguishing feature is multi‑modal interactive hybrid agent capabilities, enabling it to perceive real‑world scenes, read screens and interact with GUIs, generate code based on visual references, and perform end‑to‑end navigation within mobile apps. |
|
Text generation, Reasoning |
2026-05-21 |
Global |
|
The Max model, the largest and most capable in the Qwen3.7 series, currently offers a pure‑text‑only interface for public experimentation. Qwen3.7 is a next‑generation flagship model designed for the agent‑centric era, with its core strengths lying in the breadth and depth of its agent‑level capabilities: it excels at programming, office and productivity tasks, and long‑term autonomous execution. |
|
Text generation, Reasoning |
2026-05-11 |
Global |
|
A flagship MoE large model with 1.6 trillion parameters and 49 billion activated parameters, natively supporting context lengths of up to one million tokens. Trained on a vast corpus of high-quality data, it excels in advanced mathematical reasoning, complex logical inference, specialized coding, and deep analysis of long-form text, making it well-suited for demanding applications such as cutting-edge research, sophisticated office workflows, and advanced AI agents. |
|
Text generation, Reasoning |
2026-05-11 |
Global |
|
A highly efficient, lightweight MoE model with 284 billion parameters in total and 13 billion activated parameters, natively supporting context windows of up to one million tokens. It offers fast inference speed, low latency, and cost-effective invocation, delivering well-balanced overall performance. Designed for high-concurrency, lightweight workloads, it is ideally suited for common, essential use cases such as everyday dialogue, content creation, basic RAG applications, and batch text processing. |
|
Video generation |
2026-04-26 |
Global |
|
HappyHorse-1.0-V2V supports advanced video editing through natural language instructions. It allows for local or global editing of video elements using up to 5 reference images, precisely preserving original motion dynamics to achieve superior expressiveness. |
|
Text generation, Reasoning, Visual understanding |
2026-04-17 |
Global |
|
The Qwen3.6 native vision-language Flash model series delivers a significant performance boost over the 3.5-Flash version. This model particularly excels in agentic coding capabilities, substantially outperforming its predecessor on multiple code-agent benchmarks, as well as in mathematical and code reasoning. In terms of vision, it features markedly improved spatial intelligence, with especially notable enhancements in object localization and object detection. |
|
Text generation, Reasoning |
2026-04-14 |
Global |
|
GLM-5.1 is a model developed by Zhipu AI, specifically designed for long-horizon tasks. It has 744 billion parameters, supports an ultra-long context of 200k tokens, and can generate up to 128k tokens in a single response. GLM-5.1 excels in logical reasoning, long-text understanding, and code generation, while balancing performance with inference efficiency. It delivers outstanding results across multiple multi-task benchmarks and is well-suited for applications such as intelligent human-computer interaction, enterprise solutions, and developer assistance. |
|
Video generation |
2026-04-03 |
Global |
|
Wan2.7 video edit, supports both localized and global editing with prompt. Seamlessly replace elements using image references and replicate complex dynamic processes, including motion, special effects, and camera movements. |
|
Video generation |
2026-04-03 |
Global |
|
Wan2.7 text to video, performance fully reimagined. Delivering nuanced and organic emotional depth in narrative arcs and visceral, bone-crunching impact in action sequences. Enhanced by rhythmic cinematic cuts for unparalleled storytelling power. |
|
Video generation |
2026-04-03 |
Global |
|
Wan2.7 reference to video, enhanced consistency & performance. Delivering superior stability for characters, props, and scenes. Supports hybrid referencing of up to 5 mixed image/video inputs and audio timbre cloning. Together with core engine upgrades, it achieves unprecedented cinematic expressive power. |
|
Video generation |
2026-04-03 |
Global |
|
Wan2.7 image to video, performance fully reimagined. Delivering nuanced and organic emotional depth in narrative arcs and visceral, bone-crunching impact in action sequences. Enhanced by rhythmic cinematic cuts for unparalleled storytelling power. |
|
Image generation |
2026-04-01 |
Global |
|
Wan2.7 – image generation and editing, supports text to image, text/image to sequential images, image editing, multi-image reference generation, and interactive editing. Delivers enhanced performance in text rendering, subject consistency, and complex instruction following. |
|
Text generation, Reasoning, Visual understanding |
2026-04-01 |
Global |
|
The Qwen3.6 native vision-language Plus series models demonstrate exceptional performance on par with the current state-of-the-art models, with a significant improvement in overall results compared to the 3.5 series. The models have been markedly enhanced in code-related capabilities such as agentic coding, front-end programming, and Vibe coding, as well as in multi-modal general object recognition, OCR, and object localization. |