Recommended models

Updated at:

Alibaba Cloud Model Studio offers Qwen and third-party models for text, image, audio, and video.

Text generation

qwen3.8-maxqwen3.7-plusqwen3.7-flash
deepseek-v4-prodeepseek-v4-flashkimi/kimi-k3
glm-5.2ZHIPU/GLM-5.3kimi-k3
MiniMax-M3mimo-v2.5-pro

More

Image & video

Understanding

Extract text descriptions or structured data from images and videos

qwen3.8-maxqwen3.7-plusqwen3.5-omni-plus
kimi/kimi-k3

More

Generation

Generate images and videos from text or images, with support for editing, reference, and high-resolution output

qwen-image-3.0-prowan2.7-image-prohappyhorse-1.1-t2v
happyhorse-1.1-i2vhappyhorse-1.1-r2vhappyhorse-1.0-video-edit
wan3.0-video

More

3D model generation

Generate 3D models from text or images to build 3D assets

Tripo/Tripo-H3.1Tripo/Tripo-P1.0

More

Audio & speech

Text-to-speech

For audiobook reading, voice broadcasting, virtual avatars, and more

qwen-audio-3.0-tts-plusMiniMax/speech-2.8-hd

More

Music generation

Generate music from prompts or lyrics

fun-music-v1

More

Speech recognition

Dedicated ASR and LLM-based approaches — choose based on accuracy and flexibility

qwen-audio-3.0-asr-flash-streamingqwen-audio-3.0-asr-flash-filetransqwen3.5-omni-plus-realtime
qwen3.5-omni-plus

More

Speech-to-speech

End-to-end voice conversation without separate ASR and TTS calls

qwen-audio-3.0-realtime-plusqwen3.5-omni-plus

More

Omni

Integrates understanding and generation capabilities across text, image, audio, and video modalities

qwen3.5-omni-plus-realtimeqwen3.5-omni-plus

More

Embeddings & reranking

Convert text or multimodal content into vectors, combined with reranking to improve retrieval accuracy

text-embedding-v4qwen3.7-text-embeddingtongyi-embedding-vision-plus
qwen3-rerank

More

View all models

Go to Model Plaza to browse all Qwen, third-party, domain-specific, and legacy models.