Recommended models
Alibaba Cloud Model Studio offers Qwen and third-party models for text, image, audio, and video.
Text generation
Image & video
Understanding
Extract text descriptions or structured data from images and videos
Generation
Generate images and videos from text or images, with support for editing, reference, and high-resolution output
3D model generation
Generate 3D models from text or images to build 3D assets
Audio & speech
Text-to-speech
For audiobook reading, voice broadcasting, virtual avatars, and more
Music generation
Generate music from prompts or lyrics
Speech recognition
Dedicated ASR and LLM-based approaches — choose based on accuracy and flexibility
Speech-to-speech
End-to-end voice conversation without separate ASR and TTS calls
Omni
Integrates understanding and generation capabilities across text, image, audio, and video modalities
Embeddings & reranking
Convert text or multimodal content into vectors, combined with reranking to improve retrieval accuracy
View all models
Go to Model Plaza to browse all Qwen, third-party, domain-specific, and legacy models.