qwen3.8-flash
Qwen3.8-Flash is the latest multimodal model from the Qwen family, combining powerful reasoning and generation with remarkable speed. It natively supports a million-token context window, allowing it to process lengthy documents, entire codebases, and complex conversations in a single pass. It shines in coding assistance, agentic workflows, and visual understanding — whether it's fixing code autonomously, operating desktop applications, or analyzing charts and long videos. Fully compatible with both OpenAI and Anthropic API protocols, it integrates seamlessly with popular developer tools like Claude Code and Codex, making it easy to build high-concurrency applications and intelligent workflows. With strong performance and highly competitive inference costs, Qwen3.8-Flash is an ideal choice for developers and businesses seeking the best of both worlds in AI applications.
Model invocation ID (model parameter value): qwen3.8-flash
Model capabilities
China (Beijing)
| Capability | Status | Capability | Status |
|---|---|---|---|
Input modality | Image Text Video | Output modality | Text |
Model playground | Supported | Function Calling | Supported |
Structured output | Supported | Web search | Supported |
Prefix continuation | Supported | Context caching | Supported |
Batch inference | Supported | Fine-tuning | Not supported |
Singapore
Deployment scope: International
| Capability | Status | Capability | Status |
|---|---|---|---|
Input modality | Image Text Video | Output modality | Text |
Model playground | Supported | Function Calling | Supported |
Structured output | Supported | Web search | Supported |
Prefix continuation | Supported | Context caching | Supported |
Batch inference | Not supported | Fine-tuning | Not supported |
Context limits
| Parameter | Value | Parameter | Value |
|---|---|---|---|
Max input length | 991808 | Max output length | 131072 |
Max input length (thinking mode) | 983616 | Max output length (thinking mode) | 131072 |
Context length | 1000000 | Max thinking chain length | 262144 |
Model pricing
This page only shows the original model invocation price. For promotions and discounts, visit the Model Studio console.
China (Beijing)
| Billing item | Price (CNY) | Unit |
|---|---|---|
Input | 0.8 | Per 1 million tokens |
Output | 2.7 | Per 1 million tokens |
Input (cache hit) | 0.1 | Per 1 million tokens |
Input (Batch File) | 0.4 | Per 1 million tokens |
Output (Batch File) | 1.35 | Per 1 million tokens |
Explicit cache creation | 1.25 | Per 1 million tokens |
Explicit cache hit | 0.1 | Per 1 million tokens |
Input (Batch Chat) | 0.8 | Per 1 million tokens |
Output (Batch Chat) | 2.7 | Per 1 million tokens |
Singapore
Deployment scope: International
| Billing item | Price (CNY) | Unit |
|---|---|---|
Input | 1.094 | Per 1 million tokens |
Output | 3.427 | Per 1 million tokens |
Input (cache hit) | 0.117 | Per 1 million tokens |
Explicit cache creation | 1.458 | Per 1 million tokens |
Explicit cache hit | 0.117 | Per 1 million tokens |
Rate limits
China (Beijing)
| Parameter | Value |
|---|---|
RPM (requests per minute) | 30,000 |
TPM (tokens per minute) | 5,000,000 |
Singapore
Deployment scope: International
| Parameter | Value |
|---|---|
RPM (requests per minute) | 15,000 |
TPM (tokens per minute) | 2,000,000 |