# Alibaba Cloud Model Studio > Model Studio (Bailian): Qwen/qwen3.8-max, DeepSeek, Kimi, GLM APIs for text, vision, image/video, speech via OpenAI-compatible and DashScope endpoints — get an API Key to start. Per-token pay-as-you-go, free quota (Singapore), Coding Plan/Token Plan subscriptions. Coding Plan billing and errors (401/403/rate limit) are covered here; IDE code completion is Lingma. Doc map: model calls & Base URL, API Key & billing, agents/workflows/RAG, fine-tuning, MCP & Claude Code/OpenClaw. This documentation applies to the Alibaba Cloud China site. For international site documentation, please refer to https://www.alibabacloud.com/help/llms.txt ## Model User Guide: API Calls, Billing, Fine-tuning & Deployment - [Model User Guide: API Calls, Billing, Fine-tuning & Deployment](https://help.aliyun.com/en/model-studio/model-user-guide.md): Top hub for model docs: get started/first API call, billing and Token Plan, client setup (Claude Code and more), inference, fine-tuning, deployment, evaluation, usage statistics and monitoring. - [Get Started: First API Call to Qwen, Model Selection, Base URLs](https://help.aliyun.com/en/model-studio/get-started-with-models.md): New-user entry: product intro, first API call to Qwen (get an API Key), recommended models, dynamic rate limiting, Base URL overview, regions and access domains. Start here for your first request. - [Model Studio Overview: API, Agents, RAG, Fine-tuning & Billing](https://help.aliyun.com/en/model-studio/what-is-model-studio.md): Alibaba Cloud Model Studio is an LLM development platform with OpenAI-compatible APIs for qwen3.8-max, DeepSeek, Kimi and more. Build agents and workflows visually, use RAG knowledge bases, MCP plugins, SFT/DPO fine-tuning, deployment and evaluation. Free to activate; pay per token with a free quota in Beijing region. Coding Plan offers a fixed monthly AI coding subscription. - [First API Call to Qwen (Activate Model Studio, Get API Key)](https://help.aliyun.com/en/model-studio/first-api-call-to-qwen.md): First Qwen API call: activate Model Studio, get an API Key, set DASHSCOPE_API_KEY, call qwen via OpenAI-compatible or DashScope SDK; some regions need WorkspaceId in Base URL. - [Model Selection: Qwen, DeepSeek, Kimi, Image/Video, Audio, Embedding](https://help.aliyun.com/en/model-studio/models.md): Choose models by modality: text generation (qwen3.8-max, qwen3.7-plus, deepseek-v4-pro, kimi-k3, glm-5.2), image/video understanding and generation (qwen-image, wan2.7, happyhorse), ASR/TTS, omni-modal qwen3.5-omni-plus, embeddings text-embedding-v4 and qwen3-rerank. OpenAI/Anthropic/DashScope compatible with multi-region Base URLs and API Key setup. - [Dynamic Rate Limiting - Monthly TPM Quota Adjustment by Spend Tier](https://help.aliyun.com/en/model-studio/quota-management.md): Bailian dynamically adjusts TPM limits for models like qwen3.8-max across three monthly spend tiers (≤100K/100K-1M/>1M CNY). Soft limiting ensures actual TPM never falls below the quota. Limits apply per account+model, with new tiers effective on the 15th of each month. - [Model Studio Rate Limiting (RPM/TPM) and Auto-Recovery](https://help.aliyun.com/en/model-studio/rate-limit.md): Per-account RPM/TPM limits for qwen, DeepSeek, Kimi and other models; excess returns 429 and usually recovers within one minute. Includes per-model regional quotas, trigger diagnosis, temporary quota increase, and Batch API exemption. - [Base URL Overview](https://help.aliyun.com/en/model-studio/base-url.md): Overview of Alibaba Cloud Model Studio API base URLs. Pay-as-you-go supports Dashscope, workspace-specific, and trial domains. Token Plan is restricted to interactive AI tools like Claude Code and Codex with a dedicated API key. Base URL must match the billing scheme of the API key or a 401 error occurs. - [Model Studio - Select Region, Service Deployment Scope, and Access Domain](https://help.aliyun.com/en/model-studio/regions.md): Before calling Model Studio, select a region (determines data storage location), service deployment scope (inference scheduling boundary), and access domain. The workspace-specific domain supports HTTP/SSE/WebSocket/WebRTC protocols with 3600s timeout and 99.9% SLA, recommended for production; Dashscope domain is for legacy compatibility; trial domain has 1000 RPM limit for validation only. API Keys and model lists are region-specific and cannot be used across regions. - [Model Studio Billing and Pricing (Token Pricing/Free Quota)](https://help.aliyun.com/en/model-studio/test-1.md): How Model Studio charges: free quota for new users, per-token pricing for model calls, training and deployment billing, Savings Plans and resource packages, bill query and cost management. - [Model Studio New-User Free Quota (90-Day Validity/Quota Check/Auto-Stop)](https://help.aliyun.com/en/model-studio/new-free-quota.md): Auto-granted on first activation, Beijing region only, valid 90 days, 1M tokens per model. Quota check, 403 AllocationQuota.FreeTierOnly on exhaustion, auto-stop mode. - [Model Studio Model Call Pricing (Token Price/Tiered Pricing/Free Quota)](https://help.aliyun.com/en/model-studio/model-pricing.md): Pay-per-token pricing for Model Studio models: input/output rates for qwen-max/plus/flash/turbo/vl/audio/omni, tiered pricing rules, Batch API 50% off and context cache discounts, 1M-token free quota for new users (valid 90 days), with original prices listed by region (Beijing/Singapore/US/Germany/Japan). - [Model Studio Model Training & Deployment Billing (Token Prices/TPU/MU)](https://help.aliyun.com/en/model-studio/model-training-and-deployment-billing.md): Training fee = total tokens × epochs × unit price: qwen3.7-plus ¥0.35, qwen3-vl-8b ¥0.012, wan2.7-i2v ¥2, CosyVoice ¥0.2 per 1K tokens. Deployment: preset TPU, Model Units (MU), or per-token; Beijing/Singapore, overflow and overdue rules. - [Model Studio Savings Plans & Resource Packages (AI Universal/Model-Specific/Token Plan)](https://help.aliyun.com/en/model-studio/savings-plan-and-resource-package.md): Cost optimization options for Model Studio: AI Universal Savings Plan (commit monthly spend for tiered discounts up to 53% off, covers all Alibaba-supplied models), model-specific savings plans (e.g., speech series), and resource packages (fixed token quota for a single model). Compares with Token Plan for team use, covering deduction order, effective dates, and exclusions. - [Bill query and cost management](https://help.aliyun.com/en/model-studio/bill-query-and-cost-management.md): Query per-model inference costs, analyze bills by API key/workspace/channel, tag workspaces for cost allocation, handle overdue payments, and stop billing by deleting API keys or unsubscribing plans. - [Token Plan Guide (Credits, Team Edition, Coding Plan)](https://help.aliyun.com/en/model-studio/token-plan-guide.md): Hub for Token Plan subscriptions: overview and Credits metering, Team Edition seats, best practices, Coding Plan. Start here to subscribe or buy Credits. - [Token Plan Overview (Personal & Team Plans with Credits Quota)](https://help.aliyun.com/en/model-studio/token-plan-overview.md): Alibaba Cloud Model Studio Token Plan subscription: unified Credits billing for AI coding tools like Claude Code, Cursor, Qwen Code, Codex, and OpenClaw. Personal tier offers Lite/Standard/Pro plans (7-day quota 2,500–40,000 Credits); Team tier provides Standard/Pro/Max seats (monthly 25,000–250,000 Credits). Available only in China (Beijing) region; uses sk-sp- dedicated API Key for deduction. Self-built tools without OpenAI/Anthropic protocol compatibility should use sk-/sk-ws- pay-as-you-go keys instead. - [Token Plan Personal Edition (Overview/Quick Start/FAQ)](https://help.aliyun.com/en/model-studio/token-plan-personal.md): Hub for Token Plan personal edition: overview, quick start, FAQ. Subscribe with Credits-based plans to call qwen models and connect clients like Claude Code/OpenClaw. - [Token Plan Personal Edition - Overview](https://help.aliyun.com/en/model-studio/token-plan-personal-overview.md): AI model subscription for individual developers, billed in Credits and compatible with Claude Code, Cursor, Qwen Code, OpenClaw, and other tools. Available only in China (Beijing). Offers Lite, Standard, and Pro tiers plus add-on packs, covering text, vision, speech, and video generation models such as qwen3.8-max, deepseek-v4-pro, and wan2.7-image, along with Harness tools like web search. - [Token Plan Personal Edition - Quick Start](https://help.aliyun.com/en/model-studio/token-plan-personal-quick-start.md): Subscribe to Token Plan Personal Edition in three steps: choose a plan, obtain an sk-sp- prefixed API Key, configure the OpenAI/Anthropic-compatible Base URL, and connect AI tools such as Claude Code, OpenClaw, Cursor, Codex, Qwen Code, and Kilo CLI. RAM users require the AliyunTokenPlanFullAccess policy granted by the primary account. - [Token Plan Personal Edition FAQ](https://help.aliyun.com/en/model-studio/token-plan-personal-faq.md): Answers questions on Token Plan Personal Edition quotas, purchase, subscription, and integration. Covers Lite/Standard/Pro Credits limits (2,500–40,000 per 7 days), add-on pack rules, API Key/Base URL configuration, 401/404/429 error troubleshooting, Agent concurrency, and the prohibition of automated production calls. - [Token Plan Team Edition (Overview/Quick Start/Team Management/FAQ)](https://help.aliyun.com/en/model-studio/token-plan-team-edition.md): Hub for Token Plan Team Edition: overview, quick start (subscribe and first use), team management (seats/members), FAQ. Entry point for team-wide Token Plan subscriptions. - [Token Plan Team Edition - Overview](https://help.aliyun.com/en/model-studio/token-plan-team-overview.md): Alibaba Cloud Model Studio Token Plan Team Edition uses Credits as unified billing, supporting models like qwen3.8-max, deepseek-v4-pro, kimi-k2.6, MiniMax-M2.5, and happyhorse-1.1, compatible with Claude Code, OpenClaw, Qwen Code, and Codex. Offers Standard/Advanced/Premium seats (limited-time ¥150/¥550/¥1,398/month) with team management console and shared usage packages, available only in China North 2 (Beijing). - [Token Plan quick start](https://help.aliyun.com/en/model-studio/token-plan-team-quickstart.md): Three-step setup: subscribe, get dedicated API key (sk-sp-xxx) and Base URL, configure AI tool. OpenAI and Anthropic-compatible endpoints. Optional: tool calling and image generation setup - [Token Plan Team Management (Members, Seats, SSO & DingTalk)](https://help.aliyun.com/en/model-studio/token-plan-team-management.md): Manage Token Plan team: add members and assign seats, reset API Keys, configure SAML SSO and DingTalk integration, purchase/upgrade/reclaim seats, analyze Credits usage. RAM users need AliyunTokenPlanFullAccess plus Bailian admin permission. Removing a member who has used Credits does not reclaim the seat. - [Token Plan Team FAQ (Credits/Seats/Errors/Renewal)](https://help.aliyun.com/en/model-studio/token-plan-team-faq.md): Token Plan Team Edition FAQ: personal vs team differences, Standard/Premium/Ultimate seat Credits (25k/100k/250k), tool integration (Cursor/Claude Code/Qwen Code/OpenClaw), 401/404 error troubleshooting, RAM sub-account authorization, seat upgrade/add-on and auto-renewal cancellation. - [Token Plan Best Practices (Harness, Multimodal, Web Search MCP, Vision)](https://help.aliyun.com/en/model-studio/token-plan-best-practice.md): Token Plan/Coding Plan advanced usage: connect Harness tools, use multimodal generation models, add web-search MCP, add visual understanding to your coding assistant. - [Token Plan Harness Tools (Web Search/Code Interpreter/Web Fetch)](https://help.aliyun.com/en/model-studio/token-plan-harness-tool.md): Token Plan exclusive: qwen3.8-max/qwen3.7-max/qwen3.7-plus include built-in Harness tools—web search, code interpreter, web fetch, image-to-image search, and text-to-image search—billed per successful call against plan Credits. Triggered automatically via Responses API; clients supporting only OpenAI Chat Completions will not invoke these tools. - [Token Plan: Integrate Multimodal Models (Image/Video/TTS)](https://help.aliyun.com/en/model-studio/token-plan-multimodal-gen.md): Use Slash Commands/Skills in Claude Code or Codex to call Token Plan multimodal models: qwen-image-2.0 for images, happyhorse for video, qwen-audio for TTS; sk-sp- key auth. - [Coding Plan web search MCP](https://help.aliyun.com/en/model-studio/web-search-for-coding-plan.md): Enable web search MCP for Claude Code / Qwen Code via Streamable HTTP. Activate in MCP Marketplace, 2000 free calls/month. Endpoint: dashscope.aliyuncs.com/api/v1/mcps/WebSearch/mcp. - [Token Plan Vision Understanding (qwen3.7-plus / Skill Setup / Claude Code & OpenCode)](https://help.aliyun.com/en/model-studio/add-vision-skill.md): Enable image understanding under Token Plan: qwen3.8-max, qwen3.7-plus, and kimi-k2.5 support vision natively; text-only models like glm-5 and MiniMax-M2.5 gain vision via a local Skill or Agent backed by qwen3.7-plus. Includes setup examples for Claude Code, OpenCode, and Qwen Code, plus the OpenCode modalities parameter. Running the Skill consumes Token Plan Credits. - [Coding Plan Subscription & Usage Guide (Overview/AI Tools/Lobster)](https://help.aliyun.com/en/model-studio/coding-plan-guide.md): Coding Plan hub: plan overview & pricing, connect Claude Code and other AI tools, Bailian Lobster (cloud OpenClaw, one-click deploy, Pro plan), best practices and FAQ. - [Coding Plan Subscription and Usage (Pro Plan/API Key/Tool Integration)](https://help.aliyun.com/en/model-studio/coding-plan.md): Alibaba Cloud Model Studio Coding Plan Pro at CNY 200/month supports qwen3.7-plus, kimi-k2.5, glm-5, MiniMax-M2.5, and more. Limits: 6,000 requests per 5 hours, 45,000 per week, 90,000 per month. Get an sk-sp- prefixed API Key and coding.dashscope.aliyuncs.com Base URL to connect Claude Code, OpenClaw, Qwen Code, Cursor, Codex, and other AI coding tools. Interactive coding tools only; script or backend API calls are prohibited. No refunds. - [Coding Plan FAQ (Errors 401/Quota/Renewal)](https://help.aliyun.com/en/model-studio/coding-plan-faq.md): Coding Plan troubleshooting: 401/auth errors, Claude Code/OpenCode/Qwen Code/OpenClaw setup, thinking_budget errors, unexpected charges after subscription, exhausted quota, auto-renewal failures. - [Model Experience: Try Qwen Models Online in the Console](https://help.aliyun.com/en/model-studio/model-experience.md): Try qwen and other models online in the Model Studio console: chat and compare model outputs with no code required — test before API integration and model selection. - [Text Generation Model Selection Guide](https://help.aliyun.com/en/model-studio/text-generation-model.md): Bailian text generation models by scenario: qwen3.7-plus or qwen3.8-max for AI coding and Agent development; qwen3.7-flash for office tasks to reduce cost. Supports 1M context, thinking mode, Function Calling, structured output, and batch inference for lower request costs. - [Text Generation API (Chat Completions / Responses / DashScope SDK)](https://help.aliyun.com/en/model-studio/text-generation.md): Alibaba Cloud Model Studio text generation guide: OpenAI-compatible Chat Completions and Responses APIs, DashScope SDK samples in Python/Java/Node.js/Go/C#/PHP/curl for qwen3.8-max/qwen-plus, System/User/Assistant message construction, multimodal image/video handling, async calls, temperature/top_p tuning, and timeout continuation. - [Multi-turn conversations](https://help.aliyun.com/en/model-studio/multi-round-conversation.md): Implement multi-turn chat by appending user/assistant messages to the messages array each request. Stateless API. Context management via truncation or summarization - [Streaming output](https://help.aliyun.com/en/model-studio/stream.md): Enable token-by-token SSE streaming via stream=true. Set stream_options={"include_usage": true} for token counts. Required for Qwen3 open-source, QwQ, QVQ, Qwen-Omni. Same billing as non-streaming - [Deep Thinking Model Invocation (enable_thinking/thinking_budget/reasoning_content)](https://help.aliyun.com/en/model-studio/deep-thinking.md): Call Qwen, DeepSeek, Kimi, GLM, MiniMax and other deep thinking models via OpenAI-compatible or DashScope API: enable_thinking toggle, thinking_budget to cap reasoning tokens, reasoning_content for thought process, preserve_thinking for multi-turn. Covers hybrid vs thinking-only modes, streaming and non-streaming examples, billing notes, and slow-response troubleshooting. - [Model Studio Structured Output (JSON Object & JSON Schema via response_format)](https://help.aliyun.com/en/model-studio/qwen-structured-output.md): Use the response_format parameter to get valid JSON from LLMs: JSON Object mode ensures parseable output; JSON Schema mode enforces a strict structure. Supported models include Qwen, Kimi, DeepSeek, GLM, and Stepfun. Covers two-step fix for thinking models, Pydantic/Zod SDK integration, and production validation tips. - [Partial mode](https://help.aliyun.com/en/model-studio/partial-mode.md): Continue from a prefix by setting last message role to assistant with partial: true. Use case: code completion. Supports Qwen-Max/Plus/Flash/Coder and DeepSeek in non-thinking mode. - [Context cache](https://help.aliyun.com/en/model-studio/context-cache.md): Explicit cache (cache_control marker, 10% input price on hit, 5-min TTL, min 1024 tokens) and implicit cache (automatic, 20% price, min 256 tokens). OpenAI/DashScope/Anthropic APIs. - [Batch Inference: Async Tasks at 50% Cost](https://help.aliyun.com/en/model-studio/batch-inference.md): Model Studio batch inference: upload a JSONL file to asynchronously process large-scale requests at 50% of real-time pricing, compatible with the OpenAI Batch API. Supports qwen3.8-max, qwen-plus, deepseek-r1, and more. Max 50,000 requests or 500 MB per file; results auto-deleted after 30 days. Ideal for model evaluation and data labeling. - [Model Studio Tool Calls (Function Calling/Web Search/MCP)](https://help.aliyun.com/en/model-studio/tool-calls.md): Model Studio tool hub: Function Calling, web search, code interpreter, file search, MCP, PDF understanding — add live data and external tools to model calls. - [Model Studio Function Calling (Tool Use, Parallel Calls, Streaming)](https://help.aliyun.com/en/model-studio/qwen-function-calling.md): Enable models to call external APIs, databases, or custom functions: workflow, tools parameter definition, OpenAI/DashScope/Responses API examples, with parallel tool calls, forced tool_choice, streaming output, multi-turn conversations, and tool use for Qwen-Omni/Realtime and thinking models. - [Model Studio Web Search (enable_search / web_search & Search Strategies)](https://help.aliyun.com/en/model-studio/web-search.md): Enable real-time web search for model calls: use enable_search in Chat Completions/DashScope or the web_search tool in Responses API. Supports four strategies (turbo/max/agent/agent_max), forced search, vertical sources (stocks/weather/FX and 9 more), freshness filter, site allowlist, citation markers, text-image mixed output, and reasoning-model integration. - [Web Extractor Tool (web_extractor) Usage and Billing](https://help.aliyun.com/en/model-studio/web-extractor.md): Alibaba Cloud Model Studio web_extractor lets models fetch and extract content from URLs. Three access methods: Responses API, Chat Completions, DashScope; requires enabling web_search and enable_thinking. Supports qwen3.8-max, qwen3-max, deepseek-v4-flash. Web extraction is currently free; web search costs CNY 4 per 1,000 calls in Beijing. - [Built-in Tools – Code Interpreter (code_interpreter/enable_code_interpreter)](https://help.aliyun.com/en/model-studio/qwen-code-interpreter.md): Run Python in a sandbox for math and data analysis: code_interpreter tool via Responses API, enable_code_interpreter via Chat Completions/DashScope (thinking mode + streaming required). Best on Qwen3.8-Max/Qwen3.7-Plus/deepseek-v4-flash - [Web Search Image (web_search_image) via Responses API and Billing](https://help.aliyun.com/en/model-studio/web-search-image.md): Model Studio Responses API tool web_search_image: search internet images by text prompt and reason over image content for visual QA and image recommendations. Recommended models include Qwen3.7-Plus and Qwen3.8-Max; Responses API only. Tool calls cost CNY 24 per 1,000 in Beijing and CNY 58.71 in Singapore; image results count toward input tokens. - [Image Search Tool (image_search): Find Visually Similar Images by Input Image](https://help.aliyun.com/en/model-studio/image-search.md): image_search tool via Responses API: pass an image URL with input_image to search the web for visually similar images and reason over results (find-same-item, visual sourcing). Recommended: qwen3.8-max, Qwen3.7/3.6/3.5-Plus; Responses API only. - [File Search: Knowledge Base Retrieval via Responses API](https://help.aliyun.com/en/model-studio/file-search.md): Use the file_search tool in Responses API to retrieve private content from Model Studio knowledge bases for LLM answers. Supports Qwen3.8-Max/Plus/Flash and Qwen3.6/3.5 open-source series. Requires a pre-created knowledge base and one vector_store_ids. Includes Python/Node.js/curl examples and streaming output. - [Model Studio MCP Integration (Responses API, Supported Models, Streaming)](https://help.aliyun.com/en/model-studio/mcp.md): Configure MCP servers via the Responses API tools parameter (SSE protocol, up to 10 servers) so LLMs can call external tools and data. Supports Qwen3.8-Max/Plus/Flash and Qwen3.6/3.5 open-source series; includes Python/Node.js/curl examples, streaming output, and parameters like server_label, server_url, and headers. Billing covers model inference tokens plus MCP service fees. - [PDF Understanding - qwen3.8-max PDF Parsing](https://help.aliyun.com/en/model-studio/pdf-understanding.md): Pass PDF files via URL or Base64 through OpenAI-compatible or DashScope API for qwen3.8-max to extract and analyze text and images. Available only in China (Beijing); Responses API not supported. Max 150 MB / 500 pages per file; first-token timeout up to 300s. Billing includes model input tokens plus document parsing at CNY 0.02/page. - [Qwen Specialized Models (Long Context/Coding/Translation/Deep Research)](https://help.aliyun.com/en/model-studio/specialized-models.md): Qwen specialized models: Qwen-Long (long context), Qwen-Coder (coding), Qwen-MT (translation), Qwen-Character (role-play), Qwen-Deep-Research and 5 more, with calling guides. - [Qwen-Long (10M token context)](https://help.aliyun.com/en/model-studio/long-context-qwen-long.md): Process documents up to 10M tokens via file upload + file-id reference mechanism. Upload files once, reference by file-id in system message. File tokens counted as input per API call - [Qwen-Coder Code Models (qwen3-coder-next, Code Generation & Completion)](https://help.aliyun.com/en/model-studio/qwen-coder.md): Qwen-Coder (qwen3-coder-next) for code generation/completion via OpenAI-compatible API or DashScope SDK, tool calling supported. Latest general-purpose models are recommended instead. - [Qwen-MT translation model](https://help.aliyun.com/en/model-studio/machine-translation.md): Machine translation across 92 languages. Models: qwen-mt-plus (quality), qwen-mt-flash (general), qwen-mt-lite (speed). Supports term intervention and translation_options param - [Qwen-Character role-playing models](https://help.aliyun.com/en/model-studio/role-play.md): Character-consistent conversation: qwen-plus-character (32K context), qwen-flash-character (8K). Configure persona via system message. Supports session cache and opening remarks - [Data mining with Qwen-Doc-Turbo](https://help.aliyun.com/en/model-studio/data-mining-qwen-doc.md): Extract structured JSON from documents via file URL (up to 10 files, 253k tokens), file ID, or plain text (9k tokens). Also generates PPTs in template or creative mode. DashScope-only for file URLs. - [Qwen-Deep-Research](https://help.aliyun.com/en/model-studio/qwen-deep-research.md): Automated research agent: follow-up questions then deep search and report. Call qwen-deep-research via DashScope Python SDK (streaming only). Beijing region only, no Java/OpenAI API. - [Qwen-Math Models (qwen-math-plus/turbo Pricing & API Calls)](https://help.aliyun.com/en/model-studio/math-language-model.md): Alibaba Cloud Model Studio Qwen-Math series for math reasoning: qwen-math-plus and qwen-math-turbo, input ¥2–4 per million tokens, output ¥6–12 per million tokens, 4096 context, 1M free tokens each (valid 90 days after activation). Available only in China (Beijing); latest general models recommended as replacement. - [Tongyi Xiaomi Conversation Analysis](https://help.aliyun.com/en/model-studio/dialogue-analysis.md): Analyze conversations: info extraction, classification, satisfaction mining, quality inspection. Models: analysis-flash (low latency) and analysis-pro (complex reasoning). Chinese only. - [GUI-Plus GUI Agent Model (computer_use/mobile_use, Beijing Only)](https://help.aliyun.com/en/model-studio/gui-automation.md): GUI-Plus converts screenshots and instructions into GUI actions (click/type/scroll) via computer_use/mobile_use tools. Beijing-region API Key only; gui-plus-2026-02-26 adds thinking mode. - [Qwen3-Omni-Captioner audio understanding](https://help.aliyun.com/en/model-studio/qwen3-omni-captioner.md): Generate detailed captions for speech, music, and environmental sounds without prompts. Model: qwen3-omni-30b-a3b-captioner. Identifies emotions, instruments, sensitive content. OpenAI-compatible API - [Vision Understanding Model Selection](https://help.aliyun.com/en/model-studio/vision-model.md): Alibaba Cloud Model Studio vision models for image analysis, video understanding, and OCR. qwen3.8-max/qwen3.7-plus support 1M context, 2-hour video, Function Calling, and built-in tools; qwen3.5-ocr is optimized for document/table/handwriting extraction. Includes GPT/Claude/Gemini migration mapping and model spec comparison. - [Model Studio Vision API: Image/Video QA, OCR & Object Detection](https://help.aliyun.com/en/model-studio/vision.md): Call qwen3.8-max/qwen3-vl vision models via OpenAI-compatible or DashScope SDK with single/multi-image or video input for image captioning, visual QA, OCR, 2D/3D object detection, document parsing and video event timestamping; includes enable_thinking mode, fps sampling, vl_high_resolution_images, and local file upload via Base64 or file path. - [Qwen-OCR text extraction](https://help.aliyun.com/en/model-studio/qwen-vl-ocr.md): Extract text from images via qwen-vl-ocr. Supports documents, tables, receipts, formulas with multi-language OCR, text localization, and skew correction. 38K context, OpenAI and DashScope SDK - [Visual Reasoning Models (Qwen3-VL/QVQ/Kimi/MiniMax Thinking Mode)](https://help.aliyun.com/en/model-studio/visual-reasoning.md): Model Studio visual reasoning: qwen3.8-max, qwen3-vl-plus, qvq-max, kimi-k2.6, MiniMax-M3 and other hybrid/thinking-only models; enable_thinking toggle, streaming reasoning_content for math solving, chart analysis, and complex video understanding. - [Model Studio - Image Generation and Editing Model Selection](https://help.aliyun.com/en/model-studio/image-model.md): Model comparison for text-to-image and image editing: qwen-image-3.0-pro handles complex layouts and small-text rendering; wan2.7-image-pro supports up to 4096x4096 resolution and multi-image reference; z-image-turbo is 10x faster at ~1/5 cost for realistic portraits. Includes migration guide from Nano Banana, GPT Image, Midjourney, and Seedream. - [Model Studio Text-to-Image API (Wan/Qwen-Image/z-image)](https://help.aliyun.com/en/model-studio/text-to-image.md): Generate images from text prompts using Wan, Qwen-Image, and z-image models. Covers e-commerce banners, UI design, PPT visuals, and illustrations. Includes online demo links, prompt examples, and model output comparisons. - [Image Editing Guide (Qwen / Wan 2.7 / Wan 2.1)](https://help.aliyun.com/en/model-studio/image-edit-guide.md): Entry to image editing docs by model: Qwen Image Edit, Wan 2.7/2.6/2.5, and Wan 2.1. Pick your model's tutorial for API usage and parameters. - [Qwen Image Editing (qwen-image-3.0-pro): Multi-Image Edit Guide](https://help.aliyun.com/en/model-studio/qwen-image-edit-guide.md): qwen-image-3.0-pro image editing: multi-image I/O, edit text, objects, style in images. DashScope SDK (Python/Java) and curl examples, prompt_extend, Base64/URL input. - [Wanxiang image editing guide (2.5-2.7)](https://help.aliyun.com/en/model-studio/wan-image-edit.md): Edit images with wan2.7-image-pro, wan2.6-image, or wan2.5-i2i-preview. Multi-image fusion, subject preservation, bbox interactive editing, detection, and segmentation. Up to 2K output. - [Wanx 2.1 general image editing guide](https://help.aliyun.com/en/model-studio/wanx-image-edit.md): Edit images with wanx2.1-imageedit via text instructions. Outpainting, watermark removal, style transfer, inpainting, image restoration. Beijing region only, CNY 0.14/image. - [Portrait style repainting](https://help.aliyun.com/en/model-studio/style-repaint.md): Transform portraits into artistic styles via wanx-style-repaint-v1. Preset styles (Retro Comic, Anime, 3D) or custom style via reference image. CNY 0.12/image, HTTP-only async API - [Wanx image background generation](https://help.aliyun.com/en/model-studio/image-background-generation.md): Replace product backgrounds via text, reference image, or both. Model: wanx-background-generation-v2. Supports edge-guided foreground/background blending. Async HTTP API, CNY 0.08/image - [Image outpainting](https://help.aliyun.com/en/model-studio/image-expansion.md): Expand images by aspect ratio, scale factor, or custom pixel padding. Model: image-out-painting. Supports rotation before expansion. Async HTTP-only API, CNY 0.18/image - [Virtual model generation (Wanx)](https://help.aliyun.com/en/model-studio/virtual-model-generation.md): Replace background and human model in product photos while keeping pose. V1 wanx-virtualmodel (512/1024px), V2 virtualmodel-v2 (up to 2048px, background ref). Free trial, 500 images - [Shoe Model Generation (shoemodel-v1 AI Try-On)](https://help.aliyun.com/en/model-studio/shoes-and-boots-model.md): AI shoe try-on: input model template + multi-angle shoe images to generate product photos. Model shoemodel-v1, Beijing region only, free trial 500 images. - [Creative poster generation overview](https://help.aliyun.com/en/model-studio/creative-poster-generation-overview.md): Generate styled posters with wanx-poster-generation-v1 (free trial, 500 items). Supports lora_name styles (Paper-cut, Embroidery, Oil Painting), horizontal/vertical, max 50-char prompt. - [Image Instance Segmentation (image-instance-segmentation, Pixel Masks)](https://help.aliyun.com/en/model-studio/image-instance-segmentation.md): image-instance-segmentation: detects different people in an image and outputs a pixel-level mask per person; pairs with image erasure for person removal. Beijing-region API Key only; free trial with 500-image quota. - [Image Erasure & Inpainting (Mask-Based Removal of People/Watermarks)](https://help.aliyun.com/en/model-studio/image-erasure-and-completion.md): image-erase-completion: erase people/watermarks/objects via mask + background inpainting, no prompt. Beijing API Key only, 500 free images; alt: Qwen/Wanx 2.1 image edit - [Image inpainting (wanx-x-painting)](https://help.aliyun.com/en/model-studio/vary-region.md): Edit masked regions of an image with original image + mask + text prompt. Free trial (500 images). Async DashScope API. Max resolution 4096x4096. Formats: JPG/PNG/BMP/TIFF/WEBP - [Wanx Sketch-to-Image: Hand-Drawn Sketch + Text Prompt to Image](https://help.aliyun.com/en/model-studio/sketch-to-image.md): wanx-sketch-to-image-lite: sketch + text prompt to image, flat illustration/oil painting/anime/3D cartoon/watercolor styles, ¥0.06/image (500 free), Beijing-region API Key required - [Video Generation and Editing Model Selection](https://help.aliyun.com/en/model-studio/video-generate-edit-model.md): Model selection guide for Alibaba Cloud Model Studio video generation, covering text-to-video, image-to-video, reference-to-video, video editing, and character animation. Recommends happyhorse-1.1-t2v/i2v/r2v, wan2.7-t2v/i2v/r2v/videoedit, wan2.2-animate-move/mix, supporting 480P/720P/1080P resolution with up to 15-second clips, including custom audio, first-last-frame continuation, and effect replication capabilities. - [Model Studio Text-to-Video API (wan3.0-video, Up to 30s)](https://help.aliyun.com/en/model-studio/text-to-video-guide.md): wan3.0-video text-to-video: text/image/audio multimodal input, up to 30s, adaptive aspect ratio, 480P/720P/1080P, multi-shot narrative with audio, Python SDK sample (DashScope ≥1.25.16). - [Wan 2.7 image-to-video guide](https://help.aliyun.com/en/model-studio/wan-image-to-video-guide.md): Generate 2-15s video from images/audio/video with wan2.7-i2v. Three modes: first-frame, first-and-last-frame, video continuation. 720P/1080P, multi-shot narrative, driving_audio for lip-sync. - [Image-to-video: first and last frames](https://help.aliyun.com/en/model-studio/image-to-video-first-and-last-frames-guide.md): Generate 5-second videos from first/last frame images and optional prompt. Model: wan2.2-kf2v-flash. Supports 480P/720P/1080P resolution, prompt rewriting, and video effect templates - [Image-to-Video from First Frame (Wan i2v/Audio/Multi-shot)](https://help.aliyun.com/en/model-studio/image-to-video-guide.md): Wan i2v: turn a first-frame image + prompt into 2–15s video at 480P/720P/1080P; auto dubbing on wan2.5/2.6, multi-shot narrative on wan2.6; Python/Java SDK examples with img_url/resolution/duration. - [Reference-to-video (Wan-R2V)](https://help.aliyun.com/en/model-studio/video-to-video-guide.md): Cast characters from reference images/videos using Image N/Video N identifiers in prompts. Wan 2.7 supports voice cloning and dialogue. DashScope VideoSynthesis SDK (Python/Java/curl) - [Video Editing Models: wan Video Editing 2.7 / 2.1 Guides](https://help.aliyun.com/en/model-studio/wan-video-editing.md): Hub for Wan video editing model docs: usage guides for wan Video Editing 2.7 (wan-video-editing) and wan Video Editing 2.1 (wan-vace), for instruction-based editing and re-creation of existing videos. - [Wan 2.7 video editing guide](https://help.aliyun.com/en/model-studio/wan-video-editing-guide.md): Edit video with wan2.7-videoedit model. Two modes: instruction-based editing and instruction+reference image. Replicate actions, effects, camera movements. Style transfer, background replacement. - [Wan 2.1 VACE video editing guide](https://help.aliyun.com/en/model-studio/wan-vace-guide.md): Edit video with wanx2.1-vace-plus. Five functions: image_reference, video_repainting (pose/depth/scribble), video_edit (mask), video_extension, video_outpainting. 720P, max 5s. - [Tripo 3D model generation guide](https://help.aliyun.com/en/model-studio/tripo-3d-generation-guide.md): Generate GLB 3D models via text, single-image, or multi-image. Tripo-P1.0 (fast, 20K polygons) for previews, Tripo-H3.1 (2M polygons) for film quality. Async API, 15s polling - [TTS Model Selection](https://help.aliyun.com/en/model-studio/tts-model.md): Alibaba Cloud Model Studio TTS supports standard synthesis, voice cloning, and voice design. Recommended models include qwen-audio-3.0-tts-plus/flash, cosyvoice-v3.5-plus/flash, and MiniMax/speech-2.8-hd, with WebSocket/HTTP/AOQ access and instruction control for scenarios like smart customer service, audiobooks, and virtual streamers. - [Real-time Text-to-Speech](https://help.aliyun.com/en/model-studio/realtime-tts-user-guide.md): Alibaba Cloud Model Studio real-time TTS supports streaming I/O with low first-packet latency. Compatible with PCM, WAV, MP3, and Opus formats up to 48kHz. Features CosyVoice, Qwen-TTS, and Sambert models with voice cloning, instruction control, and emotion tags for voice assistants and smart customer service. - [Qwen speech synthesis (TTS)](https://help.aliyun.com/en/model-studio/non-realtime-tts-user-guide.md): Non-realtime TTS via qwen3-tts-flash (49 voices, 10 languages, CNY 0.8/10k chars) or qwen-tts (token-based). DashScope SDK/RESTful API. Output: 24 kHz wav. Key params: text, voice, language_type. - [Voice Cloning - Create Custom Voices and Speech Synthesis on Bailian](https://help.aliyun.com/en/model-studio/voice-cloning-user-guide.md): Create custom voices from 10-20s audio samples on Alibaba Cloud Model Studio (Bailian). Supports Qwen-Audio-TTS, CosyVoice, MiniMax, and Qwen-TTS model series. Covers voice enrollment API calls, audio requirements (WAV/MP3/M4A, ≤10MB, ≥16kHz), region availability (China North 2 Beijing / Singapore), and billing rules (MiniMax charges CNY 9.9 unlock fee on first synthesis; Qwen-TTS billed at CNY 0.01 per voice). - [Voice Design - Speech Synthesis](https://help.aliyun.com/en/model-studio/voice-design-user-guide.md): Create custom voice timbres from natural language descriptions without audio samples. Supports CosyVoice, Qwen-Audio-TTS, and Qwen-TTS model series. CosyVoice is available only in Beijing region; Qwen-TTS supports up to 2048-character voice descriptions and is available in Beijing and Singapore regions. Workflow: write a voice description, call the voice-enrollment or qwen-voice-design API to create a voice, then use the returned voice_id for speech synthesis. - [CosyVoice SSML Speech Control and LaTeX Formula Reading](https://help.aliyun.com/en/model-studio/ssml-latex-user-guide.md): Control CosyVoice rate, pauses and pronunciation with SSML tags (say-as/phoneme), and read LaTeX formulas as speech. Only cosyvoice-v3.5/v3/v2; set enable_ssml=true. - [TTS Voice List (CosyVoice/Qwen-TTS/Qwen-Audio-TTS)](https://help.aliyun.com/en/model-studio/tts-voice-list.md): Voice hub for Model Studio TTS: voice catalogs of CosyVoice, Qwen-TTS and Qwen-Audio-TTS models, browse voices per model for speech synthesis, audiobooks and voiceover. - [Qwen-Audio-TTS Voice List](https://help.aliyun.com/en/model-studio/qwen-audio-tts-voice-list.md): System voices (longanlingxin, loongmary, etc.) and 500+ cloned base voices for qwen-audio-3.0-tts-plus and qwen-audio-3.0-tts-flash, with voice parameters, scenarios, language support, InvalidParameter troubleshooting, and audio samples - [CosyVoice voice list](https://help.aliyun.com/en/model-studio/cosyvoice-voice-list.md): System voice catalog for CosyVoice models (v3-plus/v3-flash/v2/v1). Lists supported model, languages, SSML, Instruct, and timestamp per voice. Custom voices via cloning also supported - [Qwen-TTS Voice List: voice Parameter Values and Compatible Models](https://help.aliyun.com/en/model-studio/qwen-tts-voice-list.md): Voice lookup table for realtime TTS: voice parameter values Cherry/Serena/Ethan/Chelsie/Momo with style, 10-language support and compatible models like qwen3-tts-instruct-flash-realtime and qwen-tts-realtime. - [Fun-Music song generation](https://help.aliyun.com/en/model-studio/fun-music.md): Generate full songs with vocals from prompts or custom lyrics. Model fun-music-v1, male/female voices, Chinese/English, MP3/WAV output. API endpoint: /services/audio/tts/SpeechSynthesizer - [ASR Model Selection](https://help.aliyun.com/en/model-studio/asr-model.md): Alibaba Cloud Model Studio ASR model selection guide: real-time streaming qwen-audio-3.0-asr-flash-streaming (WebSocket, hotwords/Prompt context, multilingual & dialects) and offline file transcription qwen-audio-3.0-asr-flash-filetrans (HTTP, speaker diarization, 12h/2GB). Supports emotion recognition, domain terminology handling, Mandarin/Cantonese dialects and 30+ languages including English, Japanese, and Korean. - [Real-time Speech Recognition](https://help.aliyun.com/en/model-studio/real-time-speech-recognition-user-guide.md): Alibaba Cloud Model Studio real-time speech recognition transcribes audio streams to punctuated text with low latency. Supports Mandarin, Cantonese, Sichuan dialect, emotion detection, hotword customization, context enhancement, and timestamp output for live captions, online meetings, and voice chat. - [Non-real-time Speech Recognition - Async Audio Transcription and Model Invocation](https://help.aliyun.com/en/model-studio/non-realtime-speech-recognition-user-guide.md): Alibaba Cloud Model Studio non-real-time ASR supports Qwen-Audio-3.0-ASR-Flash-Filetrans, Fun-ASR, Qwen3-ASR-Flash-Filetrans, and Paraformer models for async transcription up to 2GB or 12 hours per file; sync models are limited to 5 minutes/10MB. Features include speaker diarization, emotion recognition, hotwords, context enhancement, profanity filtering, and word-level timestamps for meeting transcription, subtitle generation, and call analysis. - [Speech Recognition - Improve Recognition Accuracy](https://help.aliyun.com/en/model-studio/improve-asr-accuracy.md): Alibaba Cloud Model Studio speech recognition improves terminology accuracy via precompiled hotwords, instant hotwords, and context enhancement. Precompiled hotwords reuse vocabularies across requests; instant hotwords apply per session (Qwen-Audio-3.0-ASR-Flash series only); context enhancement corrects results using dialogue history. Hotwords are unavailable in Singapore region; merged hotword limit is 2,000. - [Speech-to-speech model selection guide](https://help.aliyun.com/en/model-studio/s2s-model.md): Compare S2S vs ASR+LLM+TTS pipeline. Models: qwen3.5-omni-plus-realtime (WebSocket), qwen3-omni-flash (HTTP, thinking mode), qwen3-livetranslate-flash. Function calling, web search, 29+ languages - [Speech-to-Speech - Real-time Voice Chat (Qwen-Audio-Realtime)](https://help.aliyun.com/en/model-studio/fun-audiochat-realtime.md): Qwen-Audio end-to-end real-time voice interaction model supporting WebSocket, AOQ, and WebRTC protocols with server_vad, smart_turn, and push-to-talk modes. smart_turn combines acoustic and semantic detection to prevent filler-word interruptions. Supports Function Calling, context management, voice cloning, and speaker enhancement. Audio input at 16 kHz PCM, output at 24 kHz PCM, with up to 50 turns or 300 seconds of context. Designed for low-latency duplex scenarios such as voice assistants, intelligent customer service, and AI companions. - [qwen3.5-livetranslate-flash-realtime Real-Time Speech and Audio-Visual Translation](https://help.aliyun.com/en/model-studio/qwen3-5-livetranslate-flash-realtime.md): Vision-enhanced real-time simultaneous interpretation model supporting 60 languages with latency as low as 2.8s. Accessible via WebSocket, AOQ, or WebRTC. Features voice cloning, hotword configuration, and VAD/Manual modes. Billed at 7 tokens/s for audio input and 12.5 tokens/s for output. - [Qwen3-LiveTranslate file translation](https://help.aliyun.com/en/model-studio/qwen3-livetranslate-flash.md): Translate audio/video files across 18 languages via OpenAI-compatible streaming API. Returns text, audio, or both. Set source_lang and target_lang via translation_options in extra_body - [Omni Model Selection: Qwen3.5-Omni / Omni-Flash / Livetranslate](https://help.aliyun.com/en/model-studio/omni.md): Pick omni models: qwen3.5-omni-plus, qwen3-omni-flash, qwen3.5-livetranslate (60 languages, ~3s latency) for real-time voice/video chat, A/V analysis and voice cloning. - [Qwen-Omni-Realtime: Real-Time Audio/Video Chat (WebSocket/WebRTC/AOQ)](https://help.aliyun.com/en/model-studio/realtime.md): Integration guide for Qwen-Omni-Realtime multimodal model: WebSocket, WebRTC and AOQ protocols with streaming audio/image input and text+audio output; covers session.update, VAD/semantic_vad turn detection, voice cloning, web search and function calling; includes qwen3.5-omni-plus/flash-realtime selection, latency tuning, and session duration/turn limits. - [Qwen-Omni](https://help.aliyun.com/en/model-studio/qwen-omni.md): Multimodal model (text/image/audio/video in, text+audio out) via OpenAI-compatible API. Call qwen3.5-omni-plus with modalities: [text, audio]. Streaming required. Beijing and Singapore. - [Qwen-Omni voice list](https://help.aliyun.com/en/model-studio/omni-voice-list.md): Voice parameter values for Qwen3.5-Omni, Qwen3-Omni-Flash, and Qwen-Omni-Turbo models (real-time and non-real-time). Default voice: Tina (3.5) or Cherry (Flash/Turbo). 50+ voices across 30 languages. - [Embedding and Rerank Model Selection (text-embedding-v4/qwen3-rerank)](https://help.aliyun.com/en/model-studio/embedding-rerank-model.md): Pick embedding and rerank models for RAG/semantic search: text-embedding-v4, qwen3-vl-embedding, qwen3-rerank; migration mapping from OpenAI/Cohere/Voyage. - [Text & Multimodal Embedding (Embedding API)](https://help.aliyun.com/en/model-studio/embedding.md): Convert text/images/video to vectors via qwen3.7-text-embedding for semantic search, recommendation, clustering. OpenAI-compatible and DashScope interfaces, batch calling. - [Rerank models](https://help.aliyun.com/en/model-studio/rerank.md): Re-score retrieved documents by relevance using qwen3-rerank. RAG workflow: retrieve 50-100 candidates, rerank to top 5-10. OpenAI-compatible /reranks endpoint and DashScope SDK - [Connect Clients and Dev Tools to Model Studio (Claude Code/Qwen Code/OpenClaw)](https://help.aliyun.com/en/model-studio/use-chat-client-or-development-tool.md): Connect Claude Code, Qwen Code, OpenClaw, Codex and OpenCode clients to Model Studio: API Key setup to call qwen models, plus the MCP marketplace. - [OpenClaw Installation and Bailian Access Config (Token Plan/Coding Plan/Pay-as-you-go)](https://help.aliyun.com/en/model-studio/openclaw.md): Install open-source AI assistant OpenClaw on macOS/Linux/Windows, and configure access to Alibaba Cloud Model Studio via four methods: pay-as-you-go, Coding Plan, Token Plan Personal/Team edition API Keys and Base URLs. - [Hermes Agent Setup and Access Configuration (Token Plan/Coding Plan/Pay-as-you-go)](https://help.aliyun.com/en/model-studio/hermes-agent.md): Install, verify, and connect the Hermes Agent terminal AI coding tool to Model Studio: covers Token Plan Personal/Team, Coding Plan, and pay-as-you-go billing, with Anthropic-compatible and OpenAI-compatible Base URL examples and API Key retrieval. - [Connect Claude Code to Model Studio (Install & 4 Billing Modes)](https://help.aliyun.com/en/model-studio/claude-code.md): Set up Claude Code with Model Studio: npm install, skip Anthropic login, API key config for pay-as-you-go, Coding Plan, Token Plan Solo/Team; context window 200K–1M. - [Connect OpenCode to Model Studio (Token Plan/Coding Plan/Pay-as-you-go)](https://help.aliyun.com/en/model-studio/opencode.md): Configure the OpenCode terminal AI coding tool for Alibaba Cloud Model Studio: npm install, opencode.json setup for Token Plan Personal/Team, Coding Plan, and pay-as-you-go with baseURL and API Key; supports qwen3.8-max, qwen3.7-plus, DeepSeek V4, Kimi, GLM, and more. - [Cursor integration](https://help.aliyun.com/en/model-studio/cursor.md): Connect Cursor IDE to Model Studio via OpenAI API Key and Base URL. Three plans: Token Plan (Team), Coding Plan, Pay-as-you-go, each with different endpoints and model name aliases. - [Connect Codex to Model Studio (Token Plan/Coding Plan/Pay-as-you-go)](https://help.aliyun.com/en/model-studio/codex.md): Configure OpenAI Codex CLI to work with Alibaba Cloud Model Studio: install @openai/codex via npm, set up ~/.codex/config.toml and OPENAI_API_KEY. Covers Token Plan Personal/Team, Coding Plan, and pay-as-you-go billing, Responses vs Chat/Completions API selection, CC Switch for multi-key rotation, and troubleshooting for 401/404/429 errors. - [Qwen Code Setup and Configuration (Token Plan/Coding Plan/Pay-as-you-go)](https://help.aliyun.com/en/model-studio/qwen-code.md): Install and configure Qwen Code, a terminal AI coding tool: connect via Token Plan Personal/Team, Coding Plan, or pay-as-you-go. Covers macOS/Linux/Windows install commands, settings.json advanced config, model switching, VS Code plugin, and Desktop app. - [QwenPaw Installation and Bailian Integration (Token Plan/Coding Plan/Pay-as-you-go)](https://help.aliyun.com/en/model-studio/qwenpaw.md): Install and configure QwenPaw (formerly CoPaw), an open-source personal AI assistant by AgentScope, for local or cloud deployment. Covers pip/one-click script/Docker installation, API Key and Base URL setup for Token Plan Personal/Team, Coding Plan, and DashScope pay-as-you-go, default model settings, and troubleshooting for 401 errors and context limit issues. - [Connect Cherry Studio to Model Studio (Token Plan/Coding Plan/Pay-as-you-go)](https://help.aliyun.com/en/model-studio/cherry-studio.md): Configure open-source AI desktop client Cherry Studio for Alibaba Cloud Model Studio: API Key and endpoint setup for Token Plan Personal/Team, Coding Plan, and pay-as-you-go billing. Includes enable_thinking error troubleshooting and free quota region notes. - [Chatbox Setup for Model Studio (Token Plan/Coding Plan/Pay-as-you-go)](https://help.aliyun.com/en/model-studio/chatbox.md): Connect Chatbox to Model Studio via OpenAI-compatible API: API Key and host setup for Token Plan Personal/Team, Coding Plan, or pay-as-you-go; error troubleshooting. - [Connect Cline to Model Studio (Token Plan/Coding Plan/Pay-as-you-go Setup)](https://help.aliyun.com/en/model-studio/cline.md): Connect Cline to Model Studio: Base URL and API Key for Token Plan, Coding Plan, and pay-as-you-go; R1 messages format for Qwen3/QwQ; Bailian CLI example. - [Qoder Integration with Model Studio (IDE/CLI/JetBrains Plugin)](https://help.aliyun.com/en/model-studio/qoder-agent.md): Connect Qoder to Model Studio: install IDE or CLI (qodercli), select provider Alibaba Cloud Bailian, choose Token Plan/Coding Plan/pay-as-you-go, add API Key. - [Qoder CN (Formerly Lingma): Connect to Model Studio (API Key & Billing Plans)](https://help.aliyun.com/en/model-studio/lingma-agent.md): Connect Qoder CN (formerly Lingma) standalone IDE to Model Studio: set provider and API Key for Token Plan, Coding Plan or pay-as-you-go. Enterprise edition unsupported. - [Connect Kilo CLI to Model Studio (Token Plan/Coding Plan/Pay-as-you-go)](https://help.aliyun.com/en/model-studio/kilo-cli.md): Configure Kilo Code CLI for Alibaba Cloud Model Studio: npm install, config.json setup for Token Plan Personal/Team, Coding Plan, and pay-as-you-go. Includes baseURL, API Key, model declarations (qwen3.8-max, deepseek-v4, kimi-k2.6), and verification steps. - [Image and video API with Postman and cURL](https://help.aliyun.com/en/model-studio/first-call-to-image-and-video-api.md): Test image/video generation APIs using Postman or cURL. Async pattern: create task (returns task_id), then poll for result. Text-to-image example with wan2.5-t2i-preview model - [Connect Dify to Model Studio (Qianwen Plugin & API Key Setup)](https://help.aliyun.com/en/model-studio/dify.md): Dify + Model Studio: install Qianwen plugin, set pay-as-you-go API Key, run qwen-plus/Qwen-VL, build chatbot/workflow/RAG apps. Token Plan/Coding Plan Keys unsupported. - [Third-party tool configuration](https://help.aliyun.com/en/model-studio/more-tools.md): Connect any OpenAI or Anthropic API-compatible tool via three billing plans: Token Plan (Team Edition), Coding Plan, or pay-as-you-go. Includes base URLs and API key setup for each plan and region - [Model Inference on Model Studio (DashScope SDK / OpenAI-compatible)](https://help.aliyun.com/en/model-studio/model-high-speed-inference.md): Call models for inference on Model Studio: invoke qwen-series models via DashScope SDK or OpenAI-compatible endpoints, with RPM/TPM rate-limit rules and invocation guidance. - [Model Studio Fast Mode (80-100 TPS Inference)](https://help.aliyun.com/en/model-studio/fast-mode.md): High-TPS inference: 1.5-2x standard speed (80-100 TPS). Set model=glm-5.2-fast-preview; per-token billing, over-TPM requests queue without throttling; now in preview. - [Model Inference - TPM Reservation](https://help.aliyun.com/en/model-studio/tpm-reservation.md): Lock dedicated TPM inference capacity for specified models to avoid public rate limiting during peak traffic. Generates a dedicated model code after creation; supports Qwen, DeepSeek, GLM, Kimi and other models with daily prepaid billing and optional auto-overflow to pay-as-you-go or reserved-only mode returning 429. - [Model Studio Fine-tuning: Qianwen / Image / Video / Speech Synthesis Models](https://help.aliyun.com/en/model-studio/fine-tuning.md): Model Studio fine-tuning hub: tune Qianwen text-generation, image-generation, video-generation and speech-synthesis models. For domain-style customization scenarios. - [Qwen Model Fine-tuning (Methods, Training Data, Console & API)](https://help.aliyun.com/en/model-studio/fine-tune-text-generation-model.md): Hub for fine-tuning Qwen models: fine-tuning overview, training data upload rules, console fine-tuning, fine-tuning via API, zero-code safety and compliance enhancement. - [Model Studio Fine-tuning Overview (SFT/CPT/DPO & Billing)](https://help.aliyun.com/en/model-studio/model-training-overview.md): Overview of Model Studio fine-tuning: SFT, CPT and DPO methods with use cases and data requirements; supports Qwen3/Qwen2.5 text and VL models; billed by training tokens (¥0.003–¥0.35 per 1K tokens); available only in China (Beijing) region. - [Model Fine-tuning - Training Data Upload Rules](https://help.aliyun.com/en/model-studio/text-generation-tuning-data-upload-rules.md): Alibaba Cloud Model Studio text generation fine-tuning data formats and upload limits: SFT/DPO/CPT jsonl structure, multimodal zip packaging rules, single file size cap 200-300 MB, API upload quota 100 GB / 10,000 files. Warning: publish and delete are irreversible; dataset type cannot be changed after creation. - [Model Tuning in Console: CPT/SFT/DPO Training & Hyperparameters](https://help.aliyun.com/en/model-studio/model-training-on-console.md): Create training jobs: CPT/SFT/DPO selection, full-parameter vs LoRA training, hyperparameters (learning_rate), data requirements (CPT 10M tokens, SFT 1000+, DPO 100+). - [Model Fine-tuning API: Create SFT/CPT/DPO Training Jobs](https://help.aliyun.com/en/model-studio/fine-tuning-api-guide.md): Fine-tune Qwen via DashScope API: upload jsonl data, create SFT/CPT/DPO jobs (/api/v1/fine-tunes), hyperparameters, OSS mount; API jobs billed per token only. - [Zero-code LLM security compliance fine-tuning](https://help.aliyun.com/en/model-studio/enhance-the-security-compliance-of-large-models.md): Fine-tune Qwen3-8B with zero-code SFT (LoRA or full-parameter) to reject unsafe content. Covers dataset preparation, training config (batch_size, lr, eval_steps), deployment, and evaluation workflow. - [Wan Image Generation Fine-tuning (wan2.7-image LoRA, Text-to-Image & Image-to-Image)](https://help.aliyun.com/en/model-studio/wan-image-generation-finetune-guide.md): Fine-tune Wan2.7 image models (wan2.7-image-pro/wan2.7-image) via SFT-LoRA: upload zip datasets, create t2i/i2i training jobs (max_steps, learning_rate). Beijing region only. - [Fine-tune Wan video generation models](https://help.aliyun.com/en/model-studio/wan-video-generation-finetune-guide.md): SFT-LoRA fine-tune wan2.6-i2v, wan2.5-i2v-preview, or wan2.2-i2v-flash for custom effects, actions, or camera movements. Upload ZIP dataset, train, deploy, generate. Singapore region only. - [Fine-tune Speech Synthesis Models (TTS Voice Fine-tuning)](https://help.aliyun.com/en/model-studio/fine-tune-speech-synthesis-model.md): Model Studio fine-tuning branch: fine-tune speech synthesis (TTS) models to customize voices; training data preparation and tuning workflow. - [CosyVoice TTS Fine-tuning (efficient_sft, API Only)](https://help.aliyun.com/en/model-studio/fine-tune-speech-synthesis-model-by-api.md): SFT fine-tune cosyvoice-v3-flash to clone one speaker's voice; API-only, Beijing only; voice locked to "default"; ¥0.2/1k training tokens plus hourly deployment fees. - [Reinforcement Learning Training - Model Studio Fine-Tuning](https://help.aliyun.com/en/model-studio/rl-training-overview.md): Optimize reasoning and tool-calling capabilities of qwen3.5-9b/qwen3.6-flash models via reward-driven RL training. MTU billing only (MTU4 prepaid ¥19,914/month or postpaid ¥41/hour). Covers environment setup, Rollout/Reward function development, job submission, trajectory observation, and checkpoint publishing. - [Reinforcement Learning - Rollout/Reward/Group Reward Function Development Guide](https://help.aliyun.com/en/model-studio/rl-function-development-guide.md): Develop custom function components for Alibaba Cloud Model Studio RL training: Rollout for Agent trajectory generation, Reward for single-output scoring (including decorator-based multi-dimension aggregation), and Group Reward for batch comparative scoring. Covers AgentOutput/TaskStatus data models, inline scoring, three-layer weight design, Reward Hacking defenses, and remote testing methods. - [Model Studio RL Training Submission and Configuration (run/YAML/submit_job)](https://help.aliyun.com/en/model-studio/rl-training-config-monitoring.md): Submit AgenticRL training via one-step run, YAML, or submit_job; config precedence, Rollout/Reward Runtime resources, MTU4 postpaid spec, and required hyperparameters. - [Observability for RL Training: OpenTelemetry Tracing Setup and Metrics Reference](https://help.aliyun.com/en/model-studio/observable-configuration-for-reinforcement-learning.md): Add OpenTelemetry Tracing to RL training: observe_processor decorators for LLM/tool calls, export to ARMS, view in console; metrics reference and FAILED triage. - [Model Studio Model Deployment (Provisioned Throughput/Model Import/API Deployment)](https://help.aliyun.com/en/model-studio/model-deployment-1.md): Model deployment hub: overview, provisioned throughput for long inputs and caching, model import, deploying models via API. Deploy models as dedicated instances. - [Model Deployment Billing: PTU Provisioned Throughput, Model Units, Token Usage](https://help.aliyun.com/en/model-studio/model-deployment-introduction.md): Dedicated inference for preset or tuned models: PTU (TPS ~1.5–2.0x), Model Units, Token Usage (LoRA); billing locked at creation, PTU overflow → pay-as-you-go or 429. - [PTU Provisioned Throughput Long Input and Cache](https://help.aliyun.com/en/model-studio/ptu-long-input-and-cache.md): PTU deployment supports up to 256K-token long input and prefix caching. Models like glm-5.1, deepseek-v4-pro, and qwen3.7-plus offer cache discounts (as low as 8%) and tiered coefficients. Covers quota calculator, overflow policies (auto fallback to pay-as-you-go or 429), and response fields service_tier, provisioned_tokens, and cached_tokens. - [Import LoRA Models from OSS to Model Studio (Auth/File Validation)](https://help.aliyun.com/en/model-studio/model-import.md): Import locally trained LoRA models from Alibaba Cloud OSS to Model Studio: first-time OSS service-linked role authorization with bailian-datahub-access=read tag, required adapter_model.safetensors/adapter_config.json/config.json files and rank 8/16/32/64 constraints, supported Qwen3/Qwen2.5/Qwen3-VL base models, status flow (Creating/Succeeded/Invalidated) and AvailableModelFileNotFound troubleshooting. LoRA only; full-parameter fine-tuning and incremental training are not supported. - [Model Studio Deployment API Quick Start (PTU/MU/Token Billing)](https://help.aliyun.com/en/model-studio/model-deployment-quick-start.md): Deploy Qwen and other models via DashScope API: create deployment with PTU, Model Unit (MU), or Token billing plans; query status; run inference; delete service. Available only in China (Beijing); billing starts immediately after deployment. - [Model Studio Model Evaluation (Auto-scoring/Human Annotation/Metrics)](https://help.aliyun.com/en/model-studio/model-evaluation-introduction.md): Evaluate fine-tuned models: auto-scoring by judge model or manual Pass/Fail annotation, custom evaluation dimensions and scorer types, compare pre/post training performance. - [Model Evaluation Overview (Custom/Baseline, Scoring Methods)](https://help.aliyun.com/en/model-studio/model-evaluation-overview.md): Custom vs baseline (Beijing-only) model evaluation: AI judge (qwen-Max), ROUGE/BLEU/Cosine rules, human scoring; dimensions, tasks and score reports for model selection. - [Evaluation dimensions](https://help.aliyun.com/en/model-studio/evaluation-metrics.md): This topic describes how to create and manage evaluation dimension templates. - [Model Compression (Quantize Fine-tuned Models)](https://help.aliyun.com/en/model-studio/model-compression.md): Quantize fine-tuned models into low-precision versions to cut deployment cost: quantization templates, calibration data; tasks free during trial, compressed models billed by MU spec. - [Model Compression - Quantize Fine-tuned Models to Reduce Deployment Cost](https://help.aliyun.com/en/model-studio/model-compression-introduction.md): Quantize full-precision fine-tuned models to low-precision versions to lower MU specs and inference costs. Available only in China (Beijing). Compression is irreversible and does not support re-compression or further fine-tuning. Currently supports qwen3.5-flash and other models, reducing deployment costs by approximately 56%. Compression tasks are currently free of charge. - [Model Call Usage Statistics & Performance Monitoring (Token/Latency)](https://help.aliyun.com/en/model-studio/model-monitoring.md): Model call monitoring hub in Model Studio: Token consumption, call volume and success rate, latency stats — reconcile costs with billing and check call quality. - [Model Studio Usage Statistics (Tokens/Images/Seconds) & Free Quota](https://help.aliyun.com/en/model-studio/model-usage-statistics.md): View Model Studio model usage and token consumption: stats by workspace with ~1-hour delay, last 30 days only; enable free-quota auto-stop (403 AllocationQuota.FreeTierOnly) to avoid overage; LLMs billed by token, images by count, video by second, audio by second/char/token. - [Model Telemetry - Call Records, Token Usage and Alert Configuration](https://help.aliyun.com/en/model-studio/model-telemetry.md): Alibaba Cloud Model Studio model telemetry tracks call records, performance metrics (RPM, TPM, first-token latency), token consumption and anomaly alerts. Standard monitoring has hourly delay while advanced monitoring offers minute-level insights. Inference logs are available only for selected models in Beijing, Singapore and Virginia regions; calls before log activation cannot be retroactively recorded. - [Model Data](https://help.aliyun.com/en/model-studio/model-data-overview.md) - [Model Studio Training & Evaluation Datasets (SFT/DPO/CPT)](https://help.aliyun.com/en/model-studio/training-set-and-evaluation-set.md): Manage training sets (SFT/DPO/CPT for text generation, multimodal understanding, image-to-video) and evaluation sets in Model Studio. Import via local upload, OSS, or log backflow. DPO/CPT available only in Beijing region; CPT recommends ≥50M tokens. Publish/delete are irreversible; data management itself is free. - [Model Studio Data Processing: Clean and Augment Training Sets via Data Flow Canvas](https://help.aliyun.com/en/model-studio/data-processing.md): Clean and augment SFT training sets on a data flow canvas: 10 operators (special-content removal, sensitive-info masking), ChatML format only, Beijing region only, no API yet. - [Model Data - Log Backflow](https://help.aliyun.com/en/model-studio/model-log-backflow.md): Converts SLS inference logs into JSONL structured datasets for Bailian model fine-tuning (SFT/DPO/CPT) or evaluation. Available only in China North 2 (Beijing) and Singapore regions, with a single-run limit of 100,000 entries. Supports platform storage and OSS mounting. Audit logs must be enabled before inference logs. - [Model Studio Security & Compliance (AI Guardrails/Filing/Privacy)](https://help.aliyun.com/en/model-studio/security-and-compliance.md): SOC 2 audit, AES-256 encryption, data never used for training; AI safety guardrails on input/output; app launch & algorithm filing; compliance-hardened model deployment. - [Bailian Permission Management (Workspaces, Roles, Model-level Access, API Keys)](https://help.aliyun.com/en/model-studio/permission-management-overview.md): Page- and model-level access control; super admin/workspace admin/user roles, model call/fine-tune/deploy grants, AliyunBailianFullAccess RAM policy, API Keys, bill splitting. - [Model Studio Transmission Security (Encrypted Access/RSA Key/PrivateLink)](https://help.aliyun.com/en/model-studio/transmission-security.md): Secure model invocation: encrypted access to model inference (AES encryption), obtain RSA public key, access model or application APIs through PrivateLink private network for high-confidentiality scenarios. - [Encrypted model inference requests](https://help.aliyun.com/en/model-studio/encrypted-access-to-model-inference.md): Encrypt input field with AES, wrap AES key with RSA via X-DashScope-EncryptionKey header. SDK shortcut: enableEncrypt=true (Java) or enable_encryption=True (Python). - [Get RSA public key API](https://help.aliyun.com/en/model-studio/model-interface-aes-encryption.md): GET /api/v1/public-keys/latest to retrieve an RSA public key ID and value for encrypting input fields before transmission. Returns public_key and public_key_id - [Access Model Studio APIs via PrivateLink](https://help.aliyun.com/en/model-studio/access-model-studio-through-privatelink.md): Create a PrivateLink endpoint (com.aliyuncs.dashscope) for VPC-only API access. Replace base_url with endpoint domain. Cross-region via CEN. Beijing and Singapore. - [Model Studio Secure Storage (Endpoint/Zone IP/MSE Gateway)](https://help.aliyun.com/en/model-studio/secure-storage.md): Secure Storage private access: configure endpoint and connection, zone IPs, private-network resources, MSE gateway. Call Model Studio from a VPC over private link. - [Secure Storage VPC endpoint configuration](https://help.aliyun.com/en/model-studio/configure-an-endpoint-and-initiate-a-connection.md): Create a PrivateLink reverse endpoint in a VPC to connect a Secure Storage workspace to Elasticsearch, ADB, and OSS. Requires Beijing region VPC spanning zones G, H, or L. - [Secure Storage zone IP configuration](https://help.aliyun.com/en/model-studio/configure-zone-ip.md): Create an MSE cloud-native gateway, obtain zone VIPs from the NLB, configure zone IPs in the Secure Storage workspace, and add NLB IPs to the security group inbound rules. - [Secure Storage VPC resource setup](https://help.aliyun.com/en/model-studio/configure-resources-in-private-network.md): Provision OSS bucket, AnalyticDB PostgreSQL, and Elasticsearch instances for a Secure Storage workspace. Includes bucket tagging, CORS rules, ADB vector engine, and ES allowlist configuration. - [Secure Storage MSE gateway configuration](https://help.aliyun.com/en/model-studio/configure-mse.md): Set up MSE Cloud Native Gateway service and route for a Secure Storage workspace. Connects to Elasticsearch via DNS domain, then activate the workspace with ES/ADB/OSS resources. - [Input and output AI guardrail](https://help.aliyun.com/en/model-studio/content-security.md): Add content moderation to model calls via X-DashScope-DataInspection header. Set input/output to 'cip' to scan for non-compliant content. Pay-as-you-go per token, min 1000 tokens per request. - [Bailian Model Filing Info (Algorithm & LLM Filing Numbers)](https://help.aliyun.com/en/model-studio/model-filing-information-publicity.md): Algorithm and LLM filing numbers for models on Bailian: Qwen, Wanxiang, DeepSeek, MiniMax, Kling, Vidu and 13 models in total, plus disclaimer on third-party filings. - [Qwen LLM application listing and compliance filing](https://help.aliyun.com/en/model-studio/compliance-and-launch-filing-guide-for-ai-apps-powered-by-the-tongyi-model.md): Obtain algorithm filing numbers and cooperation agreements to list Qwen/Wan-powered apps on marketplaces. Covers consumer apps, public-opinion apps, and internal-only scenarios. - [Model Studio Compliance & Privacy Notice (SOC 2/Data Not Used for Training)](https://help.aliyun.com/en/model-studio/privacy-notice.md): SOC 2 certified (security/availability/confidentiality). User data never used for training; transmissions AES-256 encrypted; invocation data stored per service agreement. - [Best Practices](https://help.aliyun.com/en/model-studio/use-cases.md) - [HappyHorse: Unified film and TV creation platform](https://help.aliyun.com/en/model-studio/infinite-canvas.md): Traditional AI visual creation suffers from fragmented models, disjointed workflows, and inefficient asset management. This unified film and television creation platform addresses these core pain points. Built on Alibaba Cloud Model Studio visual models, the platform integrates the image generation capabilities of Wan 2.7 with the video generation capabilities of HappyHorse. Three core engines power the platform: node-based visual orchestration, a conversational AI director, and online editing. Together they create a seamless workflow from initial concept to final output. This simplifies complex visual creation and lets your creativity flow freely. - [Text-to-text prompt guide](https://help.aliyun.com/en/model-studio/prompt-engineering-guide.md): Prompt engineering techniques: framework (context/objective/style/tone/audience/response), output examples, task steps, and separators. Includes console one-click optimization tool. - [Text-to-image prompt guide](https://help.aliyun.com/en/model-studio/text-to-image-prompt.md): Prompt formulas: basic (subject+setting+style) and advanced (+camera+atmosphere+detail). Visual vocabulary for shot size, perspective, lens, lighting. Covers prompt_extend and negative_prompt params - [Text-to-video prompt guide](https://help.aliyun.com/en/model-studio/text-to-video-prompt.md): Five prompt formulas: basic (entity+scene+motion), advanced (+aesthetic+style), image-to-video, sound (voice+SFX+BGM), and multi-shot with timestamps. Wan 2.7/2.6/2.5 examples - [Build RAG with LlamaIndex](https://help.aliyun.com/en/model-studio/build-rag-applications-based-on-llamaindex.md): Use DashScopeCloudIndex and DashScopeCloudRetriever for RAG pipelines. Covers file parsing, knowledge base creation, and query engine setup. Requires llama-index-indices-managed-dashscope. - [Custom model tuning best practices](https://help.aliyun.com/en/model-studio/model-training-best-practices.md): End-to-end workflow: data preparation (min 500 samples as Prompt-Completion pairs), fine-tuning, deployment, and evaluation. Covers iterative optimization strategy - [Convert documents to video with LLMs](https://help.aliyun.com/en/model-studio/use-llm-to-convert-document-to-video.md): Four-step pipeline: segment document with LLM, generate slides via Marp, synthesize narration with multimodal models, compile video with FFmpeg. Downloadable Python code package included - [Rate limiting best practices](https://help.aliyun.com/en/model-studio/rate-limiting-best-practices.md): Four traffic control strategies: exponential backoff retry, token bucket, dual RPM/TPM shaping, adaptive congestion control. Also covers PTU provisioned throughput, Batch API, and model fallback - [Explicit Cache (cache_control) Guide and Best Practices](https://help.aliyun.com/en/model-studio/explicit-cache-guide.md): Model Studio explicit caching: cache_control marks give 100% deterministic cache hits; first write costs +25%, hits save 90%. Setup for Claude Code/OpenCode/OpenClaw/Hermes Agent, plus API limits (min 1024 tokens, up to 4 marks, 5-minute TTL). - [Third-Party Model Invocation Tutorials (DeepSeek/Kimi/GLM/MiniMax)](https://help.aliyun.com/en/model-studio/third-party-model-integration-tutorial.md): Call DeepSeek/Kimi/GLM/MiniMax via Alibaba Cloud hosting or direct vendor endpoints (SiliconFlow/Moonshot/Zhipu), plus the Vidu video prompt guide. - [Call DeepSeek Models on Model Studio (API Key, Base URL, enable_thinking)](https://help.aliyun.com/en/model-studio/deepseek-api.md): Call deepseek-v4-pro/v4-flash/r1 on Alibaba Cloud Model Studio via OpenAI-compatible, DashScope, or Anthropic APIs. Covers regional Base URLs, enable_thinking/reasoning_effort parameters, Responses API with web search tools, model list, and per-token billing. - [DeepSeek via SiliconFlow integration](https://help.aliyun.com/en/model-studio/siliconflow-deepseek-api.md): SiliconFlow-hosted DeepSeek (deepseek-v3.2) via OpenAI-compatible API. Use enable_thinking for reasoning mode. Longer context than native provider; Beijing region API key required - [DeepSeek via Kuaishou Wanqing (vanchin/deepseek-v4-pro)](https://help.aliyun.com/en/model-studio/deepseek-api-by-vanchin.md): Call Kuaishou Wanqing-supplied DeepSeek (vanchin/deepseek-v4-pro) on Model Studio: activation, OpenAI-compatible calls, enable_thinking streaming; Beijing region only, API Key from Beijing. - [Kimi Model Invocation (OpenAI-Compatible/DashScope/Base URLs)](https://help.aliyun.com/en/model-studio/kimi-api.md): Invoke Kimi models on Model Studio: kimi-k2-thinking, kimi-k2.5, kimi-k2.6 via OpenAI-compatible or DashScope API; Base URLs for Beijing, Singapore, Tokyo, Virginia, Frankfurt. Moonshot-Kimi-K2-Instruct and kimi-k2-thinking retire on 2026-07-09; migrate to qwen3.7-plus/qwen3.8-max. - [Call Kimi Models (kimi-k3/k2.7-code, Thinking Mode, Multimodal)](https://help.aliyun.com/en/model-studio/kimi-api-by-moonshot-ai.md): Call Moonshot AI Kimi models on Model Studio in China (Beijing): activation, API Key setup, OpenAI-compatible calls. Covers kimi-k3, kimi-k2.7-code-highspeed, kimi-k2.7-code, kimi-k2.6, kimi-k2.5 with text/image/video inputs and reasoning_effort thinking mode; images/videos must be passed via public URLs. Context cache hits billed at 10%-20% of input price; thinking tokens billed as output. - [Model Studio GLM API (glm-5.2 / Thinking Mode / Streaming Tool Calls)](https://help.aliyun.com/en/model-studio/glm.md): Call Zhipu GLM models (glm-5.2/glm-5.1/glm-5, etc.) via OpenAI/DashScope/Anthropic-compatible APIs with 1M context. Supports enable_thinking, reasoning_effort, tool_stream for streaming Function Calling, and clear_thinking. 1M free tokens per model; billed by input/output tokens. glm-4.6/glm-4.7 retire on 2026-10-10. - [GLM Zhipu Direct Model Invocation](https://help.aliyun.com/en/model-studio/glm-zhipu.md): Invoke Zhipu-direct GLM-5.2/5.1/5 models on Alibaba Cloud Model Studio, available only in China (Beijing) region with workspace-specific endpoint. Supports 1M context, enable_thinking mode, reasoning_effort depth control, tool_stream for streaming function calling, and clear_thinking to exclude historical reasoning content. Billed by input/output tokens. - [MiniMax-M2.5 integration (Bailian-hosted)](https://help.aliyun.com/en/model-studio/minimax-api.md): MiniMax-M2.5 through Model Studio's OpenAI-compatible endpoint. Supports streaming with reasoning_content for chain-of-thought. Python, Node.js, Java, and curl examples - [MiniMax Models on Model Studio (MiniMax-M2.7/Thinking Mode/OpenAI Compatible)](https://help.aliyun.com/en/model-studio/minimax-api-by-minimax.md): Call MiniMax-M2.7 on Model Studio: console activation, OpenAI-compatible API, thinking mode, Beijing region only, get API Key from cn-beijing. - [Vidu Video Generation Prompt Guide (Formula/Keyword Dictionary/Tuning)](https://help.aliyun.com/en/model-studio/vidu-video-generation-prompt-guide.md): Vidu video prompt guide: formula (subject/scene/environment/style), trigger keywords (motion, camera moves, Ghibli/Shinkai styles, VFX), text/image-to-video tuning cases. - [Xiaomi MiMo Model Call (mimo-v2.5-pro, OpenAI Compatible)](https://help.aliyun.com/en/model-studio/mimo.md): Call Xiaomi mimo-v2.5-pro on Model Studio: Beijing API Key only; thinking on by default (enable_thinking:false to disable); OpenAI-compatible Python/Node.js samples. - [Stepfun Step Model Invocation (step-3.7-flash Thinking Mode, Beijing Only)](https://help.aliyun.com/en/model-studio/stepfun.md): Invoke Stepfun step-3.7-flash multimodal reasoning model: thinking off by default, enable_thinking=true to turn on, reasoning_effort sets depth; Beijing-region API Key only; OpenAI-compatible. - [WebRTC Multimodal Dialog Kit Real-Time Call Integration](https://help.aliyun.com/en/model-studio/best-practice-webrtc-multimodal-dialog.md): Integrate Bailian multimodal-dialog kit via WebRTC in browsers for low-latency audio/video interaction with AI/AR glasses, learning devices, and robots. Covers RTCPeerConnection setup, oai-events DataChannel, SDP exchange, and API Key authentication - [WebRTC Integration with qwen3.5-omni-plus-realtime for Real-Time Calls](https://help.aliyun.com/en/model-studio/best-practice-webrtc-omni-realtime.md): Integrate Bailian Realtime API via WebRTC and JavaScript in browsers to enable real-time audio/video calls with the qwen3.5-omni-plus-realtime model. Covers RTCPeerConnection setup, SDP exchange, media gating, session.update configuration, and resource cleanup, supporting server_vad and semantic_vad modes. - [Realtime API - Integrate AOQ Client SDK for qwen3.5-omni-plus-realtime Calls](https://help.aliyun.com/en/model-studio/best-practice-aoq-omni-realtime.md): Integrate AOQ Client SDK on Android/iOS/HarmonyOS for qwen3.5-omni-plus-realtime audio/video calls: token auth; enable media sending only after session.updated. - [AOQ Push-to-Talk with qwen3.5-omni-plus-realtime (Manual Mode)](https://help.aliyun.com/en/model-studio/use-aoq-to-access-qwen3-5-omni-plus-realtime-to-realize-key-voice-dialogue.md): Integrate qwen3.5-omni-plus-realtime via AOQ Client SDK in Manual mode for client-controlled push-to-talk and optional photo Q&A. Covers iOS/Android/HarmonyOS/Linux SDK setup, AppServer token auth, session.update with turn_detection=null, Audio/Video/Data track config, input_audio_buffer.commit and response.create sequencing, and key cautions. - [Real-Time Voice Chat via AOQ with qwen-audio-3.0-realtime-plus (Android/iOS/HarmonyOS/Python)](https://help.aliyun.com/en/model-studio/real-time-voice-conversation-using-aoq-access-qwen-audio-3-0-realtime-plus.md): Integrate qwen-audio-3.0-realtime-plus via AOQ SDK with server-side VAD for low-latency real-time voice conversation. Covers Android, iOS, HarmonyOS and Linux Python SDK setup, AppServer token authentication, Audio/Data dual-track configuration and session.update flow. - [Stream Speech Synthesis with qwen-audio-3.0-tts-flash via AOQ (Android/iOS/HarmonyOS/Python)](https://help.aliyun.com/en/model-studio/speech-synthesis-using-aoq-access-qwen-audio-3-0-tts-flash.md): Call qwen-audio-3.0-tts-flash over the AOQ Inference protocol for segmented real-time TTS: run-task/continue-task/finish-task event flow, AppServer token auth, Audio-track streaming playback, covering Android/iOS/HarmonyOS/Linux Python SDK integration and troubleshooting. - [Real-Time Speech Recognition with fun-asr-realtime via AOQ (Android/iOS/HarmonyOS/Python)](https://help.aliyun.com/en/model-studio/real-time-speech-recognition-using-aoq-access-fun-asr-realtime.md): Integrate the fun-asr-realtime model via AOQ SDK for live microphone transcription: AppServer token auth, separate Audio/Data tracks, run-task/result-generated/finish-task event flow, full Android Java sample and connection troubleshooting. - [Model Studio Service Support (FAQ/Agreements/After-sales)](https://help.aliyun.com/en/model-studio/support.md): Service support entry for Model Studio: FAQ (errors and usage), related agreements, after-sales instructions. Start here for troubleshooting, usage questions or support requests. - [Model Studio Model List (Qwen/DeepSeek/Kimi Selection)](https://help.aliyun.com/en/model-studio/model-studio-model-list.md): Catalog of all callable models: Qwen qwen3.8-max/plus/flash, DeepSeek, Kimi, GLM; text generation, vision, image/video generation, speech, embeddings. Pick model IDs here. - [Model Studio Text Generation Models (Qwen-max/Plus/Flash/Qwen3)](https://help.aliyun.com/en/model-studio/model-list-text-generation.md): Model Studio text generation models: qwen-max/plus/flash, open-source Qwen3 series, qwen-coder code, qwen-mt translation, qwen-long. Pick a model and start your API call here. - [qwen3.8-max Model Capabilities and Pricing](https://help.aliyun.com/en/model-studio/qwen3-8-max.md): 2.4-trillion-parameter MoE flagship model with Image/Text/Video input and Text output, supporting 1M-token context. Enables Function Calling, structured output, web search, context caching, and batch inference (Beijing region). Priced at CNY 12 per million input tokens, CNY 36 per million output tokens, and CNY 1.5 per million cached input tokens. - [qwen3.7-max Model Capabilities and Pricing](https://help.aliyun.com/en/model-studio/qwen3-7-max.md): Qwen3.7 flagship Max model with 1M-token context. Input ¥12/M tokens, output ¥36/M tokens. Supports Function Calling, web search, context caching, and batch inference for coding and agent scenarios. Includes qwen3.7-max-preview and snapshots 2026-05-17, 2026-05-20, 2026-06-08. - [qwen3.7-max-us Text Generation Model Deployed in US](https://help.aliyun.com/en/model-studio/qwen3-7-max-us.md): Qwen3.7 Max flagship model for text-in/text-out with Function Calling, structured output, and context caching. Supports 1M-token context and up to 131,072 output tokens. Deployed in US (Virginia). Input: CNY 18.736 per million tokens; output: CNY 56.207 per million tokens; RPM 600. - [qwen3.6-max-preview Qwen Max Preview Model](https://help.aliyun.com/en/model-studio/qwen3-6-max.md): Largest pure-text Max preview model in the Qwen3.6 series, supporting Function Calling, structured output, web search, and context caching. Max input 245,760 tokens, output 65,536 tokens; thinking mode chain-of-thought up to 131,072 tokens. China North 2 pricing: input ≤128k at CNY 9 per million tokens, output at CNY 54 per million; RPM 600, TPM 1,000,000. Vibe coding and frontend development capabilities significantly improved over Qwen3-Max and Qwen3.6-Plus. - [qwen3-max Qwen3 Max Model](https://help.aliyun.com/en/model-studio/model-qwen3-max.md): Qwen3 Max GA with agent programming and tool-calling upgrades, 262144 context / max output 65536 tokens, supports Function Calling, web search, structured output, context caching, batch inference; input ≤32k from CNY 2.5 per million tokens in cn-beijing - [qwen-max Qwen 2.5 Trillion-Parameter Language Model](https://help.aliyun.com/en/model-studio/qwen-max.md): Qwen 2.5 series trillion-parameter text generation model with max input 30,720 tokens and output 8,192 tokens. China North 2 supports Function Calling, web search, context caching, and batch inference; Singapore does not support Function Calling or web search. Beijing pricing: input ¥2.4/M tokens, output ¥9.6/M tokens, cache hit ¥0.48/M; RPM 1,200, TPM 1,000,000. - [qwen3.7-plus Multimodal Agent Model](https://help.aliyun.com/en/model-studio/qwen3-7-plus.md): Cost-effective Plus model in the Qwen3.7 series with Image/Text/Video input and Text output. Supports Function Calling, web search, context caching, and batch inference with up to 1M-token context, designed for multimodal agent scenarios such as screen operation, GUI navigation, and visual code generation. - [qwen3.7-plus-us Multimodal Hybrid Agent Model](https://help.aliyun.com/en/model-studio/qwen3-7-plus-us.md): Cost-effective Qwen3.7 Plus model supporting Text/Image/Video input and text output, with hybrid agent capabilities including screen reading, GUI operation, and vision-referenced code generation. 1M-token context, deployed in US Virginia, pricing from CNY 2.998 per million tokens for input ≤256k - [qwen3.6-plus Native Vision-Language Model](https://help.aliyun.com/en/model-studio/qwen3-6-plus.md): Qwen3.6 Plus multimodal model with Image/Text/Video input, 1M context, Function Calling and web search. Enhanced Agentic coding and OCR, input from CNY 2 per million tokens - [qwen3.5-plus Multimodal Large Model](https://help.aliyun.com/en/model-studio/qwen3-5-plus.md): Qwen3.5 native vision-language Plus model supporting Text/Image/Video input and Text output, with Function Calling, structured output, web search, and context caching. Max context 1M tokens; input from CNY 0.8 per million tokens. - [qwen-plus Qwen Plus Model](https://help.aliyun.com/en/model-studio/qwen-plus.md): Qwen3-series Plus text generation model with thinking/non-thinking mode switching, Function Calling, web search, context caching, and batch inference. Supports 1M context length; input pricing starts at CNY 0.8 per million tokens (≤128k). Beijing region RPM 30,000, TPM 5,000,000. Includes qwen-plus-latest and snapshots 2025-12-01, 2025-09-11, 2025-07-28. - [qwen-plus-us Qwen Enhanced Model Deployed in US](https://help.aliyun.com/en/model-studio/qwen-plus-us.md): Bailian qwen-plus-us is the enhanced Qwen large language model deployed in US (Virginia), supporting Chinese and English input with structured output, context caching, and batch inference. Context length is 1M tokens. Input pricing starts at CNY 2.936 per million tokens for ≤128k, with RPM 600 and TPM 1,000,000. - [qwen3.7-flash Multimodal Vision-Language Model](https://help.aliyun.com/en/model-studio/qwen3-7-flash.md): Qwen3.7 native vision-language Flash model with Image/Text/Video input and Text output. Supports Function Calling, web search, context caching, and structured output. Max context 1M tokens; input from CNY 0.2 per million tokens. - [qwen3.6-flash Native Vision-Language Model](https://help.aliyun.com/en/model-studio/qwen3-6-flash.md): Qwen3.6 Flash vision-language model accepting Image/Text/Video input and Text output with 1M-token context. Supports Function Calling, web search, structured output, context caching, and batch inference. Enhanced agentic coding, math/code reasoning, and spatial intelligence. Beijing region: input ≤256k at CNY 1.2 per million tokens, output at CNY 7.2; RPM 30,000, TPM 10M. - [qwen3.6-flash-us US-Deployed Vision-Language Model](https://help.aliyun.com/en/model-studio/qwen3-6-flash-us.md): Qwen3.6 native vision-language Flash model supporting Text/Image/Video input, Function Calling, and context caching. Max context 1M tokens, output 65536 tokens. Deployed in US (Virginia); input ≤256k priced at CNY 1.874 per million tokens, RPM 15000, TPM 5M. - [qwen3.5-flash Native Vision-Language Model](https://help.aliyun.com/en/model-studio/qwen3-5-flash.md): Qwen3.5 native vision-language Flash model accepting Text/Image/Video input with Text output and 1M-token context. Supports Function Calling, structured output, web search, and context caching. Pay-as-you-go pricing from CNY 0.2 per million input tokens. - [qwen-flash Model Capabilities and Tiered Pricing](https://help.aliyun.com/en/model-studio/qwen-flash.md): Qwen3 Flash model with switchable thinking/non-thinking modes, 1M context, Function Calling, web search, structured output, and batch inference. Tiered pricing by context length: input ≤128k at CNY 0.15 per million tokens, cache hit at CNY 0.03; snapshot equivalent to qwen-flash-2025-07-28. - [Model Studio qwen-flash-us Qwen3 Flash Model Deployed in US](https://help.aliyun.com/en/model-studio/qwen-flash-us.md): qwen-flash-us is a Qwen3-series Flash text model deployed in US (Virginia) with thinking/non-thinking mode switching, structured output, context caching, and prefix continuation over 1M context. Input pricing is CNY 0.367 per million tokens for ≤256k and CNY 1.835 for 256k–1M; RPM 15,000 and TPM 10 million. - [qwen-turbo Model Capabilities and Pricing](https://help.aliyun.com/en/model-studio/qwen-turbo.md): Qwen3-series Turbo text generation model with switchable thinking/non-thinking modes. Context window 131,072 tokens (98,304 input + 16,384 output). China North 2 pricing: ¥0.3 per million input tokens, ¥0.6 per million output tokens, cache hit ¥0.06; RPM 1,200, TPM 5,000,000. Supports structured output and context caching; Function Calling and web search not supported. - [qwq-plus Qwen QwQ Reasoning Model Enhanced Edition](https://help.aliyun.com/en/model-studio/qwq-plus.md): Reinforcement-learning reasoning model based on Qwen2.5, matching DeepSeek-R1 full version on AIME and LiveCodeBench benchmarks. Supports Function Calling, web search, and batch inference with 131,072-token context. Beijing region pricing: input ¥1.6 per million tokens, output ¥4 per million tokens; Batch API at half price. - [qwen-plus-character Role-Playing Model](https://help.aliyun.com/en/model-studio/qwen-plus-character.md): Qwen series role-playing model optimized for persona instruction following, topic progression, and empathy. Supports deep personalized character restoration. Context length 32K tokens. Input: ¥0.8/million tokens, output: ¥2/million tokens, cache hit: ¥0.16/million tokens. - [qwen-plus-character-ja Japanese Role-Playing Model](https://help.aliyun.com/en/model-studio/qwen-plus-character-ja.md): Alibaba Cloud Bailian Japanese anthropomorphic interaction model optimized for character consistency, context-aware dialogue progression, and empathetic engagement with enhanced dialect and honorific localization. Text-only I/O, 8,192-token context (7,680 input / 512 output); no Function Calling, web search, batch inference, or fine-tuning. Deployed in Singapore; input ¥3.67/M tokens, output ¥10.275/M tokens; RPM 120, TPM 500,000. - [qwen-flash-character Role-Playing Model](https://help.aliyun.com/en/model-studio/qwen-flash-character.md): Qwen multilingual role-playing model optimized for persona adherence, topic progression, and empathy. 8192-token context, input at CNY 0.25 per million tokens. Supports web search and context caching; Function Calling not supported. - [Model Studio-qwen-long Long-Context LLM](https://help.aliyun.com/en/model-studio/qwen-long.md): qwen-long handles up to 10 million tokens (~15 million words) of context for long-text dialogue, parsing TXT/PDF/DOCX and BMP/PNG/JPG files. HTTP requests cap at 1M tokens; larger inputs require file upload. Pricing: input ¥0.5/M tokens, output ¥2/M tokens, Batch 50% off. RPM 1,200, TPM 3M. - [qwen-math-plus Math Problem-Solving Model](https://help.aliyun.com/en/model-studio/qwen-math-plus.md): Qwen math model on Bailian for Chinese/English equations, calculations and proofs. 4096-token context, input/output at 4/12 CNY per million tokens, RPM 1200/TPM 1M, with latest dynamic and 0919 snapshot versions - [qwen-math-turbo Math Problem-Solving Model](https://help.aliyun.com/en/model-studio/qwen-math-turbo.md): Qwen series math problem-solving model with text-only input/output. Context length 4096 tokens, max input/output 3072 tokens each. Does not support Function Calling, web search, structured output, or batch inference. Pricing in China (Beijing): input CNY 2 per million tokens, output CNY 6 per million tokens; RPM 1200, TPM 1,000,000. - [qwen3-coder-flash Code Generation Model](https://help.aliyun.com/en/model-studio/qwen3-coder-flash.md): Qwen3-based code generation model with multi-turn tool interaction and repository-level code understanding, supporting up to 1M-token context. Pricing in China (Beijing) starts at CNY 1 per million input tokens, CNY 0.2 for cache hits; also deployed in Singapore, Frankfurt, and Virginia. - [qwen3-coder-next Code Generation Model](https://help.aliyun.com/en/model-studio/qwen3-coder-next.md): Qwen3 next-generation code generation model optimized for repository-level understanding, multi-turn tool interaction, and agentic coding tools. Context length 262144 tokens, max output 65536 tokens. Beijing region input ≤32k at CNY 1 per million tokens, output at CNY 4 per million tokens with tiered pricing. - [qwen3-coder-plus Code Generation Model](https://help.aliyun.com/en/model-studio/qwen3-coder-plus.md): Alibaba Cloud Model Studio Qwen3 code generation model with Coding Agent, Function Calling, and context caching. Max input 997,952 tokens, output 65,536 tokens. China (Beijing) pricing starts at CNY 4 per million input tokens, CNY 16 per million output tokens, and CNY 0.8 for cache hits. - [qwen-coder-plus Code Generation Model](https://help.aliyun.com/en/model-studio/qwen-coder-plus.md): Alibaba Cloud Bailian Qwen code model qwen-coder-plus for programming and code generation. Context window 131,072 tokens. Input ¥3.5/million tokens, output ¥7/million tokens. RPM 1,200, TPM 1,000,000. - [qwen-coder-turbo Code Generation Model](https://help.aliyun.com/en/model-studio/qwen-coder-turbo.md): Qwen Coder turbo model for code generation with fast inference and low cost. Max input 129,024 tokens, output 8,192 tokens; RPM 1,200 / TPM 1,000,000. Function Calling, web search, and context caching not supported. Pricing: ¥2 per million input tokens, ¥6 per million output tokens. - [qwen-mt-plus Translation Model](https://help.aliyun.com/en/model-studio/qwen-mt-plus.md): Qwen3-based flagship translation model supporting 92 languages with 16384-token context. Offers terminology customization, format restoration, and domain prompting. Priced at CNY 1.8 per million input tokens and CNY 5.4 per million output tokens (China North 2 Beijing), RPM 60. - [qwen-mt-flash Lightweight Translation Model](https://help.aliyun.com/en/model-studio/qwen-mt-flash.md): Qwen3-based lightweight text translation model supporting 92 languages with 16,384-token context. Pricing: Beijing/Frankfurt/Virginia input ¥0.7/M tokens, output ¥1.95/M tokens; Singapore input ¥1.174/M tokens, output ¥3.596/M tokens. RPM 60, TPM 35,000–100,000. - [qwen-mt-lite Basic Text Translation Model](https://help.aliyun.com/en/model-studio/qwen-mt-lite.md): Qwen3-based text translation model supporting 32 languages with 16,384-token context. Priced at CNY 0.6 per million input tokens and CNY 1.6 per million output tokens in Beijing/Frankfurt/Virginia, with RPM 60 and TPM 100,000. - [qwen-mt-lite-us Entry-level Translation Model (31 Languages, Pricing & Limits)](https://help.aliyun.com/en/model-studio/qwen-mt-lite-us.md): Entry-level Qwen3 translation model: 31 languages, term customization, domain prompts. US Virginia: ¥0.881/¥2.642 per M tokens in/out, 16K context, RPM 60/TPM 100K. - [qwen-mt-turbo Lightweight Text Translation Model](https://help.aliyun.com/en/model-studio/qwen-mt-turbo.md): Qwen3-based lightweight translation model supporting 92 languages with max input/output of 8,192 tokens. Offers terminology customization, format preservation, and domain prompting. Beijing region pricing: input ¥0.7/million tokens, output ¥1.95/million tokens; RPM 60, TPM 35,000. - [qwen-doc-turbo Document Understanding Model](https://help.aliyun.com/en/model-studio/qwen-doc-turbo.md): Alibaba Cloud Model Studio qwen-doc-turbo text-in/text-out model for document information extraction, tagging, content moderation, and summarization. Context window 262,144 tokens, max output 8,192 tokens, supports context caching. China North 2 (Beijing): input ¥0.6/million tokens, output ¥1/million tokens; RPM 600, TPM 3,000,000. - [qwen-deep-research Model](https://help.aliyun.com/en/model-studio/model-qwen-deep-research.md): Agent system for complex research tasks with multi-round reasoning and global planning. Automatically decomposes tasks and generates traceable research reports. Context length 1M tokens; input ¥54/M tokens, output ¥163/M tokens; RPM 120, TPM 1.2M. - [tongyi-intent-detect-v3 Intent Recognition and Slot Filling Model](https://help.aliyun.com/en/model-studio/tongyi-intent-detect-v3.md): Bailian intent detection model that returns intent and slots parameters in a single API call with standard JSON output. Supports up to 8192 input tokens and 4096 output tokens. Pricing: 0.4 CNY per million input tokens, 1 CNY per million output tokens. Rate limits: 1200 RPM, 1M TPM. - [farui-plus Legal Industry Model](https://help.aliyun.com/en/model-studio/farui-plus.md): Tongyi Farui legal LLM for legal Q&A, contract clause review, case recommendation, case analysis, and legal document generation. Context window 12,000 tokens; input/output priced at CNY 20 per million tokens; RPM 240 / TPM 1,000,000. - [tongyi-xiaomi-analysis-flash Dialogue Analysis Model](https://help.aliyun.com/en/model-studio/tongyi-xiaomi-analysis-flash.md): Tongyi Xiaomi dialogue analysis flash edition for conversation info extraction and scene classification in offline/online tasks. Context window 32,768 tokens (input 28,672 / output 4,096), RPM 600, TPM 1M. Pricing: input ¥0.2 per million tokens, output ¥0.4 per million tokens. - [Tongyi Xiaomi-tongyi-xiaomi-analysis-pro Dialogue Analysis Pro Model](https://help.aliyun.com/en/model-studio/tongyi-xiaomi-analysis-pro.md): tongyi-xiaomi-analysis-pro targets complex business logic quality inspection with custom fine-grained analysis criteria and multi-turn context modeling. Max input 28,672 tokens, output 4,096 tokens. China North 2 (Beijing): input CNY 1/million tokens, output CNY 2.7/million tokens; RPM 600, TPM 1,000,000. - [qwen3-coder-30b-a3b-instruct Code Generation Model](https://help.aliyun.com/en/model-studio/qwen3-coder-30b-a3b-instruct.md): Qwen3-based code generation model inheriting coding agent capabilities from Qwen3-Coder-480B-A35B-Instruct. Supports 262,144-token context with up to 65,536 output tokens. Input pricing starts at CNY 1.5 per million tokens for ≤32k inputs; RPM 600, TPM 1,000,000. Function Calling, structured output, and web search are not supported. - [qwen3-coder-480b-a35b-instruct Code Generation Model](https://help.aliyun.com/en/model-studio/qwen3-coder-480b-a35b-instruct.md): Qwen3-based code generation model with Coding Agent capability. 262,144-token context, up to 65,536 output tokens. Deployed in Beijing/Singapore/Frankfurt/Virginia. Input pricing from ¥6 per million tokens for ≤32k - [qwen3.8-2.4t-a95b Model (2.4T Params, 1M Context, Pricing & Limits)](https://help.aliyun.com/en/model-studio/qwen3-8-2-4t-a95b.md): Qwen flagship open-source model with sparse MoE (2.4T total / 95B active params) and 1M-token context. Text in/out; supports Function Calling, web search, structured output, and context caching. Available in Beijing and Singapore; input from CNY 12 per million tokens, RPM 5,000, TPM 5M. - [qwen3.8-27b Vision-Language Model (Capabilities/Pricing/Limits)](https://help.aliyun.com/en/model-studio/qwen3-8-27b.md): qwen3.8-27b vision-language Dense model: image/video/text input, Function Calling, web search, context cache; 1M context; Beijing ¥3/¥12 per 1M tokens (in/out), 5K RPM. - [qwen3.6-27b Native Vision-Language Dense Model](https://help.aliyun.com/en/model-studio/qwen3-6-27b.md): Qwen3.6-series 27B native vision-language Dense model with Image/Text/Video input and Text output. Context window 262,144 tokens (thinking mode: max input 258,048, output 65,536). Enhanced Agentic coding and STEM reasoning; improved spatial intelligence, object detection, video understanding, and document OCR. Supports Function Calling, structured output, web search, and prefix completion; no context caching or batch inference. Beijing region pricing: input ¥3/million tokens, output ¥18/million tokens; RPM 600, TPM 1,000,000. - [qwen3.6-35b-a3b Native Vision-Language Model](https://help.aliyun.com/en/model-studio/qwen3-6-35b-a3b.md): Qwen3.6-series 35B-A3B native vision-language model combining linear attention and sparse MoE architecture. Supports Image/Text/Video input, Function Calling, web search, and structured output. Context window 262,144 tokens with max output 65,536 tokens. Beijing region pricing: input ¥1.8/million tokens, output ¥10.8/million tokens; RPM 600, TPM 1,000,000. - [qwen3.5-122b-a10b Native Vision-Language Model](https://help.aliyun.com/en/model-studio/qwen3-5-122b-a10b.md): Qwen3.5-series 122B-A10B hybrid-architecture vision-language model accepting Text/Image/Video input and producing Text output, with Function Calling, structured output, and web search. Context window 262,144 tokens; max output 65,536 tokens. In China North 2 (Beijing), pricing is CNY 0.8 per million input tokens (≤128k) and CNY 6.4 per million output tokens; CNY 2 / CNY 16 for 128k–256k inputs. RPM 600, TPM 1,000,000. - [qwen3.5-27b Native Vision-Language Model](https://help.aliyun.com/en/model-studio/qwen3-5-27b.md): Qwen3.5-series 27B Dense vision-language model with linear attention, accepting Text/Image/Video input and Text output, 262144-token context. Supports Function Calling, structured output, web search, prefix completion, and fine-tuning across Beijing/Singapore/Frankfurt/Virginia regions, from CNY 0.6 per million input tokens - [qwen3.5-35b-a3b Native Vision-Language Model](https://help.aliyun.com/en/model-studio/qwen3-5-35b-a3b.md): Qwen3.5 series 35B-A3B hybrid-architecture vision-language model accepting Text/Image/Video input with Text output. Supports Function Calling, structured output, and web search. Context window 262144 tokens, max output 65536 tokens. Pricing in China (Beijing): CNY 0.4 per million input tokens for ≤128k context - [qwen3.5-397b-a17b Native Vision-Language Model](https://help.aliyun.com/en/model-studio/qwen3-5-397b-a17b.md): Qwen3.5 series 397B-A17B native vision-language model accepting Text/Image/Video input and producing Text output. Supports Function Calling, structured output, web search, and prefix completion; context window 262,144 tokens with max output 65,536 tokens. In China (Beijing), input ≤128k costs CNY 1.2 per million tokens and output CNY 7.2 per million tokens; RPM 600, TPM 1,000,000. - [qwen3-235b-a22b Model Capabilities and Pricing](https://help.aliyun.com/en/model-studio/qwen3-235b-a22b.md): Bailian qwen3-235b-a22b text model with thinking/non-thinking mode switching, 131072-token context. China North 2 pricing: input CNY 2/million tokens, output CNY 8/million tokens, RPM 600. - [qwen3-235b-a22b-instruct-2507 Model Capabilities and Pricing](https://help.aliyun.com/en/model-studio/qwen3-235b-a22b-instruct-2507.md): Qwen3 non-thinking open-source model with text input/output. Supports Function Calling, structured output, and prefix completion. Context length 131,072 tokens. Beijing region: input ¥2/million tokens, output ¥8/million tokens, RPM 600, TPM 1,000,000. - [qwen3-235b-a22b-thinking-2507 Thinking Mode Model](https://help.aliyun.com/en/model-studio/qwen3-235b-a22b-thinking-2507.md): Alibaba Cloud Model Studio Qwen3 thinking-mode open-source model with 131,072-token context. Supports up to 32,768 output tokens and 81,920 chain-of-thought tokens in thinking mode for complex reasoning. Priced at CNY 2 per million input tokens and CNY 20 per million output tokens (Beijing); also deployed in Singapore, Frankfurt, and Virginia. - [qwen3-30b-a3b Model Capabilities and Pricing](https://help.aliyun.com/en/model-studio/qwen3-30b-a3b.md): qwen3-30b-a3b integrates thinking and non-thinking modes with 131,072-token context and 8,192-token max output. Supports Function Calling, structured output, and prefix completion. Beijing region: input ¥0.75/M tokens, output ¥3/M tokens, thinking output ¥7.5/M tokens; RPM 600, TPM 1,000,000. - [qwen3-30b-a3b-instruct-2507 Open-Source Text Model](https://help.aliyun.com/en/model-studio/qwen3-30b-a3b-instruct-2507.md): Qwen3-based open-source model in non-thinking mode with 131,072-token context and 32,768 max output. Supports Function Calling, structured output, prefix completion, and fine-tuning (China North 2 Beijing only). Input ¥0.75/million tokens, output ¥3/million tokens. - [qwen3-30b-a3b-thinking-2507 Thinking Open-Source Model](https://help.aliyun.com/en/model-studio/qwen3-30b-a3b-thinking-2507.md): Qwen3 thinking-mode open-source model with 81,920-token context, max input 126,976 and output 32,768 tokens. Excels at logic, math, and code reasoning with multilingual translation. Priced at CNY 0.75/M input and 7.5/M output tokens in Beijing/Frankfurt/Virginia, RPM 600. - [qwen3-32b Qwen 32B Large Language Model](https://help.aliyun.com/en/model-studio/qwen3-32b.md): qwen3-32b integrates thinking and non-thinking modes, outperforming QwQ in reasoning and Qwen2.5-32B-Instruct in general tasks. Context window 131,072 tokens (max input 98,304, output 8,192). Supports Function Calling, structured output, prefix completion, and fine-tuning (Beijing only). Pricing: ¥2 per million input tokens, ¥8 per million output tokens. - [qwen3-14b Model Capabilities and Pricing](https://help.aliyun.com/en/model-studio/qwen3-14b.md): qwen3-14b text LLM with thinking/non-thinking modes, 131,072-token context (max input 98,304 / output 8,192). Supports Function Calling, structured output, prefix completion, and fine-tuning. China North 2 pricing: ¥1/M input tokens, ¥4/M output tokens; RPM 600, TPM 1,000,000. - [qwen3-8b Model Capabilities and Pricing](https://help.aliyun.com/en/model-studio/qwen3-8b.md): Alibaba Cloud Model Studio qwen3-8b text generation model with switchable thinking/non-thinking modes and 131,072-token context (max input 98,304, output 8,192). Priced at CNY 0.5 per million input tokens and CNY 2 per million output tokens in Beijing/Frankfurt/Virginia. Function Calling available only in Beijing region. RPM 600, TPM 1,000,000. - [qwen3-next-80b-a3b-instruct Text Generation Model](https://help.aliyun.com/en/model-studio/qwen3-next-80b-a3b-instruct.md): Next-generation Qwen3-based open-source text generation model in non-thinking mode. Context length 131,072 tokens with max output 32,768 tokens. Priced at CNY 1 per million input tokens and CNY 4 per million output tokens in China (Beijing). Function Calling, web search, batch inference, and fine-tuning are not supported. - [qwen3-next-80b-a3b-thinking Open-Source Thinking Model](https://help.aliyun.com/en/model-studio/qwen3-next-80b-a3b-thinking.md): Next-generation Qwen3-based open-source thinking model with improved instruction following and more concise responses. Text input/output, 131,072-token context, max 32,768 output tokens. Deployed in Beijing, Singapore, Frankfurt, Virginia; thinking input from CNY 1 per million tokens. - [deepseek-v4-pro Flagship MoE Model](https://help.aliyun.com/en/model-studio/deepseek-v4-pro.md): deepseek-v4-pro on Alibaba Cloud Model Studio: 1.6T total params, 49B active, native million-token context. Excels in math logic, complex reasoning, coding, and long-text analysis for research, office, and agent scenarios. Pricing: input ¥12/M tokens, output ¥24/M tokens, cache hit ¥1/M tokens. - [deepseek-v4-pro-us US-Deployed Model Specs and Pricing](https://help.aliyun.com/en/model-studio/deepseek-v4-pro-us.md): Alibaba Cloud Model Studio deepseek-v4-pro-us flagship MoE model with 1.6T total parameters and 49B active, supporting million-token context (1M input / 393,216 output). Delivers top-tier math logic, complex reasoning, coding, and long-text analysis for research, office, and agent scenarios. Supports Function Calling, structured output, web search, and context caching. Deployed in US (Virginia): input ¥17.986/M tokens, output ¥35.972/M tokens, cache hit ¥1.499/M tokens; RPM 10,000, TPM 1,200,000. - [deepseek-v4-flash Lightweight MoE Model](https://help.aliyun.com/en/model-studio/deepseek-v4-flash.md): Bailian-hosted DeepSeek-V4-Flash MoE model with 284B total / 13B active parameters, supporting 1M-token context, Function Calling, structured output, web search, and context caching. Ideal for high-concurrency lightweight tasks like daily chat, content creation, and basic RAG with fast inference and low cost. - [deepseek-v4-flash-us Model Overview and Pricing](https://help.aliyun.com/en/model-studio/deepseek-v4-flash-us.md): Alibaba Cloud Model Studio deepseek-v4-flash-us is a MoE text model with 284B total parameters and 13B active, natively supporting 1M-token context (max input 1,000,000; max output 393,216). It supports Function Calling, structured output, web search, and context caching. Deployed in US (Virginia), priced at CNY 1.499 per million input tokens, CNY 2.998 per million output tokens, and CNY 0.3 per million cached input tokens; RPM 10,000 and TPM 1,200,000. - [deepseek-v3.2 Model Capabilities and Pricing](https://help.aliyun.com/en/model-studio/deepseek-v3-3.md): DeepSeek-V3.2 introduces sparse attention and supports tool calling in both thinking and non-thinking modes. Context length 131,072 tokens with max output 65,536. Beijing pricing: input ¥2/million tokens, output ¥3/million tokens; Singapore: input ¥4.272, output ¥12.815. RPM up to 15,000, TPM up to 1.2 million. - [deepseek-v3.2-exp Experimental Model](https://help.aliyun.com/en/model-studio/deepseek-v3-2-exp.md): Bailian-hosted deepseek-v3.2-exp experimental text model using DeepSeek Sparse Attention for long-context training and inference optimization. Supports Function Calling and web search. Context window 131,072 tokens (input 98,304 / output 65,536), RPM 15,000, TPM 1.2M. Pricing: CNY 2 per million input tokens, CNY 3 per million output tokens. - [vanchin/deepseek-v3.2-think Model Capabilities and Pricing](https://help.aliyun.com/en/model-studio/deepseek-v3-2-think.md): DeepSeek-V3.2 thinking model served by Kuaishou Wanchin with 131,072-token context. Input ¥2/million tokens, output ¥3, cache hit ¥0.2. Supports Function Calling, structured output, prefix completion, and context caching. RPM 30, TPM 600,000 - [DeepSeek-V3.1 Hybrid Reasoning Model](https://help.aliyun.com/en/model-studio/deepseek-v3-1.md): Bailian-hosted DeepSeek-V3.1 text generation model with thinking/non-thinking modes, Function Calling, web search, and context caching. Max input 98,304 tokens, output 65,536 tokens. China (Beijing) pricing: input ¥4/M tokens, output ¥12/M tokens, cache hit ¥0.8/M tokens; RPM 15,000, TPM 1.2M. - [siliconflow/deepseek-v3.1-terminus Hybrid Agent Model](https://help.aliyun.com/en/model-studio/deepseek-v3-1-terminus.md): DeepSeek-V3.1-Terminus hybrid agent LLM with switchable thinking/non-thinking modes, enhanced code and search agent capabilities. 163,840-token context, input ¥4/million tokens, output ¥12/million tokens, RPM 500. - [DeepSeek-V3 Text Generation Model](https://help.aliyun.com/en/model-studio/deepseek-v3.md): Bailian-hosted DeepSeek-V3 MoE model with 671B parameters (37B active) and 65,536-token context. Supports Function Calling, web search, and context caching. Pricing: ¥2 per million input tokens, ¥8 per million output tokens; RPM 15,000. - [Model Studio - DeepSeek-R1 Reasoning Model API and Pricing](https://help.aliyun.com/en/model-studio/deepseek-r1.md): Model Studio-hosted DeepSeek-R1 text reasoning model with 131,072-token context. Supports Function Calling, web search, and context caching. Input ¥4/million tokens, output ¥16/million tokens, cache hit ¥0.8/million tokens; RPM 15,000, TPM 1.2M. Includes deepseek-r1-0528 snapshot. - [deepseek-r1-distill-qwen-1.5b Model Overview](https://help.aliyun.com/en/model-studio/deepseek-r1-distill-qwen-1-5b.md): Text LLM distilled from Qwen2.5-Math-1.5B. Max input 32768 tokens, max output 16384 tokens. Rate limit RPM 60 / TPM 100000. Function Calling and web search not supported. - [DeepSeek-R1-Distill-Qwen-14B Model Overview and Pricing](https://help.aliyun.com/en/model-studio/deepseek-r1-distill-qwen-14b.md): Text LLM distilled from Qwen2.5-14B with DeepSeek R1 outputs. Max input 32,768 tokens, max output 16,384 tokens. In China (Beijing): input CNY 1 per million tokens, output CNY 3 per million tokens; RPM 15,000, TPM 1.2M. Function Calling, web search, and fine-tuning are not supported. - [deepseek-r1-distill-qwen-32b Distilled Model](https://help.aliyun.com/en/model-studio/deepseek-r1-distill-qwen-32b.md): DeepSeek R1 distilled text-generation model based on Qwen2.5-32B with 32,768-token context and 16,384-token max output. Function Calling, web search, structured output, and batch inference are not supported. In China (Beijing), input is CNY 2 per million tokens and output is CNY 6 per million tokens, with RPM 15,000 and TPM 1.2 million. - [deepseek-r1-distill-qwen-7b Model Capabilities and Pricing](https://help.aliyun.com/en/model-studio/deepseek-r1-distill-qwen-7b.md): Distilled reasoning model based on Qwen2.5-Math-7B. Text-only input/output; Function Calling, web search, and context caching are not supported. Max input 32,768 tokens, max output 16,384 tokens. In China (Beijing): input ¥0.5 per million tokens, output ¥1 per million tokens; RPM 15,000, TPM 1,200,000. - [vanchin/deepseek-ocr Document Recognition and Image-to-Text Model](https://help.aliyun.com/en/model-studio/deepseek-ocr.md): Vision-text compression model inferred by Kuaishou VanChin. Supports Text+Image input and structured output with 8,192-token context. Pricing: 0.216 CNY per million tokens for both input and output; RPM 500, TPM 1,000,000 - [Model Studio - kimi/kimi-k3 Capabilities and Pricing](https://help.aliyun.com/en/model-studio/kimi-k3.md): kimi/kimi-k3 is Moonshot AI's 2.8-trillion-parameter flagship model with native Text/Image/Video input, 1M-token context window, and support for Function Calling, structured output, prefix completion, and context caching. In China North 2 (Beijing): input ¥20/M tokens, output ¥100/M tokens, cache hit ¥2/M tokens; RPM 500, TPM 3M. - [kimi-k3 Model (2.8T Parameters, 1M-Token Context, Pricing)](https://help.aliyun.com/en/model-studio/aliyun-kimi-k3.md): Kimi open-source flagship: 2.8T params, 1M-token context, vision; Function Calling, structured output, context cache. Beijing: ¥20 in / ¥100 out per million tokens. - [kimi-k2.7-code Multimodal Coding Model](https://help.aliyun.com/en/model-studio/kimi-k2-7-code.md): kimi-k2.7-code accepts text, image, and video inputs with thinking mode. Context window 262,144 tokens. Supports Function Calling, web search, structured output, prefix completion, and context caching. Beijing region: input ¥6.5/M tokens, output ¥27/M tokens, RPM 500. - [kimi-k2.7-code-highspeed High-Speed Coding Model](https://help.aliyun.com/en/model-studio/kimi-k2-7-code-highspeed.md): K2.7 Code High-Speed Edition inferred by Moonshot AI delivers ~180 Token/s output (up to 260 Token/s for short context), 5-6× faster than the standard edition. Supports Text/Image/Video input, Function Calling, structured output, prefix completion, and context caching with a 262,144-token context window. Pricing: input ¥13/M tokens, output ¥54/M tokens, cache hit ¥2.6/M tokens; RPM 500, TPM 3,000,000. - [kimi-k2.6 Multimodal Large Model](https://help.aliyun.com/en/model-studio/kimi-k2-6.md): Latest Kimi multimodal model on Bailian platform supporting text, image, and video input with Function Calling. Context window of 262,144 tokens. Pricing: input ¥6.5/M tokens, output ¥27/M tokens, cache hit ¥1.3/M tokens. Rate limits: 500 RPM, 1M TPM. - [kimi-k2.5 Multimodal Large Model](https://help.aliyun.com/en/model-studio/kimi-k2-5.md): Moonshot AI native multimodal model supporting text/image/video input and thinking mode, with 262,144-token context. Function Calling and context caching available. Input ¥4/million tokens, output ¥21/million tokens, RPM 500. - [Moonshot-Kimi-K2-Instruct Model Overview and Pricing](https://help.aliyun.com/en/model-studio/moonshot-kimi-k2-instruct.md): Moonshot open-source trillion-parameter MoE model with 32B active parameters. Supports Function Calling and context caching, 131,072-token context length. Pricing: input ¥4/M tokens, output ¥16/M tokens, cache hit ¥0.8/M tokens. Rate limits: 500 RPM, 1M TPM. - [kimi-k2-thinking Reasoning Model](https://help.aliyun.com/en/model-studio/kimi-k2-thinking.md): Agentic reasoning model from Moonshot AI with Function Calling and context caching. Max input 229,376 tokens, output 16,384 tokens, context 262,144 tokens. Beijing region pricing: input ¥4/M tokens, output ¥16/M tokens, cache hit ¥0.8/M tokens; RPM 500, TPM 1,000,000 - [glm-5.2-fast-preview Model Overview](https://help.aliyun.com/en/model-studio/glm-5-2-fast.md): High-speed variant of Zhipu AI's flagship GLM-5.2 model with 1M context window and output TPS 1.5–2× the standard version. Supports logical reasoning, long-text comprehension, and code generation, with Function Calling, structured output, web search, and context caching. Suited for real-time chat, multi-turn agent calls, and streaming code generation. Pricing: input ¥16 per million tokens, output ¥56, cache hit ¥4. - [glm-5.2 Model Capabilities and Pricing](https://help.aliyun.com/en/model-studio/glm-5-2.md): Zhipu AI flagship model glm-5.2 with 1M context (max input 1,048,576 tokens, output 131,072 tokens), featuring logical reasoning, long-text understanding, and code generation. Deployed in Beijing/Singapore/Frankfurt/Virginia/Tokyo; supports Function Calling, structured output, and context caching; pay-as-you-go pricing. - [glm-5.2-us Model Capabilities and Pricing](https://help.aliyun.com/en/model-studio/glm-5-2-us.md): Zhipu AI flagship text model with 1M context (max input 1,048,576 tokens / output 131,072 tokens), offering logical reasoning, long-text understanding, and code generation. Supports Function Calling, structured output, prefix continuation, and context caching; does not support web search or batch inference. Deployed in US (Virginia). Pricing: input ¥10.492/M tokens, output ¥32.974/M tokens, cache hit ¥2.098/M tokens; RPM 500, TPM 1,000,000. - [glm-5.1 Model Overview and Pricing](https://help.aliyun.com/en/model-studio/glm-5-1.md): Zhipu AI glm-5.1 model with 744B parameters, supporting 200K context window and up to 128K output tokens. Features logical reasoning, long-text comprehension, and code generation, with Function Calling, structured output, and context caching. Input pricing starts at CNY 6 per million tokens for ≤32k, CNY 8 per million tokens for 32k–200k, and cache hits as low as CNY 0.6 per million tokens. - [glm-5 Model Capabilities and Pricing](https://help.aliyun.com/en/model-studio/glm-9.md): A 744B-parameter model for coding and agent scenarios, supporting Function Calling and context caching with up to 169,984 input tokens and 16,384 output tokens. Pay-as-you-go pricing in China (Beijing): ≤32k input at CNY 4/million tokens, output at CNY 18/million tokens; >32k–200k input at CNY 6/million tokens, output at CNY 22/million tokens. RPM 500, TPM 1 million. - [glm-4.7 Text Generation Model](https://help.aliyun.com/en/model-studio/glm-4-7.md): Zhipu flagship LLM with 355B parameters, supporting Function Calling and context caching. Max input 169,984 tokens, output 16,384 tokens. Excels at long-horizon task planning, coding, and tool orchestration. Pricing from CNY 3 per million tokens (input ≤32k), RPM 500 / TPM 1M. - [glm-4.6 Model Capabilities and Pricing](https://help.aliyun.com/en/model-studio/glm-4-6.md): GLM next-gen flagship text model on Bailian with 355B total / 32B active parameters and 200K context. Supports context caching; no Function Calling, web search, or structured output. Input ≤32K tokens: CNY 3 per million; >32K: CNY 4 per million. RPM 60, TPM 1M. - [glm-4.5 Model Overview and Pricing](https://help.aliyun.com/en/model-studio/glm-4.md): Alibaba Cloud Model Studio glm-4.5 MoE model with 355B total / 32B active parameters, supporting Text input/output and 131,072-token context (max input 98,304, output 16,384). Pay-as-you-go pricing in China (Beijing): ≤32k input ¥3/million tokens, output ¥14; 32k–128k input ¥4, output ¥16. RPM 60, TPM 1,000,000. - [glm-4.5-air Model Overview and Pricing](https://help.aliyun.com/en/model-studio/glm-4-5-air.md): Alibaba Cloud Model Studio glm-4.5-air text generation model with MoE architecture (106B total / 12B active parameters). Context window 131,072 tokens (max input 98,304, output 16,384). Function Calling, web search, and structured output are not supported. In China (Beijing), pricing for input ≤32k is CNY 0.8 per million tokens (input) and CNY 6 (output); for 32k25,000 images. RPM limit 600. - [Model Studio-facechain-facedetect Face Detection](https://help.aliyun.com/en/model-studio/facechain-facedetect.md): facechain-facedetect checks whether portrait images meet facechain fine-tuning standards across face count, size, angle, lighting, and clarity. Supports batch image group input with per-image results. RPM 300 in China (Beijing). - [facechain-generation Portrait Photo Generation](https://help.aliyun.com/en/model-studio/model-facechain-generation.md): Generate portrait photos from trained person likenesses with preset styles including ID photos and business portraits. Image generation costs CNY 0.18 per image in China (Beijing) region, with RPM 120 and TPM 3,000. - [image-erase-completion Image Erase and Completion](https://help.aliyun.com/en/model-studio/image-erase-completion.md): Alibaba Cloud Model Studio image-erase-completion removes people, pets, objects, text, or watermarks via mask while preserving background; prompt-based removal is not supported. Beijing region RPM 120; no public pricing available. - [Bailian image-instance-segmentation Person Instance Segmentation Model](https://help.aliyun.com/en/model-studio/model-image-instance-segmentation.md): The image-instance-segmentation model detects persons in images and generates pixel-level masks outlining each object's boundary. Input and output are both Image modality. Rate-limited to 120 RPM in China (Beijing), with inference served by Alibaba Cloud Bailian. - [Bailian image-out-painting Image Outpainting Model](https://help.aliyun.com/en/model-studio/image-out-painting.md): image-out-painting freely extends input images with rotation support, expanding by ratio or pixel count via specified width/height ratios or top/bottom/left/right pixels. Rate-limited at 120 RPM and 60 TPM in China (Beijing), suitable for creative entertainment, assisted drawing, and post-production. - [Model Studio-kling-v3-image-generation Kling V3 Image Generation Model](https://help.aliyun.com/en/model-studio/kling-v3-image-generation.md): Kling AI text-to-image and image-to-image model supporting up to 10 reference images to lock subject, elements, and color tone. Integrates style transfer, portrait reference, multi-image fusion, and local repaint. Pricing: CNY 0.2 per image (1K/2K), RPM limit 300. - [kling-v3-omni-image-generation Kling Image Generation Model](https://help.aliyun.com/en/model-studio/kling-v3-omni-image-generation.md): Kling AI image generation model on Alibaba Cloud Model Studio accepts text and reference images to produce 2K/4K cinematic visuals for storyboards, concept art, and scene design. Pricing: ¥0.2 per 1K/2K image, ¥0.4 per 4K image; RPM limit 300. - [Model Studio - vidu-image_reference2image Reference-to-Image and Image Editing](https://help.aliyun.com/en/model-studio/vidu-image-reference2image.md): Vidu model accepts 0-14 reference images or text prompts for reference-to-image, text-to-image, and image editing with precise CJK/Latin text rendering and pixel-level UI/chart restoration for posters and infographics. 1K/2K/4K generation priced at CNY 0.625-1.47 per image, RPM limit 300. - [Model Studio - viduq2-fast_reference2image Reference-to-Image and Image Editing](https://help.aliyun.com/en/model-studio/viduq2-fast-reference2image.md): Vidu-supplied viduq2-fast image generation model supporting 0-14 reference images or text prompts for reference-to-image, text-to-image, and image editing. Priced at CNY 0.28125 per 1K image in China (Beijing), with RPM limit of 300. - [viduq2-pro_reference2image Reference-to-Image Model](https://help.aliyun.com/en/model-studio/viduq2-pro-reference2image.md): Vidu reference-to-image model accepts 0-14 reference images or text prompts to generate images, supporting reference-to-image, text-to-image, and image editing. Pricing: ¥0.9375 per image at 1K/2K, ¥1.71875 at 4K; RPM limit 300. - [viduq3-fast_reference2image Reference-to-Image Model](https://help.aliyun.com/en/model-studio/viduq3-fast-reference2image.md): Vidu high-speed image generation model supporting 0-14 reference images or text input for reference-to-image, text-to-image, and image editing. Cost ~50% lower than Pro; 1K/2K/4K at CNY 0.47-1.09 per image with 300 RPM limit. - [Model Studio 3D Generation Model List (Tripo Text/Image-to-3D)](https://help.aliyun.com/en/model-studio/model-list-3d-generation.md): Catalog of 3D generation models on Model Studio: Tripo text-to-3D and image-to-3D, outputting GLB-format 3D models with PBR materials and rendered preview images via task_id polling. - [Tripo-H3.1 High-Precision 3D Generation Model](https://help.aliyun.com/en/model-studio/tripo-h3-1.md): Bailian-integrated Tripo H3.1 (20B parameters) for text/image-to-3D generation with billion-voxel resolution and up to 2M polygons. Standard/ultra-HD options available, RPM limit 5/min, starting at CNY 0.7/call - [Tripo-P1.0 3D Generation Model](https://help.aliyun.com/en/model-studio/tripo-p1-0.md): Tripo P1.0 on Alibaba Cloud Model Studio generates professional-topology 3D assets from text or images in ~2 seconds for gaming and Web3D real-time scenarios. Supports Function Calling, structured output, and batch inference with RPM limit of 5. Text-to-3D without texture costs CNY 2.1/request, with HD texture CNY 3.5/request; image-to-3D without texture CNY 2.8/request, with HD texture CNY 4.2/request. - [Model Studio FAQ](https://help.aliyun.com/en/model-studio/faq-about-alibaba-cloud-model-studio.md): Billing (pricing, invoices, subscription), API/SDK errors (error code 100004, doc_reference_type), service activation, data privacy, and Qwen vs Model Studio differences - [Model Studio Service Agreement, SLA, and Model Usage Terms](https://help.aliyun.com/en/model-studio/related-agreements.md): Model Studio legal agreements: service agreement, model inference SLA, trial feature notes, open-source model terms, third-party model service terms list. - [Model Studio Release Notes (New Models, Deprecation, Platform Updates)](https://help.aliyun.com/en/model-studio/release-notes.md): Model Studio changelog hub: newly released and updated models, model deprecation policy, and platform feature updates. Check here for model changes and new features. - [Model Deprecation Policy](https://help.aliyun.com/en/model-studio/model-depreciation.md): Alibaba Cloud Model Studio deprecates legacy models: snapshot models (e.g., qwen-max-2025-01-25) with 30-day notice, mainline models with 3-month notice. After deprecation, inference, fine-tuning, and deployment are discontinued; already trained or deployed models remain unaffected. - [Model Studio Newly Released Models (Launch Timeline & Model IDs)](https://help.aliyun.com/en/model-studio/newly-released-models.md): Timeline of new models on Model Studio: launch dates and IDs for qwen3.8-max, deepseek-v4, wan3.0-video, qwen-image-3.0, kimi-k3 by region; retirement rules linked. - [Model Studio Release Notes (Price Cuts, Token Plan, Coding Plan)](https://help.aliyun.com/en/model-studio/model-release-notes.md): Latest Model Studio updates: price cuts and retirement notices for qwen/GLM models, Token Plan benefit upgrades, Coding Plan Pro first-month discount, Memory Store and Managed Agent commercialization, API Key encrypted storage, Codex client integration, knowledge retrieval service launch. ## Model Studio Application User Guide (App Dev, RAG, MCP, Calling) - [Model Studio Application User Guide (App Dev, RAG, MCP, Calling)](https://help.aliyun.com/en/model-studio/application-user-guide.md): Hub for application docs: get started, Managed Agents, app development and Prompt, knowledge base (RAG), MCP/plugins/Skill, app calling and publishing, evaluation and monitoring. - [Get Started with Model Studio Apps (Agent App/RAG/Publish)](https://help.aliyun.com/en/model-studio/start-using.md): App-side entry: create agent apps, configure Prompt and RAG knowledge base, connect MCP/plugins/Skills, publish and call via API. Build your first AI app here. - [Build a Knowledge Base Q&A App Without Coding (Agent App + RAG Tutorial)](https://help.aliyun.com/en/model-studio/build-knowledge-base-qa-assistant-without-coding.md): Tutorial: create an agent app (Qwen-Max, system prompt), upload private docs via data connector, build a standard knowledge base with smart chunking, attach it and publish—private-domain QA in about 5 minutes; model calls are billed, free trial quota available. - [Application feature updates](https://help.aliyun.com/en/model-studio/application-release-notes.md): Chronological changelog for Model Studio application features (2024-2026). Covers agent apps, workflow apps, high-code apps, knowledge bases, MCP services, evaluation, and observability updates. - [Managed Agents: Build, Configure & Delegate Tasks](https://help.aliyun.com/en/model-studio/managed-agents.md): Hub for Bailian Managed Agents: overview and quick start, build agents, configure agent environments, delegate tasks, context management, and billing. Start here to create and run a managed agent. - [Managed Agents - Overview of the hosted agent runtime](https://help.aliyun.com/en/model-studio/managed-agents-introduction.md): Alibaba Cloud Model Studio Managed Agents is a server-hosted agent runtime that maintains session state and runs each agent in an isolated sandbox. It supports long-running tasks such as multi-step tool calling, code execution, and file processing. Organized around four core concepts—Agent, Environment, Session, and Event—it provides command execution, file operations, MCP services, and Skills, with event history persisted on the server and delivered via SSE streams. - [Managed Agents-Quick Start](https://help.aliyun.com/en/model-studio/managed-agents-quick-start.md): Create a Managed Agent on Alibaba Cloud Model Studio in 4 steps: configure the agent (models like qwen3.7-plus, 7 built-in tools including bash/read/write), set up a cloud-hosted sandbox environment, bind a session, and invoke via API or preview debugging with SSE event streaming for real-time output. - [Managed Agents - Build an Agent](https://help.aliyun.com/en/model-studio/managed-agents-agent.md): Combine models, system prompts, tools, MCP services, and skills in Model Studio to create an agent, generate an agent ID for multi-session reuse and API invocation - [Managed Agents-Define Agent](https://help.aliyun.com/en/model-studio/managed-agents-agent-definition.md): Configure Bailian Managed Agents with model (qwen3-max etc.), system prompt, 7 built-in tools, MCP servers and skill packs. Each save auto-increments version. Create/archive via API; sessions lock to the creation version; archived agents remain queryable with include_archived=true. - [Managed Agents-Built-in Tool Configuration](https://help.aliyun.com/en/model-studio/managed-agents-builtin-tools.md): Alibaba Cloud Model Studio Managed Agents provides 7 built-in tools including bash, write, read, edit, glob, grep, and download_file. These tools execute shell commands, file operations, and external downloads in a session-bound sandbox environment, with generated files persisted throughout the session lifecycle. - [Agent MCP - Mount and Manage MCP Services](https://help.aliyun.com/en/model-studio/managed-agents-mcp.md): Bailian managed MCP marketplace covers web search, document processing, geospatial visualization, and more. Custom MCP services support plugin, script deployment, AI gateway, and Alibaba Cloud OpenAPI types. Mount services via the agent editor or mcp_servers field in API; execution results display in the event panel. - [Managed Agents - Skill Management](https://help.aliyun.com/en/model-studio/managed-agents-skill.md): Alibaba Cloud Model Studio Managed Agents skill management: marketplace official skills (PDF/Word/Excel/PPT/bailian-cli) and custom zip packages (≤10 MB, must contain SKILL.md). Supports versioning; specify version when attaching to agents. Deleted skills are unrecoverable and existing agent references become invalid. - [Managed Agents-Configure Agent Environment](https://help.aliyun.com/en/model-studio/managed-agents-environment.md): Define the execution sandbox runtime for tool invocation, managed independently from agents and reusable across multiple sessions with security isolation and resource control - [Managed Agents - Cloud-hosted environment configuration](https://help.aliyun.com/en/model-studio/managed-agents-configure-environment.md): Bailian-managed sandbox container supporting apt/pip/npm preinstalled packages, unrestricted network policy, organization scope, and metadata; environments can be archived or hard-deleted, managed via console or API - [Managed Agents - Delegate Tasks to Agent](https://help.aliyun.com/en/model-studio/managed-agents-session.md): A session hosts one agent run instance, recording all messages, tool calls, and state changes as events to track Managed Agents task execution. - [Managed Agents - Create Session](https://help.aliyun.com/en/model-studio/managed-agents-session-event.md): Create a Bailian Managed Agents session by binding agent, environment_id and resources for file mounting. Monitor messages, tool calls and state changes in real time via the event panel; archive to retain history or delete permanently. - [Managed Agents-Manage Sessions](https://help.aliyun.com/en/model-studio/managed-agents-session-operations.md): Bailian managed agent session state machine (idle/running/terminated), tool call approval and interruption mechanisms, plus archive (retains event history) and delete (hard delete, irreversible) API examples - [Managed Agents-Session Event Stream SSE Subscription and Push](https://help.aliyun.com/en/model-studio/managed-agents-event-stream.md): Subscribe to Managed Agents session event stream via SSE to receive message, tool_call, mcp_call, and session_status events in real time. Send messages via POST /sessions/{session_id}/events to trigger agent processing, with Bash/Python/Java SDK examples - [Managed Agents-Credential Vault Authentication](https://help.aliyun.com/en/model-studio/managed-agents-credential.md): Alibaba Cloud Model Studio Managed Agents Credential Vault centrally manages authentication for MCP services and skills. Create vault containers, add API Key/Token credentials as key-value pairs, and reference them by variable name at session runtime; plaintext is hidden after saving. - [Managed Agents-Context Management](https://help.aliyun.com/en/model-studio/managed-agents-context.md): Extend agent session data access by mounting resources that are independent of session lifecycle and reusable across sessions. Supports file upload (≤10 MB per file) mounted at /mnt/session/uploads; in-session modifications do not affect the original resource or copies in other sessions. - [Managed Agents - File Upload and Mounting](https://help.aliyun.com/en/model-studio/managed-agents-file.md): Alibaba Cloud Model Studio Managed Agents file management: max 10 MB per file, 100 GB workspace quota, 30-day retention. Upload via console or API (multipart/form-data); mount to session sandbox with /uploads/ prefix after approval; deletion is irreversible. - [Managed Agents Billing (Runtime Fee/Model Calls/Free Quota)](https://help.aliyun.com/en/model-studio/managed-agents-billing.md): Managed Agents has three separate billing items: session runtime fee at CNY 0.5/hour, model invocation fees based on Qwen and other models' pay-as-you-go rates, and tool/MCP invocation fees billed separately. A free quota of 10 runtime hours is provided (valid for 30 days), applicable only to runtime fees. Commercial billing starts from 2026-08-17. - [Build LLM Apps on Model Studio (Agent Apps & Workflow Orchestration)](https://help.aliyun.com/en/model-studio/llm-application.md): Entry for app building on Model Studio: create and orchestrate agent and workflow applications, combining Prompt, knowledge base (RAG), memory, MCP and plugins to build and debug LLM apps. - [Application types](https://help.aliyun.com/en/model-studio/application-introduction.md): Three app types: Agent (no-code, autonomous planning), Workflow (low-code visual orchestration), High-code (Python). Comparison of development approach, skill level, and use cases. - [Agent 2.0 applications](https://help.aliyun.com/en/model-studio/new-single-agent-application.md): Create Agent 2.0 apps that unify knowledge bases and MCP as tools with autonomous multi-step planning. Configure system prompts, file pre-parsing, and knowledge base label filtering - [Create an agent application](https://help.aliyun.com/en/model-studio/single-agent-application.md): No-code agent builder: configure prompts, connect RAG knowledge bases, add plugins (code interpreter, web search, image gen). Supports API calls, DingTalk/WeChat publishing, version management - [Build Workflow Apps in Model Studio (Node Orchestration & Cases)](https://help.aliyun.com/en/model-studio/workflow-application.md): Build workflow apps in Model Studio: chain Start, LLM, intent-classification and End nodes with API/function nodes; fraud-detection and shopping-guide cases. - [Model Studio High-Code App: Deploy Python AI Services (Serverless/K8s)](https://help.aliyun.com/en/model-studio/rich-code-application.md): Deploy Python AI services as public APIs via Serverless Function or K8s. Upload .whl or use templates, integrate MCP tools, set gateway with custom domain. Billing starts on deploy. - [Pro-code app MCP tool integration](https://help.aliyun.com/en/model-studio/rich-code-application-mcp.md): Integrate tools via MCP: Knowledge Base, MCP Service, Application Component, Data Connector. Connect using fastmcp.Client with StreamableHttpTransport. Includes code examples. - [Pro-code app development guide](https://help.aliyun.com/en/model-studio/rich-code-app-develop-guide.md): Develop and deploy pro-code apps with AgentScope-AI (agentscope-runtime). Entry: main.py with GET /health. Deploy via runtime-fc-deploy. Supports @trace for observability. - [Model Studio File Q&A (Full-text/RAG Retrieval/Custom Processing)](https://help.aliyun.com/en/model-studio/file-q-a.md): Chat with uploaded documents, images, audio/video in agent apps: full-text citation, RAG chunk retrieval, custom processing (plugins/MCP). Up to 10 files at 10MB each per session, Beijing region only; files lost on refresh. Qwen/DeepSeek supported. - [Model Studio Prompt Engineering (Templates/Sample Library/Auto-Optimization)](https://help.aliyun.com/en/model-studio/prompt.md): Prompt engineering hub for Model Studio apps: template overview, custom templates, sample library, auto-optimization and feedback-based optimization. - [Prompt template overview](https://help.aliyun.com/en/model-studio/prompt-template.md): Built-in and custom prompt templates separating fixed structures from variables. Manage via console or CreatePromptTemplate API. Retrieve by ID with GetPromptTemplate - [Custom Prompt Templates (Text/Image, ICIO Framework)](https://help.aliyun.com/en/model-studio/prompt-custom-template.md): Only in China (Beijing) region. Create text/image Prompt templates in Model Studio: custom or ICIO-framework input, plus AI Optimize Prompt enhancement. - [Model Studio Prompt Sample Library (Few-shot, Deprecated – Migrate to RAG Table Store)](https://help.aliyun.com/en/model-studio/prompt-sample-optimization.md): Deprecated—migrate to RAG Table Store. Few-shot: create sample libraries (manual/Excel ≤20MB), link to agent apps (≤5), multi-recall ≤10 into context. 300 max/library; free but adds Token cost. - [Migrate Prompt Example Library to RAG Table Knowledge Base](https://help.aliyun.com/en/model-studio/migrate-sample-library-prompt-to-rag-table-library.md): Export Q&A pairs from Prompt Example Library, import into RAG Table Knowledge Base, update agent config. Removes the 300-pair limit and adds configurable retrieval strategies. - [Automatic prompt optimization](https://help.aliyun.com/en/model-studio/optimize-prompt.md): Auto-rewrite prompts for better structure and clarity. Strategies: restructuring, role assignment, instruction enhancement, safeguard addition. Free to use. Save as reusable templates - [Prompt Feedback Optimization](https://help.aliyun.com/en/model-studio/prompt-feedback-optimization.md): Auto-generate optimized prompts from input-output examples. Upload 5-10 example pairs and 20+ evaluation samples, then run multi-round evaluation (Qwen-Max recommended) - [Translation Memory](https://help.aliyun.com/en/model-studio/memory-library-overview.md) - [Model Studio Memory Store (Cross-session Long-term Memory & User Profile)](https://help.aliyun.com/en/model-studio/memory-library.md): Cross-session memory for agents: auto-extracts key info from conversations and persists it. Supports memory snippets and user profiles. APIs like AddMemory/SearchMemory integrate with any app; multiple apps can share one store. Commercial billing starts 2026-08-20, with Pro/Lite tiers. - [OpenClaw Long-Term Memory Plugin (Cross-Session Capture & Recall)](https://help.aliyun.com/en/model-studio/modelstudio-memory-for-openclaw.md): Install modelstudio-memory-for-openclaw via npm, configure API Key for autoCapture/autoRecall cross-session memory; all Agents share one memory, register in openclaw.json - [Long-term memory guide](https://help.aliyun.com/en/model-studio/long-term-memory-2-0.md): Extract and store memory nodes and user profiles from conversations. Use AddMemory/SearchMemory/ListMemory APIs to write, retrieve, and manage cross-session context. Currently free. - [Data Connections & Scheduled Sync (Knowledge Base Data Sources)](https://help.aliyun.com/en/model-studio/data-connection-overview.md): Data connection overview: connect external data sources to a Bailian knowledge base, plus the scheduled data sync guide. Start here for knowledge base data sourcing. - [Bailian Data Connectors: MySQL/PostgreSQL/OSS/Yuque/Files](https://help.aliyun.com/en/model-studio/data-connection.md): Connect external data to Bailian: file, table, MySQL, PostgreSQL, PolarDB-X 2.0, Yuque and OSS connectors plus OSS/Feishu/DingTalk/SharePoint sync; searchFile-style tools for agents and workflows; MySQL cap of 10M rows. - [Knowledge Base (RAG)](https://help.aliyun.com/en/model-studio/knowledge-base.md) - [Create and use knowledge bases](https://help.aliyun.com/en/model-studio/rag-knowledge-base.md): Supplement LLMs with private data via RAG. Supports document search, data query, image Q&A, and audio/video search. Available only in China (Beijing). Billed hourly: Standard ¥0.03/h, Flagship ¥0.2/h. - [RAG performance optimization](https://help.aliyun.com/en/model-studio/rag-optimization.md): Fix incomplete retrieval or inaccurate answers. Strategies: optimize file layout, enable multi-turn rewriting, tag filtering, tune chunk size, evaluation baselines with 100+ Q&A pairs. - [Knowledge base API guide](https://help.aliyun.com/en/model-studio/rag-knowledge-base-api-guide.md): Programmatically create, upload to, and retrieve from knowledge bases via bailian SDK. Requires AliyunBailianDataFullAccess. Python/Java sample code for file upload and retrieval. - [Knowledge base logging and monitoring](https://help.aliyun.com/en/model-studio/rag-knowledge-base-log-monitoring.md): Deliver retrieval logs to SLS for auditing and alerting. Fields: request_id, pipeline_id, latency, response_code. Includes SQL examples for usage stats and error monitoring. - [Knowledge base quotas and limits](https://help.aliyun.com/en/model-studio/rag-knowledge-base-specifications.md): Quotas: unlimited KBs (100 for RDS), 100GB (Standard) / 9999GB (Ultimate) storage. Formats: PDF/DOCX/XLSX/images/audio/video. Retrieval: 1 QPS (Standard) to 10K QPS (Ultimate). - [Model Studio Knowledge Base Billing (Standard/Flagship/Free Quota/Resource Packages)](https://help.aliyun.com/en/model-studio/billing-for-knowledge-base.md): Standard 0.03 CNY/KB/hr (1 QPS), Flagship 0.2 CNY/RCU/hr (1 RCU≈50 QPS); 720-hr free quota (Standard only), RAG resource packages, model fees separate. Deleting a KB permanently erases data. - [Knowledge Base (RAG): Retrieval Service - Multi-KB Joint Search & Rerank](https://help.aliyun.com/en/model-studio/rag-knowledge-retrieval.md): Create a retrieval service binding up to 15 knowledge bases; query rewriting, hybrid vector+keyword recall, rerank via qwen3-rerank/qwen3-vl-rerank. - [Knowledge QA Service (Multi-KB RAG, Express/Agentic Retrieval)](https://help.aliyun.com/en/model-studio/rag-knowledge-qa.md): RAG QA service: bind up to 15 knowledge bases, express or Agentic multi-turn retrieval, TopK/rerank/threshold params, refusal, anti-leakage and citation controls. - [Model Studio Skill (Agent Capability Package): Official & Custom Skills](https://help.aliyun.com/en/model-studio/skill.md): Skill hub: what Skills are, official Skills (platform-preset, ready to use), creating custom Skills by uploading ZIP packs (SKILL.md spec, ≤10MB, ~2-min review), adding Skills to agent apps and testing. - [Bailian Skills: Agent Capability Packages (Official & Custom ZIP Skills)](https://help.aliyun.com/en/model-studio/introduction-to-skill.md): Capability packages for Bailian agent apps: official Skills plus custom Skills uploaded as ZIP with SKILL.md (≤10MB, ~2-min review), no coding needed. - [Model Studio MCP Hub (Official/Custom MCP Services & External Calls)](https://help.aliyun.com/en/model-studio/model-context-protocol.md): MCP docs entry: MCP introduction, official and third-party MCP services, creating custom MCP services, calling MCP externally, and FAQ. Connect tool capabilities to models and agents. - [Model Context Protocol (MCP) overview](https://help.aliyun.com/en/model-studio/mcp-introduction.md): Connect agent and workflow apps to third-party tools via MCP. Official and custom services supported. Cloud deployment free for limited time; custom billed per-second (Basic: CNY 0.000156/s). - [Official MCP services](https://help.aliyun.com/en/model-studio/official-and-third-party-mcp.md): Enable cloud-deployed MCP services (Amap Maps, Sequential Thinking, QuickChart) for agents and workflows. Up to 5 MCP services per agent. Workflow nodes use one tool each - [Custom MCP Services](https://help.aliyun.com/en/model-studio/custom-mcp.md): Deploy custom MCP servers via three methods: script-based (npx/uvx on Function Compute), AI Gateway import (wrap REST APIs), or OpenAPI import. Supports stdio and HTTP transport - [MCP external invocation](https://help.aliyun.com/en/model-studio/mcp-external-calls.md): Invoke Model Studio MCP services from Cherry Studio, Cursor, or custom projects via MCP SDK. Supports one-click config and manual JSON import. Uses Streamable HTTP protocol. - [MCP FAQ](https://help.aliyun.com/en/model-studio/mcp-faq.md): Troubleshoot MCP connection, deployment, and permission errors. Covers custom service deployment (npx/uvx/SSE), error codes (CONNECTION_REFUSED to INIT_TIMEOUT), and agent integration. - [Model Studio App Plugins (Official/Third-Party/Custom)](https://help.aliyun.com/en/model-studio/plug-in.md): Hub for Model Studio app plugins: plugin overview, official and third-party plugins, creating and configuring custom plugins to extend apps with external tool calls. - [Plug-in overview](https://help.aliyun.com/en/model-studio/plug-in-overview.md): Official plug-ins (code_interpreter, calculator, text_to_image, quark_search, generate_qrcode, github_search), third-party, and custom. Models: qwen-turbo/plus/max - [Official and third-party plug-ins](https://help.aliyun.com/en/model-studio/plugins.md): Setup guide for official plug-ins (Code Interpreter, Quark Search, Image Generation, Calculator, QR Code, GitHub Search) and third-party plug-ins. RAM service-linked role required - [Custom Plug-ins](https://help.aliyun.com/en/model-studio/custom-plug-ins.md): Create plug-ins from REST APIs or import from Marketplace. Define tool path, input/output params (LLM-recognized or business pass-through), debug, then publish as MCP service for agent apps - [Application Publishing and Sharing](https://help.aliyun.com/en/model-studio/application-publishing-and-sharing.md) - [Share and publish agent applications](https://help.aliyun.com/en/model-studio/share-an-application.md): Five publishing channels: web link (7-day expiry), DingTalk robot, WeChat Official Account, reusable component, and real-time audio/video via H5 QR code or SDK - [Publish agent or workflow as component](https://help.aliyun.com/en/model-studio/use-agent-or-workflow-as-component.md): Publish agents/workflows as reusable components with query/imageList params. Connect via Model Recognition or Business Pass-through. Weather query example included. - [UI Designer](https://help.aliyun.com/en/model-studio/ui-designer.md): Drag-and-drop builder for publishing Model Studio agents as web apps. 4 templates (Travel, Portal, Chat, Knowledge Base). DingTalk/WeCom login, OIDC/OAuth 2.0, database mapping, custom domains - [Bailian Application Invocation (Calling Published Apps via API)](https://help.aliyun.com/en/model-studio/bailian-application-calling.md): Call published Bailian apps with API Key + App ID: invoke agent, workflow, and RAG applications, streaming or non-streaming. Full API spec in API Reference (Application). - [Call Agent Application](https://help.aliyun.com/en/model-studio/call-single-agent-application.md): Integrate Bailian agent applications into business systems via DashScope SDK or HTTP requests. Requires API Key and APP_ID configuration, supports Python, Java, and curl with environment variable authentication - [Invoke a workflow application](https://help.aliyun.com/en/model-studio/invoke-workflow-application.md): This guide explains how to quickly and efficiently integrate an Alibaba Cloud Model Studio Workflow Application into your business system using the DashScope SDK or an HTTP API. - [Custom application parameter pass-through](https://help.aliyun.com/en/model-studio/pass-through-of-application-parameters.md): Pass params to plugins/workflow nodes via biz_params.user_defined_params in Agent/Workflow Application API calls. Python, Java, HTTP examples included - [Application evaluation](https://help.aliyun.com/en/model-studio/application-evaluation.md) - [Automated application evaluation](https://help.aliyun.com/en/model-studio/application-auto-evaluation.md): LLM-based automated evaluation for agent apps. Single-app or multi-app (up to 8) comparison. Auto-generates eval sets from knowledge bases, scores 1-5, and runs key driver analysis on bad cases. - [Manual evaluation](https://help.aliyun.com/en/model-studio/evaluate-manual-application.md): Create evaluation sets from .xls/.xlsx files, run batch evaluations on published agent applications, annotate responses with Poor/Fair/Good ratings, and generate evaluation reports - [Evaluation sets](https://help.aliyun.com/en/model-studio/application-evaluation-dataset.md): Manage evaluation datasets: conversation analysis (.xls/.xlsx with Prompt/Completion/SessionId) and knowledge-based Q&A (.jsonl with query/referenceAnswer/keywords). Auto-generation or manual upload. - [Application Evaluation (Evaluation Sets/Tasks/Label Management)](https://help.aliyun.com/en/model-studio/new-version-of-application-evaluation.md): Hub for Bailian app evaluation: manage new-version evaluation sets, run evaluation tasks, review results, and manage labels. Verify application answer quality before release. - [Application evaluation sets](https://help.aliyun.com/en/model-studio/new-version-of-evaluation-set.md): Create and manage evaluation datasets for agent, workflow, or custom tasks. Upload via Excel (.xls/.xlsx, max 20 MB) or import from Application Insights. Supports version management and schema editing - [Evaluation tasks](https://help.aliyun.com/en/model-studio/evaluation-task.md): Create evaluation tasks for agent and workflow apps. Combine automatic graders (LLM or Code) with human labels. Supports up to 10 evaluators per task. Token fees apply for LLM-based scoring - [Evaluation tag management](https://help.aliyun.com/en/model-studio/label-management.md): Create and manage tags for application evaluation and observability annotation. Four tag types: classification, boolean, number, and text. Supports quick annotate mode and filtering by tag values. - [Evaluation graders](https://help.aliyun.com/en/model-studio/grader.md): Preset templates (general quality, agent, text matching, format validation) and custom graders: LLM-based with configurable score range, or Python 3.10 Code graders - [Application Observability: Traces, Latency & Token Cost](https://help.aliyun.com/en/model-studio/application-monitoring.md): Monitor Bailian app runs: view call-chain traces, per-node latency and token consumption, and pinpoint errors — performance and cost troubleshooting for agent and RAG apps. - [App Performance Analytics](https://help.aliyun.com/en/model-studio/application-observation.md): End-to-end observability for agent, workflow, and high-code apps. Tracks CHAIN/LLM/RETRIEVER/RERANKER/TOOL nodes with latency and token count. Supports data annotation and JSONL/Excel export. - [Model Studio Application Gallery (Tingwu Agent/XiYan GBI/FaRui/DianJin)](https://help.aliyun.com/en/model-studio/application-gallery.md): Official app hub of Model Studio: Tingwu Agent, XiYan GBI, Tongyi FaRui, Tongyi DianJin, Qwen web-search Agent, deep search and more—find by scenario and try directly. - [Tongyi Photo Problem-Solving Coaching App (Segmentation/AnswerSSE API)](https://help.aliyun.com/en/model-studio/edu-tutor.md): Photo-solving app: per-problem segmentation of exam paper images, inference solving with key point analysis and explanation. With AnswerSSE streaming solving API. - [Photo-based Problem Solving and Tutoring](https://help.aliyun.com/en/model-studio/brief-introduction-of-edu-tutor.md): Upload test paper images for automatic segmentation and step-by-step solutions. Supports K-12, university, and adult exams. Free during public preview, 1000 calls/API/day, 5 QPS limit - [Model Studio EduTutor API Reference Entry](https://help.aliyun.com/en/model-studio/api-reference-edututor.md): Entry to Model Studio EduTutor (edututor 2025-07-07) API docs: API overview, endpoints, RAM authorization info and the API directory. Start here for EduTutor OpenAPI calls. - [Tingwu Agent Audio/Video Analysis (Service Insights/Buyer Profiling/Minutes)](https://help.aliyun.com/en/model-studio/official-application-tingwu-agent.md): Tingwu Agent hub: car-sales service insights, buyer profiling, general service insights, industrial speech transcription, smart minutes; for meetings/interviews/training. - [Tongyi Tingwu Agent overview](https://help.aliyun.com/en/model-studio/product-introduction-2.md): Speech AI + Qwen LLM for audio/video analysis: automotive sales insights, car buyer profiling, service insights, industrial transcription, and smart meeting notes - [RAM user authorization for Tingwu Agent](https://help.aliyun.com/en/model-studio/agent-ram-agent.md): Grant RAM users access to Model Studio and Tingwu: assign AliyunTingwuFullAccess policy in RAM console, then add user in Model Studio Permission Management with Administrator role - [Tingwu ASR discounted resource plans](https://help.aliyun.com/en/model-studio/purchase-and-use-discount-asr.md): Purchase discounted ASR plans (350 to 9M hours, CNY 0.028-0.60/hr) for Tingwu agents. Covers Automotive Sales, Car Buyer, Service Insights, Meeting Summary. Valid 3 months, non-refundable. - [Tingwu Automotive Sales Service Insights (Overview/Guide/API)](https://help.aliyun.com/en/model-studio/tingwu-automotive-service-insights.md): Entry for the Tingwu automotive solution on Model Studio: product overview, usage guide, business process and API reference for audio-based insight analysis of car sales and service scenarios. - [Automotive Service Insights overview](https://help.aliyun.com/en/model-studio/tingwu-automotive-service-insights-overview.md): AI agent analyzing automotive sales conversations from calls, store visits, and test drives. Scores service quality across 74 default items in 11 categories. CNY 0.73/hr audio. - [Automotive Service Insights user guide](https://help.aliyun.com/en/model-studio/tingwu-automotive-service-insights-guidelines.md): Configure the automotive service insights app: select content source (audio/text/Tingwu task), choose transcription model (automotive/paraformer), set analysis rules with ccai-pro, and publish. - [Automotive Service Insights business flow](https://help.aliyun.com/en/model-studio/tingwu-automotive-service-insights-process.md): Four-step integration flow: collect recordings from outbound calls/smart badges/in-vehicle systems, manage per-salesperson, upload via API for quality inspection, and render multi-dimensional results. - [Tingwu Automotive Service Insights API (CreateTask/GetTask)](https://help.aliyun.com/en/model-studio/tingwu-automotive-service-insights-api.md): Tingwu Automotive Service Insights API: CreateTask to create analysis tasks, GetTask to query task status and results for automotive sales service conversations. - [Automotive Service Insights CreateTask API](https://help.aliyun.com/en/model-studio/tingwu-automotive-service-insights-api-create-task.md): POST to DashScope with model=tingwu-automotive-service-insights. Accepts fileUrl, text, or dataId. Configure up to 150 insightsContents items with title, content, and score (-100 to 100). - [Automotive Service Insights GetTask API](https://help.aliyun.com/en/model-studio/tingwu-automotive-service-insights-api-get-task.md): Poll task status with task=getTask and dataId. Returns transcriptionPath (ASR paragraphs with timestamps), serviceInsightsPath (matched items with scores), and saleInsightsPath (hit/miss counts). - [Tingwu - Automotive Customer Profile (Overview, Guide, Workflow & API)](https://help.aliyun.com/en/model-studio/tingwu-automotive-customer-profile.md): Entry to Tingwu's automotive scenario: product overview, usage guide, business workflow and API reference for Automotive Customer Profile — voice-transcription-based profiling of car-buying customers. - [Automotive Customer Portrait overview](https://help.aliyun.com/en/model-studio/tingwu-automotive-customer-profile-overview.md): AI agent generating car buyer profiles from sales conversations. Outputs lead scores, test drive intent scores, one-sentence profiles, and detailed interest analysis. CNY 0.6/hr transcription. - [Automotive Customer Portrait user guide](https://help.aliyun.com/en/model-studio/tingwu-automotive-customer-profile-guidelines.md): Configure customer portrait app: select insight model (ccai-pro/qwen-plus/qwq), define up to 100 custom focus points across 15 default categories, assign up to 5 speaker roles. - [Automotive Customer Portrait business workflow](https://help.aliyun.com/en/model-studio/tingwu-automotive-customer-profile-process.md): Four-step integration: collect recordings from call systems/badges/in-car devices, match audio to buyer info, upload via API for portrait analysis, display lead scores. - [Tingwu Automotive Customer Profile APIs (CreateTask/GetTask)](https://help.aliyun.com/en/model-studio/tingwu-automotive-customer-profile-api.md): Tingwu automotive customer profile APIs: CreateTask starts a profile analysis task, GetTask queries status and results, extracting customer intent from car-buying conversations. - [Automotive Customer Portrait CreateTask API](https://help.aliyun.com/en/model-studio/tingwu-automotive-customer-profile-api-create-task.md): POST with model=tingwu-automotive-customer-profile. Accepts fileUrl, text, or dataId. Configure identityRecognition (up to 5 roles) and contentExtraction (up to 150 focus dimensions). - [Automotive Customer Portrait GetTask API](https://help.aliyun.com/en/model-studio/tingwu-automotive-customer-profile-api-get-task.md): Poll with task=getTask and dataId. Returns customerProfilePath (lead classify, test drive score, one-sentence summary), transcriptionPath, identityRecognitionPath, and contentExtractionsPath. - [Tingwu Service Insights (Customer Service & Sales Call QA)](https://help.aliyun.com/en/model-studio/tingwu-service-insights.md): QA for customer service and sales calls: spot common issues, drive data-based training improvements. Covers overview, user guide, business process, API reference. - [TingWu Service Insights overview](https://help.aliyun.com/en/model-studio/tingwu-service-insights-overview.md): ASR agent analyzing telemarketing, customer service, and sales recordings for quality inspection. Audio from phones/badges/in-vehicle systems. CNY 0.73/hr (audio), CNY 1.9/1K requests (text) - [TingWu Service Insights user guide](https://help.aliyun.com/en/model-studio/tingwu-service-insights-guidelines.md): Configure insight apps: choose source (audio/task ID/dialogue text), select ASR model (automotive/paraformer-v2), define up to 100 custom insight items with scoring, then publish for API use - [TingWu Service Insights business flow](https://help.aliyun.com/en/model-studio/tingwu-service-insights-process.md): Four-step pipeline: collect recordings from outbound calls/smart badges, manage and tag by rep, upload via API for quality analysis, render analytics results in your CRM - [Tingwu Service Insights API (CreateTask/GetTask)](https://help.aliyun.com/en/model-studio/tingwu-service-insights-api.md): Hub for Tingwu Service Insights APIs: CreateTask to create a general service-insight task, GetTask to query task status and results — for analyzing service calls and conversations. - [CreateTask — Service Insights API](https://help.aliyun.com/en/model-studio/tingwu-service-insights-api-create-task.md): POST with model=tingwu-service-insights. Input: fileUrl, text, or dataId (mutually exclusive) + appId. Define up to 150 insightsContents items with title, content, and score (-100 to 100) - [Tingwu-GetTask Query Service Insights Task Status and Results](https://help.aliyun.com/en/model-studio/tingwu-service-insights-api-get-task.md): Query Tingwu service-insights task status and results: input.dataId from createTask, task=getTask; returns transcription, insight hit/miss details and hit sentence IDs - [Tingwu Industrial Instruction Transcription: Overview, Guide & API](https://help.aliyun.com/en/model-studio/tingwu-industrial-instruction.md): Tingwu industrial instruction transcription hub: product overview, usage guide, business process and API reference. Turns production-line voice commands into structured text. - [Industrial Instruction Transcription overview](https://help.aliyun.com/en/model-studio/tingwu-industrial-instruction-overview.md): AI agent for transcribing voice commands in industrial quality inspection and device operation. Supports 10 languages. Uses LLM correction for specialized vocabulary. CNY 1.9/1000 requests. - [Industrial Instruction Transcription user guide](https://help.aliyun.com/en/model-studio/tingwu-industrial-instruction-guidelines.md): Configure industrial instruction app: upload custom instruction sets (up to 10 sets, 1000 items each), select language mode (auto-detect or single), test via microphone, publish for API use. - [Industrial Instruction Transcription business flow](https://help.aliyun.com/en/model-studio/tingwu-industrial-instruction-process.md): Three-step flow: upload instruction sets with professional terms, connect WebSocket API to industrial devices for real-time ASR, receive LLM-corrected instructions in business systems. - [Tingwu Industrial Instruction Transcription API (WebSocket Protocol)](https://help.aliyun.com/en/model-studio/tingwu-industrial-instruction-api.md): Tingwu industrial instruction transcription API: WebSocket real-time protocol for connection setup, audio streaming, and result delivery in industrial voice command scenarios. - [Industrial Instruction Transcription WebSocket API](https://help.aliyun.com/en/model-studio/tingwu-industrial-instruction-api-websocket.md): Full-duplex WebSocket at wss://dashscope.aliyuncs.com/api-ws/v1/inference. Stream 16kHz audio (pcm/wav/mp3/opus). Events: speech-listen, recognize-result, ai-result, speech-end. - [Tingwu Smart Meeting Minutes (Overview/Guide/API/JS SDK)](https://help.aliyun.com/en/model-studio/tingwu-meeting.md): Tingwu smart meeting minutes hub: product overview, usage guide, business process, API reference, open-source JS SDK. For meeting speech transcription and structured minutes. - [Intelligent Meeting Minutes overview](https://help.aliyun.com/en/model-studio/tingwu-meeting-summary-overview.md): AI agent for meeting/interview/class recordings. Real-time and offline transcription with 8-language translation. LLM features: chapter overview, summary, mind map, keywords, to-do, PPT extraction. - [Intelligent Meeting Minutes user guide](https://help.aliyun.com/en/model-studio/tingwu-meeting-summary-guidelines.md): Configure meeting minutes app: choose content source (file/real-time/task/text), select paraformer-v2 variant, enable diarization and translation, set analysis options (summary/mind map/PPT/prompt). - [Tingwu Smart Meeting Minutes Workflow (Real-time/Offline Transcription)](https://help.aliyun.com/en/model-studio/tingwu-meeting-summary-process.md): Tingwu Smart Meeting Minutes Agent flow: real-time or offline transcription, audio capture, Agent analysis, output of chapter overview, summary and speaker highlights - [Tingwu Meeting Transcription API (Offline/Real-time)](https://help.aliyun.com/en/model-studio/tingwu-meeting-api.md): Entry to Tingwu meeting transcription APIs: how to call offline-meeting and real-time-meeting interfaces to convert speech to text and generate meeting content for minutes and recordings. - [Tingwu Offline Meeting Transcription API (CreateTask/GetTask)](https://help.aliyun.com/en/model-studio/tingwu-meeting-offline-api-usage.md): Tingwu offline meeting APIs: CreateTask starts offline transcription from an audio/video URL; GetTask polls status and results. For meeting recording transcription. - [Offline Meeting Minutes CreateTask API](https://help.aliyun.com/en/model-studio/tingwu-meeting-offline-api-create-task.md): POST with model=tingwu-meeting, type=offline. Configure transcription (paraformer-v2, diarization, translation), analysis (summary, mind map, PPT, chapters, custom prompt with {Transcription}). - [Offline Meeting Minutes GetTask API](https://help.aliyun.com/en/model-studio/tingwu-meeting-offline-api-get-task.md): Poll with task=getTask. Returns OSS URLs: transcriptionPath, translationsPath, autoChaptersPath, summarizationPath, meetingAssistancePath, pptExtractionPath, textPolishPath, customPromptPath. - [Real-time meetings - API call procedure](https://help.aliyun.com/en/model-studio/tingwu-meeting-api-usage.md): This topic describes the API call procedure for real-time meeting transcription. - [Real-time Meeting Minutes CreateTask API](https://help.aliyun.com/en/model-studio/tingwu-meeting-api-create-task.md): POST with model=tingwu-meeting, type=realtime. Specify audio format (pcm/opus/aac/mp3) and sampleRate (8000/16000). Use dataId from response to end meeting and generate minutes. - [Real-time Meeting Minutes WebSocket API](https://help.aliyun.com/en/model-studio/tingwu-meeting-api-websocket.md): Full-duplex WebSocket at wss://dashscope.aliyuncs.com/api-ws/v1/inference with model=tingwu-meeting-realtime. Send run-task with dataId from CreateTask, then stream audio. - [Real-time Meeting Minutes GetTask API](https://help.aliyun.com/en/model-studio/tingwu-meeting-api-get-task.md): Same API as offline GetTask. Status values: 0=completed, 1=generating minutes, 2=failed, 3=transcribing. See offline GetTask for response parsing protocols. - [Intelligent summary JavaScript SDK demo](https://help.aliyun.com/en/model-studio/open-source-js-sdk.md): Open-source JS SDK for Tingwu intelligent summary frontend. Features: audio/video file transcription with speaker diarization, real-time microphone recording, and meeting minutes. Requires Node.js 18+ - [Tingwu Server SDK Installation (Java/Python with Real-time)](https://help.aliyun.com/en/model-studio/tingwu-sdk.md): Install Tingwu server-side SDKs for speech transcription: Java SDK, Python SDK, and Java/Python real-time SDKs. Use CreateTask/GetTask to run transcription tasks. - [TingWu Java SDK](https://help.aliyun.com/en/model-studio/tingwu-sdk-java.md): DashScope Java SDK for TingWu agents. TingWuParam config (model, input.fileUrl/text/dataId, appId, task), TingWu.call() HTTP method, DashScopeResult response. Requires SDK >= 2.21.1 - [TingWu Python SDK](https://help.aliyun.com/en/model-studio/tingwu-sdk-python.md): DashScope Python SDK for TingWu agents. TingWu.call(model, user_defined_input, parameters) with fileUrl/text/dataId inputs. Requires SDK >= 1.24.1, Python >= 3.9 - [TingWu Java SDK (Realtime)](https://help.aliyun.com/en/model-studio/tingwu-sdk-java-realtime.md): WebSocket real-time transcription SDK for Java. Call TingWuRealtime.call(), stream with sendAudioFrame(), then stop(). Params: appId, sampleRate (16000), format (pcm/wav/mp3/opus). SDK >= 2.21.7 - [TingWu Python SDK (Realtime)](https://help.aliyun.com/en/model-studio/tingwu-sdk-python-realtime.md): WebSocket real-time transcription SDK for Python. TingWuRealtime start/send_audio_data/stop flow. Callbacks: on_recognize_result, on_ai_result, on_speech_listen. SDK >= 1.24.4 - [Tingwu Agent Best Practices (Async Callback for Task Results)](https://help.aliyun.com/en/model-studio/tingwu-agent-best-practices.md): Tingwu Agent best practices: get async task results via callback instead of polling; EventBridge callback setup for sales insight and customer profiling. - [Tingwu Agent Async Task Callback Setup (HTTP POST/EventBridge)](https://help.aliyun.com/en/model-studio/tingwu-agent-async-get-task-result.md): Set up Tingwu Agent async callbacks: HTTP POST (Spring Boot receiver) or EventBridge. Only SDK/API-triggered tasks fire callbacks; retry failures via the status field. - [Tingwu FAQ (Agent Error Info Query)](https://help.aliyun.com/en/model-studio/tingwu-faq.md): FAQ hub for Tongyi Tingwu: Tingwu Agent error info query — find error causes and fixes by code or message; use when speech-to-text tasks or API calls fail. - [Tingwu Agent error codes](https://help.aliyun.com/en/model-studio/tingwu-agent-error-information-query.md): API error code reference: Arrearage, InvalidParameter (InputAppIdIllegal, AppNotExist, AppNotPublished), UnpurchasedAccess, UnsupportedOperation, and task failure codes - [Tongyi AI Podcast (Text-to-Podcast Audio/Voice Cloning/Bilingual)](https://help.aliyun.com/en/model-studio/official-application-aipodcast.md): Model Studio official app: Product Overview and API Reference. Text-to-podcast audio via PodcastTaskSubmit, voice cloning, topic expansion, topic specification, CN/EN output. For media and training. - [Qwen Podcast Generation](https://help.aliyun.com/en/model-studio/brief-introduction-of-ai-podcast.md): Convert documents (Word/PDF/TXT) into dual-host podcast audio. Supports voice cloning, topic extension, subject focus, and Chinese/English output. 100 free generations, then CNY 0.60/generation, 5 QPM - [Model Studio - AI Podcast API Reference (PodcastTaskSubmit/PodcastTaskResultQuery)](https://help.aliyun.com/en/model-studio/api-reference-aipodcast.md): AI podcast generation APIs: PodcastTaskSubmit to submit tasks, PodcastTaskResultQuery to check results. Covers API overview, endpoints (Beijing), RAM auth, API catalog. - [Multimodal Interaction Development Kit (Overview/Billing/SDK/API)](https://help.aliyun.com/en/model-studio/multimodal-products.md): Entry to the official Multimodal Interaction Development Kit on Model Studio (Bailian): product overview, billing, usage guide, SDK installation, API reference, best practices, FAQ. - [Multimodal Interaction Developer Suite overview](https://help.aliyun.com/en/model-studio/multimodal-products-overview.md): SDK/API suite for voice and vision AI on devices (glasses, robots, toys). Android/iOS/Linux/RTOS support, full-duplex conversation, voice cloning, on-device VAD/AEC, agent/plugin integration. - [Multimodal Interaction SDK Billing (Pay-as-you-go/Device Subscription/Savings Plan)](https://help.aliyun.com/en/model-studio/product-billing.md): How Multimodal Interaction SDK charges: pay-as-you-go by interaction round (ASR/TTS/intent recognition/LLM chat), prepaid annual device subscription (License), and Savings Plans with committed spend discounts. New users get a one-time CNY 10 free trial quota valid for 3 months. Lightweight ASR at CNY 0.05 per 1K calls, standard ASR at CNY 0.75 per 1K; knowledge-base Q&A at CNY 3.7 per 1K (excludes knowledge-base service fee). - [Real-time Multimodal Usage Guide (App Creation, Configuration, Timbre, A2A)](https://help.aliyun.com/en/model-studio/multimodal-guidelines.md): Real-time multimodal app setup: creation, configuration, publishing, instruction/timbre lists, multilingual dialogue, chat log integration and A2A Agent access. - [Create a multimodal interaction application](https://help.aliyun.com/en/model-studio/multimodal-app-creation.md): Create multimodal (integrated or Vision Edition) or voice (integrated or Lightweight) apps. Multimodal supports camera+mic; voice is mic-only. Includes per-app latency dashboard. - [Multimodal Interactive App Configuration (ASR/TTS/Prompt/Knowledge Base)](https://help.aliyun.com/en/model-studio/multimodal-app-configuration.md): Pick ASR (Fun-ASR, Paraformer, Qwen3-ASR) with hotword libraries, TTS (CosyVoice, Qwen3-TTS) with voice cloning, full-duplex interruption, semantic stop detection, text models, Prompt variables, knowledge base and long-term memory. - [Multimodal App Testing and Publishing (Voice/Video Chat, API Key)](https://help.aliyun.com/en/model-studio/multimodal-app-experience-and-publishing.md): Test and launch Bailian multimodal apps: run text, voice, or real-time video chat; configure API Key, system voices and voice cloning; view end-to-end first-packet latency; publish a version to get an app ID for production calls. - [Multimodal application templates](https://help.aliyun.com/en/model-studio/agent-template.md): Browse and copy templates from Model Studio marketplace for the Multimodal Interaction Developer Suite. Includes English tutor, writing coach, fitness coach, calorie analyzer, and Q&A by photo - [Multi-modal interaction instruction list](https://help.aliyun.com/en/model-studio/instruction-list.md): System and custom voice instructions for device control: brightness, volume, media, calls, camera, music, and app switching. Each instruction has name, params, and examples. - [Model Studio TTS Voice List (CosyVoice/Qwen-TTS/Sambert Models and voice Parameters)](https://help.aliyun.com/en/model-studio/multimodal-timbre-list.md): Overview of TTS voices on Model Studio: CosyVoice-v3-Flash/v3-plus/v2, Qwen-Audio-3.0-TTS-Plus/Flash, Qwen3-TTS-Flash-Realtime, Sambert, Multimodal Lite, and Qwen-TTS — with voice names, voice parameter values, supported languages (Chinese/English/dialects/Japanese/Korean), audio formats, sample rates, and voice cloning capabilities. - [Model Studio Multi-language Dialogue (Languages, ASR/TTS, Setup)](https://help.aliyun.com/en/model-studio/multi-language-dialogue.md): Supported single-language list for Model Studio voice apps: Chinese, English, Japanese, Korean, French, German, Spanish, Italian, Russian, Portuguese, Thai, Indonesian, Malay, Cantonese, etc. Available and recommended ASR/TTS models per language, plus multimodal app creation and voice interaction setup. Single-language only (Chinese-English mixing allowed). - [Ingest multimodal chat logs via EventBridge](https://help.aliyun.com/en/model-studio/multimodal-chatlog.md): Route chat logs using event source acs.multimodal, type multimodal:Dialogue:ChatLogPush. Targets: RocketMQ, SLS, Lark, WeCom, DB. Logs contain requestId, sessionId, input/output. - [Third-Party Agent Integration (Google A2A 0.2.5/AgentCard)](https://help.aliyun.com/en/model-studio/multimodal-integration-a2a.md): Integrate third-party agents via Google A2A 0.2.5: configure AgentCard (url/skills/streaming), call the agent, apiKey auth (X-API-KEY header), SSE streaming via /stream. - [A2A protocol extension for multimodal interaction](https://help.aliyun.com/en/model-studio/multimodal-integration-a2a-protocol.md): Extend A2A to access client metadata (user, device, location, images) and send commands. Declare in AgentCard capabilities.extensions. Response uses Command(name, params, commandRequestId). - [A2A intent extension for multimodal interaction](https://help.aliyun.com/en/model-studio/multimodal-integration-a2a-intent.md): Map user intents to agent skills via SkillExtension (id + inputSchema) in AgentCard capabilities.extensions. inputSchema follows MCP Tools protocol for parameter detection. - [Model Studio Multimodal SDK Installation (Server/Mobile/Linux/RTOS)](https://help.aliyun.com/en/model-studio/multimodal-sdk.md): Entry for installing Model Studio multimodal SDKs: server-side Java/Python/Go, mobile Android/iOS (incl. Lite), Linux C++, RTOS C—10 SDKs in total, incl. License mode. - [Model Studio Real-time Multimodal Server-side Java SDK (Push2Talk/Tap2Talk/Duplex)](https://help.aliyun.com/en/model-studio/multimodal-sdk-java.md): Server-side Java SDK for real-time multimodal interaction: install dashscope-sdk-java 2.21.16+, three modes Push2Talk/Tap2Talk/Duplex, WebSocket only, uplink pcm/opus, downlink pcm/mp3, MultiModalDialog API. - [Multimodal real-time interaction Python SDK](https://help.aliyun.com/en/model-studio/multimodal-sdk-python.md): Server-side Python SDK for real-time multimodal dialog via WebSocket. Key class: MultiModalDialog. Supports PCM/Opus audio uplink, PCM/MP3 downlink. Requires dashscope>=1.24.2 - [Server-side Go SDK](https://help.aliyun.com/en/model-studio/server-go-sdk.md): This document describes how to use the server-side Go SDK for the multimodal real-time interaction service from Alibaba Cloud Model Studio (Bailian). ... - [Model Studio Real-time Multimodal Android SDK (MultiModalDialog)](https://help.aliyun.com/en/model-studio/multimodal-sdk-android.md): Integrate the MultiModalDialog Android SDK: Push2Talk/Tap2Talk/Duplex modes, WebSocket link to DashScope, api_key/workspace_id/app_id params, voice chat and VQA photo recognition. - [Multimodal Android SDK (License Mode)](https://help.aliyun.com/en/model-studio/multimodal-license-android.md): Android License SDK for device auth. APIs: initialize, genRegisterReq, writeDeviceInfo, getToken. Semi-managed (your cloud + POP SDK) or fully managed (deviceRegister + getToken). - [Android SDK error codes](https://help.aliyun.com/en/model-studio/android-error-code-and-error-handling.md): Handle client-side (network/timeout) and server-side errors in the Android multimodal SDK via onErrorReceived. Header errors require restart; payload errors are recoverable. - [Model Studio Multimodal Real-time iOS SDK (MultiModalDialog)](https://help.aliyun.com/en/model-studio/multimodal-sdk-ios.md): MultiModalDialog iOS SDK: real-time audio/video multimodal interaction on Model Studio (Bailian); WebSocket/RTC links, Push2Talk/Tap2Talk/Duplex modes, integration code samples. - [Multimodal iOS SDK (License Mode)](https://help.aliyun.com/en/model-studio/multimodal-license-ios.md): iOS License SDK for device auth. APIs: initialize, genRegisterReq, writeDeviceInfo, getToken. Device ID persisted in keychain. Semi-managed or fully managed modes. - [Multimodal real-time interaction Android Lite SDK](https://help.aliyun.com/en/model-studio/multimodal-sdk-android-lite.md): Lightweight Android SDK for multimodal dialog over WebSocket (AudioOnly). Supports Push2Talk, Tap2Talk, and Duplex modes. No built-in AEC for full-duplex. Import dashscope-multimodal-dialog-lite AAR - [Multimodal real-time interaction iOS Lite SDK](https://help.aliyun.com/en/model-studio/multimodal-sdk-ios-lite.md): Lightweight iOS SDK for multimodal dialog over WebSocket (AudioOnly). Supports Push2Talk, Tap2Talk, and Duplex modes. Uses DashscopeLite.framework with SocketRocket. No built-in AEC - [Multimodal real-time interaction Linux C++ SDK](https://help.aliyun.com/en/model-studio/multimodal-sdk-linux.md): Native C++ SDK for Linux x86_64 and aarch64. Key APIs: CreateConversation, Connect, SendAudioData. Supports PCM/Opus/MP3 audio with internal Opus decoding. ConvEventType callbacks for dialog state - [RTOS C SDK (License Mode)](https://help.aliyun.com/en/model-studio/mmi-rtos-sdk.md): Embedded C SDK for multimodal interaction on RTOS. Supports 16+ chip platforms (ESP32, Allwinner, Rockchip, HiSilicon). Semi-managed or fully managed device auth modes. - [Real-time Multimodal Interaction API Reference (WebSocket/HTTP/Hot Words/Error Codes)](https://help.aliyun.com/en/model-studio/multimodal-api-references.md): API hub for real-time multimodal interaction: WebSocket/HTTP protocols, official Agents and third-party voice model calls, plugins, hot words, extra_info, error codes, long-term memory API. - [Real-time Multimodal WebSocket Protocol (tap2talk/duplex/push2talk)](https://help.aliyun.com/en/model-studio/multimodal-interaction-protocol.md): Low-latency WebSocket access: Bearer API Key auth, Paraformer/qwen3-asr ASR + CosyVoice/qwen3-tts TTS; wait for DialogStateChanged=Listening before sending audio. - [Multimodal interaction HTTP API](https://help.aliyun.com/en/model-studio/multimodal-http-protocol.md): HTTP-SSE API for text/image queries. Model: multimodal-dialog. Params: app_id, text, dialog_id (multi-turn), images (base64/URL), biz_params. No vision-only app support. - [Official agents for multimodal interaction](https://help.aliyun.com/en/model-studio/official-agent.md): Built-in agents: News Radio (dual AI streamers, tech/entertainment/social), Voice Translation (13 languages, simultaneous interpretation), Photo Q&A. Configure via SDK biz_params in the Start message - [Third-party voice model integration](https://help.aliyun.com/en/model-studio/third-party-voice-integration.md): Replace built-in ASR/TTS with third-party voice services. Use requestToRespond with type="prompt" to send recognized text, and configure text-only output for external TTS. Java/Python/C++ examples - [Call plug-ins in multimodal interactions](https://help.aliyun.com/en/model-studio/call-plugins.md): Integrate recommended, marketplace, or custom plug-ins into the Multimodal Interaction Kit. Pass plug-in params via biz_params.user_defined_params in Java/Python/Swift SDK - [Hotword OpenAPI management](https://help.aliyun.com/en/model-studio/management-hot-words.md): Manage ASR hotword vocabularies via OpenAPI. CRUD operations for custom phrase lists that improve recognition of domain-specific terms in Paraformer and Qwen-ASR models. - [Multimodal extra_info field reference](https://help.aliyun.com/en/model-studio/extra-info-description.md): Structure of extra_info in RespondingContent: commands (device control JSON), agent_info (intent domain/round), and tool_calls (plugin function results). Includes skill response configuration - [Multimodal interaction error codes](https://help.aliyun.com/en/model-studio/multimodal-error-code.md): Error codes with causes and fixes: AccessDenied.Unpurchased, Model.AccessDenied, ResponseTimeout (send HeartBeat), 421-InvalidParameter, 422-DirectiveNotSupported, 432-AppConfigError. - [Legacy long-term memory API](https://help.aliyun.com/en/model-studio/long-term-memory-api.md): API for the legacy multi-modal interaction long-term memory. Control memory mining policies and manage memory results via Java/Python SDK. Endpoint: sfmmultimodalapp.cn-beijing.aliyuncs.com. - [Multimodal Best Practices: Agent Integration, Voice Cloning & Custom Roles](https://help.aliyun.com/en/model-studio/multimodal-best-practices.md): Integrate Memo/Photo Q&A/Image Generation/Video Call/Tingwu Minutes agents, action/emotion control, custom roles, voice cloning, audio capture, RTOS SDK chat. - [Multimodal Memo Agent Integration (Voice Alarm/Schedule/Record)](https://help.aliyun.com/en/model-studio/multimodal-memo-agent.md): Voice memo management: create/query/modify/delete alarms, schedules, records via set_reminder/update_reminder/delete_reminder; recurring triggers and photo-based visual memo. - [Connect Bailian and Third-Party Agents (Direct Call/Agent App/Workflow)](https://help.aliyun.com/en/model-studio/bailian-and-tripartite-agent.md): Three ways to plug agents into the multimodal interaction suite for real-time calls: direct call of Bailian or third-party agents, Bailian agent applications, and Bailian workflow applications. - [Direct Call to Bailian and Third-Party Agents (A2A/Multimodal)](https://help.aliyun.com/en/model-studio/agent-direct-call.md): In multimodal apps, disable the text model and enable Direct Agent to forward requests to Bailian apps (Agent/workflow/orchestration) or third-party A2A agents. Covers agent_command for connection, UpdateInfo for dynamic parameters, and image passthrough (base64 ≤180KB). Not supported in voice apps. - [Connect Model Studio apps to multimodal interaction](https://help.aliyun.com/en/model-studio/multimodal-call-app.md): Import agent/workflow apps via Agent > Add > My Applications. Pass prompt variables through biz_params.user_defined_params in SDK. Update params at runtime via UpdateInfo. - [Integrate Model Studio Workflows into Multimodal Apps (Pass-through/Streaming/Timeout)](https://help.aliyun.com/en/model-studio/multimodal-call-workflow.md): Call Model Studio workflows in multimodal apps: direct connection, pass user_defined_params, set incremental output. 10s timeout limit; avoid JSON/Markdown output for TTS. - [Photo Q&A (VQA) Agent Integration via HTTP/WebSocket](https://help.aliyun.com/en/model-studio/vqa-agent.md): Integration entry for Model Studio's photo Q&A (VQA) Agent: connect via HTTP for photo question answering, or via WebSocket to combine photo Q&A with speech synthesis. - [Visual Q&A agent HTTP integration](https://help.aliyun.com/en/model-studio/vqa-agent-through-the-http-protocol.md): Send images to visual Q&A agent via HTTP, bypassing ASR/intent/TTS. Model: multimodal-dialog, agent app_id: visual_qa. Supports base64 and URL image input. - [Visual Q&A agent WebSocket integration with TTS](https://help.aliyun.com/en/model-studio/vqa-agent-via-websocket-protocol.md): Send images to visual Q&A agent via WebSocket with TTS audio output. Uses RequestToRespond directive after DialogStateChanged Listening state. Base64 images limited to 180KB. - [Connect to the Image Generation Agent (HTTP/WebSocket Audio)](https://help.aliyun.com/en/model-studio/generateimgagent.md): Image generation Agent integration: connect via HTTP, or send audio requests straight to the Agent over WebSocket. Entry point for choosing an integration protocol. - [Image Generation Agent HTTP integration](https://help.aliyun.com/en/model-studio/image-agent.md): HTTP API for text-to-image, image-to-image, and scribble-to-image via multimodal-dialog model. Passthrough mode with agent_command. Lite/Balanced/Advanced editions - [Voice passthrough to Image Generation Agent](https://help.aliyun.com/en/model-studio/audio-to-generateimgagent.md): Send voice to the Image Generation Agent via WebSocket for text-to-image. Configure agent_command with app_id=image_to_image and a custom query variable. Parse RespondingContent for image URLs. - [Music radio agent integration](https://help.aliyun.com/en/model-studio/music-agent.md): Integrate the music radio agent via SDK. Trigger with voice or requestToRespond. Parse music MP3 links from tool_calls (voice apps) or commands (multimodal). Instrumental-only library - [Integrate Video Call Agent (Multimodal Suite Android/iOS SDK, RTC)](https://help.aliyun.com/en/model-studio/live-api-integration.md): Enable Video Call Agent in Bailian's multimodal suite (voice apps excluded): Android/iOS SDK with built-in RTC, ChainMode=RTC, connect/exit via voicechat_video_channel. - [Third-party RTC video call integration](https://help.aliyun.com/en/model-studio/third-party-rtc-invoke-liveai.md): Enable LiveAI video calls by sending Base64 images (<180 KB) at 500ms intervals from your RTC server to the multimodal SDK. Java/Python/C++ WebSocket examples included - [Integrate Tingwu Meeting Minutes Agent (Recording/Offline/Real-time)](https://help.aliyun.com/en/model-studio/fast-integrate-tingwu-meeting-agent.md): Integrate Tingwu meeting-minutes Agent: Recording Summary Agent tutorial, offline and real-time transcription integration — turn meeting/call audio into minutes. - [Recording Summary Agent Tutorial (Tingwu Setup/Real-time Transcription/Wake Words)](https://help.aliyun.com/en/model-studio/recording-summary-agent-tutorial.md): Enable the Recording Summary Agent in the multimodal console: activate Tingwu pay-as-you-go, configure paraformer-v2 with speaker diarization, set start/pause/exit wake words (default Xiaoyun), and view command examples for real-time and file-based recording summaries. - [Integrate Tingwu Intelligent Summary Agent](https://help.aliyun.com/en/model-studio/fast-integrate-offline-tingwu-meeting-agent.md): SDK integration for offline meeting summarization. Commands: start_local_recording, pause, resume, end. Supports translation slots and speaker diarization via state machine - [Tingwu real-time transcription integration](https://help.aliyun.com/en/model-studio/realtime-tingwu-meeting-agent-integration.md): Integrate Tingwu real-time transcription via WebSocket/RTC. Voice commands control start/pause/resume/exit. Returns meeting_state_change events with dataId. Supports translation and multi-speaker. - [Action and Emotion Control Practice (Action Commands, Prompt Templates)](https://help.aliyun.com/en/model-studio/action-emotion-control-practice.md): Model Studio multimodal kit: one LLM call outputs robot-dog actions plus text replies. Prompt templates for JSON (action_query/response) or natural-language format. - [Custom Dialogue Roles in Multimodal Interaction (Prompt/Voice/Role Switching)](https://help.aliyun.com/en/model-studio/custom-role.md): Multimodal custom roles: prompt variables (${name}), voice config (e.g. longanhuan), voice/user_prompt_params in run-task. Role switch needs new WebSocket connection. - [Custom Instruction Practices for Device Control](https://help.aliyun.com/en/model-studio/custom-directive.md): Tool definition standards for in-vehicle/IoT multimodal interaction: JSON schema for actions, humanReadableTime/Date params, parameter merging, and device-side callback protocol - [Custom Commands for Alarms (CREATE_clock Setup & Samples)](https://help.aliyun.com/en/model-studio/custom-alarm-clock.md): Build alarms via custom commands: CREATE_clock with time_create/date_create/loop/offset params, cross-day boundary time handling and JSON samples; prefer the preset multimodal memo Agent for alarms and schedules. - [Audio capture and playback guide](https://help.aliyun.com/en/model-studio/audio-capture-and-playback-instructions.md): Audio format specs for multimodal interaction: input PCM (16-bit, 16kHz) or raw-opus; output PCM/opus/raw-opus/MP3 at 8-48kHz. FFmpeg conversion commands and RTOS config included. - [Bailian RTOS SDK (License Mode) Chat Integration and HAL Porting](https://help.aliyun.com/en/model-studio/chat-capability-based-on-rtos-sdk.md): Chat on RTOS devices via Bailian License-mode SDK: HAL porting of 5 modules (memory/random/storage/time/mutex), libqwen_sdk.a and libc_license.a linking, AppId/AppSecret/API Key init, aliyun_sdk_test() acceptance. - [RTOS SDK (License) vision module integration](https://help.aliyun.com/en/model-studio/rtos-sdk-license-mode-vision-module.md): Integrate camera-based Visual Q&A, Video Call, Visual Translation, and Express Video Call via c_visual API. Covers c_visual_config initialization, event callbacks, and camera capture implementation - [RTOS SDK (License) prompt variables](https://help.aliyun.com/en/model-studio/rtos-license-prompt-params.md): Set custom prompt variables via c_mmi_data_set_prompt_param to switch agent roles or rewrite prompts from device-side without changing App ID - [RTOS SDK device control via Skill-Command module](https://help.aliyun.com/en/model-studio/device-control.md): Load Skill-Command modules (volume, brightness, device, screen, multimedia) in RTOS/Linux SDK. Register callbacks via c_mmi_cmd_load(). ESP32 CMake and C examples - [RTOS SDK (License) volume control](https://help.aliyun.com/en/model-studio/rtos-sdk-license.md): Handle volume set/increase/decrease/mute/unmute events via c_mmi_cmd_volume_register callback. Value range 0-100, supports system/media/call volume types - [Multimodal Interaction Developer Suite FAQ](https://help.aliyun.com/en/model-studio/multimodal-products-faq.md): FAQ on TTS voices, on-device algorithms (VAD/AEC), upstream modes (push2talk/tap2talk/duplex), audio format, prompt variables, billing (pay-as-you-go vs license), and troubleshooting. - [XiyanGBI: NL2SQL and Natural Language Data Q&A](https://help.aliyun.com/en/model-studio/xiyan-gbi.md): XiyanGBI is a Tongyi-based native data assistant on Model Studio that turns natural language into NL2SQL, data Q&A and insights. Covers product intro, user guide and API reference. - [XiYan GBI update announcements](https://help.aliyun.com/en/model-studio/xiyan-update-announcement.md): Feature updates: permission management, SQL validation, price reductions (MIX CNY 499/mo, TURBO CNY 299/mo), Excel upload, VPC database access, and Q&A improvements. - [XiYan GBI product overview](https://help.aliyun.com/en/model-studio/brief-introduction-of-gbi-products.md): Natural language to SQL analytics. Three editions: Standard Turbo, Standard Mix (customizable), Custom (fine-tunable). Connects MySQL, PostgreSQL, Hologres, PolarDB. Term/synonym/business logic config - [XiYan GBI billing](https://help.aliyun.com/en/model-studio/xiyangbi-billing-description.md): Standard MIX (CNY 499/mo) and TURBO (CNY 299/mo) subscriptions, each with 5 RPM and 500 free calls/month. Extra RPM at CNY 1,000/5-RPM/month. Per-call overage: MIX CNY 0.80, TURBO CNY 0.50. - [GBI user guide](https://help.aliyun.com/en/model-studio/xiyan-gbi-user-guide.md): GBI offers a range of editions and specifications to fit your needs. This document provides guidance for common issues, helping you troubleshoot and optimize results for your selected edition. - [XiYan GBI reverse VPC access](https://help.aliyun.com/en/model-studio/reverse-network-access-to-vpc.md): Connect Xiyan GBI to a VPC database via PrivateLink reverse endpoint (epsrv-2zeczobsytmn6nv9actr) in Beijing. Cross-region needs CEN/VPC peering. Includes whitelist setup. - [Connect XiYan GBI to Hologres](https://help.aliyun.com/en/model-studio/connect-gbi-to-hologres.md): Connect via PostgreSQL protocol over public network using AccessKey credentials. Import TPC-H ORDERS table via MaxCompute foreign table, then query with natural language Q&A - [XiYan GBI - API Reference (Endpoints/RAM Auth/API Directory)](https://help.aliyun.com/en/model-studio/api-reference-4.md): Entry to XiYan GBI (dataanalysisgbi 2024-08-23) OpenAPI: API overview, service endpoints, RAM authorization, and API directory (data Q&A, NL2SQL analysis interfaces) - [XiYan GBI best practices](https://help.aliyun.com/en/model-studio/gbi-best-practices.md): Integrate Xiyan GBI API (RunDataAnalysis) for database Q&A. Java/Python SDK examples. Requires AliyunDataAnalysisGBIFullAccess policy and database connection - [XiYan GBI on-premises data best practices](https://help.aliyun.com/en/model-studio/xiyan-gbi-local-data-best-practices.md): Two-stage API for on-premises databases: RunSqlGeneration converts natural language to SQL, RunDataResultAnalysis visualizes results. Java and Python SDK examples included. - [XiYan GBI virtual data sources](https://help.aliyun.com/en/model-studio/analysis-of-gbi-virtual-data-source-best-practices.md): Set up unmanaged database connections using virtual data sources. APIs: createVirtualDatasourceInstance, saveVirtualDatasourceDdl, syncRemoteTables. MySQL/PostgreSQL. Java and Python examples. - [XiYan GBI sample database](https://help.aliyun.com/en/model-studio/official-database-sample-for-demonstration.md): Pre-built sample database for XiYan GBI with demo tables and queries. Connect to evaluate natural-language-to-SQL generation before integrating production data sources. - [XiYan GBI configuration and testing](https://help.aliyun.com/en/model-studio/xiyan-gbi-configuration-and-testing-recommendations.md): Configure table descriptions, column metadata, foreign keys, and enumerations. Add enterprise knowledge (terms, synonyms, business logic, optimization cases) to improve SQL accuracy. - [XiYan GBI permission management](https://help.aliyun.com/en/model-studio/xiyan-gbi-permission-management.md): Control data access at table, column, and value levels. Create permissions with fixed or variable filters, associate with roles (max 20), and isolate via Chat API. - [XiYan GBI DingTalk card integration](https://help.aliyun.com/en/model-studio/dingtalk-card-integrates-xiyan-gbi-question-and-answer-function.md): Build data Q&A bot in DingTalk using Xiyan GBI API. Create DingTalk app with Stream Mode robot, import card template, integrate Java backend with workspace-id and client credentials - [Tongyi Farui Legal LLM (Legal Q&A/Case Retrieval/Contract Review)](https://help.aliyun.com/en/model-studio/tongyi-farui.md): Legal LLM fine-tuned on Qwen: legal Q&A, court judgment case retrieval, statute search, legal text analysis, document generation, contract review with custom rules. - [Tongyi Farui billing](https://help.aliyun.com/en/model-studio/billing-for-tongyi-farui-1.md): Subscription: contract review free to CNY 80K/month (5 tiers). Pay-as-you-go: review CNY 4/page, consultation CNY 0.7/call, case retrieval CNY 0.8/call, regulation CNY 0.3/call - [Tongyi Farui FAQ](https://help.aliyun.com/en/model-studio/tongyi-farui-faq.md): Farui LLM (prompt-based, flexible legal Q&A) vs Farui Application API (RAG+Agent, better consultation and contract review). Includes support group info - [Quanmiao Official Apps (Miaobi/Miaoce/Proofreading/Miaosou/Miaodu)](https://help.aliyun.com/en/model-studio/quanmiao-solution-products.md): Entry to Quanmiao official apps: Miaobi, Miaoce and proofreading (AI writing/planning/review), Miaosou and Miaodu (search/reading), plus Quanmiao development docs for API integration. - [Miaobi, Miaoce & Proofreading (AI Writing, Billing, Usage Guide)](https://help.aliyun.com/en/model-studio/miaobi-miaoce-shenjiao.md): Entry to Miaobi writing, Miaoce data sources, and proofreading: billing (Miaobi/gov docs/PPT/video mixing), usage guide, writing best practices, and update notes. - [Miaobi & Miaoce Feature Update Announcements (Writing Genres/Pipeline Upgrades)](https://help.aliyun.com/en/model-studio/update-announcement.md): Entry to dated feature update announcements for Miaobi and Miaoce on Model Studio: expanded writing genres (official documents, media, marketing), writing pipeline model upgrades, AI toolbox continue-writing iterations. - [Miaobi & QuanMiao Feature Update Log (V2.1 to Feb 2025)](https://help.aliyun.com/en/model-studio/miaobi-and-miaoce-function-update.md): Dated release notes for Miaobi and QuanMiao solutions: Feb 2025 Miaobi update, Jan 2025 QuanMiao update, AI QuanMiao V2.2.2/V2.2.1/V2.2, AI Miaobi V2.1 feature changes. - [Miaobi February 2025 update](https://help.aliyun.com/en/model-studio/february-26-2025-update-miaobi.md): Adds Deepseek-r1 model support, three writing methods (direct/step-by-step/advanced), nine new official document types, template-based input, and long-form writing over 5000 words - [Quanmiao update January 24, 2025](https://help.aliyun.com/en/model-studio/2025-1-24-function-update-announcement-quanmiao-saas.md): Miaobi updates: optimized writing workflow reducing LLM inference calls, improved article style/format learning, and reduced inference quota for title generation - [AI Magic Series V2.2.2 update (March 2024)](https://help.aliyun.com/en/model-studio/update-2024-03-11-ai-quanmiao-v2-2.md): Streaming generation in Copilot mode now skips AI thinking/search steps. ASR display logic optimized for related and timeline videos. - [AI Magic Series V2.2.1 update (March 2024)](https://help.aliyun.com/en/model-studio/update-2024-03-01-ai-quanmiao-v2-2.md): AI Search adds ASR preprocessing for all library videos and ranking optimization. AMB adds Marketing and Office Documents writing scenarios. - [AI All-in-One Suite V2.2 update (Feb 2024)](https://help.aliyun.com/en/model-studio/update-2024-02-28-ai-quanmiao-v2.md): Hotspot Management upgraded to AI Strategy with topic analysis. AI Search upgraded to multimodal with copilot-based multi-agent architecture. AMB adds OSS custom storage. - [AMB V2.1 update (Dec 2023)](https://help.aliyun.com/en/model-studio/update-2023-12-19-amb-v2.md): AiMiaoBi V2.1: upgraded intervention module, Material Library image insertion, text-to-image in writing, 6 new styles (promo/PR/newspaper), Article Review with fact/duplication checks - [Miaobi Billing: Editions, Quotas and Add-ons](https://help.aliyun.com/en/model-studio/miaobi-billing.md): Pricing for Miaobi (incl. Miaoce), Deep Writing and Bid Doc Generation: Free 0 CNY/21 days, Team Light 59 CNY/mo, Team Standard 99 CNY/mo, Enterprise Pro 22,333 CNY/mo, Enterprise Custom; each edition includes inference quota. Add-ons: Deep Writing 20 CNY/use, Bid Doc Generation 30 CNY/use - [Government Document Toolkit billing](https://help.aliyun.com/en/model-studio/government-document-tool-billing.md): Pay-as-you-go pricing: Intelligent Proofreading at 0.05 CNY/1000 chars, Document Search at 0.50 CNY/request, standard templating at 0.50 CNY/application, custom templates at 8640 CNY/set - [PPT generation billing](https://help.aliyun.com/en/model-studio/ppt-generation-billing.md): Subscription editions: Team Lite (15 CNY/mo, 10 uses), Standard (100/mo, 100 uses), Enterprise Pro (1,000/mo). Add-on: 1 CNY/use. Template import: 6,000 CNY each - [Miaoce Custom Data Source billing](https://help.aliyun.com/en/model-studio/billing-document-miaoce-custom-data-source.md): Token-based pay-as-you-go: Quanmiao-Plus at CNY 0.002/1K tokens, Quanmiao-Max at CNY 0.06/1K tokens. APIs for custom-source hot topic analysis, query, and export - [Video Remixing billing](https://help.aliyun.com/en/model-studio/billing-description-video-mixing.md): Pay-as-you-go at CNY 0.5/min based on uploaded video duration. Default 2 concurrent tasks; Concurrent Task Expansion available at CNY 0.1/min for additional parallelism - [Model Studio Writing Apps Guide (AI Miaobi/Miaoce/Proofreading)](https://help.aliyun.com/en/model-studio/usage-guide.md): Entry to Model Studio writing apps: AI Miaobi (article drafting), AI Miaoce (topic planning), Intelligent Proofreading (error checking), Deep Writing (long-form creation) — usage guides for media and content workflows. - [AI Miaobi (Intelligent Writing/Article Generation/Style Imitation)](https://help.aliyun.com/en/model-studio/amb.md): Model Studio AI Miaobi writing assistant: AI Toolbox for article generation, material library, style imitation and format learning, system configuration. - [AMB (Miaobi) product overview](https://help.aliyun.com/en/model-studio/product-overview-for-amb.md): LLM text generation for media, government, and marketing: news drafts, official documents, marketing copy, video scripts, and poetry. Includes multimodal capabilities - [AMB homepage overview](https://help.aliyun.com/en/model-studio/amb-homepage-overview.md): Writing workbench with 5 modes: media, official document, marketing, office, and custom. Creative tools include style/format learning, ad generation, one-click video creation, and material search - [Miaobi Writing Interface (Step-by-step/Direct Generation/Smart Images)](https://help.aliyun.com/en/model-studio/function-interface.md): Miaobi writing UI entries: step-by-step writing, direct generation, smart image generation, material search — AI article drafting, auto illustration, asset lookup. - [AMB step-by-step article creation](https://help.aliyun.com/en/model-studio/step-by-step-generation.md): Four-stage article generation: Subject, Outline, Summary, Article. Edit/regenerate at each step. Includes AI Toolbox for title generation, rewriting, translation, and keyword extraction - [AMB direct generation](https://help.aliyun.com/en/model-studio/direct-generation.md): Generate articles in one step without outline via AMB workbench. Set subject, title, style, length, output language. Supports result intervention, rewrite/expand/shorten, and inline editing - [AMB smart image generation](https://help.aliyun.com/en/model-studio/smart-image-generation.md): Generate AI images for articles by entering content keywords and selecting a style (e.g. 3D cartoon). Save results to Document Management - [AMB material search](https://help.aliyun.com/en/model-studio/search-materials.md): Search online resources via Quark integration, preview results, and add articles to the Material Library for use in AMB content creation - [AI Toolbox](https://help.aliyun.com/en/model-studio/ai-toolbox.md): AMB text processing toolkit: title generation, summary, continue writing, keyword extraction, comment prediction, format layout, Chinese-English translation, and one-click polish - [AMB material library](https://help.aliyun.com/en/model-studio/material-library.md): Manage Public Material Library (shared within skill group) and My Material Library (personal uploads) in AMB. Search, reference, and share writing materials. - [AMB system configuration](https://help.aliyun.com/en/model-studio/system-configuration.md): Configure intervention rules (keyword, intent, and URL filtering) and writing sources (Quark Search, CCTV, Xinhuanet) with adjustable weights for AMB content generation - [AMB style and format learning](https://help.aliyun.com/en/model-studio/style-imitation.md): Upload documents (up to 50,000 chars) to train a custom writing style model. Apply learned styles to generate drafts matching your preferred tone and format - [AI Strategy (MiaoCe)](https://help.aliyun.com/en/model-studio/ai-miaoce.md): Topic analysis from web, popularity, timeliness, and novelty perspectives. Generates topic plans with subject, summary, and outline exportable as mind map or Word document - [Quanmiao Smart Review (Doc Proofreading/Custom Dictionary/Rules/Fact Check)](https://help.aliyun.com/en/model-studio/article-review.md): Bailian Quanmiao smart review: upload Word/PDF/plain text; check content accuracy, formatting, political sensitivity, compliance, legal issues, images. Configure custom dictionary, rule library and up to 10 fact-check sources; accept, ignore or add findings to dictionary. - [Deep Writing](https://help.aliyun.com/en/model-studio/deep-writing.md): Multi-agent research report generation: Project Manager, Data Collector, Report Writer, and Data Analyst sub-agents. Supports web search, local upload, OSS data, custom methodologies, and traceability - [Text Writing Guidance (Media Style/Government Docs/FAQ)](https://help.aliyun.com/en/model-studio/document-writing-best-practices.md): Hub for writing guidance on Model Studio: media-style writing best practices, government document writing guidance, and FAQs for Quanmiao-series products, for news releases and official documents. - [Media Writing Guide: One-Shot Prompt, Outlining, Title & Summary](https://help.aliyun.com/en/model-studio/media-style-writing-best-practices.md): Hub for media writing tutorials: draft an article with a one-step prompt, outline structure when out of ideas, generate titles and summaries from existing text. - [Miaobi Media Writing: One-Prompt News, Official & Marketing Drafts](https://help.aliyun.com/en/model-studio/quick-media-writing-prompt.md): Miaobi direct writing mode: enter a topic to generate news, official documents, marketing or office drafts in one step. Supports ~600-word length, auto-supplemented references, result intervention, rewrite/expand/condense, plus AI toolkit for style rewriting, full-text translation, polishing and review. - [AMB writing assistance with material search](https://help.aliyun.com/en/model-studio/use-amb-to-help-writing.md): Three AMB writing aids: material search via Quark, step-by-step outline-to-article generation with regeneration, and hot topics for content inspiration. - [Generate titles and summaries (media)](https://help.aliyun.com/en/model-studio/generate-titles-summaries-media-text.md): Use the Miaobi AI Toolbox on media manuscripts to generate titles, summaries, continuations, expansions, or shortened versions for full articles or selected sections - [Best practices for writing official government documents](https://help.aliyun.com/en/model-studio/government-document-writing-best-practices.md) - [Quick government document writing](https://help.aliyun.com/en/model-studio/quick-gov-writing-prompt.md): Generate formal government documents with Miaobi via a single prompt. Supports style, length, and language config. Features: Result Intervention, AI Toolbox, and Material Library. - [Step-by-step government document writing](https://help.aliyun.com/en/model-studio/step-by-step-gov-writing.md): Create government documents using the step-by-step workflow: enter subject and keywords, generate outline, review excerpts, and enrich content with material search - [Generate titles and summaries for government documents](https://help.aliyun.com/en/model-studio/generate-titles-summaries-gov-text.md): Use the Miaobi AI Toolbox on government documents to generate titles, summaries, continuations, expansions, or abbreviations for full articles or selected sections - [Miaobi writing guide FAQ](https://help.aliyun.com/en/model-studio/faq-for-using-quanmiao-series-products.md): Prompt-writing tips for Miaobi: determine article type/subject/stance, when to use step-by-step vs direct generation, and how to add reference materials for better control - [MiaoSou & MiaoDu: Multimodal Search and Document Reading (Billing/Guide)](https://help.aliyun.com/en/model-studio/miaosou-and-miaodu.md): MiaoSou (multimodal search: QA search, vector retrieval) and MiaoDu (doc reading: summary, mind map, Q&A): billing (post-paid Token, offline index fees) and usage guide. - [MiaoSearch and MiaoRead billing](https://help.aliyun.com/en/model-studio/miaosou-miaodu-api-billing.md): MiaoSearch online billing: quanmiao-max at CNY 0.06/1K tokens, quanmiao-plus at CNY 0.002/1K tokens. Offline: index storage, file processing, vector building. MiaoRead billed per-token by model. - [MiaoDou & MiaoDu User Guide (MiaoSou Q&A Search / Text-to-Text)](https://help.aliyun.com/en/model-studio/miaodou-and-miaodu-guidelines-for-use.md): Entry to MiaoDou and MiaoDu: MiaoSou Q&A search (LLM filters sources by query and summarizes answers), deep search, text-to-text search; accessed via the QuanMiao app in Model Studio app plaza. - [AI Miaosou: Multi-Agent Multimodal Search](https://help.aliyun.com/en/model-studio/ai-miaosou.md): Multi-agent multimodal search: Q&A search (general/deep/research/auto) and pure search; local/internet sources; doc/image/video/audio retrieval; token-based billing. - [QuanMiao Development Docs (API Reference/Best Practices)](https://help.aliyun.com/en/model-studio/ai-quan-miao-development-document.md): Dev docs hub for QuanMiao (MiaoBi/MiaoCe): API reference, best practices, and more. For developers integrating MiaoBi writing and MiaoCe reading apps. - [Miaobi API Reference (API Overview/Endpoints/RAM Authorization)](https://help.aliyun.com/en/model-studio/amb-api-reference.md): Miaobi (AI writing) OpenAPI hub: API overview, endpoints, RAM authorization, API catalog, data structures, version notes, for integrating writing APIs like RunWritingV2. - [AiMiaoBi API overview](https://help.aliyun.com/en/model-studio/api-aimiaobi-2023-08-01-overview.md): RPC-standard API catalog for AiMiaoBi (2023-08-01). Covers material library, custom text, video editing, hot topics, VOC mining, smart search, and deep writing. Multi-language SDKs. - [AiMiaoBi RAM authorization](https://help.aliyun.com/en/model-studio/api-aimiaobi-2023-08-01-ram.md): RAM policies for AiMiaoBi API access. Grant aimiaobi:* actions via custom policy or AliyunAiMiaoBiFullAccess. Required for content generation and copywriting APIs. - [Miao Bi & Miao Ce Best Practices (Proofreading/Miao Sou/Miao Du/Video/PPT)](https://help.aliyun.com/en/model-studio/miaobi-and-miaoce-best-practices.md): Best practices for Miao Bi AI writing: text generation & polishing, smart proofreading, Miao Ce ideation, Miao Sou search, Miao Du reading, video mixing, PPT generation. - [AiMiaoBi Writing API Best Practices (RunWritingV2/Web Traceability)](https://help.aliyun.com/en/model-studio/best-practices-for-miaobi-api.md): Best practices for the AiMiaoBi writing API: get a Workspace ID and AccessKey, call RunWritingV2 via Java/Python SDK to generate articles directly, with web traceability and streaming output examples. - [Automated review best practices](https://help.aliyun.com/en/model-studio/best-practices-for-smart-audit.md): Java SDK examples for AiMiaoBi Automated Review: SubmitSmartAudit, GetSmartAuditResult, ExportAuditContentResult APIs. Covers article, image, and factuality review plus ruleset-based content audit - [Miaoce API best practices](https://help.aliyun.com/en/model-studio/best-practices-for-miaoce-api.md): Java SDK examples for Miaoce: news topic clustering (SubmitDocClusterTask), hot topic mining, content scoring, and sentiment analysis via AiMiaoBi SDK - [Miaobi MiaoSou API Best Practices: Data Sources, Source Weights, Intelligent Search](https://help.aliyun.com/en/model-studio/best-practices-for-miaosou-api.md): Miaosou API best practices: data source setup (web search/API/file), source weight config, intelligent search; ListSearchTasks sample; needs AgentKey and WorkSpaceId. - [AMB third-party search API template](https://help.aliyun.com/en/model-studio/third-party-search-api-template.md): HTTP POST JSON template for integrating custom search APIs with AMB. Request params: query, current, size, includeContent. Response returns Article objects with source, title, content, url, pubTime - [MiaoDu writing pipeline best practices](https://help.aliyun.com/en/model-studio/miaodu-best-practices.md): End-to-end MiaoDu API workflow: upload docs via UploadDoc, process with the AiMiaoBi SDK (alibabacloud-aimiaobi20230801), and generate content. Includes Java code examples. - [AiMiaoBi video mixing best practices](https://help.aliyun.com/en/model-studio/best-practices-for-video-mixing-and-cutting.md): Java SDK workflow: asyncUploadVideo, asyncCreateClipsTimeLine, asyncCreateClipsTask APIs. Supports single/multi-source remix and reference-video style transfer. Max 200 MB per video, 20 min total - [PPT generation best practices](https://help.aliyun.com/en/model-studio/ppt-generation-best-practices.md): End-to-end PPT workflow: outline generation via RunPptOutlineGeneration API, then rendering. Java (aimiaobi SDK), Python, JavaScript examples with streaming SSE - [Quanmiao Light Apps: Extra Setup (SLR/Source Integration/iframe/OSS)](https://help.aliyun.com/en/model-studio/quanmiao-more.md): Quanmiao PaaS extras: service-linked role, Miaobi source integration, Miaosou API data source, iframe embedding, Logo customization, OSS setup, AgentKey retrieval. For PaaS integrators. - [Quanmiao service-linked roles](https://help.aliyun.com/en/model-studio/quanmiao-slr.md): Two SLRs: AliyunServiceRoleForAIMiaoBiAccessingOss (Material Library/OSS) and AliyunServiceRoleForAiMiaoBiAccessingIMS (multimodal/IMS). Includes permission details and deletion impact. - [Miaobi data source integration](https://help.aliyun.com/en/model-studio/miaobi-writing-source-docking.md): Connect Miaobi Writing to external data sources via CreateDataset/UpdateDataset APIs or local file upload. Configure third-party search endpoints with JSON request/response mapping. - [MiaoSearch data source API integration](https://help.aliyun.com/en/model-studio/miaosou-introduce-data-source-through-api.md): Configure a third-party search API as a MiaoSearch custom data source. Define request/response mapping in JSON using searchSourceRequestConfig and jqNodes for field extraction. - [Quanmiao iframe embedding](https://help.aliyun.com/en/model-studio/iframe-embedding-scheme.md): Embed Quanmiao SaaS into OA/CRM systems via iframe. Create RAM user with AliyunAiMiaoBiFullAccess, generate STS token, build logon-free URL with TicketType=mini - [Quanmiao SaaS logo customization](https://help.aliyun.com/en/model-studio/logo-customization-specification-and-deployment-method.md): Customize logos on the login page and menu bar for Quanmiao SaaS products. Container specs: 190x32 px (single product) or 140x32 px (multi-product). Deploy via iframe embedding. - [Quanmiao OSS authorization setup](https://help.aliyun.com/en/model-studio/oss-setup-guide.md): Configure OSS storage for Quanmiao apps. Steps: create bucket, add /aimiaobi directory, set CORS, authorize AliyunServiceRoleForAIMiaoBiAccessingOss RAM role. Region: cn-beijing - [Obtain Quanmiao AgentKey](https://help.aliyun.com/en/model-studio/quanmiao-paas-agentkey-get-guide.md): Get the AgentKey identity field for legacy Quanmiao API calls. Navigate to Miaobi/Miaosou SaaS page and hover over the account to view the AgentKey. - [Quanmiao Light Apps Series (Billing/Usage/Dev Docs/FAQ)](https://help.aliyun.com/en/model-studio/quanmiao-light-application-series.md): Quanmiao light apps hub: update announcements, billing, usage guide, developer docs, FAQ. For users activating and using Quanmiao light apps on Model Studio. - [QuanMiao Light App Release Notes (Video Understanding/VOC/Video Creation)](https://help.aliyun.com/en/model-studio/light-application-update-announcement.md): QuanMiao light app updates: video understanding (Qwen-VL-Post/OCR/12 languages), VOC mining, video segmentation, one-click video, essay correction, hot news MCP Server. - [Model Studio Video Understanding Sales & Pricing Announcements](https://help.aliyun.com/en/model-studio/light-application-sales-strategy-update.md): Video understanding sales & pricing updates: price-cut announcement (2025-02-18), new promotional video (2025-03-11), commercialization notice (2024-12-20). Review before purchase. - [Visual understanding March 2025 update](https://help.aliyun.com/en/model-studio/march-11-2025-update-video-analysis-understanding-new-announcement.md): Promotional announcement for the visual understanding application. Covers 8 use cases (video description, tag categorization, Q&A, retrieval) with 24 built-in templates for film and media. - [Video understanding price reduction (Feb 2025)](https://help.aliyun.com/en/model-studio/update-february-18-2025-video-understanding-call-price-reduction-announcement.md): Qwen-Max inference price reduced effective Feb 18, 2025. New rates: input CNY 0.0024/1K tokens, output CNY 0.0096/1K tokens. Applies to video understanding calls. - [Video understanding commercialization (Dec 2024)](https://help.aliyun.com/en/model-studio/update-2024-12-20-ai-quanmiao-video-understanding.md): Quanmiao video understanding commercialized mid-January 2025. Input video for LLM-based summaries, tags, and analysis with custom prompts in visual and text phases. - [Model Studio Quanmiao Light Apps Feature Update Announcements](https://help.aliyun.com/en/model-studio/light-application-function-update.md): Feature update announcements for Quanmiao-series light apps on Alibaba Cloud Model Studio, tracking version iterations and new feature releases over time. - [Model Studio Update Announcement: E-commerce Retail Promotion Copywriting](https://help.aliyun.com/en/model-studio/e-commerce-retail-promotion-copy-writing-update-announcement.md): Bailian release announcement: update to the e-commerce/retail promotion copywriting feature—new capabilities for generating promo and product marketing copy. - [Media/Retail Article Style & Format Learning Update Announcement (Jan 24, 2025)](https://help.aliyun.com/en/model-studio/media-retail-article-style-and-format-learning-update-announcement.md): Update to Miaobi media/retail article style and format learning: Jan 24, 2025 release improved the interaction experience and upgraded the style learning and writing pipeline. - [Article style learning update Jan 24, 2025](https://help.aliyun.com/en/model-studio/updated-january-24-2025.md): Upgraded two-step workflow: upload sample articles for style learning, then generate new articles. Improved first-package response time and reduced LLM token consumption. - [Model Studio Update: Film & TV Script Creation Announcement](https://help.aliyun.com/en/model-studio/announcement-on-update-of-film-and-television-mutual-entertainment-script-creation.md): Bailian (Model Studio) product update notice on film and television script-creation model capabilities, covering feature changes and release schedule. - [Film & TV Media Video Understanding Update Announcement](https://help.aliyun.com/en/model-studio/film-and-television-media-video-understanding-update-announcement.md): Announcement of updates to film/TV media video understanding in Model Studio: feature changes, capability upgrades, and effective dates for video understanding models. - [Video understanding update Feb 21, 2025](https://help.aliyun.com/en/model-studio/updated-february-21-2025-video-analysis.md): Adds Qwen2.5-7B-1M model for text processing in Step 3. Token billing optimized to reduce costs. Supplementary text limit increased from 400 to 600 characters. - [Video understanding update Feb 18, 2025](https://help.aliyun.com/en/model-studio/updated-february-18-2025-video-analysis.md): Adds support for FLV, MOV, TS, AVI, MKV, WMV, MPG video upload formats in addition to MP4. - [Model Studio Car/Content Platform News Hot List Interaction Update Announcement](https://help.aliyun.com/en/model-studio/car-machine-content-platform-news-hot-list-interactive-update-announcement.md): Model Studio announcement: update notice for the News Hot List Interaction feature on car machine and content platforms; review the changes before using this feature. - [Pan-Enterprise VOC Mining Update Announcement (Model Studio)](https://help.aliyun.com/en/model-studio/pan-enterprise-voc-mining-update-announcement.md): Model Studio product announcement: update notes for Pan-Enterprise VOC (Voice of Customer) Mining, extracting insights from user feedback for enterprise scenarios. - [Pan-Enterprise Lead Mining Update Announcement](https://help.aliyun.com/en/model-studio/pan-enterprise-lead-mining-update-announcement.md): Update announcement for the Pan-Enterprise Lead Mining app, which mines user profiles and upsell leads from social media and call records; token billing on Qwen-Plus/Max/Qwen-Long. - [Enterprise Lead Mining update (Jan 2025)](https://help.aliyun.com/en/model-studio/pan-enterprise-lead-mining-update-announcement-updated-january-24-2025.md): Jan 2025 updates: consumer behavior and partnership mining examples, batch content tagging, Excel-compatible output, and new batch processing API - [Model Studio Content Security Audit Update Announcements](https://help.aliyun.com/en/model-studio/network-content-security-audit-update-announcement.md): Content security audit update announcements for Model Studio (Bailian): Jan 24, 2025 update on compliance policy changes for model input/output content review. - [Quanmiao Light App - Network Content Audit Update (Proofreading/Batch API)](https://help.aliyun.com/en/model-studio/2025-1-24-function-update-announcement-quanmiao-light-application.md): Jan 24, 2025 update: added Article Proofreading scenario example (typo detection, grammar review, punctuation check) via Tongyi Qianwen-Max in Effect Debugging; added batch processing API for async multi-file/text audit tasks, Java SDK requires alibababcloud-quanmiaolightapp20240801 dependency. - [Quanmiao Light Apps Billing (E-commerce Copy/Scripts/Moderation)](https://help.aliyun.com/en/model-studio/light-application-billing-document.md): Billing for Quanmiao light apps: pricing of e-commerce retail copywriting, film & entertainment script creation, content security moderation, smart video clipping, composition correction and more. - [Quanmiao e-commerce promotional copywriting billing](https://help.aliyun.com/en/model-studio/e-commerce-retail-promotion-copywriting-billing.md): Pay-as-you-go billing by token usage. Supported models: Qwen-Max (CNY 0.0024/0.0096 per 1K tokens in/out), Qwen-Plus (0.0008/0.002), DeepSeek-R1 (0.002/0.008). Free trial on Performance Debugging page - [Quanmiao controllable e-commerce copy billing](https://help.aliyun.com/en/model-studio/e-commerce-copy-intelligent-controllable-generation-billing.md): Pay-as-you-go by token. Models: Qwen-Plus/Max, DeepSeek-R1, Quanmiao Controllable LLM Standard (0.0008/0.002) and Premium (0.003/0.009). Failed tasks not charged - [Screenplay creation billing](https://help.aliyun.com/en/model-studio/film-and-television-mutual-entertainment-script-creation-billing.md): Pay-as-you-go billing for the screenplay creation AI Agent. Uses Quanmiao-Max model at 0.06 CNY/1000 tokens for both input and output. Failed tasks (task-failed event) are not charged - [Media Video Understanding billing](https://help.aliyun.com/en/model-studio/film-and-television-media-video-understanding-billing.md): Real-time (30 min max, 1 free concurrency) and async (1 hr, 2 free, scalable to 30) modes. Processing at 2.8 CNY/hr plus token fees. Savings plans with 5-50% discount - [In-Vehicle Infotainment Hotspot Q&A Billing](https://help.aliyun.com/en/model-studio/car-machine-network-hot-information-interactive-question-and-answer-billing.md): Pay-as-you-go: broadcast list CNY 1/call, personalized recommendation CNY 0.009/call, news Q&A CNY 0.018/question. APIs: HotNewsRecommend, GetHotTopicBroadcast, RunHotTopicChat - [Cross-Enterprise VOC Mining billing](https://help.aliyun.com/en/model-studio/pan-enterprise-voc-mining-billing.md): Pay-as-you-go per output token. Models: Qwen-Plus (0.0008/0.002 CNY per 1K tokens), Qwen-Max (0.0024/0.0096), Qwen-Long (0.0005/0.002) - [Enterprise Lead Mining billing](https://help.aliyun.com/en/model-studio/pan-enterprise-lead-mining-billing.md): Pay-as-you-go per output token. Models: Qwen-Plus, Qwen-Max, Qwen-Long with per-1K-token pricing. Failed tasks (task-failed event) are not charged - [Network Content Moderation billing](https://help.aliyun.com/en/model-studio/network-content-security-audit-billing.md): Pay-as-you-go pricing based on LLM token usage. Models: Qwen-Plus (0.0008/0.002 per 1K tokens), Qwen-Max (0.0024/0.0096), Qwen-Long (0.0005/0.002). Free trial on Performance Debugging page - [Composition Correction Billing](https://help.aliyun.com/en/model-studio/composition-correction-billing.md): Token-based pay-as-you-go: Correction Model CNY 0.015/1k tokens, Lightweight Model CNY 0.0015/1k tokens, OCR Model CNY 0.01/1k tokens, OCR Lightweight CNY 0.0006/1k tokens - [Intelligent video splitting billing](https://help.aliyun.com/en/model-studio/video-smart-strip-billing.md): Real-time (1 free concurrency, <=30min video) and async (2 free, <=1hr) modes. Multimodal processing CNY 2.8/hr plus LLM token fees. Optional shot enhancement via VL/OCR models. - [Model Studio Light Apps Guide: 9 Scenario Solutions (Copywriting/Scripts/VOC/Moderation)](https://help.aliyun.com/en/model-studio/light-application-guidelines-for-use.md): 9 light apps by scenario: e-commerce copywriting, film/TV scripts, video understanding, VOC/lead mining, content moderation, composition correction. - [E-commerce copywriting generation app](https://help.aliyun.com/en/model-studio/intelligent-and-controllable-generation-of-e-commerce-copywriting.md): Generate product titles, summaries, and marketing copy with word-count control. Supports Xiaohongshu notes, WeChat Moments promos, ad slogans, and travel articles via Qwen-Max/Qwen-Plus models. - [Quanmiao: Media/Retail Article Style & Format Learning (No-code)](https://help.aliyun.com/en/model-studio/media-retail-article-style-and-format-learning.md): No-code style learning: upload up to 10 sample articles (4,000 chars total), Qwen-Max analyzes the style and generates matching articles. Limited-time free, then pay-per-token. - [Film & TV Script Writing (QuanMiao App/OpenAI-compatible/Free Trial)](https://help.aliyun.com/en/model-studio/film-and-television-script-creation.md): QuanMiao script writing app: character/scene/script prompts, organize & export, OpenAI-compatible API, agent chain for context. Limited-time free, then per-token billing. - [Media Video Understanding (Video Summaries/Tagging/Content Analysis)](https://help.aliyun.com/en/model-studio/media-video-understanding.md): Quanmiao light app: chains ASR, VL models and LLM to extract video highlights as summaries, tags and copy. 10 scenarios. Realtime ≤30 min / async ≤60 min; billed per token. - [Quanmiao In-Car Hot News Interaction & Q&A (Smart Cockpit News Broadcast)](https://help.aliyun.com/en/model-studio/car-machine-content-platform-news-hot-list-interaction.md): Quanmiao in-car news broadcast app: 8-channel hot topics clustered every 4 hours, custom playlist & summary style, news Q&A; limited-time free, then per-token billing. - [General Enterprise VOC Mining](https://help.aliyun.com/en/model-studio/pan-enterprise-voc-mining.md): LLM-powered tagging of unstructured VOC data (reviews, forums, chats). Predefined/custom tags, batch file upload, JSON output. Two APIs: real-time streaming (SSE) and batch - [Enterprise Leads Mining](https://help.aliyun.com/en/model-studio/pan-enterprise-clue-mining.md): Extract user profiles and cross-selling leads from social media and call transcripts. Custom lead tags with batch input, JSON output. Real-time streaming and batch APIs - [Network Content Moderation user guide](https://help.aliyun.com/en/model-studio/network-content-security-audit.md): Scan text for policy violations with custom moderation dimensions. Supports real-time streaming API (SSE) and batch processing API. Configure via Qwen-Max or Qwen-Plus models - [Essay Correction Assistant](https://help.aliyun.com/en/model-studio/composition-correction-assistant.md): Grade student essays with auto-scoring, grammar/spelling checks, and revision suggestions. Supports text input or OCR from photos. APIs: single real-time streaming and batch processing - [Model Studio Development Documentation (Best Practices/API Reference)](https://help.aliyun.com/en/model-studio/development-documentation.md): Development docs hub: Best Practices (samples and solutions for lightweight apps), API Reference (APIs for invoking models and applications). For integrating Model Studio APIs. - [Best practices](https://help.aliyun.com/en/model-studio/light-application-best-practices.md) - [Video understanding and one-click video creation](https://help.aliyun.com/en/model-studio/best-practices-for-applying-video-understanding-and-one-click-film.md): Combine Quanmiao video understanding with one-click video creation via CloudFlow and Function Compute. RunVideoAnalysis API generates scripts from long videos for automated short video production - [VOC mining and data analytics best practices](https://help.aliyun.com/en/model-studio/best-practices-for-mining-voc-information-and-data-analysis.md): Combine Quanmiao Enterprise VOC Mining (SubmitEnterpriseVocAnalysisTask API) with Xiyan GBI ChatBI for structured tagging and real-time analysis of customer feedback. Java SDK with DataAnalysisGBI - [Integrate video understanding in Model Studio workflows](https://help.aliyun.com/en/model-studio/best-practices-for-workflow-integration-video-understanding.md): Create Function Compute functions (SubmitVideoAnalysisTask, GetVideoAnalysisTask) and wire them into a Model Studio workflow. Supports serial/parallel component composition. Max 3-min video - [Quanmiao Light App OpenAPI Reference (API Catalog, Endpoints, RAM Auth)](https://help.aliyun.com/en/model-studio/api-reference-1.md): Hub for the quanmiaolightapp 2024-08-01 OpenAPI: API overview, service endpoints, RAM authorization, full API catalog, data structures, and version change history. - [QuanMiaoLightApp API overview](https://help.aliyun.com/en/model-studio/api-quanmiaolightapp-2024-08-01-overview.md): ROA-standard API catalog for QuanMiaoLightApp (2024-08-01). Covers e-commerce copywriting, script creation, video understanding, video segmentation, VOC mining, essay grading, and content moderation. - [Quanmiao Light App FAQ](https://help.aliyun.com/en/model-studio/quanmiao-lightapp-faq.md): Fix 403 authorization errors calling Quanmiao Light App APIs. Grant AliyunQuanMiaoLightAppFullAccess policy, add workspace permissions, or verify workspaceId. - [Lingque CCAI Dialogue Analysis AIO Docs (Billing/Guide/API/Integration)](https://help.aliyun.com/en/model-studio/official-application-lingque-ccai-dialogue-analysis-aio.md): Entry hub for Model Studio official app Lingque CCAI - Dialogue Analysis AIO: update announcements, product overview, billing, user guide, technical integration scheme, API reference, best practices, API call examples. For customer service dialogue QA and analysis. - [Tongyi Xiaomi CCAI update announcements](https://help.aliyun.com/en/model-studio/tongyi-xiaomi-ccai-update-announcement.md): This document describes product updates. - [Lingque CCAI Dialogue Analysis AIO Launch (Summarization/Quality Inspection/Analysis)](https://help.aliyun.com/en/model-studio/july-4-2024-official-application-of-lingque-ccai-dialogue-analysis-aio-online-announcement.md): Release (Jul 2024): Lingque CCAI Dialogue Analysis AIO — generative summarization, quality inspection and analysis; custom templates, dashscope Application.call API. - [Lingque CCAI Dialogue Analysis AIO Model Price Reduction (Turbo -90%, Plus -70%)](https://help.aliyun.com/en/model-studio/chai-dialogue-analysis-aio-model-calls-price-reduction-notification.md): Per-call price cuts: Lingque Dialogue Model-Turbo from CNY 0.01 to 0.001 (-90%), Plus from 0.033 to 0.01 (-70%). Full pricing in Billing. - [Lingque CCAI Dialogue Analysis AIO Billing Change Notice (Paid Speech Recognition, New Image Recognition)](https://help.aliyun.com/en/model-studio/lingque-ccai-dialogue-analysis-aio-billing-item-change-notice.md): Aug 14, 2025: Lingque CCAI Dialogue Analysis AIO adds image recognition (VLMax, CNY 0.01/call) and moves offline speech recognition from free trial to CNY 0.33/hour. - [CCAI Conversation Analysis AIO overview](https://help.aliyun.com/en/model-studio/product-overview-1.md): Generative summarization, quality inspection, and multi-instruction analysis API for call/ticket data. Models: Tongyi Xiaomi-Turbo and Plus. Pay-as-you-go per call - [Billing for Lingque CCAI Dialogue Analysis AIO (Per-Call/Per-Hour Pricing)](https://help.aliyun.com/en/model-studio/billing-description-magpie-ccai-dialogue-analysis-aio.md): Free activation. Lingque-Plus ¥0.01/call, Turbo ¥0.001/call (1 call=2000 tokens), image recognition ¥0.01/call, offline ASR ¥0.33/hr, 1000 free calls for new users. - [User guide](https://help.aliyun.com/en/model-studio/lingque-ccai-aio-user-guide.md) - [Lingque CCAI Dialog Analysis AIO Activation (Free Activation & RAM Authorization)](https://help.aliyun.com/en/model-studio/product-activation.md): Activate Lingque CCAI Dialog Analysis AIO: App Plaza > Free Activation, choose a billing plan. RAM sub-accounts need root-account authorization first or activation fails. - [CCAI application management](https://help.aliyun.com/en/model-studio/application-management.md): Copy, modify, and delete CCAI conversation analysis applications. Includes API integration for structured conversation analysis results and call volume monitoring (billed per 2000 tokens). - [Model Studio Dialogue Analysis Agent App (Summary/QA/Tagging Instructions)](https://help.aliyun.com/en/model-studio/create-an-application-based-on-dialogue-analysis-agent.md): Dialogue Analysis Agent apps: standard (summary/keywords/Q&A/solutions) & advanced (service QA/tagging) instructions, ≤15,000-char text, Lingque-Plus/Turbo. - [CCAI (Tongyi Xiaomi): Create an App via Custom Prompts](https://help.aliyun.com/en/model-studio/create-an-application-based-on-a-custom-method.md): Build a CCAI (Tongyi Xiaomi) app via custom prompts: 5 templates (summary, extraction, QA check, tagging, multi-instruction); analyze text, audio (≤40MB WAV/MP3) or images (≤10MB). - [Hotword group management](https://help.aliyun.com/en/model-studio/hot-phrase-management.md): Configure hotword groups (max 128 words each, weight 1-5) to improve speech-to-text accuracy in offline and real-time quality inspection. Chinese characters only. One group per application - [CCAI-AIO knowledge base configuration](https://help.aliyun.com/en/model-studio/using-the-knowledge-base.md): Configure RAG knowledge bases (document, table, image) for CCAI-AIO. Set invocation mode (always/intelligent), similarity threshold (0.01-1), and recall weight (0.5-2). - [CCAI technical integration overview](https://help.aliyun.com/en/model-studio/technology-integration-scheme.md): Integrate enterprise service platforms with Smart Conversation Analysis via OpenAPI (version 2019-01-15). Upload conversation data for quality inspection and receive results in console - [Tongyi Xiaomi (CCAI) ContactCenterAI OpenAPI Reference](https://help.aliyun.com/en/model-studio/api-reference-2.md): OpenAPI hub for Tongyi Xiaomi Contact Center AI: API overview, service endpoints, API catalog (BailianChatBot/SseChat/BridgeWebCall voice and service bot APIs), change history. - [Model Studio Best Practices: Service QA, Field Extraction, Summarization](https://help.aliyun.com/en/model-studio/best-practices.md): Hands-on Model Studio examples: customer service quality inspection, field information extraction, summarization (summary/title/keywords) — reference for prompt and app design. - [Customer Service Quality Inspection Best Practices](https://help.aliyun.com/en/model-studio/customer-service-quality-inspection-best-practices.md): End-to-end workflow: activate Model Studio, create CCAI-AIO app, call AnalyzeConversation API with custom inspection items (overpromising, sentiment). Java SDK example included - [CCAI-AIO field information extraction](https://help.aliyun.com/en/model-studio/best-practices-for-automatic-work-order-generation.md): Extract structured fields (region, age, date, service type) from conversations using CCAI-AIO AnalyzeConversation API. Java SDK contactcenterai20240603, sync and streaming modes - [CCAI-AIO summary generation best practices](https://help.aliyun.com/en/model-studio/summary-best-practices.md): Generate conversation summaries, titles, and keywords via CCAI-AIO AnalyzeConversation API. Includes Java SDK setup (contactcenterai20240603) and sync/streaming call examples - [伶鹊CCAI Dialogue Analysis AIO API Call Examples (Template ID/Prompt/Task Type)](https://help.aliyun.com/en/model-studio/call-tyxm-ccai-aio-api.md): 伶鹊CCAI Dialogue Analysis AIO API call examples: invoke by template ID, native Prompt, or task type; offline data upload, image analysis, ROA signature, RAM user authorization. - [CCAI-AIO template ID SDK call examples](https://help.aliyun.com/en/model-studio/use-template-id-to-call-tongyi-xiaomi-ccai-aio.md): Java/Python SDK examples for calling CCAI-AIO using templateIds parameter. Locate template IDs via Professional build mode > Instruction template management. - [CCAI-AIO native prompt SDK call examples](https://help.aliyun.com/en/model-studio/use-native-prompt-to-call-tongyi-xiaomi-ccai-aio.md): Java/Python SDK for calling CCAI-AIO with native prompts. Uses contactcenterai20240603 artifact, workspaceId and appId params. Supports async streaming. - [CCAI-AIO conversation analysis SDK examples](https://help.aliyun.com/en/model-studio/call-tongyi-xiaomi-ccai-dialogue-analysis-aio-application-through-task-type.md): Java SDK (contactcenterai20240603) for CCAI-AIO conversation analysis. Sync and async streaming call examples with AnalyzeConversation API. Configure workspaceId, appId, and resultTypes per task - [CCAI offline task analysis code examples](https://help.aliyun.com/en/model-studio/tongyi-xiaomi-ccai-dialogue-analysis-by-uploading-offline-task-data.md): Java SDK examples for CCAI-AIO: create audio/text tasks, query results. Uses contactcenterai20240603 SDK with AccessKey auth. Requires workspaceId and appId - [ROA-style request signature mechanism](https://help.aliyun.com/en/model-studio/roa-style-request-body-signature-mechanism.md): Construct and sign HTTP requests with SDK V3. ACS3-HMAC-SHA256 algorithm. Build CanonicalRequest from method, URI, query, headers, SHA256 payload hash. Required: x-acs-action, x-acs-version, host. - [CCAI Conversation Analysis image analysis example](https://help.aliyun.com/en/model-studio/picture-analysis-through-tongyi-xiaomi-ccai-dialogue-analysis-aio-application.md): Analyze images via AnalyzeImage API with contactcenterai20240603 SDK. Sync/async Java code examples with watermark detection and AccessKey auth - [Bailian CCAI Dialogue Analysis: RAM Sub-account Authorization](https://help.aliyun.com/en/model-studio/use-and-authorize-ram-users-for-ccai-dialogue-analysis.md): Authorize RAM users to manage Bailian CCAI Dialogue Analysis: create the RAM user, attach the AliyunSFMFullAccess system policy (account-level, distinct from same-named custom policies) in the RAM console, then add the user and grant admin permissions in the Bailian console. - [Lingque CCAI Voice Dialogue Robot (Overview/Operation/Billing)](https://help.aliyun.com/en/model-studio/official-application-lingque-ccai-voice-dialogue-robot.md): Lingque CCAI voice robot on Model Studio: custom voice, LLM conversation, speech playback/transcription, real-time interaction. Billed ¥0.1/min call time. - [CCAI Voice Chatbot overview](https://help.aliyun.com/en/model-studio/product-0verview.md): Build custom voice chatbots powered by LLMs. Prompt-based persona creation, configurable timbre/volume/speed/pitch, silence timeout, and multi-endpoint integration - [Lingque CCAI Voice Dialogue Robot Billing (¥0.1/Minute Call Duration)](https://help.aliyun.com/en/model-studio/billing-information-lingque-ccai-voice-dialogue-robot.md): Billing rules for Lingque CCAI voice dialogue robot: pay-as-you-go by actual call duration at ¥0.1/minute; activation is free, no charge if the service is not used. - [Tongyi Xiaomi Voice Bot Setup (Model/Voice/Instruction Templates)](https://help.aliyun.com/en/model-studio/operation-guide.md): Configure Tongyi Xiaomi CCAI voice bot in Model Studio: Qwen-Plus/Max/Turbo, 4 instruction templates, voice/silence timeout (1-60s), test call, publish, API integration - [API Reference](https://help.aliyun.com/en/model-studio/api-reference-chat6.md) - [Model Studio Official App: Lingque CCAI Customer Service Dialogue Agent](https://help.aliyun.com/en/model-studio/official-application-voicepica-ccai-beebot-agent.md): Official app on Model Studio: Lingque CCAI customer service dialogue agent for building customer service chatbots in enterprise online service scenarios. - [Lingque CCAI Customer Service Agent Overview (Build CS Chatbots)](https://help.aliyun.com/en/model-studio/product-overview-voicepica-beebot-agent.md): Lingque CCAI Customer Service Agent: build customer service chatbots on a Qwen-based customer-service LLM. Visual workflow skill orchestration (reply/branch/API plugin/parameter-collection nodes), FAQ knowledge configuration, activation and RAM authorization. - [Customer Service Dialog Agent Billing (1,000 Free Calls)](https://help.aliyun.com/en/model-studio/billing-description-beebot-agent.md): Customer Service Dialog Agent charges per call: 1,000 free calls per Alibaba Cloud UID, then pay-as-you-go after quota runs out. One user question counts as one call. Lingque dialog model: CNY 0.015/call; no fee for activation without usage. - [Tongyi Xiaomi CCAI Conversational Agent guide](https://help.aliyun.com/en/model-studio/guidelines-for-use.md): Build agents with persona prompts, knowledge base (flow skills + FAQ), and visual flow orchestration. Supports security interception and response language config - [BailianChatBot API Reference (API Overview/RAM Auth/API Catalog)](https://help.aliyun.com/en/model-studio/api-reference-5.md): BailianChatBot (2024-11-05) API reference: API overview, RAM authorization, and API catalog for Tongyi Xiaomi customer-service chat/voice bot APIs such as SseChat. - [Tongyi Dianjin Hub: Overview, Announcements, API Reference](https://help.aliyun.com/en/model-studio/tongyi-dianjin.md): Tongyi Dianjin hub: overview, announcements, API reference. Postpaid by consumed tokens: Standard ¥0.01/1K tokens, Advanced ¥0.1/1K tokens, billed hourly; activation free. - [Tongyi Dianjin overview](https://help.aliyun.com/en/model-studio/tongyi-dianjin-overview.md): Finance-specific LLM app. Standard (60 QPM, CNY 0.01/1K tokens) and Premium (300 QPM, CNY 0.1/1K tokens, Pro parsing) editions. Pay-as-you-go on input+output tokens - [Tongyi Dianjin API Reference (Endpoints/RAM/API Catalog)](https://help.aliyun.com/en/model-studio/api-reference-3.md): OpenAPI hub for Tongyi Dianjin (dianjin-2024-06-28): API overview, service endpoints, RAM authorization, API catalog and version change notes. Check endpoints and auth before coding calls. - [API overview](https://help.aliyun.com/en/model-studio/api-dianjin-2024-06-28-overview.md): API standards and multilingual preset SDKsThe OpenAPI of this product (DianJin/2024-06-28) uses the ROA signature style. We have encapsulated SDKs for... - [DianJin RAM authorization](https://help.aliyun.com/en/model-studio/api-dianjin-2024-06-28-ram.md): RAM policies for DianJin data analysis API access. Grant dianjin:* actions via custom policy or AliyunDianJinFullAccess. Required for knowledge retrieval and Q&A APIs. - [Tongyi Doc Mining: Extraction, Review, Tagging & Summarization](https://help.aliyun.com/en/model-studio/tongyi-docmining.md): Bailian official doc app: information extraction, content review, tagging, summarization via Tongyi LLMs, structured JSON output; product intro and API reference. - [Tongyi Data Mining overview](https://help.aliyun.com/en/model-studio/docmining-product-introduction.md): LLM-based document processing: extraction, content moderation, classification, and summarization. Supports PDF/DOC/XLS/images up to 100 MB. Token-based billing. Storage: 10,000 files - [Model Studio DocMining API Reference (Overview/Endpoints/Directory/Error Codes)](https://help.aliyun.com/en/model-studio/docmining-api-reference.md): DocMining API reference hub in Model Studio: API overview, service access endpoints, API directory, and error codes for Tongyi Data Mining document parsing. - [Tongyi Data Mining API overview](https://help.aliyun.com/en/model-studio/docmining-api-overview.md): DashScope HTTP API for Tongyi Data Mining. Six endpoints: document upload, information extraction, content moderation, tag categorization, summary generation, and document deletion. Requires API key - [Tongyi Data Mining service endpoints](https://help.aliyun.com/en/model-studio/docmining-service-access-point.md): Regional API endpoints for Tongyi Data Mining. China (Beijing) cn-beijing public endpoint: dashscope.aliyuncs.com/api/v2/apps/, no VPC endpoint - [Model Studio DocMind Document APIs (Upload/Extraction/Audit/Summary)](https://help.aliyun.com/en/model-studio/docmining-api-directory.md): DocMind APIs in Model Studio: document upload, information extraction, content audit, tagging/classification, summary generation and document deletion. - [Data Mining document upload API](https://help.aliyun.com/en/model-studio/document-upload.md): Three-step upload flow: (1) POST /zhiwen-file/apply_upload_lease with fileName, sizeBytes, md5 to get pre-signed URL and lease_id, (2) PUT file to OSS, (3) submit for parsing. Python and Java examples - [Data Mining information extraction API](https://help.aliyun.com/en/model-studio/document-information-extraction.md): POST /zhiwen-chat/extraction for document NER and field extraction. capabilityType values: BID_EXTRACTION, RESUME_EXTRACTION, ENTITY_EXTRACTION. Pass fileIdList and userPrompt. Supports streaming - [Data Mining content moderation API](https://help.aliyun.com/en/model-studio/document-content-audit.md): POST /zhiwen-chat/audit for document content review. Set capabilityType=CONTENT_REVIEW, pass fileIdList and userPrompt. Supports streaming responses. Python and Java examples included - [Data Mining tag categorization API](https://help.aliyun.com/en/model-studio/document-tagging.md): POST /zhiwen-chat/tagging to classify and label documents. Set capabilityType=TAG_CLASSIFY, pass fileIdList and userPrompt. Supports streaming responses. Python and Java examples included - [Data Mining summary generation API](https://help.aliyun.com/en/model-studio/document-summary-generation.md): POST /zhiwen-chat/summary to generate document summaries. Set capabilityType=SUMMARY_GEN, pass fileIdList and userPrompt. Supports streaming. Returns output with token usage stats - [Data Mining delete document API](https://help.aliyun.com/en/model-studio/document-delete.md): POST /zhiwen-file/delete_file to remove uploaded documents. Pass fileId (format: file_zhiwen_XXX). Returns success boolean and request_id. Python and Java examples included - [Tongyi Data Mining error codes](https://help.aliyun.com/en/model-studio/docmining-error-code.md): Error codes for Data Mining APIs. InternalError (500), InvalidParameter (400) for prompt/type/file issues, InvalidParameter.File for size (100 MB), format, Workspace.AccessDenied - [Tongyi Multimodal Translation Official App (API Reference/Web JSSDK)](https://help.aliyun.com/en/model-studio/official-application-tongyi-translate.md): Hub for the Tongyi Multimodal Translation official app in Model Studio: product overview, translation API reference, and web page translation JSSDK integration for multilingual webpage translation. - [Multimodal translation overview](https://help.aliyun.com/en/model-studio/official-application-tongyi-translate-overview.md): Text, image, document (DOCX/PDF/XLSX), and web page translation APIs. Turbo and Plus model tiers. 80+ languages. Glossary intervention, format preservation, batch image translation - [Tongyi Translate API Reference (Overview, Endpoints, RAM, Directory)](https://help.aliyun.com/en/model-studio/tongyi-translate-api-reference.md): Tongyi Translate (anytrans) API reference hub: API overview, service endpoints, RAM authorization, API directory, version changeset. - [API overview](https://help.aliyun.com/en/model-studio/api-anytrans-2025-07-07-overview.md): API standard and pre-built SDKs in multi-languageThe OpenAPI specification of this product (AnyTrans/2025-07-07) follows the ROA standard. Alibaba Clo... - [RAM authorization](https://help.aliyun.com/en/model-studio/api-anytrans-2025-07-07-ram.md): Resource Access Management (RAM) is a service provided by Alibaba Cloud to manage user identities and resource access permissions. Using RAM helps you... - [Web Page Translation JSSDK Integration (Page/Paragraph/Lazy Load)](https://help.aliyun.com/en/model-studio/web-page-translation-jssdk.md): Alibaba Cloud web translation JSSDK v3.0.1: pageTranslate for full-page replacement and paragraphTranslate for bilingual comparison, with lazy loading, dynamic DOM auto-translation, terminology and translation memory. Includes CDN setup, getToken signing, targetSelectors/excludeSelectors config, plus Node.js and Java signature examples. - [Tongyi Deep Search (Official App: Intro/Usage Guide/API Reference)](https://help.aliyun.com/en/model-studio/tongyi-deepsearch.md): Entry for the official Tongyi Deep Search app docs: product introduction, usage guide (how to use), and API reference. Start here for building deep search / deep research features. - [Tongyi Deep Search overview](https://help.aliyun.com/en/model-studio/tongyi-deep-search-introduction.md): Multi-round inference agent for complex research with web search and planning. General (CNY 0.35/call) and legal (CNY 0.75/call) scenarios. 30 free calls on activation - [Tongyi Deep Search user guide](https://help.aliyun.com/en/model-studio/deepsearch-guide.md): Configure Deep Search apps: internet retrieval (Max/Turbo/custom), self-owned or Model Studio knowledge base, code_interpret, dynamic file parsing (10 files, 10 MB), and report generation - [Tongyi DeepSearch API Reference (Endpoints, API List, Error Codes)](https://help.aliyun.com/en/model-studio/deepsearch-api-reference.md): API docs entry for Tongyi DeepSearch: API overview, service access endpoints, API catalog, and error codes. Read before integrating deep search via API. - [Tongyi DeepSearch API Overview](https://help.aliyun.com/en/model-studio/deepsearch-api-overview.md): Five DashScope HTTP APIs: Generate Conversation, Upload File, Manage Conversation Files, Export Report (MD/HTML/PDF), and Connect Knowledge Base. Requires default workspace API key - [Deep Search service endpoints](https://help.aliyun.com/en/model-studio/deepsearch-service-access-point.md): Regional API endpoints for Tongyi Deep Search. China (Beijing) cn-beijing public endpoint: dashscope.aliyuncs.com/api/v2/apps/, no VPC endpoint available - [Model Studio DeepSearch API Catalog (Chat/Files/Report Export/Knowledge Base)](https://help.aliyun.com/en/model-studio/deepsearch-api-list.md): DeepSearch API list: generate conversations, upload files, session file management, report export, connect own knowledge base — 5 APIs for deep-search and research apps. - [DeepSearch Generate Conversations API](https://help.aliyun.com/en/model-studio/deepsearch-chat-generate.md): POST /deep-search-agent/chat/completions. Streaming-only (stream=true required). Key params: agent_id, agent_version, session_files (max 10). Returns planning/thinking/generating phases - [DeepSearch File Upload API](https://help.aliyun.com/en/model-studio/deepsearch-file-upload.md): Three-step upload: apply lease credential, PUT to pre-signed OSS URL, submit for parsing. Poll status until FILE_IS_READY. Pass file_id via session_files in chat API. Python/Java examples - [Deep Search session file management API](https://help.aliyun.com/en/model-studio/deepsearch-session-file-management.md): Manage per-session file trees for legal case review. Endpoints: upsert, delete, get under /deep-search-agent/session/files/. Params: session_id, file_id, file_path, type (file/dir) - [Deep Search report export API](https://help.aliyun.com/en/model-studio/deepsearch-report-export.md): POST /deep-search-agent/file/expose to obtain report files. Params: writing_result_path, writing_result_type (html, html_url, md, md_url, pdf_url). Python and Java examples included - [Deep Search custom knowledge base API](https://help.aliyun.com/en/model-studio/docking-self-built-database.md): HTTP POST spec for connecting custom knowledge bases to Deep Search. Params: query, num (1-100), page, debug. Returns docs with id, title, url, snippet. 3000 ms timeout - [DeepSearch Error Codes](https://help.aliyun.com/en/model-studio/deepsearch-error-code.md): Error reference: InvalidParameter (400), NotFound (404), InternalError (500), LoadConfig.Error, FileConversion.Error, ProcessTimeout.Error, and File.NotFound with troubleshooting guidance - [Qwen Web Search Agent (Official Web Search App Hub)](https://help.aliyun.com/en/model-studio/web-search-agent.md): Hub for the official Qwen Web Search Agent app: product intro, API reference, chat generation calls, error codes. Start here to add web search with real-time info to your model calls. - [Qwen Web Search Agent overview](https://help.aliyun.com/en/model-studio/web-search-agent-guide.md): Real-time web search agent with smart query rewriting, domain-specific tools, citation tracing, and think-and-search mode. Configure search sources, time filters, and system prompts. - [Model Studio Web Search Agent API Reference (agent/agent_max Strategy)](https://help.aliyun.com/en/model-studio/web-search-agent-api.md): API specification for the agent/agent_max web search strategies: request parameters, response structure, and invocation examples for multi-turn deep search reasoning scenarios. - [Qwen Web Search Agent API Overview (Generate Conversation/File Operations)](https://help.aliyun.com/en/model-studio/web-search-agent-api-overview.md): Qwen Web Search Agent APIs (DashScope HTTP): Generate Conversation (web search Q&A) and Multimodal File Operations (image upload). Beta: default-workspace Keys only. - [Model Studio Service Access Points (Public/VPC Endpoints)](https://help.aliyun.com/en/model-studio/service-access-point.md): Service endpoints for Model Studio app calls: cn-beijing (North China 2) public endpoint https://dashscope.aliyuncs.com/api/v2/apps/; no VPC endpoint available. - [Generate a conversation](https://help.aliyun.com/en/model-studio/web-search-agent-api-chat.md): Enables web search and contextual conversation using the Qwen web search agent's agent_id and agent_version. - [Web Search Agent Multimodal File Upload API (Presigned OSS URL)](https://help.aliyun.com/en/model-studio/web-search-agent-api-chat-multimodal-file.md): Web search multimodal file upload: call upload/apply for a presigned OSS URL, PUT the image, then upload/callback with file_id to verify. upload_url expires in 1 hour. - [Qwen Web Search Agent error codes](https://help.aliyun.com/en/model-studio/web-search-agent-error-code.md): Error codes for Web Search Agent API: InvalidParameter (400), NotFound (404), InternalError (500), LoadConfig.Error, ProcessTimeout.Error, and RequestPath.NotFound with causes and solutions. - [Tongyi UI Agent (GUI Agent API Docs Entry)](https://help.aliyun.com/en/model-studio/ui-agent.md): Entry for Tongyi UI Agent on Model Studio: GUI agent API docs (gui-owl service), Agent API protocol, requires an API Key, endpoint dashscope.aliyuncs.com. - [Tongyi UI Agent API (GUI-Owl Device Control)](https://help.aliyun.com/en/model-studio/ui-agent-api.md): UI Agent (GUI-Owl) device-control API: POST gui_agent_server with image/instruction/session_id, model pre-gui_owl_7b, Mobile/PC pipelines driving cross-app automation - [Model Studio App Permission Management (Roles/API Key/Access Control)](https://help.aliyun.com/en/model-studio/application-permission-management.md): App-side permission hub: super admin / workspace admin / regular user role hierarchy, page-level access control, model call rate limiting, API Key management. Model-side permissions are on the model permission page. - [Permission management](https://help.aliyun.com/en/model-studio/application-permission-management-overview.md): Workspace-based access control: super admin, workspace admin, regular user roles. Covers model calling/fine-tuning/deployment permissions, API key scoping, and OpenAPI authorization policies. - [Bailian Assistant API (Sunset in Progress) and Migration to Responses API](https://help.aliyun.com/en/model-studio/assistant-api.md): Sunset in progress: migrate to Responses API. Built-in Thread multi-turn context and tools (code execution, text-to-image, online search); separate from Agent apps. - [Deprecated: Assistant API quick start](https://help.aliyun.com/en/model-studio/quick-start-of-assistant-api.md): Build a painting assistant with the deprecated Assistant API: create Assistant, Thread, Message, Run via DashScope SDK. Includes Quark Search and text_to_image plugins. Migrate to Responses API. - [Deprecated: Assistant API streaming output](https://help.aliyun.com/en/model-studio/streaming-output.md): Streaming output lets you retrieve the real-time running status of an assistant. This lets you display the content generated by the Large Language Model (LLM) to users word by word. - [Assistant API Tool Calling Overview (Retiring, Migrate to Responses API)](https://help.aliyun.com/en/model-studio/tool-calling-overview.md): Assistant API tool calling: Python code interpreter, text-to-image, Quark Search plugins, RAG retrieval, Function Calling. Retiring—migrate to Responses API. - [Deprecated: Assistant API RAG tool](https://help.aliyun.com/en/model-studio/retrieval-augmented-generation.md): Build a RAG agent using the deprecated Assistant API. Example: phone shopping guide with KB creation and retrieval. Requires bailian + DashScope SDK. Migrate to Responses API. - [Deprecated Assistant API function calling](https://help.aliyun.com/en/model-studio/function-calling.md): Function calling via the (being unpublished) Assistant API. Walkthrough: define a translate_text function, register it with the agent, and handle tool_call responses. Migrate to Responses API instead - [Assistant API Code Interpreter](https://help.aliyun.com/en/model-studio/code-interpreter.md): Pre-built plugin that runs Python code for math, data viz, and format conversion. Enable via tools=[{'type': 'code_interpreter'}]. Retrieve results with dashscope.Steps.list. Free during preview - [Assistant API Best Practices (Deprecated)](https://help.aliyun.com/en/model-studio/component-management.md): Production guide for Assistant/Thread/Message/Run/Step components: lifecycle management, workspace isolation, concurrency, local DB storage with TTL cleanup. Migrate to Responses API - [Model Studio Application Tutorials (Agent/RAG/MCP)](https://help.aliyun.com/en/model-studio/application-use-cases.md): Step-by-step application tutorials: agent building, knowledge base (RAG) Q&A, prompt engineering, MCP/plugin integration — build, invoke and publish your app. - [Add AI assistant to website](https://help.aliyun.com/en/model-studio/add-an-ai-assistant-to-your-website-in-10-minutes.md): Four-step tutorial: create a Model Studio agent app, build an AppFlow AI assistant, embed in your website with a few lines of code, then add a RAG knowledge base for domain Q&A - [Integrate a Bailian AI Assistant into WeChat Work (AppFlow/RAG)](https://help.aliyun.com/en/model-studio/add-an-ai-assistant-to-your-work-wechat.md): No-code guide to adding a Bailian AI assistant to WeChat Work: create a Qwen-Plus agent, get API Key and App ID, build an AppFlow connection flow, configure WebhookUrl/Token/EncodingAESKey and trusted IPs, then attach a RAG knowledge base for private-domain answers. Free quota for new users, then per-token billing. - [Add AI assistant to WeChat Official Account](https://help.aliyun.com/en/model-studio/add-an-ai-assistant-to-your-wechat-in-10-minutes.md): Connect a Model Studio RAG app to your WeChat Official Account via AppFlow. Verified accounts use Customer Service Messages API; unverified use Passive Reply with 5s timeout - [Add an AI robot to DingTalk](https://help.aliyun.com/en/model-studio/add-an-ai-assistant-to-your-dingtalk.md): On Alibaba Cloud, you can create an AI robot powered by a large model for your organization in DingTalk—all without writing code. This robot responds to user inquiries 24/7 and answers questions using your private-domain data, serving as a dedicated assistant that improves the user experience and gives your business a competitive advantage. - [Build RAG with local knowledge base](https://help.aliyun.com/en/model-studio/build-rag-application-based-on-local-retrieval.md): Local retrieval + cloud Qwen generation RAG app. Python with LlamaIndex and GTE embedding. Supports pdf/docx/txt/xlsx/csv. Streamlit UI served via uvicorn on port 7866 - [Model Studio Application Support (FAQ/Agreements/After-Sales)](https://help.aliyun.com/en/model-studio/application-support.md): Application-side support hub for Model Studio: FAQ on app building and invocation issues, service agreements, and after-sales support instructions—start here when app development or API calls fail. - [Model Studio Application FAQ: Plugins, RAG, Streaming, Filing](https://help.aliyun.com/en/model-studio/application-faq.md): App FAQs: 6 plugins (Code Interpreter, Quark/GitHub Search); RAG accuracy tuning; incremental_output streaming; error 140010, 100k-doc cap; app filing. - [Model Studio Agreements (Service Agreement/Trial Features/Open Source Terms)](https://help.aliyun.com/en/model-studio/application-related-agreements.md): Legal documents for Bailian (Model Studio): Alibaba Cloud Bailian Service Agreement, Trial Feature Special Notes, and Open Source Model License Terms. Read before activation or use. ## Model Studio Model API Reference Hub (Text/Image/Video/Audio/Realtime) - [Model Studio Model API Reference Hub (Text/Image/Video/Audio/Realtime)](https://help.aliyun.com/en/model-studio/model-api-reference.md): Top hub for model APIs: text generation (Qwen), image/video/3D generation, audio, Realtime API, vector & ranking; plus TPM-reserved DashScope API and file management. - [Prepare to Call Model Studio APIs (API Key/SDK/CLI/Error Codes)](https://help.aliyun.com/en/model-studio/preparations.md): Prep for calling Model Studio model APIs: get and configure an API Key, install SDKs, use the Model Studio CLI, and look up error codes when calls fail. - [Create an API key](https://help.aliyun.com/en/model-studio/get-api-key.md): Create API keys for Model Studio. Beijing, Singapore, Virginia regions with separate base URLs. Permissions: All or Custom (IP whitelist). Coding Plan uses sk-sp-xxxxx format. - [Install the SDK](https://help.aliyun.com/en/model-studio/install-sdk.md): Install DashScope SDK (Python pip, Java Maven/Gradle) or OpenAI SDK (Python, Node.js, Java, Go) for Model Studio API calls. Python requires >= 3.8, Java >= 8. - [Bailian CLI Installation and Authentication](https://help.aliyun.com/en/model-studio/use-model-studio-cli.md): Command-line tool bailian-cli (commands bl/bailian) for AI Agents. Install globally via npm with Node.js ≥ 22.12.0. Authenticate via browser console login or API Key to access model inference and application management on Bailian. - [Model Studio Error Codes (400/401/403/429/500) and Troubleshooting](https://help.aliyun.com/en/model-studio/error-code.md): Quick reference for Alibaba Cloud Model Studio API errors: causes and fixes organized by HTTP status code, covering InvalidParameter, InvalidApiKey, AccessDenied, Throttling, InternalError, plus how to get Request ID and Coding Plan-specific errors. - [Text Generation APIs (OpenAI Compatible/Anthropic Messages/DashScope)](https://help.aliyun.com/en/model-studio/qwen-api-reference.md): Four Model Studio text-generation APIs: OpenAI-compatible Chat Completions, Responses (web search/code interpreter), Anthropic Messages, DashScope (most parameters). - [Model Studio OpenAI-Compatible Chat API (/chat/completions Examples & Parameters)](https://help.aliyun.com/en/model-studio/qwen-api-via-openai-chat-completions.md): Call Model Studio models via the OpenAI-compatible /chat/completions endpoint: base_url and API Key setup, Python/Node.js/curl SDK samples; supports text/image/video input, streaming, Function Calling, web search (enable_search), thinking mode (enable_thinking), and key parameters like temperature, top_p, max_completion_tokens. - [Model Studio OpenAI-Compatible Responses API (Create/Retrieve/Delete)](https://help.aliyun.com/en/model-studio/openai-compatible-responses.md): OpenAI-compatible Responses endpoints on Model Studio: create, retrieve, delete a response, and list input items — call qwen models with the OpenAI SDK. - [Responses API for Qwen (OpenAI-Compatible, Built-in Tools)](https://help.aliyun.com/en/model-studio/qwen-api-via-openai-responses.md): Call qwen3.8-max/qwen3.7-plus/deepseek-v4-pro via the OpenAI-compatible Responses API: input, model, previous_response_id, reasoning.effort parameters; built-in web_search, code_interpreter, web_extractor, file_search and MCP tools; Session cache and streaming examples. - [Model Studio Responses API: Retrieve a Completed Response by ID](https://help.aliyun.com/en/model-studio/retrieve-a-response.md): Retrieve a completed response via GET /responses/{response_id}; only responses created with store=true are retrievable. Returns status and usage fields; Python/curl examples. - [Model Studio - Responses API: Delete a Response by response_id](https://help.aliyun.com/en/model-studio/delete-a-response.md): Delete a stored model response by response_id (resp_xxx, only if store=true). DELETE /responses/{response_id}; Python/curl samples; missing ID returns InvalidParameter. - [Model Studio-ListInputItems Get Response Input Items](https://help.aliyun.com/en/model-studio/list-input-items.md): Retrieve input items for a specified Response, including multi-turn conversation history via previous_response_id. Requires store=true. Paginate by response_id, after, limit, and order to fetch msg_xxx-format messages. OpenAI SDK compatible. - [Model Studio Anthropic-Compatible Messages API (Setup, Parameters, Examples)](https://help.aliyun.com/en/model-studio/anthropic-api-messages.md): Call Model Studio models via the Anthropic Messages API: migrate by replacing api_key, base_url, and model. Supports dedicated endpoints in Beijing, Singapore, US East, Frankfurt, and Tokyo. Covers basic calls, streaming, thinking mode, image/video understanding, Function Call, explicit cache (cache_control), and structured output (output_config), with Python/TypeScript/curl samples and Claude Desktop 404 troubleshooting. - [DashScope native API for text generation](https://help.aliyun.com/en/model-studio/qwen-api-via-dashscope.md): Qwen text/multimodal models via DashScope HTTP endpoints (/text-generation, /multimodal-generation). Multi-region: Beijing, Singapore, Virginia, Frankfurt. Python and Java SDK examples - [Image Generation Model APIs (Qwen/Wan/Z-Image/Kling/Vidu)](https://help.aliyun.com/en/model-studio/image-generation.md): Image generation API hub: Qwen Image, Wan, Z-Image, Kling, Vidu models, creative tools, and FAQ on failures/error codes. Start here for text-to-image API calls. - [Qwen Image API Docs: Generation & Editing 3.0, Image Translation](https://help.aliyun.com/en/model-studio/qwen-image-api-reference.md): Entry to Qwen image API docs: qwen-image generation and editing 3.0, image translation (qwen-mt-image), and early image models — call references and parameter details. - [Qwen Image Generation & Editing 3.0 API Reference](https://help.aliyun.com/en/model-studio/qwen-image-generation-and-editing-api-reference.md): API reference for qwen-image-3.0-pro and qwen-image-3.0 models supporting text-to-image (T2I) and image-to-image (I2I) editing with 1-3 reference images. Output resolution 512*512 to 2048*2048, parameters include prompt_extend, n, size, negative_prompt. Generated image URLs expire in 24 hours. - [Legacy Qwen Image Models (qwen-image-max/plus/edit)](https://help.aliyun.com/en/model-studio/legacy-qwen-image-models.md): Specs for legacy Qwen image models: qwen-image-max/plus, qwen-image, and qwen-image-edit series — capabilities, resolution limits, rate limits. For new work, use qwen-image-3.0. - [Qwen-Image text-to-image API](https://help.aliyun.com/en/model-studio/qwen-image-api.md): Text-to-image generation via DashScope API. Models: qwen-image-2.0-pro, qwen-image-max, qwen-image-plus. Params: negative_prompt, size (up to 2048x2048), prompt_extend, seed. Sync and async modes - [Qwen-Image Edit API](https://help.aliyun.com/en/model-studio/qwen-image-edit-api.md): Image editing via natural language prompts. Up to 3 input images, 1-6 outputs. Models: qwen-image-2.0-pro, qwen-image-edit-max/plus. Params: negative_prompt, size, seed - [Qwen-MT-Image API](https://help.aliyun.com/en/model-studio/qwen-mt-image-api.md): Translate text in images while preserving layout. Async HTTP API (create task + poll). Params: source_lang, target_lang, domainHint, sensitives, terminologies, imageSegment. 15 languages supported - [Wan Image Generation & Editing API Reference (wan2.7/2.6/2.5)](https://help.aliyun.com/en/model-studio/wan-image-api-reference.md): Wan image APIs: wan2.7-image-pro (4K), wan2.6-image mixed text-image output, image editing 2.5, text-to-image; async task polling and size parameters. - [Wan text-to-image V2 API](https://help.aliyun.com/en/model-studio/text-to-image-v2-api-reference.md): Generate images from text prompts. Models: wan2.6-t2i (recommended), wan2.5-t2i-preview, wan2.2-flash/plus. Supports sync and async HTTP calls, prompt_extend, negative_prompt, seed control - [Wanx text-to-image V1 API](https://help.aliyun.com/en/model-studio/text-to-image-api-reference.md): Legacy text-to-image model (wanx-v1). Supports 10 art styles, reference image transfer (repaint/refonly modes), negative prompts. Beijing region only. V2 recommended for new projects - [Wan2.7 image generation and editing API](https://help.aliyun.com/en/model-studio/wan-image-generation-and-editing-api-reference.md): Five modes: text-to-image, text-to-image-set, image-to-image-set, image editing, and multi-image reference generation. Models: wanx2.1-t2i-turbo and variants. Successor to Wan2.6. - [Wan2.6 image generation and editing API](https://help.aliyun.com/en/model-studio/wan-image-generation-api-reference.md): Multi-image input, image editing, and interleaved text-image output. Models: wanx2.0-t2i-turbo and variants. Earlier generation; see Wan2.7 for latest capabilities. - [Wanxiang 2.5 image editing API](https://help.aliyun.com/en/model-studio/wan2-5-image-edit-api-reference.md): Edit or fuse multiple images using text instructions only, with subject consistency across edits. Model: wan2.5-imageedit. No mask or region selection required. - [Wanxiang Universal Image Editing API](https://help.aliyun.com/en/model-studio/wanx-image-edit-api-reference.md): Multi-feature image editor: outpainting, inpainting, watermark removal, style transfer, super-resolution, colorization. Model: wanx2.1-imageedit. Instruction-based. - [Wanx Sketch-to-Image API (wanx-sketch-to-image-lite, Beijing Only)](https://help.aliyun.com/en/model-studio/wanx-sketch-to-image-api-reference.md): wanx-sketch-to-image-lite: generate doodle art from sketches and text prompts. Beijing-region API Key required. Async via HTTP/DashScope SDK. ¥0.06/image, 500 free images. - [Wan Image Local Repaint API (wanx-x-painting)](https://help.aliyun.com/en/model-studio/vary-region-api-reference.md): wanx-x-painting: local image repaint from original image + mask + prompt. Free trial, 500-image quota; Beijing-region API Key required. Alternatives: Qwen Image Edit, Wan 2.1. - [Z-Image Text-to-Image API (z-image-turbo Synchronous Call)](https://help.aliyun.com/en/model-studio/z-image-generation-api-reference.md): Z-Image text-to-image API: synchronous HTTP call with model fixed to z-image-turbo (6B params, 8-step inference); prompt_extend=true enables smart thinking and returns the rewritten prompt. - [Z-Image API reference](https://help.aliyun.com/en/model-studio/z-image-api-reference.md): Synchronous text-to-image via z-image-turbo model. Resolution 512x512 to 2048x2048, Chinese/English text rendering, optional prompt_extend for LLM-optimized prompts. Image URLs valid 24 hours - [Kling](https://help.aliyun.com/en/model-studio/kling-image-api-reference.md) - [Kling image generation API](https://help.aliyun.com/en/model-studio/kling-image-generation-api-reference.md): Text-to-image and reference image-to-image via kling-v3-image-generation and kling-v3-omni (storyboard mode). Output: 1k/2k/4k, 1-9 images. Supports element_list for subjects. - [Vidu Video & Image Generation Models (viduq3/viduq2)](https://help.aliyun.com/en/model-studio/vidu-image-models.md): Vidu models on Model Studio: viduq3-ad/viduq3-drama reference-to-video, viduq2-pro image-to-video, viduq3-fast/viduq2 reference2image models, with invocation steps and prompt guides. - [Vidu-Image Generation API Reference](https://help.aliyun.com/en/model-studio/vidu-image-generation-api-reference.md): The Vidu Image Generation models support text-to-image, image editing and reference image-to-image tasks. - [Model Studio Creative Tools API Reference (Portrait Repaint/Poster/Background)](https://help.aliyun.com/en/model-studio/image-creative-tools-api-reference.md): API reference for image creative tools: portrait style repaint, creative poster generation (wanx-poster-generation-v1), background generation, virtual model (virtualmodel-v2), sketch-to-image, outpainting. Requires Beijing region API Key. - [Portrait style repaint API](https://help.aliyun.com/en/model-studio/portrait-style-redraw-api-reference.md): Transform portraits into artistic styles using wanx-style-repaint-v1. 20+ presets (anime, 3D, Chinese painting) or custom via style_ref_url. HTTP-only async API. Output: 1536px short side - [Image Out-Painting API (image-out-painting)](https://help.aliyun.com/en/model-studio/image-scaling-api.md): image-out-painting outpainting API: expand images by aspect ratio, scale, or directional pixels, with rotation. Async two-step call, Beijing region only. ¥0.18/image, 500 free. - [Wanx Virtual Model API (Swap Model & Background in Product Photos)](https://help.aliyun.com/en/model-studio/virtual-model-api-details.md): wanx-virtualmodel/virtualmodel-v2: swap model & background of product photos, pose kept; Beijing only, free trial (500 images); use Qwen/Wanx2.1 image editing. - [Shoe Model API (shoemodel-v1 Async Try-on Generation)](https://help.aliyun.com/en/model-studio/shoe-model-api.md): shoemodel-v1: multi-angle shoe images + model template → async HTTP AI try-on. Beijing-region API Key only; 500 free images, unavailable after quota exhaustion. - [Creative poster generation API](https://help.aliyun.com/en/model-studio/creative-poster-generation-api.md): Async API (wanx-poster-generation-v1) to generate posters with auto text layout. Params: title, prompt_text, lora_name for 18 styles. Supports sr/hrf upscaling. Free trial, Beijing region - [Human instance segmentation API](https://help.aliyun.com/en/model-studio/image-instance-segmentation-api-reference.md): Async API to detect people in images and generate pixel-level masks per person. Model: image-instance-segmentation. Free trial only (500 images). Includes mask-splitting Python code. - [Wanxiang background generation API](https://help.aliyun.com/en/model-studio/wanx-background-generation-api-reference.md): Generate backgrounds for product images. Model: wanx-background-generation-v2. Methods: text-guided, image-guided, or combined. Supports e-commerce and poster scenarios. - [Image Erase Completion (image-erase-completion API)](https://help.aliyun.com/en/model-studio/image-erase-completion-api-reference.md): image-erase-completion API: erase people/objects/watermarks via mask_url and inpaint the background. Beijing region only; free trial, 500 images, no paid tier. - [AI Try-On OutfitAnyone (aitryon Fitting/Refiner/Segmentation)](https://help.aliyun.com/en/model-studio/outfitanyone.md): AI try-on OutfitAnyone: aitryon/aitryon-plus fitting, aitryon-refiner touch-up, aitryon-parsing-v1 segmentation (partial try-on, bbox). Beijing-region API Key only. - [OutfitAnyone Basic Edition API](https://help.aliyun.com/en/model-studio/outfitanyone-api.md): Virtual try-on from flat-lay clothing and full-body portraits. Async API: POST /api/v1/services/aigc/image2image/image-synthesis. Params: person_image_url, top/bottom_garment_url. CNY 0.20/image. - [OutfitAnyone Plus API](https://help.aliyun.com/en/model-studio/aitryon-plus-api.md): High-quality virtual try-on model (aitryon-plus). Supports single/combo garment try-on, face policy control, and custom resolution. Async HTTP API, CNY 0.50/image. Better texture than Basic edition. - [AI Virtual Try-On Image Refinement API (aitryon-refiner, Beijing Only)](https://help.aliyun.com/en/model-studio/ai-fitting-picture-finishing-api-details.md): aitryon-refiner post-processes AI try-on images to boost realism and sharpness. Async HTTP: POST to create a task, GET to poll task_id. Inputs must match the Basic/Plus try-on call; Beijing-region API Key required. 400-image free quota, result URL valid 24h. - [OutfitAnyone Parsing API](https://help.aliyun.com/en/model-studio/aitryon-parsing-api.md): Segment clothing from model images via aitryon-parsing-v1. Returns cropped garment image (crop_img_url), parsed mask (parsing_img_url), and bounding box coordinates. Use for partial try-on. - [OutfitAnyone billing and metering](https://help.aliyun.com/en/model-studio/billing-for-outfitanyone.md): Per-image pricing for AI Try-on models: aitryon CNY 0.20, aitryon-plus CNY 0.50, aitryon-refiner tiered from CNY 0.30. Free tier: 400 images per model, valid 90 days. RPS limit: 10. - [FaceChain Portrait Generation (LoRA Training, Only 2 Photos)](https://help.aliyun.com/en/model-studio/facechain-portrait-generation.md): FaceChain: train a personal portrait model from just 2 photos via LoRA, then batch-generate styled portraits. Face detection/training/generation APIs; Beijing region API Key only. - [FaceChain Portrait Quick Start (LoRA Training/Train-free)](https://help.aliyun.com/en/model-studio/facechain-quick-start.md): Wanx FaceChain portrait generation: LoRA training and trainfree modes; flow of image detection, file upload, training, generation; Java demo. Beijing-region API Key and trial application required - [Model Studio: FaceChain Face Detection API (facechain-facedetect)](https://help.aliyun.com/en/model-studio/facechain-face-detection-api.md): facechain-facedetect sync API: checks face count/size/angle/sharpness against FaceChain fine-tuning criteria; batch image input, returns is_face/failed_reason. Beijing-region API Key only. - [FaceChain training API](https://help.aliyun.com/en/model-studio/facechain-finetune-api.md): Async API (facechain-finetune) to train a character model from 1-10 face images. Upload via OSS URLs or DashScope file service. Produces resource ID for generation. QPS: 2, Beijing only - [FaceChain portrait generation API](https://help.aliyun.com/en/model-studio/facechain-generation.md): Async API for facechain-generation model. LoRA training mode (preset styles) and training-free mode (custom template). POST to /aigc/album/gen_potrait, poll via task_id. Beijing only. - [FaceChain billing](https://help.aliyun.com/en/model-studio/facechain-billing.md): FaceChain pricing: detection (facechain-facedetect) free, training (facechain-finetune) 2.5 CNY/session with 50 free, generation 0.18 CNY/image with 500 free. Beijing region, trial required - [WordArt Creative Text API (Deformation & Texture Generation)](https://help.aliyun.com/en/model-studio/wordart-quick-start.md): WordArt Jinshu creative text: deformation and wordart-texture generation APIs; 3 custom + 18 preset styles via prompts. Beijing region only, Beijing API Key required. - [Model Studio - WordArt Text Deformation API (wordart-semantic)](https://help.aliyun.com/en/model-studio/word-transformer.md): wordart-semantic: deforms text edge contours per prompt, returns black-bg white mask. Async submit-query mode; params input.text/input.prompt. Requires Beijing-region API Key. - [WordArt Text Texture Generation API (wordart-texture, Async)](https://help.aliyun.com/en/model-studio/fill-texture-effect-api.md): Async wordart-texture API: prompt-driven textured art lettering from text/image input; 3 custom + 20 preset styles, style reference image, transparent output; Beijing API Key only - [WordArt Billing (wordart-texture/wordart-semantic Pricing & Free Quota)](https://help.aliyun.com/en/model-studio/wordart-billing.md): Beijing region only: wordart-texture CNY 0.08/image, wordart-semantic CNY 0.24/image; 500 free images each (90 days). Async calls limited to QPS 2, concurrency 1. - [Model Studio Image API FAQ (Debugging, Billing, Rate Limits, Errors)](https://help.aliyun.com/en/model-studio/image-faq.md): 13 image models FAQ (wanx text2image, doodle, inpainting): curl debugging, 500-image free quota valid 90 days, pricing, rate limits, InputDownloadFailed errors. - [Video Generation](https://help.aliyun.com/en/model-studio/video-generation-api.md) - [Model Studio - HappyHorse Video Generation API (Text-to-Video/Image-to-Video/Reference-to-Video/Video Editing)](https://help.aliyun.com/en/model-studio/happyhorse-api-reference.md): HappyHorse video generation APIs: happyhorse-1.1-t2v text-to-video, i2v image-to-video, r2v reference-to-video, 1.0-video-edit video editing, DashScope async. - [HappyHorse Text-to-Video API Reference](https://help.aliyun.com/en/model-studio/happyhorse-text-to-video-api-reference.md): Generate videos from text prompts using happyhorse-1.0-t2v or happyhorse-1.1-t2v models. Supports nine aspect ratios (16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 9:21, 21:9), three resolutions (480P/720P/1080P), and 3-15s duration. Async workflow via task creation and polling; video URL valid for 24 hours. - [HappyHorse Image-to-Video API (First Frame)](https://help.aliyun.com/en/model-studio/happyhorse-image-to-video-api-reference.md): HappyHorse first-frame image-to-video: models happyhorse-1.1-i2v/1.0-i2v, async create then poll by task_id (24h); prompt, 720P, duration; tasks run 1-5 min. - [HappyHorse reference-to-video API](https://help.aliyun.com/en/model-studio/happyhorse-reference-to-video-api-reference.md): Generate video from multiple reference images plus text prompt. Combines subjects from images into a scene. Same async workflow as other HappyHorse APIs. - [HappyHorse video editing API](https://help.aliyun.com/en/model-studio/happyhorse-video-edit-api-reference.md): Edit existing videos using a reference image and text instructions. Style transfer and local replacement. Input: source video + reference image + text. Same async API pattern. - [Wan Video Generation API Reference (Text-to-Video/Image-to-Video)](https://help.aliyun.com/en/model-studio/wan-api-reference.md): Wan video model API hub: wan3.0 video generation, wan2.7 text-to-video, image-to-video, reference-to-video, video editing, motion generation, digital human, legacy 2.1-2.6 models. - [Wanx 3.0 Video Generation API Reference](https://help.aliyun.com/en/model-studio/wan3-video-generation-api-reference.md): API guide for wan3.0-video model covering text-to-video, first-frame/first-last-frame image-to-video, and reference-file-to-video generation up to 30 seconds. Async workflow includes task creation and polling; model, Endpoint URL, and API Key must be in the same region (Beijing/Singapore). task_id valid for 24 hours. - [Wan 2.7 image-to-video API](https://help.aliyun.com/en/model-studio/image-to-video-general-api-reference.md): Multimodal video generation: first-frame, first-and-last-frame, and video continuation. Accepts text, images, audio, video. Params: resolution (720P/1080P), duration (2-15s), prompt_extend. - [Wan text-to-video API](https://help.aliyun.com/en/model-studio/text-to-video-api-reference.md): Generate videos from text prompts via async HTTP API. Latest model: wan2.7. Tasks take 1-5 minutes; create task then poll for result. Multi-region: Beijing, Singapore - [Wan reference-to-video API](https://help.aliyun.com/en/model-studio/wan-video-to-video-api-reference.md): Generate videos featuring single or multiple characters from reference images/videos. Model: wan2.1-i2v and R2V variants. Multimodal input, distinct from Vidu R2V. - [Wan2.7 video editing API](https://help.aliyun.com/en/model-studio/wan-video-editing-api-reference.md): Edit existing videos via text instructions, reference images, or source videos. Supports instruction-based editing and style transfer. Async DashScope API with multimodal input. - [Wan Legacy Video Models wan2.1–wan2.6 (Early Video Generation)](https://help.aliyun.com/en/model-studio/legacy-video-models.md): Early Wan video generation models wan2.1–wan2.6, including text-to-video and image-to-video variants; for the latest video models refer to the wan2.7 series, and for current pricing see Billing. - [Wan Image-to-Video API (First-Frame, wan2.1–wan2.6)](https://help.aliyun.com/en/model-studio/legacy-image-to-video-api-reference.md): Alibaba Cloud Model Studio Wan image-to-video API reference: generate video from a first-frame image and prompt using wan2.6/wan2.5/wan2.2/wanx2.1 models. Covers HTTP and DashScope SDK calls, multi-shot narrative, auto dubbing, resolution (480P/720P/1080P), duration (2–15s), and async task polling. Wan 2.7 is recommended for new projects. - [Wanx image-to-video effect templates](https://help.aliyun.com/en/model-studio/wanx-video-effects.md): Apply preset motion effects (flying, squish, carousel, hug) to a first-frame image via the template parameter. Supports wanx2.1-i2v-plus and i2v-turbo models with 720P output - [Wan text-to-video API (legacy, 2.1-2.6)](https://help.aliyun.com/en/model-studio/legacy-wan-text-to-video-api-reference.md): Legacy async API for text-to-video generation using wan2.5, wan2.2, or wanx2.1 models. Endpoint: /api/v1/services/aigc/video-generation/video-synthesis. Processing takes 1-5 minutes. - [Wanx Reference-to-Video (wan2.6-r2v) API Reference](https://help.aliyun.com/en/model-studio/legacy-wan-reference-to-video-api-reference.md): Legacy Wanx 2.6 reference-to-video API: generate single-character or multi-character interaction videos via reference_urls (images/videos), with shot_type=multi for multi-shot and audio=true. Async workflow (create task + poll task_id); only wan2.6 models supported; Endpoint and API Key must be in the same region. - [Wan 2.2 first-and-last-frame-to-video API (legacy)](https://help.aliyun.com/en/model-studio/legacy-image-to-video-by-first-and-last-frame-api-reference.md): Legacy async API (Wan 2.2) to generate transitioning video from first frame, last frame, and text prompt. Endpoint: /api/v1/services/aigc/image2video/video-synthesis. Processing takes 1-5 minutes. - [Wan 2.1 video editing API (legacy, VACE)](https://help.aliyun.com/en/model-studio/legacy-wanx-vace-api-reference.md): Legacy video editing via wanx2.1-vace-plus. Functions: image_reference, video_repainting, video_edit (mask-based), video_extension, video_outpainting. 5-10 min processing. - [Wan image-to-action API](https://help.aliyun.com/en/model-studio/wan-animate-move-api.md): Transfer actions and expressions from a reference video to a character image. Model: wan2.2-animate-move. Two modes: wan-std (standard) and wan-pro (professional). - [Wan Video Person Swap API (wan2.2-animate-mix)](https://help.aliyun.com/en/model-studio/wan-animate-mix-api.md): wan2.2-animate-mix replaces the person in a reference video with one from an image, preserving scene and lighting. wan-std/wan-pro modes; async only (X-DashScope-Async: enable, create task, poll task_id). Beijing and Singapore use separate API Keys. - [Wan Digital Human wan2.2-s2v (Image + Audio to Talking Video)](https://help.aliyun.com/en/model-studio/wan-s2v-overview.md): wan2.2-s2v animates an image + audio into talking/singing/performing videos (full-body/half-body/portrait). Beijing only. ¥0.004/image, ¥0.5/s 480P, ¥0.9/s 720P; vs EMO. - [Wan Digital Human wan2.2-s2v Image Detection API](https://help.aliyun.com/en/model-studio/wan-s2v-detect-api.md): wan2.2-s2v-detect checks if an image meets s2v digital human input specs, returning check_pass/humanoid. Beijing region only. ¥0.004/image, 200 free - [Wanx Digital Human wan2.2-s2v Video Generation API (Image + Audio Driven)](https://help.aliyun.com/en/model-studio/wan-s2v-api.md): Generate talking/singing/performing videos from one image + audio, real or cartoon characters, 480P(¥0.5/s)/720P(¥0.9/s), 100s free quota, Beijing region API Key only - [Portrait Animation API Reference (LivePortrait Video Generation)](https://help.aliyun.com/en/model-studio/portrait-animation-api-reference.md): LivePortrait animates a portrait image with voice audio into a talking-head video; LivePortrait-detect validates input images. Requires Beijing region API Key. - [Dancing Human AnimateAnyone: Image-to-Dance Video (3-Model Pipeline)](https://help.aliyun.com/en/model-studio/animateanyone-quick-start.md): Generate dance videos from a portrait image via three models: animate-anyone-detect-gen2 (image check), -template-gen2 (motion template), animate-anyone-gen2 (video). Beijing region only with its API Key; pay-as-you-go at 0.004 CNY/image, 0.08 CNY/sec; dedicated deployment available. - [AnimateAnyone image detection API](https://help.aliyun.com/en/model-studio/animate-anyone-detect-api.md): Validate whether an image passes requirements for AnimateAnyone video generation. Model: animate-anyone-detect-gen2. Returns check_pass boolean and bodystyle (half/full). - [AnimateAnyone Motion Template Generation API (animate-anyone-template-gen2)](https://help.aliyun.com/en/model-studio/animate-anyone-template-api.md): Extract motion from video into AnimateAnyone video-generation templates via animate-anyone-template-gen2; 2-60s mp4/avi/mov, async HTTP, Beijing API Key only - [AnimateAnyone Video Generation API (animate-anyone-gen2)](https://help.aliyun.com/en/model-studio/animateanyone-video-generation-api.md): animate-anyone-gen2 generates human motion videos from a portrait image and a motion template, with image or video background. Beijing region API Key only; async HTTP call. - [EMO Portrait Image-to-Singing-Video (Image + Audio to Dynamic Video)](https://help.aliyun.com/en/model-studio/emo-quick-start.md): EMO turns a portrait image + voice audio into a dynamic portrait video. emo-detect-v1 image check + emo-v1 video generation; Beijing-region API Key only. From ¥0.08/s, 1,800s free. - [EMO Image Detection API (emo-detect-v1)](https://help.aliyun.com/en/model-studio/emo-detect-api.md): emo-detect-v1 verifies a portrait image meets the EMO video generation spec, returns face_bbox/ext_bbox coordinates; 1:1 and 3:4 ratios; Beijing-region API Key only. - [EMO video generation API](https://help.aliyun.com/en/model-studio/emo-api.md): Async API (emo-v1) to animate portrait images with voice audio into talking-head videos. Requires face_bbox/ext_bbox from EMO detection API. style_level: calm/normal/active. Beijing region - [LivePortrait Quick Start: Talking Portrait Video from Image and Audio](https://help.aliyun.com/en/model-studio/liveportrait-quick-start.md): Talking portrait video from a portrait image + voice audio: liveportrait-detect validates images, liveportrait renders video; Beijing API Key only; ¥0.004/image, ¥0.02/sec - [LivePortrait image detection API](https://help.aliyun.com/en/model-studio/liveportrait-detect-api.md): Validate portrait images before LivePortrait video generation. Endpoint: /api/v1/services/aigc/image2video/face-detect. Returns pass (bool) and failure message. Image < 10 MB, aspect ratio <= 2. - [LivePortrait video generation API](https://help.aliyun.com/en/model-studio/liveportrait-api.md): Async API to generate talking-head videos from portrait image + audio. Params: template_id (calm/normal/active), eye_move_freq, mouth_move_strength, video_fps. Requires LivePortrait-detect first. - [VideoRetalk Lip-Sync Replacement (Person Video + Audio)](https://help.aliyun.com/en/model-studio/videoretalk.md): Match lip movements in a person video to input voice audio. Beijing region only, regional API Key required. ¥0.08/sec by video length, 1800-sec free tier, API only. - [VideoRetalk Talking-Portrait Video API (Audio Lip Sync/Async)](https://help.aliyun.com/en/model-studio/videoretalk-api.md): VideoRetalk lip-syncs portrait video to voice audio for talking-head video; Beijing region API Key only; async submit+query; video mp4/avi/mov ≤300MB 2-120s, audio wav/mp3/aac. - [Emoji Video Generation from Portrait Photo (Beijing Region Only)](https://help.aliyun.com/en/model-studio/emoji-quick-start.md): Generate emoji videos from a portrait photo with preset templates: image detection API then video generation API with template ID (jingdian_xianqi). Beijing region only; free quota available. - [emoji-detect-v1 Image Detection for Emoji Video Generation](https://help.aliyun.com/en/model-studio/emoji-detect-api.md): emoji-detect-v1 validates images for Emoji video generation (frontal face, no occlusion), returns face_bbox/ext_bbox_face coordinates; Beijing region API Key only - [Emoji Meme Video Generation API (emoji-v1, Async)](https://help.aliyun.com/en/model-studio/emoji-api.md): Async emoji-v1 API: generate meme videos from portrait image + template ID (driven_id). Needs face_bbox/ext_bbox from Emoji detection. Beijing-region API Key only. - [Video Style Transform API](https://help.aliyun.com/en/model-studio/video-style-transform-api-reference.md): Convert videos into 8 artistic styles (manga, comic, 3D cartoon, ink wash). Model: video-style-transform. Params: style (0-7), video_fps, use_SR for super-resolution. 720p/540p output. - [Model Studio PixVerse Video Generation API (Text-to-Video/Image-to-Video/Lip Sync)](https://help.aliyun.com/en/model-studio/pixverse-api-reference.md): PixVerse (Aishi) video API hub: text-to-video, image-to-video (first/first-last frame), reference-to-video, lip sync, motion control, video upscaling — 7 API references. - [PixVerse text-to-video API](https://help.aliyun.com/en/model-studio/pixverse-text-to-video-api-reference.md): Generate video from text prompts only. Models: pixverse-c1-t2v (dynamic), pixverse-v6-t2v (general). Multi-shot via shot_type param. 360P-1080P, 1-15s, optional audio generation - [PixVerse image-to-video API](https://help.aliyun.com/en/model-studio/pixverse-image-to-video-api-reference.md): Generate video from one input image plus text prompt. Models: pixverse-c1-it2v (dynamic), pixverse-v6-it2v (general). Params: resolution (360P-1080P), duration (1-15s), audio, multi-shot - [PixVerse keyframe-to-video API](https://help.aliyun.com/en/model-studio/pixverse-keyframe-to-video-api-reference.md): Generate video from first and last frame images plus prompt. Models: pixverse-c1-kf2v, pixverse-v6-kf2v. Media array requires type=first_frame and type=last_frame. 360P-1080P, 1-15s - [PixVerse Reference-to-Video API (pixverse-v6-r2v-omni/pixverse-c1-r2v)](https://help.aliyun.com/en/model-studio/pixverse-reference-to-video-api-reference.md): Pass reference images/videos via media and cite subjects with @ref_name in prompts. Async: create a task and poll with task_id (1–5 min). Beijing region only with a Beijing API Key; resolution/duration parameters. - [Video Generation - PixVerse Lip-Sync API (pixverse-lipsync)](https://help.aliyun.com/en/model-studio/pixverse-lipsync-api-reference.md): pixverse-lipsync lip-sync video: video + audio_url or TTS text input; Beijing-region API Key only; async only (X-DashScope-Async: enable required), 1-5 min per task. - [PixVerse Motion Control API (pixverse-motioncontrol, Beijing Only)](https://help.aliyun.com/en/model-studio/pixverse-motioncontrol-api-reference.md): PixVerse Motion Control API: transfer actions from a reference video onto a character image to generate choreography-replica videos. Beijing region only, async create-task + poll task_id; 360P/540P/720P; image ≤20MB, video ≤100MB and ≤30s. - [PixVerse Video Upscale API (4K Super-Resolution, Beijing Region Only)](https://help.aliyun.com/en/model-studio/pixverse-upscale-api-reference.md): Enable and call pixverse/pixverse-upscale to upscale videos to 4K. Async task + task_id polling, Beijing-region API Key only; input MP4/MOV/WebM, ≤100MB, ≤30s. - [Kling Image & Video Generation API (kling-v3)](https://help.aliyun.com/en/model-studio/kling-api-reference.md): Kling API: kling-v3 image generation (text-to-image, reference-to-image) and kling-v3-omni video generation. Beijing region only; Beijing API Key required. - [Kling Video Generation API (Text-to-Video/Image-to-Video)](https://help.aliyun.com/en/model-studio/kling-video-generation-api-reference.md): Alibaba Cloud Model Studio Kling AI video generation API reference: kling-v3-omni/v3 models for text-to-video, first/last-frame image-to-video, reference-based video and video editing; async calls via task_id polling, Beijing region only, with mode/duration/audio parameters and billing notes. - [Kling entity ID list](https://help.aliyun.com/en/model-studio/kling-object-ids.md): Reference table of Kling entity IDs (element_id) for image/video generation APIs. Entities include Code Rain (101), Lightning (103), Magic Circle (104), Snowflakes (108), Wormhole (109), and more. - [Vidu Video Generation API (viduq3/viduq2 img2video & reference2video)](https://help.aliyun.com/en/model-studio/vidu-api-reference.md): Vidu video generation APIs on Model Studio: call viduq3/viduq2 img2video (image-to-video) and reference2video models, plus reference2image image generation. - [Vidu Text-to-Video API (Beijing Region Only)](https://help.aliyun.com/en/model-studio/vidu-text-to-video-api-reference.md): Beijing region only. Vidu text-to-video (viduq3-pro/turbo, viduq2): prompts to 540P/720P/1080P video, 1-16s; async task + polling, HTTP & DashScope SDK samples. - [Vidu image-to-video API](https://help.aliyun.com/en/model-studio/vidu-image-to-video-api-reference.md): Generate video from an image and text prompt. Models: viduq3-pro/turbo, viduq2-pro/turbo. Params: resolution (540P-1080P), duration (1-16s), audio, watermark, seed. - [Vidu keyframe-to-video API](https://help.aliyun.com/en/model-studio/vidu-keyframe-to-video-api-reference.md): Generate video transitioning between first and last frame images with a text prompt. Model: viduq3-turbo_start-end2video. Requires 2 images in media array. Duration 1-16s. - [Vidu Reference-to-Video API (viduq3-mix/ad/drama_reference2video)](https://help.aliyun.com/en/model-studio/vidu-reference-to-video-api-reference.md): Alibaba Cloud Model Studio Vidu reference-to-video API: input reference images plus text prompt to blend the subject into the described scene and generate smooth video. Supports viduq3-mix/ad/drama_reference2video models, China (Beijing) region only, async workflow (create task then poll task_id); parameters include duration, size, resolution, and watermark. - [MiniMax](https://help.aliyun.com/en/model-studio/minimax-video-api-reference.md) - [MiniMax Video Generation API Reference](https://help.aliyun.com/en/model-studio/minimax-video-generation-api-reference.md): The MiniMax Video Generation model supports text-to-video, image-to-video (first frame), image-to-video (last frame), image-to-video (first and last frames), and multimodal reference-to-video. - [3D Generation API Reference (Tripo Text/Image-to-3D)](https://help.aliyun.com/en/model-studio/3d-generation.md): Tripo text-to-3D, single/multi-image-to-3D APIs (Tripo-H3.1/P1.0). Async create-task-then-poll flow; China North 2 (Beijing) region only, requires a Beijing API Key. - [Tripo 3D Model Generation API (Text/Image-to-3D, Async Task)](https://help.aliyun.com/en/model-studio/tripo-3d-generation-api-reference.md): Text/single-image/multi-image to 3D (GLB+PBR). Async: poll task_id for results. Beijing region + API Key only. Models Tripo-P1.0/H3.1; texture_quality param. - [Model Studio Audio API Reference (ASR/TTS/Music/Speech Translation)](https://help.aliyun.com/en/model-studio/audio-api-references.md): Hub for audio model API references: speech recognition, speech synthesis, music generation, speech translation and voice conversation, with parameters and call examples. - [Speech Recognition APIs (Realtime/Non-realtime: Qwen-ASR/Fun-ASR/Paraformer)](https://help.aliyun.com/en/model-studio/speech-recognition-api-reference.md): Speech recognition API index: realtime streaming (Qwen-ASR-Realtime/Paraformer/Fun-ASR), non-realtime file transcription (Fun-ASR/Qwen-ASR), and custom hotwords. - [Real-time Speech Recognition API (Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime)](https://help.aliyun.com/en/model-studio/fun-asr-real-time-speech-recognition-api-reference.md): Real-time ASR API: streaming audio to Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime models, incremental text output; file transcription uses the non-real-time API. - [Realtime ASR WebSocket API (Fun-ASR-Realtime/Qwen-Audio-3.0-ASR-Flash-Streaming)](https://help.aliyun.com/en/model-studio/fun-asr-realtime-websocket-api.md): Realtime ASR WebSocket API: Beijing/Singapore wss endpoints, Authorization Bearer API Key header (401/403 if invalid), run-task/audio/finish-task flow, mono audio only. - [Real-time ASR Client Events](https://help.aliyun.com/en/model-studio/fun-asr-client-events.md): WebSocket client events for Qwen-Audio-3.0-ASR-Flash-Streaming and Fun-ASR-Realtime speech recognition, covering run-task, continue-task, and finish-task request structures and field constraints - [Bailian Fun-ASR Real-time ASR Server Events](https://help.aliyun.com/en/model-studio/fun-asr-server-events.md): WebSocket server events for Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime: task-started, result-generated, task-finished, task-failed, with fields like sentence_end, word-level timestamps, and usage.duration billing - [Real-Time ASR Python SDK (qwen-audio-3.0-asr-flash-streaming)](https://help.aliyun.com/en/model-studio/fun-asr-realtime-python-sdk.md): Recognition class: call for non-streaming, start/stop for bidirectional streaming transcription, RecognitionCallback for real-time results, 16000Hz sample rate, DashScope SDK required. - [Model Studio Real-time ASR Java SDK (Qwen-Audio-3.0/Fun-ASR)](https://help.aliyun.com/en/model-studio/fun-asr-realtime-java-sdk.md): DashScope Java SDK for Qwen-Audio-3.0-ASR-Flash-Streaming and Fun-ASR-Realtime: non-streaming call, duplex streaming via callback/Flowable, sendAudioFrame; parameters include sampleRate, format, language_hints, vocabulary, semantic_punctuation_enabled, heartbeat; dedicated endpoints in Beijing and Singapore. - [Model Studio Real-time ASR Android SDK (Qwen-Audio-3.0/Fun-ASR)](https://help.aliyun.com/en/model-studio/android-sdk-for-fun-asr-real-time-service.md): Integrate Qwen-Audio-3.0-ASR-Flash-Streaming and Fun-ASR-Realtime on Android for real-time speech-to-text: key interfaces initialize/startDialog/stopDialog/release, connection params (apikey/url/service_mode), audio formats (pcm/wav/mp3/opus/aac/amr), sample rate, semantic punctuation, instant and precompiled hotwords, sensitive word filtering. - [Model Studio Real-time ASR iOS SDK (Qwen-Audio-3.0/Fun-ASR)](https://help.aliyun.com/en/model-studio/ios-sdk-for-fun-asr-real-time-service.md): Integrate Qwen-Audio-3.0-ASR-Flash-Streaming and Fun-ASR-Realtime speech-to-text on iOS: SDK initialization, key interfaces (nui_initialize/nui_set_params/nui_dialog_start), audio format/sample rate/hotwords/semantic punctuation parameters, and onNuiEventCallback event handling. - [Qwen-ASR-Realtime Real-Time Speech Recognition API (Streaming Transcription)](https://help.aliyun.com/en/model-studio/qwen-asr-realtime-api.md): API reference for Qwen-ASR-Realtime real-time speech recognition: streaming audio-to-text, request parameters and response fields. TTS/offline ASR in sibling audio docs. - [Real-time ASR (Qwen) WebSocket API: VAD/Manual Interaction](https://help.aliyun.com/en/model-studio/qwen-asr-realtime-interaction-process.md): WebSocket API for Qwen real-time speech recognition: wss endpoint URL and model query parameter, Authorization header (401/403 handshake errors), VAD auto-segmentation vs Manual mode, and the session.finish requirement to avoid losing transcription results. - [Qwen-ASR-Realtime client events](https://help.aliyun.com/en/model-studio/qwen-asr-realtime-client-events.md): WebSocket client-to-server events: session.update (audio format, language, VAD config), input_audio_buffer.append (stream audio), input_audio_buffer.commit (Manual mode), session.finish - [Qwen-ASR-Realtime server events](https://help.aliyun.com/en/model-studio/qwen-asr-realtime-server-events.md): WebSocket server events: session.created, speech_started/stopped (VAD), transcription.text (interim with text+stash), transcription.completed (final transcript with emotion) - [Qwen-ASR-Realtime Python SDK](https://help.aliyun.com/en/model-studio/qwen-asr-realtime-python-sdk.md): DashScope Python SDK for real-time ASR. OmniRealtimeConversation class with connect(), append_audio(), commit(), end_session(). Supports VAD and Manual modes, 27 languages, pcm/opus input - [Qwen-ASR-Realtime Java SDK](https://help.aliyun.com/en/model-studio/qwen-asr-realtime-java-sdk.md): DashScope Java SDK for real-time ASR. OmniRealtimeParam builder + OmniRealtimeConfig for VAD/Manual mode. Key methods: connect(), appendAudio(), commit(), endSession(). Requires SDK 2.22.5+ - [Model Studio – Paraformer Real-time Speech Recognition API (WebSocket Streaming)](https://help.aliyun.com/en/model-studio/paraformer-real-time-speech-recognition-api-reference.md): Paraformer real-time ASR API reference: WebSocket streaming audio input, incremental intermediate and final transcription results, model/request/response parameters and sample code for Chinese/English, multilingual and 8kHz phone audio. - [WebSocket API for Paraformer Real-time Speech Recognition (Endpoint/Headers/Event Flow)](https://help.aliyun.com/en/model-studio/websocket-for-paraformer-real-time-service.md): WebSocket access to Paraformer real-time speech recognition: wss endpoint in Beijing region only, Authorization header auth (401/403), run-task/finish-task event flow. - [Paraformer Real-time ASR Client Events](https://help.aliyun.com/en/model-studio/paraformer-client-events.md): Send run-task and finish-task commands via WebSocket to the Paraformer real-time ASR service, configuring model, format, sample_rate, language_hints and other parameters to start or end recognition. opus/speex require Ogg container, wav must use PCM encoding, amr supports AMR-NB only; paraformer-realtime-v2 accepts any sample rate while v1/8k variants are limited to 16000 Hz or 8000 Hz respectively. - [Server-sent events for real-time speech recognition (Paraformer)](https://help.aliyun.com/en/model-studio/paraformer-server-events.md): Reference for the server-sent events that the Paraformer real-time speech recognition service pushes to clients over WebSocket. This topic documents the data structure and field semantics of the four event types: task-started, result-generated, task-finished, and task-failed. - [Paraformer real-time speech recognition Python SDK](https://help.aliyun.com/en/model-studio/paraformer-real-time-speech-recognition-python-sdk.md): DashScope Python SDK for Paraformer real-time ASR. Recognition class with call() for local files and start()/send_audio_frame()/stop() for streaming. Supports microphone input via pyaudio - [Paraformer Real-Time Speech Recognition Java SDK (Model Studio)](https://help.aliyun.com/en/model-studio/paraformer-real-time-speech-recognition-java-sdk.md): Paraformer real-time ASR Java SDK (DashScope): paraformer-realtime-v2/8k-v2/v1 models for live meetings and 8kHz call audio; non-streaming & bidirectional streaming calls. - [Paraformer Real-time ASR Android SDK Integration](https://help.aliyun.com/en/model-studio/android-sdk-for-paraformer-real-time-service.md): Alibaba Cloud Model Studio Paraformer real-time speech recognition Android SDK guide: obtain API Key, download AAR/CPP packages, key interfaces (initialize/startDialog/stopDialog/release), models including paraformer-realtime-v2/v1/8k-v2, configure sample rate, language hints, hotwords, and semantic punctuation. - [Real-time Speech Recognition - Paraformer iOS SDK Setup](https://help.aliyun.com/en/model-studio/ios-sdk-for-paraformer-real-time-service.md): Paraformer real-time speech recognition iOS SDK: integrate nuisdk.framework, nui_dialog_start call flow, url/apikey parameters; temporary API Key recommended - [Non-realtime Speech Recognition API (Qwen-Audio-3.0-ASR-Flash-Filetrans/Fun-ASR)](https://help.aliyun.com/en/model-studio/fun-asr-recorded-speech-recognition-api-reference.md): Recorded-file ASR HTTP API: Qwen-Audio-3.0-ASR-Flash-Filetrans/Fun-ASR async calls (X-DashScope-Async), poll task_id; hotwords, Prompt context, speaker diarization. - [Bailian Qwen-Audio ASR Offline Speech Recognition HTTP API](https://help.aliyun.com/en/model-studio/fun-asr-recorded-speech-recognition-http-api.md): Asynchronous file transcription API for Qwen-Audio-3.0-ASR-Flash-Filetrans and Fun-ASR models. Submit a task to get a task_id, then poll for results. Supports instant hotwords, context enhancement, speaker diarization (≤2 hours), and multi-track billing. The new endpoint requires the parameters object; billing is based on detected speech duration. - [Non-realtime ASR Python SDK (Recorded Audio File Transcription)](https://help.aliyun.com/en/model-studio/funauidio-asr-recorded-speech-recognition-python-sdk.md): DashScope Transcription for recorded-audio ASR: async_call submits, wait blocks, fetch polls; model=qwen-audio-3.0-asr-flash-filetrans. Result URLs expire in 24h. - [Non-realtime Speech Recognition Java SDK (Qwen-Audio-3.0-ASR-Flash-Filetrans)](https://help.aliyun.com/en/model-studio/fun-asr-recorded-speech-recognition-java-sdk.md): Java DashScope SDK for Qwen-Audio-3.0-ASR-Flash-Filetrans file transcription: submit tasks via Transcription.asyncCall, wait or fetch results (PENDING/RUNNING/SUCCEEDED/FAILED); result URLs valid 24 hours. - [Fun-ASR audio file recognition Android SDK](https://help.aliyun.com/en/model-studio/fun-asr-recorded-speech-recognition-android-sdk.md): Android SDK (AAR format) for Fun-ASR file transcription. Sync and async modes via startFileTranscriber. EVENT_FILE_TRANS_RESULT callback. C++ support via android_libs. - [Non-Realtime Speech Recognition iOS SDK (Fun-ASR/Qwen-Audio-3.0-ASR-Flash-Filetrans)](https://help.aliyun.com/en/model-studio/fun-asr-recorded-speech-recognition-ios-sdk.md): Integrate Fun-ASR non-realtime speech recognition on iOS: add nuisdk.framework, call nui_file_trans_start (sync/async), model qwen-audio-3.0-asr-flash-filetrans. - [Model Studio-Qwen-Audio-3.0-ASR-Flash Non-real-time Speech Recognition API](https://help.aliyun.com/en/model-studio/non-real-time-speech-recognition-for-fun-asr-flash.md): Qwen-Audio-3.0-ASR-Flash/Fun-ASR-Flash non-real-time speech recognition HTTP API (SDK not supported). Supports streaming (SSE) and non-streaming transcription with context. Parameters include model, input_audio, format, and sample_rate. Available via dedicated endpoints in China (Beijing) and Singapore regions. - [Model Studio Qwen-ASR Non-real-time Speech Recognition API (OpenAI/DashScope/Async)](https://help.aliyun.com/en/model-studio/qwen-asr-api-reference.md): Qwen-ASR non-real-time speech recognition API: qwen3-asr-flash (OpenAI-compatible/DashScope sync) and qwen3-asr-flash-filetrans (DashScope async submit-poll). Supports audio URL/Base64 input, asr_options for language and ITN, streaming output, and word-level timestamps; OpenAI-compatible mode is unavailable in the US region. - [Paraformer non-real-time speech recognition API reference](https://help.aliyun.com/en/model-studio/paraformer-recorded-speech-recognition-api-reference.md) - [Paraformer audio file recognition RESTful API](https://help.aliyun.com/en/model-studio/paraformer-recorded-speech-recognition-restful-api.md): HTTP REST API for Paraformer batch file transcription without SDK. POST to /services/audio/asr/transcription with X-DashScope-Async header, then poll GET /tasks/{task_id} for results - [Paraformer audio file recognition Python SDK](https://help.aliyun.com/en/model-studio/paraformer-recorded-speech-recognition-python-sdk.md): DashScope Python SDK for Paraformer batch file transcription. Transcription class with async_call/wait/fetch methods. Returns TranscriptionResponse with task_id, subtask status, and transcription_url - [Paraformer audio file recognition Java SDK](https://help.aliyun.com/en/model-studio/paraformer-recorded-speech-recognition-java-sdk.md): DashScope Java SDK for Paraformer batch file transcription. Transcription class with asyncCall/wait/fetch. Supports up to 100 file URLs, speaker diarization, multi-track audio - [Paraformer audio file recognition Android SDK](https://help.aliyun.com/en/model-studio/paraformer-recorded-speech-recognition-android-sdk.md): Native Android SDK (AAR) for Paraformer batch file transcription. Sync and async modes via startFileTranscriber. WebSocket connection to dashscope.aliyuncs.com - [Paraformer audio file recognition iOS SDK](https://help.aliyun.com/en/model-studio/paraformer-recorded-speech-recognition-ios-sdk.md): Native iOS SDK (nuisdk.framework) for Paraformer batch file transcription. Integrate via Xcode with Embed & Sign. Supports sync and async modes via nui_file_trans_start and nui_file_trans_query - [Paraformer file transcription best practices](https://help.aliyun.com/en/model-studio/paraformer-best-practices.md): Pre-process video files with ffmpeg to extract audio (16kHz opus) before submitting to Paraformer file transcription API. Reduces file size and improves throughput. - [Custom Hot Words for Speech Recognition (Boost Accuracy)](https://help.aliyun.com/en/model-studio/custom-hot-words.md): Configure hot word lists to improve speech recognition accuracy for proper nouns, product names, and industry terms in real-time and non-real-time ASR scenarios. - [Model Studio - Manage Vocabulary Lists via HTTP API](https://help.aliyun.com/en/model-studio/vocabulary-http-api.md): Create, query, update, and delete custom vocabulary lists for ASR via the speech-biasing HTTP API. Supports Paraformer and Fun-ASR models with weight range 1-5. Not available in Singapore sub-workspaces. - [Model Studio - VocabularyService Python SDK for Hotword Management](https://help.aliyun.com/en/model-studio/vocabulary-python-sdk.md): Manage ASR hotword lists via dashscope.audio.asr.VocabularyService with create_vocabulary, list_vocabularies, update_vocabulary, and delete_vocabulary methods. Hotword feature is unavailable in Singapore sub-workspaces; SDK defaults to Beijing endpoint and requires base_http_api_url switch for other regions - [Java SDK Reference - Hotword Vocabulary Management](https://help.aliyun.com/en/model-studio/vocabulary-java-sdk.md): Manage ASR hotword vocabularies via VocabularyService class with createVocabulary, listVocabulary, queryVocabulary, updateVocabulary, and deleteVocabulary operations. Hotword feature is unavailable in Singapore sub-workspaces; use Beijing or Singapore dedicated endpoints with region-specific API Keys. - [Speech Synthesis APIs (Qwen-TTS/CosyVoice/Sambert/MiniMax/Voice Cloning)](https://help.aliyun.com/en/model-studio/speech-synthesis-api-reference.md): Speech synthesis API hub: realtime TTS (Qwen-Audio-TTS/CosyVoice, Qwen-TTS-Realtime, Sambert), non-realtime TTS (Qwen-TTS, MiniMax), voice cloning and voice design. Start here for text-to-speech. - [Real-time Speech Synthesis (Qwen-Audio-TTS/CosyVoice Streaming TTS)](https://help.aliyun.com/en/model-studio/cosyvoice-large-model-for-speech-synthesis.md): Real-time TTS on Model Studio: stream low-latency speech with Qwen-Audio-TTS and CosyVoice models; voice timbre lists, voice cloning, voice design, SSML/LaTeX markup control. - [CosyVoice WebSocket API](https://help.aliyun.com/en/model-studio/cosyvoice-websocket-api.md): WebSocket protocol for CosyVoice TTS in Go/PHP/Node.js (non-Java/Python). Full-duplex via wss://dashscope.aliyuncs.com/api-ws/v1/inference. CosyVoice does NOT support HTTP REST - [Speech Synthesis - Client Events run-task/continue-task/finish-task](https://help.aliyun.com/en/model-studio/cosyvoice-client-events.md): Client-side WebSocket commands for CosyVoice/Qwen-Audio-TTS streaming speech synthesis. run-task starts a task with model, voice, sample_rate, format, volume, rate, pitch, bit_rate, enable_ssml, and word_timestamp_enabled parameters; continue-task sends text to synthesize; finish-task ends the task. cosyvoice-v1 does not support opus format or bit_rate parameter. - [Qwen-Audio-TTS/CosyVoice server-side events](https://help.aliyun.com/en/model-studio/cosyvoice-server-events.md): User guide: For model introduction and selection recommendations, see Speech synthesis.User guide:Speech synthesistask-startedAfter the client sends t... - [Bailian TTS Java SDK (CosyVoice/Qwen-Audio-TTS SpeechSynthesizer)](https://help.aliyun.com/en/model-studio/cosyvoice-java-sdk.md): DashScope Java SDK for CosyVoice/Qwen-Audio-TTS speech synthesis: SpeechSynthesizer call/streamingCall, 20000-char text limit, WorkspaceId WebSocket endpoints per region. - [Model Studio TTS Python SDK (Qwen-Audio-TTS/CosyVoice Streaming & Non-streaming)](https://help.aliyun.com/en/model-studio/cosyvoice-python-sdk.md): Call Qwen-Audio-TTS and CosyVoice real-time TTS via DashScope Python SDK: SpeechSynthesizer constructor (model/voice/format/volume/speech_rate), non-streaming call(), bidirectional streaming_call()/streaming_complete(), ResultCallback callbacks, word-level timestamps and AIGC tags. Single text ≤20,000 chars, cumulative ≤200,000 chars. - [Qwen-Audio-TTS/CosyVoice Speech Synthesis Android SDK (NativeNui)](https://help.aliyun.com/en/model-studio/cosyvoice-android-sdk.md): Android speech synthesis SDK: NativeNui singleton callbacks; one-shot (SSML) and streaming input modes; startStreamInputTts/sendStreamInputTts APIs; temp API Key recommended. - [Speech Synthesis iOS SDK (Qwen-Audio-TTS/CosyVoice via NeoNui)](https://help.aliyun.com/en/model-studio/cosyvoice-ios-sdk.md): iOS SDK for Bailian TTS (Qwen-Audio-TTS/CosyVoice): NeoNui singleton, one-shot input with SSML or streaming input, startStreamInputTts/sendStreamInputTts APIs, ticket auth with apikey. - [Real-time speech synthesis API reference (Qwen-TTS-Realtime)](https://help.aliyun.com/en/model-studio/qwen-tts-realtime-api-reference.md) - [Qwen-TTS-Realtime interaction flow](https://help.aliyun.com/en/model-studio/interactive-process-of-qwen-tts-realtime-synthesis.md): WebSocket event flow for qwen-tts real-time synthesis. Two modes: ServerCommit (auto-segmented) and Commit (manual). Key events: input_text_buffer.append/commit, response.audio.delta, session.finish. - [Qwen-TTS-Realtime client events](https://help.aliyun.com/en/model-studio/qwen-tts-realtime-client-events.md): WebSocket client events for real-time TTS: session.update (voice, mode, language, format, speech_rate), input_text_buffer.append/commit/clear, session.finish. Supports server_commit and commit modes - [Qwen-TTS-Realtime server events](https://help.aliyun.com/en/model-studio/qwen-tts-realtime-server-events.md): WebSocket server events for real-time TTS: session.created/updated, response.audio.delta (Base64 audio chunks), response.done (usage/billing), session.finished. Error event with code and message - [Qwen-TTS-Realtime Python SDK](https://help.aliyun.com/en/model-studio/qwen-tts-realtime-python-sdk.md): DashScope Python SDK for real-time TTS. QwenTtsRealtime class with connect(), update_session(), append_text(), commit(), finish(). Supports pcm/wav/mp3/opus output, instruction control, voice cloning - [Qwen-TTS-Realtime Java SDK](https://help.aliyun.com/en/model-studio/qwen-tts-realtime-java-sdk.md): DashScope Java SDK for real-time TTS. QwenTtsRealtime with appendText(), commit(), finish(). Config: voice, format (pcm/wav/mp3/opus), sample rate, speech rate. SDK 2.22.7+ - [Sambert Real-time Speech Synthesis (TTS)](https://help.aliyun.com/en/model-studio/sambert-speech-synthesis.md): Real-time text-to-speech with Sambert models on Model Studio: streaming audio output, multiple voice options, suited for narration, announcements, and voice interaction. - [Sambert TTS WebSocket API (Endpoint, Headers, Interaction Flow)](https://help.aliyun.com/en/model-studio/sambert-websocket-api.md): Sambert TTS WebSocket API: wss endpoint (Beijing only), Bearer auth (401/403), run-task/task-finished flow; all text in one run-task, no streaming input. - [Sambert TTS WebSocket Client Events (run-task)](https://help.aliyun.com/en/model-studio/sambert-client-events.md): Sambert TTS WebSocket client event run-task: Beijing region only; text sent once via input.text, no streaming; model/format/sample_rate and timestamp params. - [Sambert server-sent events](https://help.aliyun.com/en/model-studio/sambert-server-events.md): User guide: For model introduction and selection recommendations, see Speech synthesis.User guide:Speech synthesisImportantModel Studio has released a... - [Sambert TTS Java SDK (Non-streaming & Streaming Calls)](https://help.aliyun.com/en/model-studio/sambert-java-sdk.md): Sambert TTS Java SDK: SpeechSynthesizer non-streaming and streaming calls, sambert-zhichu-v1 voice, sampleRate/format params, Beijing region only. - [Sambert TTS Python SDK (Non-streaming/Streaming)](https://help.aliyun.com/en/model-studio/sambert-python-sdk.md): Sambert TTS Python SDK: SpeechSynthesizer.call for non-streaming and ResultCallback one-way streaming synthesis, text up to 10,000 chars, returns audio data and timestamps. - [Sambert TTS Android SDK](https://help.aliyun.com/en/model-studio/sambert-android-sdk.md): Native Android SDK (AAR) for Sambert speech synthesis via WebSocket. Key interfaces: tts_initialize, startTts, onTtsDataCallback. Supports pcm/wav/mp3 output formats - [Sambert TTS iOS SDK](https://help.aliyun.com/en/model-studio/sambert-ios-sdk.md): iOS framework (nuisdk.framework) for Sambert speech synthesis via WebSocket. Key methods: nui_tts_initialize, nui_tts_play. Supports pcm/wav/mp3 output, word-level timestamps - [Non-Realtime TTS API (Qwen-Audio-TTS/CosyVoice Text-to-Speech)](https://help.aliyun.com/en/model-studio/non-realtime-cosyvoice-api.md): Model Studio text-to-speech API: call Qwen-Audio-TTS/CosyVoice models to synthesize text into audio files. Distinct from real-time streaming TTS; suited to long text and batch narration. Voice list in the speech synthesis section. - [Model Studio - Non-real-time TTS HTTP API Reference](https://help.aliyun.com/en/model-studio/cosyvoice-tts-http-api.md): HTTP API for Qwen-Audio-TTS/CosyVoice non-real-time speech synthesis, supporting streaming and non-streaming calls. Available only in China (Beijing) region; use workspace-specific endpoint {WorkspaceId}.cn-beijing.maas.aliyuncs.com. Required params: model, input.text, voice; optional: format, sample_rate, volume, rate - [Non-Realtime TTS Java SDK (Qwen-Audio-TTS/CosyVoice)](https://help.aliyun.com/en/model-studio/cosyvoice-tts-java-sdk.md): Java SDK for non-realtime TTS with Qwen-Audio-TTS/CosyVoice: HttpSpeechSynthesizer.call returns audio bytes or URL, streamCall streams audio via callback, Beijing region only. - [Model Studio Non-realtime TTS Python SDK (CosyVoice/Qwen-Audio-TTS)](https://help.aliyun.com/en/model-studio/cosyvoice-tts-python-sdk.md): Non-realtime TTS via HttpSpeechSynthesizer.call(): streaming/non-streaming, SSML, voice cloning. Models qwen-audio-3.0-tts, cosyvoice-v3.5. Beijing only. - [Qwen-TTS speech synthesis API](https://help.aliyun.com/en/model-studio/qwen-tts-api.md): Non-realtime TTS via DashScope API. Models: qwen3-tts-flash, qwen3-tts-instruct-flash. Params: text, voice, language_type, instructions, stream. Returns audio URL or Base64 - [MiniMax Non-realtime Text-to-Speech (speech-2.8-hd/turbo, Beijing only)](https://help.aliyun.com/en/model-studio/minimax-speech-synthesis.md): MiniMax speech-2.8-hd/turbo text-to-speech on Model Studio (Beijing region only), with voice cloning; HD ¥3.5 per 10k characters, Turbo ¥2. - [MiniMax Synchronous Speech Synthesis API](https://help.aliyun.com/en/model-studio/minimax-synchronous-speech-synthesis-api.md): Text-to-speech with MiniMax/speech-2.8-hd/turbo models. Supports streaming (SSE), emotion control, voice blending (timbre_weights), pronunciation dicts, and configurable sample rate/bitrate/format. - [Voice Cloning on Model Studio: Clone a Custom Voice from 10-20s Audio](https://help.aliyun.com/en/model-studio/sound-reengraving.md): Voice cloning: upload a 10-20 second audio sample to create a custom voice with no model training, usable with CosyVoice/Qwen-TTS TTS models; voice creation costs CNY 0.01 each, with 1,000 free tries within 90 days of activation. - [Bailian - Voice Clone HTTP API Reference](https://help.aliyun.com/en/model-studio/voice-clone-design-http-api.md): Voice clone APIs for Qwen-Audio-TTS, CosyVoice, Qwen-TTS, and MiniMax. Supports create_voice/create, list_voice/list, query, update, and delete operations. China (Beijing) and Singapore regions use WorkspaceId-based dedicated endpoints; MiniMax only supports dashscope.aliyuncs.com. Audio URLs must be publicly accessible; Qwen-TTS also accepts Base64 Data URL. - [Voice Clone Java SDK Reference](https://help.aliyun.com/en/model-studio/voice-clone-java-sdk.md): Bailian VoiceEnrollmentService Java SDK for managing Qwen-Audio-TTS/CosyVoice cloned voice lifecycle. Includes createVoice, listVoice, queryVoice, updateVoice, deleteVoice methods with targetModel/prefix/url parameters, supporting China (Beijing) and Singapore regions - [Voice Clone - Python SDK Reference](https://help.aliyun.com/en/model-studio/voice-clone-python-sdk.md): VoiceEnrollmentService manages Qwen-Audio-TTS/CosyVoice cloned voice lifecycle with create_voice, query_voice, update_voice, and delete_voice methods. Supports target_model, prefix, url, language_hints, max_prompt_audio_length (3-30s), and enable_preprocess parameters. Requires workspace-specific endpoint and region-matched API Key. - [Model Studio Voice Design API Create Query Delete Custom Voices](https://help.aliyun.com/en/model-studio/voice-design-api-references.md): HTTP API to manage custom voices for Qwen-Audio-TTS/CosyVoice and Qwen, covering create_voice, list_voice, query_voice, and delete_voice actions with voice_prompt, preview_text, and target_model parameters; returns voice_id/voice and preview audio. Use workspace-specific endpoints in cn-beijing and ap-southeast-1 - [Music Generation (Fun-Music API: Lyrics to Song)](https://help.aliyun.com/en/model-studio/music-generation-references.md): Hub for Fun-Music (fun-music-v1/fun-music-preview) API: input lyrics or creative prompts to generate male/female vocal songs in Chinese or English; invite-only, Beijing region only. - [Fun-Music music generation API](https://help.aliyun.com/en/model-studio/fun-music-api.md): REST API for fun-music-v1. Generate songs from text prompts or lyrics, male/female vocals. Streaming (SSE) and non-streaming output, mp3/wav. Invitational preview, Beijing only. - [Speech Translation APIs (Realtime Long/Short Audio, Video)](https://help.aliyun.com/en/model-studio/speech-translation-api-reference.md): Hub for Model Studio speech translation APIs: realtime long-audio and sentence translation (Gummy), realtime AV translation (Qwen-Livetranslate-Realtime), qwen3-livetranslate-flash. For live meetings and streaming subtitles. - [Real-time audio and video translation - Qwen-Livetranslate-Realtime](https://help.aliyun.com/en/model-studio/live-translator-api.md) - [Live Translation - Client Events](https://help.aliyun.com/en/model-studio/live-translator-client-events.md): WebSocket client events for the qwen3.5-livetranslate-flash-realtime API, including session.update to configure language, voice, and VAD; input_audio_buffer.append/commit/clear to manage the audio buffer; input_image_buffer.append to add images; and session.finish to end the session. - [Live Translator Server Events (qwen3.5-livetranslate-flash-realtime)](https://help.aliyun.com/en/model-studio/live-translator-server-events.md): Server-side events for qwen3.5-livetranslate-flash-realtime API: session.created/updated/finished, response.created/done, streaming response.text/audio, input_audio_buffer VAD detection and commit, conversation.item ASR and translation results. Covers error codes, modalities, voice, translation.language, corpus.phrases hotwords, turn_detection server_vad, and enable_voice_clone parameters. - [Qwen-LiveTranslate Python SDK](https://help.aliyun.com/en/model-studio/qwen-livetranslate-python-sdk.md): DashScope Python SDK for real-time speech translation. OmniRealtimeConversation with TranslationParams (target language, custom terminology). Outputs text and/or synthesized speech audio via WebSocket - [Qwen-LiveTranslate Real-time Audio/Video Translation Java SDK](https://help.aliyun.com/en/model-studio/qwen-livetranslate-java-sdk.md): DashScope Java SDK for Qwen-LiveTranslate real-time audio/video translation: OmniRealtimeParam sets model/url, OmniRealtimeConfig sets voice and hot words. - [Qwen3 live translation API](https://help.aliyun.com/en/model-studio/qwen3-livetranslate-flash-api.md): Real-time audio/video translation via OpenAI-compatible streaming endpoint. Model: qwen3-livetranslate-flash. Set translation_options with source_lang and target_lang in extra_body - [Model Studio Voice Conversation API References (Qwen-Audio Realtime)](https://help.aliyun.com/en/model-studio/voice-conversation-api-references.md): API hub for voice conversation models: Qwen-Audio realtime voice chat (qwen-audio-3.0-realtime-plus/flash, WebSocket/AOQ access), model selection for voice assistants, customer-service dialogue and semantic VAD. - [Qwen-Audio Real-time Voice Conversation API Reference (WebSocket/Events)](https://help.aliyun.com/en/model-studio/real-time-voice-conversation-api-references.md): Qwen-Audio Realtime API hub: WebSocket API reference, server_vad/smart_turn/push-to-talk interaction modes, client events and server events data structures and fields. - [Qwen-Audio Realtime WebSocket API for Real-Time Voice Conversation](https://help.aliyun.com/en/model-studio/fun-audiochat-realtime-websocket-api.md): Call the Qwen-Audio Realtime model over WebSocket to enable real-time voice conversations with server_vad, smart_turn, and push-to-turn modes, VAD-based speech detection, streaming audio/text output, and function calling. Use the wss:// protocol with API Key authentication in headers; dedicated workspace endpoints in China (Beijing) and Singapore regions are recommended for higher stability. - [Qwen-Audio Realtime API Client Events](https://help.aliyun.com/en/model-studio/fun-audiochat-client-events.md): Reference for Qwen-Audio Realtime WebSocket client events including session.update, input_audio_buffer.append/commit/clear, conversation.item.create/retrieve/delete, and response.create/cancel with parameter details and usage constraints - [Qwen-Audio Realtime Server Events Reference (session/response/error)](https://help.aliyun.com/en/model-studio/qwen-audio-realtime-server-events.md): Qwen-Audio Realtime API server-side events: error/session.created/session.updated, input_audio_buffer speech detection and commit, conversation.item CRUD and ASR transcription, response text/audio/transcript delta and done, streaming Function Calling arguments, voiceprint registration status. Includes event_id/type common fields and smart_turn ambient audio passthrough. - [Omni Realtime API (Qwen-Omni-Realtime Audio/Video Chat)](https://help.aliyun.com/en/model-studio/omni-realtime-api.md): API hub for Qwen-Omni-Realtime realtime audio/video chat models: streaming audio and image input, realtime text and audio output; Beijing and Singapore regions. - [Qwen-Omni-Realtime client events](https://help.aliyun.com/en/model-studio/client-events.md): WebSocket client events: session.update, response.create/cancel, input_audio_buffer append/commit/clear, input_image_buffer.append. Supports VAD, function calling, web search. - [Qwen-Omni-Realtime server events](https://help.aliyun.com/en/model-studio/server-events.md): Server-sent events for real-time multimodal API: session.created, response.done, response.audio.delta, function_call. Includes VAD configuration and token usage details - [Qwen-Omni real-time Python SDK](https://help.aliyun.com/en/model-studio/omni-realtime-python-sdk.md): OmniRealtimeConversation class for Qwen-Omni WebSocket sessions. Audio/video streaming, VAD (server_vad/semantic_vad), voice interruption, tool calling, web search. Requires dashscope >= 1.25.17. - [Qwen-Omni real-time Java SDK](https://help.aliyun.com/en/model-studio/omni-realtime-java-sdk.md): OmniRealtimeConversation and OmniRealtimeConfig for Qwen-Omni WebSocket sessions in Java. Audio/video streaming, VAD, voice interruption, tool calling. Requires Java SDK >= 2.22.15. - [Qwen-Omni real-time interaction flow](https://help.aliyun.com/en/model-studio/omni-realtime-interaction-process.md): WebSocket event diagrams for VAD mode (auto speech detection) and Manual mode (push-to-talk). Covers audio buffer append/commit, response.create, voice interruption, and tool calling flows. - [Qwen-Omni voice cloning API](https://help.aliyun.com/en/model-studio/qwen-omni-voice-cloning.md): Clone a voice from 10-20s audio without training. Two-step flow: create voice via qwen-voice-enrollment, then use with qwen3.5-omni-plus/flash-realtime. Supports WAV/MP3/M4A, 30 languages - [Realtime API: Real-time Voice & Multimodal (Overview/Quick Start/AOQ SDK)](https://help.aliyun.com/en/model-studio/realtime-api-user-guide.md): Realtime API hub: overview, quick start, best practices, and AOQ client SDK. Build low-latency real-time voice/multimodal conversations with models, distinct from text-generation and standard multimodal APIs. - [Realtime API Overview](https://help.aliyun.com/en/model-studio/realtime-api-overview.md): Alibaba Cloud Model Studio Realtime API supports AOQ, WebRTC, and WebSocket protocols for realtime omni-modal, speech translation, ASR, and TTS models including qwen3.5-omni-plus-realtime, qwen3.5-livetranslate-flash-realtime, and CosyVoice. AOQ offers built-in echo cancellation and noise suppression with strong weak-network resilience; WebSocket enables quick integration via DashScope SDK. - [Model Studio Realtime API Quick Start (SDK/Token Auth/Connect Models)](https://help.aliyun.com/en/model-studio/realtime-api-quick-start-guide.md): Realtime API entry: SDK download, Token authentication, connecting models and apps. Start here to run your first real-time voice conversation on Model Studio. - [Bailian AOQ SDK Download (Android/iOS/HarmonyOS/Windows/macOS)](https://help.aliyun.com/en/model-studio/realtime-sdk-download.md): Download AOQ SDK v1.2.0 for Android, iOS, HarmonyOS, Windows, macOS, Electron and Linux, with release notes; Opus audio codec requires the separate libPluginOpus plugin. - [Realtime API-Token Authentication](https://help.aliyun.com/en/model-studio/realtime-token-authentication.md): Alibaba Cloud Model Studio Realtime API authenticates connections via HTTP Header Authorization: Bearer , covering AOQ, WebRTC, and WebSocket protocols. AOQ uses server-side proxy mode to avoid exposing the API Key on clients; WebRTC authenticates during SDP exchange and requires whitelist approval; WebSocket carries the API Key directly during handshake. - [Realtime API-Connect Models and Applications](https://help.aliyun.com/en/model-studio/realtime-connect-model.md): Connect to Bailian Realtime API models or applications via AOQ, WebRTC, or WebSocket protocols. AOQ uses QUIC for mobile apps with weak-network resilience; WebRTC supports native browser JS access. Includes connection flows, sequence diagrams, and code samples. - [Model Studio AOQ Client SDK Realtime Multimodal Integration](https://help.aliyun.com/en/model-studio/realtime-api-aoq-api.md): AOQ Client SDK integrates with Model Studio Realtime API, offering SDK overview and function reference for client-side integration of realtime voice dialogue models such as qwen3.5-omni-plus-realtime - [Model Studio-AOQ SDK for Realtime Multimodal Development](https://help.aliyun.com/en/model-studio/realtime-api-aoq-sdk-desc.md): AOQ SDK is built on Alibaba Cloud Realtime API with declarative design and automatic exception recovery, helping developers quickly build real-time audio/video interactive applications. It reports progress via callbacks and only requires intervention for unrecoverable errors such as physical limitations or invalid tokens. - [AOQ Client SDK Android API reference](https://help.aliyun.com/en/model-studio/aoq-android-sdk-reference.md): The AOQ Client SDK for Android provides a full set of real-time audio/video communication capabilities, including engine lifecycle management, audio/video capture and playback, codec configuration, external audio stream injection, audio file mixing, real-time messaging, and audio/video frame callbacks. This document is the complete Java API reference for the Android platform. - [Model Studio AOQ Client SDK for iOS API Reference (Realtime Audio/Video)](https://help.aliyun.com/en/model-studio/aoq-ios-sdk-reference.md): AOQ iOS SDK API list: engine lifecycle (createEngine/connect), audio/video capture and codec config, audio file playback, external PCM audio streams, video frame push, realtime data messages, frame callbacks and delegate protocols for qwen realtime voice/video conversations. - [Realtime API-HarmonyOS SDK (AOQ Engine, ArkTS Reference)](https://help.aliyun.com/en/model-studio/aoq-harmony-sdk-reference.md): AOQ HarmonyOS (OHOS) SDK reference in ArkTS: createEngine/connect, mic/camera capture, audio file playback, external streams, sendDataMsg, audio frame callbacks. - [Model Studio AOQ Windows SDK (C++ Real-Time Audio/Video API Reference)](https://help.aliyun.com/en/model-studio/aoq-windows-sdk-reference.md): AOQ Windows SDK C++ reference: AoqClientEngine singleton, createEngine/connect lifecycle, audio/video capture & render, file playback, external stream push, event callbacks. - [AOQ Client SDK macOS API reference](https://help.aliyun.com/en/model-studio/aoq-macos-sdk-reference.md): This topic describes the Objective-C APIs, callbacks, and data types of AOQ Client SDK for macOS. - [Bailian AOQ Electron SDK API Reference (Real-time Audio/Video Engine)](https://help.aliyun.com/en/model-studio/aoq-electron-sdk-reference.md): aoq-electron-sdk (npm) TypeScript APIs: engine lifecycle, audio/video capture & codec, external audio streams, real-time messaging. macOS x64/arm64, Windows x64, Node.js >= 16. - [AOQ Client SDK for Linux – Python API Reference (Engine/Audio/Video/Messaging)](https://help.aliyun.com/en/model-studio/aoq-linux-python-sdk-reference.md): AOQ Client SDK (Linux Python): create_engine/connect, external audio/video frame push, send_data_msg realtime messaging. Callbacks on native threads; ensure thread safety. - [AOQ Client SDK Linux C++ API Reference](https://help.aliyun.com/en/model-studio/aoq-linux-cpp-sdk-reference.md): AOQ Linux C++ SDK: createEngine/connect lifecycle, audio/video capture, external frame push, file playback, sendDataMsg. Audio capture no-op; video external-only. - [Realtime API - AOQ SDK Function Reference](https://help.aliyun.com/en/model-studio/realtime-api-aoq-sdk-function.md): Function reference for the AOQ SDK with the Realtime API: function parameters for connecting to real-time multimodal models like qwen3.5-omni-plus-realtime to build real-time voice conversation, speech recognition, and TTS scenarios - [Connection state management](https://help.aliyun.com/en/model-studio/aoq-connection-management.md): Describes the connection state machine of the AOQ Client SDK and the corresponding API calls. - [Media stream sending control](https://help.aliyun.com/en/model-studio/aoq-media-stream-control.md): enableSendMediaStream controls whether the client sends audio or video media streams to the AI service, giving you precise control over when media transmission starts in AOQ protocol scenarios. - [AOQ Client SDK Audio Features (Capture/Playback/Speaker/Mixing)](https://help.aliyun.com/en/model-studio/aoq-audio-features.md): AOQ Client SDK audio: capture, playback (pause/interrupt, fadeMs), speaker switch, mixing, external stream injection, frame callbacks — Android/iOS/Ohos API reference. - [Custom audio playback](https://help.aliyun.com/en/model-studio/aoq-custom-audio-playback.md): The AOQ Client SDK supports custom audio playback. Through the audio frame callback mechanism, decoded PCM data is delivered to the application layer, letting you implement your own audio rendering logic. - [AOQ Client SDK Custom Audio Capture (External Stream / PCM Push)](https://help.aliyun.com/en/model-studio/aoq-custom-audio-capture.md): Custom audio capture in AOQ SDK: addAudioExternalStream adds external streams, pushAudioExternalStreamData pushes PCM mixed with mic audio; enable3A for echo cancellation. - [AOQ Client SDK Video Features (Capture/Rendering/Encoding/External Input)](https://help.aliyun.com/en/model-studio/aoq-video-features.md): AOQ Client SDK video (Android/iOS/Ohos): capture 1280x720@15fps, external frame input, render modes Auto/Stretch/Fill/Crop, encoding and frame callbacks. - [AOQ Client SDK Custom Video Input (Raw Frame / Encoded Frame Modes)](https://help.aliyun.com/en/model-studio/aoq-custom-video-input.md): AOQ Client SDK custom video input: raw frames (BGRA/I420/NV12/NV21) pushed via pushExternalVideoCapturedFrame for SDK encoding; JPEG frames pushed directly. Modes are exclusive. - [Model Studio Vector & Rerank APIs (Embedding/Rerank)](https://help.aliyun.com/en/model-studio/vector-and-sort.md): Embedding and rerank model APIs: text vectorization with text-embedding-v4, multimodal qwen3-vl-embedding, reranking with qwen3-rerank/gte-rerank-v2 for RAG retrieval. - [Text Embedding Models (text-embedding-v2/v3/v4, qwen3.7-text-embedding)](https://help.aliyun.com/en/model-studio/general-text-vector.md): text-embedding-v2/v3/v4 and qwen3.7-text-embedding: text-to-vector models for RAG retrieval, clustering, classification; 50+ languages, 64–2048 dims, from ¥0.0005/1K tokens with free quota. - [Text embedding synchronous API](https://help.aliyun.com/en/model-studio/text-embedding-synchronous-api.md): Convert text to vectors via OpenAI-compatible or DashScope API. Models: text-embedding-v1 to v4 (Qwen3-Embedding). Configurable dimensions (64-2048), supports dense and sparse output - [Text embedding batch API](https://help.aliyun.com/en/model-studio/text-embedding-batch-api.md): Async batch embedding for up to 100K lines per file (200MB max). Models: text-embedding-async-v1/v2. Submit via HTTP or DashScope SDK, poll task_id for results - [Multimodal embedding](https://help.aliyun.com/en/model-studio/multimodal-vector.md) - [Multimodal embeddings API](https://help.aliyun.com/en/model-studio/multimodal-embedding-api-reference.md): Embed text, images, and videos in a shared semantic space. Models: qwen3-vl-embedding (2560d), tongyi-embedding-vision-plus/flash. Independent and fused (enable_fusion) embedding modes. - [Rerank Models on Model Studio (qwen3-rerank/gte-rerank-v2)](https://help.aliyun.com/en/model-studio/rerank-model.md): Rerank model overview: re-score retrieved documents to boost precision in RAG and semantic search; covers qwen3-rerank, qwen3-vl-rerank, gte-rerank-v2 with HTTP/SDK call entries. - [Rerank Models API (qwen3-rerank / qwen3-vl-rerank)](https://help.aliyun.com/en/model-studio/text-rerank-api.md): Rerank recalled documents for RAG and semantic search: qwen3-rerank (text, up to 500 docs), qwen3-vl-rerank (text/image/video), gte-rerank-v2; instruct parameter for task types; gte-rerank retires 2026-05-30, use qwen3-rerank instead. - [More Models](https://help.aliyun.com/en/model-studio/more-models.md) - [Tongyi Farui legal LLM API](https://help.aliyun.com/en/model-studio/tongyi-farui-api.md): Legal-domain LLM (farui-plus, 12K context). Features: legal Q&A, case analysis, document generation, contract review, similar case retrieval. DashScope SDK, single/multi-turn - [Intent recognition (tongyi-intent-detect-v3)](https://help.aliyun.com/en/model-studio/intent-detect-capability.md): Qwen model for parsing user intents and selecting tools in milliseconds. 8K context. Set system message with "Response in INTENT_MODE" for combined intent and function calling output. - [Qwen-MT text translation API](https://help.aliyun.com/en/model-studio/qwen-mt-api.md): Text translation via OpenAI-compatible or DashScope API. Models: qwen-mt-plus/flash/lite/turbo. Features: auto language detection, term intervention, translation memory, domain prompts - [Qwen-Deep-Research API](https://help.aliyun.com/en/model-studio/qwen-deep-research-api.md): Two-step DashScope API: Step 1 gets a clarifying question, Step 2 performs deep research with web search. Python SDK only. Output formats: model_detailed_report (~6000 tokens) or model_summary_report - [Qwen-OCR API reference](https://help.aliyun.com/en/model-studio/qwen-vl-ocr-api-reference.md): Extract text and structured data from images via OpenAI-compatible or DashScope API. Model: qwen-vl-ocr. Params: min_pixels, max_pixels for resolution control. Supports streaming - [GUI-Plus API reference](https://help.aliyun.com/en/model-studio/gui-plus-interface-interaction-model.md): API for GUI-Plus interface interaction model. OpenAI-compatible and DashScope endpoints. Enables computer_use tool for mouse/keyboard automation via screenshots at 1000x1000 resolution. - [Toolkit/Framework](https://help.aliyun.com/en/model-studio/toolkits-and-frameworks.md) - [OpenAI-compatible Chat API](https://help.aliyun.com/en/model-studio/compatibility-of-openai-with-dashscope.md): Migrate OpenAI Chat Completions code: set base_url to dashscope.aliyuncs.com/compatible-mode/v1, model to qwen-plus/max/turbo. Supports commercial and open-source Qwen models. - [Model Studio OpenAI-Compatible Responses API (Multi-turn/Built-in Tools/Reasoning)](https://help.aliyun.com/en/model-studio/compatibility-with-openai-responses-api.md): Qwen OpenAI-compatible Responses API: auto-link context via previous_response_id, control reasoning depth with reasoning.effort, and use built-in web_search/code_interpreter/web_extractor tools. Supports qwen3-max/qwen3.8-max/DeepSeek-V4 models, with Session cache and Chat Completions migration guide. - [Completions API](https://help.aliyun.com/en/model-studio/completions.md): Text/code completion for Qwen Coder models (qwen2.5-coder-7b/14b/32b). Prefix-only and fill-in-middle (FIM) with fim_prefix/fim_suffix/fim_middle tokens. Beijing region only - [OpenAI Vision API compatibility](https://help.aliyun.com/en/model-studio/qwen-vl-compatible-with-openai.md): Migrate OpenAI vision apps by changing base_url, api_key, and model. Supported models: Qwen3-VL, QVQ, Qwen-OCR series. Regions: Beijing, Singapore, Virginia - [OpenAI compatible file interface](https://help.aliyun.com/en/model-studio/openai-file-interface.md): Upload files for Qwen-Long document Q&A (purpose=file-extract) or batch tasks (purpose=batch). Max 10,000 files, 100 GB total, 150 MB per file. Supports txt, docx, pdf, xlsx, images, and more. - [OpenAI-Compatible Batch File Inference API](https://help.aliyun.com/en/model-studio/batch-interfaces-compatible-with-openai.md): Alibaba Cloud Model Studio OpenAI-compatible Batch File API for asynchronous bulk requests via JSONL files at 50% of real-time cost. Supports qwen3.8-max, qwen3.7-plus, deepseek-r1 and other text/multimodal models with up to 256K token context, suitable for data analysis and model evaluation - [OpenAI compatible Batch Chat API](https://help.aliyun.com/en/model-studio/openai-compatible-batch-chat.md): Low-cost batch inference for data annotation and content generation. base_url: batch.dashscope.aliyuncs.com. Queues requests with up to 3600s timeout. Supports qwen3.6-plus, qwen-plus, deepseek-v3.2. - [OpenAI-compatible Embedding API](https://help.aliyun.com/en/model-studio/embedding-interfaces-compatible-with-openai.md): OpenAI-compatible text embedding endpoint. Models: text-embedding-v4 (2048d, 100+ languages), v3, v2, v1. Swap base_url/api_key/model to migrate from OpenAI. Configurable dimensions - [OpenAI compatible Conversations API](https://help.aliyun.com/en/model-studio/openai-compatible-conversations.md): Server-side conversation state management via /compatible-mode/v1/conversations. CRUD for sessions and message items. Use with Responses API to auto-inject history across devices. - [LangChain integration](https://help.aliyun.com/en/model-studio/use-bailian-in-langchain.md): Integrate Model Studio into LangChain via langchain_openai (ChatOpenAI) or langchain_community (ChatTongyi). Set base_url to DashScope endpoint. Chat, embedding, tool calling examples - [Model Production](https://help.aliyun.com/en/model-studio/model-production.md) - [Model tuning](https://help.aliyun.com/en/model-studio/fine-tuning-jobs-api.md): Create custom models by fine-tuning. - [Text Generation APIs (OpenAI/DashScope/Anthropic Compatible)](https://help.aliyun.com/en/model-studio/model-fine-tuning-text-generation-api.md): Alibaba Cloud Model Studio text generation access: OpenAI-compatible Chat Completions, Responses (built-in web search and code interpreter), Anthropic Messages (thinking and tool calling), and native DashScope API. - [Model Studio-CreateFineTuningJob Create Fine-Tuning Job](https://help.aliyun.com/en/model-studio/create-fine-tuning-job-api.md): POST /api/v1/fine-tunes creates a text-generation model fine-tuning job. Supports file_id upload or OSS-mounted datasets. Available only in China (Beijing) region; requires DASHSCOPE_API_KEY. training_type options: sft, efficient_sft, cpt, dpo_full, dpo_lora. hyper_parameters n_epochs, batch_size, and max_length are required and affect billing. Returns job_id for status tracking. - [Model Studio-Custom Models Import API Create Query Delete Import Jobs](https://help.aliyun.com/en/model-studio/custom-models-api.md): Import full-parameter or LoRA fine-tuned models from OSS into Model Studio via DashScope API, available only in Beijing Region. Includes create, get, list, and delete import job endpoints. Required parameters: model_name, weight_type, storage_info.bucket_name, and object_key. Job statuses: PENDING, RUNNING, SUCCESSED, FAILED. - [Model Compression API](https://help.aliyun.com/en/model-studio/model-compression-api.md): Compress custom full-parameter fine-tuned models (qwen3.5-flash-2026-02-23) via quantization to reduce inference memory usage. Provides RESTful APIs for template listing, job creation/polling/logs/cancellation/deletion. Available only in Beijing Region; currently free of charge. - [Image Generation Model Fine-Tuning API](https://help.aliyun.com/en/model-studio/model-fine-tuning-image-generation-api.md): API entry for image generation fine-tuning on Alibaba Cloud Model Studio. Create fine-tuning jobs to customize wan/qwen-image models with request parameters and call examples - [Model Studio-CreateFineTuningJob Create Image Generation Fine-Tuning Job](https://help.aliyun.com/en/model-studio/image-generation-create-fine-tuning-job-api.md): POST /api/v1/fine-tunes to create LoRA fine-tuning jobs for wan2.7-image-pro/wan2.7-image models, supporting text-to-image (t2i) and image-to-image (i2i). Datasets can be uploaded via API or mounted from OSS. Available only in China (Beijing); max_steps affects training billing. - [Video Generation Model Fine-tuning](https://help.aliyun.com/en/model-studio/model-fine-tuning-video-generation-api.md): Entry for fine-tuning video generation models on Bailian. Create tuning tasks to customize text-to-video and image-to-video models. - [Video Generation-CreateFineTuningJob Create Fine-Tuning Job](https://help.aliyun.com/en/model-studio/video-generation-create-fine-tuning-job-api.md): Alibaba Cloud Model Studio video generation fine-tuning API. POST /api/v1/fine-tunes to create a training job. Available only in China (Beijing) region. Supports wan2.7-i2v and wan2.2-kf2v-flash models for first-frame or first-and-last-frame video generation with efficient_sft LoRA. Datasets via training_file_ids or OSS mount. - [Speech Synthesis Model Fine-tuning API](https://help.aliyun.com/en/model-studio/model-fine-tuning-speech-synthesis-api.md): Entry point for Alibaba Cloud Model Studio speech synthesis model fine-tuning APIs. Create fine-tuning tasks to customize CosyVoice and other TTS models for improved synthesis quality in specific scenarios. - [Model Studio-CreateFineTuningJob Create Voice Synthesis Fine-Tuning Job](https://help.aliyun.com/en/model-studio/voice-synthesis-create-fine-tuning-job-api.md): Creates an efficient_sft fine-tuning job for CosyVoice via POST /api/v1/fine-tunes. Supports file_id or oss_mount dataset sources. Available only in China (Beijing) region with API Key configured. Hyperparameters include lm_max_epoch, lm_step, fm_max_epoch, and fm_batch_size. Response returns job_id and finetuned_output for subsequent query and deployment. - [Model Studio Fine-Tuning Job Management API](https://help.aliyun.com/en/model-studio/get-fine-tuning-job-api.md): REST APIs to query, cancel, delete Model Studio fine-tuning jobs and retrieve logs. Available only in China (Beijing) region with DASHSCOPE_API_KEY configured. Supports status queries for text (qwen3-14b), video (wan2.5-i2v), and image (wan2.7-image-pro) models via SFT or efficient_sft training, returning job_id, status, finetuned_output, hyper_parameters, and usage fields. - [Model Studio-Checkpoint Management API List Export and Validation Query](https://help.aliyun.com/en/model-studio/list-checkpoints-api.md): Manage checkpoints produced by fine-tuning jobs via four APIs: list, export to deployable model, query validation list, and query validation details. Available only in China (Beijing) region with a Beijing-region API Key. Supports text generation, video generation, image generation, and speech synthesis fine-tuned models; validation APIs are limited to video and image generation models. - [Model Studio Deployments API: Online Inference Deployment](https://help.aliyun.com/en/model-studio/deployments-api.md): Use the Deployments API to deploy fine-tuned or imported custom models as online inference services, with provisioned throughput, long-input caching, and dedicated quota management. - [Model Studio - Text Generation Model Deployment APIs](https://help.aliyun.com/en/model-studio/model-deployment-text-generation-api.md): API catalog for publishing text generation models as online inference services: CreateDeployment, ListDeployments, GetDeployment, DeleteDeployment, UpdateDeployment. Available only in China (Beijing); RAM users need model call/training/deployment permissions. - [Deploy Model - CreateDeployment API](https://help.aliyun.com/en/model-studio/create-deployment-api.md): POST /api/v1/deployments publishes a fine-tuned text model as an online service with three billing plans: mu (model unit), lora (token usage), and ptu (provisioned throughput). Required parameters: model_name and plan. For mu plan, deploy_spec, capacity, and billing_method are required; optional fields include enable_thinking, max_context_length, rpm_limit, and tpm_limit. Available only in China (Beijing) region; billing starts immediately after successful deployment. - [Model Deployment - Image Generation API](https://help.aliyun.com/en/model-studio/model-deployment-image-generation-api.md): Entry for invoking and deploying image generation models on Bailian, covering qwen-image, wan2.7-image, wanx text-to-image and image editing series with access methods and request examples - [Model Studio-Deploy Model Publish Fine-tuned Image Generation Model as Online API](https://help.aliyun.com/en/model-studio/image-generation-deploy-model-api.md): Publish a trained image generation model as an online API service via POST /api/v1/deployments, requiring model_name, capacity, and plan parameters. Currently available only in China (Beijing) region. After async deployment, invoke the model using the deployed_model identifier. Billed on a pay-as-you-go (post_paid) basis. - [Model Deployment - Video Generation](https://help.aliyun.com/en/model-studio/model-deployment-video-generation-api.md): Dedicated deployment entry for video generation models on Model Studio, covering wan2.7/wan2.6/happyhorse/vidu/pixverse/kling text-to-video, image-to-video, reference-to-video and video editing models, deployable via API with private or reserved throughput. - [Model Studio-Deployments Deploy Fine-tuned Video Generation Model](https://help.aliyun.com/en/model-studio/video-generation-deploy-model-api.md): POST /api/v1/deployments publishes a fine-tuned video generation model as an online API service, available only in China (Beijing). Required parameters: model_name, capacity, plan, aigc_config. Supports use_input_prompt to control prompt source. Async deployment; invoke via deployed_model identifier. - [Model Studio TTS API Directory (CosyVoice/Qwen-TTS/MiniMax)](https://help.aliyun.com/en/model-studio/model-deployment-speech-synthesis-api.md): Hub for Model Studio text-to-speech APIs: real-time and non-real-time TTS, Qwen-TTS/CosyVoice/MiniMax speech-2.8-hd models, voice cloning and voice design, instruction control, streaming output, Base64/WAV/MP3/Opus audio formats. - [Model Studio-Deploy Model Publish Fine-tuned TTS Model as Online API](https://help.aliyun.com/en/model-studio/tts-deploy-model-api.md): Deploy fine-tuned CosyVoice TTS models as online API services, available only in China (Beijing) region. Use POST /api/v1/deployments with model_name, deploy_spec (MU5/MU2), capacity, and POST_PAY billing; asynchronously returns deployed_model ID for subsequent invocation. - [Model Studio-Deployment Management API](https://help.aliyun.com/en/model-studio/get-deployment-api.md): DashScope deployment management APIs for querying deployment status, listing deployable and deployed models, modifying RPM/TPM rate limits, scaling capacity, and deleting deployments. Available only in China (Beijing) region; deployment takes 5–10 minutes; deletion is irreversible. - [Model Studio Files API (Upload/Query/List/Delete Files)](https://help.aliyun.com/en/model-studio/file-management-api.md): Bailian (Model Studio) Files API (/files endpoint): upload a file, retrieve file info by file_id, list files, and delete files — manage file objects used in model calls. - [Model Studio Upload File API Multi-file Upload and Quotas](https://help.aliyun.com/en/model-studio/upload-file-api.md): POST /api/v1/files uploads files to Model Studio. Supports fine-tune, file-extract, and batch purposes with max 500 MB per file (batch). Total quota: 100 GB storage and 10,000 active files. Returns file_id for model tuning, content analysis, or Batch tasks. - [File Management: Query and Manage Files (Get/List/Delete)](https://help.aliyun.com/en/model-studio/get-file-api.md): DashScope Files API: get file info by file_id, list all files with pagination (page_no/page_size, max 100 per page), delete files; returns url, size, md5 fields. - [Model Studio Advanced Operations (Temp API Key/Async Tasks/Rate Limits)](https://help.aliyun.com/en/model-studio/more-about-models.md): Temporary API Keys, async tasks and callbacks, sub-workspace model calls, temporary file URLs, DashScope SDK connection reuse, model list/rate-limit/permission queries and updates. - [Generate a temporary API key (model)](https://help.aliyun.com/en/model-studio/generate-temporary-api-key.md): Create short-lived API keys (1-1800s, default 60s) for browser/mobile apps. POST /api/v1/tokens with expire_in_seconds. Inherits parent key permissions. Cannot be manually revoked. - [Asynchronous task management API](https://help.aliyun.com/en/model-studio/manage-asynchronous-tasks.md): Query, batch-list, and cancel async tasks for image/video generation models. GET /api/v1/tasks/{task_id} returns status (PENDING/RUNNING/SUCCEEDED/FAILED) and results. 20 QPS rate limit per account. - [Async task completion notifications](https://help.aliyun.com/en/model-studio/async-task-api.md): Receive task completion events via EventBridge instead of polling. Configure an HTTP callback URL or RocketMQ queue. Polling is rate-limited to 20 QPS; EventBridge push has no limit. - [Model calls in a sub-workspace](https://help.aliyun.com/en/model-studio/model-calling-in-sub-workspace.md): Invoke models using a sub-workspace API key for per-user access control or cost allocation. Examples in Python, Java, Node.js, Go, C#, PHP for OpenAI-compatible and DashScope methods. - [Upload files for temporary URLs](https://help.aliyun.com/en/model-studio/get-temporary-file-url.md): Get oss:// URLs (48h expiry) for multimodal model inputs. Files bound to specific model and account. 100 QPS limit. HTTP calls require X-DashScope-OssResourceResolve: enable header. - [DashScope SDK connection reuse](https://help.aliyun.com/en/model-studio/connection-multiplexing-configuration.md): Configure TCP connection pooling for high-concurrency DashScope calls. Java SDK: connectionPoolSize, maximumAsyncRequests (default 32). Python SDK: custom Session for sync/async reuse - [Model Studio ListModels API (GET /api/v1/models)](https://help.aliyun.com/en/model-studio/list-models.md): Call GET /api/v1/models to list available models on Model Studio. Filter by providers (qwen, DeepSeek, MiniMax, etc.), capabilities (TG, VU, IG, VG, ASR, TTS, etc.), features, context_window, service_site, and deployment_methods. Returns model ID, pricing, context length, and modality info. Requires DASHSCOPE_API_KEY. - [Query Model Rate Limits (GET /api/v1/models/limits)](https://help.aliyun.com/en/model-studio/list-quotas.md): Call GET /api/v1/models/limits to list RPM/QPS, token usage caps, and async task concurrency/queue limits per model under an API Key. Filter by name/model; returns account-level and workspace-level quotas for troubleshooting 429 errors and capacity planning. - [Update Model Rate Limits (Set or Delete QPM/TPM Quotas)](https://help.aliyun.com/en/model-studio/update-model-rate-limits.md): POST /api/v1/models/limits: set or modify QPM (request_limit), TPM (usage_limit) and period for models in a workspace. Supports OVERLAY (merge) and DELETE operations; includes TPM exemption two-step method and InvalidParameter error handling. - [Model Studio ListModelPermissions Query Model Authorizations](https://help.aliyun.com/en/model-studio/list-model-permissions.md): GET /api/v1/models/permissions lists authorizable or authorized models and permission details in a workspace. Supports name/model filtering, authorization_scope (AUTHORIZABLE/AUTHORIZED), page_no/page_size pagination; returns inference/fine_tune/deploy booleans. Requires DASHSCOPE_API_KEY. - [Model Studio UpdateModelPermissions API](https://help.aliyun.com/en/model-studio/update-model-permissions.md): POST /api/v1/models/permissions to grant or revoke inference, finetune, and deploy permissions for models like qwen-plus and qwen3-max in a workspace, either per-model (models array, 1–20 items) or one-click via access_all_entities (OPEN/CLOSE/KEEP). Requires DASHSCOPE_API_KEY Bearer auth. ## Model Studio Application API Reference (App Calls/RAG/Managed Agents) - [Model Studio Application API Reference (App Calls/RAG/Managed Agents)](https://help.aliyun.com/en/model-studio/applicantion-api-reference.md): App-side OpenAPI hub: application call, knowledge retrieval & QA, Managed Agents, components, long-term memory, frameworks. Assistant API deprecated. - [What is a Bucket? Naming Rules, Region, Storage Classes and ACL](https://help.aliyun.com/en/model-studio/managed-agents-api.md): OSS container for Objects. Globally unique name required (3-63 chars). Region fixed after creation. Storage classes: Standard/IA/Archive/Cold Archive; ACL configurable. - [Overview and authentication](https://help.aliyun.com/en/model-studio/managed-agents-api-overview.md): The Managed Agents API is the agent runtime hosted by Model Studio: sessions, sandboxes, tool execution, and event streams are all managed by the platform. - [Quickstart](https://help.aliyun.com/en/model-studio/managed-agents-quickstart.md): Get the full Managed Agents workflow running in five minutes: create an Agent and Environment, open a Session, submit a task, and stream the results. - [Managed Agents - Agent API (Create/Get/List/Update/Archive)](https://help.aliyun.com/en/model-studio/agent-api.md): Agent = reusable config: model, prompt, tools, skills. POST/GET /agents: full-replace updates with version lock (409), soft archive, ?version=N history, sessions pin version. - [Managed Agents - Create Agent (POST /agents)](https://help.aliyun.com/en/model-studio/agent-create.md): POST /agents creates an agent: required name/model (e.g. qwen3-max), optional system prompt, tools, mcp_servers, skills. Returns agent_id and version. - [Managed Agents - Get Agent (GET /agents/{agent_id})](https://help.aliyun.com/en/model-studio/agent-get.md): GET /agents/{agent_id} retrieves an agent's full config (system prompt, model, tools, mcp_servers, skills). Defaults to latest; pass version for a past one. Auth via Bearer DASHSCOPE_API_KEY. - [Managed Agents - List Agents (GET /agents)](https://help.aliyun.com/en/model-studio/agent-list.md): GET /agents paginates workspace agents in reverse created_at order. Params: limit (default 20, max 100), page cursor, include_archived; returns data array and next_page. - [Managed Agents – Update Agent API (Full Replacement, version Optimistic Lock)](https://help.aliyun.com/en/model-studio/agent-update.md): Update Agent via POST /agents/{agent_id}, full replacement: version required (optimistic lock, 409 on mismatch), omitted fields cleared, version auto-increments. - [Managed Agents - Archive Agent (POST /agents/{agent_id}/archive)](https://help.aliyun.com/en/model-studio/agent-archive.md): Soft-archive an agent: archived_at is set; agent hidden from lists unless include_archived=true and cannot start new sessions; existing sessions unaffected. Bash/Python/Java examples. - [Managed Agents - List Agent Versions (GET /agents/{agent_id}/versions)](https://help.aliyun.com/en/model-studio/agent-versions.md): GET /agents/{agent_id}/versions: paginated historical version snapshots of an agent in descending order; limit default 20 (max 100), cursor paging via next_page. Sessions lock the version at creation. - [Managed Agents – Environment API (Create/Update/Archive/Delete)](https://help.aliyun.com/en/model-studio/environment-api.md): Managed Agents Environment API: POST /environments to create sandbox and dependencies, GET retrieve/list, full-replacement update (sessions keep bound snapshot), soft archive, irreversible DELETE. - [Model Studio-POST /environments Create Runtime Environment (Sandbox)](https://help.aliyun.com/en/model-studio/environment-create.md): POST /environments creates a sandbox runtime: name required, config.type=cloud immutable, apt/pip/npm preinstall, egress unrestricted only. Returns an Environment object. - [Managed Agents – Get Environment API (Sandbox Config)](https://help.aliyun.com/en/model-studio/environment-get.md): GET /environments/{environment_id}: returns the full sandbox config by environment ID (env_) — sandbox type, apt/pip packages, networking policy, scope and archived status. Bearer API Key auth. - [Managed Agents-List Environments](https://help.aliyun.com/en/model-studio/environment-list.md): Paginate Environments in a workspace via GET /environments, sorted by created_at desc, archived excluded by default. Params: limit (max 100), page, include_archived; response has data and next_page. - [Model Studio – Update Environment (POST /environments/{environment_id})](https://help.aliyun.com/en/model-studio/environment-update.md): POST /environments/{environment_id} partial update: only passed fields apply; config.type immutable; packages/metadata replaced wholesale; running sessions keep the snapshot. - [Managed Agents - Delete Environment (Hard Delete)](https://help.aliyun.com/en/model-studio/environment-delete.md): DELETE /environments/{environment_id} hard-deletes the environment and config (unrecoverable). Use archive API to retain settings. Bearer DASHSCOPE_API_KEY auth. - [Managed Agents - Archive Environment (Soft Archive API)](https://help.aliyun.com/en/model-studio/environment-archive.md): Soft-archive an Environment via POST /environments/{environment_id}/archive: retained and still queryable via GET, bound sessions stay usable; hidden from lists unless include_archived=true, archived_at records the archive time. Bash/Python/Java samples. - [Managed Agents – Session and Event API (State Machine & SSE Events)](https://help.aliyun.com/en/model-studio/session-api.md): Managed Agents Session/Event API: state machine (idle/running/terminated), session create/archive/delete, event injection, SSE stream subscription; hard delete erases event history permanently. - [Managed Agents - Create Session (POST /sessions)](https://help.aliyun.com/en/model-studio/session-create.md): POST /sessions creates a session: required agent and environment_id bind agent and environment, locking the agent's latest snapshot. Returns sesn_ ID, status idle. - [Managed Agents - Get Session](https://help.aliyun.com/en/model-studio/session-get.md): GET /sessions/{session_id}: returns session metadata, status (idle/running/terminated), agent snapshot frozen at creation, environment_id; events/messages via Events API. - [Managed Agents - List Sessions (GET /sessions)](https://help.aliyun.com/en/model-studio/session-list.md): GET /sessions lists agent sessions paginated: filter by agent_id, statuses[], created_at range; limit≤100, next_page cursor; returns sesn_ session objects. - [Managed Agents-Update Session (title/metadata)](https://help.aliyun.com/en/model-studio/session-update.md): POST /sessions/{session_id} modifies only title and metadata — the only operation that fires session.updated. Response returns updated fields, updated_at, request_id. - [Managed Agents-Delete Session (DELETE /sessions/{session_id})](https://help.aliyun.com/en/model-studio/session-delete.md): Hard-delete a session: DELETE /sessions/{session_id} removes metadata, event history and copied resources, irreversible; use the archive API to keep history. Returns session_deleted. - [Managed Agents-Archive Session (POST /sessions/{session_id}/archive)](https://help.aliyun.com/en/model-studio/session-archive.md): POST /sessions/{session_id}/archive: archive a session; status becomes terminated (final), archived_at set, event history still queryable. Returns full session object. - [Managed Agents - Send Event API (Write Events to Session)](https://help.aliyun.com/en/model-studio/event-post.md): POST /sessions/{session_id}/events writes events into a session: user message, interrupt, tool approval, result backfill. input holds 1-50 events; Session required. - [Model Studio - List Events: Paginated Session History Events](https://help.aliyun.com/en/model-studio/event-list.md): Paginates session history events GET /sessions/{session_id}/events; response same as SSE frame data. Filter by created_at window + order, limit max 100, next_page cursor. - [Subscribe to Event SSE Stream](https://help.aliyun.com/en/model-studio/event-sse-stream.md): Subscribes to a real-time session event stream via SSE: assistant output messages, tool calls, and status changes. Upon connection, the server immediately sends a : connected comment line. - [File API: Upload, Mount to Session Sandbox, Quota and Status](https://help.aliyun.com/en/model-studio/files-api.md): File resources: upload once, mount to session sandboxes or use as message attachments; 20 MB/file, 100 GB/workspace, 30-day retention; only status=available usable. Endpoints: POST/GET/DELETE /files. - [Upload File](https://help.aliyun.com/en/model-studio/file-upload.md): Uploads a single file using multipart/form-data, with a maximum size of 10 MB. The file enters a security review after upload; only files with available status can be mounted or used as message content. - [Get a file](https://help.aliyun.com/en/model-studio/file-get.md): Returns file metadata (content excluded). Commonly used to confirm that the post-upload safety scan status has transitioned to available. - [Model Studio File Management - List Files (GET /files)](https://help.aliyun.com/en/model-studio/file-list.md): GET /files paginates workspace files by created_at desc. Params: limit (max 100), page (cursor), scope_id (sandbox filter). Response: filename, size_bytes, status. - [Delete File](https://help.aliyun.com/en/model-studio/file-delete.md): Hard deletes the file metadata and original content; this action cannot be undone. Internal copies already mounted to session sandboxes are not affected; new sessions can no longer mount this file_id. - [Model Studio Skill API (Create/Versions/Security Scan/Mount to Agents)](https://help.aliyun.com/en/model-studio/skills-api.md): Skill API: mount zip skills to agents after security scan (checking/active/rejected); exact version required (no latest). Skill/version CRUD and zip download endpoints. - [Create a skill](https://help.aliyun.com/en/model-studio/skill-create.md): Create a skill from an uploaded zip package. The request body contains only file_id; name and description are parsed by the server from SKILL.md. The new skill enters checking and becomes active after passing the scan. - [Model Studio - GetSkill: Retrieve Skill Metadata and Latest Version](https://help.aliyun.com/en/model-studio/skill-get.md): GET /skills/{skill_id} returns skill metadata and latest_version. Only status=active skills are mountable; mounting requires a specific version, never latest. - [Skill-ListSkills: List Skills (GET /skills)](https://help.aliyun.com/en/model-studio/skill-list.md): GET /skills lists skills via cursor pagination; filter source=customer or official, limit max 100; only status=active skills are mountable and require a pinned version. - [Skill - Delete Skill (DELETE /skills/{skill_id})](https://help.aliyun.com/en/model-studio/skill-delete.md): DELETE /skills/{skill_id} deletes a skill and all versions: can't be attached to new agents; agents already on an old version keep working. Auth: Bearer DASHSCOPE_API_KEY. - [Upload New Skill Version](https://help.aliyun.com/en/model-studio/skill-version-create.md): Uploads a new version zip package for an existing skill. The version number is generated by the server; the new version also undergoes a security scan; agents mounted with older versions are not affected. - [Model Studio: List Skill Versions (GET /skills/{skill_id}/versions)](https://help.aliyun.com/en/model-studio/skill-version-list.md): GET /skills/{skill_id}/versions lists skill versions in descending order; limit/page pagination; rejected versions include error_info with security review issues. - [Model Studio – Get Skill Version: Metadata & Scan Status](https://help.aliyun.com/en/model-studio/skill-version-get.md): GET /skills/{skill_id}/versions/{version}: metadata & scan status of a Skill version; status checking/active/rejected, only active mountable; rejected returns error_info. - [Skill - Download a Skill Package Version (Presigned URL)](https://help.aliyun.com/en/model-studio/skill-version-download.md): GET /skills/{skill_id}/versions/{version}/content downloads a Skill package version: returns an OSS presigned file_url valid for 2 hours; GET it to fetch the zip. Requires Endpoint and auth setup. - [Model Studio App Invocation API (Call Agent/Workflow Apps)](https://help.aliyun.com/en/model-studio/application-call.md): Call Bailian agent/workflow apps: sync and async APIs, app_id/prompt/session_id/biz_params, OpenAI-compatible endpoint, streaming; apps free, only inference billed. - [Get app ID and workspace ID](https://help.aliyun.com/en/model-studio/obtain-the-app-id-and-workspace-id.md): Copy APP ID and workspace ID from the Model Studio console for API calls to agents and workflows. Sub-workspace apps require both IDs; default workspace apps need only APP ID. - [Model Studio Application DashScope API (New Agent / Workflow & Legacy Agent)](https://help.aliyun.com/en/model-studio/application-dashscope-api-reference.md): Entry for calling Model Studio apps via DashScope API: new agent application API and workflow/legacy agent application API—choose the endpoint matching your app type. - [New Agent Application API](https://help.aliyun.com/en/model-studio/new-agent-application-api-reference.md): DashScope API for new agent (Agent 2.0) applications. Endpoint: POST /api/v1/apps/APP_ID/completion. Supports streaming and non-streaming modes, with Python/Java/cURL examples - [Model Studio – Workflow & Agent App API Reference (DashScope)](https://help.aliyun.com/en/model-studio/agent-and-workflow-application-api-reference.md): DashScope API for Bailian workflow and legacy agent apps: POST /api/v1/apps/{APP_ID}/completion, app_id and request/response params, streaming, multi-turn dialog, file upload and RAG examples, Python/Java SDK. Beijing region only. - [Model Studio Responses API (Sync & Async Call Reference)](https://help.aliyun.com/en/model-studio/openai-responses-api.md): Responses API entry: synchronous and asynchronous call API references. OpenAI-compatible model invocation interface on Model Studio for qwen and other models. - [Responses API -- synchronous invocation](https://help.aliyun.com/en/model-studio/synchronous-call-api-reference.md): OpenAI-compatible Responses API for synchronous calls to agent/workflow apps. Supports text, image, file input and streaming. Endpoint: /api/v2/apps/agent/{APP_ID}/compatible-mode/v1/responses. - [Responses API asynchronous calls](https://help.aliyun.com/en/model-studio/asynchronous-call-api-reference.md): OpenAI-compatible async calls for Agent/Workflow apps. Set background=true to get a task ID, then poll with responses.retrieve(). Statuses: completed, failed, cancelled. Python and Java examples - [Model Studio Application Component API Reference (Overview/Endpoints/Auth)](https://help.aliyun.com/en/model-studio/application-component-api-reference.md): Entry to Application Component OpenAPI: API overview, service endpoints, RAM authorization, API catalog, and version notes (bailian 2023-12-29). - [Bailian component API overview](https://help.aliyun.com/en/model-studio/api-bailian-2023-12-29-overview.md): ROA-standard OpenAPI (bailian/2023-12-29) for application data, knowledge base, and prompts. Key APIs: CreateIndex, Retrieve, AddFile, ListChunks, CreatePromptTemplate, memory CRUD - [RAM authorization for Model Studio](https://help.aliyun.com/en/model-studio/api-bailian-2023-12-29-ram.md): Grant RAM sub-users access to Model Studio resources. Covers AliyunBailianFullAccess managed policy, custom policy syntax for workspace/app/model scopes, and cross-account delegation. - [API change records](https://help.aliyun.com/en/model-studio/api-bailian-2023-12-29-changeset.md): Change time: 2025-12-03Change time:APIChangeActionCreateIndexThe request parameters have changed.View change detailsView API documentationChange time:... - [Long-Term Memory APIs (AddMemory/SearchMemory/DeleteMemory)](https://help.aliyun.com/en/model-studio/long-term-memory-new.md): Long-term memory API catalog: AddMemory to write memory fragments, SearchMemory for semantic retrieval, DeleteMemory and more; Pro/Lite tiers by Rerank; commercial billing starts 2026-08-20. - [Long-term memory API reference](https://help.aliyun.com/en/model-studio/long-term-memory-api-reference.md): REST APIs: AddMemory, SearchMemory, ListMemory, DeleteMemory, UpdateMemory, plus ProfileSchema CRUD. Base URL: /api/v2/apps/memory/. Rate limits: 3000 QPM total, 120 QPM for add, 300 QPM for search - [Model Studio Framework Integration (LlamaIndex/Spring AI Alibaba)](https://help.aliyun.com/en/model-studio/frameworks.md): Integrate Model Studio models and apps via LlamaIndex and Spring AI Alibaba — dev guides for building RAG retrieval and Agent applications. - [Build a Cloud RAG App with LlamaIndex (Knowledge Base Q&A)](https://help.aliyun.com/en/model-studio/llamaindex.md): Build a cloud RAG app with LlamaIndex on Model Studio: upload txt/docx/pdf to DashScopeCloudIndex, answer with qwen-max; custom chunking/embeddings not supported. - [LLMs from Model Studio in LlamaIndex](https://help.aliyun.com/en/model-studio/dashscopellm-in-llamaindex.md): Two integration methods: OpenAILike wrapper (llama-index-llms-openai-like) for OpenAI-compatible models, and DashScope native SDK (llama-index-llms-dashscope) for all text generation models - [DashScopeEmbedding in LlamaIndex](https://help.aliyun.com/en/model-studio/dashscopeembedding-in-llamaindex.md): Build vector indexes with DashScopeEmbedding (text-embedding-v2, 1536 dims). Install llama-index-embeddings-dashscope. Batch embedding via get_text_embedding_batch() - [DashScopeRerank](https://help.aliyun.com/en/model-studio/dashscopererank.md): GTE-Rerank text reranking for LlamaIndex. Params: model (gte-rerank), top_n, return_documents. Install llama-index-postprocessor-dashscope-rerank. Use postprocess_nodes() to reorder retrieval results - [DashScopeParse](https://help.aliyun.com/en/model-studio/dashscopeparse.md): LlamaIndex document parser using Document Mind. Supports PDF/DOC/DOCX (up to 100 MB, 1000 pages). Params: result_type, num_workers, check_interval, category_id. Install llama-index-readers-dashscope - [DashScopeJsonNodeParser](https://help.aliyun.com/en/model-studio/dashscopejsonnodeparser.md): LlamaIndex text chunking for DashScopeParse output. Params: chunk_size (default 500), overlap_size (100), separator, language (cn/en/any). Install llama-index-node-parser-dashscope. Python 3.8-3.12 - [DashScopeCloudIndex and DashScopeCloudRetriever](https://help.aliyun.com/en/model-studio/dashscopecloudindex-and-dashscopecloudretriever.md): Build a cloud knowledge base from local files using LlamaIndex. DashScopeCloudIndex.from_documents() creates indexes; DashScopeCloudRetriever queries them. Supports DashScopeParse for PDF/DOC parsing - [DashScopeCloudRetriever](https://help.aliyun.com/en/model-studio/dashscopecloudretriever.md): LlamaIndex retriever SDK for cloud knowledge bases. Key params: dense_similarity_top_k, sparse_similarity_top_k, enable_reranking, rerank_model_name (gte-rerank-hybrid). Python 3.8-3.12 - [Spring AI Alibaba for Model Studio (App Calls/Knowledge Base RAG)](https://help.aliyun.com/en/model-studio/spring-ai-alibaba.md): Entry to Spring AI Alibaba, the Spring AI integration for Model Studio: invoke Bailian applications and retrieve from Bailian knowledge bases (RAG) in Java/Spring Boot apps. - [Spring AI Alibaba -- integrate Model Studio apps](https://help.aliyun.com/en/model-studio/spring-ai-alibaba-integrate-llm-application.md): Model Studio agent/workflow apps from Spring Boot using DashScopeAgent class. Requires JDK 17+, Spring Boot 3.x. Dependency: spring-ai-alibaba-starter-dashscope. - [Spring AI Alibaba -- knowledge base retrieval](https://help.aliyun.com/en/model-studio/spring-ai-alibaba-integrate-knowledge-base.md): Query a Model Studio knowledge base from Spring Boot via DashScopeDocumentRetriever. Uses DashScopeApi for RAG retrieval with qwen-max as default LLM. JDK 17+. - [Assistant API (Deprecating) - Model Studio Application API Reference](https://help.aliyun.com/en/model-studio/assistantapi.md): Model Studio Assistant API reference entry. This suite is being deprecated; new development should use Managed Agents and Application Invocation APIs in the same catalog. - [Deprecated Assistants API](https://help.aliyun.com/en/model-studio/assistant.md): Deprecated. Migrate to Responses API. CRUD operations for assistant objects via DashScope SDK (Assistants.create/list/retrieve/update/delete). Supports model, tools, temperature, and metadata params - [Deprecated: Assistant API threads](https://help.aliyun.com/en/model-studio/thread.md): CRUD operations on Thread objects for multi-turn conversations. Threads persist on server and do not expire. Deprecated -- migrate to Responses API. - [Deprecated Assistant API Messages](https://help.aliyun.com/en/model-studio/message.md): Being unpublished. Message CRUD for the Assistant API: Messages.create/list/retrieve/modify within a thread. Migrate to Responses API for built-in multi-turn context management - [Deprecated: Assistant API runs](https://help.aliyun.com/en/model-studio/runs.md): Create, retrieve, and cancel agent runs in a thread. Supports streaming. Deprecated -- migrate to Responses API. Key params: thread_id, assistant_id, stream. - [Deprecated: Assistant API run steps](https://help.aliyun.com/en/model-studio/run-steps.md): List run steps for an agent execution, including model and tool call details. Deprecated -- migrate to Responses API. SDK: Steps.list(run_id, thread_id). - [Deprecated: Assistant API streaming events](https://help.aliyun.com/en/model-studio/event-streaming.md): Being deprecated. Event stream format for Assistant API: thread.message.delta and thread.run.step.delta objects. Tool call types: code_interpreter, quark_search, text_to_image, calculator - [Deprecated Assistant API call examples](https://help.aliyun.com/en/model-studio/call-example.md): Being unpublished. Python/Java examples for the Assistant API: simple assistant, tool calling (code_interpreter, quark_search), and streaming runs. Migrate to Responses API - [Model Studio Miscellaneous (Service-Linked Role/Temporary API Key/SearchFilters)](https://help.aliyun.com/en/model-studio/more.md): Model Studio misc entry: service-linked role authorization, generating temporary API Keys (temporary auth tokens), and how to use knowledge base SearchFilters. Check here for permissions and temporary credentials. - [Service-linked roles](https://help.aliyun.com/en/model-studio/bailian-service-linked-role.md): Auto-created SLRs for accessing FC, OSS, ADB-PG, MNS, SLS, CMS, and DTS. Lists 11 roles (e.g. AliyunServiceRoleForSFMAccessFC) with permissions and deletion instructions - [Generate a temporary API key (application)](https://help.aliyun.com/en/model-studio/application-obtain-temporary-authentication-token.md): Exchange API key for a short-lived token (default 60s) via POST /api/v1/tokens for client-side apps. Custom expiry with expire_in_seconds. Python, Node.js, cURL examples - [Model Studio Knowledge Base SearchFilters (Single/Multi/Range/Wildcard/Tag Query)](https://help.aliyun.com/en/model-studio/how-to-use-search-filters.md): Pass searchFilters to the Retrieve API to filter semantic retrieval results on structured fields. Subgroups use AND logic; supports single-value, multi-value, range (gte/lte), wildcard (like % _), and tag queries, with Python/Java SDK samples and request/response payloads.