Agent Application (Agent 2.0)
The new agent application (Agent 2.0) unifies knowledge bases, MCP, and other capabilities as tools, and invokes them through autonomous reasoning and planning to solve complex tasks.
Version Comparison and Selection Guide
The new agent offers better performance and development experience for most use cases. If you have no dependencies on the legacy version, the new version is recommended.
| Dimension | Legacy Version (Agent 1.0) | New Version (Agent 2.0) |
|---|---|---|
| Planning and Scheduling | The agent retrieves from the knowledge base first, then decides whether to invoke other tools such as MCP. | Knowledge bases and MCP are unified as tools, and the agent autonomously plans when and in what order to invoke them. |
| Process Transparency | Only the final result is displayed, with no way to fully trace intermediate decisions. | The entire "Plan - Execute - Reflect" chain of each step is fully visible. |
| Applicable Scenarios | Suitable for simple tasks with a single intent and a fixed process. | Capable of handling tasks ranging from simple Q&A to complex planning. |
Quick Start: Create a Basic Agent
- Go to the Application Management page in the Model Studio console. Click Create Application, and select Agent Application > Agent 2.0.
- Enter an application name and click Create Now. After creation, you are automatically redirected to the application configuration page.
- Select a model from the model selector dropdown, for example
Qwen-Plus-Latest. - Enter a question in the dialog box on the right, such as
Who are you?.
Capability Configuration
Model Selection
The model selector dropdown offers two types of options:
- Smart Mode: The platform automatically balances cost and performance, so you don't need to pick a model manually.
- Specified Model: Directly select a specific model. For optimal multi-step planning, we recommend selecting a model with strong tool-calling capabilities, such as the
Qwen-Maxseries. Click More Models to see additional models.
Smart Mode includes four tiers — AUTO, Performance, Balanced, and Economy — each selected manually by you. The difference between AUTO and the other three tiers is: when AUTO is selected, the platform automatically assesses task complexity and dynamically routes to the appropriate tier; when one of the other three tiers is selected, that tier is used consistently.
| Mode | Positioning | Features |
|---|---|---|
| AUTO | Auto routing, multi-model, context-adaptive | Intelligently matches the optimal thinking depth and model |
| Performance | High-quality models, balancing speed and quality | Fast response, 1M context, tool calling |
| Balanced | Advanced reasoning, high-quality output | Enhanced reasoning, 1M context, stable output |
| Economy | Standard reasoning, high cost-efficiency | Standard reasoning, 1M context, high cost-efficiency |
AUTO assesses the complexity of each request and dynamically routes to the appropriate tier:
- Simple tasks (single-step, short Q&A, chit-chat, etc.) → Economy
- Medium tasks (single-file code, medium-length writing, multi-step reasoning, etc.) → Balanced
- Complex tasks (multi-file architecture, long-form writing, multi-tool orchestration, etc.) → Performance
NoteComplexity assessment considers both the user's latest message and the current context length: when the context exceeds 200K tokens, requests are routed to at least the Performance tier to ensure quality on long-context tasks.
When AUTO is selected, billing uses a unified list price regardless of which model the request is actually routed to. When you directly select the Performance / Balanced / Economy tier, billing follows that tier. See Billing.
Click Settings on the right side of the model selector to modify the following parameters:
- Maximum Response Length: The maximum length of model-generated content, excluding the prompt.
- Temperature: Controls the randomness and diversity of the output. Higher values increase randomness.
- enable_thinking: Whether to enable thinking mode. Enabling it helps improve the agent's reflection capability. Models that do not support thinking mode cannot configure this parameter.
Prompts
The system prompt defines the agent's role, behavioral instructions, and capability boundaries, ensuring consistency, controllability, and task compliance during interactions.
-
Configure the System Prompt: Set the system prompt to
Please answer my questions in the style of "One Hundred Years of Solitude". The following is an effect comparison:- Without a system prompt: The model replies using its default persona.
- With the system prompt configured: The model replies in the specified literary style.
-
Using Custom Variables in the System Prompt (Optional): In addition to static text, the system prompt also supports embedding custom variables.
- Click Custom variables in the upper right corner of Prompts, configure the custom variable, and click OK to save.
- Type
/to use a configured variable.
File Pre-parsing
The file pre-parsing feature controls how uploaded files are processed.
- Pre-parsing disabled: When disabled, the system does not proactively parse files. The file URL is passed as context information to the agent, which can then decide whether to invoke tools and pass the URL as a parameter in subsequent steps.
- Pre-parsing enabled: When enabled, the system uses built-in parsers to process uploaded documents, images, videos, audio files, and other formats, returning parsed text content to the model as reference.
NoteQwen-VL series models have multimodal capabilities and can directly parse images and video files even when file pre-parsing is disabled.
In all other cases (for example, when using a text-only model without multimodal capabilities, or when using a Qwen-VL series model to process non-image/video files), the agent's file processing capability strictly follows the enabled or disabled logic described above.
Built-in Tools
Built-in tools run in an isolated sandbox environment, providing code execution and file operation capabilities. All tools are disabled by default and can be enabled as needed.
| Tool | Description |
|---|---|
bash | Executes shell commands, supporting bash, python3, pip, and other command-line operations. |
write | Creates or overwrites files, automatically creating parent directories. |
read | Reads file content, supporting line-range reading for large files. |
edit | Performs exact text replacements (find and replace) in files. |
glob | Searches for files by pattern, returning a list of matching file paths. |
grep | Searches for content within files, supporting regular expressions. |
download_file | Exports files generated during execution as downloadable links. |
Knowledge Base
The knowledge base enables the agent to query external information and use the retrieved content as the basis for generating answers. This proactive approach to knowledge retrieval improves answer accuracy and reduces hallucinations when dealing with private knowledge or domain-specific Q&A. To create and manage knowledge bases, see Create and Use a Knowledge Base.
Agent 2.0 applications use knowledge bases by associating a knowledge retrieval service. A knowledge retrieval service can contain one or more knowledge bases and supports both single-knowledge-base retrieval and multi-knowledge-base joint retrieval. An application can be associated with only one knowledge retrieval service at a time; associating another service replaces the existing one. Before you associate a service, make sure that you have created a knowledge retrieval service, added the required knowledge bases to the service, and published the service. See Knowledge Retrieval.
NoteEnable Show Source in the Response section to display knowledge sources and source file URLs.
Associate a knowledge retrieval service
-
Go to the configuration page of the Agent 2.0 application and find the Retrieval Strategy area under Knowledge Base. On first configuration, the page indicates that no knowledge retrieval service is associated. Click Configure Now to open the service selection panel. You can also click Configure in the upper-right corner of the area.
-
In the Configure Retrieval Service panel, find the target service. You can use the search box to find published retrieval services, and confirm the target based on the service name, version, included knowledge bases, and the number of knowledge bases.
-
Click Associate on the right of the target service to associate the knowledge retrieval service with the current application.
After the association is complete, the Retrieval Strategy area displays the service name, invocation settings, and the knowledge bases and weights included in the service.
Adjust invocation settings
You can adjust Invocation Method and Multi-turn Rewrite on the service card to control how the application uses the retrieval service:
| Item | Option | Description |
|---|---|---|
| Invocation method | Intelligent invocation | The model autonomously decides whether to invoke the knowledge retrieval service based on the question and conversation context. |
| Invocation method | Always invoke | In each request, a knowledge retrieval is performed before the model's first inference. |
| Multi-turn rewrite | On | When the model autonomously invokes the knowledge retrieval service, it generates a retrieval query based on the current question and conversation context. |
| Multi-turn rewrite | Off | The original input is used for retrieval, and the rewritten query generated by the model is not used. |
Manage the associated service
The upper-right corner of the service card provides the following actions:
- Settings: Click the gear icon to go to the corresponding knowledge retrieval service page and modify the knowledge bases included in the service and the retrieval strategy.
- Delete: Click the delete icon to remove the association between the current application and the knowledge retrieval service. This removes the association in the application only, not the knowledge retrieval service or the knowledge bases themselves.
Modify the retrieval service configuration and publish
After you click Settings on the service card, you are directed to the Service Configuration page of the corresponding knowledge retrieval service.
- Add a knowledge base: Click Add in the upper-right corner of the knowledge base area to add a new knowledge base to the service.
- Adjust the retrieval strategy: Adjust the knowledge base weights, knowledge base routing, hybrid ranking model, and other settings as needed. The available settings are subject to the service page.
- Publish the changes: After the configuration is complete, click Publish in the upper-right corner of the page to publish the knowledge retrieval service configuration. Make sure you complete this step after modifying the configuration.
- Return to the application to confirm: Return to the Agent 2.0 application page and click the refresh icon in the upper-right corner of the Retrieval Strategy area to confirm that the knowledge bases included in the service and the related information have been updated.
NoteLegacy knowledge bases are integrated as MCP tools and appear under Tools → MCP Services on the new Model Studio page. To add a legacy knowledge base, go to the legacy Model Studio page.
The new agent supports using tags to narrow the query scope of knowledge bases. By assigning tags to knowledge base files and defining usage rules in the system prompt, you can guide the agent to search within a smaller, more precise set of files based on user intent, significantly improving answer accuracy and relevance. For more information, see Knowledge Base Tag Filtering for New Agent.
MCP
In the new agent, external tools are integrated via the MCP protocol and incorporated into the scheduling system, including official MCP services from the MCP Market and custom MCP services. The agent can dynamically invoke MCP in a non-fixed order during multi-step reasoning to handle more complex tasks. Additionally, plugins can be converted to MCP services with one click.
Data Connectors
Data connectors serve as bridges for agents, workflows, and knowledge bases on the Model Studio platform to access external data. By configuring data connectors, agents can directly query and operate on data from external data sources.
Application Components
Integrate existing agents or workflows as tools. The agent or workflow application must first be published as a component.
Skills
Skills are capability packages that can be added to agents, enabling the agent to automatically handle specific types of tasks during conversations. After adding skills, the agent automatically identifies matching tasks and invokes the corresponding skill for processing, without requiring additional code. For more information, see Skill.
Memory
- Short-term Memory: The new agent supports short-term memory, which provides context information across multi-turn conversations. You can set 0 to 30 turns of context (0 means no multi-turn conversation history is passed). More turns increase conversational relevance but also increase input length.
- Long-term Memory: This feature is planned for future iterations.
Environment
The environment section is used to configure authentication credentials and environment variables required by skills. Once configured, the agent automatically injects the corresponding credentials and parameters when invoking skills, eliminating the need to hardcode sensitive information in skill code.
Response
The response section supports displaying answer sources. When enabled, knowledge sources and source file/web page URLs are displayed as footnote markers. This feature is recommended for use in combination with knowledge bases and the web search MCP.
Run and Result Analysis
After completing the application configuration, you can run the agent in the dialog window on the right side of the page. For complex requests requiring multi-step planning, the new agent displays its decision-making process and execution trace as a card flow. The process mainly includes the following two types of steps:
- Thinking: This step displays the model's reasoning logic, which helps analyze its decision path and locate the root cause of unexpected behavior (only appears when a model that supports thinking mode is selected).
- Tool Invocation: This step records the specific tool invocation parameters and returned results executed by the model.
ReAct Maximum Rounds (range: 1-50) limits the maximum number of tool invocations the agent can make in a single session. When this limit is exceeded, the agent automatically exits the tool invocation loop and generates the final response.
Publishing and Integration
WarningPublishing the application is a prerequisite for all subsequent agent application invocations and integrations.
Publish Application
In the upper-right corner of the application configuration page, click Publish. A pop-up window displays the configuration changes since the last publication. After confirming the release information, click Confirm Publish to complete the publication.
API Invocation
In the Publish Channel tab of the agent application, click View API on the right side of API Call to view the API invocation method for the new agent application. For more information, see New agent application API.
Application Management
Version Management
The version management feature allows you to edit the description of historical versions or roll back to a previously published version.
- On the application configuration page, click Version Management on the right side of the top navigation bar.
- Select the historical version you want to roll back to, hover over the card, and click the edit icon in the upper-right corner. In the Edit Version Description dialog box, make the necessary modifications and click OK to save. Click Overwrite Current Draft to roll back to that version.
Billing
Agent billing mainly includes the following aspects:
-
Model Invocation
- Agents incur model invocation fees, which depend on the model type and the number of input and output tokens.
- For specific model types and corresponding billing rules, see Recommended models. When using Smart Mode, the list price of each tier is detailed below.
-
Knowledge Base
- Knowledge base is billed on a pay-as-you-go basis. For more information, see Knowledge base pricing.
- Text chunks retrieved from the knowledge base increase the number of model input tokens, which may lead to higher model inference (invocation) costs.
-
MCP
- Some official MCP services are billed based on model invocations, such as text-to-image, text-to-video, and speech synthesis MCP services.
- Some MCP services involve third-party API calls, which may incur fees. These fees are charged by the third party, and Alibaba Cloud Model Studio does not charge any additional fees.
Smart Mode Pricing Details
Prices are in CNY per million tokens. AUTO is billed at a unified price regardless of which model the request is actually routed to. When you directly select the Performance / Balanced / Economy tier, billing follows the selected tier:
| Mode | Input | Output |
|---|---|---|
| AUTO | ¥3 | ¥12 |
| Performance | ¥6.4 | ¥22.4 |
| Balanced | ¥1.6 | ¥6.4 |
| Economy | ¥0.5 | ¥2 |
Cache prices for each tier are as follows (also in CNY per million tokens):
| Mode | Implicit Cache Hit | Explicit Cache Creation | Explicit Cache Hit |
|---|---|---|---|
| AUTO | ¥0.75 | — | — |
| Performance | ¥1.6 | — | — |
| Balanced | ¥0.32 | ¥2.5 | ¥0.2 |
| Economy | ¥0.12 | ¥0.75 | ¥0.06 |
FAQ
Can I upgrade a legacy agent to the new version?
No. The legacy agent and the new agent are built on different technical architectures and are not compatible. Direct version switching, upgrading, or downgrading is not possible.
If you are currently using a legacy agent and want to experience the features of the new agent, go to the console and create a new agent application.
Why does the agent not invoke the configured tools as expected?
You can troubleshoot from the following four aspects:
- Skill configuration and binding: Verify that the skill has been successfully created and correctly bound to the current agent application.
- System prompt guidance: Check whether the system prompt clearly describes the skill's functionality, parameters, and applicable scenarios. The model relies on these descriptions to decide when to invoke a skill.
- Intent-skill relevance: Evaluate whether the question is clearly formulated and whether its intent clearly points to a specific skill. If the intent is vague or unrelated to the skill's functionality, the model may choose not to invoke it.
- Execution round limit: Check whether the ReAct round limit has been reached. The agent may have planned to invoke the skill but was forced to terminate before reaching that step because the round limit was exhausted.
Does the agent application support context cache?
It supports implicit cache, but configuring explicit cache within an application is not yet supported.
- Implicit cache: It takes effect automatically when the agent invokes a model that supports implicit cache, requiring no configuration and cannot be disabled. The system automatically identifies and caches common prefixes of requests (such as identical system prompts, multi-turn conversation history, and text chunks retrieved from the knowledge base). Input tokens that hit the cache are billed at 20% of the standard input price, which can reduce model invocation costs accordingly.
- Explicit cache: It requires actively creating cache markers for the specified content in the model invocation request. Because the agent application constructs model requests uniformly on the platform side, configuring explicit cache is not yet supported.
For the working modes, supported models, and billing details of context cache, see Context Cache.