Agent applications (Agent 1.0)
Alibaba Cloud Model Studio agent applications let you connect an LLM to external tools and knowledge bases without code, extending the model's capabilities beyond its built-in limits.
How it works
An agent uses prompts to orchestrate external tools and complete complex tasks. When the agent receives a request, the LLM identifies the user's intent, plans the tasks, calls the necessary external tools, and integrates the results to generate a final response.
Use an agent when you want the model to autonomously decide how to complete a task using available tools, rather than following a manually designed workflow. For example, you can build a customer support assistant that automatically queries a private knowledge base, calls real-time data APIs, and consolidates the results into a unified reply.
Model Studio agents support the following core capabilities:
- Knowledge base (RAG): Connects the agent to an external knowledge base to answer questions based on your private data. This is useful for Q&A scenarios in specialized domains not covered by the model's training data.
- Plugin: Calls pre-built platform tools for tasks such as code execution, image generation, or weather queries. This is useful when the agent must perform actions rather than just converse.
Quick start
Create a basic agent
- Go to the Application Management page in the Model Studio console, click Create Application, switch to the Agent Application tab, and click Create Now.
- In the application configuration interface, select a model from the model dropdown, such as
Qwen-Plus. You can leave the other parameters at their defaults. - After creation, type
Helloin the left-side chat window to test your agent.
Agent capabilities
You can extend an agent's capabilities by selecting a model, optimizing the system prompt, adding a knowledge base (RAG), and calling plugins.
Model
The model is the core component that drives an agent's reasoning and decision-making. Model Studio agents support Qwen series models and custom-deployed models.
- Model selection: Choose a model from the model dropdown, such as
Qwen-Plus. Click More Models to browse additional options. - Parameter configuration: Click the gear icon next to the model dropdown to configure the following parameters:
- Maximum response length: The upper limit on the length of generated content, excluding the prompt. The maximum varies by model.
- Context turns retained: The maximum number of prior conversation turns passed to the model. More turns improve contextual relevance.
- Temperature: Controls output randomness. Higher values increase diversity; lower values increase consistency. The range is [0, 2).
- Reasoning mode: Whether to enable reasoning mode. This parameter cannot be configured for models that do not support reasoning mode.
System prompt
A system prompt defines an agent's role, behavior, and constraints to ensure that its responses are consistent and task-oriented. To write an effective prompt, consider the following points:
-
Define a persona: Specify the role the model should adopt and the expertise it should have.
-
Specify an output format: Describe the desired structure, length, or style of the response.
-
Set constraints: Instruct the model on what content to avoid or which rules to follow.
-
Guide tool usage: Explicitly mention tool names and explain when they should be used.
-
Configure a prompt
Set the system prompt to
Please answer my questions in the style of 'One Hundred Years of Solitude'. The following is a comparison of the results:- Without the system prompt: When the user asks "Who are you?", the model provides its default introduction: "I am Qwen, a large language model developed by the Tongyi Lab of Alibaba Group. I can help you answer questions, provide information, create content, write code, and perform logical reasoning."
- With the system prompt: In the text conversation interface, when the user asks "Who are you?", the AI introduces itself in the literary style of "One Hundred Years of Solitude," referencing elements from the novel such as Macondo and the Buendía family. This confirms the system prompt has taken effect. The interaction statistics show 234 words, 245 input tokens, and 177 output tokens.
-
Knowledge base (RAG)
Retrieval-augmented generation (RAG) lets an agent query an external knowledge base and use the retrieved content as grounding for answers. For private or domain-specific Q&A, RAG can significantly improve accuracy and reduce hallucinations. For more information, see Create and use a knowledge base.
NoteText retrieved from a knowledge base consumes the model's context window. You may need to adjust your retrieval strategy and the length of the retrieved text to use the context window efficiently and avoid exceeding its limit.
Plugins
Agents call plugins to perform specific tasks, such as code execution, web searches, and text-to-image generation. Plugins are useful when an agent needs to perform operations or generate content beyond its native capabilities. Model Studio provides a variety of official plugins and also supports custom plugins. For more information, see Plugin overview.
Agent interaction
Text conversation
Text conversation is the primary method for interacting with an agent and supports multi-turn conversations.
Text conversation supports two input methods:
- Text input: Enter text to chat with the agent.
- File upload: Upload files, such as documents, images, videos, or audio clips, as attachments.
Publish and call
After publishing, agents can be called through API/SDK, published to third-party platforms such as DingTalk and WeChat Official Accounts, or published as reusable components for integration into your business systems.
Publish an application
WarningYou must publish an application before you can call or integrate it.
In the upper-right corner of the agent application management page, click Publish and then click Confirm Publish.
NoteWhen republishing the application, a dialog box appears and displays the changes made since the last publication. If the application was created by a RAM user, make sure you have the ram:CreateServiceLinkedRole permission before publishing. For more information, see service-linked role.
API call
On the Publish Channel tab of your agent application, click View API next to API Call to view the API methods.
NoteReplace YOUR_API_KEY with your actual Model Studio API key before you invoke the API.
Agent management
Copy and delete
On the My Applications page, find the application card. From the More > Copy Application/Delete Application menu, you can copy, delete, or rename the agent.
Common scenarios for copying an application include:
- Creating test versions that use different prompts or models.
- Customizing an agent for different audiences or use cases.
- Creating a backup before making major configuration changes.
Version management
Version management lets you edit historical version descriptions or revert to a previously published version.
-
On the Configure tab of the agent application, click Version Management in the upper-right corner of the top navigation bar.
-
In the list of historical versions, select a target version. The Version History panel opens and displays a timeline that includes the current draft, the online version (marked with a Latest tag), and other historical version entries. Each entry lists the version ID, publication time, publisher, and version information.
-
To edit the version information, hover over the edit icon and click it. In the Edit Version Description dialog box, make your changes and then click OK.
-
To use this historical version, click Overwrite Current Draft, and then click Confirm in the confirmation dialog box.
This action overwrites the current draft with the content of the selected historical version.
-
Billing
Agent billing includes the following:
-
Model calls
Agents incur model call fees, which depend on the model type and token usage.
For billing details, see the Model Studio console.
-
Knowledge base
- Text chunks retrieved from the knowledge base increase the number of input tokens, which can increase model call fees.
-
MCPs
- Some official Model Customization Platforms (MCPs) are billed based on model calls, such as MCPs for text-to-image, text-to-video, and speech synthesis.
- Some MCP services involve third-party API calls, which may incur fees charged by the third-party service provider. Model Studio does not charge additional fees for these calls.
-
Long-term memory
- Data storage for long-term memory is free of charge.
- During a Q&A session, the system merges content from memory into the prompt and passes it to the large language model, which increases token consumption. Tokens consumed by content from memory are currently not charged.
Supported models
NoteThis list may not be up-to-date. For the latest list of supported models, refer to the agent application interface.
- Qwen-Plus
- Qwen-Max
- Qwen-VL-Max
- Qwen-VL-Plus
- Qwen-Turbo
FAQ
How are Model Studio applications billed?
Creating an application is free. When you call an application for a Q&A session, you are charged for model calls based on the model type used.
Is there a timeout limit for custom plugins?
Yes, the timeout limit is 5 seconds.
Can I create agent applications by using an API?
You can use the Assistant API to create large language model applications similar to agent applications. However, applications created with the Assistant API cannot be managed on the console. For more information, see the Assistant API documentation.