Knowledge base billing

Updated at:

The Alibaba Cloud Model Studio knowledge base service officially starts billing on January 4, 2026. This topic describes the billing methods, cost components, billing examples, and bill management of the knowledge base.

WarningData in knowledge bases created before January 4, 2026 for which the service has not been activated will be retained until June 30, 2026. If the service is not activated by then, the data will be permanently deleted. Activate the knowledge base service promptly.


Billing methods

The knowledge base offers two billing methods: pay-as-you-go and prepaid resource packages. After you activate the knowledge base service, pay-as-you-go is used by default. You can purchase a Standard Edition resource package or an Enterprise Edition resource package in the console.

  • Service activation: Free of charge.
  • Billing start point: Billing starts when a knowledge base is successfully created.
  • Deduction logic:
    • Deduction order: Free quota > resource package > pay-as-you-go. If you run multiple knowledge bases, the free quota or resource package is deducted based on the number of knowledge bases.
    • Overage handling: After your free quota and resource package are exhausted, billing automatically switches to pay-as-you-go.
  • Billing modes:
    • Free quota: When the knowledge base service is activated, the validity period of the free quota (for Standard Edition knowledge bases) starts.
    • Resource package: Resource packages are available in different specifications based on business volume and can be purchased with a one-time payment.
    • Pay-as-you-go: The system automatically deducts fees from your Alibaba Cloud account based on the specification and cumulative runtime, and bills hourly. Make sure your account balance is sufficient (you can top up on the Billing and Costs page).
  • Stop billing: If you no longer need a knowledge base, delete it promptly to stop billing.

WarningDeletion permanently clears all data in the knowledge base and cannot be undone. Proceed with caution.


Cost components

The total cost of a knowledge base consists of two parts: specification fees and model invocation fees.

Specification fees

The specification fee is the runtime fee of the knowledge base. Alibaba Cloud Model Studio offers two knowledge base specifications: Standard Edition and Enterprise Edition.

If you choose to use your own ADB-PG instance as vector storage, additional ADB-PG fees apply.

  • Standard Edition: Suitable for individual or small-scale use and PoC environments.
  • Enterprise Edition: Suitable for high-concurrency, production-grade environments.
SpecificationMax concurrency (knowledge base retrieval)StoragePrice
Standard Edition1 QPS (fixed, not adjustable)Platform storage ≤ 100 GB¥0.03 / knowledge base / hour
Enterprise Edition50–10,000 QPS (adjustable, corresponding to 1–200 RCU)Platform storage ≤ 9,999 GB (for larger storage, select and configure your own ADB-PG instance when creating the knowledge base)¥0.2 / RCU / hour

Note

  • RCU: RCU (Retrieval Compute Unit) measures the concurrent retrieval capacity of a knowledge base. 1 RCU ≈ supports up to 50 QPS for online retrieval. The larger the RCU, the higher the supported concurrency.
  • How to estimate the required RCU: Required RCU = round up (peak retrieval QPS requirement ÷ 50). For example, a peak demand of 80 QPS requires at least 2 RCU.
  • Max concurrency: Refers to the core retrieval performance of the knowledge base itself (excluding dependent links, such as calls to the rerank model).

    During retrieval from a knowledge base (Enterprise Edition), if model dependency links (such as Embedding and Rerank models) are throttled in extreme cases, we will scale up the related services as quickly as possible, but a small number of retrieval requests may be degraded in the meantime, with a brief decline in performance.

  • Storage fees: The prices above already include the cost of platform storage. If you choose to use your own ADB-PG, additional fees apply. The price is subject to the ADB-PG pricing page.
  • Specification change: Billed in segments based on when the change takes effect. The change operation itself is free of charge. A knowledge base can be modified at most once per calendar day.

Free quota

Alibaba Cloud Model Studio provides all users with a one-time 720-hour free quota for knowledge bases. After the free quota is exhausted, pay-as-you-go billing applies.

Warning

  • For existing users, the free quota is valid until 23:59 on February 3, 2026. After that, pay-as-you-go billing applies automatically.
  • For new users, the free quota is valid for 30 days from the activation date. Any remaining quota expires and can no longer be used.

Existing users are those who activated the service before knowledge base billing officially started on January 4, 2026. New users are those who activate the service for the first time on or after that date.

On the knowledge base page, click View Bill in the upper-right corner to view the remaining free hours and validity period.

Usage rules

  • Scope: Only offsets the specification fees of Standard Edition knowledge bases. It does not apply to the Enterprise Edition.
  • Deduction method: Deducted cumulatively based on actual runtime. If you run multiple knowledge bases, the free quota is deducted based on the number of knowledge bases.

    For example, if you run 4 Standard Edition knowledge bases at the same time, 4 hours of quota are deducted every hour.

  • Exclusions: Model invocation fees are not covered by the free quota and follow the billing policy of each model.

Examples

  • One Standard Edition knowledge base: 720 hours ÷ 1 = 720 hours of free runtime
  • Two Standard Edition knowledge bases running at the same time: 720 hours ÷ 2 = 360 hours of free runtime

Resource packages

A resource package is valid for one year after purchase. Unused hours in the package expire after one year.

Standard Edition resource package specifications and prices

UnitInstances × hoursUse casePackage price (CNY)
1 instance / month7201 knowledge base for 1 month20
1 instance / quarter2,1601 knowledge base for 1 quarter59
1 instance / year8,7601 knowledge base for 1 year239
10 instances / year87,60010 knowledge bases for 1 year2,099
50 instances / year438,00050 knowledge bases for 1 year9,999
100 instances / year876,000100 knowledge bases for 1 year18,999

Enterprise Edition resource package specifications and prices

UnitRCU × hoursUse casePackage price (CNY)
1 RCU / month7201 knowledge base for 1 month139
1 RCU / quarter2,1601 knowledge base for 1 quarter399
1 RCU / year8,7601 knowledge base for 1 year1,599
10 RCU / year87,60010 knowledge bases for 1 year14,999
30 RCU / year262,80030 knowledge bases for 1 year41,999
50 RCU / year438,00050 knowledge bases for 1 year65,999

Usage notes

  • Effective time: The resource package takes effect automatically after purchase.
  • Validity period: Determined by the purchased package. Remaining hours in the package expire automatically after the validity period.
  • Deduction logic:
    • Deduction order: Free quota > resource package > pay-as-you-go.
    • Multiple resource packages of the same type: The package that expires first is deducted first. If the expiration time is the same, the package purchased first is deducted first.
    • If you run multiple knowledge bases, the free quota or resource package is deducted based on the number of knowledge bases.
    • Overage handling: If all resource packages of the same type expire or are fully deducted, the excess is automatically billed on a pay-as-you-go basis.
  • Balance monitoring and alerts:
    • Check balance: Click Resource Packages to view the remaining balance, and click Statistics to view usage information.
    • Set alerts: We recommend that you set budget alerts. When the resource package usage falls below the preset threshold, the system automatically sends notifications by text message, email, and internal message.
  • Unsubscription: According to the Alibaba Cloud unsubscription rules, you can request a refund for the unused portion of a prepaid product based on the unused quota fee. The used portion cannot be refunded. After unsubscription, the knowledge base switches to pay-as-you-go. Delete the knowledge base to stop billing.

Model invocation fees

When you create, update, or retrieve from a knowledge base, or use the knowledge Q&A service, the following models are called. These calls incur model invocation fees independent of the specification fees:

Model categoryModel namePurpose
Embedding modelstext-embedding-v4, etc.Text vectorization for document knowledge bases
qwen3-vl-embeddingMultimodal vectorization for image Q&A and audio/video search knowledge bases
Rerank modelsqwen3-rerankSecondary ranking of retrieval results for document knowledge bases (optional)
qwen3-vl-rerankSecondary ranking of retrieval results for image Q&A and audio/video search knowledge bases (optional)
Routing modelqwen-plusWhen knowledge base routing is enabled, the system calls qwen-plus to determine which knowledge bases the query should be routed to
Q&A modelqwen3.7-plus, etc.The large language model that generates answers in the knowledge Q&A service, selected by the user in the application

WarningModel invocation fees are independent billable items, calculated based on the actual input token volume. The prices and free quota policies follow the invocation billing standards of the corresponding models in the model marketplace, and are not included in the specification fees of the knowledge base.

Billing formula: Model fee = (total input tokens / 1000) × model unit price (CNY / 1,000 tokens)

Billing rule for multiple knowledge bases: When an Alibaba Cloud Model Studio application is mounted with multiple knowledge bases, retrieval is performed in each of them. Token consumption (Query vectorization and Rerank ranking) increases by a multiple of the number of knowledge bases (N knowledge bases means consumption × N).

Knowledge management (create and update knowledge bases)

  • Invocation scenario: When you upload new files or perform incremental updates, the system calls an embedding model to vectorize the text content.
  • Billing: Billed based on the number of tokens in the newly added content. Deleting files does not incur model invocation fees.
  • Models called:

Knowledge retrieval

  • Invocation scenarios:

    1. Vectorization: The system calls an embedding model to vectorize the user query.
    2. Knowledge base routing (optional): If the application is associated with multiple knowledge bases and knowledge base routing is enabled, the system calls qwen-plus to determine which knowledge bases the user query should be routed to. This call is billed based on the token usage of qwen-plus.
    3. Reranking (optional): The system calls a rerank model to rerank the initially retrieved results to improve the accuracy of the final answer. Document search knowledge bases use qwen3-rerank. Image Q&A and audio/video search knowledge bases use qwen3-vl-rerank.
  • Billing:

    • Query vectorization fee: Billed based on the number of tokens in the user input.
    • Rerank fee (can be disabled): This is the main part of the retrieval fee. The fee depends on the total number of chunks in the initial recall.

Retrieval process and billing relationship

User inputs a query
    │
    ▼
Query vectorization (embedding model, billed by input tokens)
    │
    ▼
Initial recall (based on initial vector retrieval TopK + initial keyword retrieval TopK)
    │
    ▼
Reranking (rerank model, billed by tokens of all initially recalled chunks)
    │
    ▼
Final recall (returns the number of chunks specified by the max recall count parameter)
  1. Initial recall: The system recalls text chunks from the knowledge base based on the following parameters:

    • Initial vector retrieval TopK: Controls the number of relevant chunks recalled based on semantic similarity (default: 50)
    • Initial keyword retrieval TopK: Controls the number of relevant chunks recalled based on exact text matching (default: 50)
  2. Reranking:

    1. All initially recalled chunks are sent to the rerank model for ranking.
    2. Fee = total initially recalled chunks × average tokens per chunk × model unit price

    WarningThe rerank model fee depends on the total number of initially recalled chunks, not the number of chunks returned in the final recall.

  3. Final recall: After the rerank model ranks the chunks, the system returns the number of chunks specified by the max recall count parameter (for example, 5).

NoteClicking Hit Test on the knowledge base card to enter the configuration and debugging page incurs invocation fees for the corresponding models (embedding model and rerank model).

Knowledge Q&A

When you use knowledge bases for Q&A through Alibaba Cloud Model Studio applications (agent applications and workflow applications), the following model invocation fees are incurred in addition to the model fees in the retrieval stage:

  • Q&A generation model: The system generates answers based on the Q&A model you select in the application (such as qwen-plus), billed based on the token usage of that model.
  • Pre-file parsing (optional): When a user uploads a file in a conversation and enables the pre-file parsing feature, the system calls qwen3-rerank to rank the file content, billed based on the token usage of the rerank model.
  • Knowledge base routing (optional): If the application is associated with multiple knowledge bases and routing is enabled, the system calls qwen-plus to make routing decisions.

WarningTotal fee of the knowledge Q&A service = specification fee (knowledge base runtime) + model fees in the retrieval stage (vectorization + reranking + routing) + model fees in the Q&A stage (answer generation + pre-file parsing). Each model fee is calculated independently based on the actual token consumption.

Cost optimization recommendations

Disable rerankingAdjust initial recall parameters
If your application scenario does not require high search precision, you can disable the reranking feature to eliminate rerank model fees.By lowering the values of Initial vector retrieval TopK and Initial keyword retrieval TopK, you can reduce the number of tokens sent to the rerank model, significantly lowering costs.
Impact: This operation reduces the relevance ranking of search results.Impact: This operation may affect the final retrieval performance. You can adjust it yourself to balance cost and performance.
Actions:Actions:
- Legacy agent and workflow applications: Click the Debug button to the right of the knowledge base in the application to enter the page, and turn off the Re-ranking Strategy switch.Adjust on the Edit or Hit Test page of the knowledge base, and click Save.
- New agent applications: Click Hit Test on the knowledge base card, select Do Not Use Model, and click Save.For example, set both Initial vector retrieval TopK and Initial keyword retrieval TopK to 50 (value range: 10–100), and then click Save.

NoteClicking Hit Test on the knowledge base card to enter the configuration and debugging page incurs invocation fees for the corresponding models (embedding model and rerank model).

Savings plan deduction

The embedding models (such as text-embedding-v4) and rerank models (such as qwen3-rerank) used by knowledge bases are Class A models on the Alibaba Cloud Model Studio platform. Their invocation fees can be offset by the following savings plans:

  • AI general-purpose savings plan (recommended): Covers all Class A models (including text embedding, multimodal embedding, and rerank models), with tiered discounts for monthly committed spending.
  • Embedding and rerank model savings plan: A savings plan dedicated to embedding and rerank models, purchased as a one-time fixed amount.

NoteSavings plans can only offset model invocation fees. They cannot offset the specification fees (runtime fees) of knowledge bases. To optimize specification fees, see the Resource packages section above.


Billing examples

Continuous runtime for 1 day

SpecificationConfigurationQuantityDaily specification fee
Standard EditionPlatform storage124 hours × ¥0.03 / hour = ¥0.72
Enterprise EditionPlatform storage, 1 RCU124 hours × 1 RCU × ¥0.2 / RCU / hour = ¥4.80

Create, update, and retrieve from knowledge bases

The following examples are based on a document search knowledge base using text-embedding-v4 (embedding model) and qwen3-rerank (rerank model), both priced at ¥0.0005 / 1,000 tokens.

Billing logic: Fee = token consumption (in units of "1,000 tokens") × model unit price

Create a knowledge base

  • Operation: Upload a file containing 50,000 tokens for vectorization.
  • Fee: 50 × ¥0.0005 / 1,000 tokens = ¥0.025

Update a knowledge base

  • Operation: Add a file containing 20,000 tokens.
  • Fee: 20 × ¥0.0005 / 1,000 tokens = ¥0.01

Retrieve from a knowledge base (single)

  • Operation: Enter a 100-token query, which recalls 150 relevant chunks (500 tokens per chunk on average) for reranking.
  • Fee:
    • Query vectorization: 0.1 × ¥0.0005 / 1,000 tokens = ¥0.00005
    • Tokens for reranking: 150 chunks × 500 tokens / chunk = 75,000 tokens
    • Rerank fee (if any): 75 × ¥0.0005 / 1,000 tokens = ¥0.0375
    • Total: ¥0.00005 (query vectorization) + ¥0.0375 (reranking) = ¥0.03755

Retrieve from knowledge bases (multiple)

  • Operation: An Alibaba Cloud Model Studio agent application is associated with 4 knowledge bases. The same query is executed once in each knowledge base by default (this cannot be changed).
  • Fee: ¥0.03755 / request × 4 = ¥0.1502

Specification change (prorated billing)

  • Scenario: During 14:40–15:40, the knowledge base is upgraded from Standard Edition to Enterprise Edition (2 RCU) at 15:10. Both the Standard Edition and the Enterprise Edition run for 30 minutes (that is, 0.50 hours, rounded to 2 decimal places).
  • Specification fee (during 14:40–15:40):
    • Standard Edition: 0.50 hours × ¥0.03 / hour = ¥0.015
    • Enterprise Edition: 0.50 hours × 2 RCU × ¥0.2 / RCU / hour = ¥0.20
    • Total: ¥0.215

Runtime less than 1 hour

  • Scenario: A Standard Edition knowledge base is created at 14:12 and deleted at 14:21, for a total runtime of 9 minutes (0.15 hours, rounded to 2 decimal places).
  • Specification fee (during 14:12–14:21): 0.15 hours × ¥0.03 / hour = ¥0.0045

Fees and bills

View bills and usage

Query the specification fees of a knowledge base

Export the bill on the Bill Details page. In the bill (aggregated by hour), you can view the specification fee of the specified knowledge base in the corresponding period (the List Total Price column).

The instance ID in the figure is the knowledge base ID.

Filter Product Name to Large Model Service Platform Alibaba Cloud Bailian and Commodity Name to Alibaba Cloud Bailian Knowledge Base (RAG) - Pay-as-you-go to view the billable items of each knowledge base instance (such as Standard Edition - Computing Resources) and the corresponding list price (such as ¥0.03 / (instance × hour)).

Query token consumption and corresponding amounts in the detailed bill

Export the bill on the Bill Details page. In the bill (aggregated by hour), you can view the token usage (the Usage column) and the corresponding amount (the List Total Price column) for the corresponding period.

View embedding model usage

Hover over the instance ID on the bill: If the instance ID is in the format llm-xxx;xxx-embedding-xxx;embedding_token;RAG;0, the bill is generated by an embedding model.

On the Bill Details page, set Billing Date to Monthly, filter Product Name to Large Model Service Platform Alibaba Cloud Bailian, filter Commodity Name to Alibaba Cloud Bailian Large Model Inference, and select Yes for Include Items with Payable Amount of 0. The bill details table shows token consumption at an hourly granularity, including the service start and end time, billable item name, usage, usage unit (1,000 tokens), official list price, and list total price.

View rerank model usage

Hover over the instance ID on the bill: If the instance ID is in the format llm-xxx;xxx-rerank;embedding_token;RAG;0, the bill is generated by a rerank model.

On the Bill Details page, filter Product Name to Large Model Service Platform Alibaba Cloud Bailian and Commodity Name to Alibaba Cloud Bailian Large Model Inference. In the bill details table, the Billable Item Name is "Text Embedding Usage". You can view the usage (unit: 1,000 tokens) and the corresponding list total price for each period.

Cost allocation

If you need to attribute costs to different departments or projects, you can use the tag feature to mark workspaces.

Step 1: Obtain workspace information

In Workspace Management, determine the Workspace ID of the workspace to which the tag is bound (example: llm-xxx).

Step 2: Bind tags

  1. On the Tag Management page, select Bind Tags to Resources.
  2. For the resource selection method, select Enter Multiple Resource IDs. On the Product tab, search for and select Large Model Service Platform Alibaba Cloud Bailian: Workspace, select the region of the workspace, enter the Workspace ID in the resource ID input box, and then click the Bind Tags button.
  3. On the Bind Tags page, create a tag key-value pair or use an existing preset tag to bind to the workspace. After entering the key-value pair or selecting a preset tag, click Confirm to complete the tag binding for the workspace.

After the operation is complete, a Bind Resources confirmation dialog box appears, showing the resource ID, operation status, and failure reason of each resource in a table. After confirming that everything is correct, click Got It to close the dialog box.

Step 3: Verify

You have now bound tags to your Alibaba Cloud Model Studio workspace. You can verify and query the tags bound to the workspace in the Instance Tag column on the Bill Details page.

NoteNewly created instance tags have a certain delay (at the hour level).

When filtering, you can set Product Name to Large Model Service Platform Alibaba Cloud Bailian and Commodity Name to Alibaba Cloud Bailian Knowledge Base (RAG) - Pay-as-you-go.

Overdue payments

After your Alibaba Cloud account has an overdue payment, all its knowledge bases enter the Service Suspended state (you cannot retrieve, update, or create knowledge bases through the console or API), and billing stops.

Vector storage using platform storage

StageDescription
0–14 daysYou cannot retrieve, update, or create knowledge bases through the console or API, but existing data is retained. After you pay all overdue bills within the first 14 days, the service automatically returns to normal.
≥ 15 daysOn the 15th day after the overdue payment, you are deemed to have voluntarily given up the pay-as-you-go knowledge base service. Alibaba Cloud Model Studio will release the related knowledge bases and permanently delete their data, which cannot be recovered.

Vector storage using self-purchased ADB-PG

StageDescription
0–7 daysYou cannot retrieve, update, or create knowledge bases through the console or API, but existing data is retained. After you pay all overdue bills within the first 7 days, the service automatically returns to normal.
≥ 8 daysOn the 8th day after the overdue payment, you are deemed to have voluntarily given up the pay-as-you-go ADB-PG service. ADB-PG will clean up the instances related to the knowledge bases and permanently delete their data, which cannot be recovered.

NoteWhen you use self-purchased ADB-PG, the data retention period follows the overdue payment policy of ADB-PG, which is 7 days (not 14 days).

Refunds

Pay-as-you-go generates bills based on the specification of the knowledge base and the actual usage duration, so no refund is involved.


FAQ

Can a RAM user activate a knowledge base or view bills?

Yes. An authorized RAM user (with the AliyunBailianFullAccess or AliyunSFMFullAccess system policy) can activate a knowledge base. The fees are attributed to the Alibaba Cloud account.

What does "free storage" for the Standard Edition and Enterprise Edition specifically mean?

It only means that platform storage is free. Self-purchased ADB-PG is billed by the ADB-PG service and is not included in the knowledge base bill.

What should I do if my knowledge base has a large amount of data and the platform storage of the Enterprise Edition is insufficient?

When creating a knowledge base, you can choose to use your own ADB-PG instance as vector storage. For specific configuration methods, see the Create a knowledge base section.

How are specification changes billed across hours?

Billing is segmented based on when the change occurs, and charges within the same hour are accumulated based on the proportion of each time segment. For an example, see Specification change (prorated billing) above.

Why is my rerank fee particularly high? How can I reduce model invocation fees?

The fee of the rerank model is not related to the number of results you finally receive. Instead, it is determined by the total number of text chunks in the initial recall. To reduce model invocation fees, see the Cost optimization recommendations section above.

How do I completely stop billing for a knowledge base? Can I delete the files in it?

No. The only way to stop billing is to delete the entire knowledge base instance.

  • Incorrect operation: Deleting only the files in the knowledge base only clears the data, but the knowledge base instance (as the billing entity) is still running, so specification fees continue to be incurred.
  • Correct operation: Find the corresponding knowledge base instance in the console and delete it.

WarningDeletion permanently clears all data in the knowledge base and cannot be undone. Proceed with caution.

Why is the number of rerank model invocations greater than the number of application calls?

This is an automatic optimization performed by the system to improve performance. When a single request sent to the rerank model contains a large number of chunks, the system splits it into multiple batches to call the rerank model to speed up processing.

This increases the recorded number of rerank model invocations, but the total fee remains unchanged, because billing is only related to the total token consumption, not the number of invocations.