Scheduled data synchronization guide for knowledge bases

Updated at:

Best practices for data synchronization rules, helping you understand the synchronization mechanisms of different data sources, choose an appropriate synchronization interval, and answer common questions.

Overview

Data synchronization rules are an automated data synchronization feature provided by the Alibaba Cloud Model Studio platform. By creating synchronization rules, you can automatically synchronize files and documents from external data sources (OSS, Lark, DingTalk, Yuque, and SharePoint) to the Model Studio platform without manual uploads.

Each synchronization rule contains the following core elements:

  • Synchronization category: The category under which synchronized data is stored on the Model Studio platform.
  • Synchronization source: The type of the data source (OSS, Lark, DingTalk, Yuque, or SharePoint).
  • Source credentials: The authentication information required to access the data source, such as the OSS path, the Lark App ID/Secret, or the DingTalk application credentials.
  • Synchronization interval: The frequency of automatic data synchronization (one minute, one hour, or one day).
  • Data tags: Tags configured for the synchronized data to facilitate subsequent classified management and retrieval.

For detailed steps on how to create a synchronization rule, see Dataset — Data synchronization rules.

Entry path

The data synchronization rule feature is located on the file management page of a file connector. Follow these steps to go to the synchronization rule creation page:

  1. Go to the Model Studio console — Dataset page to view the list of created datasets.
  2. In the dataset list, find the target file connector and click Details on the card to go to the file management page.
  3. On the file management page, click the Data Synchronization Rules button in the upper-right corner to open the synchronization rule list.
  4. In the Synchronization Rule List dialog box that appears, click the Create Synchronization Rule button to go to the rule creation form.

After creation, the synchronization rule is displayed in the synchronization rule list. You can enable, disable, or delete rules in the list, and view synchronized files on the file management page.

Synchronization mechanism

Synchronization process

Synchronization rules run automatically according to the configured synchronization interval. During each synchronization, the system connects to the specified data source, checks the status of the files under the source path, and synchronizes new or updated files to the Model Studio platform.

The synchronization process includes the following stages:

  1. Connect: Connect to the data source using the configured credentials.
  2. Scan: Check the file list under the source path and identify new and updated files.
  3. Import: Synchronize new or updated files to the Model Studio platform and trigger parsing and vectorization.
  4. Confirm: After the synchronization is complete, view the synchronization results on the file management page.

NoteSynchronized files are stored as independent copies in the secure storage of the Model Studio platform and are not associated with the original data. Even if the source file is deleted, the copy on the Model Studio platform is retained and must be deleted manually.

Synchronization behavior by source

Different synchronization sources differ in file update detection, synchronization scope, and special limits:

Synchronization source

File update detection

Synchronization scope

Description

OSS

Checks the specified path at each synchronization interval and automatically synchronizes new and updated files.

All files under the specified OSS object path (a directory or a single file).

Access is based on the Alibaba Cloud service-linked role (SLR), so no AccessKey pair is required. You must add the bailian-datahub-access tag to the target bucket.

Lark

Accesses the specified knowledge base, directory, or document at each synchronization interval and synchronizes new and updated content.

The specified Lark knowledge base, directory, or document.

You must create a Lark custom app and configure permissions. MindNote document export is not supported.

DingTalk

Accesses the specified knowledge base, folder, or document at each synchronization interval and synchronizes new and updated content.

The specified DingTalk knowledge base, folder, or document.

Each synchronization consumes DingTalk API quota. We recommend that you evaluate the quota in advance. Supports synchronizing DingTalk documents, tables, and AI tables.

Yuque

Accesses the specified URL list at each synchronization interval and synchronizes new and updated documents.

The specified Yuque document or knowledge base URLs.

Supports filtering by document permission level. Only the public network version of Yuque is supported.

SharePoint

Accesses the specified sharing link at each synchronization interval and synchronizes new and updated documents.

The specified SharePoint document or folder sharing link.

You must register an Azure AD application and configure permissions. Supports both the global public cloud and the China 21Vianet cloud.

Synchronization interval recommendations

The synchronization interval determines how often data is synchronized from the source to the Model Studio platform. Choosing an appropriate synchronization interval requires a balance between data timeliness and resource consumption:

Synchronization interval

Scenario

Advantages

Notes

One minute

Scenarios that require near-real-time updates, such as frequently changing online collaborative documents.

Lowest data update latency.

High synchronization frequency and high resource consumption. For DingTalk synchronization, the API quota is consumed faster.

One hour (default)

Most business scenarios with a moderate data change frequency.

Balances data timeliness and resource consumption. Suitable for the vast majority of scenarios.

Maximum latency of 1 hour.

One day

Low data change frequency, such as periodically updated documents or reports.

Lowest resource consumption and minimal impact on API quotas.

Maximum latency of 1 day. Not suitable for scenarios that require timely updates.

ImportantFor DingTalk synchronization, each synchronization consumes DingTalk Open Platform API quota. The higher the synchronization frequency, the faster the quota is consumed. We recommend that you evaluate whether the quota is sufficient based on the number of documents in the knowledge base, and refer to DingTalk API call consumption for estimation.

FAQ

OSS file updates and incremental synchronization

After a file in OSS is updated, will the Model Studio knowledge base be updated accordingly?

Yes. The synchronization rule runs periodically according to the configured synchronization interval, and checks the status of the files under the specified OSS path during each run. After a file in OSS is added or modified, the updated file is automatically detected and re-synchronized to the Model Studio platform when the next synchronization interval is triggered.

The synchronization interval determines the latency of update detection: If set to one minute, the update takes effect within about 1 minute at most. If set to one hour, the update takes effect within about 1 hour at most.

NoteAfter a file is synchronized to the Model Studio platform, parsing and vectorization are triggered. During peak request hours, this process may take several hours. Please wait patiently.

Does OSS support incremental synchronization?

Yes. During each run, the synchronization rule scans the files under the specified OSS path, identifies new and updated files, and synchronizes them. Existing files that have not changed are not synchronized again, which reduces unnecessary data transfer and processing.

The synchronization scope is determined by the OSS object path configured in the rule:

  • If the path is a directory (such as my-bucket/docs/): All files under the directory are synchronized, and new and updated files are continuously monitored.
  • If the path is a single file (such as my-bucket/docs/foo.md): Only that file is synchronized, and it is automatically re-synchronized after the file is updated.

Behavior when source files are deleted

What happens to the data in the Model Studio knowledge base after the source files are deleted?

Files synchronized to the Model Studio platform through synchronization rules are stored as independent copies and are not associated with the original data. Therefore, even if a source file (such as a file in OSS or a Lark document) is deleted, the synchronized copy on the Model Studio platform is not automatically deleted.

To delete a synchronized file on the Model Studio platform, manually delete the corresponding file on the file management page of the file connector.

NoteOnly files imported within the last 90 days can be viewed. After this time range, imported files cannot be viewed, but they are not deleted.

Synchronization source selection recommendations

What are the differences between the synchronization sources? How should I choose?

When choosing a synchronization source, mainly consider where your data is stored and how it is updated:

  • OSS: Suitable for users who already store a large number of files in Alibaba Cloud OSS. Access is authorized based on the SLR, so no AccessKey pair is required, which is highly secure.
  • Lark: Suitable for users whose enterprise knowledge bases are mainly managed on the Lark platform. You must create a Lark custom app and configure permissions.
  • DingTalk: Suitable for users whose enterprise knowledge bases are mainly managed on the DingTalk platform. You must create a DingTalk application and activate the MCP service. Each synchronization consumes API quota.
  • Yuque: Suitable for users who use Yuque to manage documents and knowledge bases. Only a token and URLs are required, so configuration is simple.
  • SharePoint: Suitable for users who use Microsoft 365 / SharePoint to manage enterprise documents. You must register an Azure AD application. Currently available only to whitelist users.

If you do not need the scheduled synchronization feature, you can also directly upload local files through a file connector. Directly uploaded files do not pass through the public network and are written directly to the secure storage on the Alibaba Cloud internal network. For details, see Import files.

Synchronization troubleshooting

Files fail to synchronize after a synchronization rule is created. How do I troubleshoot?

  1. On the Synchronization Rule List page, check whether the effective status of the rule is "Enabled". If it is "Disabled", click Enable in the Actions column.
  2. On the rule creation or edit page, click Connection Test to verify that the credentials and path are accessible. If the test fails, check the credential information and source path configuration.
  3. Confirm that the OSS object path (or the ID/URL of Lark, DingTalk, Yuque, or SharePoint) is correct. The path must include the bucket name (for OSS) or the complete document link.
  4. Confirm the data source permission configuration: Whether the bailian-datahub-access tag has been added to the OSS bucket, and whether the Lark or DingTalk application permissions are correctly configured.
  5. If the file has been synchronized but cannot be retrieved, parsing may be taking a long time. During peak traffic hours, file parsing may take several hours. Please wait patiently or retry during off-peak hours.

How do I manage synchronization rules (enable, disable, or delete)?

On the file management page of the file connector, click the Data Synchronization Rules button in the upper-right corner to go to the synchronization rule list page. On this page, you can:

  • Enable/Disable: Click "Enable" or "Disable" in the Actions column to control whether the rule runs.
  • Delete: Click "Delete" in the Actions column to remove a synchronization rule that is no longer needed. Deleting a rule does not delete the files that have been synchronized.
  • Search: Enter keywords in the search box to quickly find a specific synchronization rule.