Parse documents with DTS-AI
This topic describes how to use the DTS-AI service API to parse documents in formats such as PDF and Word into Markdown. It also provides complete examples for four calling methods: curl, SDK, Alibaba Cloud CLI, and DTS Skill.
Overview
The DTS-AI service provides a document parsing feature that converts documents in formats such as PDF, Word, PPT, plain text, Markdown, and images into structured Markdown content.
Document parsing supports synchronous and asynchronous calling methods. The synchronous method is simpler — you just set the ResponseMode parameter to sync.
This topic uses the asynchronous approach as the main example. For the synchronous approach, see Method 3: Use the SDK for synchronous calls (AccessKey authentication). The overall asynchronous workflow is as follows:
Create a parsing job: Call the
CreateDocParserJobAPI to submit a document parsing job and get a job ID (JobId).Query the job status: Call the
DescribeDocParserJobStatusAPI to poll the job status until it becomessuccess.Get the parsing result: Call the
DescribeDocParserJobResultAPI to get the parsed Markdown content.
Quick trial
If you want to test the document parsing capability before integrating the SDK or CLI, you can call the DTS-AI service APIs directly from the Alibaba Cloud OpenAPI Developer Portal:
Create a document parsing job: Try CreateDocParserJob.
Query the job status: Try DescribeDocParserJobStatus.
Get the parsing result: Try DescribeDocParserJobResult.
After you sign in and grant permissions, enter the request parameters on the left side of the page and click Debug to view the response. Once your test succeeds, refer to the SDK or CLI examples below to integrate the API into your application.
Prerequisites
Before you call the DTS-AI service APIs, complete the following preparations.
Prepare authentication credentials
The DTS-AI service supports two types of authentication credentials. Choose either one. API Key is recommended because its permissions are restricted to the DTS service scope, resulting in more controllable security risks.
API Key (recommended): An API Key is a service credential issued by RAM for a single cloud service. It can only call the OpenAPIs of the bound cloud service, resulting in a smaller permission scope. In the RAM console, locate the target user, go to the Credential Management section, click Create API Key, and set Cloud Service to Data Transmission Service / DTS. After the API Key is created, copy the plaintext credential immediately and store it securely (the plaintext is unavailable after the dialog is closed). For more information, see Create an API key.
ImportantAlibaba Cloud accounts (root accounts) cannot create API Keys. Only RAM users can create them. Each RAM user can create up to 2 API Keys per cloud service.
An API Key is a credential used to access cloud services. If it is leaked, anyone can call your APIs, resulting in unpredictable costs and security risks. Keep your API Key secure and do not write it in plaintext in code, commit it to code repositories, or publish it to public platforms. We recommend that you store it in environment variables. For more information, see the "" section below.
AccessKey ID and AccessKey Secret: Suitable for scenarios where you want to reuse existing Alibaba Cloud AccessKey credentials. For security, we recommend using the AccessKey of a RAM user and configuring a permission policy based on the principle of least privilege. For more information, see Authorize a RAM user to manage DTS.
The following examples read credentials from environment variables:
API Key: read from
DTS_AI_API_KEY.AccessKey: read from
ALIBABA_CLOUD_ACCESS_KEY_IDandALIBABA_CLOUD_ACCESS_KEY_SECRET.
Before you run the examples, set the corresponding environment variables in your operating system to avoid hard-coding credentials in code. For secure credential usage, see Manage credentials.
Prepare a calling tool (choose as needed)
Use curl: No additional installation is required. You only need an API Key to initiate calls. Suitable for quick API verification or lightweight integration scenarios.
Use the SDK: You have installed the DtsAI SDK. You can obtain SDKs for various languages from the SDK download page.
Use the Alibaba Cloud CLI: You have installed and configured the Alibaba Cloud CLI. For more information, see Install, update, and uninstall Alibaba Cloud CLI and Configure and manage credentials.
Usage notes
The document parsing API is currently available only in the China (Beijing) region. Set the RegionId to
cn-beijing.The FileUrl parameter must be a valid OSS URL. If your file is not stored in OSS, we recommend using the
CreateDocParserJobAdvancemethod of the SDK to pass a local file stream directly.The output format (OutputFormat) currently supports only
markdown.The parsing result is retained for 72 hours. Download and save it promptly, as it cannot be retrieved after this period.
Supported file formats include PDF, DOCX, DOC, PPTX, PPT, TXT, Markdown, PNG, JPG, and JPEG.
The SDKs for different languages provide the
CreateDocParserJobAdvancemethod, which lets you pass a local file stream directly without first uploading the file to OSS.
Method 1: Use curl (recommended, API Key authentication)
Pass the API Key through the X-Acs-ApiKey request header to call the DTS-AI service APIs directly. No signature computation is required, which is suitable for quick verification or lightweight integration scenarios.
Step 1: Configure the API Key environment variable
Configure the API Key as an environment variable to avoid exposing the plaintext in commands.
export DTS_AI_API_KEY='<your API Key>'Step 2: Create a document parsing job
Call the CreateDocParserJob operation to create a document parsing job.
curl -X POST \
"https://dtsai.cn-beijing.aliyuncs.com?Action=CreateDocParserJob&Version=2026-04-01&SignatureNonce=$(uuidgen)" \
--header "X-Acs-ApiKey: ${DTS_AI_API_KEY}" \
--header "Content-Type: application/json" \
--header "Accept: application/json" \
--data-raw '{
"RegionId": "cn-beijing",
"FileUrl": "https://<BucketName>.oss-cn-beijing.aliyuncs.com/document.pdf",
"FileName": "document.pdf",
"FileFormat": "pdf",
"OutputFormat": "markdown"
}'Sample response:
{
"RequestId": "019F6482-DD04-510E-A212-A21B3C59CB8C",
"Success": true,
"HttpStatusCode": 200,
"JobId": "job_abc123"
}Note the JobId from the response. Use this ID in the following steps to query the job status and get the parsing result.
Step 3: Query the job status
Call the DescribeDocParserJobStatus operation to poll the job status until it returns success.
curl -X POST \
"https://dtsai.cn-beijing.aliyuncs.com?Action=DescribeDocParserJobStatus&Version=2026-04-01&SignatureNonce=$(uuidgen)" \
--header "X-Acs-ApiKey: ${DTS_AI_API_KEY}" \
--header "Content-Type: application/json" \
--data-raw '{
"RegionId": "cn-beijing",
"JobId": "job_abc123"
}'Step 4: Get the parsing result
After the job status is success, call the DescribeDocParserJobResult operation to get the parsing result.
curl -X POST \
"https://dtsai.cn-beijing.aliyuncs.com?Action=DescribeDocParserJobResult&Version=2026-04-01&SignatureNonce=$(uuidgen)" \
--header "X-Acs-ApiKey: ${DTS_AI_API_KEY}" \
--header "Content-Type: application/json" \
--data-raw '{
"RegionId": "cn-beijing",
"JobId": "job_abc123"
}'The Result field in the response contains the parsed content in Markdown format.
Method 2: Use the SDK for asynchronous calls (AccessKey authentication)
The following example uses the Java SDK to demonstrate the document parsing workflow. The process is similar for SDKs in other languages. For more information, see the SDK download page.
Step 1: Install the SDK
Add the following dependency to the pom.xml file of your Maven project:
<dependency>
<groupId>com.aliyun</groupId>
<artifactId>dtsai20260401</artifactId>
<version>1.0.0</version>
</dependency>Step 2: Create a document parsing job
Call the CreateDocParserJob API to create a document parsing job. If the file is stored locally, you can use the CreateDocParserJobAdvance method to pass the file stream directly.
File URL
import com.aliyun.dtsai20260401.Client;
import com.aliyun.dtsai20260401.models.*;
import com.aliyun.teaopenapi.models.Config;
public class DocParserExample {
public static void main(String[] args) throws Exception {
// Initialize the client.
Config config = new Config()
.setAccessKeyId(System.getenv("ALIBABA_CLOUD_ACCESS_KEY_ID"))
.setAccessKeySecret(System.getenv("ALIBABA_CLOUD_ACCESS_KEY_SECRET"));
config.endpoint = "dtsai.cn-beijing.aliyuncs.com";
Client client = new Client(config);
// Create a parsing job.
CreateDocParserJobRequest request = new CreateDocParserJobRequest()
.setRegionId("cn-beijing")
.setFileUrl("https://example.oss-cn-beijing.aliyuncs.com/document.pdf")
.setFileName("document.pdf")
.setFileFormat("pdf")
.setOutputFormat("markdown");
CreateDocParserJobResponse response = client.createDocParserJob(request);
String jobId = response.getBody().getJobId();
System.out.println("Job created successfully. JobId: " + jobId);
}
}Local file stream
import com.aliyun.dtsai20260401.Client;
import com.aliyun.dtsai20260401.models.*;
import com.aliyun.teaopenapi.models.Config;
import java.io.FileInputStream;
public class DocParserAdvanceExample {
public static void main(String[] args) throws Exception {
Config config = new Config()
.setAccessKeyId(System.getenv("ALIBABA_CLOUD_ACCESS_KEY_ID"))
.setAccessKeySecret(System.getenv("ALIBABA_CLOUD_ACCESS_KEY_SECRET"));
config.endpoint = "dtsai.cn-beijing.aliyuncs.com";
Client client = new Client(config);
// Use the Advance method to directly pass a local file stream.
CreateDocParserJobAdvanceRequest request = new CreateDocParserJobAdvanceRequest()
.setRegionId("cn-beijing")
.setFileUrlObject(new FileInputStream("/path/to/document.pdf"))
.setFileName("document.pdf")
.setFileFormat("pdf")
.setOutputFormat("markdown");
CreateDocParserJobResponse response = client.createDocParserJobAdvance(request,
new com.aliyun.teautil.models.RuntimeOptions());
String jobId = response.getBody().getJobId();
System.out.println("Job created successfully. JobId: " + jobId);
}
}Step 3: Query the job status
Call the DescribeDocParserJobStatus operation to poll the job status until it returns success.
// Poll the job status
DescribeDocParserJobStatusRequest statusRequest = new DescribeDocParserJobStatusRequest()
.setRegionId("cn-beijing")
.setJobId(jobId);
String status = "";
while (!"success".equals(status) && !"failed".equals(status)) {
Thread.sleep(5000); // Query every 5 seconds
DescribeDocParserJobStatusResponse statusResponse =
client.describeDocParserJobStatus(statusRequest);
status = statusResponse.getBody().getStatus();
System.out.println("Current status: " + status);
if ("failed".equals(status)) {
System.out.println("Job failed: " + statusResponse.getBody().getFailureMessage());
return;
}
}The following table describes the job status values.
Status value | Description |
| Created. The job is being prepared. |
| Queued and waiting for execution. |
| Processing. The document is being parsed. |
| Parsing is complete. Call DescribeDocParserJobResult to get the result. |
| Parsing failed. The failure message is available in the FailureMessage parameter. |
| The job is canceled. |
Step 4: Get the parsing result
After the job status changes to success, call the DescribeDocParserJobResult API to get the parsing result.
// Get the parsing result
DescribeDocParserJobResultRequest resultRequest = new DescribeDocParserJobResultRequest()
.setRegionId("cn-beijing")
.setJobId(jobId);
DescribeDocParserJobResultResponse resultResponse =
client.describeDocParserJobResult(resultRequest);
String markdownContent = resultResponse.getBody().getResult();
System.out.println("Parsing result:\n" + markdownContent);The parsing result is retained for only 72 hours. Get and save the result promptly after the job is complete.
Method 3: Use the SDK for synchronous calls (AccessKey authentication)
The following example uses the Java SDK to demonstrate the complete document parsing workflow. For usage of other languages, see SDK download page.
Step 1: Install the SDK
Add the following dependency to the pom.xml file of your Maven project:
<dependency>
<groupId>com.aliyun</groupId>
<artifactId>dtsai20260401</artifactId>
<version>1.0.0</version>
</dependency>Step 2: Create a synchronous document parsing job
Call the CreateDocParserJob API to create a document parsing job. To parse a local file, use the CreateDocParserJobAdvance method to pass the file stream directly.
Create a parsing job with a file URL
import com.aliyun.dtsai20260401.Client;
import com.aliyun.dtsai20260401.models.*;
import com.aliyun.teaopenapi.models.Config;
public class DocParserExample {
public static void main(String[] args) throws Exception {
// Initialize client
Config config = new Config()
.setAccessKeyId(System.getenv("ALIBABA_CLOUD_ACCESS_KEY_ID"))
.setAccessKeySecret(System.getenv("ALIBABA_CLOUD_ACCESS_KEY_SECRET"));
config.endpoint = "dtsai.cn-beijing.aliyuncs.com";
Client client = new Client(config);
// Create a parsing job
CreateDocParserJobRequest request = new CreateDocParserJobRequest()
.setRegionId("cn-beijing")
.setFileUrl("https://example.oss-cn-beijing.aliyuncs.com/document.pdf")
.setFileName("document.pdf")
.setFileFormat("pdf")
.setOutputFormat("markdown")
.setResultType("content")
.setResponseMode("sync");
CreateDocParserJobResponse response = client.createDocParserJob(request);
String jobId = response.getBody().getJobId();
System.out.println("Job created successfully. JobId: " + jobId);
}
}Create a parsing job with a local file stream (Advance method)
import com.aliyun.dtsai20260401.Client;
import com.aliyun.dtsai20260401.models.*;
import com.aliyun.teaopenapi.models.Config;
import java.io.FileInputStream;
public class DocParserAdvanceExample {
public static void main(String[] args) throws Exception {
Config config = new Config()
.setAccessKeyId(System.getenv("ALIBABA_CLOUD_ACCESS_KEY_ID"))
.setAccessKeySecret(System.getenv("ALIBABA_CLOUD_ACCESS_KEY_SECRET"));
config.endpoint = "dtsai.cn-beijing.aliyuncs.com";
Client client = new Client(config);
// Use the Advance method to pass a local file stream directly
CreateDocParserJobAdvanceRequest request = new CreateDocParserJobAdvanceRequest()
.setRegionId("cn-beijing")
.setFileUrlObject(new FileInputStream("/path/to/document.pdf"))
.setFileName("document.pdf")
.setFileFormat("pdf")
.setOutputFormat("markdown")
.setResultType("content")
.setResponseMode("sync");
CreateDocParserJobResponse response = client.createDocParserJobAdvance(request,
new com.aliyun.teautil.models.RuntimeOptions());
String jobId = response.getBody().getJobId();
System.out.println("Job created successfully. JobId: " + jobId);
}
}The returned result is as follows:
{
"Status": "success",
"RequestId": "01A055D1-178A-5292-9230-****",
"HttpStatusCode": 200,
"ResultType": "content",
"Success": true,
"JobId": "01a055d1-17bd-7ed2-ba59-****",
"Result": "<!-- page 1 -->\n\n# DTS FORMULA MARKDOWN 2026\n\n$$\nx = \\frac{-b \\pm \\sqrt{b^2 - 4ac}}{2a}\n$$"
}Method 4: Use the Alibaba Cloud CLI
This example shows how to complete the document parsing process using the Alibaba Cloud CLI. Before you start, complete the following preparations:
Upgrade the Alibaba Cloud CLI to the latest version and install the
aliyun-cli-dtsaiplug-in by running the following command:aliyun plugin install --names aliyun-cli-dtsaiConfigure the identity credentials for the Alibaba Cloud CLI. OAuth mode is recommended (no need to maintain AccessKey): run
aliyun configure --mode OAuthand follow the prompts to sign in. You can also use other authentication modes such as AccessKey or STS Token. For more information, see Configure and manage credentials.
DTS-AI service CLI commands are provided by the aliyun-cli-dtsai plug-in. All commands and parameters use kebab-case (hyphen-separated), for example aliyun dtsai web-search --biz-region-id cn-beijing --query "...".
Step 1: Create a document parsing job
Run the following command to call the create-doc-parser-job command and create a document parsing job.
aliyun dtsai create-doc-parser-job \
--endpoint dtsai.cn-beijing.aliyuncs.com \
--biz-region-id cn-beijing \
--file-url "https://<BucketName>.oss-cn-beijing.aliyuncs.com/document.pdf" \
--file-name document.pdf \
--file-format pdf \
--output-format markdownSample response:
{
"HttpStatusCode": 200,
"JobId": "019f64f2-565d-7d70-95d0-3b97d784a113",
"RequestId": "019F64F2-561D-522C-8F56-71D8D3E8F62A",
"Success": true
}Note the JobId from the response. Use this ID in the following steps to query the job status and get the parsing result.
Step 2: Query the job status
Run the following command to call the describe-doc-parser-job-status command and query the status of the parsing job.
aliyun dtsai describe-doc-parser-job-status \
--endpoint dtsai.cn-beijing.aliyuncs.com \
--biz-region-id cn-beijing \
--job-id 019f64f2-565d-7d70-95d0-3b97d784a113Sample response:
{
"HttpStatusCode": 200,
"RequestId": "019F64F2-F9AC-5F55-8EFC-65729321A64C",
"Status": "success",
"Success": true
}Run this command repeatedly until the Status is success. If the status is failed, check the FailureMessage field for the failure message.
Step 3: Get the parsing result
After the job status is success, run the following command to get the parsing result.
aliyun dtsai describe-doc-parser-job-result \
--endpoint dtsai.cn-beijing.aliyuncs.com \
--biz-region-id cn-beijing \
--job-id 019f64f2-565d-7d70-95d0-3b97d784a113Sample response:
{
"HttpStatusCode": 200,
"RequestId": "019F64F3-75D0-583C-B54D-7E0522689216",
"Result": "# What is Data Transmission Service (DTS)?\n\nData Transmission Service (DTS) is a one-stop data transmission and processing platform provided by Alibaba Cloud...",
"Success": true
}The Result field in the response contains the parsed content in Markdown format.
Method 5: Use a DTS Skill in a coding agent
DTS provides official Skills that allow you to invoke the document parsing capability of the DTS-AI service in natural language within coding agents such as Qoder, Claude Code, Cursor, and Codex, without writing API calls manually. The Skills reuse the APIs described in this topic under the hood, and are suitable for scenarios such as batch parsing of local documents and interactive data processing.
Step 1: Install the Skill
Run the following command to install all DTS Skills (three Skills in total: document parsing, web search, and DTS task management) in one go.
curl -fsSL 'https://aliyun-dts-skills.oss-cn-hangzhou.aliyuncs.com/install.sh' | bash -s -- --agent <Agent Param>If you only need the document parsing capability, use the --skill parameter to install the corresponding Skill individually.
curl -fsSL 'https://aliyun-dts-skills.oss-cn-hangzhou.aliyuncs.com/install.sh' | bash -s -- --agent <Agent Param> --skill aliyun-dts-doc-parseThe --agent parameter supports the following coding agents. The default installation directory of each agent is also listed.
Value | Coding agent | Default installation directory |
| Codex |
|
| Claude Code |
|
| Cursor |
|
| ZCode |
|
| Qoder |
|
| QoderWork |
|
| Qianwen Office |
|
| OpenCode |
|
| Kimi Code |
|
The Skills require Python 3 to run. Make sure that Python 3 is installed on your system before the installation.
To upgrade the Skills to the latest version, re-run the installation command above. The installer automatically backs up the existing Skills and completes the upgrade.
Step 2: Configure the API Key
The document parsing Skill uses API Key authentication. For how to create an API Key, see prepare authentication credentials. When you use the Skill for the first time, the Skill automatically guides you through the API Key configuration. After the configuration succeeds, the credential is saved in the ~/.aliyun-dts/dts-ai/credentials.json file, and no repeated configuration is needed in subsequent sessions. You can also pass the API Key through the DTS_AI_API_KEY environment variable.
Step 3: Use the Skill in a coding agent
After the installation and configuration are complete, describe your requirement in natural language in the coding agent to invoke the Skill for document parsing. Example:
Use the dts skill to parse /path/to/document.pdf into MarkdownThe coding agent automatically calls the document parsing API to complete the task and returns the parsed Markdown content.