Parse documents with DTS-AI

Updated at:

This topic describes how to use the DTS-AI service API to parse documents in formats such as PDF and Word into Markdown. It also provides complete examples for four calling methods: curl, SDK, Alibaba Cloud CLI, and DTS Skill.

Overview

The DTS-AI service provides a document parsing feature that converts documents in formats such as PDF, Word, PPT, plain text, Markdown, and images into structured Markdown content.

Document parsing supports synchronous and asynchronous calling methods. The synchronous method is simpler — you just set the ResponseMode parameter to sync.

This topic uses the asynchronous approach as the main example. For the synchronous approach, see Method 3: Use the SDK for synchronous calls (AccessKey authentication). The overall asynchronous workflow is as follows:

  1. Create a parsing job: Call the CreateDocParserJob API to submit a document parsing job and get a job ID (JobId).

  2. Query the job status: Call the DescribeDocParserJobStatus API to poll the job status until it becomes success.

  3. Get the parsing result: Call the DescribeDocParserJobResult API to get the parsed Markdown content.

Quick trial

If you want to test the document parsing capability before integrating the SDK or CLI, you can call the DTS-AI service APIs directly from the Alibaba Cloud OpenAPI Developer Portal:

After you sign in and grant permissions, enter the request parameters on the left side of the page and click Debug to view the response. Once your test succeeds, refer to the SDK or CLI examples below to integrate the API into your application.

Prerequisites

Before you call the DTS-AI service APIs, complete the following preparations.

Prepare authentication credentials

The DTS-AI service supports two types of authentication credentials. Choose either one. API Key is recommended because its permissions are restricted to the DTS service scope, resulting in more controllable security risks.

  • API Key (recommended): An API Key is a service credential issued by RAM for a single cloud service. It can only call the OpenAPIs of the bound cloud service, resulting in a smaller permission scope. In the RAM console, locate the target user, go to the Credential Management section, click Create API Key, and set Cloud Service to Data Transmission Service / DTS. After the API Key is created, copy the plaintext credential immediately and store it securely (the plaintext is unavailable after the dialog is closed). For more information, see Create an API key.

    Important
    • Alibaba Cloud accounts (root accounts) cannot create API Keys. Only RAM users can create them. Each RAM user can create up to 2 API Keys per cloud service.

    • An API Key is a credential used to access cloud services. If it is leaked, anyone can call your APIs, resulting in unpredictable costs and security risks. Keep your API Key secure and do not write it in plaintext in code, commit it to code repositories, or publish it to public platforms. We recommend that you store it in environment variables. For more information, see the "" section below.

  • AccessKey ID and AccessKey Secret: Suitable for scenarios where you want to reuse existing Alibaba Cloud AccessKey credentials. For security, we recommend using the AccessKey of a RAM user and configuring a permission policy based on the principle of least privilege. For more information, see Authorize a RAM user to manage DTS.

Note

The following examples read credentials from environment variables:

  • API Key: read from DTS_AI_API_KEY.

  • AccessKey: read from ALIBABA_CLOUD_ACCESS_KEY_ID and ALIBABA_CLOUD_ACCESS_KEY_SECRET.

Before you run the examples, set the corresponding environment variables in your operating system to avoid hard-coding credentials in code. For secure credential usage, see Manage credentials.

Prepare a calling tool (choose as needed)

Usage notes

  • The document parsing API is currently available only in the China (Beijing) region. Set the RegionId to cn-beijing.

  • The FileUrl parameter must be a valid OSS URL. If your file is not stored in OSS, we recommend using the CreateDocParserJobAdvance method of the SDK to pass a local file stream directly.

  • The output format (OutputFormat) currently supports only markdown.

  • The parsing result is retained for 72 hours. Download and save it promptly, as it cannot be retrieved after this period.

  • Supported file formats include PDF, DOCX, DOC, PPTX, PPT, TXT, Markdown, PNG, JPG, and JPEG.

  • The SDKs for different languages provide the CreateDocParserJobAdvance method, which lets you pass a local file stream directly without first uploading the file to OSS.

Method 1: Use curl (recommended, API Key authentication)

Pass the API Key through the X-Acs-ApiKey request header to call the DTS-AI service APIs directly. No signature computation is required, which is suitable for quick verification or lightweight integration scenarios.

Step 1: Configure the API Key environment variable

Configure the API Key as an environment variable to avoid exposing the plaintext in commands.

export DTS_AI_API_KEY='<your API Key>'

Step 2: Create a document parsing job

Call the CreateDocParserJob operation to create a document parsing job.

curl -X POST \
  "https://dtsai.cn-beijing.aliyuncs.com?Action=CreateDocParserJob&Version=2026-04-01&SignatureNonce=$(uuidgen)" \
  --header "X-Acs-ApiKey: ${DTS_AI_API_KEY}" \
  --header "Content-Type: application/json" \
  --header "Accept: application/json" \
  --data-raw '{
    "RegionId": "cn-beijing",
    "FileUrl": "https://<BucketName>.oss-cn-beijing.aliyuncs.com/document.pdf",
    "FileName": "document.pdf",
    "FileFormat": "pdf",
    "OutputFormat": "markdown"
  }'

Sample response:

{
  "RequestId": "019F6482-DD04-510E-A212-A21B3C59CB8C",
  "Success": true,
  "HttpStatusCode": 200,
  "JobId": "job_abc123"
}

Note the JobId from the response. Use this ID in the following steps to query the job status and get the parsing result.

Step 3: Query the job status

Call the DescribeDocParserJobStatus operation to poll the job status until it returns success.

curl -X POST \
  "https://dtsai.cn-beijing.aliyuncs.com?Action=DescribeDocParserJobStatus&Version=2026-04-01&SignatureNonce=$(uuidgen)" \
  --header "X-Acs-ApiKey: ${DTS_AI_API_KEY}" \
  --header "Content-Type: application/json" \
  --data-raw '{
    "RegionId": "cn-beijing",
    "JobId": "job_abc123"
  }'

Step 4: Get the parsing result

After the job status is success, call the DescribeDocParserJobResult operation to get the parsing result.

curl -X POST \
  "https://dtsai.cn-beijing.aliyuncs.com?Action=DescribeDocParserJobResult&Version=2026-04-01&SignatureNonce=$(uuidgen)" \
  --header "X-Acs-ApiKey: ${DTS_AI_API_KEY}" \
  --header "Content-Type: application/json" \
  --data-raw '{
    "RegionId": "cn-beijing",
    "JobId": "job_abc123"
  }'

The Result field in the response contains the parsed content in Markdown format.

Method 2: Use the SDK for asynchronous calls (AccessKey authentication)

The following example uses the Java SDK to demonstrate the document parsing workflow. The process is similar for SDKs in other languages. For more information, see the SDK download page.

Step 1: Install the SDK

Add the following dependency to the pom.xml file of your Maven project:

<dependency>
  <groupId>com.aliyun</groupId>
  <artifactId>dtsai20260401</artifactId>
  <version>1.0.0</version>
</dependency>

Step 2: Create a document parsing job

Call the CreateDocParserJob API to create a document parsing job. If the file is stored locally, you can use the CreateDocParserJobAdvance method to pass the file stream directly.

File URL

import com.aliyun.dtsai20260401.Client;
import com.aliyun.dtsai20260401.models.*;
import com.aliyun.teaopenapi.models.Config;

public class DocParserExample {
    public static void main(String[] args) throws Exception {
        // Initialize the client.
        Config config = new Config()
            .setAccessKeyId(System.getenv("ALIBABA_CLOUD_ACCESS_KEY_ID"))
            .setAccessKeySecret(System.getenv("ALIBABA_CLOUD_ACCESS_KEY_SECRET"));
        config.endpoint = "dtsai.cn-beijing.aliyuncs.com";
        Client client = new Client(config);

        // Create a parsing job.
        CreateDocParserJobRequest request = new CreateDocParserJobRequest()
            .setRegionId("cn-beijing")
            .setFileUrl("https://example.oss-cn-beijing.aliyuncs.com/document.pdf")
            .setFileName("document.pdf")
            .setFileFormat("pdf")
            .setOutputFormat("markdown");

        CreateDocParserJobResponse response = client.createDocParserJob(request);
        String jobId = response.getBody().getJobId();
        System.out.println("Job created successfully. JobId: " + jobId);
    }
}

Local file stream

import com.aliyun.dtsai20260401.Client;
import com.aliyun.dtsai20260401.models.*;
import com.aliyun.teaopenapi.models.Config;
import java.io.FileInputStream;

public class DocParserAdvanceExample {
    public static void main(String[] args) throws Exception {
        Config config = new Config()
            .setAccessKeyId(System.getenv("ALIBABA_CLOUD_ACCESS_KEY_ID"))
            .setAccessKeySecret(System.getenv("ALIBABA_CLOUD_ACCESS_KEY_SECRET"));
        config.endpoint = "dtsai.cn-beijing.aliyuncs.com";
        Client client = new Client(config);

        // Use the Advance method to directly pass a local file stream.
        CreateDocParserJobAdvanceRequest request = new CreateDocParserJobAdvanceRequest()
            .setRegionId("cn-beijing")
            .setFileUrlObject(new FileInputStream("/path/to/document.pdf"))
            .setFileName("document.pdf")
            .setFileFormat("pdf")
            .setOutputFormat("markdown");

        CreateDocParserJobResponse response = client.createDocParserJobAdvance(request,
            new com.aliyun.teautil.models.RuntimeOptions());
        String jobId = response.getBody().getJobId();
        System.out.println("Job created successfully. JobId: " + jobId);
    }
}

Step 3: Query the job status

Call the DescribeDocParserJobStatus operation to poll the job status until it returns success.

// Poll the job status
DescribeDocParserJobStatusRequest statusRequest = new DescribeDocParserJobStatusRequest()
    .setRegionId("cn-beijing")
    .setJobId(jobId);

String status = "";
while (!"success".equals(status) && !"failed".equals(status)) {
    Thread.sleep(5000); // Query every 5 seconds
    DescribeDocParserJobStatusResponse statusResponse =
        client.describeDocParserJobStatus(statusRequest);
    status = statusResponse.getBody().getStatus();
    System.out.println("Current status: " + status);

    if ("failed".equals(status)) {
        System.out.println("Job failed: " + statusResponse.getBody().getFailureMessage());
        return;
    }
}

The following table describes the job status values.

Status value

Description

init

Created. The job is being prepared.

pending

Queued and waiting for execution.

running

Processing. The document is being parsed.

success

Parsing is complete. Call DescribeDocParserJobResult to get the result.

failed

Parsing failed. The failure message is available in the FailureMessage parameter.

cancelled

The job is canceled.

Step 4: Get the parsing result

After the job status changes to success, call the DescribeDocParserJobResult API to get the parsing result.

// Get the parsing result
DescribeDocParserJobResultRequest resultRequest = new DescribeDocParserJobResultRequest()
    .setRegionId("cn-beijing")
    .setJobId(jobId);

DescribeDocParserJobResultResponse resultResponse =
    client.describeDocParserJobResult(resultRequest);
String markdownContent = resultResponse.getBody().getResult();
System.out.println("Parsing result:\n" + markdownContent);
Important

The parsing result is retained for only 72 hours. Get and save the result promptly after the job is complete.

Method 3: Use the SDK for synchronous calls (AccessKey authentication)

The following example uses the Java SDK to demonstrate the complete document parsing workflow. For usage of other languages, see SDK download page.

Step 1: Install the SDK

Add the following dependency to the pom.xml file of your Maven project:

<dependency>
  <groupId>com.aliyun</groupId>
  <artifactId>dtsai20260401</artifactId>
  <version>1.0.0</version>
</dependency>

Step 2: Create a synchronous document parsing job

Call the CreateDocParserJob API to create a document parsing job. To parse a local file, use the CreateDocParserJobAdvance method to pass the file stream directly.

Create a parsing job with a file URL

import com.aliyun.dtsai20260401.Client;
import com.aliyun.dtsai20260401.models.*;
import com.aliyun.teaopenapi.models.Config;

public class DocParserExample {
    public static void main(String[] args) throws Exception {
        // Initialize client
        Config config = new Config()
            .setAccessKeyId(System.getenv("ALIBABA_CLOUD_ACCESS_KEY_ID"))
            .setAccessKeySecret(System.getenv("ALIBABA_CLOUD_ACCESS_KEY_SECRET"));
        config.endpoint = "dtsai.cn-beijing.aliyuncs.com";
        Client client = new Client(config);

        // Create a parsing job
        CreateDocParserJobRequest request = new CreateDocParserJobRequest()
            .setRegionId("cn-beijing")
            .setFileUrl("https://example.oss-cn-beijing.aliyuncs.com/document.pdf")
            .setFileName("document.pdf")
            .setFileFormat("pdf")
            .setOutputFormat("markdown")
            .setResultType("content")
            .setResponseMode("sync");

        CreateDocParserJobResponse response = client.createDocParserJob(request);
        String jobId = response.getBody().getJobId();
        System.out.println("Job created successfully. JobId: " + jobId);
    }
}

Create a parsing job with a local file stream (Advance method)

import com.aliyun.dtsai20260401.Client;
import com.aliyun.dtsai20260401.models.*;
import com.aliyun.teaopenapi.models.Config;
import java.io.FileInputStream;

public class DocParserAdvanceExample {
    public static void main(String[] args) throws Exception {
        Config config = new Config()
            .setAccessKeyId(System.getenv("ALIBABA_CLOUD_ACCESS_KEY_ID"))
            .setAccessKeySecret(System.getenv("ALIBABA_CLOUD_ACCESS_KEY_SECRET"));
        config.endpoint = "dtsai.cn-beijing.aliyuncs.com";
        Client client = new Client(config);

        // Use the Advance method to pass a local file stream directly
        CreateDocParserJobAdvanceRequest request = new CreateDocParserJobAdvanceRequest()
            .setRegionId("cn-beijing")
            .setFileUrlObject(new FileInputStream("/path/to/document.pdf"))
            .setFileName("document.pdf")
            .setFileFormat("pdf")
            .setOutputFormat("markdown")
            .setResultType("content")
            .setResponseMode("sync");

        CreateDocParserJobResponse response = client.createDocParserJobAdvance(request,
            new com.aliyun.teautil.models.RuntimeOptions());
        String jobId = response.getBody().getJobId();
        System.out.println("Job created successfully. JobId: " + jobId);
    }
}

The returned result is as follows:

{
  "Status": "success",
  "RequestId": "01A055D1-178A-5292-9230-****",
  "HttpStatusCode": 200,
  "ResultType": "content",
  "Success": true,
  "JobId": "01a055d1-17bd-7ed2-ba59-****",
  "Result": "<!-- page 1 -->\n\n# DTS FORMULA MARKDOWN 2026\n\n$$\nx = \\frac{-b \\pm \\sqrt{b^2 - 4ac}}{2a}\n$$"
}

Method 4: Use the Alibaba Cloud CLI

This example shows how to complete the document parsing process using the Alibaba Cloud CLI. Before you start, complete the following preparations:

  • Upgrade the Alibaba Cloud CLI to the latest version and install the aliyun-cli-dtsai plug-in by running the following command:

    aliyun plugin install --names aliyun-cli-dtsai
  • Configure the identity credentials for the Alibaba Cloud CLI. OAuth mode is recommended (no need to maintain AccessKey): run aliyun configure --mode OAuth and follow the prompts to sign in. You can also use other authentication modes such as AccessKey or STS Token. For more information, see Configure and manage credentials.

Note

DTS-AI service CLI commands are provided by the aliyun-cli-dtsai plug-in. All commands and parameters use kebab-case (hyphen-separated), for example aliyun dtsai web-search --biz-region-id cn-beijing --query "...".

Step 1: Create a document parsing job

Run the following command to call the create-doc-parser-job command and create a document parsing job.

aliyun dtsai create-doc-parser-job \
  --endpoint dtsai.cn-beijing.aliyuncs.com \
  --biz-region-id cn-beijing \
  --file-url "https://<BucketName>.oss-cn-beijing.aliyuncs.com/document.pdf" \
  --file-name document.pdf \
  --file-format pdf \
  --output-format markdown

Sample response:

{
  "HttpStatusCode": 200,
  "JobId": "019f64f2-565d-7d70-95d0-3b97d784a113",
  "RequestId": "019F64F2-561D-522C-8F56-71D8D3E8F62A",
  "Success": true
}

Note the JobId from the response. Use this ID in the following steps to query the job status and get the parsing result.

Step 2: Query the job status

Run the following command to call the describe-doc-parser-job-status command and query the status of the parsing job.

aliyun dtsai describe-doc-parser-job-status \
  --endpoint dtsai.cn-beijing.aliyuncs.com \
  --biz-region-id cn-beijing \
  --job-id 019f64f2-565d-7d70-95d0-3b97d784a113

Sample response:

{
  "HttpStatusCode": 200,
  "RequestId": "019F64F2-F9AC-5F55-8EFC-65729321A64C",
  "Status": "success",
  "Success": true
}

Run this command repeatedly until the Status is success. If the status is failed, check the FailureMessage field for the failure message.

Step 3: Get the parsing result

After the job status is success, run the following command to get the parsing result.

aliyun dtsai describe-doc-parser-job-result \
  --endpoint dtsai.cn-beijing.aliyuncs.com \
  --biz-region-id cn-beijing \
  --job-id 019f64f2-565d-7d70-95d0-3b97d784a113

Sample response:

{
  "HttpStatusCode": 200,
  "RequestId": "019F64F3-75D0-583C-B54D-7E0522689216",
  "Result": "# What is Data Transmission Service (DTS)?\n\nData Transmission Service (DTS) is a one-stop data transmission and processing platform provided by Alibaba Cloud...",
  "Success": true
}

The Result field in the response contains the parsed content in Markdown format.

Method 5: Use a DTS Skill in a coding agent

DTS provides official Skills that allow you to invoke the document parsing capability of the DTS-AI service in natural language within coding agents such as Qoder, Claude Code, Cursor, and Codex, without writing API calls manually. The Skills reuse the APIs described in this topic under the hood, and are suitable for scenarios such as batch parsing of local documents and interactive data processing.

Step 1: Install the Skill

Run the following command to install all DTS Skills (three Skills in total: document parsing, web search, and DTS task management) in one go.

curl -fsSL 'https://aliyun-dts-skills.oss-cn-hangzhou.aliyuncs.com/install.sh' | bash -s -- --agent <Agent Param>

If you only need the document parsing capability, use the --skill parameter to install the corresponding Skill individually.

curl -fsSL 'https://aliyun-dts-skills.oss-cn-hangzhou.aliyuncs.com/install.sh' | bash -s -- --agent <Agent Param> --skill aliyun-dts-doc-parse

The --agent parameter supports the following coding agents. The default installation directory of each agent is also listed.

Value

Coding agent

Default installation directory

codex

Codex

~/.codex/skills

claude

Claude Code

~/.claude/skills

cursor

Cursor

~/.cursor/skills

zcode

ZCode

~/.zcode/skills

qoder

Qoder

~/.qoder/skills

qoderwork

QoderWork

~/.qoderwork/skills

qwenworkcn

Qianwen Office

~/.qwenworkcn/skills

opencode

OpenCode

~/.config/opencode/skills

kimi

Kimi Code

~/.kimi/skills

Note
  • The Skills require Python 3 to run. Make sure that Python 3 is installed on your system before the installation.

  • To upgrade the Skills to the latest version, re-run the installation command above. The installer automatically backs up the existing Skills and completes the upgrade.

Step 2: Configure the API Key

The document parsing Skill uses API Key authentication. For how to create an API Key, see prepare authentication credentials. When you use the Skill for the first time, the Skill automatically guides you through the API Key configuration. After the configuration succeeds, the credential is saved in the ~/.aliyun-dts/dts-ai/credentials.json file, and no repeated configuration is needed in subsequent sessions. You can also pass the API Key through the DTS_AI_API_KEY environment variable.

Step 3: Use the Skill in a coding agent

After the installation and configuration are complete, describe your requirement in natural language in the coding agent to invoke the Skill for document parsing. Example:

Use the dts skill to parse /path/to/document.pdf into Markdown

The coding agent automatically calls the document parsing API to complete the task and returns the parsed Markdown content.