Features

更新时间:
复制 MD 格式

This topic describes the features, supported scenarios, limits, and API call methods of the Tongyi Tingwu service.

Audio and video file service parameter table

Service

Real-time recording

Audio and video file transcription

Mode

Real-time

Offline

File type

Audio stream

Audio stream

Audio file

Audio file

Video file

Audio sampling rate

8k

16k

8k

16k/24k/48k

16k/24k/48k

File format

PCM, OPUS, WAV

PCM, OPUS, WAV

MP3, WAV, M4A, WMA, AAC, OGG, AMR, FLAC, AIFF

MP3, WAV, M4A, WMA, AAC, OGG, AMR, FLAC, AIFF

MP4, WMV, M4V, FLV, RMVB, DAT, MOV, MKV, WEBM, AVI, MPEG, 3GP, OGG

Size limit

24 hours

24 hours

6 GB & 6 hours

6 GB & 6 hours

6 GB & 6 hours

Sound channel / ingest endpoint

Three-way

Three-way

Stereo

First channel

First channel

Language

Chinese

Chinese, English, Cantonese, Japanese, Korean, mixed Chinese/English/Japanese/Korean/Cantonese/German/French/Russian

Chinese, English

Chinese, English, Cantonese, Japanese, Korean, mixed Chinese/English/Japanese/Korean/Cantonese/German/French/Russian

Chinese, English, Cantonese, Japanese, Korean, mixed Chinese/English/Japanese/Korean/Cantonese/German/French/Russian

Hotword-supported languages

Chinese

Chinese, English

Chinese, English

Chinese, English

Chinese, English

Offline speaker diarization

No separation

No separation, 2 speakers, multiple speakers

No separation, 2 speakers

No separation, 2 speakers, multiple speakers

No separation, 2 speakers, multiple speakers

Recognition result return method

By status: return words during sentence; update entire sentence upon completion

By status: return words during sentence; update entire sentence upon completion

Return full transcription result with timestamps

Return full transcription result with timestamps

Return full transcription result with timestamps

Call SDK

Java, Python, Go

Java, Python, Go

Java, Python, Go

Java, Python, Go

Java, Python, Go

Source file transfer method

Establish WebSocket connection and stream ingest in real time

Establish WebSocket connection and stream ingest in real time

OSS URL

OSS URL

OSS URL

Large Language Model (LLM)-related capabilities (prerequisite: speech-to-text)

Feature

Minimum word count

limit

Corresponding minimum

audio duration

Optimal audio duration

Response content

limit

Supported languages

Full-text summary

Full text of 250 characters

The information above

Approximately 70 seconds or more

Up to 4 hours

Up to 1000 words

Chinese, English,

mixed Chinese/English

Chapter overview

Section: 250 characters

The preceding

Approximately 70 seconds or more

Up to 4 hours

Summary per segment

Fewer than 1,000 characters

Chinese, English,

mixed Chinese/English

Speaker summary

Speech Content

250 words or more

Approximately 70 seconds or more

Up to 4 hours

Up to 1000 words per speaker

Chinese, English,

mixed Chinese/English

Q&A review

Full text of 300 characters

Above

Approximately 90 seconds or more

Up to 4 hours

About 30–50 Q&A pairs per hour of audio

Average length of each Q&A pair: 90 words

Chinese, English,

mixed Chinese/English

Action items

No limit

No limit

Over 90 seconds

Up to 4 hours

Up to 6 action items

5 to 30 characters

Chinese, English

Keywords

Full text of 200 characters

The preceding

Approximately 60 seconds or more

Up to 70 minutes

Up to 20 keywords

Chinese, English, Cantonese,

mixed Chinese/English

Spoken-to-written conversion

No limit

No limit

Up to 4 hours

None

Chinese, English,

mixed Chinese/English

Mind map

No limit

No limit

Up to 90 minutes

Up to 4 levels deep

Chinese

Custom prompt

No limit

No limit

Up to 4 hours

Up to 1000 words

Chinese, English

Service quality inspection

No limit

No limit

Up to 4 hours

Based on inspection requirements

Chinese

Content extraction

No limit

No limit

Up to 4 hours

Based on extraction requirements

Chinese

PPT extraction and summarization (prerequisite: audio and video file transcription; file type: video)

Feature

Extractable graphics

Description

Summary supported languages

Video PPT extraction

Full PPT or lecture mode

After upload completes, about 2–5 minutes per hour of video; up to 200 PPT slides extracted

Unlimited

PPT explanation summary

Full PPT or lecture mode

About 1 minute after transcription completes

Chinese, English

Note: This feature supports only screen-shared PPTs or PPTs displayed with a speaker in a picture-in-picture layout. It does not support scenarios where a person walks in front of the PPT or gives a live presentation.

You can test this feature on the Tongyi Tingwu website. Test now

Tongyi Tingwu translation (prerequisite: speech-to-text)

Service

File type

Audio sampling rate

Translation

Translation Support

Real-time speech translation

Audio stream

8k

Real-time

Bidirectional translation among Chinese, English, Japanese, Korean, German, French, and Russian;

Mixed Chinese/English speech translated to Chinese, English, or both

Audio stream

16k

Real-time

Audio and video file translation

Audio file

8k

Offline

Audio file

16k/24k/48k

Offline

Video file

16k/24k/48k

Offline