This topic describes the features, supported scenarios, limits, and API call methods of the Tongyi Tingwu service.
Audio and video file service parameter table
Service | Real-time recording | Audio and video file transcription | |||
Mode | Real-time | Offline | |||
File type | Audio stream | Audio stream | Audio file | Audio file | Video file |
Audio sampling rate | 8k | 16k | 8k | 16k/24k/48k | 16k/24k/48k |
File format | PCM, OPUS, WAV | PCM, OPUS, WAV | MP3, WAV, M4A, WMA, AAC, OGG, AMR, FLAC, AIFF | MP3, WAV, M4A, WMA, AAC, OGG, AMR, FLAC, AIFF | MP4, WMV, M4V, FLV, RMVB, DAT, MOV, MKV, WEBM, AVI, MPEG, 3GP, OGG |
Size limit | 24 hours | 24 hours | 6 GB & 6 hours | 6 GB & 6 hours | 6 GB & 6 hours |
Sound channel / ingest endpoint | Three-way | Three-way | Stereo | First channel | First channel |
Language | Chinese | Chinese, English, Cantonese, Japanese, Korean, mixed Chinese/English/Japanese/Korean/Cantonese/German/French/Russian | Chinese, English | Chinese, English, Cantonese, Japanese, Korean, mixed Chinese/English/Japanese/Korean/Cantonese/German/French/Russian | Chinese, English, Cantonese, Japanese, Korean, mixed Chinese/English/Japanese/Korean/Cantonese/German/French/Russian |
Hotword-supported languages | Chinese | Chinese, English | Chinese, English | Chinese, English | Chinese, English |
Offline speaker diarization | No separation | No separation, 2 speakers, multiple speakers | No separation, 2 speakers | No separation, 2 speakers, multiple speakers | No separation, 2 speakers, multiple speakers |
Recognition result return method | By status: return words during sentence; update entire sentence upon completion | By status: return words during sentence; update entire sentence upon completion | Return full transcription result with timestamps | Return full transcription result with timestamps | Return full transcription result with timestamps |
Call SDK | Java, Python, Go | Java, Python, Go | Java, Python, Go | Java, Python, Go | Java, Python, Go |
Source file transfer method | Establish WebSocket connection and stream ingest in real time | Establish WebSocket connection and stream ingest in real time | OSS URL | OSS URL | OSS URL |
Large Language Model (LLM)-related capabilities (prerequisite: speech-to-text)
Feature | Minimum word count limit | Corresponding minimum audio duration | Optimal audio duration | Response content limit | Supported languages |
Full-text summary | Full text of 250 characters The information above | Approximately 70 seconds or more | Up to 4 hours | Up to 1000 words | Chinese, English, mixed Chinese/English |
Chapter overview | Section: 250 characters The preceding | Approximately 70 seconds or more | Up to 4 hours | Summary per segment Fewer than 1,000 characters | Chinese, English, mixed Chinese/English |
Speaker summary | Speech Content 250 words or more | Approximately 70 seconds or more | Up to 4 hours | Up to 1000 words per speaker | Chinese, English, mixed Chinese/English |
Q&A review | Full text of 300 characters Above | Approximately 90 seconds or more | Up to 4 hours | About 30–50 Q&A pairs per hour of audio Average length of each Q&A pair: 90 words | Chinese, English, mixed Chinese/English |
Action items | No limit | No limit | Over 90 seconds Up to 4 hours | Up to 6 action items 5 to 30 characters | Chinese, English |
Keywords | Full text of 200 characters The preceding | Approximately 60 seconds or more | Up to 70 minutes | Up to 20 keywords | Chinese, English, Cantonese, mixed Chinese/English |
Spoken-to-written conversion | No limit | No limit | Up to 4 hours | None | Chinese, English, mixed Chinese/English |
Mind map | No limit | No limit | Up to 90 minutes | Up to 4 levels deep | Chinese |
Custom prompt | No limit | No limit | Up to 4 hours | Up to 1000 words | Chinese, English |
Service quality inspection | No limit | No limit | Up to 4 hours | Based on inspection requirements | Chinese |
Content extraction | No limit | No limit | Up to 4 hours | Based on extraction requirements | Chinese |
PPT extraction and summarization (prerequisite: audio and video file transcription; file type: video)
Feature | Extractable graphics | Description | Summary supported languages |
Video PPT extraction | Full PPT or lecture mode | After upload completes, about 2–5 minutes per hour of video; up to 200 PPT slides extracted | Unlimited |
PPT explanation summary | Full PPT or lecture mode | About 1 minute after transcription completes | Chinese, English |
Note: This feature supports only screen-shared PPTs or PPTs displayed with a speaker in a picture-in-picture layout. It does not support scenarios where a person walks in front of the PPT or gives a live presentation.
You can test this feature on the Tongyi Tingwu website. Test now
Tongyi Tingwu translation (prerequisite: speech-to-text)
Service | File type | Audio sampling rate | Translation | Translation Support |
Real-time speech translation | Audio stream | 8k | Real-time | Bidirectional translation among Chinese, English, Japanese, Korean, German, French, and Russian; Mixed Chinese/English speech translated to Chinese, English, or both |
Audio stream | 16k | Real-time | ||
Audio and video file translation | Audio file | 8k | Offline | |
Audio file | 16k/24k/48k | Offline | ||
Video file | 16k/24k/48k | Offline |