This topic describes the core benefits of the Tingwu API.
Multimodal capabilities for voice, language, and vision with 17 flexible AI features
Tingwu supports a wide range of AI capabilities for processing and extracting information from audio and video. In addition to speech recognition, translation, and speaker diarization, it offers features such as chapter overviews, Large Language Model (LLM)-powered summarization (full-text summaries, speaker summaries, Q&A reviews, and mind maps), key point extraction (keywords, action items, highlights, and scenario recognition), service quality inspection, PPT extraction and summarization, spoken-to-written text conversion, and custom prompts.
Module breakdown | Capabilities |
Speech-to-text: Transcribes speech from real-time audio streams or audio and video files into text. It supports transcription for Chinese, English, Cantonese, mixed Chinese-English, Japanese, and Korean. The results can include paragraph and sentence divisions, along with word-level start and end times for displaying captions. Speaker diarization: Distinguishes between speakers in a conversation. You can specify whether there are two or multiple speakers. This feature can be enabled or disabled. | |
Custom prompts let you define your own LLM prompts to guide the model in performing various user-defined tasks. If the standard AI model capabilities provided by Tingwu do not meet your business needs, you can use this feature to leverage the LLM with greater flexibility. | |
This feature combines the following three AI capabilities to segment and summarize chapters in audio and video content: Chapter segmentation: Divides audio and video content into chapters based on different conversation topics. Chapter title: A one-sentence summary of the chapter (32 characters maximum). Chapter summaries: Summarizes chapter content in under 1,000 characters. | |
Summarization (full-text summary, speaker summary, Q&A review, mind map) | Full-text summary: Summarizes the entire audio or video content. Speaker summary: Summarizes the speech of different speakers. The speaker diarization feature must be enabled first. Q&A review: Analyzes the conversation to extract explicit questions, summarize implicit questions, and refine answers based on the dialogue. Mind map: Summarizes the audio or video content and generates the data structure needed to draw a mind map. You can pass the result to a frontend framework to render the mind map image. Mind maps can be generated with a tree structure of up to four levels of depth. |
Keywords: Extracts keywords from the conversation. Action items: Extracts action items from the conversation. Highlights: Extracts key sentences from the conversation. Scenario recognition: Analyzes the content type to identify the scenario. It can recognize interviews, speeches, or meetings. | |
Video PPT extraction: Extracts PPT slides that appear in a video file. PPT narration summary: Summarizes the narration corresponding to each PPT slide. This feature matches the narration to the corresponding PPT slide. It returns the start and end times and a summary for each slide. | |
Real-time speech translation: Supports real-time bidirectional translation between Chinese, English, Japanese, and Korean. It also supports free-form mixed Chinese-English speech translated into Chinese, English, or both. Offline file translation: Transcribes speech from audio and video files and supports bidirectional translation between Chinese, English, Japanese, and Korean. It also supports free-form mixed Chinese-English speech translated into Chinese, English, or both. | |
Spoken-to-written text conversion: Rewrites and polishes the speech-to-text results to produce a formal, written version of the transcript. |
Fast and easy integration
A single set of API parameters lets you enable the AI capabilities needed for different scenarios. This reduces the cost and effort of integrating APIs to build scenario-based AI services.
Stable service
Tingwu supports custom push notifications and status queries. It provides multiple mechanisms for handling abnormal situations, making it easier for you to manage your business logic.