Performance

更新时间:
复制 MD 格式

This topic lists frequently asked questions about the performance of the Tingwu service.

What is the recording duration for audio and video files?

Transcription is usually completed within 3 hours.

What is the latency for real-time transcription?

The end-point latency for real-time transcription is about 300 ms. This can vary slightly based on the video model and audio quality.

Can a meeting support Mandarin, English, and Cantonese at the same time?

No. The service supports meetings in Mandarin with mixed English.

How is speech recognition accuracy calculated, and what is the character accuracy rate?

The industry typically uses an error rate to measure recognition performance. Character Error Rate (CER) is common for Chinese, and Word Error Rate (WER) is common for English. The formula is: (Insertions + Deletions + Substitutions) / Total Characters. For example, in the figure below:错误率 The accuracy for this data is (14365 - 74 - 385 - 1706) / 14365 = 84.93%. A quick calculation is 100% - 15.07% = 84.93%.

DAMO Academy's Intelligent Speech Interaction has been certified by the China National Accreditation Service for Conformity Assessment (CNAS) Software Testing Center. In a controlled test environment, our recognition accuracy for Mandarin exceeds 98%.

What is the maximum lifecycle of a meeting? How long after creating a real-time meeting is it automatically destroyed?

A real-time meeting has a lifecycle of 24 hours. If you do not actively destroy the meeting, it is automatically destroyed after 24 hours.

If there is no audio data for a long time in a meeting, will it automatically disconnect?

The connection automatically disconnects if no audio data is received for 10 seconds.

After a meeting automatically disconnects, do I need to create a new meeting, or can I rejoin the previous one?

If a meeting disconnects, you can rejoin it using its ingest URL. You do not need to create a new one.