API reference

Updated at:

Short sentence recognition converts audio clips of up to 60 seconds into text for chat conversations, voice commands, voice input, and voice search. Clients send audio streams over WebSocket and receive recognition events and results.

Billing and concurrency limits

Short sentence recognition is available in trial and commercial versions. For billable items, see Billable items. For billing methods and upgrade instructions, see Billing methods.

For concurrency quotas and how to adjust them, see Concurrency and QPS.

Usage notes

The audio encoding must match the request parameters. A mismatch can cause recognition to fail or return an empty result.

  • Channels and bit depth: mono, 16-bit.

  • Audio formats: PCM, PCM-encoded WAV, OPUS in an OGG container, SPEEX in an OGG container, and AMR. MP3 and AAC are also supported.

  • Sample rate: 8000 Hz or 16000 Hz. The project model must support the audio sample rate and language.

  • Audio duration: up to 60 seconds.

  • Audio size: up to 2 MB.

Request parameters control features such as intermediate results, punctuation, and inverse text normalization (ITN).

Sentiment analysis is available only with the Chinese 8 kHz sentiment recognition model.

Note

For Android and iOS SDKs, see Mobile SDK API reference.

Select a recognition model

Language and dialect models cannot be specified in request parameters. On the All Projects page of the Intelligent Speech Interaction console, find the project and click Configure Project Features. Select a model that matches the audio language and sample rate. For configuration instructions, see Manage projects.

Language and dialect models

The tables list language and dialect models and their capabilities. Feature availability also depends on the API and model in use.

Languages

Language

Model name

Sample rate

Punctuation

ITN

Output smoothing

Semantic segmentation

Audio-text alignment

English

General - English, Education Livestream - English, Educational Content Analysis - English

16 kHz

Supported

Supported

Supported

Not supported

Supported

General Customer Service - English

8 kHz

Supported

Supported

Supported

Not supported

Not supported

Southeast Asian multilingual

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Japanese

General - Japanese

16 kHz

Supported

Supported

Not supported

Not supported

Supported

Spanish

General - Spanish

16 kHz

Supported

Supported

Not supported

Not supported

Not supported

General Customer Service - Spanish

8 kHz

Supported

Supported

Not supported

Not supported

Not supported

Arabic

General - Arabic

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Kazakh

General - Kazakh

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Korean

General - Korean

16 kHz

Supported

Supported

Not supported

Not supported

Not supported

Thai

General - Thai

16 kHz

Not supported

Not supported

Not supported

Not supported

Not supported

General Customer Service - Thai

8 kHz

Not supported

Not supported

Not supported

Not supported

Not supported

Southeast Asian multilingual

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Indonesian

General - Indonesian

16 kHz

Supported

Supported

Not supported

Not supported

Not supported

General Customer Service - Indonesian

8 kHz

Supported

Supported

Not supported

Not supported

Not supported

Southeast Asian multilingual

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Russian

General - Russian

16 kHz

Supported

Supported

Not supported

Not supported

Not supported

Vietnamese

General - Vietnamese

16 kHz

Supported

Supported

Not supported

Not supported

Not supported

General Customer Service - Vietnamese

8 kHz

Supported

Supported

Not supported

Not supported

Not supported

Southeast Asian multilingual

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

French

General - French

16 kHz

Supported

Supported

Not supported

Not supported

Not supported

German

General - German

16 kHz

Supported

Supported

Not supported

Not supported

Not supported

Italian

General - Italian

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Hindi

General - Hindi

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Malay

General - Malay

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

General Customer Service - Malay

8 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Southeast Asian multilingual

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Filipino

General - Filipino

16 kHz

Supported

Supported

Not supported

Not supported

Not supported

General Customer Service - Filipino

8 kHz

Supported

Supported

Not supported

Not supported

Not supported

Southeast Asian multilingual

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Tamil

General - Tamil

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Portuguese

General - Portuguese

16 kHz

Supported

Supported

Not supported

Not supported

Not supported

Turkish

General - Turkish

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Polish

General - Polish

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Ukrainian

General - Ukrainian

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Romanian

General - Romanian

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Dutch

General - Dutch

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Greek

General - Greek

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Hungarian

General - Hungarian

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Javanese

General - Javanese

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Bengali

General - Bengali

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Burmese

General - Burmese

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Lao

General - Lao

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Swahili

General - Swahili

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Azerbaijani

General - Azerbaijani

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Persian

General - Persian

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Sinhala

General - Sinhala

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Catalan

General - Catalan

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Khmer

General - Khmer

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Hebrew

General - Hebrew

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Croatian

General - Croatian

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Hausa

General - Hausa

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Marathi

General - Marathi

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Telugu

General - Telugu

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Punjabi

General - Punjabi

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Swedish

General - Swedish

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Bulgarian

General - Bulgarian

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Danish

General - Danish

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Norwegian

General - Norwegian

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Kannada

General - Kannada

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Malayalam

General - Malayalam

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Czech

General - Czech

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Urdu

General - Urdu

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Nepali

General - Nepali

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Mongolian (Cyrillic)

General - Mongolian (Cyrillic)

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Uzbek

General - Uzbek

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Dialects

Language

Model name

Sample rate

Punctuation

ITN

Output smoothing

Semantic segmentation

Audio-text alignment

Cantonese

General - Cantonese

16 kHz

Supported

Supported

Supported

Not supported

Supported

Telephony Customer Service (General)

8 kHz

Supported

Supported

Supported

Not supported

Supported

Cantonese-Mandarin Code-switching

8 kHz

Supported

Supported

Supported

Not supported

Not supported

Southeast Asian Multilingual

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Cantonese (Traditional Chinese)

General - Cantonese (Traditional Chinese)

8 kHz

Supported

Not supported

Not supported

Not supported

Not supported

General - Cantonese (Traditional Chinese)

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Sichuanese

General - Sichuanese

16 kHz

Supported

Supported

Supported

Supported

Supported

Telephony Customer Service (General)

8 kHz

Supported

Supported

Supported

Supported

Supported

Hubei dialect

General - Hubei dialect

16 kHz

Supported

Supported

Supported

Supported

Supported

General - Hubei dialect

8 kHz

Supported

Supported

Supported

Supported

Supported

Shanghainese

General - Shanghainese

16 kHz

Supported

Supported

Supported

Supported

Not supported

Hunan dialect

General - Hunan dialect

16 kHz

Supported

Supported

Supported

Supported

Supported

Henan dialect

General - Henan dialect

16 kHz

Supported

Supported

Supported

Supported

Supported

General - Henan dialect

8 kHz

Supported

Supported

Supported

Supported

Supported

Zhejiang dialect

General - Zhejiang dialect

16 kHz

Supported

Supported

Supported

Supported

Not supported

Northeastern Mandarin

General - Northeastern Mandarin

16 kHz

Supported

Supported

Supported

Supported

Supported

Shandong dialect

General - Shandong dialect

16 kHz

Supported

Supported

Supported

Supported

Supported

Tianjin dialect

General - Tianjin dialect

16 kHz

Supported

Supported

Supported

Supported

Supported

Shaanxi dialect

General - Shaanxi dialect

16 kHz

Supported

Supported

Supported

Supported

Supported

Shanxi dialect

General - Shanxi dialect

16 kHz

Supported

Supported

Supported

Supported

Supported

Guizhou dialect

General - Guizhou dialect

16 kHz

Supported

Supported

Supported

Supported

Supported

Yunnan dialect

General - Yunnan dialect

16 kHz

Supported

Supported

Supported

Supported

Supported

Gansu dialect

General - Gansu dialect

16 kHz

Supported

Supported

Supported

Supported

Supported

Uyghur

General - Uyghur

16 kHz

Not supported

Not supported

Not supported

Not supported

Not supported

General - Uyghur

8 kHz

Not supported

Not supported

Not supported

Not supported

Not supported

Suzhou dialect

General - Suzhou dialect

16 kHz

Supported

Supported

Supported

Supported

Not supported

Minnan

General - Minnan

16 kHz

Supported

Supported

Supported

Supported

Not supported

Jiangxi dialect

General - Jiangxi dialect

16 kHz

Supported

Supported

Supported

Supported

Supported

Ningxia dialect

General - Ningxia dialect

16 kHz

Supported

Supported

Supported

Supported

Supported

Guangxi dialect

General - Guangxi dialect

16 kHz

Supported

Supported

Supported

Supported

Supported

General - Guangxi dialect

8 kHz

Supported

Supported

Supported

Supported

Supported

Mandarin Chinese

Shiyinshi V1 - end-to-end model; Education Content Analysis; Medical Content Analysis; News Media Content Analysis; Entertainment Video Content Analysis; Offline Audio/Video Transcription (Upgraded); New Retail Domain Recognition Model; Travel Domain Recognition Model; Automotive Domain

16 kHz

Supported

Supported

Supported

Supported

Supported

Chinese-English Code-switching

16 kHz

Supported

Supported

Supported

Supported

Not supported

Shiyinshi V1 - end-to-end model

8 kHz

Supported

Supported

Supported

Supported

Supported

Southeast Asian Multilingual

16 kHz

Supported

Not supported

Not supported

Not supported

Not supported

Endpoints

Access type

Description

URL

Public network access (default: China (Shanghai) region)

All servers can use the public network access URL. The public network access URL is set by default in the SDK.

  • China (Shanghai): wss://nls-gateway-cn-shanghai.aliyuncs.com/ws/v1

  • China (Beijing): wss://nls-gateway-cn-beijing.aliyuncs.com/ws/v1

  • China (Shenzhen): wss://nls-gateway-cn-shenzhen.aliyuncs.com/ws/v1

ECS private network access

If you use Alibaba Cloud ECS instances in the China (Shanghai), China (Beijing), or China (Shenzhen) regions, you can use the private network access URL. The classic network of an ECS instance cannot access AnyTunnel, which means you cannot access the speech recognition service over the private network. If you want to use AnyTunnel, create a VPC and access the service from within it.

Important
  • Using the private network access method does not incur public data transfer costs for the ECS instance.

  • China (Shanghai): ws://nls-gateway-cn-shanghai-internal.aliyuncs.com:80/ws/v1

  • China (Beijing): ws://nls-gateway-cn-beijing-internal.aliyuncs.com:80/ws/v1

  • China (Shenzhen): ws://nls-gateway-cn-shenzhen-internal.aliyuncs.com:80/ws/v1

Intelligent access to the nearest region

Short sentence recognition supports access to the nearest region through nls-gateway.aliyuncs.com. The domain name resolves to a nearby server based on the client's location. For example, a request from Beijing resolves to a server in the China (Beijing) region, the same as using nls-gateway-cn-beijing.aliyuncs.com.

Interaction process

The following process applies to WebSocket and SDKs that use this protocol. For RESTful requests, see RESTful API.

image

Every server event includes the recognition task's task_id in its header. Use this ID to correlate requests and responses and troubleshoot issues.

1. Authenticate

The client uses an NLS token to authenticate when establishing a WebSocket connection.

For instructions, see Obtain an access token.

2. Start recognition

The client sends a StartRecognition directive with the recognition parameters. The server validates the request and returns a RecognitionStarted event. Wait for this event before sending audio data.

3. Send data

The client sends binary audio in chunks while receiving server events.

  • If enable_intermediate_result is true, the server can return multiple RecognitionResultChanged events with intermediate results.

  • If enable_intermediate_result is false, the server does not return intermediate results. The client must still handle events such as RecognitionCompleted and TaskFailed.

    Important

    The last intermediate result can differ from the final result. Use the result in RecognitionCompleted as the final result.

4. Stop recognition

After sending the audio, the client sends a StopRecognition directive. The server returns RecognitionCompleted with the final result on success, or TaskFailed on failure. Wait for the server's completion or failure event before closing the connection.

If voice activity detection is enabled, the server can complete recognition when trailing silence exceeds max_end_silence. Audio sent after recognition completes is not recognized.

Request parameters

In an SDK, use the methods of the SpeechRecognizer object to set these parameters. For direct WebSocket requests, set appkey in the request header and the other recognition parameters in the StartRecognition payload.

Parameter

Type

Required

Description

appkey

String

Yes

The Appkey of the project created in the console.

format

String

No

The audio format: pcm, wav, opus, speex, or amr. mp3 and aac are also supported. WAV must use PCM encoding. OPUS and SPEEX must use OGG containers.

sample_rate

Integer

No

The sample rate in Hz. Default: 16000. Valid values: 8000 and 16000. The value must match the audio and project model.

enable_intermediate_result

Boolean

No

Whether to return intermediate recognition results. Default: false.

enable_punctuation_prediction

Boolean

No

Whether to add punctuation during post-processing. Default: false.

enable_inverse_text_normalization

Boolean

No

Whether to enable inverse text normalization (ITN). Setting this parameter to true converts Chinese numerals to Arabic numerals in the output. Default: false.

disfluency

Boolean

No

Whether to filter filler words (output smoothing). Default: false.

customization_id

String

No

The ID of the custom language model. For configuration instructions, see Customize language models.

vocabulary_id

String

No

The ID of the custom hotword list. For configuration instructions, see Customize hotwords.

enable_voice_detection

Boolean

No

Whether to enable voice activity detection (VAD). VAD detects the start and end of speech and excludes noise. Default: false.

max_start_silence

Integer

No

Takes effect only when enable_voice_detection is true. The maximum duration of leading silence, in milliseconds. Recommended range: (0, 60000]. If no speech is detected within this period, the server returns TaskFailed and ends recognition.

max_end_silence

Integer

No

Takes effect only when enable_voice_detection is true. The maximum duration of trailing silence, in milliseconds. Valid range: 200–6000. When trailing silence exceeds this value, the server returns RecognitionCompleted and ends recognition. Subsequent audio is not recognized.

special_word_filter

Object (JSON)

No

The custom word filter configuration. Specify up to 32 words in total. Matching words can be replaced with an empty string or *. See the custom word filtering example below.

enable_multi_thresh_mod

Boolean

No

Takes effect only when enable_voice_detection is true. Set to true to prevent VAD from producing overly long segments. Default: false (disabled).

To recognize audio from a file URL, use the RESTful API.

Custom word filtering

special_word_filter is a JSON object, not a serialized JSON string. Specify up to 32 custom words in total. filter_with_empty replaces matching words with an empty string, and filter_with_signed replaces them with *. Use either or both fields.

The following StartRecognition payload fragment filters words in Chinese recognition results:

{
  "format": "pcm",
  "sample_rate": 16000,
  "special_word_filter": {
    "filter_with_empty": {
      "word_list": ["北京"]
    },
    "filter_with_signed": {
      "word_list": ["测试", "苹果"]
    }
  }
}

Response events

The header contains the following common fields.

Parameter

Type

Description

namespace

String

The namespace: SpeechRecognizer.

name

String

The event name. See the event descriptions below.

status

Integer

The status code. 20000000 indicates success. For other values, see Status codes.

status_text

String

The status message.

task_id

String

The globally unique task ID, which corresponds to the task ID in the client request. Record this value for troubleshooting.

message_id

String

The ID of this server response message.

RecognitionStarted

The server has accepted the start request, and the client can send audio data. This event does not contain recognition results.

RecognitionResultChanged

The payload.result field contains the intermediate recognition result as a string. This event is returned only when enable_intermediate_result is true.

{
  "header": {
    "namespace": "SpeechRecognizer",
    "name": "RecognitionResultChanged",
    "status": 20000000,
    "message_id": "8b756a91809843619c920e176d26****",
    "task_id": "af41104b6806410e9d3f192532aa****",
    "status_text": "Gateway:SUCCESS:Success."
  },
  "payload": {
    "result": "hello today is a beautiful day i would like to buy for two apples thank you"
  }
}

RecognitionCompleted

Recognition has completed successfully. The payload.result field contains the final recognition result as a string.

{
  "header": {
    "namespace": "SpeechRecognizer",
    "name": "RecognitionCompleted",
    "status": 20000000,
    "message_id": "25227a2a43a74259a45685fc1633****",
    "task_id": "af41104b6806410e9d3f192532aa****",
    "status_text": "Gateway:SUCCESS:Success."
  },
  "payload": {
    "result": "hello today is a beautiful day i would like to buy for two apples thank you"
  }
}

The Chinese 8 kHz sentiment recognition model also returns the following fields. Other recognition models are not guaranteed to return them.

Parameter

Type

Description

emo_tag

String

The sentence's sentiment: positive (such as happiness or satisfaction), negative (such as anger, gloom, or disappointment), or neutral (no clear sentiment).

emo_confidence

Double

The confidence score for sentiment recognition, ranging from 0.0 to 1.0. A higher value indicates greater confidence.

TaskFailed

The recognition task failed. Check header.status and header.status_text for the cause. For example, stopping a request without sending audio returns 40000000 and Gateway:CLIENT_ERROR:Empty audio data!.

Status codes

Check both status and status_text when troubleshooting. A status code can have multiple causes, so use the specific status message to identify the issue.

Common error codes

Status code

Status message

Cause

Solution

40000000

The default client error code. This code corresponds to multiple error messages.

Invalid parameters or call sequence.

Check the request parameters, directive order, and specific status message.

40000001

The token 'xxx' has expired.

The token 'xxx' is invalid

The token has expired or is invalid.

Obtain a valid NLS token and retry.

40000002

Gateway:MESSAGE_INVALID:Can't process message in state'FAILED'!

The message is invalid or is not accepted in the current task state.

Check the message structure and directive order. Start a new recognition task after a task fails.

40000003

PARAMETER_INVALID

Failed to decode url params

Invalid parameters.

Check the parameter names, types, and values against the API requirements.

40000005

Gateway:TOO_MANY_REQUESTS:Too many requests!

Too many concurrent requests.

Reduce concurrent requests and check the available concurrency quota.

40000009

Invalid wav header!

The message header is invalid.

If you send a WAV audio file and set the format parameter to wav, check whether the WAV header of the audio file is correct. If the header is incorrect, the server may reject the request.

40000009

Too large wav header!

The WAV header of the transmitted audio is invalid.

You can send the audio stream in a format such as PCM or OPUS. If you use the WAV format, make sure that the WAV header of the audio file contains the correct data length.

40000010

Gateway:FREE_TRIAL_EXPIRED:The free trial has expired!

The trial has expired without a commercial upgrade, or the account has overdue payments.

Check the activation status of short sentence recognition and the account balance in the console.

40010001

Gateway:NAMESPACE_NOT_FOUND:RESTful url path illegal

An unsupported API or parameter was used.

Check the endpoint, namespace, and parameters against the API requirements.

40010003

Gateway:DIRECTIVE_INVALID:[xxx]

An invalid parameter or directive was sent.

Use the specific status message to check the parameters and directive.

40010004

Gateway:CLIENT_DISCONNECT:Client disconnected before task finished!

The client disconnected before the task completed.

Wait for the server's completion or failure event before closing the connection.

40010005

Gateway:TASK_STATE_ERROR:Got stop directive while task is stopping!

The directive is not supported in the current task state.

Check the directive order and avoid sending duplicate stop directives.

40020105

Meta:APPKEY_NOT_EXIST:Appkey not exist!

The Appkey does not exist.

Check the Appkey in the project configuration in the console.

40020106

Meta:APPKEY_UID_MISMATCH:Appkey and user mismatch!

The Appkey and token belong to different accounts.

Use a project Appkey and NLS token from the same account.

403

Forbidden

The token does not exist, has expired, or is invalid.

Use a valid NLS token and obtain a new token before it expires.

41000003

MetaInfo doesn't have end point info

Routing information for the Appkey could not be obtained.

Check the Appkey and make sure it belongs to the same account as the token.

41010101

UNSUPPORTED_SAMPLE_RATE

The sample rate is not supported by the current configuration.

Use 8000 Hz or 16000 Hz audio and make sure the audio, request parameter, and project model match.

50000000

GRPC_ERROR:Grpc error!

A server-side call failed, possibly due to load or network conditions.

Retry the request. If the issue persists, contact technical support and provide the task_id.

50000001

GRPC_ERROR:Grpc error!

A server-side call failed, possibly due to load or network conditions.

Retry the request. If the issue persists, contact technical support and provide the task_id.

52010001

GRPC_ERROR:Grpc error!

A server-side call failed, possibly due to load or network conditions.

Retry the request. If the issue persists, contact technical support and provide the task_id.

Short sentence recognition error codes

Status code

Status message

Cause

Solution

40000000

Gateway:CLIENT_ERROR:Empty audio data!

No audio data was sent.

Send non-empty binary audio data before sending the stop directive.

40000004

Gateway:IDLE_TIMEOUT:Websocket session is idle for too long time

No data was sent after the WebSocket connection was established, and the connection was idle for more than 10 seconds.

Send recognition directives and audio promptly after connecting. Send the stop directive when all audio has been sent.

40010002

Gateway:DIRECTIVE_NOT_SUPPORTED:Directive'SpeechRecognizer.EnhanceRecognition'isnotsupported!

The server does not support the directive.

Check the directive name and use a directive supported by short sentence recognition.

40010003

Gateway:DIRECTIVE_INVALID:Too many items for ‘vocabulary'!(173)

Too many hotwords were specified.

Reduce the hotword count to the limit for the hotword configuration method in use.

40270002

NO_VALID_AUDIO_ERROR

The audio is invalid and no valid text was recognized.

Check that the audio contains clear speech and that its encoding, sample rate, and model match.

41010104

TOO_LONG_SPEECH

The audio exceeds the duration limit for short sentence recognition.

Use audio of up to 60 seconds. For longer audio, use real-time speech recognition.

41010105

SILENT_SPEECH

No speech was detected in silent or noisy audio.

Check the audio. If VAD is enabled, also check the leading silence setting.