Sambert client events

Updated at:

For model details and recommendations, see Speech synthesis.

ImportantSambert is available only in the China (Beijing) region.

ImportantAlibaba Cloud Model Studio has released a workspace-specific domain for the China (Beijing) region: https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com. The new dedicated domain delivers superior performance and higher stability for inference requests. We recommend migrating from https://dashscope.aliyuncs.com to the new domain.

{WorkspaceId} is your workspace ID, which can be found on the Workspace Details page in the Alibaba Cloud Model Studio console. The existing domain remains fully functional.

run-task

Description: Starts a speech synthesis task and sends the text to synthesize in a single request.

When to send: Send this event immediately after establishing the WebSocket connection.

Response event: The server returns a task-started event.

ImportantSambert doesn't support streaming input or the continue-task and finish-task events. Send the full text in the input.text field of the run-task event.

headerobject(required)

Attributes

actionstring(required)

The event type. Must be run-task.

task_idstring(required)

A client-generated task ID in UUID format, used to correlate subsequent events.

streamingstring(required)

Must be out.

payloadobject(required)

Attributes

task_groupstring(required)

The task group. Must be audio.

taskstring(required)

The task type. Must be tts.

functionstring(required)

The function type. Must be SpeechSynthesizer.

modelstring(required)

The model name, such as sambert-zhichu-v1.

inputobject(required)

The input data. Contains a text field for the text to synthesize.

textstring(required)

The text to synthesize.

parametersobject(required)

Speech synthesis parameters.

Attributes

text_typestring(required)

Must be PlainText.

formatstring(optional)

The audio format.

Valid values:

  • pcm
  • wav (default)
  • mp3

sample_rateinteger(optional)

The audio sample rate, in Hz.

Valid values: 8000, 16000 (default), 22050, and 24000.

volumeinteger(optional)

The volume level.

Default value: 50.

Valid values: 0 to 100.

ratefloat(optional)

The speech rate.

Default value: 1.0.

Valid values: 0.5 to 2.0.

pitchfloat(optional)

The pitch level.

Default value: 1.0.

Valid values: 0.5 to 2.0.

word_timestamp_enabledboolean(optional)

Specifies whether to enable word-level timestamps.

Default value: false.

Applies to all Sambert models.

phoneme_timestamp_enabledboolean(optional)

Specifies whether to enable phoneme-level timestamps.

Default value: false.

word_timestamp_enabled must be true before you enable this parameter.

{
    "header": {
        "action": "run-task",
        "task_id": "2bf83b9a-baeb-4fda-8d9a-xxxxxxxxxxxx",
        "streaming": "out"
    },
    "payload": {
        "task_group": "audio",
        "task": "tts",
        "function": "SpeechSynthesizer",
        "model": "sambert-zhichu-v1",
        "parameters": {
            "text_type": "PlainText",
            "format": "wav",
            "sample_rate": 16000,
            "volume": 50,
            "rate": 1.0,
            "pitch": 1.0,
            "word_timestamp_enabled": true,
            "phoneme_timestamp_enabled": true
        },
        "input": {
            "text": "Before my bed, moonlight shines bright, I suspect it's frost upon the ground."
        }
    }
}