Qwen-Audio-TTS non-real-time speech synthesis Java SDK reference

Updated at:

Use the Java SDK to call the Qwen-Audio-TTS speech synthesis API in non-streaming or streaming mode.

ImportantThe features described in this document are available only in the China (Beijing) region.

ImportantAlibaba Cloud Model Studio has released a workspace-specific domain for the China (Beijing) region. The new dedicated domain delivers superior performance and higher stability for inference requests. We recommend migrating from dashscope.aliyuncs.com to {WorkspaceId}.cn-beijing.maas.aliyuncs.com.

Prerequisites

HttpSpeechSynthesizer class

Package: com.alibaba.dashscope.audio.http_tts.HttpSpeechSynthesizer

Description: HTTP-based speech synthesis that supports non-streaming and streaming modes.

Constructor

public HttpSpeechSynthesizer()

Creates an HttpSpeechSynthesizer instance with the default configuration. The SDK reads the API key from the DASHSCOPE_API_KEY environment variable. If that variable is not set, it falls back to Constants.apiKey.

callAndReturnAudio() — non-streaming (returns audio data)

Method signature:

public ByteBuffer callAndReturnAudio(HttpSpeechSynthesisParam param) throws ApiException, NoApiKeyException, InputRequiredException

Parameters:

ParameterTypeDescription
paramHttpSpeechSynthesisParamThe speech synthesis parameter object that contains model, text, voice, and other configurations.

Returns: ByteBuffer containing the complete audio data. Call remaining() to get the audio size in bytes.

call() — non-streaming (returns audio URL)

Method signature:

public HttpSpeechSynthesisResult call(HttpSpeechSynthesisParam param) throws ApiException, NoApiKeyException, InputRequiredException

Parameters:

ParameterTypeDescription
paramHttpSpeechSynthesisParamThe speech synthesis parameter object.

Returns: An HttpSpeechSynthesisResult object. Call getAudioInfo().getUrl() to get the audio download URL. The URL expires after a fixed period. Call getAudioInfo().getExpiresAt() to get the expiration time.

streamCall() — streaming

Method signature:

public void streamCall(HttpSpeechSynthesisParam param, ResultCallback<HttpSpeechSynthesisResult> callback) throws ApiException, NoApiKeyException, InputRequiredException

Parameters:

ParameterTypeDescription
paramHttpSpeechSynthesisParamThe speech synthesis parameter object.
callbackResultCallback<HttpSpeechSynthesisResult>The callback object. Implement onEvent (receives audio chunks), onComplete (synthesis finished), and onError (error handling).

This method is asynchronous. Audio data arrives in chunks through the callback as synthesis progresses, which reduces first-packet latency. Use this mode for real-time playback scenarios.

ResultCallback methods:

com.alibaba.dashscope.common.ResultCallback is a generic callback interface provided by the DashScope SDK. Implement the following three methods:

MethodParameterDescription
onEventHttpSpeechSynthesisResult resultTriggered on each audio chunk. Call result.hasAudioData() to check for audio data and result.getAudioDataSize() to get the chunk size.
onCompleteNoneTriggered when synthesis completes and all audio chunks have been received.
onErrorException eTriggered on synthesis error. Call e.getMessage() to get the error details.

HttpSpeechSynthesisParam class

Package: com.alibaba.dashscope.audio.http_tts.HttpSpeechSynthesisParam

Use the Builder pattern to construct parameter objects.

Some parameters don't have dedicated Builder methods. Set them using the parameter(String key, Object value) method or the parameters(Map<String, Object>) method inherited from the parent class.

MethodTypeRequiredDescription
model(String)StringYesThe speech synthesis model.
text(String)StringYesThe text to synthesize.

SSML and LaTeX format inputs are supported. Replace the text content with the corresponding format.

voice(String)StringYesThe voice for synthesis.

Valid values:

format(String)StringNoThe audio format.

Default value: mp3.

Valid values:

  • mp3
  • pcm
  • wav
  • opus
sampleRate(int)intNoThe audio sample rate in Hz.

Valid values: 8000, 16000, 22050 (default), 24000, 44100, 48000.

volume(int)intNoThe volume level.

Default value: 50.

Valid values: [0, 100].

rate(float)floatNoThe speech rate.

Default value: 1.0.

Valid values: [0.5, 2.0].

pitch(float)floatNoThe pitch.

Default value: 1.0.

Valid values: [0.5, 2.0].

enable_ssmlbooleanNoEnables SSML. When set to true, pass SSML-formatted text in the text parameter. For supported SSML tags and usage, see SSML. For the SSML usage restrictions (supported models, voices, and APIs), see Limitations.

Default: false.

Set enable_ssml using the HttpSpeechSynthesisParam instance's parameter method or parameters method:

HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("<speak>Hello</speak>")
    .voice("longanhuan_v3.6")
    .parameter("enable_ssml", true)
    .build();
HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("<speak>Hello</speak>")
    .voice("longanhuan_v3.6")
    .parameters(Collections.singletonMap("enable_ssml", true))
    .build();
word_timestamp_enabledbooleanNoSpecifies whether to enable word-level timestamps.

Default value: false.

Available only in streaming output mode. Cloned voices are supported. For supported system voices, see Qwen-Audio-TTS voice list.

Set word_timestamp_enabled using the HttpSpeechSynthesisParam instance's parameter method or parameters method:

HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("Behind my house is a big garden.")
    .voice("longanhuan_v3.6")
    .parameter("word_timestamp_enabled", true)
    .build();
HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("Behind my house is a big garden.")
    .voice("longanhuan_v3.6")
    .parameters(Collections.singletonMap("word_timestamp_enabled", true))
    .build();
seedintNoThe random seed for synthesis. Different seeds produce different outputs. With identical model version, text, voice, and other parameters, the same seed reproduces the same output.

Default value: 0.

Valid values: [0, 65535].

Set seed using the HttpSpeechSynthesisParam instance's parameter method or parameters method:

HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("Behind my house is a big garden.")
    .voice("longanhuan_v3.6")
    .parameter("seed", 1234)
    .build();
HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("Behind my house is a big garden.")
    .voice("longanhuan_v3.6")
    .parameters(Collections.singletonMap("seed", 1234))
    .build();
language_hintsListNo

Important

  • This parameter is an array, but the current version processes only the first element. Pass a single value.
  • This parameter specifies the target language for speech synthesis and is unrelated to the language of the sample audio used in voice cloning. To set the source language for a voice cloning task, see the Voice Cloning API reference.

Specifies the target language for speech synthesis to improve output quality.

Use this parameter when numbers, abbreviations, or symbols are not pronounced as expected, or when synthesis quality for secondary languages is poor. Examples:

  • A number isn't read as expected: "hello, this is 110" is read as "hello, this is one one zero" instead of "hello, this is yao-yao-ling"
  • A symbol isn't read correctly: "@" is read as "ai-te" instead of "at"
  • Minor language synthesis sounds unnatural

Valid values

  • zh: Chinese
  • en: English
  • fr: French
  • de: German
  • ja: Japanese
  • ko: Korean
  • ru: Russian
  • pt: Portuguese
  • th: Thai
  • id: Indonesian
  • vi: Vietnamese
  • es: Spanish
  • it: Italian
  • ms: Malaysian
  • fil: Filipino
  • ar: Arabic

Set language_hints using the HttpSpeechSynthesisParam instance's parameter method or parameters method:

HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("Behind my house is a big garden.")
    .voice("longanhuan_v3.6")
    .parameter("language_hints", Arrays.asList("zh"))
    .build();
HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("Behind my house is a big garden.")
    .voice("longanhuan_v3.6")
    .parameters(Collections.singletonMap("language_hints", Arrays.asList("zh")))
    .build();
instructionStringNoAn instruction that controls the synthesis behavior, such as dialect, emotion, or role.

For usage details, see Non-real-time speech synthesis.

Set instruction using the HttpSpeechSynthesisParam instance's parameter method or parameters method:

HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("Behind my house is a big garden.")
    .voice("longanhuan_v3.6")
    .parameter("instruction", "Speak in a very happy tone.")
    .build();
HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("Behind my house is a big garden.")
    .voice("longanhuan_v3.6")
    .parameters(Collections.singletonMap("instruction", "Speak in a very happy tone."))
    .build();
bit_rateintNoThe audio bitrate in kbps.

Default value: 32.

Valid values: [6, 510].

ImportantThis parameter is supported only when format is set to opus.

Set bit_rate using the HttpSpeechSynthesisParam instance's parameter method or parameters method:

HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("Behind my house is a big garden.")
    .voice("longanhuan_v3.6")
    .format("opus")
    .parameter("bit_rate", 32)
    .build();
HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("Behind my house is a big garden.")
    .voice("longanhuan_v3.6")
    .format("opus")
    .parameters(Collections.singletonMap("bit_rate", 32))
    .build();
enable_aigc_tagbooleanNoSpecifies whether to embed an AIGC watermark in the generated audio. When set to true, the watermark is embedded in audio files of supported formats (wav/mp3/opus).

Default value: false.

Set enable_aigc_tag using the HttpSpeechSynthesisParam instance's parameter method or parameters method:

HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("Behind my house is a big garden.")
    .voice("longanhuan_v3.6")
    .parameter("enable_aigc_tag", true)
    .build();
HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("Behind my house is a big garden.")
    .voice("longanhuan_v3.6")
    .parameters(Collections.singletonMap("enable_aigc_tag", true))
    .build();
aigc_propagatorStringNoSets the ContentPropagator field in the AIGC watermark, identifying the content propagator. Takes effect only when enable_aigc_tag is true.

Default value: Alibaba Cloud UID.

Set aigc_propagator using the HttpSpeechSynthesisParam instance's parameter method or parameters method:

HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("Behind my house is a big garden.")
    .voice("longanhuan_v3.6")
    .parameter("enable_aigc_tag", true)
    .parameter("aigc_propagator", "xxxx")
    .build();
Map<String, Object> map = new HashMap<>();
map.put("enable_aigc_tag", true);
map.put("aigc_propagator", "xxxx");

HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("Behind my house is a big garden.")
    .voice("longanhuan_v3.6")
    .parameters(map)
    .build();
aigc_propagate_idStringNoSets the PropagateID field in the AIGC watermark, uniquely identifying a specific propagation action. Takes effect only when enable_aigc_tag is true.

Default value: The request ID of the current speech synthesis request.

Set aigc_propagate_id using the HttpSpeechSynthesisParam instance's parameter method or parameters method:

HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("Behind my house is a big garden.")
    .voice("longanhuan_v3.6")
    .parameter("enable_aigc_tag", true)
    .parameter("aigc_propagate_id", "xxxx")
    .build();
Map<String, Object> map = new HashMap<>();
map.put("enable_aigc_tag", true);
map.put("aigc_propagate_id", "xxxx");

HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("Behind my house is a big garden.")
    .voice("longanhuan_v3.6")
    .parameters(map)
    .build();
hot_fixMapNoText hot-fix configuration. Customizes the pronunciation of specified words or substitutes text before speech synthesis.

Parameters:

  • pronunciation: Custom pronunciation. Provides pinyin for given words to correct cases where the default pronunciation is inaccurate.
  • replace: Text substitution. Replaces specified words with target text before speech synthesis; the substituted text is what gets synthesized.

Example:

"hot_fix": {
  "pronunciation": [
    {"weather": "tian1 qi4"}
  ],
  "replace": [
    {"today": "gold day"}
  ]
}

Set hot_fix using the HttpSpeechSynthesisParam instance's parameter method or parameters method:

Map<String, Object> hotFix = new HashMap<>();

List<Map<String, String>> pronunciation = new ArrayList<>();
Map<String, String> pronItem = new HashMap<>();
pronItem.put("weather", "tian1 qi4");
pronunciation.add(pronItem);
hotFix.put("pronunciation", pronunciation);

List<Map<String, String>> replace = new ArrayList<>();
Map<String, String> replaceItem = new HashMap<>();
replaceItem.put("today", "gold day");
replace.add(replaceItem);
hotFix.put("replace", replace);

HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("The weather is really nice today.")
    .voice("longanhuan_v3.6")
    .parameter("hot_fix", hotFix)
    .build();
// Build the hotFix object as shown above
HttpSpeechSynthesisParam param = HttpSpeechSynthesisParam.builder()
    .model("qwen-audio-3.0-tts-flash")
    .text("The weather is really nice today.")
    .voice("longanhuan_v3.6")
    .parameters(Collections.singletonMap("hot_fix", hotFix))
    .build();

Code examples

The following examples show non-streaming and streaming calls using the Qwen-Audio-TTS speech synthesis API. Before running the examples, set the DASHSCOPE_API_KEY environment variable.

ImportantUse a voice that matches the model and supports the target language. When switching models, update the voice accordingly. See Qwen-Audio-TTS voice list.

Non-streaming

Non-streaming calls wait for synthesis to complete, then return the full result. Two return types are available:

  • callAndReturnAudio(): Returns audio binary data (ByteBuffer). Use this to save or process the audio directly.
  • call(): Returns an audio URL. Use this to download the audio from a URL.
import com.alibaba.dashscope.audio.http_tts.AudioInfo;
import com.alibaba.dashscope.audio.http_tts.HttpSpeechSynthesisParam;
import com.alibaba.dashscope.audio.http_tts.HttpSpeechSynthesisResult;
import com.alibaba.dashscope.audio.http_tts.HttpSpeechSynthesizer;
import com.alibaba.dashscope.exception.ApiException;
import com.alibaba.dashscope.exception.InputRequiredException;
import com.alibaba.dashscope.exception.NoApiKeyException;
import com.alibaba.dashscope.utils.Constants;

import java.io.FileOutputStream;
import java.io.IOException;
import java.nio.ByteBuffer;

public class QwenAudioTtsSyncExample {
    static {
        // The following is the configuration for the China (Beijing) region. Replace "{WorkspaceId}" with your actual workspace ID. The configuration varies by region.
        Constants.baseHttpApiUrl = "https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1";
    }

    /**
     * Non-streaming example 1: Returns audio data (ByteBuffer)
     * Calls callAndReturnAudio, which blocks until synthesis completes and returns the full audio data.
     */
    public static void syncCallReturnAudio() {
        HttpSpeechSynthesizer synthesizer = new HttpSpeechSynthesizer();

        HttpSpeechSynthesisParam param =
            HttpSpeechSynthesisParam.builder()
                .model("qwen-audio-3.0-tts-flash")  // When switching models, also switch to a compatible voice version
                .text("Behind my house is a big garden.")
                .voice("longanhuan_v3.6")
                .format("wav")
                .sampleRate(24000)
                // If the environment variable is not set, replace with: apiKey("sk-xxx")
                .apiKey(System.getenv("DASHSCOPE_API_KEY"))
                // Set additional parameters using the parameter method
                // .parameter("seed", 1234)
                // .parameter("enable_ssml", true)
                .build();

        try {
            ByteBuffer audioData = synthesizer.callAndReturnAudio(param);
            if (audioData != null && audioData.hasRemaining()) {
                byte[] bytes = new byte[audioData.remaining()];
                audioData.get(bytes);

                try (FileOutputStream fos = new FileOutputStream("sync_output.wav")) {
                    fos.write(bytes);
                    System.out.println("Audio saved to sync_output.wav, size: " + bytes.length + " bytes");
                } catch (IOException e) {
                    System.err.println("Failed to save audio: " + e.getMessage());
                }
            }
        } catch (ApiException | NoApiKeyException | InputRequiredException e) {
            System.err.println("Synthesis failed: " + e.getMessage());
        }
        System.exit(0);
    }

    /**
     * Non-streaming example 2: Returns an audio URL
     * Calls call, which returns a result object containing an audio URL for downloading.
     */
    public static void syncCallReturnUrl() {
        HttpSpeechSynthesizer synthesizer = new HttpSpeechSynthesizer();

        HttpSpeechSynthesisParam param =
            HttpSpeechSynthesisParam.builder()
                .model("qwen-audio-3.0-tts-flash")  // When switching models, also switch to a compatible voice version
                .text("Behind my house is a big garden.")
                .voice("longanhuan_v3.6")
                .format("wav")
                .sampleRate(24000)
                // If the environment variable is not set, replace with: apiKey("sk-xxx")
                .apiKey(System.getenv("DASHSCOPE_API_KEY"))
                .build();

        try {
            HttpSpeechSynthesisResult result = synthesizer.call(param);
            System.out.println("Request ID: " + result.getRequestId());

            if (result.hasAudioUrl()) {
                AudioInfo audioInfo = result.getAudioInfo();
                System.out.println("Audio URL: " + audioInfo.getUrl());
                System.out.println("Expires At: " + audioInfo.getExpiresAt());
                System.out.println("Remaining Time: " + audioInfo.getRemainingSeconds() + " seconds");
            }
        } catch (ApiException | NoApiKeyException | InputRequiredException e) {
            System.err.println("Synthesis failed: " + e.getMessage());
        }
    }

    public static void main(String[] args) {
        // Non-streaming example 1: Returns audio data (ByteBuffer)
        // syncCallReturnAudio();
        // Non-streaming example 2: Returns an audio URL
        syncCallReturnUrl();
    }
}

Streaming

Streaming calls return audio data in chunks through a callback as synthesis progresses. Use this mode for real-time playback where low first-packet latency matters.

import com.alibaba.dashscope.audio.http_tts.HttpSpeechSynthesisParam;
import com.alibaba.dashscope.audio.http_tts.HttpSpeechSynthesisResult;
import com.alibaba.dashscope.audio.http_tts.HttpSpeechSynthesizer;
import com.alibaba.dashscope.common.ResultCallback;
import com.alibaba.dashscope.exception.ApiException;
import com.alibaba.dashscope.exception.InputRequiredException;
import com.alibaba.dashscope.exception.NoApiKeyException;
import com.alibaba.dashscope.utils.Constants;

import java.io.FileOutputStream;
import java.io.IOException;
import java.util.concurrent.CountDownLatch;

public class QwenAudioTtsStreamExample {
    static {
        // The following is the configuration for the China (Beijing) region. Replace "{WorkspaceId}" with your actual workspace ID. The configuration varies by region.
        Constants.baseHttpApiUrl = "https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1";
    }

    public static void streamCallWithCallback() {
        HttpSpeechSynthesizer synthesizer = new HttpSpeechSynthesizer();

        HttpSpeechSynthesisParam param =
                HttpSpeechSynthesisParam.builder()
                        .model("qwen-audio-3.0-tts-flash")  // When switching models, also switch to a compatible voice version
                        .text("The weather is great today, perfect for going out.")
                        .voice("longanhuan_v3.6")
                        .format("wav")
                        .sampleRate(24000)
                        // If the environment variable is not set, replace with: apiKey("sk-xxx")
                        .apiKey(System.getenv("DASHSCOPE_API_KEY"))
                        .build();

        CountDownLatch latch = new CountDownLatch(1);

        // Create the output file stream
        try (FileOutputStream fos = new FileOutputStream("output.wav")) {

            synthesizer.streamCall(param,
                    new ResultCallback<HttpSpeechSynthesisResult>() {
                        private int chunkCount = 0;

                        @Override
                        public void onEvent(HttpSpeechSynthesisResult result) {
                            chunkCount++;
                            if (result.hasAudioData()) {
                                System.out.println("Received chunk #" + chunkCount
                                        + ", size: " + result.getAudioDataSize() + " bytes");
                                try {
                                    fos.write(result.getAudioData());
                                } catch (IOException e) {
                                    System.err.println("Failed to write audio data: " + e.getMessage());
                                }
                            }
                            if (result.getRequestId() != null) {
                                System.out.println("Request ID: " + result.getRequestId());
                            }
                        }

                        @Override
                        public void onComplete() {
                            System.out.println("Synthesis completed, total chunks: " + chunkCount);
                            System.out.println("Audio saved to output.wav");
                            latch.countDown();
                        }

                        @Override
                        public void onError(Exception e) {
                            System.err.println("Error during synthesis: " + e.getMessage());
                            latch.countDown();
                        }
                    });

            latch.await();

        } catch (ApiException | NoApiKeyException | InputRequiredException | InterruptedException e) {
            System.err.println("Failed: " + e.getMessage());
        } catch (IOException e) {
            System.err.println("Failed to create output file: " + e.getMessage());
        }
    }

    public static void main(String[] args) {
        streamCallWithCallback();
        System.exit(0);
    }
}