Qwen-Audio-TTS non-real-time speech synthesis Java SDK reference
Use the Java SDK to call the Qwen-Audio-TTS speech synthesis API in non-streaming or streaming mode.
ImportantThe features described in this document are available only in the China (Beijing) region.
ImportantAlibaba Cloud Model Studio has released a workspace-specific domain for the China (Beijing) region. The new dedicated domain delivers superior performance and higher stability for inference requests. We recommend migrating from dashscope.aliyuncs.com to {WorkspaceId}.cn-beijing.maas.aliyuncs.com.
Prerequisites
- An Obtain an API key, configured as an environment variable
- The DashScope Java SDK, version 2.22.15 or later. Install the latest version.
HttpSpeechSynthesizer class
Package: com.alibaba.dashscope.audio.http_tts.HttpSpeechSynthesizer
Description: HTTP-based speech synthesis that supports non-streaming and streaming modes.
Constructor
public HttpSpeechSynthesizer()
Creates an HttpSpeechSynthesizer instance with the default configuration. The SDK reads the API key from the DASHSCOPE_API_KEY environment variable. If that variable is not set, it falls back to Constants.apiKey.
callAndReturnAudio() — non-streaming (returns audio data)
Method signature:
public ByteBuffer callAndReturnAudio(HttpSpeechSynthesisParam param) throws ApiException, NoApiKeyException, InputRequiredException
Parameters:
| Parameter | Type | Description |
|---|---|---|
param | HttpSpeechSynthesisParam | The speech synthesis parameter object that contains model, text, voice, and other configurations. |
Returns: ByteBuffer containing the complete audio data. Call remaining() to get the audio size in bytes.
call() — non-streaming (returns audio URL)
Method signature:
public HttpSpeechSynthesisResult call(HttpSpeechSynthesisParam param) throws ApiException, NoApiKeyException, InputRequiredException
Parameters:
| Parameter | Type | Description |
|---|---|---|
param | HttpSpeechSynthesisParam | The speech synthesis parameter object. |
Returns: An HttpSpeechSynthesisResult object. Call getAudioInfo().getUrl() to get the audio download URL. The URL expires after a fixed period. Call getAudioInfo().getExpiresAt() to get the expiration time.
streamCall() — streaming
Method signature:
public void streamCall(HttpSpeechSynthesisParam param, ResultCallback<HttpSpeechSynthesisResult> callback) throws ApiException, NoApiKeyException, InputRequiredException
Parameters:
| Parameter | Type | Description |
|---|---|---|
param | HttpSpeechSynthesisParam | The speech synthesis parameter object. |
callback | ResultCallback<HttpSpeechSynthesisResult> | The callback object. Implement onEvent (receives audio chunks), onComplete (synthesis finished), and onError (error handling). |
This method is asynchronous. Audio data arrives in chunks through the callback as synthesis progresses, which reduces first-packet latency. Use this mode for real-time playback scenarios.
ResultCallback methods:
com.alibaba.dashscope.common.ResultCallback is a generic callback interface provided by the DashScope SDK. Implement the following three methods:
| Method | Parameter | Description |
|---|---|---|
onEvent | HttpSpeechSynthesisResult result | Triggered on each audio chunk. Call result.hasAudioData() to check for audio data and result.getAudioDataSize() to get the chunk size. |
onComplete | None | Triggered when synthesis completes and all audio chunks have been received. |
onError | Exception e | Triggered on synthesis error. Call e.getMessage() to get the error details. |
HttpSpeechSynthesisParam class
Package: com.alibaba.dashscope.audio.http_tts.HttpSpeechSynthesisParam
Use the Builder pattern to construct parameter objects.
Some parameters don't have dedicated Builder methods. Set them using the parameter(String key, Object value) method or the parameters(Map<String, Object>) method inherited from the parent class.
| Method | Type | Required | Description |
|---|---|---|---|
| model(String) | String | Yes | The speech synthesis model. |
| text(String) | String | Yes | The text to synthesize. SSML and LaTeX format inputs are supported. Replace the text content with the corresponding format.
|
| voice(String) | String | Yes | The voice for synthesis. Valid values:
|
| format(String) | String | No | The audio format. Default value: mp3. Valid values:
|
| sampleRate(int) | int | No | The audio sample rate in Hz. Valid values: 8000, 16000, 22050 (default), 24000, 44100, 48000. |
| volume(int) | int | No | The volume level. Default value: 50. Valid values: [0, 100]. |
| rate(float) | float | No | The speech rate. Default value: 1.0. Valid values: [0.5, 2.0]. |
| pitch(float) | float | No | The pitch. Default value: 1.0. Valid values: [0.5, 2.0]. |
enable_ssml | boolean | No | Enables SSML. When set to true, pass SSML-formatted text in the text parameter. For supported SSML tags and usage, see SSML. For the SSML usage restrictions (supported models, voices, and APIs), see Limitations.Default: false. Set |
word_timestamp_enabled | boolean | No | Specifies whether to enable word-level timestamps. Default value: false. Available only in streaming output mode. Cloned voices are supported. For supported system voices, see Qwen-Audio-TTS voice list. Set |
seed | int | No | The random seed for synthesis. Different seeds produce different outputs. With identical model version, text, voice, and other parameters, the same seed reproduces the same output. Default value: 0. Valid values: [0, 65535]. Set |
language_hints | List | No | Important
Specifies the target language for speech synthesis to improve output quality. Use this parameter when numbers, abbreviations, or symbols are not pronounced as expected, or when synthesis quality for secondary languages is poor. Examples:
Valid values
Set |
instruction | String | No | An instruction that controls the synthesis behavior, such as dialect, emotion, or role. For usage details, see Non-real-time speech synthesis. Set |
bit_rate | int | No | The audio bitrate in kbps. Default value: 32. Valid values: [6, 510]. ImportantThis parameter is supported only when Set |
enable_aigc_tag | boolean | No | Specifies whether to embed an AIGC watermark in the generated audio. When set to true, the watermark is embedded in audio files of supported formats (wav/mp3/opus). Default value: false. Set |
aigc_propagator | String | No | Sets the ContentPropagator field in the AIGC watermark, identifying the content propagator. Takes effect only when enable_aigc_tag is true.Default value: Alibaba Cloud UID. Set |
aigc_propagate_id | String | No | Sets the PropagateID field in the AIGC watermark, uniquely identifying a specific propagation action. Takes effect only when enable_aigc_tag is true.Default value: The request ID of the current speech synthesis request. Set |
hot_fix | Map | No | Text hot-fix configuration. Customizes the pronunciation of specified words or substitutes text before speech synthesis. Parameters:
Example: Set |
Code examples
The following examples show non-streaming and streaming calls using the Qwen-Audio-TTS speech synthesis API. Before running the examples, set the DASHSCOPE_API_KEY environment variable.
ImportantUse a voice that matches the model and supports the target language. When switching models, update the voice accordingly. See Qwen-Audio-TTS voice list.
Non-streaming
Non-streaming calls wait for synthesis to complete, then return the full result. Two return types are available:
callAndReturnAudio(): Returns audio binary data (ByteBuffer). Use this to save or process the audio directly.call(): Returns an audio URL. Use this to download the audio from a URL.
import com.alibaba.dashscope.audio.http_tts.AudioInfo;
import com.alibaba.dashscope.audio.http_tts.HttpSpeechSynthesisParam;
import com.alibaba.dashscope.audio.http_tts.HttpSpeechSynthesisResult;
import com.alibaba.dashscope.audio.http_tts.HttpSpeechSynthesizer;
import com.alibaba.dashscope.exception.ApiException;
import com.alibaba.dashscope.exception.InputRequiredException;
import com.alibaba.dashscope.exception.NoApiKeyException;
import com.alibaba.dashscope.utils.Constants;
import java.io.FileOutputStream;
import java.io.IOException;
import java.nio.ByteBuffer;
public class QwenAudioTtsSyncExample {
static {
// The following is the configuration for the China (Beijing) region. Replace "{WorkspaceId}" with your actual workspace ID. The configuration varies by region.
Constants.baseHttpApiUrl = "https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1";
}
/**
* Non-streaming example 1: Returns audio data (ByteBuffer)
* Calls callAndReturnAudio, which blocks until synthesis completes and returns the full audio data.
*/
public static void syncCallReturnAudio() {
HttpSpeechSynthesizer synthesizer = new HttpSpeechSynthesizer();
HttpSpeechSynthesisParam param =
HttpSpeechSynthesisParam.builder()
.model("qwen-audio-3.0-tts-flash") // When switching models, also switch to a compatible voice version
.text("Behind my house is a big garden.")
.voice("longanhuan_v3.6")
.format("wav")
.sampleRate(24000)
// If the environment variable is not set, replace with: apiKey("sk-xxx")
.apiKey(System.getenv("DASHSCOPE_API_KEY"))
// Set additional parameters using the parameter method
// .parameter("seed", 1234)
// .parameter("enable_ssml", true)
.build();
try {
ByteBuffer audioData = synthesizer.callAndReturnAudio(param);
if (audioData != null && audioData.hasRemaining()) {
byte[] bytes = new byte[audioData.remaining()];
audioData.get(bytes);
try (FileOutputStream fos = new FileOutputStream("sync_output.wav")) {
fos.write(bytes);
System.out.println("Audio saved to sync_output.wav, size: " + bytes.length + " bytes");
} catch (IOException e) {
System.err.println("Failed to save audio: " + e.getMessage());
}
}
} catch (ApiException | NoApiKeyException | InputRequiredException e) {
System.err.println("Synthesis failed: " + e.getMessage());
}
System.exit(0);
}
/**
* Non-streaming example 2: Returns an audio URL
* Calls call, which returns a result object containing an audio URL for downloading.
*/
public static void syncCallReturnUrl() {
HttpSpeechSynthesizer synthesizer = new HttpSpeechSynthesizer();
HttpSpeechSynthesisParam param =
HttpSpeechSynthesisParam.builder()
.model("qwen-audio-3.0-tts-flash") // When switching models, also switch to a compatible voice version
.text("Behind my house is a big garden.")
.voice("longanhuan_v3.6")
.format("wav")
.sampleRate(24000)
// If the environment variable is not set, replace with: apiKey("sk-xxx")
.apiKey(System.getenv("DASHSCOPE_API_KEY"))
.build();
try {
HttpSpeechSynthesisResult result = synthesizer.call(param);
System.out.println("Request ID: " + result.getRequestId());
if (result.hasAudioUrl()) {
AudioInfo audioInfo = result.getAudioInfo();
System.out.println("Audio URL: " + audioInfo.getUrl());
System.out.println("Expires At: " + audioInfo.getExpiresAt());
System.out.println("Remaining Time: " + audioInfo.getRemainingSeconds() + " seconds");
}
} catch (ApiException | NoApiKeyException | InputRequiredException e) {
System.err.println("Synthesis failed: " + e.getMessage());
}
}
public static void main(String[] args) {
// Non-streaming example 1: Returns audio data (ByteBuffer)
// syncCallReturnAudio();
// Non-streaming example 2: Returns an audio URL
syncCallReturnUrl();
}
}
Streaming
Streaming calls return audio data in chunks through a callback as synthesis progresses. Use this mode for real-time playback where low first-packet latency matters.
import com.alibaba.dashscope.audio.http_tts.HttpSpeechSynthesisParam;
import com.alibaba.dashscope.audio.http_tts.HttpSpeechSynthesisResult;
import com.alibaba.dashscope.audio.http_tts.HttpSpeechSynthesizer;
import com.alibaba.dashscope.common.ResultCallback;
import com.alibaba.dashscope.exception.ApiException;
import com.alibaba.dashscope.exception.InputRequiredException;
import com.alibaba.dashscope.exception.NoApiKeyException;
import com.alibaba.dashscope.utils.Constants;
import java.io.FileOutputStream;
import java.io.IOException;
import java.util.concurrent.CountDownLatch;
public class QwenAudioTtsStreamExample {
static {
// The following is the configuration for the China (Beijing) region. Replace "{WorkspaceId}" with your actual workspace ID. The configuration varies by region.
Constants.baseHttpApiUrl = "https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1";
}
public static void streamCallWithCallback() {
HttpSpeechSynthesizer synthesizer = new HttpSpeechSynthesizer();
HttpSpeechSynthesisParam param =
HttpSpeechSynthesisParam.builder()
.model("qwen-audio-3.0-tts-flash") // When switching models, also switch to a compatible voice version
.text("The weather is great today, perfect for going out.")
.voice("longanhuan_v3.6")
.format("wav")
.sampleRate(24000)
// If the environment variable is not set, replace with: apiKey("sk-xxx")
.apiKey(System.getenv("DASHSCOPE_API_KEY"))
.build();
CountDownLatch latch = new CountDownLatch(1);
// Create the output file stream
try (FileOutputStream fos = new FileOutputStream("output.wav")) {
synthesizer.streamCall(param,
new ResultCallback<HttpSpeechSynthesisResult>() {
private int chunkCount = 0;
@Override
public void onEvent(HttpSpeechSynthesisResult result) {
chunkCount++;
if (result.hasAudioData()) {
System.out.println("Received chunk #" + chunkCount
+ ", size: " + result.getAudioDataSize() + " bytes");
try {
fos.write(result.getAudioData());
} catch (IOException e) {
System.err.println("Failed to write audio data: " + e.getMessage());
}
}
if (result.getRequestId() != null) {
System.out.println("Request ID: " + result.getRequestId());
}
}
@Override
public void onComplete() {
System.out.println("Synthesis completed, total chunks: " + chunkCount);
System.out.println("Audio saved to output.wav");
latch.countDown();
}
@Override
public void onError(Exception e) {
System.err.println("Error during synthesis: " + e.getMessage());
latch.countDown();
}
});
latch.await();
} catch (ApiException | NoApiKeyException | InputRequiredException | InterruptedException e) {
System.err.println("Failed: " + e.getMessage());
} catch (IOException e) {
System.err.println("Failed to create output file: " + e.getMessage());
}
}
public static void main(String[] args) {
streamCallWithCallback();
System.exit(0);
}
}