Audio detection and splitting
This topic describes the silence detection, voice activity detection, and audio splitting operators in the Python DataFrame API.
Limits
-
Supported only by Realtime Compute engine VVR 11.8 and later.
-
All three audio splitting operators are UDTFs and must be used with
DataFrame.join_lateral. -
This topic uses audio segment reference (AUDIO_CLIP_REF_TYPE) and decoded audio waveform (AUDIO_WAVEFORM_TYPE) to refer to the DataFrame API types defined in
pyflink.multimodal.types.
Operator summary
|
Category |
Operator |
Description |
|
Audio detection |
Detects low-amplitude silence intervals in an audio waveform. |
|
|
Uses WebRTC VAD to detect voice activity intervals in an audio waveform. |
||
|
Audio splitting |
Splits audio sequentially into segments of a fixed duration and outputs actual audio waveforms or segment references based on |
|
|
Splits audio by start and end time ranges provided by the caller and outputs audio waveforms or segment references in the original order. Suitable when subtitles, logs, meeting minutes, or manually annotated timelines are already available, for example extracting audio by subtitle timestamps or locating utterance segments by meeting minutes. |
||
|
Based on existing voice activity ranges, pads, merges, filters out short segments, and splits long ones, then outputs audio waveforms or segment references. Suitable for calls, interviews, or podcasts that contain a lot of silence, for example combined with |
Time range type
Silence detection and voice activity detection return the following DataFrame API type:
DataType.list(
DataType.struct({
"start_ms": DataType.int64(),
"end_ms": DataType.int64(),
"duration_ms": DataType.int64()
})
)
Time ranges use the half-open interval [start_ms, end_ms), in milliseconds.
Common runtime parameters
Function signatures retain the runtime parameters that each operator actually supports. To avoid repetition, the operator parameter tables describe only business parameters. Runtime parameters are described uniformly as follows.
|
Parameter |
Type |
Default |
Scope |
Description |
|
|
|
|
All operators on this page |
The parallelism of the UDF or UDTF. |
Audio detection
audio_silence_detection
Detects low-amplitude silence intervals in an audio waveform.
Input type: AUDIO_WAVEFORM_TYPE, a decoded audio waveform column.
Function signature:
audio_silence_detection(
*columns,
threshold_db,
min_silence_ms=0,
concurrency=None
)
|
Parameter |
Type |
Default |
Description |
|
|
|
Required |
Silence threshold relative to full scale. Must be less than or equal to 0. Frames with energy at or below this threshold are considered silent. |
|
|
|
|
Minimum silence duration, in milliseconds. Shorter intervals are not returned. |
The operator determines silence as follows:
-
Multi-channel audio is first averaged across channels at each point in time to produce mono audio.
-
The audio is divided into short overlapping frames, and RMS (root mean square) is used to estimate the average volume of each frame.
-
The average volume is expressed in dBFS. Values closer to
0indicate louder audio; smaller values indicate quieter audio. -
Stretches that stay at or below
threshold_dbfor at leastmin_silence_msare returned as silence intervals.
For example, with threshold_db=-45.0, a frame with an average volume of -30 dBFS is not considered silent, while -50 dBFS is.
The operator detects silence based on volume alone. It does not use VAD or ASR models, so a "silence interval" is not equivalent to a "non-speech interval". A null input returns None.
Return type:
DataType.list(
DataType.struct({
"start_ms": DataType.int64(), # Start time of the interval.
"end_ms": DataType.int64(), # End time of the interval.
"duration_ms": DataType.int64() # Duration of the interval.
})
)
from pyflink.dataframe import col
from pyflink.multimodal.operators import audio_silence_detection
result = df.with_column(
"silence_ranges",
audio_silence_detection(
col("waveform"),
threshold_db=-45.0,
min_silence_ms=300
)
)
audio_detect_speech
Uses WebRTC VAD to detect voice activity intervals in an audio waveform.
Input type: AUDIO_WAVEFORM_TYPE, a mono audio waveform column. The sample rate must be 8000, 16000, 32000, or 48000 Hz.
Function signature:
audio_detect_speech(
*columns,
aggressiveness=0,
frame_ms=30,
min_speech_ms=0,
merge_gap_ms=0,
max_segments=1024,
concurrency=None
)
|
Parameter |
Type |
Default |
Description |
|
|
|
|
WebRTC VAD aggressiveness mode. Valid values: |
|
|
|
|
VAD frame length. Must be 10, 20, or 30 milliseconds. |
|
|
|
|
Discards speech intervals shorter than this duration, in milliseconds. |
|
|
|
|
Merges adjacent voice activity intervals when the gap between them is at most this value, in milliseconds. |
|
|
|
|
Maximum number of voice activity intervals to return. Must be a positive integer. |
Results are sorted by time. A null input returns None.
Return type:
DataType.list(
DataType.struct({
"start_ms": DataType.int64(), # Start timestamp of the interval.
"end_ms": DataType.int64(), # End timestamp of the interval.
"duration_ms": DataType.int64() # Duration of the interval.
})
)
from pyflink.multimodal.operators import (
audio_detect_speech,
audio_standardize,
)
# Standardize the audio first.
speech_ready = audio_standardize(
col("waveform"),
sample_rate=16000,
channels=1
)
# Detect speech intervals.
result = df.with_column(
"speech_ranges",
audio_detect_speech(
speech_ready,
aggressiveness=2,
min_speech_ms=200,
merge_gap_ms=100
)
)
Audio splitting
audio_split_by_duration
Splits audio into segments of a fixed duration.
Input type:
-
segment_type="audio":AUDIO_WAVEFORM_TYPE; -
segment_type="ref":DataType.string()(audio URL) orAUDIO_CLIP_REF_TYPE, the audio column to split.
The input type must match segment_type. No rows are emitted when the input is null.
Function signature:
audio_split_by_duration(
*columns,
segment_duration_ms,
segment_type="audio",
max_segments=1024,
concurrency=None
)
|
Parameter |
Type |
Default |
Description |
|
|
|
Required |
Target duration per segment, in milliseconds. Must be greater than 0. |
|
|
|
|
|
|
|
|
|
Maximum number of segments allowed per input. Must be greater than 0. An exception is thrown when the limit is exceeded. |
Return type: The UDTF outputs a single column named segment.
-
segment_type="audio":AUDIO_WAVEFORM_TYPE; -
segment_type="ref":AUDIO_CLIP_REF_TYPE.
Behavior:
-
segment_type="audio"requires the input to be a decoded audio waveform. -
segment_type="ref"requires the input to be a URI or an audio segment reference. The operator probes the actual audio duration and clips against the bounds of existing segment references. The output keeps segment references only, without copying audio content. -
The last segment may be shorter than
segment_duration_ms. -
When the required number of segments exceeds
max_segments, an exception is raised rather than silently truncating.
from pyflink.multimodal.operators import audio_split_by_duration
segments = df.join_lateral(
audio_split_by_duration(
col("audio_uri"),
segment_duration_ms=30_000,
segment_type="ref",
).alias("segment")
)
audio_split_by_timestamp
Splits audio using time ranges provided by the caller.
Input type:
-
First input column
-
segment_type="audio":AUDIO_WAVEFORM_TYPE; -
segment_type="ref":DataType.string()(audio URL) orAUDIO_CLIP_REF_TYPE.
-
-
Without a second input column: pass fixed time ranges as Python literals through
timestamps. -
With a second input column: do not set
timestamps. The second input column provides time ranges row by row. The DataFrame API type of this column is:
DataType.list(
DataType.struct({
"start_ms": DataType.int64(),
"end_ms": DataType.int64(),
})
)
Function signature:
audio_split_by_timestamp(
*columns,
timestamps=None,
segment_type="audio",
concurrency=None
)
|
Parameter |
Type |
Default |
Description |
|
|
|
|
Fixed time ranges. The outer container must be a Python list or tuple. Each item can be a |
|
|
|
|
|
timestamps parameter and the second time-range input column cannot be used at the same time.Ranges on URI inputs are not clipped against the actual audio duration. Bounded audio segment references are clipped only against their existing reference bounds.No rows are emitted when the audio or time-range input is null.
Return type: The UDTF outputs a single column named segment.
-
segment_type="audio":AUDIO_WAVEFORM_TYPE; -
segment_type="ref":AUDIO_CLIP_REF_TYPE.
from pyflink.dataframe import col
from pyflink.multimodal.operators import audio_split_by_timestamp
ranges = [
{"start_ms": 1_000, "end_ms": 3_500},
{"start_ms": 8_000, "end_ms": 12_000},
]
# Without a second input column: provide a fixed time range through Python parameters.
literal_segments = df.join_lateral(
audio_split_by_timestamp(
col("audio_uri"),
timestamps=ranges,
segment_type="ref"
).alias("segment")
)
# With a second input column: each row can provide a different time range.
column_segments = df.join_lateral(
audio_split_by_timestamp(
col("audio_uri"),
col("ranges"),
segment_type="ref"
).alias("segment")
)
audio_split_by_speech
Splits audio based on voice activity ranges already generated upstream. The operator itself does not perform speech detection. You can pass the output of audio_detect_speech as the voice activity ranges.
Input type:
-
segment_type="audio":AUDIO_WAVEFORM_TYPE+DataType.list(DataType.struct(...))(a voice activity interval column); -
segment_type="ref":DataType.string()orAUDIO_CLIP_REF_TYPE+DataType.list(DataType.struct(...))(a voice activity interval column).
The first input column is an audio column that matches segment_type. The second input column is an array column of voice activity ranges. The operator splits the audio column based on this voice-range array column.
Function signature:
audio_split_by_speech(
*columns,
segment_type="audio",
pre_padding_ms=0,
post_padding_ms=0,
merge_gap_ms=0,
min_segment_ms=0,
max_segment_ms=None,
max_segments=1024,
concurrency=None
)
|
Parameter |
Type |
Default |
Description |
|
|
|
|
|
|
|
|
|
Duration added before the start of each segment, in milliseconds. Must be greater than or equal to 0. |
|
|
|
|
Duration added after the end of each segment, in milliseconds. Must be greater than or equal to 0. |
|
|
|
|
Merges intervals when the gap between them is at most this value, in milliseconds. |
|
|
|
|
Discards intervals shorter than this value, in milliseconds. |
|
|
|
|
Further splits overly long intervals into segments no longer than this value. |
|
|
|
|
Maximum number of resulting segments. Must be greater than 0. |
Ranges that exceed the actual audio duration are clipped. Ranges that fall entirely out of bounds are discarded.No rows are emitted when the audio or voice activity range input is null.
Return type: The UDTF outputs a single column named segment.
-
segment_type="audio":AUDIO_WAVEFORM_TYPE; -
segment_type="ref":AUDIO_CLIP_REF_TYPE.
from pyflink.multimodal.operators import audio_split_by_speech
# Generate speech intervals.
detected_ranges = audio_detect_speech(
col("speech_ready"),
aggressiveness=2,
min_speech_ms=100,
merge_gap_ms=50
)
# Split the audio into intervals.
expression_segments = df.join_lateral(
audio_split_by_speech(
col("speech_ready"),
detected_ranges,
pre_padding_ms=200,
post_padding_ms=200,
merge_gap_ms=100,
).alias("segment")
)