Audio detection and splitting

Updated at:

This topic describes the silence detection, voice activity detection, and audio splitting operators in the Python DataFrame API.

Limits

Operator summary

Category

Operator

Description

Audio detection

audio_silence_detection

Detects low-amplitude silence intervals in an audio waveform.

audio_detect_speech

Uses WebRTC VAD to detect voice activity intervals in an audio waveform.

Audio splitting

audio_split_by_duration

Splits audio sequentially into segments of a fixed duration and outputs actual audio waveforms or segment references based on segment_type. Suitable for long audio that exceeds the per-request input limit of downstream models, for example splitting a one-hour recording into 30-second segments for parallel ASR processing. The last segment may be shorter.

audio_split_by_timestamp

Splits audio by start and end time ranges provided by the caller and outputs audio waveforms or segment references in the original order. Suitable when subtitles, logs, meeting minutes, or manually annotated timelines are already available, for example extracting audio by subtitle timestamps or locating utterance segments by meeting minutes.

audio_split_by_speech

Based on existing voice activity ranges, pads, merges, filters out short segments, and splits long ones, then outputs audio waveforms or segment references. Suitable for calls, interviews, or podcasts that contain a lot of silence, for example combined with audio_detect_speech to skip gaps and produce only speech segments suitable for ASR.

Time range type

Silence detection and voice activity detection return the following DataFrame API type:

DataType.list(
    DataType.struct({
        "start_ms": DataType.int64(),
        "end_ms": DataType.int64(),
        "duration_ms": DataType.int64()
    })
)

Time ranges use the half-open interval [start_ms, end_ms), in milliseconds.

Common runtime parameters

Function signatures retain the runtime parameters that each operator actually supports. To avoid repetition, the operator parameter tables describe only business parameters. Runtime parameters are described uniformly as follows.

Parameter

Type

Default

Scope

Description

concurrency

Optional[int]

None

All operators on this page

The parallelism of the UDF or UDTF. None uses the framework default.

Audio detection

audio_silence_detection

Detects low-amplitude silence intervals in an audio waveform.

Input type: AUDIO_WAVEFORM_TYPE, a decoded audio waveform column.

Function signature:

audio_silence_detection(
    *columns,
    threshold_db,
    min_silence_ms=0,
    concurrency=None
)

Parameter

Type

Default

Description

threshold_db

float

Required

Silence threshold relative to full scale. Must be less than or equal to 0. Frames with energy at or below this threshold are considered silent.

min_silence_ms

int

0

Minimum silence duration, in milliseconds. Shorter intervals are not returned.

The operator determines silence as follows:

  1. Multi-channel audio is first averaged across channels at each point in time to produce mono audio.

  2. The audio is divided into short overlapping frames, and RMS (root mean square) is used to estimate the average volume of each frame.

  3. The average volume is expressed in dBFS. Values closer to 0 indicate louder audio; smaller values indicate quieter audio.

  4. Stretches that stay at or below threshold_db for at least min_silence_ms are returned as silence intervals.

For example, with threshold_db=-45.0, a frame with an average volume of -30 dBFS is not considered silent, while -50 dBFS is.

The operator detects silence based on volume alone. It does not use VAD or ASR models, so a "silence interval" is not equivalent to a "non-speech interval". A null input returns None.

Return type:

DataType.list(
    DataType.struct({
        "start_ms": DataType.int64(),    # Start time of the interval.
        "end_ms": DataType.int64(),      # End time of the interval.
        "duration_ms": DataType.int64()  # Duration of the interval.
    })
)
from pyflink.dataframe import col
from pyflink.multimodal.operators import audio_silence_detection

result = df.with_column(
    "silence_ranges",
    audio_silence_detection(
        col("waveform"),
        threshold_db=-45.0,
        min_silence_ms=300
    )
)

audio_detect_speech

Uses WebRTC VAD to detect voice activity intervals in an audio waveform.

Input type: AUDIO_WAVEFORM_TYPE, a mono audio waveform column. The sample rate must be 8000, 16000, 32000, or 48000 Hz.

Function signature:

audio_detect_speech(
    *columns,
    aggressiveness=0,
    frame_ms=30,
    min_speech_ms=0,
    merge_gap_ms=0,
    max_segments=1024,
    concurrency=None
)

Parameter

Type

Default

Description

aggressiveness

int

0

WebRTC VAD aggressiveness mode. Valid values: [0, 3].

frame_ms

int

30

VAD frame length. Must be 10, 20, or 30 milliseconds.

min_speech_ms

int

0

Discards speech intervals shorter than this duration, in milliseconds.

merge_gap_ms

int

0

Merges adjacent voice activity intervals when the gap between them is at most this value, in milliseconds.

max_segments

int

1024

Maximum number of voice activity intervals to return. Must be a positive integer.

Results are sorted by time. A null input returns None.

Return type:

DataType.list(
    DataType.struct({
        "start_ms": DataType.int64(),    # Start timestamp of the interval.
        "end_ms": DataType.int64(),      # End timestamp of the interval.
        "duration_ms": DataType.int64()  # Duration of the interval.
    })
)
from pyflink.multimodal.operators import (
    audio_detect_speech,
    audio_standardize,
)

# Standardize the audio first.
speech_ready = audio_standardize(
    col("waveform"),
    sample_rate=16000,
    channels=1
)

# Detect speech intervals.
result = df.with_column(
    "speech_ranges",
    audio_detect_speech(
        speech_ready,
        aggressiveness=2,
        min_speech_ms=200,
        merge_gap_ms=100
    )
)

Audio splitting

audio_split_by_duration

Splits audio into segments of a fixed duration.

Input type:

The input type must match segment_type. No rows are emitted when the input is null.

Function signature:

audio_split_by_duration(
    *columns,
    segment_duration_ms,
    segment_type="audio",
    max_segments=1024,
    concurrency=None
)

Parameter

Type

Default

Description

segment_duration_ms

int

Required

Target duration per segment, in milliseconds. Must be greater than 0.

segment_type

str

"audio"

"audio" outputs materialized waveforms; "ref" outputs audio segment references.

max_segments

int

1024

Maximum number of segments allowed per input. Must be greater than 0. An exception is thrown when the limit is exceeded.

Return type: The UDTF outputs a single column named segment.

Behavior:

  • segment_type="audio" requires the input to be a decoded audio waveform.

  • segment_type="ref" requires the input to be a URI or an audio segment reference. The operator probes the actual audio duration and clips against the bounds of existing segment references. The output keeps segment references only, without copying audio content.

  • The last segment may be shorter than segment_duration_ms.

  • When the required number of segments exceeds max_segments, an exception is raised rather than silently truncating.

from pyflink.multimodal.operators import audio_split_by_duration

segments = df.join_lateral(
    audio_split_by_duration(
        col("audio_uri"),
        segment_duration_ms=30_000,
        segment_type="ref",
    ).alias("segment")
)

audio_split_by_timestamp

Splits audio using time ranges provided by the caller.

Input type:

  • First input column

  • Without a second input column: pass fixed time ranges as Python literals through timestamps.

  • With a second input column: do not set timestamps. The second input column provides time ranges row by row. The DataFrame API type of this column is:

DataType.list(
    DataType.struct({
        "start_ms": DataType.int64(),
        "end_ms": DataType.int64(),
    })
)

Function signature:

audio_split_by_timestamp(
    *columns,
    timestamps=None,
    segment_type="audio",
    concurrency=None
)

Parameter

Type

Default

Description

timestamps

Optional[Union[List[Any], Tuple[Any, ...]]]

None

Fixed time ranges. The outer container must be a Python list or tuple. Each item can be a Row, an object with start_ms and end_ms attributes, a dict with the same keys, or a two-element list or tuple. Start and end times must be non-negative int values, and end_ms must not be less than start_ms. None means time ranges are provided row by row through the second input column. It does not mean splitting is disabled.

segment_type

str

"audio"

"audio" returns materialized waveforms; "ref" returns audio segment references.

timestamps parameter and the second time-range input column cannot be used at the same time.Ranges on URI inputs are not clipped against the actual audio duration. Bounded audio segment references are clipped only against their existing reference bounds.No rows are emitted when the audio or time-range input is null.

Return type: The UDTF outputs a single column named segment.

from pyflink.dataframe import col
from pyflink.multimodal.operators import audio_split_by_timestamp

ranges = [
    {"start_ms": 1_000, "end_ms": 3_500},
    {"start_ms": 8_000, "end_ms": 12_000},
]

# Without a second input column: provide a fixed time range through Python parameters.
literal_segments = df.join_lateral(
    audio_split_by_timestamp(
        col("audio_uri"),
        timestamps=ranges,
        segment_type="ref"
    ).alias("segment")
)

# With a second input column: each row can provide a different time range.
column_segments = df.join_lateral(
    audio_split_by_timestamp(
        col("audio_uri"),
        col("ranges"),
        segment_type="ref"
    ).alias("segment")
)

audio_split_by_speech

Splits audio based on voice activity ranges already generated upstream. The operator itself does not perform speech detection. You can pass the output of audio_detect_speech as the voice activity ranges.

Input type:

  • segment_type="audio": AUDIO_WAVEFORM_TYPE + DataType.list(DataType.struct(...)) (a voice activity interval column);

  • segment_type="ref": DataType.string() or AUDIO_CLIP_REF_TYPE + DataType.list(DataType.struct(...)) (a voice activity interval column).

The first input column is an audio column that matches segment_type. The second input column is an array column of voice activity ranges. The operator splits the audio column based on this voice-range array column.

Function signature:

audio_split_by_speech(
    *columns,
    segment_type="audio",
    pre_padding_ms=0,
    post_padding_ms=0,
    merge_gap_ms=0,
    min_segment_ms=0,
    max_segment_ms=None,
    max_segments=1024,
    concurrency=None
)

Parameter

Type

Default

Description

segment_type

str

"audio"

"audio" returns materialized waveforms; "ref" returns audio segment references.

pre_padding_ms

int

0

Duration added before the start of each segment, in milliseconds. Must be greater than or equal to 0.

post_padding_ms

int

0

Duration added after the end of each segment, in milliseconds. Must be greater than or equal to 0.

merge_gap_ms

int

0

Merges intervals when the gap between them is at most this value, in milliseconds.

min_segment_ms

int

0

Discards intervals shorter than this value, in milliseconds.

max_segment_ms

Optional[int]

None

Further splits overly long intervals into segments no longer than this value. None means no limit.

max_segments

int

1024

Maximum number of resulting segments. Must be greater than 0.

Ranges that exceed the actual audio duration are clipped. Ranges that fall entirely out of bounds are discarded.No rows are emitted when the audio or voice activity range input is null.

Return type: The UDTF outputs a single column named segment.

from pyflink.multimodal.operators import audio_split_by_speech

# Generate speech intervals.
detected_ranges = audio_detect_speech(
    col("speech_ready"),
    aggressiveness=2,
    min_speech_ms=100,
    merge_gap_ms=50
)

# Split the audio into intervals.
expression_segments = df.join_lateral(
    audio_split_by_speech(
        col("speech_ready"),
        detected_ranges,
        pre_padding_ms=200,
        post_padding_ms=200,
        merge_gap_ms=100,
    ).alias("segment")
)