Rate this Page
★ ★ ★ ★ ★

AudioConverter#

class torchcodec.decoders.AudioConverter(sample_rate: int | None = None, num_channels: int | None = None)[source]#

Turn RawAudioSamples into normalised float32 AudioSamples, optionally resampling and remixing channels.

This is a low-level API: for straightforward decoding, use AudioDecoder instead.

converter = AudioConverter(sample_rate=16_000)

for packet in demuxer:
    for raw_samples in packet_decoder.decode(packet):
        samples = converter.convert(raw_samples)
for raw_samples in packet_decoder.drain():
    samples = converter.convert(raw_samples)
samples = converter.drain()

Unlike a ColorConverter, this object is a stateful stream processor, and it is bound to an audio stream: feed it the stream’s samples in order. When resampling, it holds samples back between calls, so you won’t necessarily get the same number of samples out as you put in for a given call to convert().

Parameters:
  • sample_rate (int, optional) – The output sample rate. Defaults to the source’s own, i.e. no resampling.

  • num_channels (int, optional) – The output number of channels. Defaults to the source’s own.

Examples using AudioConverter:

Build your own decoding pipeline

Build your own decoding pipeline
convert(raw_samples: RawAudioSamples) → AudioSamples[source]#

Convert one RawAudioSamples into normalised float32 AudioSamples.

You may not get the same number of samples out as you put in, especially if you are resampling.

Returns:

The converted samples, normalised float32 in [-1, 1]. When resampling, fewer than were passed in, possibly none.

drain() → AudioSamples[source]#

Return the samples the resampler was still holding on to.

Empty unless you are resampling. This is not the codec’s own buffer, which AudioPacketDecoder.drain() takes care of.

reset() → None[source]#

Drop the resampler’s state and start over.

Needed after a Demuxer.seek(), after drain(), and before converting a different stream.