Rate this Page
★ ★ ★ ★ ★

RawAudioSamples#

class torchcodec.decoders.RawAudioSamples(data: Tensor, sample_rate: int, pts_seconds: float, duration_seconds: float, _generation: int = 0)[source]#

One decoded audio frame’s samples, exactly as the decoder produced them.

You cannot build one yourself: an AudioPacketDecoder creates them. Use an AudioConverter to turn them into normalised float32 AudioSamples, or read data directly:

for packet in demuxer:
    for raw_samples in audio_packet_decoder.decode(packet):
        print(raw_samples.data.shape)   # e.g. [2, 1024]
        print(raw_samples.data.dtype)   # e.g. float32, for an fltp source
        print(raw_samples.sample_rate)  # e.g. 16000

Examples using RawAudioSamples:

Raw frames and raw audio samples

Raw frames and raw audio samples
data: Tensor#

Always a contiguous [num_channels, num_samples] tensor, whatever the source’s sample format. Planar and packed sources alike come out with that same shape and layout (they are copied).

The dtype is whichever one holds the source’s samples exactly: uint8 for u8, int16 for s16, int32 for s32, int64 for s64, float32 for flt and float64 for dbl. The integer ones are not normalised to [-1, 1]; that is what an AudioConverter does.

duration_seconds: float#

How long these samples last, in seconds.

property num_channels: int#

The number of channels, i.e. data.shape[0].

property num_samples: int#

The number of samples per channel, i.e. data.shape[1].

pts_seconds: float#

The pts of the first sample, in seconds.

sample_rate: int#

The source’s sample rate, in Hz.