Rate this Page
★ ★ ★ ★ ★

Demuxer#

class torchcodec.decoders.Demuxer(source: str | Path | bytes | Tensor | RawIOBase | BufferedReader, *, streams: str | int | tuple[str | int, ...] = 'video')[source]#

Reads one or more video and audio streams from a container, and produces their compressed Packets.

This is a low-level API: for straightforward decoding, use VideoDecoder or AudioDecoder instead.

Packets come out interleaved, and Packet.stream_index says which stream each one belongs to:

demuxer = Demuxer("video.mp4", streams=("video", "audio"))
decoders = {s.index: s.make_decoder(device="cuda") for s in demuxer.streams}

for packet in demuxer:
    for output in decoders[packet.stream_index].decode(packet):
        ...
Parameters:
  • source (str, Pathlib.path, bytes, torch.Tensor or file-like object) –

    The source of the media:

    • If str: a local path or a URL to a media file.

    • If Pathlib.path: a path to a local media file.

    • If bytes object or torch.Tensor: the raw encoded data.

    • If file-like object: we read data from the object on demand. The object must expose the methods read(self, size: int) -> bytes and seek(self, offset: int, whence: int) -> int.

  • streams (str, int or tuple, optional) – Which streams to follow, as a single selector or a tuple of them. A selector is either "video" or "audio" for the best stream of that type, or an int for a stream index, absolute across all media types. "all" follows every audio and video stream in container order, skipping the rest, and can only be used on its own. Default: "video".

Variables:
  • streams (tuple) – The VideoStream and AudioStream objects being followed, in the order the streams parameter named them. Packet decoders are built from these.

  • metadata (DemuxerMetadata) – What the container header says about the container itself. What it says about a given stream is on demuxer.streams[i].metadata.

Examples using Demuxer:

Build your own decoding pipeline

Build your own decoding pipeline

Low-level APIs and CUDA streams synchronization

Low-level APIs and CUDA streams synchronization

Multi-threaded decoding pipelines

Multi-threaded decoding pipelines

Raw frames and raw audio samples

Raw frames and raw audio samples
__next__() → Packet[source]#

Read and return the next Packet.

Packets come out interleaved across the streams being followed, in the order the container stores them, so this is where Packet.stream_index matters: it is what routes each packet to the decoder of its own stream.

Returns:

The next packet.

Return type:

Packet

seek(seconds: float, *, stream: VideoStream | AudioStream | None = None) → None[source]#

Move the demuxer to seconds.

This moves every stream being followed. For videos, this lands on the keyframe at or before seconds. For audio, a lossy codec’s first frames after a seek are typically slightly wrong until the codec re-primes. This is especially true when resampling is involved (via an AudioConverter). Pre-rolling a margin of audio before the target is up to you.

There is no seek_mode to choose from: seeking straight to seconds is what VideoDecoder calls seek_mode="approximate". To get the seek_mode="exact" behavior, scan the stream and seek to FrameIndex.key_frame_seconds_for() of your target instead.

Important

You must call VideoPacketDecoder.reset() or AudioPacketDecoder.reset() on every decoder fed by this demuxer afterwards, and AudioConverter.reset() on every converter too: a seek invalidates a codec and resampler states.

Parameters:
  • seconds (float) – The position to seek to.

  • stream (VideoStream or AudioStream, optional) – The stream the target seconds is resolved against. FFmpeg resolves a seek in a single stream’s time base and lands on that stream’s keyframes, the other streams merely resuming from wherever the container ends up - so a second video stream may land mid-GOP and decode garbage until its next keyframe. Defaults to the first of streams, as passed to the constructor.