Demuxer#
- class torchcodec.decoders.Demuxer(source: str | Path | bytes | Tensor | RawIOBase | BufferedReader, *, streams: str | int | tuple[str | int, ...] = 'video')[source]#
Reads one or more video and audio streams from a container, and produces their compressed
Packets.This is a low-level API: for straightforward decoding, use
VideoDecoderorAudioDecoderinstead.Packets come out interleaved, and
Packet.stream_indexsays which stream each one belongs to:demuxer = Demuxer("video.mp4", streams=("video", "audio")) decoders = {s.index: s.make_decoder(device="cuda") for s in demuxer.streams} for packet in demuxer: for output in decoders[packet.stream_index].decode(packet): ...
- Parameters:
source (str,
Pathlib.path, bytes,torch.Tensoror file-like object) –The source of the media:
If
str: a local path or a URL to a media file.If
Pathlib.path: a path to a local media file.If
bytesobject ortorch.Tensor: the raw encoded data.If file-like object: we read data from the object on demand. The object must expose the methods read(self, size: int) -> bytes and seek(self, offset: int, whence: int) -> int.
streams (str, int or tuple, optional) – Which streams to follow, as a single selector or a tuple of them. A selector is either
"video"or"audio"for the best stream of that type, or anintfor a stream index, absolute across all media types."all"follows every audio and video stream in container order, skipping the rest, and can only be used on its own. Default:"video".
- Variables:
streams (tuple) – The
VideoStreamandAudioStreamobjects being followed, in the order thestreamsparameter named them. Packet decoders are built from these.metadata (DemuxerMetadata) – What the container header says about the container itself. What it says about a given stream is on
demuxer.streams[i].metadata.
Examples using
Demuxer:- __next__() Packet[source]#
Read and return the next
Packet.Packets come out interleaved across the streams being followed, in the order the container stores them, so this is where
Packet.stream_indexmatters: it is what routes each packet to the decoder of its own stream.- Returns:
The next packet.
- Return type:
- seek(seconds: float, *, stream: VideoStream | AudioStream | None = None) None[source]#
Move the demuxer to
seconds.This moves every stream being followed. For videos, this lands on the keyframe at or before
seconds. For audio, a lossy codec’s first frames after a seek are typically slightly wrong until the codec re-primes. This is especially true when resampling is involved (via anAudioConverter). Pre-rolling a margin of audio before the target is up to you.There is no
seek_modeto choose from: seeking straight tosecondsis whatVideoDecodercallsseek_mode="approximate". To get theseek_mode="exact"behavior, scan the stream and seek toFrameIndex.key_frame_seconds_for()of your target instead.Important
You must call
VideoPacketDecoder.reset()orAudioPacketDecoder.reset()on every decoder fed by this demuxer afterwards, andAudioConverter.reset()on every converter too: a seek invalidates a codec and resampler states.- Parameters:
seconds (float) – The position to seek to.
stream (VideoStream or AudioStream, optional) – The stream the target
secondsis resolved against. FFmpeg resolves a seek in a single stream’s time base and lands on that stream’s keyframes, the other streams merely resuming from wherever the container ends up - so a second video stream may land mid-GOP and decode garbage until its next keyframe. Defaults to the first ofstreams, as passed to the constructor.