Rate this Page
★ ★ ★ ★ ★

RawFrame#

class torchcodec.decoders.RawFrame(handle: Tensor, pts_seconds: float, duration_seconds: float, storage: Tensor | None = None)[source]#

One decoded video frame, exactly as the decoder produced it.

You cannot build one yourself: a VideoPacketDecoder creates them. Use a ColorConverter to turn them into RGB Frames, or you can read and transform the raw samples directly from planes:

for packet in demuxer:
    for raw_frame in packet_decoder.decode(packet):
        y, u, v = raw_frame.planes
        # For a 480x270 yuv420p frame, y is [270, 480] uint8, and the
        # chroma is subsampled: u and v are [135, 240] each.
        print(y.shape, u.shape, v.shape)

Nothing here has been converted. The samples are in the codec’s own pixel format (typically YUV), on the device that was passed to VideoStream.make_decoder(), and width, height and planes are all pre-rotation: rotation is what a ColorConverter applies for you, and what you have to apply yourself if you convert planes on your own.

Important

On CUDA, the samples are produced on the CUDA stream that was current when you called VideoPacketDecoder.decode(), and all of that work is enqueued by the time the call returns. If you consume them on any other stream (via ColorConverter or by reading planes), you must handle synchronization yourself. See Low-level APIs and CUDA streams synchronization.

Examples using RawFrame:

Low-level APIs and CUDA streams synchronization

Low-level APIs and CUDA streams synchronization

Raw frames and raw audio samples

Raw frames and raw audio samples
property bit_depth: int#

How many bits of each planes sample are meaningful.

In almost every case this is just the bit depth of the source: 8 for an 8-bit video, 10 for a 10-bit one. It is worth having because a plane’s dtype only tells you its storage width, uint8 or uint16, while this tells you the range of the values held in it. So it is what you shift or scale by to reach a range of your own - y >> (frame.bit_depth - 8) for 8 bits, or y / (2 ** frame.bit_depth - 1) to normalise.

Two CUDA surface formats report more than their source: a 10-bit 4:4:4 source is uploaded as yuv444p16le, and a 12-bit source is tagged p016le on FFmpeg < 6, which has no p012le. Both report 16 where CPU decoding would report 10 and 12. Their samples are msb-aligned, so they genuinely are 16-bit values with zeroed low bits, and the arithmetic above still holds.

property color_primaries: str#

The FFmpeg color primaries name, e.g. "bt709", "bt2020", or "unspecified".

property color_range: str#

"tv" for limited range, "pc" for full range.

property color_space: str#

The FFmpeg color space name, e.g. "bt709", or "unspecified".

This describes planes, which is not always how the source is tagged: a CUDA decoder that falls back to the CPU converts an RGB frame into a YUV surface format, and this reports the color space of that conversion. Prefer it over VideoStreamHeaderMetadata.color_space when you convert the samples yourself.

property color_transfer_characteristic: str#

The FFmpeg transfer characteristic name, e.g. "bt709", "smpte2084" (PQ), "arib-std-b67" (HLG), or "unspecified".

duration_seconds: float#

How long this frame is displayed for, in seconds.

property height: int#

The height of the decoded samples, before rotation.

property pixel_format: str#

The FFmpeg pixel-format name, e.g. "yuv420p".

On CPU this is the source’s own format. On CUDA it is always one of the NVDEC surface formats: "nv12", "p010le", "p012le", "p016le", "yuv444p" or "yuv444p16le".

property planes: tuple[Tensor, ...]#

The decoder’s own samples, as 2D tensor views.

There is exactly one tensor per component of pixel_format, in the order that format describes, of dtype uint8 or uint16 depending on bit_depth. So yuv420p and nv12 both give three (y, u, v = planes), yuva420p four (y, u, v, a = planes) and gray one ((y,) = planes).

They are always on the device that was passed to VideoStream.make_decoder(), including when a CUDA decoder has to fall back to decoding on the CPU: it uploads those frames before handing them out.

Only the luma and alpha components are height by width. The chroma ones are subsampled by whatever pixel_format says: half in both directions for a 4:2:0 format, half the width for 4:2:2, full size for 4:4:4 and for the RGB formats. Odd sizes round up, so the chroma of a 4:2:0 frame 481 samples wide is 241 wide.

Note

None of these views is contiguous in general. FFmpeg pads each row out to a line size of its own choosing, so even a luma plane is usually strided, and the semi-planar formats (nv12, p010le, and the other NVDEC surface formats) store U and V interleaved in a single allocation, which forces those two to be strided views into it.

Raises:

RuntimeError – For the pixel formats that can’t be viewed without a copy - sub-byte-packed, palettised and float ones - and for frames stored bottom-up. Check pixel_format first if you are decoding something exotic.

pts_seconds: float#

The pts of this frame, in seconds.

property rotation: float#

How many degrees counter-clockwise the frame has to be rotated to be upright, or 0 if the container asks for no rotation.

This is not applied to planes. A ColorConverter applies it, rounded to the nearest multiple of 90, to its output.

storage_cuda: Tensor | None#

The CUDA allocation backing planes, or None on CPU.

This tensor is exposed for one purpose, which is to let you call torch.Tensor.record_stream() if you want to. See Low-level APIs and CUDA streams synchronization for more details. You shouldn’t read or write this tensor directly.

property width: int#

The width of the decoded samples, before rotation.