RawFrame#
- class torchcodec.decoders.RawFrame(handle: Tensor, pts_seconds: float, duration_seconds: float, storage: Tensor | None = None)[source]#
One decoded video frame, exactly as the decoder produced it.
You cannot build one yourself: a
VideoPacketDecodercreates them. Use aColorConverterto turn them into RGBFrames, or you can read and transform the raw samples directly fromplanes:for packet in demuxer: for raw_frame in packet_decoder.decode(packet): y, u, v = raw_frame.planes # For a 480x270 yuv420p frame, y is [270, 480] uint8, and the # chroma is subsampled: u and v are [135, 240] each. print(y.shape, u.shape, v.shape)
Nothing here has been converted. The samples are in the codec’s own pixel format (typically YUV), on the device that was passed to
VideoStream.make_decoder(), andwidth,heightandplanesare all pre-rotation:rotationis what aColorConverterapplies for you, and what you have to apply yourself if you convertplaneson your own.Important
On CUDA, the samples are produced on the CUDA stream that was current when you called
VideoPacketDecoder.decode(), and all of that work is enqueued by the time the call returns. If you consume them on any other stream (viaColorConverteror by readingplanes), you must handle synchronization yourself. See Low-level APIs and CUDA streams synchronization.Examples using
RawFrame:- property bit_depth: int#
How many bits of each
planessample are meaningful.In almost every case this is just the bit depth of the source: 8 for an 8-bit video, 10 for a 10-bit one. It is worth having because a plane’s dtype only tells you its storage width,
uint8oruint16, while this tells you the range of the values held in it. So it is what you shift or scale by to reach a range of your own -y >> (frame.bit_depth - 8)for 8 bits, ory / (2 ** frame.bit_depth - 1)to normalise.Two CUDA surface formats report more than their source: a 10-bit 4:4:4 source is uploaded as
yuv444p16le, and a 12-bit source is taggedp016leon FFmpeg < 6, which has nop012le. Both report 16 where CPU decoding would report 10 and 12. Their samples are msb-aligned, so they genuinely are 16-bit values with zeroed low bits, and the arithmetic above still holds.
- property color_primaries: str#
The FFmpeg color primaries name, e.g.
"bt709","bt2020", or"unspecified".
- property color_space: str#
The FFmpeg color space name, e.g.
"bt709", or"unspecified".This describes
planes, which is not always how the source is tagged: a CUDA decoder that falls back to the CPU converts an RGB frame into a YUV surface format, and this reports the color space of that conversion. Prefer it overVideoStreamHeaderMetadata.color_spacewhen you convert the samples yourself.
- property color_transfer_characteristic: str#
The FFmpeg transfer characteristic name, e.g.
"bt709","smpte2084"(PQ),"arib-std-b67"(HLG), or"unspecified".
- property pixel_format: str#
The FFmpeg pixel-format name, e.g.
"yuv420p".On CPU this is the source’s own format. On CUDA it is always one of the NVDEC surface formats:
"nv12","p010le","p012le","p016le","yuv444p"or"yuv444p16le".
- property planes: tuple[Tensor, ...]#
The decoder’s own samples, as 2D tensor views.
There is exactly one tensor per component of
pixel_format, in the order that format describes, of dtypeuint8oruint16depending onbit_depth. Soyuv420pandnv12both give three (y, u, v = planes),yuva420pfour (y, u, v, a = planes) andgrayone ((y,) = planes).They are always on the device that was passed to
VideoStream.make_decoder(), including when a CUDA decoder has to fall back to decoding on the CPU: it uploads those frames before handing them out.Only the luma and alpha components are
heightbywidth. The chroma ones are subsampled by whateverpixel_formatsays: half in both directions for a 4:2:0 format, half the width for 4:2:2, full size for 4:4:4 and for the RGB formats. Odd sizes round up, so the chroma of a 4:2:0 frame 481 samples wide is 241 wide.Note
None of these views is contiguous in general. FFmpeg pads each row out to a line size of its own choosing, so even a luma plane is usually strided, and the semi-planar formats (
nv12,p010le, and the other NVDEC surface formats) store U and V interleaved in a single allocation, which forces those two to be strided views into it.- Raises:
RuntimeError – For the pixel formats that can’t be viewed without a copy - sub-byte-packed, palettised and float ones - and for frames stored bottom-up. Check
pixel_formatfirst if you are decoding something exotic.
- property rotation: float#
How many degrees counter-clockwise the frame has to be rotated to be upright, or 0 if the container asks for no rotation.
This is not applied to
planes. AColorConverterapplies it, rounded to the nearest multiple of 90, to its output.
- storage_cuda: Tensor | None#
The CUDA allocation backing
planes, orNoneon CPU.This tensor is exposed for one purpose, which is to let you call
torch.Tensor.record_stream()if you want to. See Low-level APIs and CUDA streams synchronization for more details. You shouldn’t read or write this tensor directly.