Documentation
¶
Overview ¶
Package decode turns a video file into a stream of frames.
Index ¶
- Variables
- func VideoToolboxExact(ctx context.Context, bin string) (bool, error)
- type AudioRequest
- type FFmpeg
- func (d *FFmpeg) Decode(ctx context.Context, req Request, fn func(*frame.Frame) error) error
- func (d *FFmpeg) DecodeAudio(ctx context.Context, req AudioRequest, fn func(samples [][]float32) error) error
- func (d *FFmpeg) HWAccel() HWAccel
- func (d *FFmpeg) SegmentDecoders(req Request) int
- func (d *FFmpeg) SegmentHWAccel(req Request) HWAccel
- type HWAccel
- type HWAccelReporter
- type Option
- type Request
- type SegmentDecoders
- type SegmentHWAccelReporter
- type Source
- type ToneMap
Constants ¶
This section is empty.
Variables ¶
var ErrHWAccel = errors.New("unknown hwaccel")
ErrHWAccel is returned for an unknown hardware decoding mode.
Functions ¶
func VideoToolboxExact ¶
VideoToolboxExact reports whether ffmpeg (bin) decodes with VideoToolbox exactly as on the CPU: every plane of every frame of reference H.264 8-bit and HEVC 10-bit clips. It holds on Apple silicon, where H.264 and HEVC decoding is bit-exact by specification; a virtualised GPU (a macOS virtual machine, a CI runner) was seen to return other chroma. Where VideoToolbox is unavailable, ffmpeg decodes on the CPU and the frames match.
Types ¶
type AudioRequest ¶
type AudioRequest struct {
Path string
// Stream is the index of the stream in the file (ffprobe's index).
Stream int
// SampleRate and Channels are the stream's, as probed: the samples
// come at that rate (ffmpeg resamples only a decoder that disagrees
// with the probe) with that many channels, in stream order.
SampleRate int
Channels int
// ChunkFrames is the number of frames (samples per channel) handed
// over per call; 0 means 100 ms.
ChunkFrames int
}
AudioRequest describes an audio stream to decode.
type FFmpeg ¶
type FFmpeg struct {
// contains filtered or unexported fields
}
FFmpeg decodes with an ffmpeg subprocess writing raw planes to a pipe, as 8-bit samples or 16-bit little-endian ones for high bit depth pools. Luma-only pools only receive the Y plane: a third of the bytes of a 4:2:0 frame, enough for every pixel analyzer. WithHWAccel decodes on an NVIDIA GPU or with VideoToolbox.
func NewFFmpeg ¶
NewFFmpeg returns a Source backed by the ffmpeg binary. threads=0 lets ffmpeg decide.
func (*FFmpeg) Decode ¶
Decode implements Source. A hardware decode failing before its first frame is retried with the next mode down (see WithHWAccel).
func (*FFmpeg) DecodeAudio ¶
func (d *FFmpeg) DecodeAudio( ctx context.Context, req AudioRequest, fn func(samples [][]float32) error, ) error
DecodeAudio decodes an audio stream to 32-bit float samples (full scale ±1) and calls fn with chunks of planar samples, one slice per channel, ChunkFrames long (the last one shorter). The slices are reused: they are valid only during the call.
ffmpeg writes interleaved samples (its raw muxers take no planar format) through the same socket as frames; they are de-interleaved here, which costs less than a nanosecond per sample.
func (*FFmpeg) HWAccel ¶
HWAccel returns the hardware decoding mode of decodes that are not segments (Request.Segment): the configured one, none for HWAccelAuto.
func (*FFmpeg) SegmentDecoders ¶
SegmentDecoders implements SegmentDecoders. When segments decode on VideoToolbox, it is its sessions (see WithVideoToolboxSessions). Otherwise it is 1: ffmpeg's multithreaded CPU decoder already keeps every core busy, and segments would only add the GOP each one decodes before its first frame.
func (*FFmpeg) SegmentHWAccel ¶
SegmentHWAccel returns the mode the segments of the video of req start decoding in, before any fallback: with HWAccelAuto, VideoToolbox on macOS for the codecs and formats it decodes exactly.
type HWAccel ¶
type HWAccel string
HWAccel selects hardware decoding: NVIDIA NVDEC or Apple VideoToolbox.
const ( // HWAccelNone decodes on the CPU (the zero value). HWAccelNone HWAccel = "" // HWAccelCUDA decodes with NVDEC (ffmpeg -hwaccel cuda) and lets ffmpeg // download each frame to system memory: selection, scaling and format // conversion stay on the CPU, so frames are identical to a CPU decode // (decoders are bit-exact by specification) and VMAF is unchanged. HWAccelCUDA HWAccel = "cuda" // HWAccelCUDAScale also scales and converts on the GPU (scale_cuda, // bicubic) before downloading the frames: the least CPU, but the GPU // scaler does not round like swscale, so scores differ slightly from a // CPU decode (see docs/gpu.md). HWAccelCUDAScale HWAccel = "cuda-scale" // HWAccelVideoToolbox decodes with Apple VideoToolbox (ffmpeg -hwaccel // videotoolbox) and lets ffmpeg download each frame: like HWAccelCUDA, // frames are identical to a CPU decode. Only the codecs whose decoding // is bit-exact by specification, in the 4:2:0 formats VideoToolbox // outputs as they are, are decoded by it (see vtCodecs). One session // decodes slower than ffmpeg's multithreaded CPU decoder but uses // almost no CPU, and concurrent sessions add up: it pays when a video // is decoded in concurrent segments (Request.Segment). HWAccelVideoToolbox HWAccel = "videotoolbox" // HWAccelAuto uses the platform's exact hardware decoder where it pays: // on macOS, VideoToolbox for concurrent decodes (Request.Segment: the // segments of a frame analysis, the runs and segments of a VMAF // measurement), the CPU for every other decode. Elsewhere it is the // CPU (NVDEC needs an explicit cuda mode, which the GPU preflight // checks). HWAccelAuto HWAccel = "auto" )
Hardware decoding modes, from the most to the least offloaded.
func ParseHWAccel ¶
ParseHWAccel reads a hardware decoding mode: none (or empty), auto, videotoolbox, cuda or cuda-scale.
func (HWAccel) CUDA ¶
CUDA reports whether h decodes on an NVIDIA GPU, which the GPU preflight must check.
type HWAccelReporter ¶
type HWAccelReporter interface {
HWAccel() HWAccel
}
HWAccelReporter is implemented by sources that decode with a hardware decoding mode (*FFmpeg): measurements report it next to their results.
type Option ¶
type Option func(*FFmpeg)
Option configures an FFmpeg source.
func WithHWAccel ¶
WithHWAccel decodes with hardware in the given mode (HWAccelAuto is resolved for the running system). When hardware decoding of a file fails before its first frame (no device, an unsupported profile), the file is decoded again with the next mode down and later decodes of that file start there; a warning is logged.
func WithLogger ¶
WithLogger logs hardware decoding fallbacks to logger.
func WithVideoToolboxSessions ¶
WithVideoToolboxSessions lets at most n decodes use VideoToolbox at once (NumCPU by default): the others decode on the CPU, with identical frames. On an M2 Max, a dozen sessions reach the throughput of the decoding engine; a CPU decode next to them only pays on an idle machine (docs/analysis.md).
type Request ¶
type Request struct {
Path string
Pool *frame.Pool
// SourceWidth and SourceHeight are the coded dimensions of the video.
SourceWidth int
SourceHeight int
// Start seeks to the frame presented at Start (0 = beginning), relative
// to the first frame of the video.
Start media.Duration
// Origin is the presentation time of the first frame of the video on
// the container's timeline (bitstream.Report.Start), which ffmpeg seeks
// on: a seek goes to Origin+Start. Without it, ffmpeg would count Start
// from the container's start, that of its earliest stream, and land
// one frame or more early in a video starting after its audio.
Origin media.Duration
// FirstIndex is the index in the video of the first decoded frame (the
// frame presented at Start). Returned frames are numbered from it.
FirstIndex int
// MaxFrames stops after that many frames (0 = until the end).
MaxFrames int
// Select keeps only the frames in these sorted, disjoint [from, to) index
// ranges, counted from the first decoded frame (nil = every frame).
// Frames are still decoded but only selected ones are scaled and piped.
Select [][2]int
// Threads overrides the decoder thread count (0 = the Source default).
Threads int
// PTS gives the presentation time of every frame of the video, indexed
// by frame index. Frames beyond it are timestamped from Start and
// FrameRate.
PTS []media.Duration
// FrameRate is the nominal frame rate, used to seek and to timestamp
// frames missing from PTS.
FrameRate media.Rational
// Codec is the ffprobe codec name of the video, when known: hardware
// decoding is only tried for codecs the hardware decodes exactly.
Codec string
// PixelFormat is the ffprobe pixel format of the video, when known:
// VideoToolbox only decodes the formats it outputs as they are.
PixelFormat string
// Segment marks the decode as one of several concurrent decodes of
// parts of a video: HWAccelAuto then decodes with VideoToolbox, whose
// concurrent sessions add up while a single one is slower than the
// CPU.
Segment bool
// ToneMap, when set, converts HDR frames to SDR while scaling (pools
// with chroma only, see ToneMap).
ToneMap *ToneMap
}
Request describes what to decode. The pool sets the output geometry and planes: frames are scaled (bicubic) to it, those already at its size pass through.
type SegmentDecoders ¶
type SegmentDecoders interface {
// SegmentDecoders returns how many segments of the video of req to
// decode concurrently: 1 when a single decode is the fastest.
SegmentDecoders(req Request) int
}
SegmentDecoders is implemented by sources that know how many segments of a video are worth decoding at once (*FFmpeg).
type SegmentHWAccelReporter ¶
SegmentHWAccelReporter is implemented by sources whose segment decodes (Request.Segment) may use another hardware decoding mode than their other decodes (*FFmpeg with HWAccelAuto): measurements decoding segments report it.
type Source ¶
type Source interface {
// Decode calls fn for every decoded frame, in presentation order. fn owns
// the reference it receives and must Release it.
Decode(
ctx context.Context,
req Request,
fn func(*frame.Frame) error,
) error
}
Source decodes the frames of a video.
type ToneMap ¶
type ToneMap struct {
// Input is the colour description of the frames. It is given
// explicitly: raw digests carry no colour tags, and a wrong guess
// would silently skip the mapping.
Input media.Color
}
ToneMap asks the decoder for SDR frames tone mapped from an HDR (PQ or HLG) signal: ffmpeg's scaler converts them to BT.709 primaries, transfer and matrix with its perceptual intent, a tone and gamut mapping ported from libplacebo (FFmpeg ≥ 8). Both sides of a comparison go through the same mapping.