如何通过HEVC方法获取运动向量?及x265运动估计代码提取求助
Hey there! I’ve worked with x265 and HEVC motion estimation (ME) quite a bit, so let’s break down how to tackle your problem step by step—from understanding HEVC’s MV logic to pulling usable ME code from x265’s dense source tree.
1. HEVC Motion Estimation: Core Basics
First, it helps to grasp how HEVC generates motion vectors (MVs) since that’s the foundation of x265’s implementation:
- HEVC uses hierarchical motion estimation: starts with large search ranges (e.g., ±64 pixels) on downsampled frames, then refines to smaller ranges on full-resolution frames for precision.
- It supports multiple block sizes (4x4 up to 64x64) to capture both large, sweeping motion and small, detailed movement.
- Two key MV derivation modes:
- Merge Mode: Reuses MVs from neighboring blocks or reference frames to cut down on computation.
- AMVP (Advanced Motion Vector Prediction): Predicts MVs using a list of candidate vectors, then searches around those candidates for the best match.
- HEVC also does fractional-pixel refinement (down to 1/4-pixel accuracy) using interpolation filters to get sub-pixel precision.
2. Extracting Motion Estimation Code from x265
x265’s codebase is big, but the ME logic is concentrated in a few key spots. Here’s how to zero in on it:
Step 1: Locate Core ME Files
The heart of x265’s ME lives in these files:
source/common/motion_estimation.cpp/motion_estimation.h: Implements low-level integer and fractional pixel search routines.source/encoder/me.cpp: Handles encoder-level ME scheduling (e.g., selecting block sizes, managing reference frames).- Key data structures to familiarize yourself with:
MotionVector(insource/common/common.h): Stores the x/y components of the MV and its precision.PicYuv: Represents a YUV frame—the primary input to ME functions.MEParam: Holds ME configuration (search range, block size, reference frame index, etc.).
Step 2: Understand the Core ME Workflow
In x265, the encoder calls ME::search() (in me.cpp), which invokes MotionEstimation::estimate() (in motion_estimation.cpp) to perform the actual search for each block. The process goes roughly like this:
- Take the current frame’s block and the corresponding region in the reference frame.
- Run integer-pixel search to find the best matching block.
- Refine the result with fractional-pixel search to get sub-pixel accuracy.
- Output the
MotionVectorfor that block.
Step 3: Strip Down the ME Module for Standalone Use
If you want to use x265’s ME independently (outside the full encoder), you’ll need to:
- Extract the
MotionEstimationclass and its dependencies (e.g., YUV buffer handling, interpolation functions fromsource/common/filter.cpp). - Write a wrapper function that:
- Loads your two input frames into
PicYuvstructures. - Initializes
MEParamwith your desired settings (e.g., search range = ±32, block size = 16x16). - Calls
MotionEstimation::estimate()for each block in the frame. - Collects the resulting
MotionVectorobjects for all blocks.
- Loads your two input frames into
Step 4: Validate Your Implementation
To make sure you’re getting correct MVs, use x265’s built-in MV dump feature as a reference:
- Run x265 with this command to generate a text file of MVs:
x265 --dump-mvs output_mvs.txt input.yuv - Compare this output to your standalone ME code’s results to verify accuracy.
3. Alternative: Get HEVC Motion Vectors from Encoded Streams
If you don’t need to run ME from scratch, you can extract MVs from existing HEVC streams using FFmpeg:
- Command Line: Dump MVs directly to your console with:
ffmpeg -i input.hevc -flags2 +export_mvs -f rawvideo /dev/null - Programmatic Access: Use FFmpeg’s libavcodec API. When decoding HEVC frames, access the
motion_vectorsfield in theAVFramestruct (available in FFmpeg 4.0+) to get per-block MV data directly.
内容的提问来源于stack exchange,提问作者Gary

