为何平移是本质矩阵的零向量?本质矩阵与平移正交性原理探究
Awesome question—this is one of those results that clicks once you connect the mathematical definitions to the geometry of camera motion. Let's break it down into two parts: the formal math behind it, and the intuitive geometric reason that makes it make sense.
Mathematical Breakdown
First, let's recall the core definition of the essential matrix ( E ): it encodes the relative rotation ( R ) and translation ( T ) between two cameras. There are two common conventions for writing ( E ), depending on which camera's coordinate system you prioritize:
- ( E = [T]\times R ), where ( [T]\times ) is the skew-symmetric matrix of the translation vector ( T )
- ( E = R [T]_\times )
The key property here comes from skew-symmetric matrices: by definition, ( [T]\times v = T \times v ) (the cross product of ( T ) and ( v )). And a fundamental rule of cross products is that any vector cross product with itself is the zero vector—so ( T \times T = 0 ), which translates to ( [T]\times T = 0 ) in matrix form.
Let's verify both conventions:
- For ( E = R [T]\times ): Multiply ( E ) by ( T ), and you get ( E T = R [T]\times T = R \times 0 = 0 ). Straightforward—we're just rotating the zero vector, which stays zero.
- For ( E = [T]\times R ): Here, we look at ( E^T T ) instead. Since skew-symmetric matrices satisfy ( [T]\times^T = -[T]_\times ), and rotation matrices are orthogonal (( R^T R = I )):
E^T T = (R^T [T]_\times^T) T = R^T (-[T]_\times) T = R^T (-(T \times T)) = R^T 0 = 0
That's why you'll see either ( E T = 0 ) or ( E^T T = 0 ) cited—both stem from the same cross product property, just depending on how ( E ) is defined.
Intuitive Geometric Explanation
The essential matrix exists to enforce the epipolar constraint: for any 3D point ( P ), its projection onto camera 1 (( x_1 )) and camera 2 (( x_2 )) must satisfy ( x_2^T E x_1 = 0 ). This constraint boils down to a simple fact: the two camera centers (( O_1, O_2 )) and the point ( P ) must lie on the same plane.
Now, ( T ) is the vector from ( O_1 ) to ( O_2 )—the "baseline" between the two cameras. Why is this vector a null vector of ( E )? Think about what ( T ) represents in terms of the epipolar constraint:
- If we treat ( T ) as a "point" in camera 1's coordinate system, it's actually the position of camera 2's center ( O_2 ) relative to ( O_1 ).
- The projection of ( O_2 ) onto camera 2's image plane is just the origin (since ( O_2 ) is camera 2's own center).
- Plugging into the epipolar constraint: ( 0^T E T = 0 ), which holds trivially. But more intuitively: the essential matrix maps a point's projection to the epipolar line (the line where the point's projection must lie in the other camera). For ( T ), the epipolar line collapses to a single point (camera 2's center), which corresponds to the zero vector.
In short: the translation vector represents the line between the two camera centers. When you feed this line into the essential matrix (which maps points to epipolar lines), you get a degenerate line (a point) because the line is the baseline itself—there's no "spread" of possible projections, just the fixed center of the second camera.
内容的提问来源于stack exchange,提问作者melon Z

