基于投影深度图计算法线图N(i,j)的技术问询
Alright, let's walk through exactly how to compute your normal map N(i,j) from a projected depth map D(i,j) when working with a camera intrinsic matrix K (the [K|0] setup you mentioned). Here's a step-by-step breakdown tailored to your scenario:
K First, let's formalize the intrinsic matrix structure—this is critical for converting 2D pixels back to 3D coordinates. Your K matrix will look like this (in pixel units):
K = [ fx 0 cx ] [ 0 fy cy ] [ 0 0 1 ]
Where:
fx,fy: Focal lengths along the X and Y axescx,cy: Coordinates of the camera's principal point (usually the center of the image)
Since D(i,j) is the projected depth (the Z-value in camera space for the 3D point that projects to pixel (i,j)), we can reverse the projection formula to get the full 3D point P(i,j) = (X, Y, Z) in camera space:
X = (i - cx) * D(i,j) / fx Y = (j - cy) * D(i,j) / fy Z = D(i,j)
Note: Double-check your pixel coordinate origin—most image formats use top-left as (0,0), which aligns with this formula as long as cx/cy are defined relative to that origin.
To calculate the surface normal at (i,j), we need two non-parallel vectors lying on the surface at that point. A simple approach is to use adjacent pixels—for example, the pixel to the right (i+1,j) and the pixel below (i,j+1):
- Compute vector
V1 = P(i+1,j) - P(i,j)(points from(i,j)to(i+1,j)in 3D space) - Compute vector
V2 = P(i,j+1) - P(i,j)(points from(i,j)to(i,j+1)in 3D space)
Expanding these vectors explicitly (to make implementation easier):
For V1:
V1.X = [(i+1 - cx)*D(i+1,j)/fx] - [(i - cx)*D(i,j)/fx] V1.Y = [(j - cy)*D(i+1,j)/fy] - [(j - cy)*D(i,j)/fy] V1.Z = D(i+1,j) - D(i,j)
For V2:
V2.X = [(i - cx)*D(i,j+1)/fx] - [(i - cx)*D(i,j)/fx] V2.Y = [(j+1 - cy)*D(i,j+1)/fy] - [(j - cy)*D(i,j)/fy] V2.Z = D(i,j+1) - D(i,j)
The surface normal N(i,j) is the cross product of V1 and V2—this gives a vector perpendicular to both surface vectors:
N.X = V1.Y * V2.Z - V1.Z * V2.Y N.Y = V1.Z * V2.X - V1.X * V2.Z N.Z = V1.X * V2.Y - V1.Y * V2.X
The cross product result won't necessarily be a unit vector, so we need to normalize it to ensure consistent magnitude across the normal map:
norm = sqrt(N.X² + N.Y² + N.Z²) N_normalized = (N.X / norm, N.Y / norm, N.Z / norm)
If norm is 0 (this happens if V1 and V2 are parallel, e.g., flat depth or invalid depth), you can skip normalization and assign a default value like (0, 0, 1) (pointing along the camera's forward axis).
- Boundary Pixels: Pixels on the image edges don't have full neighbors—you can either copy the nearest valid normal, assign a default, or pad the depth map with mirrored values before processing.
- Depth Discontinuities: Object edges will produce noisy normals. To mitigate this, apply a Gaussian blur to the depth map first, or use a larger 3x3 neighborhood (average cross products from multiple adjacent vector pairs).
- Normal Direction: In camera space, the Z-axis points toward the scene. If you need normals to point outward from the object (instead of toward the camera), multiply the normalized normal by
-1. - Invalid Depth Values: If
D(i,j)is 0, NaN, or outside a valid range, mark those normals as invalid (e.g., set to(0,0,0)).
内容的提问来源于stack exchange,提问作者ASML

