基于Mediapipe与OpenCV-Python求解人脸姿态角的疑问
人脸姿态角(Pitch/Yaw/Roll)计算方案
核心问题解决:Mediapipe关键点与solvePnP点的对应
你提到的6个关键点位,对应Mediapipe Face Mesh的官方固定关键点索引(基于v0.8+版本):
- 鼻尖:1
- 下巴尖:152
- 左眼左角:33
- 右眼右角:263
- 左嘴角:61
- 右嘴角:291
3D目标点与2D图像点的准备规则
3D目标点
用通用标准人脸的相对3D坐标(单位:毫米)即可,不需要实际深度数据,示例如下:
# 3D参考点(顺序:鼻尖、下巴、左眼左角、右眼右角、左嘴角、右嘴角) model_points = np.array([ (0.0, 0.0, 0.0), # 鼻尖 (0.0, -330.0, -65.0), # 下巴 (-225.0, 170.0, -135.0), # 左眼左角 (225.0, 170.0, -135.0), # 右眼右角 (-150.0, -150.0, -125.0), # 左嘴角 (150.0, -150.0, -125.0) # 右嘴角 ], dtype=np.float32)
2D图像点
从Mediapipe检测结果中提取对应索引的x、y像素坐标:
- Mediapipe返回的是归一化坐标(范围0-1),需要乘以图像的实际宽高转换为像素值
- 仅需提取x、y坐标,z坐标(深度)不需要传入solvePnP
示例代码:
# 假设已通过Mediapipe获取face_landmarks对象,img_w、img_h为图像宽高 image_points = np.array([ (face_landmarks.landmark[1].x * img_w, face_landmarks.landmark[1].y * img_h), (face_landmarks.landmark[152].x * img_w, face_landmarks.landmark[152].y * img_h), (face_landmarks.landmark[33].x * img_w, face_landmarks.landmark[33].y * img_h), (face_landmarks.landmark[263].x * img_w, face_landmarks.landmark[263].y * img_h), (face_landmarks.landmark[61].x * img_w, face_landmarks.landmark[61].y * img_h), (face_landmarks.landmark[291].x * img_w, face_landmarks.landmark[291].y * img_h) ], dtype=np.float32)
完整计算流程代码示例
import cv2 import numpy as np import mediapipe as mp # 初始化Mediapipe人脸网格检测器 mp_face_mesh = mp.solutions.face_mesh face_mesh = mp_face_mesh.FaceMesh( static_image_mode=True, max_num_faces=1, min_detection_confidence=0.5 ) # 预设3D参考点 model_points = np.array([ (0.0, 0.0, 0.0), (0.0, -330.0, -65.0), (-225.0, 170.0, -135.0), (225.0, 170.0, -135.0), (-150.0, -150.0, -125.0), (150.0, -150.0, -125.0) ], dtype=np.float32) def calculate_head_pose(img): img_h, img_w = img.shape[:2] # 简化相机内参(焦距设为图像宽度,主点在图像中心) camera_matrix = np.array([ [img_w, 0, img_w / 2], [0, img_w, img_h / 2], [0, 0, 1] ], dtype=np.float32) dist_coeffs = np.zeros((4, 1), dtype=np.float32) # 默认无畸变 # 检测人脸关键点 rgb_img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB) results = face_mesh.process(rgb_img) if not results.multi_face_landmarks: return None face_landmarks = results.multi_face_landmarks[0] # 提取2D图像点 image_points = np.array([ (face_landmarks.landmark[1].x * img_w, face_landmarks.landmark[1].y * img_h), (face_landmarks.landmark[152].x * img_w, face_landmarks.landmark[152].y * img_h), (face_landmarks.landmark[33].x * img_w, face_landmarks.landmark[33].y * img_h), (face_landmarks.landmark[263].x * img_w, face_landmarks.landmark[263].y * img_h), (face_landmarks.landmark[61].x * img_w, face_landmarks.landmark[61].y * img_h), (face_landmarks.landmark[291].x * img_w, face_landmarks.landmark[291].y * img_h) ], dtype=np.float32) # 求解PnP获取旋转向量 success, rotation_vec, translation_vec = cv2.solvePnP( model_points, image_points, camera_matrix, dist_coeffs ) if not success: return None # 旋转向量转旋转矩阵 rotation_matrix, _ = cv2.Rodrigues(rotation_vec) # 分解旋转矩阵得到欧拉角 proj_matrix = np.hstack((rotation_matrix, translation_vec)) euler_angles = cv2.decomposeProjectionMatrix(proj_matrix)[6] pitch, yaw, roll = euler_angles.flatten() # 调整角度方向符合常规定义(低头负/抬头正,左转负/右转正,左倾负/右倾正) pitch = -pitch yaw = -yaw roll = -roll return round(pitch, 2), round(yaw, 2), round(roll, 2) # 测试示例 if __name__ == "__main__": test_img = cv2.imread("test_face.jpg") pose_angles = calculate_head_pose(test_img) if pose_angles: print(f"Pitch: {pose_angles[0]}°, Yaw: {pose_angles[1]}°, Roll: {pose_angles[2]}°")
关键说明
- 相机内参:示例用的是简化参数,若需更高精度,可替换为实际相机的标定参数
- 角度方向:OpenCV默认输出的欧拉角方向与常规定义有差异,需通过符号调整对齐
- 关键点稳定性:上述6个点都是Mediapipe中检测稳定性较高的点位,适合姿态计算
内容的提问来源于stack exchange,提问作者Kanan Mahammadli
相关产品推荐
相关产品推荐

