基于OpenPose提取的2D人体关键点,如何从2D图像生成3D人体关键点?
Hey there! So you've already got 2D keypoints from OpenPose, and now you want to turn those into 3D ones—great question. Let's walk through a few practical ways to do this, depending on your setup:
Option 1: Use OpenPose's Native 3D Reconstruction
OpenPose actually has built-in support for 3D keypoint reconstruction if you have multi-view data (same person captured from multiple synchronized angles). Here's how to integrate this into your current workflow:
- First, if you haven't already, extract 2D keypoints for each camera view separately using your existing OpenPose script (just repeat the process for each view's ROI folder).
- Run OpenPose's 3D reconstruction module, pointing it to your multi-view 2D results. For example, if you have two views with 2D JSON outputs in
/content/output_json_view1and/content/output_json_view2, run this command:
!cd openpose && ./build/examples/openpose/openpose.bin --3d \ --image_dir /content/extracted_roi_view1 \ --image_dir /content/extracted_roi_view2 \ --write_json /content/output_json_3d \ --write_keypoints_3d /content/output_3d.json \ --display 0 --render_pose 0
- Parse the resulting
output_3d.jsonfile—look for thepose_keypoints_3dfield, which contains each keypoint's (x, y, z, accuracy) values. You can adapt your existing JSON parsing code to handle this 3D data.
Note: This works best with calibrated cameras, but OpenPose can do basic reconstruction even without explicit calibration if the views are synchronized.
Option 2: Single-Image 3D Pose Lift Models
If you only have single-view data (just one image/angle), you'll need a "lift" model that predicts 3D keypoints from 2D ones. Here are two easy-to-implement options:
MediaPipe Pose 3D
MediaPipe has a pre-trained model that can directly output 3D keypoints from images. You can either process your original images directly, or map your existing OpenPose 2D keypoints to MediaPipe's format (though processing the image is simpler):
import cv2 import mediapipe as mp import numpy as np # Initialize MediaPipe Pose mp_pose = mp.solutions.pose pose = mp_pose.Pose(static_image_mode=True, min_detection_confidence=0.5) # Load your image (replace with your actual image path) img = cv2.imread("/content/extracted_roi/your_image.jpg") rgb_img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB) # Process the image to get 3D keypoints results = pose.process(rgb_img) if results.pose_landmarks: # Extract 3D landmarks (normalized coordinates; scale back if needed) kps_3d = [] for lm in results.pose_landmarks.landmark: # Convert normalized coordinates to image pixel scale if desired x = lm.x * img.shape[1] y = lm.y * img.shape[0] z = lm.z * img.shape[1] # Z is relative to image width kps_3d.append([x, y, z]) # Save to text file (matching your existing output format) with open("/content/output_3d_kps.txt", "w") as f: for kp in kps_3d: f.write(",".join(map(str, kp)) + "\n")
Open-Source Lift Models
If you want to use your existing OpenPose 2D keypoints directly, check out models like HRNet3D or SimpleBaseline3D. These take 2D keypoint sequences (or single frames) and output 3D coordinates. You can load a pre-trained model and feed your all_kp.txt data into it.
Option 3: Multi-View Geometry (Structure from Motion)
For more precise 3D reconstruction with calibrated cameras, use classic computer vision techniques like triangulation. Here's a quick outline using OpenCV:
- Calibrate your cameras: Use OpenCV's
calibrateCamerafunction to get intrinsic parameters (focal length, distortion) and extrinsic parameters (rotation/translation between cameras). - Triangulate 3D points: For each pair of matching 2D keypoints across views, use
cv2.triangulatePointsto compute the 3D coordinate:
import cv2 import numpy as np # Example parameters (replace with your actual calibrated values) K1 = np.array([[fx1, 0, cx1], [0, fy1, cy1], [0, 0, 1]]) # Camera 1 intrinsic matrix K2 = np.array([[fx2, 0, cx2], [0, fy2, cy2], [0, 0, 1]]) # Camera 2 intrinsic matrix R = np.array([[r11, r12, r13], [r21, r22, r23], [r31, r32, r33]]) # Rotation from cam1 to cam2 T = np.array([tx, ty, tz]) # Translation from cam1 to cam2 # Projection matrices P1 = K1 @ np.hstack((np.eye(3), np.zeros((3,1)))) P2 = K2 @ np.hstack((R, T.reshape(3,1))) # Load your OpenPose 2D keypoints for each view (shape: (18, 2)) kps_view1 = np.loadtxt("/content/output_list/view1/all_kp.txt", delimiter=",") kps_view2 = np.loadtxt("/content/output_list/view2/all_kp.txt", delimiter=",") # Triangulate each keypoint kps_3d = [] for i in range(18): # Convert to homogeneous coordinates pt1 = kps_view1[i].reshape(2,1) pt2 = kps_view2[i].reshape(2,1) # Compute 3D point pt_3d_hom = cv2.triangulatePoints(P1, P2, pt1, pt2) pt_3d = pt_3d_hom[:3,0] / pt_3d_hom[3,0] kps_3d.append(pt_3d) # Save results np.savetxt("/content/output_3d_sfm.txt", np.array(kps_3d), delimiter=",")
Quick Recommendation
- If you have multi-view data: Go with OpenPose's native 3D tool (it's the easiest to integrate with your current workflow).
- If you only have single-view data: Use MediaPipe Pose 3D for fast, pre-trained results, or a dedicated lift model for better accuracy.
内容的提问来源于stack exchange,提问作者k_p

