请求协助:OpenPose与Kinect骨骼数据结构互转实现
Hey there! I’ve tackled skeleton format conversions for academic pose analysis projects before, so I can walk you through how to build bidirectional mapping between Kinect’s NUI_Skeleton_Data and OpenPose’s output structures. Let’s break this down step by step.
First, you need to align the joint definitions of both systems. Here’s a quick breakdown of their core counts and key joints:
- Kinect: 20 tracked joints (e.g., Head, ShoulderCenter, HipCenter, Left/Right Elbow/Knee, etc.) with 3D coordinates (x, y, z) plus a tracking state (tracked, inferred, not tracked).
- OpenPose: Supports multiple skeletons (e.g., BODY_25 with 25 joints, COCO with 17 joints). For most academic work, BODY_25 is the most comprehensive, including joints like ears, ankles, and toes that Kinect doesn’t track natively.
1. Kinect → OpenPose Conversion
Follow these steps to map Kinect data to OpenPose’s format:
Step 1: Define a joint mapping table
Create a direct mapping between Kinect’s joint indices and their closest OpenPose counterparts. For example (using BODY_25 indices):Kinect Joint Index Kinect Joint Name OpenPose BODY_25 Index OpenPose Joint Name 0 Head 0 Nose 1 ShoulderCenter 1 Neck 2 Spine 8 MidSpine 4 ShoulderLeft 14 L_Shoulder 5 ElbowLeft 16 L_Elbow 6 WristLeft 18 L_Wrist ... ... ... ... Step 2: Handle missing joints
OpenPose has joints Kinect doesn’t track (like ears, toes). For these, you can:- Estimate their position using adjacent joints (e.g., offset ears from the head joint by a small pixel value).
- Mark them as low-confidence (0.0) if you can’t infer them reliably.
Step 3: Convert coordinate systems
Kinect uses a 3D camera coordinate system (x: left-right, y: up-down, z: depth). OpenPose outputs 2D image coordinates (x: left-right, y: top-bottom) plus a confidence score. To convert:- Use Kinect’s camera calibration parameters to project 3D points to 2D image space.
- Map Kinect’s tracking state to OpenPose’s confidence score (e.g., tracked = 1.0, inferred = 0.5, not tracked = 0.0).
2. OpenPose → Kinect Conversion
Reverse mapping requires filling in Kinect’s unique joints (like ShoulderCenter, HipCenter) that OpenPose doesn’t track directly:
- Step 1: Reverse the joint mapping
Map OpenPose joints back to Kinect’s 20 joints. For example:- OpenPose’s
Neck→ Kinect’sShoulderCenter - OpenPose’s
L_Hip+R_Hipmidpoint → Kinect’sHipCenter
- OpenPose’s
- Step 2: Estimate missing Kinect joints
Kinect’sShoulderCenterandHipCenterare midpoints of their respective left/right joints. Calculate these midpoints using OpenPose’s left/right shoulder/hip coordinates. - Step 3: Add 3D data (if needed)
If you need Kinect-style 3D data, use OpenPose’s 3D output (if available) or infer depth using camera calibration and 2D positions.
Here’s a simplified function for Kinect to OpenPose BODY_25 conversion:
# Kinect to OpenPose BODY_25 joint index mapping KINECT_TO_OP = { 0: 0, # Head → Nose 1: 1, # ShoulderCenter → Neck 2: 8, # Spine → MidSpine 3: 9, # HipCenter → Hip (midpoint of L/R Hip) 4: 14, # ShoulderLeft → L_Shoulder 5: 16, # ElbowLeft → L_Elbow 6: 18, # WristLeft → L_Wrist 7: 20, # HandLeft → L_Hand 8: 13, # ShoulderRight → R_Shoulder 9: 15, # ElbowRight → R_Elbow 10:17, # WristRight → R_Wrist 11:19, # HandRight → R_Hand 12:22, # HipLeft → L_Hip 13:24, # KneeLeft → L_Knee 14:26, # AnkleLeft → L_Ankle 15:28, # FootLeft → L_Foot 16:21, # HipRight → R_Hip 17:23, # KneeRight → R_Knee 18:25, # AnkleRight → R_Ankle 19:27 # FootRight → R_Foot } def convert_kinect_to_openpose(kinect_skeleton, img_width=640, img_height=480): # Initialize OpenPose BODY_25 skeleton (25 joints: (x, y, confidence)) op_skeleton = [(0.0, 0.0, 0.0) for _ in range(25)] for kinect_idx, op_idx in KINECT_TO_OP.items(): x_kinect, y_kinect, z_kinect, tracking_state = kinect_skeleton[kinect_idx] # Convert Kinect 3D to 2D image coordinates (simplified projection) x_op = (x_kinect + 1) * (img_width / 2) # Kinect x ranges from -1 to 1 y_op = (1 - y_kinect) * (img_height / 2) # Kinect y is up, OpenPose y is down # Map tracking state to confidence if tracking_state == 2: # Tracked confidence = 1.0 elif tracking_state == 1: # Inferred confidence = 0.5 else: # Not tracked confidence = 0.0 op_skeleton[op_idx] = (x_op, y_op, confidence) # Estimate ears from head position head_x, head_y, head_conf = op_skeleton[0] op_skeleton[15] = (head_x - 20, head_y, head_conf * 0.8) # L_Ear op_skeleton[16] = (head_x + 20, head_y, head_conf * 0.8) # R_Ear return op_skeleton
- Document your mapping: Clearly outline your conversion logic in your paper/code to ensure reproducibility.
- Validate accuracy: Test your conversion on a small labeled dataset to check joint position errors.
- Handle occlusion: Both systems handle occlusion differently—make sure to flag joints with low confidence/tracking state in your analysis.
- Choose the right OpenPose skeleton: If you don’t need extra joints, use COCO (17 joints) for a simpler mapping to Kinect’s 20 joints.
内容的提问来源于stack exchange,提问作者Dvir

