iOS平台下基于Vision与ARKit获取检测人脸的深度信息
Hey there! Let's tackle this issue you're facing with Vision face landmarks and ARKit node placement — I've been in a similar spot before, so I know how tricky that depth data gap can be. Here's how to fix it:
The core issue here is that Vision gives you 2D screen-based CGPoint landmarks, but ARKit needs 3D SCNVector3 coordinates with valid depth data to place nodes accurately. Let's break down the most reliable solutions:
1. Skip separate Vision detection — use ARKit's built-in face tracking directly
ARKit's ARFaceTrackingConfiguration already integrates Vision's face landmark detection, and it gives you 3D spatial coordinates out of the box. This is the simplest and most accurate approach:
- First, configure your AR session for face tracking:
let configuration = ARFaceTrackingConfiguration() configuration.isLightEstimationEnabled = true arView.session.run(configuration) - Then, in the
ARSCNViewDelegatemethod, grab the 3D landmarks directly from the face anchor:func renderer(_ renderer: SCNSceneRenderer, nodeFor anchor: ARAnchor) -> SCNNode? { guard let faceAnchor = anchor as? ARFaceAnchor else { return nil } // Get 3D position of predefined landmarks like left eye let leftEyeWorldPosition = faceAnchor.leftEyeTransform.columns.3 let eyeNode = SCNNode(geometry: SCNSphere(radius: 0.01)) eyeNode.position = SCNVector3(leftEyeWorldPosition.x, leftEyeWorldPosition.y, leftEyeWorldPosition.z) return eyeNode }
This method eliminates the need for manual 2D-to-3D conversion entirely, since ARKit handles depth and spatial positioning for you.
2. If you must use standalone Vision, add depth data manually
If you have to use Vision's 2D landmarks for some reason, you need to source depth data to convert them to 3D:
Option 1: Use ARKit's hitTest with depth
Pass the VisionCGPointto ARKit's hit test to get a corresponding 3D position:func convertVisionPointTo3D(_ point: CGPoint, in arView: ARSCNView) -> SCNVector3? { // Prioritize depth-based plane hits for accuracy let planeHits = arView.hitTest(point, types: .existingPlaneUsingDepth) if let planeResult = planeHits.first { return SCNVector3(planeResult.worldTransform.columns.3.x, planeResult.worldTransform.columns.3.y, planeResult.worldTransform.columns.3.z) } // Fall back to feature point hits if no planes are available let featureHits = arView.hitTest(point, types: .featurePoint) return featureHits.first?.worldTransform.columns.3 }Note: This can be unreliable if there are no planes or enough feature points near the face.
Option 2: Map Vision landmarks to ARKit's 3D face mesh
- Use Vision to get the face's bounding box and the relative position of each landmark within that box.
- Grab the 3D face mesh from ARKit's
ARFaceAnchor.geometry. - Calculate the 3D coordinate of each landmark by mapping its relative 2D position to the 3D mesh.
3. Why that CGPoint-to-SCNVector3 post might have failed
Most generic conversion posts skip the critical step of including depth data. Without a valid depth value, converting a 2D screen point to 3D will just place the node on the camera's far plane — which is totally useless for face-related placement. The fix is always to source depth either from ARKit's face anchor or hit test results.
My top recommendation is to use ARKit's native face tracking — it's designed exactly for this use case and saves you all the headache of manual conversion.
内容的提问来源于stack exchange,提问作者Zღk

