如何将ARKit返回的matrix_float4x4转换为CGRect?技术咨询
Hey there! Great question—let's tackle this because converting a 3D transform matrix (matrix_float4x4) to a 2D CGRect isn't a direct one-step process, but it's totally doable with a bit of context about your use case.
Key Context First
The worldTransform/localTransform matrices from ARHitTestResult describe a 3D spatial position, rotation, and scale in ARKit's coordinate system. A CGRect, on the other hand, is a 2D screen-based rectangle. So we need to bridge these two by:
- Extracting the 3D position from the matrix
- Projecting that 3D space (or the object's bounds) onto the 2D screen
- Calculating the rectangle based on either your Vision detection results or the AR object's physical bounds
Step-by-Step Approaches
Approach 1: Combine AR Hit Test with Vision Detection Results
Since you're already using Vision for object detection, you can pair the AR hit test's 3D position with Vision's VNDetectedObjectObservation bounding box to create a screen-aligned CGRect:
// Assume you have your Vision detection result handy guard let visionObservation = yourVisionDetectionResult as? VNDetectedObjectObservation else { return } // 1. Extract the 3D world position from the AR hit test result let worldPosition = closestResult.worldTransform.columns.3 let simdWorldPosition = SIMD3<Float>(worldPosition.x, worldPosition.y, worldPosition.z) // 2. Project the 3D point to 2D screen coordinates (ARSCNView uses bottom-left origin) let screenPoint = arSceneView.projectPoint(simdWorldPosition) // Convert to UIKit's top-left origin let uiKitScreenPoint = CGPoint(x: screenPoint.x, y: arSceneView.bounds.height - screenPoint.y) // 3. Convert Vision's normalized bounding box to screen pixel dimensions let normalizedBounds = visionObservation.boundingBox let screenWidth = arSceneView.bounds.width let screenHeight = arSceneView.bounds.height let objectScreenWidth = normalizedBounds.width * screenWidth let objectScreenHeight = normalizedBounds.height * screenHeight // 4. Create the CGRect centered on the projected screen point let finalRect = CGRect( x: uiKitScreenPoint.x - (objectScreenWidth / 2), y: uiKitScreenPoint.y - (objectScreenHeight / 2), width: objectScreenWidth, height: objectScreenHeight )
Approach 2: For ARKit Virtual Objects (SCNNode)
If you're tracking a virtual 3D object in ARKit, use its SCNBoundingBox to project all corners to the screen and calculate the enclosing CGRect:
// Assume you have the SCNNode linked to your hit test result guard let targetNode = closestResult.node else { return } // 1. Get the object's local bounding box let boundingBox = targetNode.boundingBox let minLocal = boundingBox.min let maxLocal = boundingBox.max // 2. Define all 8 corners of the object's bounding box (local space) let localCorners: [SIMD3<Float>] = [ SIMD3(minLocal.x, minLocal.y, minLocal.z), SIMD3(maxLocal.x, minLocal.y, minLocal.z), SIMD3(minLocal.x, maxLocal.y, minLocal.z), SIMD3(maxLocal.x, maxLocal.y, minLocal.z), SIMD3(minLocal.x, minLocal.y, maxLocal.z), SIMD3(maxLocal.x, minLocal.y, maxLocal.z), SIMD3(minLocal.x, maxLocal.y, maxLocal.z), SIMD3(maxLocal.x, maxLocal.y, maxLocal.z) ] // 3. Project each corner to screen coordinates (convert to UIKit origin) var screenCorners = [CGPoint]() for corner in localCorners { let worldCorner = targetNode.convertPosition(SIMD3<Float>(corner), to: nil) let screenPoint = arSceneView.projectPoint(worldCorner) let uiKitPoint = CGPoint(x: screenPoint.x, y: arSceneView.bounds.height - screenPoint.y) screenCorners.append(uiKitPoint) } // 4. Calculate the minimal enclosing CGRect from all screen corners guard let firstCorner = screenCorners.first else { return } var minX = firstCorner.x var maxX = firstCorner.x var minY = firstCorner.y var maxY = firstCorner.y for corner in screenCorners { minX = min(minX, corner.x) maxX = max(maxX, corner.x) minY = min(minY, corner.y) maxY = max(maxY, corner.y) } let finalRect = CGRect(x: minX, y: minY, width: maxX - minX, height: maxY - minY)
Important Notes
- Coordinate System Difference:
ARSCNView.projectPoint(_:)returns a point with origin at the bottom-left, while UIKit uses top-left—don't forget the conversion step! - Depth Check: The
zvalue fromprojectPoint(_:)ranges from 0 (near camera plane) to 1 (far camera plane). Ifz > 1, the point is outside the camera's view, so skip it. - Real-World Object Accuracy: For physical objects, the Vision bounding box gives you the 2D size, but depth from ARKit will affect how large it appears on screen. You can refine accuracy by adding size estimation logic (e.g., using known object dimensions).
内容的提问来源于stack exchange,提问作者Alexandre Odet

