You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Create ML训练的Action ML Classifier集成iOS应用后仅输出单一动作结果的问题排查求助

Hey there, let's break down why your action classifier is stuck predicting only one label, even though it worked flawlessly in Create ML. The issues are likely tied to how you're preparing input data for the model—let's go through key fixes step by step:

1. You're Feeding Single Frames + Empty Padding Instead of Continuous Sequences

This is probably the biggest culprit. Your current code processes one frame at a time, takes that single pose observation, and pads it with 59 empty frames to hit your 60-frame sequence length. But Create ML trained on continuous sequences of 60 action frames, not a single frame plus empty data. The model sees mostly zero values and can't learn the temporal pattern of your actions.

Fix: Implement a Sliding Pose Window

Maintain a buffer to collect the last 60 pose observations, and only run predictions when the buffer is full:

class YourPoseClassifier {
    private var poseWindow: [VNRecognizedPointsObservation] = []
    private let requiredSequenceLength = 60 // Match the sequence length you used in Create ML
    private var fitnessClassifier: PlayerExcercise?

    init() {
        // Initialize the model ONCE (not every prediction!)
        do {
            fitnessClassifier = try PlayerExcercise(configuration: MLModelConfiguration())
        } catch {
            print("Failed to load model: \(error)")
        }
    }

    func didOutput(pixelBuffer: CVPixelBuffer) {
        extractPoses(from: pixelBuffer)
    }

    private func extractPoses(from pixelBuffer: CVPixelBuffer) {
        let handler = VNImageRequestHandler(cvPixelBuffer: pixelBuffer)
        let request = VNDetectHumanBodyPoseRequest { [weak self] request, error in
            guard let self = self, error == nil else { return }
            
            if let observations = request.results as? [VNRecognizedPointsObservation], let latestPose = observations.first {
                // Add new pose to the window
                self.poseWindow.append(latestPose)
                // Trim the window to keep only the last N frames
                if self.poseWindow.count > self.requiredSequenceLength {
                    self.poseWindow.removeFirst()
                }
                // Only predict when we have a full sequence
                if self.poseWindow.count == self.requiredSequenceLength {
                    if let prediction = try? self.makePrediction(with: self.poseWindow, imageSize: CGSize(width: CVPixelBufferGetWidth(pixelBuffer), height: CVPixelBufferGetHeight(pixelBuffer))) {
                        print("\(prediction.label), confidence: \(prediction.confidence)")
                    }
                }
            }
        }

        do {
            try handler.perform([request])
        } catch {
            print("Pose detection failed: \(error)")
        }
    }
}

2. Verify Keypoint Conversion Matches Your Training Data

Your keypointsMultiArray() method (custom, since Vision doesn't include this natively) must produce data in the exact same format as the training data in Create ML. Common mismatches include:

  • Coordinate normalization: If you trained on normalized coordinates (0-1 relative to image size) but are feeding raw pixel values, the model will misinterpret the data.
  • Keypoint order: Ensure the order of keypoints (e.g., nose, left shoulder, right shoulder) matches what Create ML used.
  • Shape: Confirm the per-frame array shape matches (e.g., [3, 18] for x/y/confidence per 18 keypoints, not reversed).

Fix: Standardize Keypoint Normalization and Shape

Update your keypoint conversion extension to explicitly normalize coordinates and match training shape:

extension VNRecognizedPointsObservation {
    func keypointsMultiArray(normalizedTo imageSize: CGSize) throws -> MLMultiArray {
        let allKeypoints = try recognizedPoints(forGroupKey: .all)
        let expectedKeypointCount = 18 // Match your Create ML training setup
        let multiArray = try MLMultiArray(shape: [3, expectedKeypointCount] as [NSNumber], dataType: .double)

        for (index, keypoint) in VNRecognizedPoint.allBodyPoints.enumerated() {
            guard index < expectedKeypointCount, let point = allKeypoints[keypoint] else { continue }
            
            // Normalize x/y to 0-1 range (matches Create ML's default preprocessing)
            multiArray[[0, index] as [NSNumber]] = (point.location.x / imageSize.width) as NSNumber
            multiArray[[1, index] as [NSNumber]] = (point.location.y / imageSize.height) as NSNumber
            // Include confidence score
            multiArray[[2, index] as [NSNumber]] = point.confidence as NSNumber
        }

        return multiArray
    }
}

3. Fix Model Initialization Efficiency

Your original code initializes the PlayerExcercise model every time you run a prediction. This is slow and can introduce unexpected behavior. Initialize the model once as a class property (as shown in the sliding window code above).

4. Validate Input Shape Against the Model

Double-check that the final input MLMultiArray matches the model's expected input shape. Add this line right before making the prediction to confirm:

print("Input shape: \(modelInput.shape)")

Compare this to the input shape shown in Create ML (you can find this in the model's metadata or by inspecting the model file). It should look like [60, 3, 18] (sequence length × 3 values per keypoint × 18 keypoints).

Final Check: Compare Training vs. App Data

Take a sample sequence from your Create ML training data, export its raw keypoint values, and compare them to the MLMultiArray output from your app. If the values (range, order, shape) don't match, that's the root cause.

内容的提问来源于stack exchange,提问作者Senthil Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 11:23:08