You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Vision Framework应用于视频播放:实现姿态检测与绘制

问题:本地视频播放集成Vision人体姿态检测与绘制

我已经实现了基于实时摄像头采集的Vision Framework应用,现在想在本地视频播放时也实现人体姿态检测与绘制功能。目前我有一段播放本地视频的代码:

let videoURL = URL(fileURLWithPath: NSString.path(withComponents: [documentsDirectory, path]) as String)
let player = AVPlayer(url: videoURL)
let vc = AVPlayerViewController()
vc.player = player

present(vc, animated: true) {
    vc.player?.play()
}

请问如何对视频每一帧先通过以下代码用Vision Framework检测人体姿态:

let visionRequestHandler = VNImageRequestHandler(cgImage: frame)

// Use Vision to find human body poses in the frame.
do { try visionRequestHandler.perform([humanBodyPoseRequest]) } catch {
    assertionFailure("Human Pose Request failed: \(error)")
}

let poses = Pose.fromObservations(humanBodyPoseRequest.results)

并通过pose.drawWireframeToContext(cgContext, applying: pointTransform)绘制姿态后,再将修改后的视频帧传给AVPlayer播放?


解决方案

直接使用AVPlayerViewController无法拦截并修改视频帧,需要改用自定义视频渲染流程,通过AVPlayerItemVideoOutput获取原始帧,处理后再渲染到自定义视图中。以下是具体实现步骤:

1. 准备自定义播放视图

创建继承自UIView的自定义视图,承载视频播放层和姿态绘制层:

class PoseDetectVideoView: UIView {
    // 视频播放层
    let playerLayer = AVPlayerLayer()
    // 姿态绘制层
    let poseDrawLayer = CAShapeLayer()
    
    override init(frame: CGRect) {
        super.init(frame: frame)
        setupLayers()
    }
    
    required init?(coder: NSCoder) {
        super.init(coder: coder)
        setupLayers()
    }
    
    private func setupLayers() {
        playerLayer.frame = bounds
        layer.addSublayer(playerLayer)
        
        poseDrawLayer.frame = bounds
        poseDrawLayer.fillColor = UIColor.clear.cgColor
        poseDrawLayer.strokeColor = UIColor.red.cgColor
        poseDrawLayer.lineWidth = 2
        layer.addSublayer(poseDrawLayer)
    }
    
    override func layoutSubviews() {
        super.layoutSubviews()
        playerLayer.frame = bounds
        poseDrawLayer.frame = bounds
    }
}

2. 配置AVPlayer与视频输出

初始化AVPlayer和AVPlayerItemVideoOutput,通过输出获取每一帧:

// 初始化视频资源
let videoURL = URL(fileURLWithPath: NSString.path(withComponents: [documentsDirectory, path]) as String)
let asset = AVURLAsset(url: videoURL)
let playerItem = AVPlayerItem(asset: asset)
let player = AVPlayer(playerItem: playerItem)

// 配置视频输出,指定像素格式
let pixelBufferAttributes: [String: Any] = [
    kCVPixelBufferPixelFormatTypeKey as String: kCVPixelFormatType_32BGRA,
    kCVPixelBufferWidthKey as String: NSNumber(value: Int(UIScreen.main.bounds.width)),
    kCVPixelBufferHeightKey as String: NSNumber(value: Int(UIScreen.main.bounds.height))
]
let videoOutput = AVPlayerItemVideoOutput(pixelBufferAttributes: pixelBufferAttributes)
playerItem.add(videoOutput)

// 创建自定义视图并添加到当前控制器
let videoView = PoseDetectVideoView(frame: view.bounds)
videoView.playerLayer.player = player
view.addSubview(videoView)

// 用CADisplayLink同步帧处理节奏
let displayLink = CADisplayLink(target: self, selector: #selector(handleFrameUpdate))
displayLink.add(to: .main, forMode: .common)

3. 处理每一帧并绘制姿态

实现帧处理逻辑,将原始帧转换为CGImage,用Vision检测姿态后绘制到图层:

// 提前初始化Vision姿态请求
let humanBodyPoseRequest = VNDetectHumanBodyPoseRequest()

@objc private func handleFrameUpdate() {
    guard let playerItem = player.currentItem,
          let videoOutput = playerItem.outputs.first as? AVPlayerItemVideoOutput,
          let time = playerItem.currentTime().isValid ? playerItem.currentTime() : nil else {
        return
    }
    
    // 获取当前帧的像素缓冲区
    guard let pixelBuffer = videoOutput.copyPixelBuffer(forItemTime: time, itemTimeForDisplay: nil) else {
        return
    }
    
    // 将像素缓冲区转换为CGImage
    let ciImage = CIImage(cvImageBuffer: pixelBuffer)
    guard let cgImage = CIContext().createCGImage(ciImage, from: ciImage.extent) else {
        return
    }
    
    // 执行Vision姿态检测
    let visionRequestHandler = VNImageRequestHandler(cgImage: cgImage)
    do {
        try visionRequestHandler.perform([humanBodyPoseRequest])
    } catch {
        assertionFailure("Human Pose Request failed: \(error)")
        return
    }
    
    // 获取检测到的姿态
    guard let poses = Pose.fromObservations(humanBodyPoseRequest.results) as? [Pose], !poses.isEmpty else {
        videoView.poseDrawLayer.path = nil
        return
    }
    
    // 创建绘图上下文并绘制姿态
    UIGraphicsBeginImageContextWithOptions(cgImage.size, false, UIScreen.main.scale)
    guard let cgContext = UIGraphicsGetCurrentContext() else {
        UIGraphicsEndImageContext()
        return
    }
    
    // 计算帧到视图的适配变换
    let scale = min(videoView.bounds.width / cgImage.width, videoView.bounds.height / cgImage.height)
    let pointTransform = CGAffineTransform(scaleX: scale, y: scale)
        .translatedBy(x: (videoView.bounds.width - cgImage.width * scale)/2, y: (videoView.bounds.height - cgImage.height * scale)/2)
    
    // 绘制姿态骨架
    for pose in poses {
        pose.drawWireframeToContext(cgContext, applying: pointTransform)
    }
    
    // 更新绘制图层的路径
    videoView.poseDrawLayer.path = cgContext.path
    UIGraphicsEndImageContext()
}

// 启动播放
player.play()

4. 性能优化提示

  • 降低处理分辨率:在pixelBufferAttributes中设置更小的宽高,减少Vision计算量
  • 异步处理:将Vision检测逻辑放到后台队列执行,避免阻塞主线程
  • 跳帧处理:比如每2帧处理一次,平衡性能与流畅度

内容的提问来源于stack exchange,提问作者Philipp Dobrigkeit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 01:45:33