如何使用GCP Video Intelligence为视频每一帧标注姿态
GCP Video Intelligence 逐帧人体姿态标注解决方案
问题背景
GCP Video Intelligence框架可检测视频中的人体关键点,但默认仅返回间隔0.1秒的时间戳标注,无法直接生成视频每一帧的姿态数据。不想通过手动延长视频时长的方式获取逐帧结果,寻求合理替代方案。
可行解决方法
插值补全帧数据
利用API返回的0.1秒间隔关键点数据,通过线性插值或姿态平滑算法(如卡尔曼滤波)补全中间帧的关键点位置。无需重新调用API,成本低,适合对精度要求适中的场景。
实现思路:遍历相邻两个时间戳的关键点集合,根据视频帧率计算中间需补全的帧数,对每个关键点的x、y坐标按时间比例做线性插值,同时对置信度做平滑处理。单帧调用Vision API分析
将视频拆分为单帧图片,调用GCP Vision API的人体姿态检测功能对每一帧单独处理。这种方法能获取真正的逐帧精准标注,但需要额外处理视频拆帧和API批量调用,成本会相应增加。
注意:Vision API与Video Intelligence的姿态检测模型可能存在差异,需验证结果一致性。向官方提交功能需求
直接向GCP官方提交功能请求,建议在PersonDetectionConfig中添加自定义采样帧率的参数。若有足够多用户需求,官方可能在后续版本中支持该功能。
原示例代码(中文注释)
import io from google.cloud import videointelligence_v1 as videointelligence def detect_person(local_file_path="path/to/your/video-file.mp4"): """从本地文件检测视频中的人体""" client = videointelligence.VideoIntelligenceServiceClient() with io.open(local_file_path, "rb") as f: input_content = f.read() # 配置请求参数 config = videointelligence.types.PersonDetectionConfig( include_bounding_boxes=True, include_attributes=True, include_pose_landmarks=True, ) context = videointelligence.types.VideoContext(person_detection_config=config) # 发起异步请求 operation = client.annotate_video( request={ "features": [videointelligence.Feature.PERSON_DETECTION], "input_content": input_content, "video_context": context, } ) print("\n正在处理视频人体检测标注...") result = operation.result(timeout=300) print("\n处理完成。\n") # 获取首个结果(仅处理单个视频) annotation_result = result.annotation_results[0] for annotation in annotation_result.person_detection_annotations: print("检测到人体:") for track in annotation.tracks: print( "时间片段:{}秒 至 {}秒".format( track.segment.start_time_offset.seconds + track.segment.start_time_offset.microseconds / 1e6, track.segment.end_time_offset.seconds + track.segment.end_time_offset.microseconds / 1e6, ) ) # 每个时间片段包含带时间戳的目标,包含人物的特征(如服装、姿态等) # 获取第一个带时间戳的目标 timestamped_object = track.timestamped_objects[0] box = timestamped_object.normalized_bounding_box print(" bounding box:") print("\t左边界 : {}".format(box.left)) print("\t上边界 : {}".format(box.top)) print("\t右边界 : {}".format(box.right)) print("\t下边界 : {}".format(box.bottom)) # 属性包括服装、姿态、发色等信息 print(" 属性:") for attribute in timestamped_object.attributes: print( "\t{}: {} 置信度{}".format( attribute.name, attribute.value, attribute.confidence ) ) # 人体关键点包括左肩、右耳、右脚踝等身体部位 print(" 关键点:") for landmark in timestamped_object.landmarks: print( "\t{}: 置信度{} (x={}, y={})".format( landmark.name, landmark.confidence, landmark.point.x, # 归一化坐标x landmark.point.y, # 归一化坐标y ) )
PersonDetectionConfig 属性说明
""" 属性说明: include_bounding_boxes (bool): 是否在人体检测标注结果中包含bounding box。 include_pose_landmarks (bool): 是否开启人体关键点检测。若`include_bounding_boxes`设为false,则此参数无效。 include_attributes (bool): 是否开启人体属性检测,如服装颜色(黑、蓝等)、类型(外套、连衣裙等)、图案(纯色、碎花等)、发色等。若`include_bounding_boxes`设为false,则此参数无效。 """
示例输出
关键点: 2.5025秒 nose: 0.7880418300628662 (x=0.10155333578586578, y=0.29470884799957275) left_eye: 0.8498712182044983 (x=0.10444415360689163, y=0.2844345271587372) right_eye: 0.05180135369300842 (x=0.1029987558722496, y=0.28700312972068787) left_ear: 0.8830078840255737 (x=0.11745283752679825, y=0.28700312972068787) right_ear: 0.03634342923760414 (x=0.12757070362567902, y=0.28957170248031616) left_shoulder: 0.8504171371459961 (x=0.1145620197057724, y=0.34094318747520447) right_shoulder: 0.7254488468170166 (x=0.14925184845924377, y=0.34351176023483276) left_elbow: 0.7874324321746826 (x=0.10444415360689163, y=0.4205690324306488) right_elbow: 0.873414158821106 (x=0.18249624967575073, y=0.4154318571090698) left_wrist: 0.8134297132492065 (x=0.07987220585346222, y=0.44882336258888245) right_wrist: 0.5310596227645874 (x=0.1579243242740631, y=0.4693719446659088) left_hip: 0.48307809233665466 (x=0.12901613116264343, y=0.5130376815795898) right_hip: 0.4054966866970062 (x=0.14347021281719208, y=0.507900595664978) left_knee: 0.707266092300415 (x=0.14057938754558563, y=0.6363293528556824) right_knee: 0.536503791809082 (x=0.1029987558722496, y=0.615780770778656) left_ankle: 0.6422659158706665 (x=0.21429529786109924, y=0.6620150804519653) right_ankle: 0.7963647842407227 (x=0.08276302367448807, y=0.7467780113220215) 关键点: 2.6026秒 nose: 0.6816550493240356 (x=0.030738139525055885, y=0.3216055631637573) left_eye: 0.7059812545776367 (x=0.03222046047449112, y=0.3110688030719757) right_eye: 0.038844481110572815 (x=0.030738139525055885, y=0.3110688030719757) left_ear: 0.7951924800872803 (x=0.045561421662569046, y=0.3110688030719757) right_ear: 0.03984677046537399 (x=0.05742005258798599, y=0.3110688030719757) left_shoulder: 0.812483012676239 (x=0.047043751925230026, y=0.36638668179512024) right_shoulder: 0.7965729236602783 (x=0.08113731443881989, y=0.36375248432159424) left_elbow: 0.614456295967102 (x=0.05000840872526169, y=0.45331475138664246) right_elbow: 0.589381992816925 (x=0.10040760785341263, y=0.45331475138664246) left_wrist: 0.6238939762115479 (x=0.029255803674459457, y=0.49809587001800537) right_wrist: 0.31396669149398804 (x=0.0796549841761589, y=0.49282753467559814) left_hip: 0.3904641270637512 (x=0.05890238285064697, y=0.5376086235046387) right_hip: 0.3786430358886719 (x=0.0796549841761589, y=0.5323402285575867) left_knee: 0.46311870217323303 (x=0.035185132175683975, y=0.648244321346283) right_knee: 0.33408626914024353 (x=0.0411144383251667, y=0.6429758667945862)
内容的提问来源于stack exchange,提问作者Martin Brisiak
相关产品推荐
相关产品推荐

