You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用GCP Video Intelligence为视频每一帧标注姿态

GCP Video Intelligence 逐帧人体姿态标注解决方案

问题背景

GCP Video Intelligence框架可检测视频中的人体关键点,但默认仅返回间隔0.1秒的时间戳标注,无法直接生成视频每一帧的姿态数据。不想通过手动延长视频时长的方式获取逐帧结果,寻求合理替代方案。

可行解决方法

  • 插值补全帧数据
    利用API返回的0.1秒间隔关键点数据,通过线性插值或姿态平滑算法(如卡尔曼滤波)补全中间帧的关键点位置。无需重新调用API,成本低,适合对精度要求适中的场景。
    实现思路:遍历相邻两个时间戳的关键点集合,根据视频帧率计算中间需补全的帧数,对每个关键点的x、y坐标按时间比例做线性插值,同时对置信度做平滑处理。

  • 单帧调用Vision API分析
    将视频拆分为单帧图片,调用GCP Vision API的人体姿态检测功能对每一帧单独处理。这种方法能获取真正的逐帧精准标注,但需要额外处理视频拆帧和API批量调用,成本会相应增加。
    注意:Vision API与Video Intelligence的姿态检测模型可能存在差异,需验证结果一致性。

  • 向官方提交功能需求
    直接向GCP官方提交功能请求,建议在PersonDetectionConfig中添加自定义采样帧率的参数。若有足够多用户需求,官方可能在后续版本中支持该功能。


原示例代码(中文注释)

import io

from google.cloud import videointelligence_v1 as videointelligence


def detect_person(local_file_path="path/to/your/video-file.mp4"):
    """从本地文件检测视频中的人体"""

    client = videointelligence.VideoIntelligenceServiceClient()

    with io.open(local_file_path, "rb") as f:
        input_content = f.read()

    # 配置请求参数
    config = videointelligence.types.PersonDetectionConfig(
        include_bounding_boxes=True,
        include_attributes=True,
        include_pose_landmarks=True,
    )
    context = videointelligence.types.VideoContext(person_detection_config=config)

    # 发起异步请求
    operation = client.annotate_video(
        request={
            "features": [videointelligence.Feature.PERSON_DETECTION],
            "input_content": input_content,
            "video_context": context,
        }
    )

    print("\n正在处理视频人体检测标注...")
    result = operation.result(timeout=300)

    print("\n处理完成。\n")

    # 获取首个结果(仅处理单个视频)
    annotation_result = result.annotation_results[0]

    for annotation in annotation_result.person_detection_annotations:
        print("检测到人体:")
        for track in annotation.tracks:
            print(
                "时间片段:{}秒 至 {}秒".format(
                    track.segment.start_time_offset.seconds
                    + track.segment.start_time_offset.microseconds / 1e6,
                    track.segment.end_time_offset.seconds
                    + track.segment.end_time_offset.microseconds / 1e6,
                )
            )

            # 每个时间片段包含带时间戳的目标,包含人物的特征(如服装、姿态等)
            # 获取第一个带时间戳的目标
            timestamped_object = track.timestamped_objects[0]
            box = timestamped_object.normalized_bounding_box
            print(" bounding box:")
            print("\t左边界  : {}".format(box.left))
            print("\t上边界  : {}".format(box.top))
            print("\t右边界  : {}".format(box.right))
            print("\t下边界  : {}".format(box.bottom))

            # 属性包括服装、姿态、发色等信息
            print(" 属性:")
            for attribute in timestamped_object.attributes:
                print(
                    "\t{}: {} 置信度{}".format(
                        attribute.name, attribute.value, attribute.confidence
                    )
                )

            # 人体关键点包括左肩、右耳、右脚踝等身体部位
            print(" 关键点:")
            for landmark in timestamped_object.landmarks:
                print(
                    "\t{}: 置信度{} (x={}, y={})".format(
                        landmark.name,
                        landmark.confidence,
                        landmark.point.x,  # 归一化坐标x
                        landmark.point.y,  # 归一化坐标y
                    )
                )

PersonDetectionConfig 属性说明

"""
属性说明:
    include_bounding_boxes (bool):
        是否在人体检测标注结果中包含bounding box。
    include_pose_landmarks (bool):
        是否开启人体关键点检测。若`include_bounding_boxes`设为false,则此参数无效。
    include_attributes (bool):
        是否开启人体属性检测,如服装颜色(黑、蓝等)、类型(外套、连衣裙等)、图案(纯色、碎花等)、发色等。若`include_bounding_boxes`设为false,则此参数无效。
"""

示例输出

关键点: 2.5025秒
    nose: 0.7880418300628662 (x=0.10155333578586578, y=0.29470884799957275)
    left_eye: 0.8498712182044983 (x=0.10444415360689163, y=0.2844345271587372)
    right_eye: 0.05180135369300842 (x=0.1029987558722496, y=0.28700312972068787)
    left_ear: 0.8830078840255737 (x=0.11745283752679825, y=0.28700312972068787)
    right_ear: 0.03634342923760414 (x=0.12757070362567902, y=0.28957170248031616)
    left_shoulder: 0.8504171371459961 (x=0.1145620197057724, y=0.34094318747520447)
    right_shoulder: 0.7254488468170166 (x=0.14925184845924377, y=0.34351176023483276)
    left_elbow: 0.7874324321746826 (x=0.10444415360689163, y=0.4205690324306488)
    right_elbow: 0.873414158821106 (x=0.18249624967575073, y=0.4154318571090698)
    left_wrist: 0.8134297132492065 (x=0.07987220585346222, y=0.44882336258888245)
    right_wrist: 0.5310596227645874 (x=0.1579243242740631, y=0.4693719446659088)
    left_hip: 0.48307809233665466 (x=0.12901613116264343, y=0.5130376815795898)
    right_hip: 0.4054966866970062 (x=0.14347021281719208, y=0.507900595664978)
    left_knee: 0.707266092300415 (x=0.14057938754558563, y=0.6363293528556824)
    right_knee: 0.536503791809082 (x=0.1029987558722496, y=0.615780770778656)
    left_ankle: 0.6422659158706665 (x=0.21429529786109924, y=0.6620150804519653)
    right_ankle: 0.7963647842407227 (x=0.08276302367448807, y=0.7467780113220215)
关键点: 2.6026秒
    nose: 0.6816550493240356 (x=0.030738139525055885, y=0.3216055631637573)
    left_eye: 0.7059812545776367 (x=0.03222046047449112, y=0.3110688030719757)
    right_eye: 0.038844481110572815 (x=0.030738139525055885, y=0.3110688030719757)
    left_ear: 0.7951924800872803 (x=0.045561421662569046, y=0.3110688030719757)
    right_ear: 0.03984677046537399 (x=0.05742005258798599, y=0.3110688030719757)
    left_shoulder: 0.812483012676239 (x=0.047043751925230026, y=0.36638668179512024)
    right_shoulder: 0.7965729236602783 (x=0.08113731443881989, y=0.36375248432159424)
    left_elbow: 0.614456295967102 (x=0.05000840872526169, y=0.45331475138664246)
    right_elbow: 0.589381992816925 (x=0.10040760785341263, y=0.45331475138664246)
    left_wrist: 0.6238939762115479 (x=0.029255803674459457, y=0.49809587001800537)
    right_wrist: 0.31396669149398804 (x=0.0796549841761589, y=0.49282753467559814)
    left_hip: 0.3904641270637512 (x=0.05890238285064697, y=0.5376086235046387)
    right_hip: 0.3786430358886719 (x=0.0796549841761589, y=0.5323402285575867)
    left_knee: 0.46311870217323303 (x=0.035185132175683975, y=0.648244321346283)
    right_knee: 0.33408626914024353 (x=0.0411144383251667, y=0.6429758667945862)

内容的提问来源于stack exchange,提问作者Martin Brisiak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 20:15:47