You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

16:9转9:16视频ROI跟踪画面卡顿问题的解决方法问询

解决视频16:9转9:16时的帧间卡顿问题

你的核心问题是裁剪框更新间隔太大且无过渡,导致画面突然跳变。以下是针对性的优化方案和修改后的代码:

优化思路

  1. 逐帧插值过渡:在两次YOLO检测的帧之间,对前后两个裁剪框做线性插值,让画面平滑移动,而非突然切换
  2. 缩短检测间隔:把120帧的间隔降到30帧(1秒),既保证检测频率,又不会过度消耗性能(YOLOv8n速度足够快)
  3. 边界安全校验:确保裁剪区域不会超出原视频的画面范围,避免报错
  4. 持续平滑更新:将每次检测后的平滑框作为后续插值的基准,而非只保留单次检测结果

修改后的完整代码

from moviepy.editor import VideoFileClip
from ultralytics import YOLO
import numpy as np
import cv2

model = YOLO("yolov8n.pt")
clip = VideoFileClip("Mack Falls Off Cliff.mp4")

# 全局变量:记录过渡相关信息
prev_bbox = None          # 上一帧使用的裁剪框
target_bbox = None        # 目标裁剪框(最新检测结果)
last_detect_frame = 0     # 上一次检测的帧号
detect_interval = 30      # 检测间隔改为30帧(1秒)


def apply_mask(frame, bbox):
    height, width, _ = frame.shape
    x1, y1, x2, y2 = [int(val) for val in bbox]

    # 边界安全校验:确保裁剪区域不超出原视频范围
    x1 = max(0, x1)
    y1 = max(0, y1)
    x2 = min(width, x2)
    y2 = min(height, y2)

    bbox_width = x2 - x1
    bbox_height = y2 - y1
    bbox_aspect_ratio = bbox_width / bbox_height

    # 根据9:16比例计算最终裁剪区域
    if bbox_aspect_ratio > 9 / 16:
        new_width = int(bbox_height * (9 / 16))
        # 确保裁剪后中心对齐原bbox
        center_x = (x1 + x2) // 2
        x1 = center_x - new_width // 2
        x2 = center_x + new_width // 2
        # 再次校验边界
        x1 = max(0, x1)
        x2 = min(width, x2)
    else:
        new_height = int(bbox_width * (16 / 9))
        center_y = (y1 + y2) // 2
        y1 = center_y - new_height // 2
        y2 = center_y + new_height // 2
        # 再次校验边界
        y1 = max(0, y1)
        y2 = min(height, y2)

    cropped_frame = frame[y1:y2, x1:x2]
    masked_frame = cv2.resize(cropped_frame, (720, 1280))
    return masked_frame


def smooth_bbox(prev_bbox, curr_bbox, smoothing_factor=0.8):
    """对两次检测的bbox做平滑处理,避免突变"""
    return [
        int(prev_val * smoothing_factor + curr_val * (1 - smoothing_factor))
        for prev_val, curr_val in zip(prev_bbox, curr_bbox)
    ]


def process_frame(frame):
    global prev_bbox, target_bbox, last_detect_frame

    frame_count = int(process_frame.counter)
    process_frame.counter += 1

    # 到达检测间隔,执行YOLO检测
    if frame_count - last_detect_frame >= detect_interval:
        results = model(frame)
        bboxes = results[0].boxes.xyxy.cpu().numpy()

        if len(bboxes) > 0:
            # 取面积最大的目标作为ROI
            curr_bbox = max(bboxes, key=lambda x: (x[2]-x[0])*(x[3]-x[1]))
            if prev_bbox is not None:
                # 对新检测的bbox做平滑,避免突然跳变
                target_bbox = smooth_bbox(prev_bbox, curr_bbox)
            else:
                target_bbox = curr_bbox.tolist()
            last_detect_frame = frame_count
        # 如果没有检测到目标,保持原裁剪框
        else:
            target_bbox = prev_bbox

    # 计算当前帧的插值进度
    if target_bbox is not None and prev_bbox is not None:
        # 计算当前帧在两次检测之间的比例
        progress = (frame_count - last_detect_frame) / detect_interval
        # 线性插值得到当前帧的bbox
        current_bbox = [
            int(prev + (target - prev) * progress)
            for prev, target in zip(prev_bbox, target_bbox)
        ]
        masked_frame = apply_mask(frame, current_bbox)
        # 更新prev_bbox为当前使用的框,持续平滑
        prev_bbox = current_bbox
    elif prev_bbox is not None:
        masked_frame = apply_mask(frame, prev_bbox)
    else:
        # 初始帧没有检测结果,直接按中心裁剪9:16区域
        height, width = frame.shape[:2]
        new_width = int(height * 9/16)
        x1 = (width - new_width) // 2
        x2 = x1 + new_width
        masked_frame = cv2.resize(frame[:, x1:x2], (720, 1280))

    return masked_frame


# 初始化帧计数器
process_frame.counter = 0

processed_clip = clip.fl_image(process_frame)
processed_clip.write_videofile("output_smooth.mp4", codec="libx264")

关键改动说明

  1. 帧间插值逻辑:在两次检测之间,根据当前帧的位置计算progress,对prev_bbox和target_bbox做线性插值,实现平滑过渡
  2. 检测间隔调整:从120帧改为30帧,让裁剪框更新更频繁,跳变幅度更小
  3. 边界校验:在apply_mask中加入多次边界检查,避免裁剪区域超出原视频范围导致的错误
  4. 初始帧处理:当还没有检测结果时,直接按原视频中心裁剪9:16区域,避免画面异常
  5. 持续平滑更新:每次插值后的框都会更新prev_bbox,让整个过渡过程更自然

内容的提问来源于stack exchange,提问作者Michael Abrams

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 13:43:20