16:9转9:16视频ROI跟踪画面卡顿问题的解决方法问询
解决视频16:9转9:16时的帧间卡顿问题
你的核心问题是裁剪框更新间隔太大且无过渡,导致画面突然跳变。以下是针对性的优化方案和修改后的代码:
优化思路
- 逐帧插值过渡:在两次YOLO检测的帧之间,对前后两个裁剪框做线性插值,让画面平滑移动,而非突然切换
- 缩短检测间隔:把120帧的间隔降到30帧(1秒),既保证检测频率,又不会过度消耗性能(YOLOv8n速度足够快)
- 边界安全校验:确保裁剪区域不会超出原视频的画面范围,避免报错
- 持续平滑更新:将每次检测后的平滑框作为后续插值的基准,而非只保留单次检测结果
修改后的完整代码
from moviepy.editor import VideoFileClip from ultralytics import YOLO import numpy as np import cv2 model = YOLO("yolov8n.pt") clip = VideoFileClip("Mack Falls Off Cliff.mp4") # 全局变量:记录过渡相关信息 prev_bbox = None # 上一帧使用的裁剪框 target_bbox = None # 目标裁剪框(最新检测结果) last_detect_frame = 0 # 上一次检测的帧号 detect_interval = 30 # 检测间隔改为30帧(1秒) def apply_mask(frame, bbox): height, width, _ = frame.shape x1, y1, x2, y2 = [int(val) for val in bbox] # 边界安全校验:确保裁剪区域不超出原视频范围 x1 = max(0, x1) y1 = max(0, y1) x2 = min(width, x2) y2 = min(height, y2) bbox_width = x2 - x1 bbox_height = y2 - y1 bbox_aspect_ratio = bbox_width / bbox_height # 根据9:16比例计算最终裁剪区域 if bbox_aspect_ratio > 9 / 16: new_width = int(bbox_height * (9 / 16)) # 确保裁剪后中心对齐原bbox center_x = (x1 + x2) // 2 x1 = center_x - new_width // 2 x2 = center_x + new_width // 2 # 再次校验边界 x1 = max(0, x1) x2 = min(width, x2) else: new_height = int(bbox_width * (16 / 9)) center_y = (y1 + y2) // 2 y1 = center_y - new_height // 2 y2 = center_y + new_height // 2 # 再次校验边界 y1 = max(0, y1) y2 = min(height, y2) cropped_frame = frame[y1:y2, x1:x2] masked_frame = cv2.resize(cropped_frame, (720, 1280)) return masked_frame def smooth_bbox(prev_bbox, curr_bbox, smoothing_factor=0.8): """对两次检测的bbox做平滑处理,避免突变""" return [ int(prev_val * smoothing_factor + curr_val * (1 - smoothing_factor)) for prev_val, curr_val in zip(prev_bbox, curr_bbox) ] def process_frame(frame): global prev_bbox, target_bbox, last_detect_frame frame_count = int(process_frame.counter) process_frame.counter += 1 # 到达检测间隔,执行YOLO检测 if frame_count - last_detect_frame >= detect_interval: results = model(frame) bboxes = results[0].boxes.xyxy.cpu().numpy() if len(bboxes) > 0: # 取面积最大的目标作为ROI curr_bbox = max(bboxes, key=lambda x: (x[2]-x[0])*(x[3]-x[1])) if prev_bbox is not None: # 对新检测的bbox做平滑,避免突然跳变 target_bbox = smooth_bbox(prev_bbox, curr_bbox) else: target_bbox = curr_bbox.tolist() last_detect_frame = frame_count # 如果没有检测到目标,保持原裁剪框 else: target_bbox = prev_bbox # 计算当前帧的插值进度 if target_bbox is not None and prev_bbox is not None: # 计算当前帧在两次检测之间的比例 progress = (frame_count - last_detect_frame) / detect_interval # 线性插值得到当前帧的bbox current_bbox = [ int(prev + (target - prev) * progress) for prev, target in zip(prev_bbox, target_bbox) ] masked_frame = apply_mask(frame, current_bbox) # 更新prev_bbox为当前使用的框,持续平滑 prev_bbox = current_bbox elif prev_bbox is not None: masked_frame = apply_mask(frame, prev_bbox) else: # 初始帧没有检测结果,直接按中心裁剪9:16区域 height, width = frame.shape[:2] new_width = int(height * 9/16) x1 = (width - new_width) // 2 x2 = x1 + new_width masked_frame = cv2.resize(frame[:, x1:x2], (720, 1280)) return masked_frame # 初始化帧计数器 process_frame.counter = 0 processed_clip = clip.fl_image(process_frame) processed_clip.write_videofile("output_smooth.mp4", codec="libx264")
关键改动说明
- 帧间插值逻辑:在两次检测之间,根据当前帧的位置计算
progress,对prev_bbox和target_bbox做线性插值,实现平滑过渡 - 检测间隔调整:从120帧改为30帧,让裁剪框更新更频繁,跳变幅度更小
- 边界校验:在
apply_mask中加入多次边界检查,避免裁剪区域超出原视频范围导致的错误 - 初始帧处理:当还没有检测结果时,直接按原视频中心裁剪9:16区域,避免画面异常
- 持续平滑更新:每次插值后的框都会更新
prev_bbox,让整个过渡过程更自然
内容的提问来源于stack exchange,提问作者Michael Abrams
相关产品推荐
相关产品推荐

