You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于Python的DearPyGUI实现音视频同步播放?

实现方向指引与具体落地步骤

一、OpenCV渲染视频 + miniaudio播放音频(新手友好方案)

步骤1:提取音频流并解码

用ffmpeg-python直接从视频中提取PCM格式的原始音频数据,同时记录音频参数(采样率、通道数、位深):

import ffmpeg

# 从视频文件提取16位双声道、44100采样率的PCM音频
audio_process = ffmpeg.input("video.mp4").audio.output(
    "-", format="s16le", ac=2, ar=44100
).run_async(pipe_stdout=True)
audio_params = {"channels": 2, "sample_rate": 44100}

步骤2:用miniaudio播放音频

miniaudio轻量且支持PCM流播放,还能实时获取播放进度,适合同步:

import miniaudio

# 初始化音频播放器,通过回调从管道取音频数据
def audio_callback(frame_count):
    # 计算需要读取的字节数:帧数量 × 通道数 × 2字节/采样(16位)
    return audio_process.stdout.read(frame_count * 2 * 2)

stream = miniaudio.stream_playback(
    sample_rate=audio_params["sample_rate"],
    channels=audio_params["channels"],
    sample_format=miniaudio.SampleFormat.SIGNED16,
    callback=audio_callback
)
stream.start()

步骤3:音视频同步逻辑

OpenCV读取视频帧时,获取帧的时间戳,和音频播放进度对比,通过等待或跳帧实现同步:

import cv2
import time

cap = cv2.VideoCapture("video.mp4")
fps = cap.get(cv2.CAP_PROP_FPS)
frame_interval = 1 / fps

while cap.isOpened():
    ret, frame = cap.read()
    if not ret:
        break
    
    # 获取当前视频帧的时间戳(毫秒)
    frame_ts = cap.get(cv2.CAP_PROP_POS_MSEC)
    # 获取音频当前播放的时间戳(毫秒)
    audio_ts = stream.get_current_frame() * 1000 / audio_params["sample_rate"]
    
    # 同步逻辑:允许50ms误差,避免频繁卡顿
    if frame_ts > audio_ts + 50:
        time.sleep((frame_ts - audio_ts) / 1000)
    elif audio_ts > frame_ts + 100:
        # 音频超前时,跳过对应数量的视频帧
        skip_frames = int((audio_ts - frame_ts) / (1000 / fps))
        for _ in range(skip_frames):
            cap.read()
    
    # 把OpenCV的BGR帧转成RGB(DearPyGUI纹理要求RGB格式)
    frame_rgb = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)
    # 更新DearPyGUI的raw texture(替换成你创建的texture_id)
    dpg.set_value(texture_id, frame_rgb.flatten())
    
    time.sleep(frame_interval)

二、解决ffpyplayer的蓝紫色 tint 和帧率问题

  • 蓝紫色问题:ffpyplayer返回的帧通道顺序是BGR,需要翻转成RGB再传给DearPyGUI
  • 帧率低问题:单独开线程读取帧并放到队列,避免主线程阻塞

调整后的核心代码:

from ffpyplayer.player import MediaPlayer
import numpy as np
import threading
import time

player = MediaPlayer("video.mp4")
frame_queue = []

# 后台线程读取视频帧
def read_ffpyplayer_frames():
    fps = player.get_meta_data()["fps"]
    while True:
        frame, val = player.get_frame()
        if val == "eof":
            break
        if frame is not None:
            img, ts = frame
            # 转成numpy数组并翻转通道解决蓝紫色
            img_np = np.asarray(img.to_bytearray()[0]).reshape(img.get_size()[1], img.get_size()[0], 3)
            frame_rgb = img_np[:, :, ::-1]
            frame_queue.append((frame_rgb, ts))
        time.sleep(1/fps)

threading.Thread(target=read_ffpyplayer_frames, daemon=True).start()

# DearPyGUI渲染循环
while dpg.is_dearpygui_running():
    if frame_queue:
        frame, ts = frame_queue.pop(0)
        # 同步音频:用player.get_pts()获取音频时间戳,和帧ts对比调整
        audio_ts = player.get_pts()
        if abs(audio_ts - ts) > 0.1:
            continue  # 跳过不同步的帧
        dpg.set_value(texture_id, frame.flatten())
    dpg.render_dearpygui_frame()

三、ffmpeg-python管道方案落地

通过ffmpeg同时输出视频帧和音频流的管道,用多线程分别处理,同步逻辑和方案一致:

import ffmpeg
import threading
import numpy as np

# 同时输出RGB视频帧和PCM音频的管道
process = ffmpeg.input("video.mp4").output(
    "pipe:1", format="rawvideo", pix_fmt="rgb24", s="1920x1080",
    "pipe:2", format="s16le", ac=2, ar=44100
).run_async(pipe_stdout=True, pipe_stderr=True)

frame_queue = []
audio_buffer = b""

# 视频线程读取帧
def read_video_pipe():
    frame_size = 1920 * 1080 * 3
    while True:
        frame_data = process.stdout.read(frame_size)
        if not frame_data:
            break
        frame = np.frombuffer(frame_data, dtype=np.uint8).reshape(1080, 1920, 3)
        frame_queue.append(frame)

# 音频线程读取数据
def read_audio_pipe():
    global audio_buffer
    while True:
        data = process.stderr.read(4096)
        if not data:
            break
        audio_buffer += data

threading.Thread(target=read_video_pipe, daemon=True).start()
threading.Thread(target=read_audio_pipe, daemon=True).start()

# 后续播放音频(用miniaudio从audio_buffer取数据)、渲染视频、同步逻辑参考方案一

关键注意事项

  • DearPyGUI的纹理更新必须在主线程执行,若在子线程操作,需用dpg.invoke_callback把更新逻辑切换到主线程
  • 同步时不要追求绝对精准,保留50-100ms的误差范围,避免频繁的等待或跳帧导致卡顿
  • 所有IO、解码操作尽量放到子线程,保证主线程专注于UI渲染,避免界面无响应

内容的提问来源于stack exchange,提问作者Vi Tiet

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 14:17:50