You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用LiveKit Python SDK推流时如何同步音视频(无需单独音轨)

如何用LiveKit Python SDK从MP4同时推流音视频(避免音视频不同步)

要实现从MP4文件同步推送音视频,不能仅依赖OpenCV(它只能处理视频流),需要用PyAV(基于FFmpeg的音视频处理库)来同时解析音视频帧,并通过时间戳对齐来保证同步。LiveKit会根据帧的时间戳自动处理音视频同步,无需额外操作。

步骤1:安装依赖

pip install av livekit livekit-api opencv-python

步骤2:完整实现代码

import asyncio
import av
import cv2
import livekit as lk
import livekit.rtc as rtc

WIDTH = 640
HEIGHT = 480

async def main(room: rtc.Room):
    # 连接房间逻辑(此处省略)
    
    # 创建音视频源和Track
    video_source = rtc.VideoSource(WIDTH, HEIGHT)
    video_track = rtc.LocalVideoTrack.create_video_track("video-track", video_source)
    
    audio_source = rtc.AudioSource()
    audio_track = rtc.LocalAudioTrack.create_audio_track("audio-track", audio_source)
    
    # 发布Track
    await room.local_participant.publish_track(video_track, rtc.TrackPublishOptions(source=rtc.TrackSource.SOURCE_CAMERA))
    await room.local_participant.publish_track(audio_track, rtc.TrackPublishOptions(source=rtc.TrackSource.SOURCE_MICROPHONE))
    
    # 同步推送音视频
    await stream_mp4(video_source, audio_source, "video.mp4")

async def stream_mp4(video_source: rtc.VideoSource, audio_source: rtc.AudioSource, video_path: str):
    container = av.open(video_path)
    video_stream = container.streams.video[0]
    audio_stream = container.streams.audio[0]
    
    # 音频参数适配LiveKit要求(PCM 16位,双声道,采样率48000)
    audio_resampler = av.AudioResampler(
        format="s16",
        layout="stereo",
        rate=48000,
    )
    
    last_pts = 0
    for packet in container.demux((video_stream, audio_stream)):
        for frame in packet.decode():
            if isinstance(frame, av.VideoFrame):
                # 处理视频帧:转换为RGBA格式,调整尺寸
                frame = frame.reformat(width=WIDTH, height=HEIGHT, format="rgba")
                frame_data = frame.to_ndarray().tobytes()
                
                # 创建LiveKit VideoFrame,设置正确的时间戳(微秒)
                lk_frame = rtc.VideoFrame(WIDTH, HEIGHT, rtc.VideoBufferType.RGBA, frame_data)
                lk_frame.timestamp = int(frame.pts * video_stream.time_base * 1e6)
                video_source.capture_frame(lk_frame)
                
                # 根据帧间隔计算sleep时间,保证播放速度
                if last_pts != 0:
                    delta = (frame.pts - last_pts) * video_stream.time_base
                    await asyncio.sleep(delta)
                last_pts = frame.pts
                
            elif isinstance(frame, av.AudioFrame):
                # 重采样音频到LiveKit兼容格式
                resampled_frames = audio_resampler.resample(frame)
                for resampled_frame in resampled_frames:
                    # 转换为PCM字节数据
                    pcm_data = resampled_frame.to_ndarray().tobytes()
                    
                    # 创建LiveKit AudioFrame,设置时间戳
                    lk_audio_frame = rtc.AudioFrame(
                        sample_rate=48000,
                        num_channels=2,
                        samples_per_channel=resampled_frame.samples,
                        data=pcm_data
                    )
                    lk_audio_frame.timestamp = int(frame.pts * audio_stream.time_base * 1e6)
                    audio_source.capture_frame(lk_audio_frame)
    
    container.close()

# 房间连接逻辑示例(可根据实际情况调整)
async def run():
    token = "YOUR_TOKEN"
    url = "YOUR_LIVEKIT_URL"
    
    room = rtc.Room()
    await room.connect(url, token)
    await main(room)
    await asyncio.Future()  # 保持连接

if __name__ == "__main__":
    asyncio.run(run())

关键说明

  • 同步原理:通过PyAV获取帧的原始时间戳(pts),转换为LiveKit要求的微秒级时间戳后推送,LiveKit会自动根据时间戳对齐音视频,避免不同步。
  • 音频处理:LiveKit要求音频为PCM 16位、采样率48000Hz,因此需要用PyAV的音频重采样器转换格式。
  • 帧率控制:根据视频帧的时间间隔计算sleep时长,替代固定的1/30,保证播放速度与原视频一致。

内容的提问来源于stack exchange,提问作者Sid Anand

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 05:35:17