You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Ubuntu 22.04下如何用PyAV捕获音频配合aiortc流传输?

解决Ubuntu 22.04下PyAV捕获音频配合aiortc的问题

问题根源

av.open(format='alsa', file='default')报错是因为PyAV的av.open默认针对容器格式(如MP4、MKV),而alsa/pulse是音频设备的输入格式,并非容器,无法直接按容器方式调用。另外需要确保系统FFmpeg带有对应音频设备的支持,且PyAV编译时关联了这些组件。

解决步骤

1. 安装系统依赖(关键)

Ubuntu 22.04下先安装FFmpeg的音频设备支持包,确保系统能识别alsa/pulse设备:

sudo apt update
sudo apt install ffmpeg libasound2-dev libpulse-dev

2. 重新安装PyAV

PyAV基于系统FFmpeg编译,安装完依赖后需要重新安装PyAV,确保它能调用alsa/pulse设备:

pip uninstall -y av
pip install av

3. 正确的音频捕获代码

使用PyAV打开音频设备时,需要指定mode='r'(只读模式),并明确设备格式和参数,示例如下:

import av

# 用ALSA捕获音频
container = av.open(
    file='default',
    format='alsa',
    mode='r',
    options={
        'sample_rate': '48000',
        'channels': '2'
    }
)

# 读取音频帧(后续可处理后传给aiortc)
for frame in container.decode(audio=0):
    print(f"捕获到音频帧:{frame.sample_rate}Hz, {frame.channels}通道")

如果用PulseAudio,只需把format='alsa'改成format='pulse':

container = av.open(
    file='default',
    format='pulse',
    mode='r',
    options={
        'sample_rate': '48000',
        'channels': '2'
    }
)

4. 配合aiortc的处理

要把PyAV捕获的音频帧转换成aiortc能识别的MediaStreamTrack,可以自定义Track类:

from aiortc import MediaStreamTrack
import av

class AudioCaptureTrack(MediaStreamTrack):
    kind = "audio"

    def __init__(self):
        super().__init__()
        self.container = av.open(
            file='default',
            format='alsa',
            mode='r',
            options={'sample_rate': '48000', 'channels': '2'}
        )
        self.audio_stream = next(s for s in self.container.streams if s.type == 'audio')

    async def recv(self):
        for frame in self.container.decode(self.audio_stream):
            frame.pts = None
            return frame

替代方案:使用sounddevice库

如果PyAV仍无法识别设备,可以用sounddevice捕获原始音频,再封装成PyAV帧:

pip install sounddevice soundfile
import sounddevice as sd
import av
from aiortc import MediaStreamTrack

class SounddeviceAudioTrack(MediaStreamTrack):
    kind = "audio"

    def __init__(self):
        super().__init__()
        self.sample_rate = 48000
        self.channels = 2
        self.stream = sd.InputStream(
            samplerate=self.sample_rate,
            channels=self.channels,
            dtype='int16'
        )
        self.stream.start()

    async def recv(self):
        data, _ = self.stream.read(1024)
        # 封装成PyAV的AudioFrame
        frame = av.AudioFrame.from_ndarray(data, format='s16')
        frame.sample_rate = self.sample_rate
        frame.channels = self.channels
        frame.pts = None
        return frame

内容的提问来源于stack exchange,提问作者AndreyPr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 16:04:53