You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

OpenAI Jukebox音频加载时Audio frame无法转ndarray报错

问题根因

报错由PyAV版本API变更与Jukebox旧代码不兼容导致:

  • 旧版PyAV中AudioResampler.resample()直接返回单个处理后的AudioFrame对象,可直接调用to_ndarray()方法
  • 当前Colab环境预装的是新版PyAV,resample()返回值为AudioFrame对象构成的列表,你打印frame得到的带方括号的输出就是直接证据。列表本身不存在to_ndarray()方法,因此触发第一个AttributeError。直接将返回值当作单帧对象传入后续逻辑,就会触发后续类型错误。
修复方法

将原load_audio函数中的解码循环段替换为兼容新版PyAV的实现即可,核心是遍历重采样返回的帧列表逐帧处理,同时补全重采样器缓存帧的读取逻辑,避免音频末尾截断:

原待替换代码段:

for frame in container.decode(audio=0): # Only first audio stream
    if resample:
        frame.pts = None
        frame = resampler.resample(frame)
    frame = frame.to_ndarray(format='fltp') # Convert to floats and not int16
    read = frame.shape[-1]
    if total_read + read > duration:
        read = duration - total_read
    sig[:, total_read:total_read + read] = frame[:, :read]
    total_read += read
    if total_read == duration:
        break

替换为以下代码:

for frame in container.decode(audio=0): # Only first audio stream
    if resample:
        frame.pts = None
        resampled_frames = resampler.resample(frame)
        for resampled_frame in resampled_frames:
            frame_nd = resampled_frame.to_ndarray(format='fltp')
            read = frame_nd.shape[-1]
            if total_read + read > duration:
                read = duration - total_read
            sig[:, total_read:total_read + read] = frame_nd[:, :read]
            total_read += read
            if total_read == duration:
                break
    else:
        frame_nd = frame.to_ndarray(format='fltp')
        read = frame_nd.shape[-1]
        if total_read + read > duration:
            read = duration - total_read
        sig[:, total_read:total_read + read] = frame_nd[:, :read]
        total_read += read
    if total_read == duration:
        break
# 读取重采样器缓存中剩余的帧,避免末尾音频丢失
if resample and total_read < duration:
    resampled_frames = resampler.resample(None)
    for resampled_frame in resampled_frames:
        frame_nd = resampled_frame.to_ndarray(format='fltp')
        read = frame_nd.shape[-1]
        if total_read + read > duration:
            read = duration - total_read
        sig[:, total_read:total_read + read] = frame_nd[:, :read]
        total_read += read
        if total_read == duration:
            break
注意事项
  • 无需手动降级PyAV版本,新版PyAV解码性能更优,降级易触发其他依赖冲突
  • 新增的重采样器缓存读取逻辑为新版PyAV必需,缺失会导致音频读取长度不足,触发末尾的长度断言报错

内容的提问来源于stack exchange,提问作者walker_4

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 12:39:24