OpenAI Jukebox音频加载时Audio frame无法转ndarray报错
问题根因
报错由PyAV版本API变更与Jukebox旧代码不兼容导致:
- 旧版PyAV中
AudioResampler.resample()直接返回单个处理后的AudioFrame对象,可直接调用to_ndarray()方法 - 当前Colab环境预装的是新版PyAV,
resample()返回值为AudioFrame对象构成的列表,你打印frame得到的带方括号的输出就是直接证据。列表本身不存在to_ndarray()方法,因此触发第一个AttributeError。直接将返回值当作单帧对象传入后续逻辑,就会触发后续类型错误。
修复方法
将原load_audio函数中的解码循环段替换为兼容新版PyAV的实现即可,核心是遍历重采样返回的帧列表逐帧处理,同时补全重采样器缓存帧的读取逻辑,避免音频末尾截断:
原待替换代码段:
for frame in container.decode(audio=0): # Only first audio stream if resample: frame.pts = None frame = resampler.resample(frame) frame = frame.to_ndarray(format='fltp') # Convert to floats and not int16 read = frame.shape[-1] if total_read + read > duration: read = duration - total_read sig[:, total_read:total_read + read] = frame[:, :read] total_read += read if total_read == duration: break
替换为以下代码:
for frame in container.decode(audio=0): # Only first audio stream if resample: frame.pts = None resampled_frames = resampler.resample(frame) for resampled_frame in resampled_frames: frame_nd = resampled_frame.to_ndarray(format='fltp') read = frame_nd.shape[-1] if total_read + read > duration: read = duration - total_read sig[:, total_read:total_read + read] = frame_nd[:, :read] total_read += read if total_read == duration: break else: frame_nd = frame.to_ndarray(format='fltp') read = frame_nd.shape[-1] if total_read + read > duration: read = duration - total_read sig[:, total_read:total_read + read] = frame_nd[:, :read] total_read += read if total_read == duration: break # 读取重采样器缓存中剩余的帧,避免末尾音频丢失 if resample and total_read < duration: resampled_frames = resampler.resample(None) for resampled_frame in resampled_frames: frame_nd = resampled_frame.to_ndarray(format='fltp') read = frame_nd.shape[-1] if total_read + read > duration: read = duration - total_read sig[:, total_read:total_read + read] = frame_nd[:, :read] total_read += read if total_read == duration: break
注意事项
- 无需手动降级PyAV版本,新版PyAV解码性能更优,降级易触发其他依赖冲突
- 新增的重采样器缓存读取逻辑为新版PyAV必需,缺失会导致音频读取长度不足,触发末尾的长度断言报错
内容的提问来源于stack exchange,提问作者walker_4
相关产品推荐
相关产品推荐

