librosa加载CREMA数据集wav报未知格式及NoBackendError错误
问题场景
对CREMA、RAVDESS、TESS、SAVEE四个语音情感数据集标注情绪标签,将每个音频的存储路径、对应情绪标签整合为DataFrame后导出为csv文件。后续从csv读取音频路径调用librosa.load()接口加载音频时,仅CREMA数据集的wav文件加载失败,其余三个数据集的wav文件均可正常打开。已核查确认文件路径配置正确,当前使用的librosa版本为0.9.1。
报错信息
--------------------------------------------------------------------------- RuntimeError Traceback (most recent call last) File ~\.conda\envs\nhashemi\lib\site-packages\librosa\core\audio.py:155, in load(path, sr, mono, offset, duration, dtype, res_type) 153 else: 154 # Otherwise, create the soundfile object --> 155 context = sf.SoundFile(path) 157 with context as sf_desc: File ~\.conda\envs\nhashemi\lib\site-packages\soundfile.py:629, in SoundFile.__init__(self, file, mode, samplerate, channels, subtype, endian, format, closefd) 627 self._info = _create_info_struct(file, mode, samplerate, channels, 628 format, subtype, endian) --> 629 self._file = self._open(file, mode_int, closefd) 630 if set(mode).issuperset('r+') and self.seekable(): 631 # Move write position to 0 (like in Python file objects) File ~\.conda\envs\nhashemi\lib\site-packages\soundfile.py:1183, in SoundFile._open(self, file, mode_int, closefd) 1182 raise TypeError("Invalid file: {0!r}".format(self.name)) -> 1183 _error_check(_snd.sf_error(file_ptr), 1184 "Error opening {0!r}: ".format(self.name)) 1185 if mode_int == _snd.SFM_WRITE: 1186 # Due to a bug in libsndfile version <= 1.0.25, frames != 0 1187 # when opening a named pipe in SFM_WRITE mode. 1188 # See http://github.com/erikd/libsndfile/issues/77. File ~\.conda\envs\nhashemi\lib\site-packages\soundfile.py:1357, in _error_check(err, prefix) 1356 err_str = _snd.sf_error_number(err) -> 1357 raise RuntimeError(prefix + _ffi.string(err_str).decode('utf-8', 'replace')) RuntimeError: Error opening 'C:/Users/external_dipf/Documents/Dataset/CREMA/AudioWAV/1001_IEO_FEA_HI.wav': File contains data in an unknown format. During handling of the above exception, another exception occurred: NoBackendError Traceback (most recent call last) Input In [553], in <cell line: 3>() 1 emotion='fear' 2 path = np.array(data_path.Path[data_path.Emotions==emotion])[1] ----> 3 data, sampling_rate = librosa.load(path) 4 create_waveplot(data, sampling_rate, emotion) 5 create_spectrogram(data, sampling_rate, emotion) File ~\.conda\envs\nhashemi\lib\site-packages\librosa\util\decorators.py:88, in deprecate_positional_args.<locals>._inner_deprecate_positional_args.<locals>.inner_f(*args, **kwargs) 86 extra_args = len(args) - len(all_args) 87 if extra_args <= 0: ---> 88 return f(*args, **kwargs) 90 # extra_args > 0 91 args_msg = [ 92 "{}={}".format(name, arg) 93 for name, arg in zip(kwonly_args[:extra_args], args[-extra_args:]) 94 ] File ~\.conda\envs\nhashemi\lib\site-packages\librosa\core\audio.py:174, in load(path, sr, mono, offset, duration, dtype, res_type) 172 if isinstance(path, (str, pathlib.PurePath)): 173 warnings.warn("PySoundFile failed. Trying audioread instead.", stacklevel=2) --> 174 y, sr_native = __audioread_load(path, offset, duration, dtype) 175 else: 176 raise (exc) File ~\.conda\envs\nhashemi\lib\site-packages\librosa\core\audio.py:198, in __audioread_load(path, offset, duration, dtype) 192 """Load an audio buffer using audioread. 193 194 This loads one block at a time, and then concatenates the results. 195 """ 197 y = [] --> 198 with audioread.audio_open(path) as input_file: 199 sr_native = input_file.samplerate 200 n_channels = input_file.channels File ~\.conda\envs\nhashemi\lib\site-packages\audioread\__init__.py:116, in audio_open(path, backends) 113 pass 115 # All backends failed! --> 116 raise NoBackendError() NoBackendError:
现有实现代码
音频情绪标签标注代码
CREMA ="C:/Users/external_dipf/Documents/Dataset/CREMA/AudioWAV/" crema_directory_list = os.listdir(CREMA) file_emotion = [] file_path = [] for file in crema_directory_list: # storing file paths file_path.append(CREMA + file) # storing file emotions part=file.split('_') if part[2] == 'SAD': file_emotion.append('sad') elif part[2] == 'ANG': file_emotion.append('angry') elif part[2] == 'DIS': file_emotion.append('disgust') elif part[2] == 'FEA': file_emotion.append('fear') elif part[2] == 'HAP': file_emotion.append('happy') elif part[2] == 'NEU': file_emotion.append('neutral') else: file_emotion.append('Unknown') # dataframe for emotion of files emotion_df = pd.DataFrame(file_emotion, columns=['Emotions']) # dataframe for path of files. path_df = pd.DataFrame(file_path, columns=['Path']) CREMA_df = pd.concat([emotion_df, path_df], axis=1) CREMA_df.head()
数据集整合与csv存储代码
data_path = pd.concat([CREMA_df, RAVDESS_df, TESS_df, SAVEE_df], axis = 0) data_path.to_csv("data_path.csv",index=False) data_path.head()
音频加载与可视化代码
emotion='fear' path = np.array(data_path.Path[data_path.Emotions==emotion])[1] data, sampling_rate = librosa.load(path) create_waveplot(data, sampling_rate, emotion) create_spectrogram(data, sampling_rate, emotion) Audio(path)
故障原因与解决方案
报错核心原因有两个:
- CREMA数据集原生WAV文件采用WAVEFORMATEXTENSIBLE头封装,旧版本libsndfile(soundfile依赖的底层解码库)无法识别该格式,因此抛出「未知数据格式」错误
- soundfile解码失败后librosa会自动降级调用audioread解码,但当前环境缺少ffmpeg等可用解码后端,因此抛出NoBackendError错误
可按优先级选择以下方案修复:
- 方案1(零代码修改,最快生效):在当前conda环境安装ffmpeg,为audioread提供解码后端,执行命令
conda install -c conda-forge ffmpeg,安装完成后重启Python解释器即可正常加载音频,不需要修改原有路径或业务代码。 - 方案2(彻底解决格式兼容问题):批量将CREMA数据集音频转码为标准16bit PCM WAV格式,转码后原有路径无需改动,所有音频库均可正常读取。转码参考代码:
from pydub import AudioSegment import os from tqdm import tqdm CREMA_PATH = "C:/Users/external_dipf/Documents/Dataset/CREMA/AudioWAV/" for file in tqdm(os.listdir(CREMA_PATH)): if file.endswith(".wav"): full_path = os.path.join(CREMA_PATH, file) audio = AudioSegment.from_file(full_path) audio.export(full_path, format="wav", parameters=["-acodec", "pcm_s16le"])
- 前置校验:如果单个文件用系统自带播放器也无法打开,说明是压缩包解压不完整导致文件头损坏,重新解压CREMA数据集压缩包或重新下载损坏文件即可。
内容的提问来源于stack exchange,提问作者Niboo
相关产品推荐
相关产品推荐

