使用RAVDESS数据集微调wav2vec模型时遇pyarrow.lib.ArrowInvalid错误
问题
使用HuggingFace的jonatasgrosman/wav2vec2-large-xlsr-53-english模型进行语音情感识别时,该模型在CREMA、TESS、SAVEE及自定义数据集上训练均正常,但使用RAVDESS数据集时触发pyarrow.lib.ArrowInvalid错误,错误提示如下:
Map: 0%| | 0/1152 [00:00<?, ? examples/s]C:\Users\XTEND\anaconda3\envs\pytorch_gpu\lib\site-packages\transformers\feature_extraction_utils.py:165: VisibleDeprecationWarning: Creating an ndarray from ragged nested sequences (which is a list-or-tuple of lists-or-tuples-or ndarrays with different lengths or shapes) is deprecated. If you meant to do this, you must specify 'dtype=object' when creating the ndarray. tensor = as_tensor(value) Traceback (most recent call last): File "C:\Users\XTEND\anaconda3\envs\pytorch_gpu\lib\site-packages\datasets\arrow_dataset.py", line 3004, in map for rank, done, content in Dataset._map_single(**dataset_kwargs): File "C:\Users\XTEND\anaconda3\envs\pytorch_gpu\lib\site-packages\datasets\arrow_dataset.py", line 3397, in _map_single writer.write_batch(batch) File "C:\Users\XTEND\anaconda3\envs\pytorch_gpu\lib\site-packages\datasets\arrow_writer.py", line 551, in write_batch arrays.append(pa.array(typed_sequence)) File "pyarrow\array.pxi", line 236, in pyarrow.lib.array File "pyarrow\array.pxi", line 110, in pyarrow.lib._handle_arrow_array_protocol File "C:\Users\XTEND\anaconda3\envs\pytorch_gpu\lib\site-packages\datasets\arrow_writer.py", line 186, in __arrow_array__ out = list_of_np_array_to_pyarrow_listarray(data) File "C:\Users\XTEND\anaconda3\envs\pytorch_gpu\lib\site-packages\datasets\features\features.py", line 1395, in list_of_np_array_to_pyarrow_listarray return list_of_pa_arrays_to_pyarrow_listarray( File "C:\Users\XTEND\anaconda3\envs\pytorch_gpu\lib\site-packages\datasets\features\features.py", line 1388, in list_of_pa_arrays_to_pyarrow_listarray values = pa.concat_arrays(l_arr) File "pyarrow\array.pxi", line 3039, in pyarrow.lib.concat_arrays File "pyarrow\error.pxi", line 144, in pyarrow.lib.pyarrow_internal_check_status File "pyarrow\error.pxi", line 100, in pyarrow.lib.check_status pyarrow.lib.ArrowInvalid: arrays to be concatenated must be identically typed, but float and list<item: float> were encountered.
相关处理代码:
# RAVDESS DATASET RAV = "D:/program/Audio_SA/Dataset/RAVDESS/" dir_list = os.listdir(RAV) print(dir_list.sort()) print(dir_list) emotion = [] name = [] path = [] for i in dir_list: fname = os.listdir(RAV + i) for f in fname: part = f.split('.')[0].split('-') emotion.append(int(part[2])) path.append(RAV + i + '/' + f) name.append(f) emotion_df = pd.DataFrame(emotion, columns=['Emotion']) emotion_df = emotion_df.replace( {1: 'neutral', 2: 'neutral', 3: 'happy', 4: 'sad', 5: 'angry', 6: 'fear', 7: 'disgust', 8: 'surprise'}) name_df = pd.DataFrame(name, columns=['Name']) RAV_df = pd.concat([name_df, pd.DataFrame(path, columns=['Path']), emotion_df], axis=1) print(RAV_df.head()) # shuffle the DataFrame rows df = RAV_df.sample(frac=1) df.to_csv('RAVDESS/Ravdess_df.csv', index=False) # Filter broken and non-existed paths print(f"Step 0: {len(df)}") df["status"] = df["path"].apply(lambda speech_path: True if os.path.exists(speech_path) else None) df = df.dropna(subset=["path"]) df = df.drop("status", axis=1) print(f"Step 1: {len(df)}") df = df.sample(frac=1) df = df.reset_index(drop=True) print("labels: ", df["emotion"].unique()) print() print(df.groupby("emotion").count()[["path"]]) idx = np.random.randint(0, len(df)) sample = df.iloc[idx] path = sample["path"] emotion = sample["emotion"] print(f"ID Location: {idx}") print(f" emotion: {emotion}") print() print(df.head()) save_path = r"C:\Users\XTEND\PycharmProjects\AER_ENGLISH" use_auth_token = True train_df, test_df = train_test_split(df, test_size=0.2, random_state=101, stratify=df["emotion"]) train_df = train_df.reset_index(drop=True) test_df = test_df.reset_index(drop=True) test_df.to_csv("test_df_new.csv", sep="\t", encoding="utf-8", index=False) train_df.to_csv("train_df_new.csv", sep="\t", encoding="utf-8", index=False) print(train_df.shape) print(test_df.shape) print(train_df) print(test_df) # Prepare Data for Training # Loading the created dataset using datasets data_files = {"train": "C:/Users/XTEND/PycharmProjects/custom_AER/RAVDESS/train_df.csv", "validation": "C:/Users/XTEND/PycharmProjects/custom_AER/RAVDESS/test_df.csv", } # data_files = {"train": R"C:\Users\XTEND\PycharmProjects\custom_AER\Main2_files\train_df.csv", # "validation": R"C:\Users\XTEND\PycharmProjects\custom_AER\Main2_files\test_df.csv", } datasets = load_dataset("csv", data_files=data_files, delimiter="\t", ) train_dataset = datasets["train"] eval_dataset = datasets["validation"] print(train_dataset) print(eval_dataset) # We need to specify the input and output column input_column = "Path" output_column = "Emotion" # we need to distinguish the unique labels in our SER dataset label_list = train_dataset.unique(output_column) label_list.sort() # Let's sort it for determinism num_labels = len(label_list) print(f'A classification problem with {num_labels} classes: {label_list}')
解决方案
1. 统一列名大小写
代码中创建DataFrame时使用大写列名Path、Emotion,但后续处理却用小写的path、emotion,导致CSV文件列名与后续引用不匹配,引发数据类型识别异常。
- 修复:全程统一列名大小写,示例改为小写:
# 修改DataFrame创建代码 emotion_df = pd.DataFrame(emotion, columns=['emotion']) path_df = pd.DataFrame(path, columns=['path']) name_df = pd.DataFrame(name, columns=['name']) RAV_df = pd.concat([name_df, path_df, emotion_df], axis=1) # 后续列引用统一用小写 df["status"] = df["path"].apply(lambda speech_path: True if os.path.exists(speech_path) else None) input_column = "path" output_column = "emotion"
2. 校验并修复RAVDESS音频文件
错误提示中的ragged nested sequences说明部分音频处理后特征形状不一致,RAVDESS可能存在损坏、长度异常或采样率不统一的音频。
- 处理:加载音频前加入校验逻辑:
import librosa def check_audio_validity(audio_path): try: y, sr = librosa.load(audio_path, sr=16000) # 统一采样率为模型要求的16kHz if len(y) == 0: return False return True except: return False # 过滤无效音频 df = df[df["path"].apply(check_audio_validity)]
3. 手动指定数据集特征类型
使用load_dataset时自动识别的列类型可能有误,导致后续map处理时数据类型混乱。
- 修复:手动定义特征类型:
from datasets import Features, Value features = Features({ "name": Value("string"), "path": Value("string"), "emotion": Value("string") }) datasets = load_dataset("csv", data_files=data_files, delimiter="\t", features=features)
4. 确保特征提取输出结构一致
错误发生在map阶段,说明特征提取函数返回的结果存在混合单个float和float列表的情况。
- 修复:在特征提取函数中统一输出结构,对所有音频特征做固定长度的截断或填充:
import numpy as np def extract_features(batch): audio, sr = librosa.load(batch["path"], sr=16000) # 统一截断/填充到固定长度(示例为10秒音频对应的采样点数) max_length = 160000 if len(audio) > max_length: audio = audio[:max_length] else: audio = np.pad(audio, (0, max_length - len(audio)), mode='constant') # 使用wav2vec2特征提取器处理 batch["input_values"] = feature_extractor(audio, sampling_rate=sr)["input_values"][0] return batch # 应用map时指定非批量处理,避免形状不匹配 train_dataset = train_dataset.map(extract_features, batched=False)
内容的提问来源于stack exchange,提问作者Sneha T S
相关产品推荐
相关产品推荐

