You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用RAVDESS数据集微调wav2vec模型时遇pyarrow.lib.ArrowInvalid错误

问题

使用HuggingFace的jonatasgrosman/wav2vec2-large-xlsr-53-english模型进行语音情感识别时,该模型在CREMA、TESS、SAVEE及自定义数据集上训练均正常,但使用RAVDESS数据集时触发pyarrow.lib.ArrowInvalid错误,错误提示如下:

Map:   0%|          | 0/1152 [00:00<?, ? examples/s]C:\Users\XTEND\anaconda3\envs\pytorch_gpu\lib\site-packages\transformers\feature_extraction_utils.py:165: VisibleDeprecationWarning: Creating an ndarray from ragged nested sequences (which is a list-or-tuple of lists-or-tuples-or ndarrays with different lengths or shapes) is deprecated. If you meant to do this, you must specify 'dtype=object' when creating the ndarray.
  tensor = as_tensor(value)
Traceback (most recent call last):
  File "C:\Users\XTEND\anaconda3\envs\pytorch_gpu\lib\site-packages\datasets\arrow_dataset.py", line 3004, in map
    for rank, done, content in Dataset._map_single(**dataset_kwargs):
  File "C:\Users\XTEND\anaconda3\envs\pytorch_gpu\lib\site-packages\datasets\arrow_dataset.py", line 3397, in _map_single
    writer.write_batch(batch)
  File "C:\Users\XTEND\anaconda3\envs\pytorch_gpu\lib\site-packages\datasets\arrow_writer.py", line 551, in write_batch
    arrays.append(pa.array(typed_sequence))
  File "pyarrow\array.pxi", line 236, in pyarrow.lib.array
  File "pyarrow\array.pxi", line 110, in pyarrow.lib._handle_arrow_array_protocol
  File "C:\Users\XTEND\anaconda3\envs\pytorch_gpu\lib\site-packages\datasets\arrow_writer.py", line 186, in __arrow_array__
    out = list_of_np_array_to_pyarrow_listarray(data)
  File "C:\Users\XTEND\anaconda3\envs\pytorch_gpu\lib\site-packages\datasets\features\features.py", line 1395, in list_of_np_array_to_pyarrow_listarray
    return list_of_pa_arrays_to_pyarrow_listarray(
  File "C:\Users\XTEND\anaconda3\envs\pytorch_gpu\lib\site-packages\datasets\features\features.py", line 1388, in list_of_pa_arrays_to_pyarrow_listarray
    values = pa.concat_arrays(l_arr)
  File "pyarrow\array.pxi", line 3039, in pyarrow.lib.concat_arrays
  File "pyarrow\error.pxi", line 144, in pyarrow.lib.pyarrow_internal_check_status
  File "pyarrow\error.pxi", line 100, in pyarrow.lib.check_status
pyarrow.lib.ArrowInvalid: arrays to be concatenated must be identically typed, but float and list<item: float> were encountered.

相关处理代码:

# RAVDESS DATASET
RAV = "D:/program/Audio_SA/Dataset/RAVDESS/"
dir_list = os.listdir(RAV)
print(dir_list.sort())
print(dir_list)

emotion = []
name = []
path = []
for i in dir_list:
    fname = os.listdir(RAV + i)

    for f in fname:
        part = f.split('.')[0].split('-')
        emotion.append(int(part[2]))
        path.append(RAV + i + '/' + f)
        name.append(f)

emotion_df = pd.DataFrame(emotion, columns=['Emotion'])
emotion_df = emotion_df.replace(
    {1: 'neutral', 2: 'neutral', 3: 'happy', 4: 'sad', 5: 'angry', 6: 'fear', 7: 'disgust', 8: 'surprise'})
name_df = pd.DataFrame(name, columns=['Name'])
RAV_df = pd.concat([name_df, pd.DataFrame(path, columns=['Path']), emotion_df], axis=1)
print(RAV_df.head())

# shuffle the DataFrame rows
df = RAV_df.sample(frac=1)
df.to_csv('RAVDESS/Ravdess_df.csv', index=False)

# Filter broken and non-existed paths
print(f"Step 0: {len(df)}")
df["status"] = df["path"].apply(lambda speech_path: True if os.path.exists(speech_path) else None)
df = df.dropna(subset=["path"])
df = df.drop("status", axis=1)
print(f"Step 1: {len(df)}")

df = df.sample(frac=1)
df = df.reset_index(drop=True)

print("labels: ", df["emotion"].unique())
print()
print(df.groupby("emotion").count()[["path"]])

idx = np.random.randint(0, len(df))
sample = df.iloc[idx]
path = sample["path"]
emotion = sample["emotion"]

print(f"ID Location: {idx}")
print(f"      emotion: {emotion}")
print()
print(df.head())
save_path = r"C:\Users\XTEND\PycharmProjects\AER_ENGLISH"
use_auth_token = True
train_df, test_df = train_test_split(df, test_size=0.2, random_state=101, stratify=df["emotion"])

train_df = train_df.reset_index(drop=True)
test_df = test_df.reset_index(drop=True)

test_df.to_csv("test_df_new.csv", sep="\t", encoding="utf-8", index=False)
train_df.to_csv("train_df_new.csv", sep="\t", encoding="utf-8", index=False)

print(train_df.shape)
print(test_df.shape)
print(train_df)
print(test_df)


# Prepare Data for Training
# Loading the created dataset using datasets
data_files = {"train": "C:/Users/XTEND/PycharmProjects/custom_AER/RAVDESS/train_df.csv",
              "validation": "C:/Users/XTEND/PycharmProjects/custom_AER/RAVDESS/test_df.csv", }

# data_files = {"train": R"C:\Users\XTEND\PycharmProjects\custom_AER\Main2_files\train_df.csv",
#               "validation": R"C:\Users\XTEND\PycharmProjects\custom_AER\Main2_files\test_df.csv", }

datasets = load_dataset("csv", data_files=data_files, delimiter="\t", )
train_dataset = datasets["train"]
eval_dataset = datasets["validation"]

print(train_dataset)
print(eval_dataset)

# We need to specify the input and output column
input_column = "Path"
output_column = "Emotion"

# we need to distinguish the unique labels in our SER dataset
label_list = train_dataset.unique(output_column)
label_list.sort()  # Let's sort it for determinism
num_labels = len(label_list)
print(f'A classification problem with {num_labels} classes: {label_list}')

解决方案

1. 统一列名大小写

代码中创建DataFrame时使用大写列名Path、Emotion,但后续处理却用小写的path、emotion,导致CSV文件列名与后续引用不匹配,引发数据类型识别异常。

  • 修复:全程统一列名大小写,示例改为小写:
    # 修改DataFrame创建代码
    emotion_df = pd.DataFrame(emotion, columns=['emotion'])
    path_df = pd.DataFrame(path, columns=['path'])
    name_df = pd.DataFrame(name, columns=['name'])
    RAV_df = pd.concat([name_df, path_df, emotion_df], axis=1)
    
    # 后续列引用统一用小写
    df["status"] = df["path"].apply(lambda speech_path: True if os.path.exists(speech_path) else None)
    input_column = "path"
    output_column = "emotion"
    

2. 校验并修复RAVDESS音频文件

错误提示中的ragged nested sequences说明部分音频处理后特征形状不一致,RAVDESS可能存在损坏、长度异常或采样率不统一的音频。

  • 处理:加载音频前加入校验逻辑:
    import librosa
    
    def check_audio_validity(audio_path):
        try:
            y, sr = librosa.load(audio_path, sr=16000)  # 统一采样率为模型要求的16kHz
            if len(y) == 0:
                return False
            return True
        except:
            return False
    
    # 过滤无效音频
    df = df[df["path"].apply(check_audio_validity)]
    

3. 手动指定数据集特征类型

使用load_dataset时自动识别的列类型可能有误,导致后续map处理时数据类型混乱。

  • 修复:手动定义特征类型:
    from datasets import Features, Value
    
    features = Features({
        "name": Value("string"),
        "path": Value("string"),
        "emotion": Value("string")
    })
    
    datasets = load_dataset("csv", data_files=data_files, delimiter="\t", features=features)
    

4. 确保特征提取输出结构一致

错误发生在map阶段,说明特征提取函数返回的结果存在混合单个float和float列表的情况。

  • 修复:在特征提取函数中统一输出结构,对所有音频特征做固定长度的截断或填充:
    import numpy as np
    
    def extract_features(batch):
        audio, sr = librosa.load(batch["path"], sr=16000)
        # 统一截断/填充到固定长度(示例为10秒音频对应的采样点数)
        max_length = 160000
        if len(audio) > max_length:
            audio = audio[:max_length]
        else:
            audio = np.pad(audio, (0, max_length - len(audio)), mode='constant')
        # 使用wav2vec2特征提取器处理
        batch["input_values"] = feature_extractor(audio, sampling_rate=sr)["input_values"][0]
        return batch
    
    # 应用map时指定非批量处理,避免形状不匹配
    train_dataset = train_dataset.map(extract_features, batched=False)
    

内容的提问来源于stack exchange,提问作者Sneha T S

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 07:57:02