You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

语音情感识别(SER)代码触发ValueError:训练集为空的解决问询

语音情感识别代码运行触发ValueError:训练集为空

问题代码

#DataFlair - Load the data and extract features for each sound file
def load_data(test_size=0.2):
    x,y=[],[]
    for file in glob.glob("D:\archive\Actor_01\03-01-01-01-01-01-01.wav"):
        file_name=os.path.basename(file)
        emotion=emotions[file_name.split("-")[2]]
        if emotion not in observed_emotions:
            continue
        feature=extract_feature(file, mfcc=True, chroma=True, mel=True)
        x.append(feature)
        y.append(emotion)
    return train_test_split(np.array(x), y, test_size=test_size, random_state=9)

#DataFlair - Split the dataset
x_train,x_test,y_train,y_test=load_data(test_size=0.25)

错误信息

ValueError                                Traceback (most recent call last)
Input In [10], in <cell line: 2>()
      1 #DataFlair - Split the dataset
----> 2 x_train,x_test,y_train,y_test=load_data(test_size=0.25)

Input In [9], in load_data(test_size)
     10     x.append(feature)
     11     y.append(emotion)
---&gt; 12 return train_test_split(np.array(x), y, test_size=test_size, random_state=9)

File ~\anaconda3\lib\site-packages\sklearn\model_selection\_split.py:2420, in train_test_split(test_size, train_size, random_state, shuffle, stratify, *arrays)
   2417 arrays = indexable(*arrays)
   2419 n_samples = _num_samples(arrays[0])
-&gt; 2420 n_train, n_test = _validate_shuffle_split(
   2421     n_samples, test_size, train_size, default_test_size=0.25
   2422 )
   2424 if shuffle is False:
   2425     if stratify is not None:

File ~\anaconda3\lib\site-packages\sklearn\model_selection\_split.py:2098, in _validate_shuffle_split(n_samples, test_size, train_size, default_test_size)
   2095 n_train, n_test = int(n_train), int(n_test)
   2097 if n_train == 0:
-&gt; 2098     raise ValueError(
   2099         "With n_samples={}, test_size={} and train_size={}, the "
   2100         "resulting train set will be empty. Adjust any of the "
   2101         "aforementioned parameters.".format(n_samples, test_size, train_size)
   2102     )
   2104 return n_train, n_test

ValueError: With n_samples=0, test_size=0.25 and train_size=None, the resulting train set will be empty. Adjust any of the aforementioned parameters.

错误原因

核心原因:glob.glob()没有匹配到任何有效音频文件,导致x和y为空列表,np.array(x)的样本数为0,train_test_split无法对空数据集执行分割操作。

具体触发场景:

  1. 路径格式错误:Windows系统中,路径里的反斜杠\会被解析为转义字符(比如\a、\0),导致实际路径与预期不符,无法找到文件。
  2. 仅匹配单个文件且被过滤:代码中指定了单个具体文件名,即使路径正确,若该文件对应的情感不在observed_emotions列表中,会被continue语句跳过,最终x、y仍为空。
  3. 目标文件不存在:指定路径下的wav文件实际未存放,或路径拼写错误。

解决方法

  1. 修复路径格式
    使用原始字符串(在路径前加r)或转义反斜杠,避免转义字符干扰:

    # 原始字符串写法
    glob.glob(r"D:\archive\Actor_01\03-01-01-01-01-01-01.wav")
    # 转义反斜杠写法
    glob.glob("D:\\archive\\Actor_01\\03-01-01-01-01-01-01.wav")
    
  2. 匹配目录下所有音频文件
    原教程应该是批量处理目录中的所有wav文件,而非单个文件,改用通配符匹配:

    # 匹配Actor_01目录下所有wav文件
    glob.glob(r"D:\archive\Actor_01\*.wav")
    # 匹配所有Actor子目录下的wav文件(更符合教程场景)
    glob.glob(r"D:\archive\Actor_*\*.wav")
    
  3. 检查情感过滤逻辑
    确认emotions字典的映射关系正确,且目标文件对应的情感在observed_emotions列表内,避免有效文件被误过滤。比如:

    # 示例:确保情感列表包含目标文件对应的情感
    observed_emotions = ['neutral', 'happy', 'sad']
    emotions = {
        '01': 'neutral',
        '02': 'calm',
        '03': 'happy',
        # 其他情感映射...
    }
    
  4. 验证文件存在性
    手动打开文件资源管理器,检查指定路径下的wav文件是否存在,确认路径拼写、大小写完全匹配(Windows路径不区分大小写,但文件名需一致)。

内容的提问来源于stack exchange,提问作者Dem

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 09:10:35