语音情感识别(SER)代码触发ValueError:训练集为空的解决问询
语音情感识别代码运行触发ValueError:训练集为空
问题代码
#DataFlair - Load the data and extract features for each sound file def load_data(test_size=0.2): x,y=[],[] for file in glob.glob("D:\archive\Actor_01\03-01-01-01-01-01-01.wav"): file_name=os.path.basename(file) emotion=emotions[file_name.split("-")[2]] if emotion not in observed_emotions: continue feature=extract_feature(file, mfcc=True, chroma=True, mel=True) x.append(feature) y.append(emotion) return train_test_split(np.array(x), y, test_size=test_size, random_state=9) #DataFlair - Split the dataset x_train,x_test,y_train,y_test=load_data(test_size=0.25)
错误信息
ValueError Traceback (most recent call last) Input In [10], in <cell line: 2>() 1 #DataFlair - Split the dataset ----> 2 x_train,x_test,y_train,y_test=load_data(test_size=0.25) Input In [9], in load_data(test_size) 10 x.append(feature) 11 y.append(emotion) ---> 12 return train_test_split(np.array(x), y, test_size=test_size, random_state=9) File ~\anaconda3\lib\site-packages\sklearn\model_selection\_split.py:2420, in train_test_split(test_size, train_size, random_state, shuffle, stratify, *arrays) 2417 arrays = indexable(*arrays) 2419 n_samples = _num_samples(arrays[0]) -> 2420 n_train, n_test = _validate_shuffle_split( 2421 n_samples, test_size, train_size, default_test_size=0.25 2422 ) 2424 if shuffle is False: 2425 if stratify is not None: File ~\anaconda3\lib\site-packages\sklearn\model_selection\_split.py:2098, in _validate_shuffle_split(n_samples, test_size, train_size, default_test_size) 2095 n_train, n_test = int(n_train), int(n_test) 2097 if n_train == 0: -> 2098 raise ValueError( 2099 "With n_samples={}, test_size={} and train_size={}, the " 2100 "resulting train set will be empty. Adjust any of the " 2101 "aforementioned parameters.".format(n_samples, test_size, train_size) 2102 ) 2104 return n_train, n_test ValueError: With n_samples=0, test_size=0.25 and train_size=None, the resulting train set will be empty. Adjust any of the aforementioned parameters.
错误原因
核心原因:glob.glob()没有匹配到任何有效音频文件,导致x和y为空列表,np.array(x)的样本数为0,train_test_split无法对空数据集执行分割操作。
具体触发场景:
- 路径格式错误:Windows系统中,路径里的反斜杠
\会被解析为转义字符(比如\a、\0),导致实际路径与预期不符,无法找到文件。 - 仅匹配单个文件且被过滤:代码中指定了单个具体文件名,即使路径正确,若该文件对应的情感不在
observed_emotions列表中,会被continue语句跳过,最终x、y仍为空。 - 目标文件不存在:指定路径下的wav文件实际未存放,或路径拼写错误。
解决方法
修复路径格式
使用原始字符串(在路径前加r)或转义反斜杠,避免转义字符干扰:# 原始字符串写法 glob.glob(r"D:\archive\Actor_01\03-01-01-01-01-01-01.wav") # 转义反斜杠写法 glob.glob("D:\\archive\\Actor_01\\03-01-01-01-01-01-01.wav")匹配目录下所有音频文件
原教程应该是批量处理目录中的所有wav文件,而非单个文件,改用通配符匹配:# 匹配Actor_01目录下所有wav文件 glob.glob(r"D:\archive\Actor_01\*.wav") # 匹配所有Actor子目录下的wav文件(更符合教程场景) glob.glob(r"D:\archive\Actor_*\*.wav")检查情感过滤逻辑
确认emotions字典的映射关系正确,且目标文件对应的情感在observed_emotions列表内,避免有效文件被误过滤。比如:# 示例:确保情感列表包含目标文件对应的情感 observed_emotions = ['neutral', 'happy', 'sad'] emotions = { '01': 'neutral', '02': 'calm', '03': 'happy', # 其他情感映射... }验证文件存在性
手动打开文件资源管理器,检查指定路径下的wav文件是否存在,确认路径拼写、大小写完全匹配(Windows路径不区分大小写,但文件名需一致)。
内容的提问来源于stack exchange,提问作者Dem
相关产品推荐
相关产品推荐

