使用glob加载音频数据集报错及PyTorch适配问题求助
音频情感识别数据集构建与PyTorch适配问题
我正在为CNN模型训练构建音频情感识别数据集,目前遇到以下问题:
- 不确定
glob读取文件的逻辑是否正确 - 不清楚
load_data函数的输出结果是否符合预期 - 数据集整理环节出现报错,同时需要指导如何将数据集转换为PyTorch可训练张量
我的代码
emotions={ '01':'neutral', '02':'calm', '03':'happy', '04':'sad', '05':'angry', '06':'fearful', '07':'disgust', '08':'surprised' } # 关注的情感类别 observed_emotions=['neutral', 'happy', 'sad', 'angry'] def extract_feature(emo_file, mfcc): with soundfile.SoundFile(emo_file) as emo_file: X, sr = librosa.load(emo_file, sr=22050, mono=True, offset=1.0, duration=2.0) mfccs = np.mean(librosa.feature.mfcc(y=X, sr=sr, n_mfcc=40).T, axis=0) result = np.hstack((mfccs)) return result def load_data(test_size=0.2): # 预留20%作为测试集 x,y=[],[] for file in glob.glob('/Users/.../NeuralNetworks/SER_dataset/Actor*/*'): emo_file=os.path.basename(file) plt.figure(figsize=(18, 3)) # librosa绘图 matplotlib.waveplot(y, sr=sr) plt.ylim([-0.1, 0.1]) emotion=emotions[emo_file.split("-")[2]] # 根据文件名解析标签 if emotion not in observed_emotions: continue feature=extract_feature(emo_file, mfcc=True) x.append(feature) y.append(emotion) return train_test_split(np.ndarray(x), y, test_size=test_size, random_state=9) numpy_dataset = load_data print(np.ndarray.shape(numpy_dataset))
运行报错
TypeError Traceback (most recent call last) Input In [127], in <cell line: 2>() 1 numpy_dataset = load_data ----> 2 print(np.ndarray.shape(numpy_dataset)) TypeError: 'getset_descriptor' object is not callable
我的PyTorch转换尝试
emo_dataset = torch.from_numpy(numpy_dataset)
问题解决与优化方案
一、修复当前报错
函数调用错误
原代码中numpy_dataset = load_data只是赋值函数对象,未执行函数。正确写法:# 执行函数并接收返回的划分后数据集 X_train, X_test, y_train, y_test = load_data()numpy数组转换错误
np.ndarray(x)是错误的构造方式,应该用np.array(x)将列表转为numpy数组:return train_test_split(np.array(x), y, test_size=test_size, random_state=9)未定义变量报错
matplotlib.waveplot(y, sr=sr)中的y和sr未定义,直接删除该行代码(或补全音频读取逻辑后再绘图)。文件路径传递错误
extract_feature(emo_file, mfcc=True)中的emo_file是文件名而非完整路径,会导致文件找不到,应传递file变量:feature = extract_feature(file, mfcc=True)
修正后的load_data函数:
def load_data(test_size=0.2): x,y=[],[] for file in glob.glob('/Users/.../NeuralNetworks/SER_dataset/Actor*/*'): emo_file_name = os.path.basename(file) emotion = emotions[emo_file_name.split("-")[2]] if emotion not in observed_emotions: continue feature = extract_feature(file, mfcc=True) x.append(feature) y.append(emotion) return train_test_split(np.array(x), y, test_size=test_size, random_state=9)
二、验证glob读取正确性
在load_data的循环中添加打印,检查读取的文件路径是否符合预期:
for file in glob.glob('/Users/.../NeuralNetworks/SER_dataset/Actor*/*'): print("读取文件路径:", file) # 确认是否为目标音频文件 # 后续代码...
如果输出均为音频文件(如.wav),说明glob路径正确。
三、转换为PyTorch可训练张量
标签数值化
PyTorch分类任务需要整数标签,先将情感字符串映射为索引:emotion_to_idx = {emo:i for i, emo in enumerate(observed_emotions)} y_train_np = np.array([emotion_to_idx[emo] for emo in y_train]) y_test_np = np.array([emotion_to_idx[emo] for emo in y_test])转为张量
import torch # 特征转为FloatTensor,标签转为LongTensor(适配CrossEntropyLoss) X_train_tensor = torch.FloatTensor(X_train) X_test_tensor = torch.FloatTensor(X_test) y_train_tensor = torch.LongTensor(y_train_np) y_test_tensor = torch.LongTensor(y_test_np)构建Dataset与DataLoader(推荐)
方便训练时批量处理、打乱数据:from torch.utils.data import Dataset, DataLoader class EmotionDataset(Dataset): def __init__(self, features, labels): self.features = features self.labels = labels def __len__(self): return len(self.features) def __getitem__(self, idx): return self.features[idx], self.labels[idx] # 创建数据集实例 train_dataset = EmotionDataset(X_train_tensor, y_train_tensor) test_dataset = EmotionDataset(X_test_tensor, y_test_tensor) # 创建DataLoader train_loader = DataLoader(train_dataset, batch_size=32, shuffle=True) test_loader = DataLoader(test_dataset, batch_size=32, shuffle=False)
内容的提问来源于stack exchange,提问作者rawonith
相关产品推荐
相关产品推荐

