You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用glob加载音频数据集报错及PyTorch适配问题求助

音频情感识别数据集构建与PyTorch适配问题

我正在为CNN模型训练构建音频情感识别数据集,目前遇到以下问题:

  • 不确定glob读取文件的逻辑是否正确
  • 不清楚load_data函数的输出结果是否符合预期
  • 数据集整理环节出现报错,同时需要指导如何将数据集转换为PyTorch可训练张量

我的代码

emotions={
  '01':'neutral',
  '02':'calm',
  '03':'happy',
  '04':'sad',
  '05':'angry',
  '06':'fearful',
  '07':'disgust',
  '08':'surprised'
}
# 关注的情感类别
observed_emotions=['neutral', 'happy', 'sad', 'angry']

def extract_feature(emo_file, mfcc):
    with soundfile.SoundFile(emo_file) as emo_file:
         X, sr = librosa.load(emo_file, sr=22050, mono=True, offset=1.0, duration=2.0)
         mfccs = np.mean(librosa.feature.mfcc(y=X, sr=sr, n_mfcc=40).T, axis=0)
         result = np.hstack((mfccs))
         return result

def load_data(test_size=0.2): # 预留20%作为测试集
    x,y=[],[]
    for file in glob.glob('/Users/.../NeuralNetworks/SER_dataset/Actor*/*'):
        emo_file=os.path.basename(file)

        plt.figure(figsize=(18, 3))  # librosa绘图
        matplotlib.waveplot(y, sr=sr)
        plt.ylim([-0.1, 0.1])

        emotion=emotions[emo_file.split("-")[2]] # 根据文件名解析标签
        if emotion not in observed_emotions:
            continue
        feature=extract_feature(emo_file, mfcc=True)
        x.append(feature)
        y.append(emotion)
    return train_test_split(np.ndarray(x), y, test_size=test_size, random_state=9)

numpy_dataset = load_data
print(np.ndarray.shape(numpy_dataset))

运行报错

TypeError                                 Traceback (most recent call last)
Input In [127], in <cell line: 2>()
      1 numpy_dataset = load_data
----> 2 print(np.ndarray.shape(numpy_dataset))

TypeError: 'getset_descriptor' object is not callable

我的PyTorch转换尝试

emo_dataset = torch.from_numpy(numpy_dataset)

问题解决与优化方案

一、修复当前报错

  1. 函数调用错误
    原代码中numpy_dataset = load_data只是赋值函数对象,未执行函数。正确写法:

    # 执行函数并接收返回的划分后数据集
    X_train, X_test, y_train, y_test = load_data()
    
  2. numpy数组转换错误
    np.ndarray(x)是错误的构造方式,应该用np.array(x)将列表转为numpy数组:

    return train_test_split(np.array(x), y, test_size=test_size, random_state=9)
    
  3. 未定义变量报错
    matplotlib.waveplot(y, sr=sr)中的y和sr未定义,直接删除该行代码(或补全音频读取逻辑后再绘图)。

  4. 文件路径传递错误
    extract_feature(emo_file, mfcc=True)中的emo_file是文件名而非完整路径,会导致文件找不到,应传递file变量:

    feature = extract_feature(file, mfcc=True)
    

修正后的load_data函数:

def load_data(test_size=0.2):
    x,y=[],[]
    for file in glob.glob('/Users/.../NeuralNetworks/SER_dataset/Actor*/*'):
        emo_file_name = os.path.basename(file)
        emotion = emotions[emo_file_name.split("-")[2]]
        if emotion not in observed_emotions:
            continue
        feature = extract_feature(file, mfcc=True)
        x.append(feature)
        y.append(emotion)
    return train_test_split(np.array(x), y, test_size=test_size, random_state=9)

二、验证glob读取正确性

在load_data的循环中添加打印,检查读取的文件路径是否符合预期:

for file in glob.glob('/Users/.../NeuralNetworks/SER_dataset/Actor*/*'):
    print("读取文件路径:", file) # 确认是否为目标音频文件
    # 后续代码...

如果输出均为音频文件(如.wav),说明glob路径正确。

三、转换为PyTorch可训练张量

  1. 标签数值化
    PyTorch分类任务需要整数标签,先将情感字符串映射为索引:

    emotion_to_idx = {emo:i for i, emo in enumerate(observed_emotions)}
    y_train_np = np.array([emotion_to_idx[emo] for emo in y_train])
    y_test_np = np.array([emotion_to_idx[emo] for emo in y_test])
    
  2. 转为张量

    import torch
    
    # 特征转为FloatTensor,标签转为LongTensor(适配CrossEntropyLoss)
    X_train_tensor = torch.FloatTensor(X_train)
    X_test_tensor = torch.FloatTensor(X_test)
    y_train_tensor = torch.LongTensor(y_train_np)
    y_test_tensor = torch.LongTensor(y_test_np)
    
  3. 构建Dataset与DataLoader(推荐)
    方便训练时批量处理、打乱数据:

    from torch.utils.data import Dataset, DataLoader
    
    class EmotionDataset(Dataset):
        def __init__(self, features, labels):
            self.features = features
            self.labels = labels
        
        def __len__(self):
            return len(self.features)
        
        def __getitem__(self, idx):
            return self.features[idx], self.labels[idx]
    
    # 创建数据集实例
    train_dataset = EmotionDataset(X_train_tensor, y_train_tensor)
    test_dataset = EmotionDataset(X_test_tensor, y_test_tensor)
    
    # 创建DataLoader
    train_loader = DataLoader(train_dataset, batch_size=32, shuffle=True)
    test_loader = DataLoader(test_dataset, batch_size=32, shuffle=False)
    

内容的提问来源于stack exchange,提问作者rawonith

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 13:24:46