You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将n个文件夹内的PNG图像转换为可用于CNN与sklearn拆分的数据集

构建自定义Mel谱图像数据集(类似MNIST格式)

核心思路

通过遍历类别文件夹,将每个文件夹名映射为整数标签,加载所有PNG图像并转换为结构化数组,最终得到和MNIST格式一致的X(特征数据)与y(标签),方便使用sklearn的工具进行拆分和训练。

实现代码

首先导入所需依赖:

import os
import numpy as np
from PIL import Image
from sklearn.model_selection import train_test_split

定义数据集加载函数:

def load_mel_spectrogram_dataset(root_dir):
    # 获取排序后的类别文件夹,保证标签映射的一致性
    class_names = sorted(os.listdir(root_dir))
    class_to_label = {name: idx for idx, name in enumerate(class_names)}
    
    data = []
    targets = []
    
    # 遍历每个类别目录
    for class_name in class_names:
        class_dir = os.path.join(root_dir, class_name)
        if not os.path.isdir(class_dir):
            continue
        
        # 遍历目录下的所有PNG文件
        for img_name in os.listdir(class_dir):
            if not img_name.lower().endswith('.png'):
                continue
            
            img_path = os.path.join(class_dir, img_name)
            # 加载图像并转为灰度(若你的Mel谱是彩色,可改为img.convert('RGB'))
            with Image.open(img_path) as img:
                img_gray = img.convert('L')
                # 转为浮点型数组并扁平化,匹配MNIST的(样本数, 特征数)结构
                img_array = np.array(img_gray).flatten().astype('float32')
                # 归一化到0-1区间(可选,MNIST原始数据为0-255整数,可按需调整)
                img_array /= 255.0
                
                data.append(img_array)
                targets.append(class_to_label[class_name])
    
    # 转换为numpy数组,符合sklearn输入要求
    X = np.array(data)
    y = np.array(targets).astype('int64')
    
    return X, y, class_names

加载数据集并拆分训练/测试集:

# 替换为你的数据集根目录
root_dir = "/path/to/your/mel_spectrogram_classes"

# 加载数据
X, y, class_names = load_mel_spectrogram_dataset(root_dir)

# 拆分训练集与测试集,参数可按需调整
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)

# 验证数据集结构
print(f"训练集数据形状: {X_train.shape}, 训练集标签形状: {y_train.shape}")
print(f"测试集数据形状: {X_test.shape}, 测试集标签形状: {y_test.shape}")
print(f"类别与标签映射: {dict(enumerate(class_names))}")

可选调整

  • 若需保留图像的2D结构(如用于CNN模型),去掉.flatten()即可,此时X的形状为(样本数, 高度, 宽度)
  • 彩色Mel谱处理:将img.convert('L')改为img.convert('RGB'),扁平化后特征数为高度×宽度×3
  • 归一化策略:可保留原始0-255的整数范围,删除img_array /= 255.0即可

内容的提问来源于stack exchange,提问作者Rookie91

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 17:46:27