如何将n个文件夹内的PNG图像转换为可用于CNN与sklearn拆分的数据集
构建自定义Mel谱图像数据集(类似MNIST格式)
核心思路
通过遍历类别文件夹,将每个文件夹名映射为整数标签,加载所有PNG图像并转换为结构化数组,最终得到和MNIST格式一致的X(特征数据)与y(标签),方便使用sklearn的工具进行拆分和训练。
实现代码
首先导入所需依赖:
import os import numpy as np from PIL import Image from sklearn.model_selection import train_test_split
定义数据集加载函数:
def load_mel_spectrogram_dataset(root_dir): # 获取排序后的类别文件夹,保证标签映射的一致性 class_names = sorted(os.listdir(root_dir)) class_to_label = {name: idx for idx, name in enumerate(class_names)} data = [] targets = [] # 遍历每个类别目录 for class_name in class_names: class_dir = os.path.join(root_dir, class_name) if not os.path.isdir(class_dir): continue # 遍历目录下的所有PNG文件 for img_name in os.listdir(class_dir): if not img_name.lower().endswith('.png'): continue img_path = os.path.join(class_dir, img_name) # 加载图像并转为灰度(若你的Mel谱是彩色,可改为img.convert('RGB')) with Image.open(img_path) as img: img_gray = img.convert('L') # 转为浮点型数组并扁平化,匹配MNIST的(样本数, 特征数)结构 img_array = np.array(img_gray).flatten().astype('float32') # 归一化到0-1区间(可选,MNIST原始数据为0-255整数,可按需调整) img_array /= 255.0 data.append(img_array) targets.append(class_to_label[class_name]) # 转换为numpy数组,符合sklearn输入要求 X = np.array(data) y = np.array(targets).astype('int64') return X, y, class_names
加载数据集并拆分训练/测试集:
# 替换为你的数据集根目录 root_dir = "/path/to/your/mel_spectrogram_classes" # 加载数据 X, y, class_names = load_mel_spectrogram_dataset(root_dir) # 拆分训练集与测试集,参数可按需调整 X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42 ) # 验证数据集结构 print(f"训练集数据形状: {X_train.shape}, 训练集标签形状: {y_train.shape}") print(f"测试集数据形状: {X_test.shape}, 测试集标签形状: {y_test.shape}") print(f"类别与标签映射: {dict(enumerate(class_names))}")
可选调整
- 若需保留图像的2D结构(如用于CNN模型),去掉
.flatten()即可,此时X的形状为(样本数, 高度, 宽度) - 彩色Mel谱处理:将
img.convert('L')改为img.convert('RGB'),扁平化后特征数为高度×宽度×3 - 归一化策略:可保留原始0-255的整数范围,删除
img_array /= 255.0即可
内容的提问来源于stack exchange,提问作者Rookie91
相关产品推荐
相关产品推荐

