You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用ImageDataGenerator处理8分类图像数据集实现数据标准化的问题咨询

问题1:从分类文件夹中生成x_train的实现方法

你可以通过遍历目录读取图像的方式一次性加载全部数据,内存足够的场景下有两种常用实现方式:

手动遍历读取实现

import os
import numpy as np
from PIL import Image

train_root = "./train" # 替换为你的train文件夹实际路径
img_size = (224, 224) # 按你的实际需求调整图像统一尺寸
x_train = []
y_train = []
class_names = sorted(os.listdir(train_root)) # 自动读取8个类别文件夹名,顺序按文件名排序

for cls_idx, cls in enumerate(class_names):
    cls_path = os.path.join(train_root, cls)
    if not os.path.isdir(cls_path):
        continue
    # 遍历当前类别下的所有图像
    for img_name in os.listdir(cls_path):
        img_path = os.path.join(cls_path, img_name)
        # 读取图像并预处理
        img = Image.open(img_path).convert("RGB") # 单通道灰度图请改为.convert("L")
        img = img.resize(img_size)
        img_arr = np.array(img)
        # 存入数组
        x_train.append(img_arr)
        y_train.append(cls_idx)

# 转换为numpy数组,x_train形状为(总样本数, 高度, 宽度, 通道数),y_train形状为(总样本数,)
x_train = np.array(x_train)
y_train = np.array(y_train)

Keras快速实现

如果使用TensorFlow框架,可以直接调用内置接口快速加载:

import tensorflow as tf
import numpy as np

dataset = tf.keras.utils.image_dataset_from_directory(
    "./train",
    image_size=(224,224),
    batch_size=1000000 # 设为大于总样本数的数值即可一次性加载全部数据
)
x_train = np.concatenate([x for x, y in dataset], axis=0)
y_train = np.concatenate([y for x, y in dataset], axis=0)
问题2:标准化统计量是否需要按类别单独计算

绝大多数多分类场景下不需要按类别单独计算统计量。
逐特征标准化的核心是把输入特征的全局分布调整为均值0、标准差1,降低特征尺度差异对训练的影响,统计量基于全训练集计算即可。如果按类别单独计算,相当于在预处理阶段引入了类别信息,会造成数据泄露,破坏特征全局分布的一致性,反而会降低模型泛化能力。

仅在不同类别特征分布差异极大、且明确需要保留类别内分布特征的特殊场景下,才需要按类别计算统计量,实现方式如下:

from tensorflow.keras.preprocessing.image import ImageDataGenerator

datagen = ImageDataGenerator(featurewise_center=True, featurewise_std_normalization=True)
normalized_x_train = np.zeros_like(x_train)

# 按类别拆分数据分别拟合、标准化
for cls_id in range(8):
    # 提取当前类别的所有样本
    cls_x = x_train[y_train == cls_id]
    # 基于当前类别样本计算统计量
    datagen.fit(cls_x)
    # 对当前类别样本做标准化
    for idx, sample in enumerate(cls_x):
        normalized_x_train[y_train == cls_id][idx] = datagen.standardize(sample)

再次提醒:无特殊需求的场景下直接全局调用datagen.fit(x_train)即可,不要按类别单独计算。

内容的提问来源于stack exchange,提问作者AAAA

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 13:15:02