You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让sklearn.datasets.load_files加载指定比例的图像数据?

嘿,太懂你这种GPU跑几小时瓶颈特征的痛苦了!完全可以用指定比例的数据先训练自定义顶层,之后再整合到ResNet上,甚至后续还能做微调,下面给你一步步说怎么实现:

1. 让load_files加载指定比例的数据

sklearn.datasets.load_files本身没有直接的采样参数,但我们可以手动对加载后的数据集做分层采样,保证类别分布和原数据一致,避免采样偏差。这里用sklearn.model_selection.train_test_split来实现:

from sklearn.datasets import load_files
from sklearn.model_selection import train_test_split
from keras.utils import np_utils
import numpy as np

def load_dataset(path, sample_fraction=0.2):
    # 加载完整数据集
    data = load_files(path)
    files = np.array(data['filenames'])
    targets = np_utils.to_categorical(np.array(data['target']))
    
    # 按比例分层采样,stratify参数保证类别分布不变
    sampled_files, _, sampled_targets, _ = train_test_split(
        files, targets, 
        train_size=sample_fraction, 
        stratify=targets, 
        random_state=42  # 固定随机种子,保证结果可复现
    )
    return sampled_files, sampled_targets

调用这个函数时,传入sample_fraction=0.2就能得到20%的样本啦。

2. 用采样数据训练自定义顶层

接下来我们先加载预训练的ResNet(去掉顶层分类器),提取采样数据的瓶颈特征,再训练自己的顶层:

步骤1:提取瓶颈特征

from keras.applications.resnet50 import ResNet50, preprocess_input
from keras.preprocessing.image import img_to_array, load_img

def extract_bottleneck_features(file_list, base_model):
    features = []
    for file_path in file_list:
        # 加载并预处理图像(ResNet要求输入尺寸是224x224)
        img = load_img(file_path, target_size=(224, 224))
        img_array = img_to_array(img)
        img_array = np.expand_dims(img_array, axis=0)
        img_array = preprocess_input(img_array)
        
        # 提取特征
        feat = base_model.predict(img_array, verbose=0)
        features.append(feat.flatten())  # 把特征展平成一维数组
    return np.array(features)

# 加载预训练ResNet,去掉顶层,用avg pooling输出特征
base_model = ResNet50(weights='imagenet', include_top=False, pooling='avg')

# 加载20%的采样数据并提取特征
sampled_files, sampled_targets = load_dataset('你的数据路径', sample_fraction=0.2)
bottleneck_features = extract_bottleneck_features(sampled_files, base_model)

步骤2:训练自定义顶层

现在用提取到的瓶颈特征训练你的顶层分类器:

from keras.models import Sequential
from keras.layers import Dense, Dropout

# 定义顶层模型(根据你的类别数调整最后一层的units)
top_model = Sequential([
    Dense(256, activation='relu', input_shape=bottleneck_features.shape[1:]),
    Dropout(0.5),  # 防止过拟合
    Dense(sampled_targets.shape[1], activation='softmax')
])

# 编译并训练
top_model.compile(
    optimizer='adam',
    loss='categorical_crossentropy',
    metrics=['accuracy']
)

top_model.fit(
    bottleneck_features, 
    sampled_targets,
    epochs=20,
    batch_size=32,
    validation_split=0.1  # 留10%做验证
)
3. 合并顶层与ResNet(可选,用于后续微调)

训练好顶层后,你可以把它和ResNet合并成完整模型,甚至解冻ResNet的部分底层做微调,提升精度:

# 合并成完整模型
full_model = Sequential([
    base_model,
    top_model
])

# 先冻结ResNet的所有层,避免训练时破坏预训练权重
base_model.trainable = False
full_model.compile(
    optimizer='adam',
    loss='categorical_crossentropy',
    metrics=['accuracy']
)

# (可选)解冻ResNet的最后几层做微调
# base_model.trainable = True
# 只解冻最后10层,前面的层保持冻结(保留通用特征)
# for layer in base_model.layers[:-10]:
#     layer.trainable = False
# 微调时用更小的学习率,避免破坏预训练权重
# full_model.compile(
#     optimizer=keras.optimizers.Adam(learning_rate=1e-5),
#     loss='categorical_crossentropy',
#     metrics=['accuracy']
# )
# 然后用完整数据集训练微调
小技巧:加速瓶颈特征提取

如果手动循环提取特征还是慢,你可以用ImageDataGenerator直接批量生成特征,效率更高:

from keras.preprocessing.image import ImageDataGenerator

datagen = ImageDataGenerator(preprocessing_function=preprocess_input)
# 从目录加载图像(注意目录结构要符合Keras的要求:每个类别一个子目录)
generator = datagen.flow_from_directory(
    '你的数据路径',
    target_size=(224, 224),
    batch_size=32,
    class_mode='categorical',
    shuffle=False  # 保证输出顺序和文件名对应
)

# 提取20%的数据特征,通过steps参数控制
total_samples = generator.samples
sampled_steps = int(total_samples * 0.2) // generator.batch_size + 1
bottleneck_features = base_model.predict(generator, steps=sampled_steps, verbose=1)

这样就能快速拿到你需要的比例的特征了!

内容的提问来源于stack exchange,提问作者Adrian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 10:04:42