如何从MNIST数据集中随机抽取约2000条样本?
解决方案:从MNIST训练集中随机抽取2000条样本
首先先修正你代码里的一个小笔误:x_train.shape的正确输出应该是(60000, 784),不是748,应该是手滑打错啦😉
下面给你几种简单可靠的方法来实现随机抽取2000条样本的需求:
方法1:使用NumPy的随机索引选择
这是最直接的方式,利用numpy.random.choice生成不重复的随机索引,精准提取目标样本:
import numpy as np import matplotlib.pyplot as plt from keras.datasets import mnist # 加载并预处理MNIST数据 (x_train, y_train), (x_test, y_test) = mnist.load_data() x_train = x_train.reshape(60000, 784) / 255 x_test = x_test.reshape(10000, 784) / 255 # 定义要抽取的样本数量 sample_size = 2000 # 生成不重复的随机索引,避免抽到同一条数据 random_indices = np.random.choice(x_train.shape[0], sample_size, replace=False) # 根据索引抽取样本和对应的标签 x_train_sample = x_train[random_indices] y_train_sample = y_train[random_indices] # 验证结果 print(x_train_sample.shape) # 输出:(2000, 784) print(y_train_sample.shape) # 输出:(2000,)
方法2:使用Scikit-learn的train_test_split
如果你的项目已经用到了scikit-learn,这个方法更简洁,还能保证样本的类别分布和原数据集一致:
from sklearn.model_selection import train_test_split import matplotlib.pyplot as plt from keras.datasets import mnist # 加载并预处理数据步骤同上... # 拆分出2000条样本,stratify参数保证类别分布与原数据一致 x_train_sample, _, y_train_sample, _ = train_test_split( x_train, y_train, test_size=2000, random_state=42, stratify=y_train ) # random_state固定随机种子,方便实验复现 print(x_train_sample.shape) # 输出:(2000, 784)
方法3:先打乱数据集再取前2000条
这种方式逻辑直观,先打乱整个数据集的顺序,再截取前2000条即可:
import numpy as np import matplotlib.pyplot as plt from keras.datasets import mnist # 加载并预处理数据步骤同上... # 生成打乱后的索引数组 shuffled_indices = np.random.permutation(x_train.shape[0]) # 截取前2000条样本 x_train_sample = x_train[shuffled_indices[:2000]] y_train_sample = y_train[shuffled_indices[:2000]] print(x_train_sample.shape) # 输出:(2000, 784)
小提醒:
- 如果需要固定随机结果(方便复现实验),可以给随机函数设置种子,比如
np.random.seed(42)。 replace=False参数(方法1中)确保不会重复抽取同一条样本,这是大多数场景下的合理选择。
内容的提问来源于stack exchange,提问作者Harlinton Palacios Mosquera
相关产品推荐
相关产品推荐

