MNIST数据集尺寸调整疑问:为何需新增轴及无轴调整方法
MNIST数据集尺寸调整方法与问题解答
已实现的调整方法
我找到了将MNIST训练数据集从(60000, 28, 28)调整为(60000, 14, 14)的方法,代码及运行结果如下:
import tensorflow as tf import numpy as np (x_train, y_train), (x_test, y_test) = tf.keras.datasets.mnist.load_data() x_train, x_test = x_train[..., np.newaxis], x_test[..., np.newaxis] x_train_small = tf.image.resize(x_train, (14,14)).numpy() x_test_small = tf.image.resize(x_test, (14,14)).numpy() print(x_train.shape) print(x_test.shape) print(x_train_small.shape) print(x_test_small.shape)
运行结果:
(60000, 28, 28, 1) (10000, 28, 28, 1) (60000, 14, 14, 1) (10000, 14, 14, 1)
技术问题解答
1. 为何必须新增一个维度才能完成目标尺寸的调整?
这是因为tf.image.resize的输入要求是4维张量,标准格式为(batch_size, height, width, channels)(样本数、高度、宽度、通道数)。原始MNIST数据集加载后是3维的(样本数, 28, 28),它省略了单通道灰度图的通道轴。如果直接把3维张量传给tf.image.resize,TensorFlow会因为输入维度不符合API规范报错,所以必须新增一个通道维度(变成(60000,28,28,1)),才能让函数正确识别输入结构。
2. 是否存在无需新增维度即可完成该尺寸调整的其他方法?
当然有,以下几种方法都不需要提前新增通道维度就能完成尺寸调整:
方法一:用NumPy批量调整尺寸
import numpy as np from tensorflow.keras.datasets import mnist (x_train, _), _ = mnist.load_data() # 遍历每个样本,用np.resize调整尺寸 x_train_small = np.array([np.resize(img, (14,14)) for img in x_train]) print(x_train_small.shape) # 输出:(60000, 14, 14)
方法二:用OpenCV的resize函数
import cv2 import numpy as np from tensorflow.keras.datasets import mnist (x_train, _), _ = mnist.load_data() # 利用cv2.resize逐个处理样本 x_train_small = np.array([cv2.resize(img, (14,14)) for img in x_train]) print(x_train_small.shape) # 输出:(60000, 14, 14)
方法三:手动实现下采样(以均值降采样为例)
如果想避免依赖第三方库,可以手动对图像做2倍降采样,比如取每个2x2像素块的均值:
import numpy as np from tensorflow.keras.datasets import mnist (x_train, _), _ = mnist.load_data() # 重塑数组结构,按2x2块分组后取均值 x_train_small = x_train.reshape(60000, 14, 2, 14, 2).mean(axis=(2,4)) print(x_train_small.shape) # 输出:(60000, 14, 14)
注意:如果之后要把调整后的数据集用于TensorFlow的CNN模型,通常还是需要补加通道维度,因为大部分卷积层都要求输入为4维张量。
内容的提问来源于stack exchange,提问作者Minnie
相关产品推荐
相关产品推荐

