You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用sklearn的fetch_lfw_people时reshape报错:无法调整数组形状

解决fetch_lfw_people数据Reshape时的尺寸不匹配问题

嘿,我来帮你搞定这个reshape报错的问题!

首先,你遇到的核心问题是硬编码了图像尺寸(64x47),但实际获取到的数据集尺寸和这个完全不匹配。虽然你参考了文档,但大概率是误解了里面的尺寸说明——文档里提到的尺寸可能对应默认参数(比如默认resize=0.5),而你设置了resize=1,图像会保留原始大小,再加上min_faces_per_person=200的筛选条件,最终的图像尺寸和你预期的不一样。

快速修复方案:别硬写尺寸,从数据集里拿真实尺寸

fetch_lfw_people返回的对象自带一个images属性,它直接存储了二维的图像数据,我们可以从它的形状里提取真实的高度和宽度,这样就不会出错了:

from sklearn.datasets import fetch_lfw_people
from sklearn.model_selection import train_test_split

# 加载数据集
lfw_people = fetch_lfw_people(min_faces_per_person=200, resize=1)
X = lfw_people.data
y = lfw_people.target

# 获取图像的真实高度和宽度
img_height, img_width = lfw_people.images.shape[1], lfw_people.images.shape[2]
print(f"当前数据集的真实图像尺寸:{img_height}x{img_width}")

# 划分训练集和测试集
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25)

# 使用真实尺寸执行reshape
X_train = X_train.reshape(X_train.shape[0], 1, img_height, img_width).astype('float32')
X_test = X_test.reshape(X_test.shape[0], 1, img_height, img_width).astype('float32')

为什么会出现这个错误?

你计算的574*1*64*47=1,726,592,和报错里的总数据量6,744,500差了好几倍,这说明你假设的64x47完全不符合实际数据的像素数。

补充个小知识点:LFW数据集的原始图像是250x250像素,当你设置resize=1时,图像不会被缩放,所以每个样本的像素数是250*250=62500。你拿到的总数据量6,744,500除以62500,刚好对应训练集的样本数(大概108个),这也能验证你之前的尺寸假设是错误的。

内容的提问来源于stack exchange,提问作者afs_

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:13:52