MNIST数据集处理报错:'Series'无'reshape'属性,求Python3.9适配方案
解决Python 3.9中MNIST数据集处理的AttributeError问题
在准备MNIST数据集用于神经网络训练时,代码在Python 3.6.8中可正常运行,但在Python 3.9.12中执行报错,错误信息如下:
--------------------------------------------------------------------------- AttributeError Traceback (most recent call last) Input In [3], in <cell line: 14>() 10 examples = y.shape[0] 11 #print(y.shape) 12 #print(y) ---> 14 y = y.reshape(1, examples) 15 Y_new = np.eye(digits)[y.astype('int32')] 16 Y_new = Y_new.T.reshape(digits, examples) File ~\Anaconda3\lib\site-packages\pandas\core\generic.py:5575, in NDFrame.__getattr__(self, name) 5568 if ( 5569 name not in self._internal_names_set 5570 and name not in self._metadata 5571 and name not in self._accessors 5572 and self._info_axis._can_hold_identifiers_and_holds_name(name) 5573 ): 5574 return self[name] -> 5575 return object.__getattribute__(self, name) AttributeError: 'Series' object has no attribute 'reshape'
使用的代码如下:
# load MNIST dataset X, y = fetch_openml('mnist_784', version=1, return_X_y=True) # prepare dataset X = X / 255 digits = 10 examples = y.shape[0] #print(y.shape) #print(y) y = y.reshape(1, examples) Y_new = np.eye(digits)[y.astype('int32')] Y_new = Y_new.T.reshape(digits, examples) # set train test split f = 60000 m_test = X.shape[0] - f # split dataset into train and test X_train, X_test = X[:f].T, X[f:].T Y_train, Y_test = Y_new[:,:f], Y_new[:,f:] np.random.seed(1) shuffle_index = np.random.permutation(f) X_train, Y_train = X_train[:, shuffle_index], Y_train[:, shuffle_index] print(X_train.shape[0]) print(X_train.shape[1])
问题原因
报错核心是:Python 3.9对应的新版本pandas中,fetch_openml返回的y是pandas Series对象,而Series没有reshape方法;但在Python 3.6的旧环境中,返回的y是numpy数组,因此可以直接调用reshape。
修改方案
有两种可靠的修改方式,任选其一即可:
方式一:加载数据时直接返回numpy数组
修改fetch_openml调用,添加as_frame=False参数,强制返回numpy数组而非pandas对象:
X, y = fetch_openml('mnist_784', version=1, return_X_y=True, as_frame=False)
方式二:手动将Series转为numpy数组
如果不想调整加载逻辑,在调用reshape前,用.to_numpy()(pandas官方推荐)将Series转为numpy数组:
# 替换原代码中的 y = y.reshape(1, examples) y = y.to_numpy().reshape(1, examples)
完整修改后的代码(方式二示例)
# load MNIST dataset X, y = fetch_openml('mnist_784', version=1, return_X_y=True) # prepare dataset X = X / 255 digits = 10 examples = y.shape[0] # 将Series转为numpy数组后执行reshape y = y.to_numpy().reshape(1, examples) Y_new = np.eye(digits)[y.astype('int32')] Y_new = Y_new.T.reshape(digits, examples) # set train test split f = 60000 m_test = X.shape[0] - f # split dataset into train and test X_train, X_test = X[:f].T, X[f:].T Y_train, Y_test = Y_new[:,:f], Y_new[:,f:] np.random.seed(1) shuffle_index = np.random.permutation(f) X_train, Y_train = X_train[:, shuffle_index], Y_train[:, shuffle_index] print(X_train.shape[0]) print(X_train.shape[1])
内容的提问来源于stack exchange,提问作者jlocke
相关产品推荐
相关产品推荐

