逻辑回归中NumPy一维/二维数组掩码处理的索引错误解决
问题:Logistic Regression掩码处理中的形状不匹配错误
我在使用Logistic Regression模型时,训练前会对x、y数组进行掩码处理,预测阶段也需要执行掩码操作,但遇到了和x、y及掩码所用h、survey数组形状相关的错误。
原代码
# 训练阶段初始掩码 regmask = np.isfinite(x) & np.isfinite(y) & (survey > 2010) & (h == 40) print("初始变量形状 ", x.shape, y.shape, survey.shape, h.shape) print() # 划分训练集与测试集 X_train, X_test, y_train, y_test = train_test_split(x[regmask].reshape(-1,1), y[regmask], test_size=0.2, random_state=42) # 创建逻辑回归模型 model = LogisticRegression() # 训练模型 model.fit(X_train, y_train) print("测试变量形状 ", X_test.shape, y_test.shape, survey.shape, h.shape) print(np.sum(np.isfinite(x)), np.sum(np.isfinite(y)), np.sum(survey > 2010), np.sum((h >= 28) & (h <= 36)), np.sum((h >= 20) & (h <= 28))) # 预测阶段新掩码 mask = np.isfinite(X_test) & np.isfinite(y_test) h_masked = h[mask] survey_masked = survey[mask] predmask = mask.flatten() & np.isfinite(survey_masked) & (((h_masked >= 28) & (h_masked <= 36)) | ((h_masked >= 20) & (h_masked <= 28)))
执行结果与报错
Initial variables shape (293240,) (293240,) (293240,) (293240,) Test variables shape (1754, 1) (1754,) (293240,) (293240,) 147196 147196 146930 14570 10730 --------------------------------------------------------------------------- IndexError Traceback (most recent call last) Cell In[50], line 61 59 # 预测阶段的小时掩码 60 mask = np.isfinite(X_test) & np.isfinite(y_test) ---> 61 h_masked = h[mask] 62 survey_masked = survey[mask] 63 predmask = mask.flatten() & np.isfinite(survey_masked) & (((h_masked >= 28) & (h_masked <= 36)) | ((h_masked >= 20) & (h_masked <= 28))) IndexError: too many indices for array: array is 1-dimensional, but 2 were indexed
疑问
- 如何解决该问题?
- 再次对h和survey变量做掩码的操作是否合理?
解决方案
以下是修正后的代码:
# 训练阶段初始掩码 regmask = np.isfinite(x) & np.isfinite(y) & (survey > 2010) & (h == 40) print("初始变量形状 ", x.shape, y.shape, survey.shape, h.shape) X_train, y_train = x[regmask].reshape(-1,1), y[regmask] # 创建逻辑回归模型 model = LogisticRegression() model.fit(X_train, y_train) # 生成预测阶段的掩码 predmask = np.isfinite(x) & np.isfinite(y) & (survey > 2010) & ((h >= 20) & (h <= 28) | (h >= 28) & (h <= 36)) X_test, y_test = x[predmask].reshape(-1, 1), y[predmask] # 检查变量形状 print("预测变量形状 ", x[predmask].shape, y[predmask].shape, survey.shape, h.shape) print("X_pred中的样本数:", X_test.shape[0]) y_pred = model.predict(X_test)
问题原因与说明
- 错误根源:原代码中
X_test是(1754, 1)的二维数组,生成的mask也是二维数组,但h是一维数组,用二维掩码索引一维数组会触发维度不匹配的错误。 - 掩码合理性:对h和survey做掩码是合理的,因为需要筛选符合特定条件的样本进行预测,确保预测数据的基础筛选逻辑(如有限值、年份条件)和训练数据一致,仅调整h的范围来验证模型在不同区间的表现。
内容的提问来源于stack exchange,提问作者spoolito
相关产品推荐
相关产品推荐

