You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

逻辑回归中NumPy一维/二维数组掩码处理的索引错误解决

问题:Logistic Regression掩码处理中的形状不匹配错误

我在使用Logistic Regression模型时,训练前会对x、y数组进行掩码处理,预测阶段也需要执行掩码操作,但遇到了和x、y及掩码所用h、survey数组形状相关的错误。

原代码

# 训练阶段初始掩码
regmask = np.isfinite(x) & np.isfinite(y) & (survey > 2010) & (h == 40)
print("初始变量形状 ", x.shape, y.shape, survey.shape, h.shape)

print()

# 划分训练集与测试集
X_train, X_test, y_train, y_test = train_test_split(x[regmask].reshape(-1,1), y[regmask], test_size=0.2, random_state=42)

# 创建逻辑回归模型
model = LogisticRegression()

# 训练模型
model.fit(X_train, y_train)

print("测试变量形状 ", X_test.shape, y_test.shape, survey.shape, h.shape)
print(np.sum(np.isfinite(x)), np.sum(np.isfinite(y)), np.sum(survey > 2010), np.sum((h >= 28) & (h <= 36)), np.sum((h >= 20) & (h <= 28)))

# 预测阶段新掩码
mask = np.isfinite(X_test) & np.isfinite(y_test)
h_masked = h[mask]
survey_masked = survey[mask]
predmask = mask.flatten() & np.isfinite(survey_masked) & (((h_masked >= 28) & (h_masked <= 36)) | ((h_masked >= 20) & (h_masked <= 28)))

执行结果与报错

Initial variables shape  (293240,) (293240,) (293240,) (293240,)

Test variables shape  (1754, 1) (1754,) (293240,) (293240,)
147196 147196 146930 14570 10730
---------------------------------------------------------------------------
IndexError                                Traceback (most recent call last)
Cell In[50], line 61
     59 # 预测阶段的小时掩码
     60 mask = np.isfinite(X_test) & np.isfinite(y_test)
---&gt; 61 h_masked = h[mask]
     62 survey_masked = survey[mask]
     63 predmask = mask.flatten() & np.isfinite(survey_masked) & (((h_masked >= 28) & (h_masked <= 36)) | ((h_masked >= 20) & (h_masked <= 28)))

IndexError: too many indices for array: array is 1-dimensional, but 2 were indexed

疑问

  1. 如何解决该问题?
  2. 再次对h和survey变量做掩码的操作是否合理?

解决方案

以下是修正后的代码:

# 训练阶段初始掩码
regmask = np.isfinite(x) & np.isfinite(y) & (survey > 2010) & (h == 40)
print("初始变量形状 ", x.shape, y.shape, survey.shape, h.shape)

X_train, y_train = x[regmask].reshape(-1,1), y[regmask]

# 创建逻辑回归模型
model = LogisticRegression()
model.fit(X_train, y_train)

# 生成预测阶段的掩码
predmask = np.isfinite(x) & np.isfinite(y) & (survey > 2010) & ((h >= 20) & (h <= 28) | (h >= 28) & (h <= 36))
X_test, y_test = x[predmask].reshape(-1, 1), y[predmask]

# 检查变量形状
print("预测变量形状 ", x[predmask].shape, y[predmask].shape, survey.shape, h.shape)
print("X_pred中的样本数:", X_test.shape[0])

y_pred = model.predict(X_test)

问题原因与说明

  • 错误根源:原代码中X_test是(1754, 1)的二维数组,生成的mask也是二维数组,但h是一维数组,用二维掩码索引一维数组会触发维度不匹配的错误。
  • 掩码合理性:对h和survey做掩码是合理的,因为需要筛选符合特定条件的样本进行预测,确保预测数据的基础筛选逻辑(如有限值、年份条件)和训练数据一致,仅调整h的范围来验证模型在不同区间的表现。

内容的提问来源于stack exchange,提问作者spoolito

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 07:14:56