You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用MAPIE为GradientBoostingRegressor求置信区间遇维度错误求助

解决MAPIE与GradientBoostingRegressor的维度错误问题

问题场景

尝试使用MAPIE库为GradientBoostingRegressor模型计算置信区间时,触发维度错误,原代码及报错信息如下:

原代码

import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.ensemble import GradientBoostingRegressor
from mapie.regression import MapieQuantileRegressor

Model = GradientBoostingRegressor(
 n_estimators = 500,
 max_depth = 4,
 min_samples_split = 5,
 learning_rate = 0.01,
 loss = "quantile")

X, y = make_regression(n_samples=5000, n_features=1, noise=20, random_state=59)

X = dict(enumerate(X.flatten(), 1))
y = dict(enumerate(y.flatten(), 1))

df = pd.DataFrame({'X':X, 'y':y})

X_train, X_tmp,  y_train, y_tmp  = train_test_split(df.X, df.y, test_size=2000, random_state=42)
X_calib, X_test, y_calib, y_test = train_test_split(X_tmp, y_tmp, test_size=1000, random_state=42)

alpha = 0.1
mapie = MapieQuantileRegressor(estimator=Model, cv="split", alpha=alpha)
mapie.fit(X_train, y_train, X_calib=X_calib, y_calib=y_calib)
y_pred, y_pis = mapie.predict(X_test)

predictions = y_test.to_frame()
predictions.columns = ['y_true']
predictions["point prediction"] = y_pred
predictions["lower"] = y_pis.reshape(-1,2)[:,0]
predictions["upper"] = y_pis.reshape(-1,2)[:,1]
predictions

报错信息

---------------------------------------------------------------------------
ValueError                                Traceback (most recent call last)
<ipython-input-1-2ca3f38dea46> in <cell line: 31>()
     29 alpha = 0.1
     30 mapie = MapieQuantileRegressor(estimator=Model, cv="split", alpha=alpha)
---> 31 mapie.fit(X_train, y_train, X_calib=X_calib, y_calib=y_calib)
     32 y_pred, y_pis = mapie.predict(X_test)
     33 

5 frames
/usr/local/lib/python3.10/dist-packages/sklearn/utils/validation.py in check_array(array, accept_sparse, accept_large_sparse, dtype, order, copy, force_all_finite, ensure_2d, allow_nd, ensure_min_samples, ensure_min_features, estimator, input_name)
    900             # If input is 1D raise error
    901             if array.ndim == 1:
---> 902                 raise ValueError(
    903                     "Expected 2D array, got 1D array instead:\narray={}.\n"
    904                     "Reshape your data either using array.reshape(-1, 1) if "

ValueError: Expected 2D array, got 1D array instead:
array=[ 0.5914918 -1.9605857  1.3109057 ... -1.1724522 -1.8717328 -2.839449 ].
Reshape your data either using array.reshape(-1, 1) if your data has a single feature or array.reshape(1, -1) if it contains a single sample.

报错原因

Scikit-learn(包括MAPIE封装的模型)要求输入特征必须是2D数组(形状为(n_samples, n_features)),但原代码中:

  1. 将make_regression生成的2D数组X通过flatten()转为1D后再转成字典,导致DataFrame中的X列是1D Series;
  2. 拆分数据集时直接使用df.X获取1D Series作为特征输入,不符合模型要求。

修复后的代码

import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.ensemble import GradientBoostingRegressor
from mapie.regression import MapieQuantileRegressor

# 定义模型
Model = GradientBoostingRegressor(
    n_estimators=500,
    max_depth=4,
    min_samples_split=5,
    learning_rate=0.01,
    loss="quantile"
)

# 生成数据集,保持X的2D结构
X, y = make_regression(n_samples=5000, n_features=1, noise=20, random_state=59)
df = pd.DataFrame(X, columns=["X"])
df["y"] = y

# 拆分数据集,用df[['X']]获取2D特征矩阵
X_train, X_tmp, y_train, y_tmp = train_test_split(df[['X']], df.y, test_size=2000, random_state=42)
X_calib, X_test, y_calib, y_test = train_test_split(X_tmp, y_tmp, test_size=1000, random_state=42)

alpha = 0.1
mapie = MapieQuantileRegressor(estimator=Model, cv="split", alpha=alpha)
mapie.fit(X_train, y_train, X_calib=X_calib, y_calib=y_calib)
y_pred, y_pis = mapie.predict(X_test)

# 整理结果,y_pis已为(n_samples, 2),无需reshape
predictions = y_test.to_frame(name="y_true")
predictions["point prediction"] = y_pred
predictions["lower"] = y_pis[:, 0]
predictions["upper"] = y_pis[:, 1]

print(predictions.head())

关键修改点

  • 移除X、y转字典的冗余操作,直接用make_regression生成的数组构造DataFrame,保留X的2D结构;
  • 拆分数据集时使用df[['X']](返回DataFrame,2D)替代df.X(返回Series,1D),确保特征输入符合模型要求;
  • 简化y_pis的处理:当alpha=0.1时,y_pis的形状为(n_samples, 2),直接通过索引提取上下界即可,无需reshape。

内容的提问来源于stack exchange,提问作者Karol

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 13:31:07