You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将测试集记录与模型预测结果关联并生成误差分析表?

如何关联客户编号与回归模型预测误差并生成指定结构的DataFrame?

我用scikit-learn搭建了回归模型,预测客户单笔最高消费额。数据集包含customer_number(客户编号)、metric_1、metric_2及目标列target(去年单笔最高消费),结构如下:

customer_number | metric_1 | metric_2 | target
----------------|----------|----------|-------
111             | A        | X        | 15
222             | A        | Y        | 20
333             | B        | Y        | 30

我将数据集拆分为训练集与测试集,对特征独热编码后训练RandomForestRegressor,代码如下:

target = pd.DataFrame(dataset, columns = ["target"])
features = dataset.drop("target", axis = 1)
train_features, test_features, train_target, test_target = train_test_split(features, target, test_size = 0.25)

train_features = pd.get_dummies(train_features)
test_features = pd.get_dummies(test_features)

model = RandomForestRegressor()
model.fit(X = train_features, y = train_target)

test_prediction = model.predict(X = test_features)

现在我能计算预测误差(error = abs(test_target - test_prediction)),但不知道如何把误差和客户编号关联起来,想要生成如下结构的DataFrame,用于分析特征与误差的相关性:

customer_number | target | prediction | error
----------------|--------|----------- |------
111             | 15     | 17         | 2
222             | 20     | 19         | 1
333             | 30     | 50         | 20

解决方案

核心是要在数据拆分过程中保留客户编号与测试集样本的对应关系,同时注意客户编号是标识符,不能作为特征参与模型训练(否则会导致模型过拟合)。以下是两种可行方法:

方法一:拆分时单独分离客户编号

直接在拆分前把客户编号列提取出来,和特征、目标一起拆分,确保测试集的客户编号和预测结果一一对应:

# 1. 分离客户编号、特征和目标
customer_ids = dataset["customer_number"]
features = dataset.drop(["target", "customer_number"], axis=1)  # 移除目标和客户编号
target = dataset["target"]

# 2. 同时拆分特征、目标和客户编号
train_features, test_features, train_target, test_target, train_ids, test_ids = train_test_split(
    features, target, customer_ids, test_size=0.25
)

# 3. 独热编码与模型训练(和原代码一致)
train_features = pd.get_dummies(train_features)
test_features = pd.get_dummies(test_features)

model = RandomForestRegressor()
model.fit(X=train_features, y=train_target)
test_prediction = model.predict(X=test_features)

# 4. 构建结果DataFrame
result_df = pd.DataFrame({
    "customer_number": test_ids.values,
    "target": test_target.values.flatten(),  # 转成一维数组避免维度不匹配
    "prediction": test_prediction,
    "error": abs(test_target.values.flatten() - test_prediction)
})

# 可选:按误差降序排序,快速定位误差最大的客户
result_df = result_df.sort_values(by="error", ascending=False)

方法二:利用原数据集索引关联客户编号

如果不想修改原拆分代码,可以通过测试集的索引从原数据集中回溯客户编号(前提是原数据集的索引未被打乱):

# 原拆分、编码、训练代码保持不变
target = pd.DataFrame(dataset, columns = ["target"])
features = dataset.drop("target", axis = 1)
train_features, test_features, train_target, test_target = train_test_split(features, target, test_size = 0.25)

train_features = pd.get_dummies(train_features)
test_features = pd.get_dummies(test_features)

model = RandomForestRegressor()
model.fit(X = train_features, y = train_target)
test_prediction = model.predict(X = test_features)

# 1. 通过测试集索引获取对应客户编号
test_ids = dataset.loc[test_features.index, "customer_number"]

# 2. 构建结果DataFrame
result_df = pd.DataFrame({
    "customer_number": test_ids,
    "target": test_target.values.flatten(),
    "prediction": test_prediction,
    "error": abs(test_target.values.flatten() - test_prediction)
})

关键注意点

  • 不要将customer_number作为特征输入模型:客户编号是唯一标识,若参与编码会生成大量独热编码列,导致模型过拟合,无法泛化到新客户。
  • 处理维度问题:train_target如果是DataFrame格式,其values是二维数组,需要用flatten()或ravel()转成一维,才能和一维的test_prediction计算误差。

内容的提问来源于stack exchange,提问作者SRJCoding

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 08:57:11