You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中合并预测结果与RepID的DataFrame异常问题排查

问题:合并预测结果与RepID时格式混乱的原因及解决方法

问题背景

原始DataFrame df 结构如下:

Weight    Height    Depth RepID       Code
0           18         3        14    257428      0
1            6         0         6    214932      0
2           21         6        16     17675      0
3           45         6        20     60819      0
4           30         6        16    262530      0
       ...       ...       ...       ...    ...
4223        36         6        28    331596      1
4224        24         9         0    331597      1
4225        36        12         8    331632      1
4226        24        24         0    331633      1
4227        30         9         0    331634      1

[4228 rows x 5 columns]

拆分训练集和测试集的代码:

y = df["Code"]
X = df.drop("Code", axis=1, errors='ignore')
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=TestSize, random_state=56)

预测代码:

clf.fit(X_train, y_train)
y_pred = clf.predict(X_test)

尝试保存预测结果与对应RepID的代码:

dfCSV = X_test["RepID"]
dfCSV["Code"] = pd.DataFrame(y_pred)
dfCSV.to_csv(PredictionFile)

预期得到的格式:

RepID       Code
0        84833      0
1        38388      1
2         2848      0
3         2992      1
4        28279      0
       ....    ...
423     74993      1
424     39924      1
425     55339      0
426     33882      1
427     64490      1

实际得到的混乱格式:

dfCSV
Out[15]: 
3792                                                262578
482                                                 129648
62                                                    7144
2998                                                127711
840                                                 157391
                       
207                                                 277899
569                                                  89965
2895                                                116296
570                                                 279183
ICD10         0
0    1
1    1
2    0
3    1
4    0
..  ...
Name: RepID, Length: 847, dtype: object

问题原因

  1. 数据类型错误:X_test["RepID"] 返回的是Series(一维结构),而非DataFrame。直接给Series添加新列的操作不符合其数据结构特性,会导致新数据以异常方式附加。
  2. 索引不匹配:pd.DataFrame(y_pred) 会生成默认从0开始的连续索引,但X_test的索引是原DataFrame拆分后的随机索引(如3792、482),两者索引不匹配,导致赋值时数据错位、结构混乱。

解决方法

方法1:先将RepID转为DataFrame再赋值

通过双层方括号获取RepID,得到DataFrame后直接添加预测结果列:

# 用双层方括号获取DataFrame,避免得到Series
dfCSV = X_test[["RepID"]].copy()
# 直接将y_pred数组赋值给新列,自动匹配X_test的索引
dfCSV["Code"] = y_pred
# 保存时可选去掉索引列,符合预期格式
dfCSV.to_csv(PredictionFile, index=False)

方法2:用pd.concat合并两个Series

将y_pred转为与X_test索引一致的Series,再按列合并:

# 将y_pred转为带索引的Series,确保索引与X_test匹配
pred_series = pd.Series(y_pred, name="Code", index=X_test.index)
# 合并RepID的Series和预测结果的Series
dfCSV = pd.concat([X_test["RepID"], pred_series], axis=1)
dfCSV.to_csv(PredictionFile, index=False)

内容的提问来源于stack exchange,提问作者asmgx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 20:40:31