You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Random Forest在测试集获100%准确率,模型或代码是否有问题?

随机森林模型测试集100%准确率的问题排查

我使用scikit-learn中的Random Forest模型在测试集上取得了100%的准确率,怀疑模型或代码存在问题。相关代码如下:

ds = pd.read_csv('census-income.test(no unk.).csv')

df = pd.read_csv('census-income.data(no unk.).csv')

X = df 
y = df['income']

X_T = ds 
y_T = ds['income']

categorical_preprocessor = Pipeline(steps=[ ("onehot", OneHotEncoder(handle_unknown="ignore")) ])

preprocessor = ColumnTransformer([ ("categorical", categorical_preprocessor, ['workclass','education','martial-status','occupation','relationship','race','sex', 'native-country','income']), ],remainder='passthrough')

pipe = Pipeline(steps=[ ("preprocessor", preprocessor), ("classifier", RandomForestClassifier(n_estimators=128, max_depth=7)) ])

X_train = X 
X_test = X_T 
y_train = y 
y_test = y_T

pipe.fit(X_train, y_train) 
y_pred = pipe.predict(X_test)

print(classification_report(y_test, y_pred, digits=4)) 
print(confusion_matrix(y_test, y_pred)) 

训练数据包含workclass、education、martial-status等分类特征,以及作为目标变量的income列;测试数据结构与训练数据一致。

得到的混淆矩阵:

[[11360     0]
 [    0  3700]]

问题根源

核心错误在于预处理阶段将目标列income也加入了分类特征的处理列表。训练时模型会把income列作为特征学习,相当于直接把答案喂给了模型;测试时,测试集的income列同样会被转换成特征输入模型,模型自然能完全匹配到正确标签,从而得到100%的准确率。

修正方案

  1. 移除特征列表中的目标列:在ColumnTransformer的分类特征参数里删掉'income',它是需要预测的目标,不能作为特征。
  2. 正确划分特征与目标:确保训练集和测试集的特征数据不包含目标列,通过drop方法移除income列。

修正后的关键代码片段:

# 正确分离特征和目标变量
X = df.drop('income', axis=1)
y = df['income']

X_T = ds.drop('income', axis=1)
y_T = ds['income']

# 修正预处理的特征列表,移除income
preprocessor = ColumnTransformer([ 
    ("categorical", categorical_preprocessor, ['workclass','education','martial-status','occupation','relationship','race','sex', 'native-country']), 
], remainder='passthrough')

内容的提问来源于stack exchange,提问作者hre0

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 00:05:15