You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

修复RandomForestClassifier调用predict_proba缺失X参数报错

问题描述

使用RandomForestClassifier构建二分类器,建模前先基于单特征AUC得分完成特征选择,后续调用封装的模型训练函数获取AUC指标时运行失败,暂未提供配套数据集。

原始代码

依赖导入与数据处理流程

import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score, roc_auc_score
from sklearn.feature_selection import VarianceThreshold

df_process_label1 = 'AAA'
X = df_process.iloc[:,200:500]
y = df_process[df_process_label1].values

import sklearn
from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(X,y, test_size = 0.2, random_state = 0)

constant_filter = VarianceThreshold(threshold = 0.01)
constant_filter.fit(X_train)
X_train_filter = constant_filter.transform(X_train)
X_test_filter = constant_filter.transform(X_test)


roc_auc = []
for features in X_train.columns:
    clf = RandomForestClassifier(n_estimators = 100, random_state=0)
    clf.fit(X_train[features].to_frame(), y_train)
    y_pred = clf.predict(X_test[features].to_frame())
    roc_auc.append(roc_auc_score(y_test, y_pred))


roc_values = pd.Series(roc_auc)
roc_values.index = X_train.columns
roc_values.sort_values(ascending = False, inplace =True)


sel = roc_values[roc_values>0.5]
sel


X_train_roc = X_train[sel.index]
X_test_roc = X_test[sel.index]

def run_randomForest(X_train, X_test, y_train, y_test):
    clf = RandomForestClassifier(n_estimators=100, random_state=0, n_jobs=1)
    clf.fit(X_train, y_train)
    y_pred1 = clf.predict(X_test)
    print('Accuracy on test set: ', accuracy_score(y_test, y_pred))
    print(roc_auc_score(y_test, RandomForestClassifier.predict_proba(X_test)[:,1]))

函数调用代码

%time
run_randomForest(X_train_roc, X_test_roc, y_train, y_test)

运行报错信息

TypeError: predict_proba() missing 1 required positional argument: 'X'
修复方案

代码存在两处明确错误,直接导致报错和结果异常:

  • predict_proba调用方式错误:直接通过类名RandomForestClassifier调用实例方法,没有绑定训练完成的模型实例clf,方法无法定位要使用的训练好的模型,因此抛出参数缺失错误。正确调用方式为clf.predict_proba(X_test)。
  • 准确率计算引用变量错误:函数内当前模型的测试集预测结果被赋值给y_pred1,但计算准确率时传入的是单特征筛选循环中定义的全局变量y_pred,打印的准确率和当前训练的模型无关联,需要统一变量名。
    另外前期方差过滤生成的X_train_filter、X_test_filter后续未被使用,属于冗余代码,不影响运行可自行清理。

修复后的函数代码

def run_randomForest(X_train, X_test, y_train, y_test):
    clf = RandomForestClassifier(n_estimators=100, random_state=0, n_jobs=1)
    clf.fit(X_train, y_train)
    y_pred = clf.predict(X_test)
    y_pred_proba = clf.predict_proba(X_test)[:, 1]
    print('Accuracy on test set: ', accuracy_score(y_test, y_pred))
    print('AUC on test set: ', roc_auc_score(y_test, y_pred_proba))

替换原有函数后重新运行,即可正常输出测试集准确率和AUC指标。


内容的提问来源于stack exchange,提问作者Shu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 07:06:09