You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用GridSearchCV调优决策树超参数时遇ValueError错误求助

解决GridSearchCV调优Pipeline中决策树超参数的无效参数问题

错误原因

你定义的超参数字典param_dist里的参数命名不符合Sklearn Pipeline的规则。Pipeline中的参数需要使用[步骤名称]__[模型参数名]的格式(双下划线分隔),而你用了自定义的dec_tree_xxx命名,导致GridSearchCV无法识别这些参数属于Pipeline中的哪个步骤。

修正方案

  1. 修正超参数字典的参数名
    你的Pipeline中决策树分类器的步骤名称是Classifier,所以参数名需要改为Classifier__criterion、Classifier__max_depth、Classifier__min_samples_leaf。
  2. 补充GridSearchCV的初始化代码
    你的代码中缺少grid对象的定义,需要先实例化GridSearchCV,传入Pipeline和修正后的参数网格。

完整修正代码

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import MinMaxScaler
from sklearn.tree import DecisionTreeClassifier
from sklearn.model_selection import GridSearchCV
from sklearn.metrics import accuracy_score

# 定义Pipeline
pipe_tree = Pipeline([
    ('Scaler', MinMaxScaler()),
    ('Classifier', DecisionTreeClassifier(max_depth=8))
])

# 训练基础模型
pipe_tree.fit(X_train, y_train)
pred_tree = pipe_tree.predict(X_test)
print("基础模型准确率:", accuracy_score(y_test, pred_tree))

# 修正后的超参数字典
param_dist={
    "Classifier__criterion": ["gini", "entropy"], 
    "Classifier__max_depth": [1,2,3,4,5,6,7, None],
    "Classifier__min_samples_leaf": [1, 2, 3]
}

# 初始化GridSearchCV
grid = GridSearchCV(estimator=pipe_tree, param_grid=param_dist, cv=5, scoring='accuracy')

# 拟合模型
grid.fit(X_train, y_train)

# 查看最优参数和最优得分
print("最优参数:", grid.best_params_)
print("最优交叉验证得分:", grid.best_score_)

# 用最优模型预测
best_pred = grid.predict(X_test)
print("最优模型测试集准确率:", accuracy_score(y_test, best_pred))

额外验证方法

如果不确定Pipeline支持哪些参数,可以运行以下代码查看所有可用参数:

print(pipe_tree.get_params().keys())

输出结果里会包含类似Classifier__criterion、Classifier__max_depth这样的参数名,直接参考这些命名即可。

内容的提问来源于stack exchange,提问作者Masudul Islam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 16:30:58