You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何调整Decision Tree Classifier参数以达到最佳准确率?

如何让决策树模型的测试集准确率接近验证集最佳准确率

你的决策树模型在验证集上达到了0.8878504672897196的最佳准确率,对应模型参数为DecisionTreeClassifier(criterion='entropy', max_depth=4, min_samples_leaf=6, min_samples_split=5),但在测试集上仅能达到0.8333333333333334的准确率,同时分类报告显示类别1的召回率仅为0.32,存在明显的类别不平衡问题。以下是具体的优化方案:

一、先修复参数调优的代码错误

你当前的循环逻辑存在核心问题:虽然遍历了max_depth从1到19,但在创建clf时固定写死了max_depth=4,等于完全没对这个参数进行调优,自然找不到真正最优的参数组合。

修改后的参数循环代码:

best_clf = None
best_accuracy = 0.0

# 同时调优多个关键参数
for max_depth in range(1, 20):  
    for min_samples_split in [3,5,7,9]:
        for min_samples_leaf in [4,6,8]:
            # 使用循环变量作为参数值
            clf = DecisionTreeClassifier(
                criterion="entropy", 
                max_depth=max_depth, 
                min_samples_split=min_samples_split, 
                min_samples_leaf=min_samples_leaf,
                random_state=42
            )
            clf.fit(X_train, Y_train)
            
            Y_val_pred = clf.predict(X_val)
            accuracy = accuracy_score(Y_val, Y_val_pred)
            if accuracy > best_accuracy:
                best_accuracy = accuracy
                best_clf = clf

二、用交叉验证替代单一验证集

单一验证集的结果存在随机性,建议使用网格搜索+交叉验证来寻找最优参数,这样得到的模型泛化能力更强:

from sklearn.model_selection import GridSearchCV

# 定义参数网格
param_grid = {
    'max_depth': range(1,15),
    'min_samples_split': [3,5,7],
    'min_samples_leaf': [4,6,8],
    'criterion': ['entropy', 'gini']
}

# 初始化网格搜索,用5折交叉验证
grid_search = GridSearchCV(
    estimator=DecisionTreeClassifier(random_state=42),
    param_grid=param_grid,
    cv=5,
    scoring='accuracy',
    n_jobs=-1
)

grid_search.fit(X_train, Y_train)
best_clf = grid_search.best_estimator_
best_accuracy = grid_search.best_score_

print("交叉验证最佳准确率:", best_accuracy)
print("最佳模型参数:", best_clf.get_params())

三、处理类别不平衡问题

从分类报告可以看出,类别1的样本量仅为25,远少于类别0的83,模型偏向于预测多数类,导致少数类召回率极低,拉低了整体泛化能力:

  • 添加类别权重:在创建分类器时加入class_weight='balanced',让模型自动给少数类更高的权重
clf = DecisionTreeClassifier(
    criterion="entropy",
    max_depth=...,
    min_samples_split=...,
    min_samples_leaf=...,
    class_weight='balanced',
    random_state=42
)
  • 可选:对少数类进行过采样(如SMOTE)或对多数类进行欠采样,平衡数据集分布

四、优化后剪枝策略

你当前的剪枝函数仅移除重复叶子,建议使用sklearn内置的代价复杂度剪枝(ccp_alpha),这是更系统的树剪枝方法,能有效降低过拟合:

# 生成ccp_alpha候选值
path = best_clf.cost_complexity_pruning_path(X_train, Y_train)
ccp_alphas, impurities = path.ccp_alphas, path.impurities

# 遍历不同的ccp_alpha,选择验证集表现最好的
clfs = []
for ccp_alpha in ccp_alphas:
    clf = DecisionTreeClassifier(
        criterion="entropy",
        max_depth=best_clf.max_depth,
        min_samples_split=best_clf.min_samples_split,
        min_samples_leaf=best_clf.min_samples_leaf,
        ccp_alpha=ccp_alpha,
        random_state=42
    )
    clf.fit(X_train, Y_train)
    clfs.append(clf)

# 选择验证集准确率最高的剪枝后模型
val_scores = [accuracy_score(Y_val, clf.predict(X_val)) for clf in clfs]
best_pruned_clf = clfs[val_scores.index(max(val_scores))]

五、精简特征维度

查看特征重要性结果,移除那些重要性极低的特征,减少模型的噪声输入,提升泛化能力:

# 保留重要性大于0.01的特征
selected_features = feature_importance[feature_importance['Importance'] > 0.01]['Feature'].tolist()
X_train = X_train[selected_features]
X_val = X_val[selected_features]
X_test = X_test[selected_features]

参数翻译说明

  • max_depth=4:最大树深=4
  • min_samples_split=5:最小分裂样本数=5
  • min_samples_leaf=6:最小叶子样本数=6

内容的提问来源于stack exchange,提问作者Elisa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 18:34:59