You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

决策树模型训练准确率高但测试准确率低的原因及优化咨询

决策树分类模型训练/测试准确率差距过大的问题

我正在构建一个**决策树分类(Decision Tree Classification)**模型,数据集包含11个特征列+1个目标列(共12列)。采用70%训练集、30%测试集的划分方式,对X_train执行Z-score标准化((x-均值)/标准差)。模型训练完成后,训练准确率达0.995,但测试准确率仅0.501。

起初对如此大的差距感到困惑,查阅资料得知应先划分数据集再做标准化(我正是这么做的)。尝试先标准化再划分数据(明知操作错误),结果训练与测试分数反而更接近。也怀疑过拟合,但认为70%的训练集占比不会导致过拟合。

注:决策树的超参数是通过GridSearchCV得到的最优参数。我怀疑测试分数差是因为未对测试集执行变换?请问有什么方法可以提升测试分数?


数据集示例

特征集X前5行:

{'col1': {0: 1, 1: 2, 2: 3, 3: 4, 4: 5},
 'col2': {0: 50, 1: 37, 2: 29, 3: 47, 4: 18},
 'col3': {0: 1, 1: 1, 2: 1, 3: 1, 4: 0},
 'col4': {0: 57, 1: 38, 2: 27, 3: 51, 4: 1},
 'col5': {0: 7, 1: 26, 2: 18, 3: 7, 4: 14},
 'col6': {0: 0, 1: 7, 2: 2, 3: 1, 4: 3},
 'col7': {0: 7, 1: 11, 2: 12, 3: 4, 4: 0},
 'col8': {0: 0, 1: 2, 2: 2, 3: 0, 4: 1},
 'col9': {0: 2, 1: 0, 2: 1, 3: 1, 4: 0},
 'col10': {0: 142, 1: 465, 2: 679, 3: 916, 4: 195},
 'col11': {0: 21, 1: 23, 2: 27, 3: 5, 4: 1}}

目标集y前5行:

{0: 1, 1: 1, 2: 0, 3: 0, 4: 1}

代码实现

训练测试划分

# Separate into training (70%) and testing (30%) sets
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size = 0.3, random_state = 0)
X_train.shape, X_test.shape

Z-score标准化函数

# Create a function to load data
def transform_data(data):
    
    # Find mean of data
    mean=data.mean()
    
    # Find standard deviation of data
    sd=data.std()
    
    # Transform the data using a Z-score transformation
    data_tf = (data-mean)/sd
    
    # Returns dataset
    return data_tf

# Transform the train dataset
X_tf = transform_data(X_train)

决策树分类器与评估

# Create improved model
dtc = DecisionTreeClassifier(criterion='entropy', max_depth = None, min_samples_leaf=4, 
                             min_samples_split=10, splitter='best', random_state = 0)

# Fit the model
dtc.fit(X_tf,y_train)

# Create function to print model accuracy scores
def print_accuracy(model, x_train, y_train, x_test, y_test):
    
    # Get train and test accuracy scores
    train_acc = model.score(x_train, y_train)
    test_acc = model.score(x_test, y_test)
    
    # Print score
    print(f'Training accuracy: {train_acc}')
    print(f'Testing accuracy: {test_acc}')

# Accuracy score
print_accuracy(dtc, X_tf, y_train, X_test, y_test)

问题分析与解决方案

核心问题:测试集未使用训练集的统计量做标准化

你的猜测完全正确——测试集没有做标准化,或者说没有用训练集计算出的均值和标准差做标准化,这是导致测试准确率极低的直接原因。

决策树在训练时用的是经过Z-score变换后的X_tf,但评估时直接传入原始的X_test,特征尺度完全不匹配,模型自然无法正确预测。

修正方法:

改写标准化流程,保存训练集的均值和标准差,用它们来变换测试集:

# 改写标准化函数,返回变换后的数据+均值+标准差
def transform_data(data):
    mean = data.mean()
    sd = data.std()
    data_tf = (data - mean)/sd
    return data_tf, mean, sd

# 变换训练集,保存统计量
X_tf, train_mean, train_sd = transform_data(X_train)

# 用训练集的统计量变换测试集
X_test_tf = (X_test - train_mean)/train_sd

# 重新评估模型
print_accuracy(dtc, X_tf, y_train, X_test_tf, y_test)

更规范的做法是使用sklearn的标准工具类,它会自动保存训练集的统计量:

from sklearn.preprocessing import StandardScaler

scaler = StandardScaler()
X_tf = scaler.fit_transform(X_train)
X_test_tf = scaler.transform(X_test)

# 重新训练模型并评估
dtc.fit(X_tf, y_train)
print_accuracy(dtc, X_tf, y_train, X_test_tf, y_test)

其他可能的优化方向

  1. 重新审视过拟合问题
    尽管训练集占比70%,但决策树本身极易过拟合,尤其是当max_depth=None时,树会完全拟合训练数据。即使你用了GridSearchCV,也要确认交叉验证的设置是否合理(比如是否用了分层交叉验证stratified K-fold,避免因类别不平衡导致的参数偏差)。

    • 可以尝试限制max_depth(比如设置为5-10),进一步增大min_samples_leaf或min_samples_split的值,强制模型简化。
  2. 检查数据集的类别平衡
    从y.head()看目标列有0和1两类,需要确认整个数据集的类别分布是否均匀。如果某类占比过高(比如90%都是1),0.5的测试准确率可能远差于随机猜测,这时候需要用F1-score、混淆矩阵等指标替代准确率来评估模型。

  3. 特征工程优化

    • 检查特征相关性:用热力图查看特征间的共线性,删除高度相关的特征,减少冗余信息。
    • 筛选有效特征:用卡方检验、互信息等方法,保留对目标预测有显著贡献的特征,剔除无关噪声。

内容的提问来源于stack exchange,提问作者Wee Liang Kelven Lim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 18:03:16