You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何sklearn DecisionTreeClassifier未选出信息增益更高的最优分割?

sklearn DecisionTreeClassifier 最优分割点疑问

问题描述

用户对sklearn的DecisionTreeClassifier存在理解误区,相关代码如下:

from sklearn.tree import DecisionTreeClassifier, plot_tree
from scipy.stats import entropy
import numpy as np

train_x = np.array([[-0.62144528,  0.37728335],
 [-0.46795808,  0.2464509 ],
 [-0.31221227, -0.61933418],
 [-0.37111143, -0.37863888],
 [-0.4473217,  -0.15192771],
 [-0.55939442,  0.59526016],
 [-0.37522823, -0.2779457 ],
 [-0.39228952, -0.37050653],
 [-0.43533553, -0.02755128],
 [-0.45524276,  0.39507087],
 [-0.50147608,  0.58797464],
 [-0.47677197, -0.64571978],
 [-0.41001417, -0.16771494],
 [-0.39795968, -0.27224625],
 [-0.46032929,  0.24087007],
 [-0.50722624,  0.51068014],
 [-0.44299732,  0.00477296],
 [-0.37282845, -0.68609962],
 [-0.40829113, -0.26251665],
 [-0.46950366,  0.14817891],
 [-0.58785758,  0.25280204],
 [-0.45326652,  0.0034019 ],
 [-0.41441818,  0.14027937]])

train_y = np.array([[ 7], [ 2], [12], [11], [ 4], [ 7], [10], [11], [ 1], [ 3], [ 7], [10], [ 1], [ 8], [ 3], [ 2], [ 4], [ 5], [ 8], [ 4], [ 7], [ 4], [ 1]])

clf = DecisionTreeClassifier(random_state=0, criterion="entropy", max_depth=1)
clf = clf.fit(train_x, train_y)

values_root, counts_root = np.unique(train_y, return_counts=True)
counts_root = counts_root / len(train_x)
entropy_root = entropy(counts_root, base=2)

train_left = train_y[train_x[:, 1] <= -0.65]
train_right = train_y[train_x[:, 1] > -0.65]

values1, counts1 = np.unique(train_left, return_counts=True)
counts1 = counts1 / len(train_left)

values2, counts2 = np.unique(train_right, return_counts=True)
counts2 = counts2 / len(train_right)

entropy_1 = entropy(counts1, base=2)
entropy_2 = entropy(counts2, base=2)

print('Information gain manual split: ', entropy_root - (entropy_1 + entropy_2) / 2)

train_left = train_y[train_x[:, 1] <= clf.tree_.threshold[0]]
train_right = train_y[train_x[:, 1] > clf.tree_.threshold[0]]

values1, counts1 = np.unique(train_left, return_counts=True)
counts1 = counts1 / len(train_left)

values2, counts2 = np.unique(train_right, return_counts=True)
counts2 = counts2 / len(train_right)

entropy_1 = entropy(counts1, base=2)
entropy_2 = entropy(counts2, base=2)

print('Information gain sklearn split: ', entropy_root - (entropy_1 + entropy_2) / 2)

plot_tree(clf)

用户发现DecisionTreeClassifier找到的最优分割为y <= -0.215,对应信息增益为0.933,但手动设置分割点为-0.65时,计算出的信息增益可达1.656,远高于前者,且min_samples_leaf为默认值1,并非限制因素,因此询问哪里出错了。


问题原因及修正

你计算信息增益的核心错误在于子节点熵的加权方式:信息增益需要用左右样本占总样本的比例作为权重计算加权熵,而非简单对两个子节点的熵取平均。

错误计算方式

原代码中手动计算信息增益的错误逻辑:

# 错误:用平均代替加权平均
print('Information gain manual split: ', entropy_root - (entropy_1 + entropy_2) / 2)

正确计算方式

修改为按样本占比加权:

# 计算左右样本占总样本的比例
p_left = len(train_left) / len(train_y)
p_right = len(train_right) / len(train_y)
# 加权计算子节点熵总和
weighted_entropy = p_left * entropy_1 + p_right * entropy_2
# 信息增益 = 根节点熵 - 加权子节点熵总和
information_gain = entropy_root - weighted_entropy
print('Information gain manual split (corrected): ', information_gain)

验证修正结果

当分割点为-0.65时:

  • train_left仅有3个样本(对应train_y中的[12,10,5]),熵约为1.585
  • train_right有20个样本,熵约为2.641
  • 加权熵 = (3/23)*1.585 + (20/23)*2.641 ≈ 2.503
  • 根节点熵entropy_root约为2.641
  • 修正后的信息增益 ≈ 2.641 - 2.503 ≈ 0.138,远低于sklearn找到的0.933,符合实际最优分割的逻辑。

为什么sklearn选择-0.215?

当分割点为-0.215时,左右样本的加权熵更低,信息增益更高。用修正后的计算方式验证该分割点,会得到约0.933的信息增益,这确实是当前数据集下的最优分割。


内容的提问来源于stack exchange,提问作者Lara

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 05:12:09