为何sklearn DecisionTreeClassifier未选出信息增益更高的最优分割?
sklearn DecisionTreeClassifier 最优分割点疑问
问题描述
用户对sklearn的DecisionTreeClassifier存在理解误区,相关代码如下:
from sklearn.tree import DecisionTreeClassifier, plot_tree from scipy.stats import entropy import numpy as np train_x = np.array([[-0.62144528, 0.37728335], [-0.46795808, 0.2464509 ], [-0.31221227, -0.61933418], [-0.37111143, -0.37863888], [-0.4473217, -0.15192771], [-0.55939442, 0.59526016], [-0.37522823, -0.2779457 ], [-0.39228952, -0.37050653], [-0.43533553, -0.02755128], [-0.45524276, 0.39507087], [-0.50147608, 0.58797464], [-0.47677197, -0.64571978], [-0.41001417, -0.16771494], [-0.39795968, -0.27224625], [-0.46032929, 0.24087007], [-0.50722624, 0.51068014], [-0.44299732, 0.00477296], [-0.37282845, -0.68609962], [-0.40829113, -0.26251665], [-0.46950366, 0.14817891], [-0.58785758, 0.25280204], [-0.45326652, 0.0034019 ], [-0.41441818, 0.14027937]]) train_y = np.array([[ 7], [ 2], [12], [11], [ 4], [ 7], [10], [11], [ 1], [ 3], [ 7], [10], [ 1], [ 8], [ 3], [ 2], [ 4], [ 5], [ 8], [ 4], [ 7], [ 4], [ 1]]) clf = DecisionTreeClassifier(random_state=0, criterion="entropy", max_depth=1) clf = clf.fit(train_x, train_y) values_root, counts_root = np.unique(train_y, return_counts=True) counts_root = counts_root / len(train_x) entropy_root = entropy(counts_root, base=2) train_left = train_y[train_x[:, 1] <= -0.65] train_right = train_y[train_x[:, 1] > -0.65] values1, counts1 = np.unique(train_left, return_counts=True) counts1 = counts1 / len(train_left) values2, counts2 = np.unique(train_right, return_counts=True) counts2 = counts2 / len(train_right) entropy_1 = entropy(counts1, base=2) entropy_2 = entropy(counts2, base=2) print('Information gain manual split: ', entropy_root - (entropy_1 + entropy_2) / 2) train_left = train_y[train_x[:, 1] <= clf.tree_.threshold[0]] train_right = train_y[train_x[:, 1] > clf.tree_.threshold[0]] values1, counts1 = np.unique(train_left, return_counts=True) counts1 = counts1 / len(train_left) values2, counts2 = np.unique(train_right, return_counts=True) counts2 = counts2 / len(train_right) entropy_1 = entropy(counts1, base=2) entropy_2 = entropy(counts2, base=2) print('Information gain sklearn split: ', entropy_root - (entropy_1 + entropy_2) / 2) plot_tree(clf)
用户发现DecisionTreeClassifier找到的最优分割为y <= -0.215,对应信息增益为0.933,但手动设置分割点为-0.65时,计算出的信息增益可达1.656,远高于前者,且min_samples_leaf为默认值1,并非限制因素,因此询问哪里出错了。
问题原因及修正
你计算信息增益的核心错误在于子节点熵的加权方式:信息增益需要用左右样本占总样本的比例作为权重计算加权熵,而非简单对两个子节点的熵取平均。
错误计算方式
原代码中手动计算信息增益的错误逻辑:
# 错误:用平均代替加权平均 print('Information gain manual split: ', entropy_root - (entropy_1 + entropy_2) / 2)
正确计算方式
修改为按样本占比加权:
# 计算左右样本占总样本的比例 p_left = len(train_left) / len(train_y) p_right = len(train_right) / len(train_y) # 加权计算子节点熵总和 weighted_entropy = p_left * entropy_1 + p_right * entropy_2 # 信息增益 = 根节点熵 - 加权子节点熵总和 information_gain = entropy_root - weighted_entropy print('Information gain manual split (corrected): ', information_gain)
验证修正结果
当分割点为-0.65时:
train_left仅有3个样本(对应train_y中的[12,10,5]),熵约为1.585train_right有20个样本,熵约为2.641- 加权熵 = (3/23)*1.585 + (20/23)*2.641 ≈ 2.503
- 根节点熵
entropy_root约为2.641 - 修正后的信息增益 ≈ 2.641 - 2.503 ≈ 0.138,远低于sklearn找到的0.933,符合实际最优分割的逻辑。
为什么sklearn选择-0.215?
当分割点为-0.215时,左右样本的加权熵更低,信息增益更高。用修正后的计算方式验证该分割点,会得到约0.933的信息增益,这确实是当前数据集下的最优分割。
内容的提问来源于stack exchange,提问作者Lara
相关产品推荐
相关产品推荐

