Python实现ID3算法计算信息增益返回Series而非数值问题
问题根因
返回pandas.Series而非单个数值的核心原因是维度操作不匹配,常见触发场景有两种:
- 计算条件熵时未指定仅对标签列做聚合统计,pandas自动对所有列执行分组计数,返回携带所有列名的序列
- 调用增益计算函数时传入了整个DataFrame而非单个特征列,或是用
apply全局遍历后未做单列索引,直接输出了全列的计算结果
修复方案
核心修正逻辑
- 条件熵计算时,分组后仅针对标签列做统计,避免返回多列结果
- 明确传入单个特征列计算增益,如需遍历全列再额外做循环处理
- 所有聚合操作后优先提取数值,避免携带pandas的索引结构
可运行修正代码
import pandas as pd import numpy as np # 计算数据集经验熵 def calc_root_entropy(df, label_col='label'): label_cnt = df[label_col].value_counts().values total = len(df) ent = 0.0 for cnt in label_cnt: prob = cnt / total ent -= prob * np.log2(prob) return ent # 计算单个特征的条件熵 def calc_condition_entropy(df, feature_col, label_col='label'): total = len(df) # 仅对标签列做分组统计,避免多列返回 feature_group = df.groupby(feature_col)[label_col].count() cond_ent = 0.0 for feat_val, cnt in feature_group.items(): prob = cnt / total # 计算当前特征取值分组内的标签熵 sub_df = df[df[feature_col] == feat_val] sub_label_cnt = sub_df[label_col].value_counts().values sub_ent = 0.0 for sub_cnt in sub_label_cnt: sub_prob = sub_cnt / cnt sub_ent -= sub_prob * np.log2(sub_prob) cond_ent += prob * sub_ent return cond_ent # 计算单个特征的信息增益,直接返回float标量 def calc_single_feature_gain(df, feature_col, label_col='label'): root_ent = calc_root_entropy(df, label_col) return root_ent - calc_condition_entropy(df, feature_col, label_col)
调用示例
# 计算单个特征的增益,返回单个数值 age_gain = calc_single_feature_gain(train_df, 'age') print(age_gain) # 输出为单个float值 # 遍历所有特征计算增益,按需存储 feature_gain = {} for col in train_df.columns: if col == 'label': continue feature_gain[col] = calc_single_feature_gain(train_df, col) print(feature_gain)
如果之前是用apply批量计算得到的Series,直接按列名索引即可拿到单个值:
all_gain = train_df.drop(columns='label').apply(lambda x: calc_single_feature_gain(train_df, x.name)) # 取单个特征的增益 single_gain = all_gain['age']
内容的提问来源于stack exchange,提问作者Evan Gertis
相关产品推荐
相关产品推荐

