Pandas实战:统计目标变量取特定值时属性的出现次数
问题描述
现有如下Pandas数据集df:
Age Income Student Credit Rating Loan 0 <=30 high no fair no 1 <=30 high no excellent no 2 31-40 high no fair yes 3 >40 medium no fair yes 4 >40 low yes excellent no 5 31-40 low yes excellent yes
已定义属性列表:
attributes = ["Age", "Income", "Student", "Credit Rating"]
以及属性值字典:
attribute_values = { "Age": ["<=30", "31-40", ">40"], "Income": ["low", "medium", "high"], "Student": ["yes", "no"], "Credit Rating": ["fair", "excellent"] }
目前已统计出各属性值的出现次数:
attribute Age <=30 2 31-40 2 >40 2 attribute Income low 2 medium 1 high 3 attribute Student yes 4 no 2 attribute Credit Rating fair 3 excellent 3
需要进一步统计每个属性值对应的目标变量Loan为yes或no的次数,例如Age<=30的2条数据中Loan均为no,Credit Rating为fair的3条数据中1条Loan为no、2条为yes。
实现方法
方法1:使用groupby + value_counts
直接对每个属性分组后,统计Loan的取值频次,自动补全缺失结果为0,输出规整:
import pandas as pd for attr in attributes: print(f"attribute {attr}") # 分组后统计Loan的频次,unstack将结果转为宽表,fill_value补0 print(df.groupby(attr)["Loan"].value_counts().unstack(fill_value=0)) print()
输出示例(以Age为例):
attribute Age Loan no yes Age <=30 2 0 31-40 0 2 >40 1 1
方法2:使用pd.crosstab
交叉表可以直接生成属性与目标变量的频次统计,代码更简洁:
import pandas as pd for attr in attributes: print(f"attribute {attr}") print(pd.crosstab(df[attr], df["Loan"])) print()
输出格式与方法1完全一致,适合快速生成统计表格。
方法3:结合预定义属性值遍历统计
如果需要严格按照attribute_values中定义的顺序输出,可手动遍历筛选统计:
for attr in attributes: print(f"attribute {attr}") for val in attribute_values[attr]: # 筛选当前属性值的子集 subset = df[df[attr] == val]["Loan"] # 统计no和yes的次数 no_count = (subset == "no").sum() yes_count = (subset == "yes").sum() print(f"{val} no:{no_count} yes:{yes_count}") print()
输出示例:
attribute Age <=30 no:2 yes:0 31-40 no:0 yes:2 >40 no:1 yes:1
内容的提问来源于stack exchange,提问作者LeGOATJames23
相关产品推荐
相关产品推荐

