如何用pandas实现单列与其余列对比统计计数及占比
实现方法
首先你需要先指定作为对比基准的核心列,以下提供两种常用的实现方案:
方案1:基于groupby实现
输出为长表结构,适合后续进一步数据处理:
# 指定基准对比列,可按需修改 base_col = "gender" compare_cols = [col for col in df.columns if col != base_col] result = {} for col in compare_cols: # 分组统计计数 count_df = df.groupby([base_col, col]).size().reset_index(name='计数') # 按基准列分组计算对应占比 count_df['占比'] = count_df.groupby(base_col)['计数'].transform(lambda x: x/x.sum()).round(4) result[col] = count_df # 查看指定列的对比结果 print(result["homework finishing"])
方案2:基于crosstab实现(更推荐)
pd.crosstab是pandas专门用于交叉统计的函数,输出为宽表结构,可读性更高:
base_col = "gender" compare_cols = [col for col in df.columns if col != base_col] for col in compare_cols: # 统计计数 count = pd.crosstab(df[base_col], df[col]) # 统计基准列维度下的占比,normalize参数设为'columns'可改为按对比列维度计算占比 ratio = pd.crosstab(df[base_col], df[col], normalize='index').round(4) # 合并计数和占比 cross_res = count.join(ratio, rsuffix='_占比') print(f"=== {base_col} 与 {col} 统计结果 ===") print(cross_res, "\n")
结果说明
以示例数据中gender与under 15的统计结果为例,输出如下:
=== gender 与 under 15 统计结果 === under 15 False True False_占比 True_占比 gender Female 2 2 0.5 0.5 Male 0 2 0.0 1.0
如果需要计算整体占比而非基准列内的占比,将normalize参数改为'all'即可。
内容的提问来源于stack exchange,提问作者Cooper
相关产品推荐
相关产品推荐

