多列独立分组统计用户占比并整合可视化的高效实现问询
高效统计多分组独立用户占比并统一可视化
问题背景
我有如下结构的DataFrame:
user_id segment device operating_system 0 51958733 small and above desktop Chrome OS 1 48983182 unfunded desktop Chrome OS 2 54011662 unfunded desktop (not set) 3 53932081 unfunded desktop (not set) 4 51537380 unfunded desktop Chrome OS ... ... ... ... ... 503657 53898078 unfunded desktop Macintosh 503658 52169624 long tail desktop Macintosh 503659 53965505 unfunded desktop Macintosh 503660 50678194 unfunded desktop Macintosh 503661 52143912 unfunded desktop Macintosh
需要高效统计各分组列的独立用户数(实际场景分组列更多),并将结果以占比形式展示在同一可视化图表中。目前我是针对每列单独写代码:
groupby_segment = eda_df.groupby('segment').ahid.nunique() groupby_segment.plot.bar(x="Segment", y="ahid", rot=70, title="Segment Distribution") plt.show(block=True);
这种方式效率极低,要手动创建/更新每个单元格,而且生成的图表相互独立,不利于对比展示。
解决方案
1. 批量统计多分组的独立用户占比
通过循环遍历所有目标分组列,一次性完成每个分组的独立用户数和占比计算:
import pandas as pd import matplotlib.pyplot as plt # 定义需要统计的分组列列表 group_cols = ['segment', 'device', 'operating_system'] # 初始化字典存储各分组结果 group_results = {} # 计算总独立用户数(用于占比计算) total_users = eda_df['user_id'].nunique() for col in group_cols: # 统计分组内独立用户数 user_count = eda_df.groupby(col)['user_id'].nunique() # 计算占比并转百分比格式 user_ratio = (user_count / total_users).round(4) * 100 # 合并结果存入字典 group_results[col] = pd.DataFrame({ '独立用户数': user_count, '占比(%)': user_ratio }).sort_values('占比(%)', ascending=False)
2. 统一子图可视化对比
将所有分组的占比柱状图放在同一画布的子图中,方便横向对比:
# 根据分组列数量设置子图布局 n_cols = 2 n_rows = (len(group_cols) + n_cols - 1) // n_cols fig, axes = plt.subplots(n_rows, n_cols, figsize=(15, 8)) axes = axes.flatten() # 转为一维数组便于遍历 for idx, (col, df) in enumerate(group_results.items()): ax = axes[idx] # 绘制占比柱状图 df['占比(%)'].plot(kind='bar', ax=ax, rot=70, color='#1f77b4') # 设置图表标题与标签 ax.set_title(f'{col} 分组用户占比') ax.set_ylabel('占比(%)') ax.set_xlabel(col) # 在柱子上标注具体数值 for p in ax.patches: height = p.get_height() ax.text(p.get_x() + p.get_width()/2., height, f'{height:.1f}%', ha='center', va='bottom') # 隐藏多余的空白子图 for ax in axes[len(group_results):]: ax.axis('off') plt.tight_layout() plt.show()
3. 可选:堆叠柱状图紧凑展示
如果希望在单个图表中对比所有分组维度的占比,可以将结果重塑后绘制堆叠柱状图:
# 合并所有分组的占比数据并重塑格式 all_ratios = pd.concat( [df['占比(%)'].rename(col) for col, df in group_results.items()], axis=1 ).fillna(0) # 绘制堆叠柱状图 all_ratios.plot(kind='bar', stacked=True, figsize=(12, 6), rot=70) plt.title('各分组维度用户占比对比') plt.ylabel('占比(%)') plt.legend(title='分组维度') plt.tight_layout() plt.show()
内容的提问来源于stack exchange,提问作者Simon Breton
相关产品推荐
相关产品推荐

