在R中绘制「比例内嵌套比例」条形图的技术求助
实现需求的堆叠条形图方案
首先先把你提供的数据集整理成清晰的格式:
ID# group category1 category2 category3 category4 category5 1 a 1 0 1 0 0 2 a 0 0 0 0 1 3 b 1 1 1 0 0
你的需求是绘制堆叠条形图:横轴是各个category,每个条形的高度代表该category占总观测的比例;条形内部按group拆分,展示每个group在该category中的占比(最终堆叠起来的总和就是该category的总比例)。下面我分别用Python和R给出实现方案:
Python 实现(使用pandas + matplotlib)
步骤1:数据预处理
首先我们需要把宽格式的数据集转换为长格式,然后计算每个category的总占比,以及每个group在对应category中的占比(也就是该group在该category的计数除以总观测数)。
import pandas as pd import matplotlib.pyplot as plt # 读取数据 data = pd.DataFrame({ 'ID#': [1,2,3], 'group': ['a','a','b'], 'category1': [1,0,1], 'category2': [0,0,1], 'category3': [1,0,1], 'category4': [0,0,0], 'category5': [0,1,0] }) # 转换为长格式:保留group,把category列转为行 long_data = data.melt(id_vars=['group'], value_vars=[col for col in data.columns if 'category' in col], var_name='category', value_name='is_in') # 筛选出属于该类别的观测 long_data = long_data[long_data['is_in'] == 1] # 计算总观测数 total_obs = len(data) # 计算每个category下各group的计数,再转为占总观测的比例 summary = long_data.groupby(['category', 'group']).size().reset_index(name='count') summary['proportion'] = summary['count'] / total_obs # 计算每个category的总比例(用于排序或参考) category_totals = summary.groupby('category')['proportion'].sum().reset_index(name='total_proportion')
步骤2:绘制堆叠条形图
# 获取所有category和group的唯一值 categories = sorted(summary['category'].unique()) groups = sorted(summary['group'].unique()) # 设置绘图参数 plt.figure(figsize=(10, 6)) bottom = [0]*len(categories) # 逐个group绘制堆叠部分 for group in groups: # 获取当前group在每个category的比例 group_props = [] for cat in categories: prop = summary[(summary['category'] == cat) & (summary['group'] == group)]['proportion'].values group_props.append(prop[0] if len(prop) > 0 else 0) # 绘制堆叠条形 plt.bar(categories, group_props, bottom=bottom, label=group) # 更新bottom值,用于下一个group的堆叠 bottom = [b + p for b, p in zip(bottom, group_props)] # 添加标签和标题 plt.xlabel('Category') plt.ylabel('Proportion of Total Observations') plt.title('Category Proportions by Group') plt.legend(title='Group') plt.xticks(rotation=45) plt.tight_layout() # 显示每个条形的总比例(可选) for i, cat in enumerate(categories): total = category_totals[category_totals['category'] == cat]['total_proportion'].values[0] plt.text(i, total + 0.01, f'{total:.1%}', ha='center') plt.show()
这段代码会生成你需要的图:每个条形的高度是该category占总观测的比例,内部堆叠的部分是对应group在该category中的贡献(占总观测的比例),标注的总比例能清晰展示每个category的整体占比。
R 实现(使用tidyverse + ggplot2)
步骤1:数据预处理
library(tidyverse) # 构建数据集 data <- tibble( ID = c(1,2,3), group = c("a","a","b"), category1 = c(1,0,1), category2 = c(0,0,1), category3 = c(1,0,1), category4 = c(0,0,0), category5 = c(0,1,0) ) # 转换为长格式并筛选有效观测 long_data <- data %>% pivot_longer(cols = starts_with("category"), names_to = "category", values_to = "is_in") %>% filter(is_in == 1) # 计算总观测数 total_obs <- nrow(data) # 计算每个category下各group的占总观测比例 summary <- long_data %>% count(category, group) %>% mutate(proportion = n / total_obs) # 计算每个category的总比例 category_totals <- summary %>% group_by(category) %>% summarise(total_proportion = sum(proportion))
步骤2:绘制堆叠条形图
ggplot(summary, aes(x = category, y = proportion, fill = group)) + geom_col(position = position_stack()) + # 添加每个条形的总比例标签 geom_text(data = category_totals, aes(x = category, y = total_proportion, label = sprintf("%.1f%%", total_proportion*100)), vjust = -0.5, inherit.aes = FALSE) + labs( x = "Category", y = "Proportion of Total Observations", title = "Category Proportions by Group", fill = "Group" ) + theme_minimal() + theme(axis.text.x = element_text(angle = 45, hjust = 1))
这个R版本的代码同样能实现你的需求,ggplot2的语法更简洁,适合快速调整样式。
内容的提问来源于stack exchange,提问作者Kim
相关产品推荐
相关产品推荐

