如何用ggplot可视化面板数据的加权百分比?
问题
我有如下结构的面板数据,包含v1-v4的重复二元测量值,还有调查权重(custom_wt)和性别(gender)列。目标是:
- 计算并可视化任意年份/波次中v1取值为1的受访者的加权百分比,同理处理v2、v3、v4
- 横轴为v1/v2/v3/v4,纵轴为百分比
示例数据:
d <- data.frame( respondent_id = c(1, 1, 1, 2, 2, 2, 3, 3, 3, 4, 4, 4, 5, 5, 5, 6, 6, 6, 7, 7, 7, 8, 8, 8, 10), custom_wt = c(24, 24, 24, 26, 26, 26, 12, 12, 12, 14, 14, 14, 33, 33, 33, 32, 32, 32, 9, 9, 9, 10, 10, 10, 10), year = c(2002, 2004, 2006, 2002, 2004, 2006, 2002, 2004, 2006, 2002, 2004, 2006, 2002, 2004, 2006, 2002, 2004, 2006, 2002, 2004, 2006, 2002, 2004, 2006, 2006), v1 = c(0, 0, 1, NA, 1, 1, 0, NA, NA, 1, NA, NA, 1, 0, 0, 0, 0, 0, 1, 1, NA, 1, 0, 0, 0), v2 = c(0, 0, 1, NA, 1, 1, 0, NA, 1, NA, 1, 1, 0, NA, 1, 0, 1, 1, NA, 1, 1, 0, NA, 1, 1), v3 = c(0, 0, NA, 1, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, NA, 0, 0, 1, 0, 0, 0, 0, 0, 0), v4 = c(NA, 1, 0, 0, 0, 0, 0, 1, 0, 0, 0, 1, 0, 1, 0, 0, NA, 0, 0, 0, 1, 0, 0, NA, NA), gender = c(0, 0, 0, 1, 1, 1, 1, 1, 1, 0, 0, 0, 1, 1, 1, 0, 0, 0, 1, 1, 1, 0, 0, 0, 0) )
解决步骤
1. 数据整理(宽转长格式)
先把宽格式数据转为长格式,方便对v1-v4统一处理:
library(tidyverse) d_long <- d %>% pivot_longer(cols = starts_with("v"), names_to = "variable", values_to = "value") %>% filter(!is.na(value)) # 剔除缺失值,若需将缺失值视为0,可替换为mutate(value = replace_na(value, 0))
2. 计算加权百分比
分两种场景计算:
场景1:不区分年份,计算每个变量的整体加权百分比
weighted_summary <- d_long %>% group_by(variable) %>% summarise( weighted_pct = sum(custom_wt[value == 1]) / sum(custom_wt) * 100, .groups = "drop" )
场景2:按年份分组,计算每个变量在各年份的加权百分比
weighted_summary_by_year <- d_long %>% group_by(year, variable) %>% summarise( weighted_pct = sum(custom_wt[value == 1]) / sum(custom_wt) * 100, .groups = "drop" )
3. 可视化
场景1:整体加权百分比柱状图
ggplot(weighted_summary, aes(x = variable, y = weighted_pct)) + geom_col(fill = "#2c3e50", width = 0.6) + geom_text(aes(label = sprintf("%.1f%%", weighted_pct)), vjust = -0.5) + labs( title = "各变量取值为1的加权百分比", x = "变量", y = "加权百分比 (%)" ) + theme_minimal() + ylim(0, max(weighted_summary$weighted_pct) + 10) # 调整纵轴范围避免文字截断
场景2:按年份分组的分组柱状图
ggplot(weighted_summary_by_year, aes(x = variable, y = weighted_pct, fill = factor(year))) + geom_col(position = position_dodge(width = 0.7), width = 0.6) + geom_text(aes(label = sprintf("%.1f%%", weighted_pct)), position = position_dodge(width = 0.7), vjust = -0.5) + labs( title = "各年份下变量取值为1的加权百分比", x = "变量", y = "加权百分比 (%)", fill = "年份" ) + theme_minimal() + ylim(0, max(weighted_summary_by_year$weighted_pct) + 10)
补充说明
计算逻辑:加权百分比 =(取值为1的受访者权重之和)/(该组所有受访者权重之和)×100,若需要将缺失的二元值视为0参与计算,只需修改数据整理步骤的缺失值处理方式即可。
内容的提问来源于stack exchange,提问作者a_todd12
相关产品推荐
相关产品推荐

