R语言使用ggplot将10个非互斥变量绘制成单张分面条形图
实现方案
你需要先将宽格式的数据集转换为长格式,再统计每个年龄组下各选项的选择占比,最后用ggplot绘图即可,完整实现代码如下:
# 加载所需依赖包 library(tidyverse) # 导入示例数据集 df <- structure(list(improvement___1 = c(1, 0, 1, 1, 1, 1, 1, 1, 0, 1, 1, 1, 1, 1, 1, 1, 1, 1, 0, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1), improvement___2 = c(0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 1, 0, 0, 1, 0, 0, 0, 1, 1, 0, 1, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 1), improvement___3 = c(0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 1, 0, 1, 0, 0, 1, 0, 0, 1, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0), improvement___4 = c(1, 0, 1, 0, 0, 1, 1, 1, 0, 1, 1, 0, 1, 1, 0, 1, 0, 0, 1, 1, 1, 1, 1, 1, 1, 0, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 0, 1, 1, 1, 1, 1, 0, 1, 1, 1), improvement___5 = c(0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0), improvement___6 = c(0, 1, 0, 0, 1, 0, 0, 1, 0, 0, 0, 1, 0, 0, 0, 1, 0, 1, 0, 0, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 0, 1, 0, 1, 1, 1, 0, 1, 1, 1, 0, 1, 1), improvement___7 = c(0, 1, 1, 0, 0, 1, 0, 1, 1, 1, 1, 0, 1, 1, 0, 1, 1, 1, 1, 0, 1, 1, 0, 0, 0, 1, 0, 0, 0, 1, 0, 1, 1, 1, 1, 0, 1, 1, 1, 0, 1, 1, 0, 1, 1, 0), improvement___8 = c(0, 1, 0, 1, 1, 1, 0, 0, 1, 1, 1, 1, 1, 0, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 0, 1, 1, 1, 1, 1, 1, 0, 1, 1, 1, 0, 0, 1, 0, 0, 1, 0, 0, 0, 1, 1), improvement___9 = c(1, 0, 1, 0, 1, 0, 1, 1, 0, 1, 0, 1, 1, 1, 0, 1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0, 1, 1, 1, 1, 1, 0, 1, 1, 1, 0, 1, 0, 0, 0, 0, 0, 0, 1, 1, 1), improvement___10 = c(0, 0, 0, 0, 0, 0, 1, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0), AgeGroups = structure(c(2L, 1L, 2L, 1L, 2L, 1L, 3L, 1L, 2L, 2L, 3L, 2L, 2L, 2L, 2L, 3L, 2L, 2L, 3L, 1L, 2L, 2L, 1L, 2L, 2L, 1L, 1L, 1L, 1L, 2L, 1L, 3L, 3L, 2L, 2L, 3L, 2L, 2L, 3L, 3L, 1L, 2L, 3L, 1L, 2L, 1L), .Label = c("Young", "Middle", "Old" ), class = "factor")), row.names = c(NA, -46L), class = c("tbl_df", "tbl", "data.frame")) # 数据预处理:宽表转长表,统计各年龄组各选项的选择比例 plot_data <- df %>% # 把所有improvement开头的列整合为长格式 pivot_longer(cols = starts_with("improvement___"), names_to = "option", values_to = "selected") %>% # 提取选项编号用于排序 mutate(option_num = as.integer(str_extract(option, "\\d+$"))) %>% # 按年龄组、选项分组计算选择率 group_by(AgeGroups, option_num) %>% summarise(select_rate = mean(selected), .groups = "drop") # 绘制分面条形图 ggplot(plot_data, aes(x = factor(option_num), y = select_rate)) + geom_col(fill = "#2c7fb8") + # 按年龄组分面,可调整nrow参数修改排布方式 facet_wrap(~AgeGroups, nrow = 1) + labs(x = "选项编号", y = "选择比例") + # 把y轴转换为百分比格式 scale_y_continuous(labels = scales::percent_format(accuracy = 1)) + # 优化显示主题 theme_bw() + theme( panel.grid.major.x = element_blank(), strip.background = element_rect(fill = "#f0f0f0") )
关键说明
- 宽表转长表是核心操作,将分散在10列的选项数据整合为结构化的长格式,才能统一映射到ggplot的坐标轴
- 直接对0/1的选中状态求均值即可得到对应组的选择比例,无需额外计数
- 如果需要将x轴的选项编号替换为实际的问题选项文本,只需在
mutate步骤新增对应关系的option_label列,再把x轴映射改为option_label即可
内容的提问来源于stack exchange,提问作者Joe Crozier
相关产品推荐
相关产品推荐

