在R中按回答占比排序参与者并绘制ggplot柱状图
问题描述
我有一个数据集,其中Participant是实验参与者的分类名称(每个参与者对应多行重复记录),Answer是二元变量(记录任务回答)。数据示例如下:
| Participant | Answer |
|---|---|
| P1 | Yes |
| P1 | Yes |
| P1 | Yes |
| P1 | No |
| P2 | Yes |
| P2 | No |
| P2 | No |
| P3 | No |
需要用R语言完成以下操作:
- 计算每个参与者的
Yes回答占比 - 按占比对参与者排序(占比最高在左,最低在右)
- 用ggplot绘制柱状图
尝试过两段代码均失败:
- 第一段代码:
data_df %>% group_by(Participant,Answer) %>% mutate(proportion = summarize(n = n()))
错误信息:
Error: Problem with
mutate()columnproportion.
ℹproportion = summarize(n = n()).
x no applicable method for 'summarise' applied to an object of class "c('integer', 'numeric')"
ℹ The error occurred in group 1: participant = "UIT001", answer = "Yes".
str(data_df)显示两个变量均为chr类型,但不清楚错误中整数/数值类型的来源。
- 第二段代码:
data_df <- data_df %>% group_by(participant) %>% mutate(Yes_ratio = n_distinct(participant[answer == "yes"])/n_distinct(participant))
结果每行都返回1,不符合预期。
解决方案
错误原因分析
- 第一段代码错误:
mutate()用于给原数据集添加列,而summarize()是聚合数据生成汇总表,不能嵌套使用;同时按Participant和Answer分组,无法直接得到单个参与者的Yes占比。 - 第二段代码错误:
n_distinct(participant[answer == "yes"])统计的是符合条件的参与者唯一值数量(每个分组仅对应一个参与者,结果永远为1),除以n_distinct(participant)(同样为1),最终结果必然全为1。实际需要统计的是Yes的回答次数,而非参与者的唯一值数量。
正确实现步骤
1. 计算每个参与者的Yes占比
使用dplyr聚合数据,按Participant分组后,计算Yes回答次数与总回答数的比值:
library(dplyr) # 生成占比汇总表 summary_df <- data_df %>% group_by(Participant) %>% summarize( Yes_ratio = sum(Answer == "Yes") / n(), # Yes次数/总回答次数 .groups = "drop" # 取消分组 )
2. 按占比排序并转换为有序因子
为保证ggplot绘图时按占比从高到低排列,需将Participant转换为有序因子:
summary_df <- summary_df %>% arrange(desc(Yes_ratio)) %>% # 按Yes占比降序排列 mutate(Participant = factor(Participant, levels = Participant)) # 锁定排序顺序
3. 绘制柱状图
使用ggplot2绘制柱状图,并将Y轴格式化为百分比:
library(ggplot2) library(scales) ggplot(summary_df, aes(x = Participant, y = Yes_ratio)) + geom_col(fill = "#2E86AB") + # 设置柱状图填充色 scale_y_continuous(labels = percent_format()) + # Y轴显示百分比 labs( title = "参与者Yes回答占比", x = "参与者", y = "Yes回答占比" ) + theme_minimal()
完整可运行代码
library(dplyr) library(ggplot2) library(scales) # 示例数据集(可替换为你的实际数据) data_df <- tibble( Participant = c("P1", "P1", "P1", "P1", "P2", "P2", "P2", "P3"), Answer = c("Yes", "Yes", "Yes", "No", "Yes", "No", "No", "No") ) # 计算占比并排序 summary_df <- data_df %>% group_by(Participant) %>% summarize( Yes_ratio = sum(Answer == "Yes") / n(), .groups = "drop" ) %>% arrange(desc(Yes_ratio)) %>% mutate(Participant = factor(Participant, levels = Participant)) # 绘制柱状图 ggplot(summary_df, aes(x = Participant, y = Yes_ratio)) + geom_col(fill = "#2E86AB") + scale_y_continuous(labels = percent_format()) + labs(title = "参与者Yes回答占比", x = "参与者", y = "Yes回答占比") + theme_minimal()
内容的提问来源于stack exchange,提问作者Paulina
相关产品推荐
相关产品推荐

