如何用dplyr将含因子的R数据框转换为占比统计表格?
解决方案:使用dplyr生成年份-项目的占比表格
测试数据生成
先运行以下代码生成测试数据:
library(tibble) library(dplyr) library(tidyr) A = sample(c('1-Strongly disagree','2-Disagree','3-so-so','4-Agree','5-Strongly agree'),12933,replace = TRUE) B = sample(c('1-Strongly disagree','2-Disagree','3-so-so','4-Agree','5-Strongly agree'),12933,replace = TRUE) C = sample(c('1-Strongly disagree','2-Disagree','3-so-so','4-Agree','5-Strongly agree'),12933,replace = TRUE) Year = sample(c("2019","2020","2022","2023"),12933,replace = TRUE) df = tibble(A,B,C,Year)%>% mutate(across(everything(),as.factor))
转换为目标格式的代码
结合dplyr和tidyr完成格式转换,代码如下:
result_df <- df %>% # 宽转长:提取项目(Item)和对应选项(Response) pivot_longer(cols = c(A, B, C), names_to = "Item", values_to = "Response") %>% # 分组统计并计算占比 group_by(Year, Item, Response) %>% summarise(Count = n(), .groups = "drop_last") %>% mutate(Percentage = round((Count / sum(Count)) * 100, 2)) %>% select(-Count) %>% # 长转宽:将选项转为列,填充占比 pivot_wider(names_from = Response, values_from = Percentage, values_fill = 0) %>% # 重命名列匹配目标格式 rename(Group = Year) # 查看结果 head(result_df)
代码说明
- pivot_longer:把原数据中A/B/C三列转换为
Item(存储A/B/C)和Response(存储对应选项)的长格式,便于分组统计。 - group_by + summarise:按年份和项目分组,统计每个选项的出现次数,再计算该选项在组内的占比(保留两位小数)。
- pivot_wider:将长格式转回宽格式,把每个选项作为列名,对应单元格填充占比;
values_fill=0确保所有选项都能显示(即使某组无该选项也填充0)。 - rename:将
Year列重命名为Group,匹配目标表格的列名。
示例输出
转换后的表格结构如下(示例数据):
| Group | Item | 1-Strongly disagree | 2-Disagree | 3-so-so | 4-Agree | 5-Strongly agree |
|---|---|---|---|---|---|---|
| 2019 | A | 20.12 | 20.05 | 19.87 | 20.01 | 19.95 |
| 2019 | B | 19.98 | 20.10 | 20.03 | 19.97 | 19.92 |
| 2019 | C | 20.04 | 19.96 | 20.02 | 20.05 | 19.93 |
内容的提问来源于stack exchange,提问作者Homer Jay Simpson
相关产品推荐
相关产品推荐

