如何在R中按性别统计不同因子值并可视化?
问题描述
我有一个包含2787行、259列的样本数据,是西班牙研究机构CIS开展的关于民众对当前与6个月前经济状况感知的调查数据。其中:
- P3列有5个选项:
"better"、"worse"、"same"、"don't know"、"don't answer" - P19列代表性别,
"1"为男性,"2"为女性
我希望按性别统计不同群体对经济状况的评价分布并绘制可视化图表(统计男性和女性分别选择各选项的人数),尝试运行以下代码:
CVSPastIndividualSituationMales<- aggregate(CIS$P3 ~ CIS$P19 == "1", CIS, sum) CVSPastIndividualSituationFemales<- aggregate(CIS$P3 ~ CIS$P19 == "2", CIS, sum) CurrentVSPastIndividualSituationMales<-ggplot(CIS,mapping=aes(x=CVSPastIndividualSituationMales))+geom_bar(fill="LightGreen")+xlab("Current VS Past Individual Situation for Males") CurrentVSPastIndividualSituationFemales <-ggplot(CIS,mapping=aes(CVSPastIndividualSituationFemales))+geom_bar(fill="Green") + xlab("Current VS Past Individual Situation for Females") ggarrange(CurrentVSPastIndividualSituationMales, CVSPastIndividualSituationFemales, ncol = 1, nrow = 1)
但出现报错:
Don't know how to automatically pick scale for object of type data.frame. Defaulting to continuous. Error in `check_aesthetics()`: ! Aesthetics must be either length 1 or the same as the data (2787): x Backtrace: 1. ggpubr::ggarrange(...) 2. purrr::map(...) 3. ggpubr (local) .f(.x[[i]], ...) 4. cowplot::plot_grid(plotlist = plotlist, ...) 5. cowplot::align_plots(...) ... 16. ggplot2 (local) by_layer(function(l, d) l$compute_aesthetics(d, plot)) 17. ggplot2 (local) f(l = layers[[i]], d = data[[i]]) 18. l$compute_aesthetics(d, plot) 19. ggplot2 (local) f(..., self = self) 20. ggplot2:::check_aesthetics(evaled, n) Error in check_aesthetics(evaled, n) :
请问我哪里出错了?如何实现预期需求?另外是否有dplyr解决方案?
错误原因
- 聚合逻辑错误:用
sum处理分类变量P3完全错误,sum用于数值求和,而P3是文本分类,应该统计各选项的出现次数;同时分组条件CIS$P19 == "1"生成逻辑列,导致统计结果结构混乱。 - ggplot映射错误:ggplot中x美学应该对应数据集中的某一列,你却把整个数据框
CVSPastIndividualSituationMales赋值给x,违反了美学映射的规则,导致长度不匹配报错。
基础修复方案(无需dplyr)
先正确统计各性别下P3的选项分布,再绘图:
# 统计性别与经济状况感知的交叉表 gender_p3_table <- table(CIS$P19, CIS$P3) # 转换为数据框适配ggplot gender_p3_df <- as.data.frame(gender_p3_table) colnames(gender_p3_df) <- c("Gender", "Economic_Situation", "Count") # 替换性别编码为易懂标签 gender_p3_df$Gender <- ifelse(gender_p3_df$Gender == "1", "Male", "Female") library(ggplot2) library(ggpubr) # 方案1:绘制分面图(更简洁) ggplot(gender_p3_df, aes(x = Economic_Situation, y = Count, fill = Economic_Situation)) + geom_bar(stat = "identity") + facet_wrap(~Gender) + labs(x = "经济状况感知", y = "人数", title = "不同性别对经济状况的感知分布") + theme(axis.text.x = element_text(angle = 45, hjust = 1)) # 方案2:分开绘制两个图再拼接 male_plot <- ggplot(subset(gender_p3_df, Gender == "Male"), aes(x = Economic_Situation, y = Count)) + geom_bar(stat = "identity", fill = "LightGreen") + labs(x = "经济状况感知", y = "人数", title = "男性对经济状况的感知分布") + theme(axis.text.x = element_text(angle = 45, hjust = 1)) female_plot <- ggplot(subset(gender_p3_df, Gender == "Female"), aes(x = Economic_Situation, y = Count)) + geom_bar(stat = "identity", fill = "Green") + labs(x = "经济状况感知", y = "人数", title = "女性对经济状况的感知分布") + theme(axis.text.x = element_text(angle = 45, hjust = 1)) ggarrange(male_plot, female_plot, ncol = 2)
dplyr解决方案
用dplyr的管道语法可以更清晰地完成数据处理与统计:
library(dplyr) library(ggplot2) library(ggpubr) # 数据统计 gender_p3_summary <- CIS %>% filter(P19 %in% c("1", "2")) %>% # 筛选有效性别数据 group_by(P19, P3) %>% # 按性别和经济状况分组 summarise(Count = n(), .groups = "drop") %>% # 统计每组人数 mutate(Gender = case_when( # 替换性别编码为标签 P19 == "1" ~ "Male", P19 == "2" ~ "Female" )) %>% rename(Economic_Situation = P3) # 重命名列名 # 方案1:分面图 ggplot(gender_p3_summary, aes(x = Economic_Situation, y = Count, fill = Economic_Situation)) + geom_col() + # 等价于geom_bar(stat="identity") facet_wrap(~Gender) + labs(x = "经济状况感知", y = "人数", title = "不同性别对经济状况的感知分布") + theme(axis.text.x = element_text(angle = 45, hjust = 1)) # 方案2:分开绘图再拼接 male_plot_dplyr <- gender_p3_summary %>% filter(Gender == "Male") %>% ggplot(aes(x = Economic_Situation, y = Count)) + geom_col(fill = "LightGreen") + labs(x = "经济状况感知", y = "人数", title = "男性对经济状况的感知分布") + theme(axis.text.x = element_text(angle = 45, hjust = 1)) female_plot_dplyr <- gender_p3_summary %>% filter(Gender == "Female") %>% ggplot(aes(x = Economic_Situation, y = Count)) + geom_col(fill = "Green") + labs(x = "经济状况感知", y = "人数", title = "女性对经济状况的感知分布") + theme(axis.text.x = element_text(angle = 45, hjust = 1)) ggarrange(male_plot_dplyr, female_plot_dplyr, ncol = 2)
内容的提问来源于stack exchange,提问作者ArtUr693
相关产品推荐
相关产品推荐

