使用ggplot绘制按会议年份分面的二分类变量散点图求助
解决方案
你需要先对原始数据集做汇总统计,计算每个unit对应每个会议年份的女性占比,再进行绘图即可实现预期效果,完整代码如下:
步骤1:加载依赖包
library(dplyr) library(ggplot2)
步骤2:导入数据集
df <- structure(list(gender = c("Male", "Male", "Female", "Male", "Female", "Female", "Male", "Female", "Female", "Unknown"), race_ethnicity = c("Latino or Hispanic American", "Black, Afro-Caribbean, or African American", "Latino or Hispanic American", "East Asian or Asian American", "Latino or Hispanic American", "Non-Hispanic White or Euro-American", "Non-Hispanic White or Euro-American", "Non-Hispanic White or Euro-American", "Non-Hispanic White or Euro-American", "No Response"), year_of_birth = c("1979", "1976", "1981", "1977", "1985", "No Response", "No Response", "1961", "1978", "No Response" ), primary_field = c("American Politics", "American Politics", "American Politics", "American Politics", "American Politics", "American Politics", "American Politics", "American Politics", "International Politics", "No Response"), role_s = c("Chair Presenter Author", "Discussant", "Author", "Author", "Author", "Discussant", "Chair", "Discussant", "Author", "Author"), unit = c("Elections, Public Opinion, and Voting Behavior", "Elections, Public Opinion, and Voting Behavior", "Elections, Public Opinion, and Voting Behavior", "Elections, Public Opinion, and Voting Behavior", "Elections, Public Opinion, and Voting Behavior", "Political Communication", "Political Communication", "Political Communication", "Political Communication", "Political Communication"), conference_year = c(2017L, 2017L, 2017L, 2017L, 2017L, 2017L, 2017L, 2017L, 2017L, 2017L )), row.names = c(NA, 10L), class = "data.frame")
步骤3:汇总计算各分组女性占比
df_summary <- df %>% filter(gender %in% c("Male", "Female")) %>% # 可根据需求调整是否排除性别未知的样本 group_by(conference_year, unit) %>% summarise( female_pct = sum(gender == "Female") / n(), .groups = "drop" )
步骤4:绘制分面板散点图
ggplot(df_summary, aes(x = female_pct, y = unit)) + geom_point(size = 3, color = "#2c3e50") + facet_wrap(~conference_year) + # 按会议年份自动拆分面板 labs( x = "女性占比", y = "Unit" ) + theme_bw()
你之前直接筛选女性子集绘图失败的核心原因是没有做分组聚合:原始数据为单条人员记录,无法直接对应到单位维度的女性占比指标。如果你的全量数据集包含多个会议年份,代码会自动生成对应数量的展示面板。
内容的提问来源于stack exchange,提问作者Priana
相关产品推荐
相关产品推荐

