使用gather与ggplot保留因子水平的技术求助
解决gather后保留因子水平的问题
我明白你遇到的困扰啦——用gather()整理数据后,原本为各列设置好的完整因子水平没保留下来,导致图表只显示了有数据的选项,而不是所有预设的因子水平。咱们可以通过调整因子水平的处理逻辑来解决这个问题,让每个分面都显示完整的选项!
问题根源
当你用gather()把多列合并成key-value结构时,value列会自动只保留所有列中实际出现的因子水平,而不是每列原本设置的完整水平集合。所以我们需要在gather()之后,根据每个key(也就是原来的问题)重新为value设置对应的完整因子水平。
修改后的完整代码
首先我们先创建一个问题与对应因子水平的映射表,把每个问题的完整选项统一存起来,这样后续调用更方便:
# 定义每个问题对应的完整因子水平 level_mapping <- list( "How was the pace of the class?" = c("Too slow", "Slow", "About right", "Fast","Too fast"), "How much of the suggested reading list did you get through?" = c("1-25%", "26-50%", "51-75%", "76-90%","91-100%"), "How much of the course material was new to you?" = c("1-25%", "26-50%", "51-75%", "76-90%","91-100%"), "Considering your background, how did you find the level of the course?" = c("Very easy", "Somewhat too easy", "About right", "Somewhat challenging","Very challenging") )
然后修改你的绘图代码,关键是在gather()之后根据问题重置因子水平:
library(tidyverse) library(scales) # 用于percent()函数 # 你的可复现数据 data <- structure(list(`difficulty_pre_How was the pace of the class?` = c("Fast", "About right", "About right", "About right", "About right", "About right", "Fast"), `difficulty_pre_How much of the suggested reading list did you get through?` = c("26-50%", "51-75%", "91-100%", "76-90%", "91-100%", "51-75%", "76-90%"), `difficulty_pre_How much of the course material was new to you?` = c("76-90%", "51-75%", "76-90%", "51-75%", "76-90%", "91-100%", "51-75%" ), `difficulty_pre_Considering your background, how did you find the level of the course?` = c("Somewhat challenging", "About right", "About right", "About right", "Somewhat challenging", "Somewhat challenging", "Somewhat challenging")), .Names = c("difficulty_pre_How was the pace of the class?", "difficulty_pre_How much of the suggested reading list did you get through?", "difficulty_pre_How much of the course material was new to you?", "difficulty_pre_Considering your background, how did you find the level of the course?" ), row.names = c(NA, -7L), class = c("tbl_df", "tbl", "data.frame" )) data %>% na.omit() %>% # 先将各列转换为对应有序因子(和你原来的逻辑一致) mutate_at(vars(1), ~factor(., levels = level_mapping[["How was the pace of the class?"]], ordered = TRUE)) %>% mutate_at(vars(2:3), ~factor(., levels = level_mapping[["How much of the suggested reading list did you get through?"]], ordered = TRUE)) %>% mutate_at(vars(4), ~factor(., levels = level_mapping[["Considering your background, how did you find the level of the course?"]], ordered = TRUE)) %>% # 整理成key-value结构,保留key的因子属性 gather(key = "question", value = "response", factor_key = TRUE) %>% # 截取问题名称(去掉前缀difficulty_pre_) mutate(question = str_sub(question, start = 16)) %>% # 关键步骤:根据当前问题,为response设置完整的因子水平 mutate(response = factor(response, levels = level_mapping[[question]], ordered = TRUE)) %>% # 绘图 ggplot(aes(x = response)) + geom_bar(aes(y = (..count..)/tapply(..count.., ..PANEL.., sum)[..PANEL..])) + scale_y_continuous(labels = percent, limits = c(0, 1)) + scale_x_discrete(drop = FALSE) + # 强制显示所有因子水平,哪怕没有数据 ylab("Relative Frequencies (%)") + xlab("") + facet_wrap(~ question, scale = 'free_x') + theme_bw() + theme(axis.text.x = element_text(angle = 90, hjust = 0.9), strip.text = element_text(size = 6))
关键步骤说明
- 创建level_mapping:把每个问题的完整选项集合统一管理,既让代码更清晰,也方便后续快速调用对应的因子水平。
- gather后重置因子水平:通过
mutate(response = factor(response, levels = level_mapping[[question]], ordered = TRUE)),根据当前的问题(也就是原来的列名),为response设置对应的完整因子水平,这样每个分面的x轴就会包含该问题的所有预设选项。 - 保留
scale_x_discrete(drop=FALSE):这个参数是让空因子水平显示在图表里的关键——它强制ggplot显示所有因子水平,哪怕该水平没有对应的计数。
这样修改后,你的图表就会显示每个问题的所有预设因子水平啦,比如课程节奏分面会显示从“Too slow”到“Too fast”的所有选项,阅读完成率分面会显示所有百分比区间,完全符合你的期望!
内容的提问来源于stack exchange,提问作者Alberto Stefanelli
相关产品推荐
相关产品推荐

