You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用gather与ggplot保留因子水平的技术求助

解决gather后保留因子水平的问题

我明白你遇到的困扰啦——用gather()整理数据后,原本为各列设置好的完整因子水平没保留下来,导致图表只显示了有数据的选项,而不是所有预设的因子水平。咱们可以通过调整因子水平的处理逻辑来解决这个问题,让每个分面都显示完整的选项!

问题根源

当你用gather()把多列合并成key-value结构时,value列会自动只保留所有列中实际出现的因子水平,而不是每列原本设置的完整水平集合。所以我们需要在gather()之后,根据每个key(也就是原来的问题)重新为value设置对应的完整因子水平。

修改后的完整代码

首先我们先创建一个问题与对应因子水平的映射表,把每个问题的完整选项统一存起来,这样后续调用更方便:

# 定义每个问题对应的完整因子水平
level_mapping <- list(
  "How was the pace of the class?" = c("Too slow", "Slow", "About right", "Fast","Too fast"),
  "How much of the suggested reading list did you get through?" = c("1-25%", "26-50%", "51-75%", "76-90%","91-100%"),
  "How much of the course material was new to you?" = c("1-25%", "26-50%", "51-75%", "76-90%","91-100%"),
  "Considering your background, how did you find the level of the course?" = c("Very easy", "Somewhat too easy", "About right", "Somewhat challenging","Very challenging")
)

然后修改你的绘图代码,关键是在gather()之后根据问题重置因子水平:

library(tidyverse)
library(scales) # 用于percent()函数

# 你的可复现数据
data <- structure(list(`difficulty_pre_How was the pace of the class?` = c("Fast", "About right", "About right", "About right", "About right", "About right", "Fast"), `difficulty_pre_How much of the suggested reading list did you get through?` = c("26-50%", "51-75%", "91-100%", "76-90%", "91-100%", "51-75%", "76-90%"), `difficulty_pre_How much of the course material was new to you?` = c("76-90%", "51-75%", "76-90%", "51-75%", "76-90%", "91-100%", "51-75%" ), `difficulty_pre_Considering your background, how did you find the level of the course?` = c("Somewhat challenging", "About right", "About right", "About right", "Somewhat challenging", "Somewhat challenging", "Somewhat challenging")), .Names = c("difficulty_pre_How was the pace of the class?", "difficulty_pre_How much of the suggested reading list did you get through?", "difficulty_pre_How much of the course material was new to you?", "difficulty_pre_Considering your background, how did you find the level of the course?" ), row.names = c(NA, -7L), class = c("tbl_df", "tbl", "data.frame" ))

data %>% 
  na.omit() %>%
  # 先将各列转换为对应有序因子(和你原来的逻辑一致)
  mutate_at(vars(1), ~factor(., levels = level_mapping[["How was the pace of the class?"]], ordered = TRUE)) %>%
  mutate_at(vars(2:3), ~factor(., levels = level_mapping[["How much of the suggested reading list did you get through?"]], ordered = TRUE)) %>%
  mutate_at(vars(4), ~factor(., levels = level_mapping[["Considering your background, how did you find the level of the course?"]], ordered = TRUE)) %>%
  # 整理成key-value结构,保留key的因子属性
  gather(key = "question", value = "response", factor_key = TRUE) %>%
  # 截取问题名称(去掉前缀difficulty_pre_)
  mutate(question = str_sub(question, start = 16)) %>%
  # 关键步骤:根据当前问题,为response设置完整的因子水平
  mutate(response = factor(response, levels = level_mapping[[question]], ordered = TRUE)) %>%
  # 绘图
  ggplot(aes(x = response)) +
  geom_bar(aes(y = (..count..)/tapply(..count.., ..PANEL.., sum)[..PANEL..])) +
  scale_y_continuous(labels = percent, limits = c(0, 1)) +
  scale_x_discrete(drop = FALSE) + # 强制显示所有因子水平,哪怕没有数据
  ylab("Relative Frequencies (%)") +
  xlab("") +
  facet_wrap(~ question, scale = 'free_x') +
  theme_bw() +
  theme(axis.text.x = element_text(angle = 90, hjust = 0.9),
        strip.text = element_text(size = 6))

关键步骤说明

  1. 创建level_mapping:把每个问题的完整选项集合统一管理,既让代码更清晰,也方便后续快速调用对应的因子水平。
  2. gather后重置因子水平:通过mutate(response = factor(response, levels = level_mapping[[question]], ordered = TRUE)),根据当前的问题(也就是原来的列名),为response设置对应的完整因子水平,这样每个分面的x轴就会包含该问题的所有预设选项。
  3. 保留scale_x_discrete(drop=FALSE):这个参数是让空因子水平显示在图表里的关键——它强制ggplot显示所有因子水平,哪怕该水平没有对应的计数。

这样修改后,你的图表就会显示每个问题的所有预设因子水平啦,比如课程节奏分面会显示从“Too slow”到“Too fast”的所有选项,阅读完成率分面会显示所有百分比区间,完全符合你的期望!

内容的提问来源于stack exchange,提问作者Alberto Stefanelli

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:45:14