You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R ggplot2绘图时fct_relevel删除NA后重排因子水平失效问题

问题原因

  • 过滤NA的写法错误:NA是R内置缺失值标记,不是字符串"NA",filter(X!="NA")无法正确识别缺失值,需替换为filter(!is.na(X))或drop_na(X)
  • 图层数据源不统一:ggplot主图用的是过滤后的数据,但geom_crossbar调用的仍是未过滤的原始数据集,包含NA分组,导致x轴因子水平被异常覆盖,重排逻辑失效
  • 过滤缺失值后没有先丢弃因子中残留的未使用水平,原X包含的NA水平过滤后仍残留在因子属性中,干扰顺序调整

修复后完整代码

library(ggplot2)
library(forcats)
library(dplyr)
library(ggpubr)

# 生成示例数据
df <- data.frame(Y = rnorm(20, -6, 1),
               X = sample(c("yes", "no", NA), 20, replace = TRUE))

# 统一预处理数据:过滤缺失值 → 丢弃因子空水平 → 重排因子顺序
df_clean <- df %>% 
  filter(!is.na(X)) %>% 
  mutate(X = droplevels(X),
         X = fct_relevel(X, "yes"))

# 所有图层统一使用预处理后的数据集
dfplot <- df_clean %>% 
  ggplot(aes(x = X, y = Y, fill = X)) +
  geom_boxplot(size = 1, width = 0.2, show.legend = FALSE, outlier.shape = NA,
               position = position_nudge(x = 0.3)) +
  geom_jitter(show.legend = TRUE, shape = 21, width = 0.2, size = 2) +
  geom_crossbar(data = df_clean %>% group_by(X) %>% summarise(mean = mean(Y), .groups = "keep"),
                aes(x = X, ymin = mean, ymax = mean, y = mean), width = 0.2, show.legend = FALSE) +
  labs(x = "", y = "%")

dfplot

内容的提问来源于stack exchange,提问作者FGP

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 19:18:03