使用ggplot2的facet_grid分面时出现行缺失问题的解决求助
问题描述
需要绘制20名参与者(分为a、b两组各10人)40天内三类事件(missing、no、yes)的数量分布。未分面时绘图显示正常,但使用facet_grid按group分面后出现行缺失情况。
示例代码:
library(dplyr) library(tidyr) library(ggplot2) set.seed(1234) d <- data.frame(id = NULL, present = NULL) # create data. 20 ids with forty obs each for(i in 0:19) { d[i*40+(1:40), "present"] <- sample(x = c("no","yes","missing"), size = 40, prob = c(0.25, 0.4, 0.35), replace = T) d[i*40+(1:40), "id"] <- rep(paste0("id0",1+i),40) } # add group d$group <- factor(rep(letters[1:2],each=400)) # now get summaries of proportions of each of the three variables, then uncount to recreate the individual rows but this time sorted d %>% group_by(id, group, present) %>% count %>% uncount(weights = n) %>% ungroup %>% tibble::add_column(obsNum = rep(1:40,times = 20)) -> d # now rank d %>% group_by(id) %>% add_count(wt = present == "missing", # this counts the number of missing for each person and adds that count as a column name = "miss_n") %>% add_count(wt = present == "no", # does the same thing as above but for number of no values name = "no_n") %>% ungroup %>% arrange(-miss_n, -no_n, present) %>% mutate(id = factor(id) %>% fct_inorder()) %>% mutate(obsNum = row_number(), # adds a new set of row numbers within each id numbers. .by = id) -> d_ranked
未分面绘图代码(效果正常):
d_ranked %>% ggplot(mapping = aes(x = obsNum, y = id, colour = present)) + geom_point(shape = 15) + scale_colour_manual(values = c("yellow", "red", "black")) + labs(x = "observation number")
分面后出现行缺失的代码:
d_ranked %>% ggplot(mapping = aes(x = obsNum, y = id, colour = present)) + geom_point(shape = 15) + facet_grid(.~group) + scale_colour_manual(values = c("yellow", "red", "black")) + labs(x = "observation number")
问题原因
你的id因子是全局排序后生成的(按所有参与者的miss_n和no_n排序),导致a、b两组的id在因子水平中交错分布。当使用facet_grid分面后,每个面板仅包含对应组的id数据,但y轴仍会显示所有20个id的因子水平,不属于当前组的id位置就会呈现“空行”,看起来像是行缺失。
解决方案
方式1:按组内排序生成id因子
修改数据处理的排序步骤,在group分组内对id进行排序,这样每个组的id在因子水平中是连续的,分面后每个面板的y轴只会显示当前组的id(无空行):
d %>% group_by(id, group, present) %>% count %>% uncount(weights = n) %>% ungroup %>% tibble::add_column(obsNum = rep(1:40,times = 20)) -> d # 修改排序逻辑:按group分组后,在组内排序id d %>% group_by(id) %>% add_count(wt = present == "missing", name = "miss_n") %>% add_count(wt = present == "no", name = "no_n") %>% ungroup %>% # 先按group分组,再组内按-miss_n、-no_n排序 arrange(group, -miss_n, -no_n, present) %>% # 按排序后的顺序生成因子,组内id连续 mutate(id = factor(id) %>% fct_inorder()) %>% mutate(obsNum = row_number(), .by = id) -> d_ranked
使用原分面代码绘图后,每个面板会显示对应组的10个连续id,无空行。
方式2:分面时适配y轴显示
如果你希望每个面板仅显示当前组的id且无空行,也可以不修改数据,直接在分面时设置scales = "free_y"和space = "free_y":
d_ranked %>% ggplot(mapping = aes(x = obsNum, y = id, colour = present)) + geom_point(shape = 15) + facet_grid(.~group, scales = "free_y", space = "free_y") + scale_colour_manual(values = c("yellow", "red", "black")) + labs(x = "observation number")
scales = "free_y"让每个面板的y轴只显示有数据的id,space = "free_y"让每个面板的y轴高度适配id数量,避免空行占用空间。
内容的提问来源于stack exchange,提问作者llewmills
相关产品推荐
相关产品推荐

