You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中编写适配不同数据集变量的中位数函数用于ggplot绘图

解决方案

你遇到的核心问题是tidyverse系列函数采用非标准评估机制,直接将函数参数作为列名传入时,会被解析为字面量而非数据框的实际列,无法正确计算统计值。我们可以用tidyverse推荐的{{ }}(curly-curly)语法传递变量参数,修正后的完整代码如下:

前置依赖

首先确保加载所需包:

library(dplyr)
library(ggplot2)
library(forcats)

自定义绘图函数

bar_plot <- function(item, 
                     target_center, 
                     df = df5, 
                     ref_center = 713,
                     y_max = 20) {
  # 按中心分组计算目标变量的中位数
  a <- df %>% 
    group_by(center) %>%
    summarise(med_x = median({{item}}, na.rm = TRUE)) %>% 
    ungroup() %>%
    mutate(center = fct_reorder(factor(center), med_x))
  
  # 提取参考中心的中位数
  ref_med <- a$med_x[a$center == ref_center]
  
  # 生成标签:仅标注目标中心和参考中心
  a1 <- a %>% 
    mutate(my_label = ifelse(center %in% c(target_center, ref_center),
                             paste(center, med_x, sep = ":"), NA_character_))
  
  # 绘图
  ggplot(a1, aes(x = center, y = med_x,
                 fill = case_when(
                   center == target_center ~ "target",
                   center == ref_center ~ "Reference",
                   TRUE ~ "all"
                 ))) +
    geom_bar(stat = "identity") +
    scale_fill_manual(name = "center", 
                      values = c("target" = "gold", 
                                 "Reference" = "cadetblue", 
                                 "all" = "orange")) +
    xlab("TitelX") +
    ylab("Median") +
    ggtitle("Titelgraph") +
    geom_hline(yintercept = ref_med, color = "black", linetype = "dashed") +
    ylim(0, y_max) +
    theme(axis.text.x = element_blank(), 
          axis.ticks.x = element_blank(),
          legend.position = "none") +
    geom_label(aes(label = my_label), vjust = -0.1, na.rm = TRUE)
}

调用示例

直接传入变量名和目标中心即可生成对应图表:

# 生成变量P54a、目标中心206的图表
bar_plot(item = P54a, target_center = 206)

# 生成变量P79、目标中心206的图表,因P79最大值超过20调整y轴上限
bar_plot(item = P79, target_center = 206, y_max = 30)

批量计算中位数的方案

你之前的循环出错有两个原因:一是数据框包含非数值的分类列center,直接计算中位数无意义;二是不需要手写循环,用dplyr的across语法即可批量计算所有数值变量的中位数:

df5 %>% 
  summarise(across(where(is.numeric), ~median(.x, na.rm = TRUE)))

内容的提问来源于stack exchange,提问作者Sunshine_student

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 19:39:01