You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于dplyr的可灵活分组(含无分组)的R汇总函数实现需求

Tidy评估核心原理说明

dplyr的核心函数(如group_by、mutate)默认使用非标准求值(NSE):它们直接识别代码中的变量名,而非字符串。要让函数接收外部传入的字符向量作为分组变量,需要两个关键操作:

  1. 符号转换:rlang::syms()将字符向量转换为「符号对象」——相当于把字符串"homeworld"转换成代码里的homeworld,让dplyr能识别为变量。
  2. 拼接展开:!!!(反引号拼接)把符号对象列表展开到group_by的参数中。如果传入空向量,!!!会展开为空,此时group_by()等价于不分组,直接对全数据集汇总。

summarise中的.groups = "drop"是可选配置,作用是汇总后自动取消分组状态,避免后续操作因分组残留出现意外行为。


汇总结果绘图示例

按母星的BMI均值条形图

# 过滤无有效BMI数据的母星
homeworld_BMI_filtered <- homeworld_BMI %>% 
  filter(!is.na(BMI_mean)) %>% 
  arrange(desc(BMI_mean))

ggplot(homeworld_BMI_filtered, aes(x = reorder(homeworld, BMI_mean), y = BMI_mean)) +
  geom_bar(stat = "identity", fill = "#2E86AB") +
  coord_flip() +
  labs(
    title = "各母星生物BMI均值",
    x = "母星",
    y = "BMI均值"
  ) +
  theme_minimal()

按母星+物种的分组BMI均值图

# 取BMI均值最高的10组数据
species_homeworld_BMI_filtered <- species_homeworld_BMI %>% 
  filter(!is.na(BMI_mean)) %>% 
  slice_max(BMI_mean, n = 10)

ggplot(species_homeworld_BMI_filtered, aes(x = homeworld, y = BMI_mean, fill = species)) +
  geom_bar(stat = "identity", position = "dodge") +
  labs(
    title = "部分母星不同物种的BMI均值",
    x = "母星",
    y = "BMI均值",
    fill = "物种"
  ) +
  theme_minimal() +
  theme(axis.text.x = element_text(angle = 45, hjust = 1))

内容的提问来源于stack exchange,提问作者AdrianD

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 19:05:41