You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中按ID分组批量汇总求和以cat_开头的变量?

批量对前缀为cat_的变量按ID分组求和

问题描述

实际数据集比示例更复杂,需要在R中按ID分组后,对所有以cat_为前缀的变量进行求和汇总,目前只能逐个指定变量处理,求高效解决方案。

示例数据集:

df <- structure(list(ID = c("A", "B", "C", "D", "A", "B", "C", "D", 
"A", "B", "C", "D"), year = c(1900, 1900, 1900, 1900, 1901, 1901, 
1901, 1901, 1902, 1902, 1902, 1902), val = c(2635L, 8573L, 5942L, 
7390L, 8762L, 7871L, 7848L, 1928L, 6772L, 6487L, 6005L, 5341L
), cat_TS = c(1L, 0L, 0L, 1L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L), 
    cat_1 = c(0L, 0L, 0L, 0L, 1L, 1L, 0L, 0L, 0L, 0L, 0L, 0L), 
    cat_2 = c(0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 1L, 0L, 0L, 0L)), row.names = c(NA, 
-12L), class = c("tbl_df", "tbl", "data.frame"))

原手动处理代码:

df <- df %>% group_by(ID) %>% 
  summarise(cat_TS = sum(cat_TS), cat_1 = sum(cat_1), cat_2 = sum(cat_2))

解决方案

使用dplyr包的across()函数配合starts_with()选择器,实现批量处理所有cat_前缀的变量,无需逐个指定:

# 先加载dplyr(未安装则先执行 install.packages("dplyr"))
library(dplyr)

# 批量分组求和
df_summary <- df %>%
  group_by(ID) %>%
  summarise(across(starts_with("cat_"), sum))

代码说明

  • starts_with("cat_"):自动匹配所有列名以cat_开头的变量,无论数量多少
  • across():将sum函数应用到所有选中的变量上
  • 执行后会得到按ID分组,每个cat_变量求和后的结果

内容的提问来源于stack exchange,提问作者tchoup

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 21:20:43