You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于列名部分匹配分组计算R数据框的行向求和

解决列分组求和问题

正则问题修正

你之前的正则代码存在两个问题:

  • R中反向引用需要用\\1而非\1,单反斜杠会被当作转义字符解析,无法正确引用捕获组。
  • 正则模式^qw[12]_(.*)_$的结尾_$会截断列名的最后一个下划线,导致生成的新列名缺失末尾的_,不符合预期格式。

修正后的正则处理代码:

vector <- c("qw1_A_B_", "qw2_A_B_", "qw1_C_D_", "qw2_C_D_")
vector <- gsub("^qw[12](.*)$", "qw1qw2\\1", vector)

执行后会得到正确的分组列名:"qw1qw2_A_B_" "qw1qw2_A_B_" "qw1qw2_C_D_" "qw1qw2_C_D_"

完整实现方法

方法一:Base R 原生实现

直接通过列名分组后逐行求和:

# 示例数据集
df <- data.frame(qw1_A_B_ = c(2, 4, 6, 8, 10),   
                 qw2_A_B_ = c(3, 6, 9, 12, 15),   
                 qw1_C_D_ = c(5, 10, 15, 20, 25),   
                 qw2_C_D_ = c(7, 14, 21, 28, 35))

# 生成分组标识
group_ids <- gsub("^qw[12](.*)$", "qw1qw2\\1", colnames(df))

# 按分组求和并转成数据框
result_df <- as.data.frame(lapply(split.default(df, group_ids), rowSums))

方法二:tidyverse 工具链实现

如果习惯使用dplyr和tidyr,可以用长表转宽表的方式实现:

library(dplyr)
library(tidyr)

df %>%
  # 转成长格式
  pivot_longer(everything(), names_to = "col", values_to = "val") %>%
  # 生成分组列名
  mutate(group = gsub("^qw[12](.*)$", "qw1qw2\\1", col)) %>%
  # 按行和分组求和
  group_by(row_number(), group) %>%
  summarise(total = sum(val), .groups = "drop_last") %>%
  # 转回宽格式
  pivot_wider(names_from = group, values_from = total) %>%
  # 移除行号列
  select(-row_number())

两种方法最终都会得到你期望的结果:

data.frame(qw1qw2_A_B_ = c(5, 10, 15, 20, 25),   
           qw1qw2_C_D_ = c(12, 24, 36, 48, 60))

内容的提问来源于stack exchange,提问作者12666727b9

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 10:16:28