You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Tidyverse实现多列按行计算乘积和的通用优雅方案

问题描述

需要实现不依赖列位置、可通用推广的Tidyverse方案,按行计算m组共n个列的乘积之和:即同组的列逐行求乘积,再将所有组的乘积相加得到每行的结果。
此前尝试用purrr::pmap_dbl(select(., ends_with(i)), prod)实现该逻辑,未得到符合预期的结果。

示例场景(m=3组、每组2列)

示例数据构造:

library(tidyverse)

df <- tibble(
  x_0 = c(5,6),
  x_1 = c(9,1),
  x_2 = c(2,1),
  y_0 = c(3,2),
  y_1 = c(3,2),
  y_2 = c(1,3)
)
df

数据预览:

# A tibble: 2 × 6
    x_0   x_1   x_2   y_0   y_1   y_2
  <dbl> <dbl> <dbl> <dbl> <dbl> <dbl>
1     5     9     2     3     3     1
2     6     1     1     2     2     3

计算规则:同下标的x列和y列相乘,再将所有乘积相加,公式为sum_of_products = x_0*y_0 + x_1*y_1 + x_2*y_2(原提问公式笔误写成了x_2 + y_2)。

  • 第二行计算验证:6*2 + 1*2 +1*3 = 17,和预期一致
  • 第一行按示例数据计算应为5*3 +9*3 +2*1 = 44,原提问写的46为笔误

预期输出:

# A tibble: 2 × 7
    x_0   x_1   x_2   y_0   y_1   y_2 sum_of_products
  <dbl> <dbl> <dbl> <dbl> <dbl> <dbl>           <dbl>
1     5     9     2     3     3     1              44
2     6     1     1     2     2     3              17
通用解决方案

以下方案完全不依赖列排列顺序,支持任意组数、任意每组列数(只要列名遵循前缀_组序号的命名规则,同组序号的列归为一组求乘积),可直接推广到其他同类场景。

方法1:长表转换法(易读、适合大数据量)

核心逻辑是先将宽表转为长表,拆分列名得到前缀和组序号,按行+组序号分组求乘积,再聚合求和后合并回原表:

df |> 
  # 给每行加唯一标识用于后续合并
  mutate(row_id = row_number()) |> 
  # 宽转长,拆分列名为前缀、组序号两部分
  pivot_longer(cols = -row_id, names_to = c("prefix", "group_idx"), names_sep = "_") |> 
  # 按行+组分组,计算每组的乘积
  group_by(row_id, group_idx) |> 
  summarise(group_prod = prod(value), .groups = "drop") |> 
  # 按行分组,计算所有组乘积的和
  group_by(row_id) |> 
  summarise(sum_of_products = sum(group_prod), .groups = "drop") |> 
  # 合并回原表
  left_join(mutate(df, row_id = row_number()), ., by = "row_id") |> 
  select(-row_id)

方法2:行内计算法(代码更简洁)

基于rowwise和列名匹配直接在每行内计算,代码更短:

df |> 
  rowwise() |> 
  mutate(sum_of_products = {
    # 提取当前行所有列名
    all_cols <- names(cur_data())
    # 提取所有不重复的组序号
    group_ids <- str_extract(all_cols, "(?<=_)\\d+$") |> unique()
    # 逐组求乘积再求和
    map_dbl(group_ids, ~prod(c_across(ends_with(str_c("_", .x))))) |> sum()
  }) |> 
  ungroup()
原尝试失败原因

之前用pmap_dbl的方案存在两个问题:

  • pmap默认按列位置对齐传入参数,一旦列顺序调整计算结果就会出错,不符合不依赖列位置的要求
  • 循环取列时没有做完整的列名后缀匹配,容易匹配到不符合要求的列,导致乘积计算错误

内容的提问来源于stack exchange,提问作者chamaoskurumi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 22:36:22