You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为R数据框每行计算指定列的加权均值?

按行计算指定列的加权均值问题解决

问题描述

现有数据集indexIPI_branques_EU_wide,需为每行新增alimentació列,该列是C10、C11、C12的加权均值,计算公式为:

alimentació = C10*[C10/(C10+C11+C12)] + C11*[C11/(C10+C11+C12)] + C12*[C12/(C10+C11+C12)]

使用dplyr的mutate结合weighted.mean函数时,得到的每行结果完全相同,不符合预期(因每行geo和time不同,对应列值不同,结果应逐行变化)。

数据集结构:

structure(list(geo = c("Alemanya", "Espanya"), time = structure(c(19417, 11961), class = "Date"), C10 = c(115.7, 103), C11 = c(104.9, 120.6), C12 = c(70.2, NA), C13 = c(98.3, 232), C14 = c(77.3, 275), C15 = c(127.6, 268.8), C16 = c(105.7, 255.7), C17 = c(91.2, 106.5), C18 = c(70.7, 151.6), C19 = c(90.3, 84.9), C20 = c(87.9, 106.5), C21 = c(135.5, 83.2), C22 = c(105.2, 126.8), C23 = c(101.8, 266.8), C24 = c(96.2, 137.5), C25 = c(113.2, 183.8), C26 = c(145, 176.7), C27 = c(126, 170.4), C28 = c(108.7, 129.1), C29 = c(104.5, 128.1), C30 = c(153.8, 177), C31 = c(94.3, 300.9), C32 = c(127.4, 138.1), C33 = c(111.9, 92.4), D35 = c(89, 96.1), E36 = c(NA_real_, NA_real_)), row.names = c(NA, -2L), class = c("tbl_df", "tbl", "data.frame"))

错误代码:

branques_Idescat_EU <- indexIPI_branques_EU_wide %>% 
  mutate(alimentació = weighted.mean(indexIPI_branques_EU_wide$C10, indexIPI_branques_EU_wide$C11, indexIPI_branques_EU_wide$C12))

问题原因

weighted.mean()函数默认对整个输入向量计算全局加权均值,而非逐行计算;同时直接引用原数据集的列(indexIPI_branques_EU_wide$C10)会让函数一次性计算整个列的均值,再将结果重复填充到每一行,导致所有行结果相同。

解决方案

方案1:直接按公式逐行计算

观察公式可简化为:(C10² + C11² + C12²) / (C10 + C11 + C12),直接在mutate中编写逐行计算的表达式,同时用rowSums简化计算,还可处理NA值:

library(dplyr)

branques_Idescat_EU <- indexIPI_branques_EU_wide %>% 
  mutate(
    sum_weights = rowSums(select(., C10, C11, C12), na.rm = TRUE),
    sum_squares = rowSums(select(., C10, C11, C12)^2, na.rm = TRUE),
    alimentació = ifelse(sum_weights == 0, NA, sum_squares / sum_weights)
  ) %>% 
  select(-sum_weights, -sum_squares) # 可选:删除中间计算列

方案2:用rowwise()实现逐行weighted.mean

如果坚持使用weighted.mean,需用rowwise()让dplyr逐行处理数据,注意处理NA值:

branques_Idescat_EU <- indexIPI_branques_EU_wide %>% 
  rowwise() %>% 
  mutate(
    alimentació = weighted.mean(
      x = c(C10, C11, C12),
      w = c(C10, C11, C12),
      na.rm = TRUE
    )
  ) %>% 
  ungroup() # 记得取消逐行分组,避免后续操作受影响

结果验证

运行上述代码后,alimentació列会逐行返回对应结果:

  • 第一行(Alemanya):(115.7² + 104.9² +70.2²)/(115.7+104.9+70.2) ≈ 100.23
  • 第二行(Espanya):因C12为NA,计算(103² +120.6²)/(103+120.6) ≈ 112.37

内容的提问来源于stack exchange,提问作者Maria

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 23:43:31