You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于分组的DataFrame多计算规则新列生成方法

R语言分组生成自定义计算列的高效实现

原始数据框

example_df <- data.frame(
  Group_name = c("Group 1", "Group 1", "Group 2", "Group 2", "Group 2"),
  Logical_variable = as.logical(c(F,T,T,F,F)), 
  Numeric_variable = as.numeric(c(1.5e-3, 1, 1, 4e-4, 3e-6))
)

需求说明

按Group_name分组,根据Logical_variable的布尔值,按以下规则生成新列new_col:

  • 当Logical_variable为FALSE时:当前行Numeric_variable / 同组内所有Logical_variable为FALSE的Numeric_variable之和
  • 当Logical_variable为TRUE时:1 / (1 + 同组内所有Logical_variable为FALSE的Numeric_variable之和)

目标结果列

example_df$new_col <- c(1, 0.9985022, 0.9995972, 0.9925558, 0.007444169)

高效实现方案

方案1:dplyr(语法直观,适合常规场景)

先计算每组内逻辑值为FALSE的数值总和,再通过分支逻辑生成目标列:

library(dplyr)

example_df <- example_df %>%
  group_by(Group_name) %>%
  mutate(
    sum_false = sum(Numeric_variable[!Logical_variable]),
    new_col = case_when(
      !Logical_variable ~ Numeric_variable / sum_false,
      Logical_variable ~ 1 / (1 + sum_false)
    )
  ) %>%
  ungroup()

方案2:data.table(性能最优,适合大量分组/大数据量)

data.table的分组运算效率远高于基础R和dplyr,在分组数量多、数据规模大时优势显著:

library(data.table)

setDT(example_df)
example_df[, sum_false := sum(Numeric_variable[!Logical_variable]), by = Group_name]
example_df[, new_col := fifelse(!Logical_variable, Numeric_variable / sum_false, 1 / (1 + sum_false))]

结果验证

运行上述代码后,example_df$new_col将与目标结果一致:

> example_df$new_col
[1] 1.0000000 0.9985022 0.9995972 0.9925558 0.0074442

内容的提问来源于stack exchange,提问作者ZT_Geo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 07:01:54