You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Base R与data.table规范实现分组过滤后聚合操作?

Base R与规范data.table实现分组过滤后聚合的方法

需求场景

以包含country、amount、discount列的数据集为例,需求为:

  • 按country分组
  • 过滤出每组内amount ≤ 组内median(amount)*10的行
  • 计算每组的total = sum(amount - discount)

dplyr的参考实现如下(可得到正确结果):

purchases |>
  group_by(country) |>
  filter(amount <= median(amount) * 10) |>
  summarize(total = sum(amount - discount))

Base R的正确实现

使用by()函数按组处理后,将结果转换为标准数据框格式:

# 按country分组,每组内完成过滤与聚合计算
group_calculations <- by(purchases, purchases$country, function(group_data) {
  # 过滤当前组符合条件的行
  filtered_rows <- group_data[group_data$amount <= median(group_data$amount)*10, ]
  # 计算total值
  sum(filtered_rows$amount - filtered_rows$discount)
})

# 转换为规范的数据框
base_r_result <- data.frame(
  country = names(group_calculations),
  total = as.vector(group_calculations),
  row.names = NULL
)

# 按country排序
base_r_result <- base_r_result[order(base_r_result$country), ]

规范的data.table实现

通过在分组内定义临时变量存储中位数,避免重复书写过滤条件,同时保证逻辑清晰:

library(data.table)
setDT(purchases)

dt_result <- purchases[, {
  # 计算当前组的amount中位数,仅计算一次
  group_median <- median(amount)
  # 筛选当前组符合条件的行
  filtered_data <- .SD[amount <= group_median * 10]
  # 聚合计算total
  .(total = sum(filtered_data$amount - filtered_data$discount))
}, by = country][order(country)]

另一种写法是先计算各组中位数,再通过连接匹配后过滤聚合:

dt_result <- purchases[, .(group_median = median(amount)), by = country] |>
  merge(purchases, by = "country") |>
  .[amount <= group_median * 10, .(total = sum(amount - discount)), by = country] |>
  .[order(country)]

内容的提问来源于stack exchange,提问作者user12256545

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 16:42:54