You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中按类别与组快速计算多变量均值差的方法

在R中高效计算多组均值差的实现方法

需求说明

按fruit分组,分别计算color(如yellow与green)、taste(如sweet与sour)两个类别下,size和weight的均值差,最终整理为目标格式的宽表。

方法1:使用tidyverse(易读性优先)

适合常规数据规模,代码逻辑清晰,便于维护:

# 加载包
library(tidyverse)

# 构造示例数据
df <- tibble(
  fruit = rep(c("apple", "pear"), each = 8),
  color = rep(rep(c("green", "yellow"), each = 4), 2),
  taste = rep(c("sweet", "sour", "sweet", "sour"), 4),
  size = c(5,7,6,9,6,7,4,9,3,2,9,4,8,5,7,2),
  weight = c(18,12,17,11,18,12,11,19,17,18,15,11,17,18,19,13)
)

# 批量处理多个类别
categories <- c("color", "taste")
# 定义每组的对比顺序(根据需求调整)
contrast_pairs <- list(color = c("yellow", "green"), taste = c("sweet", "sour"))

# 生成每个类别的均值差结果
result_list <- map2(categories, contrast_pairs, function(cat, pair) {
  df %>%
    group_by(fruit, .data[[cat]]) %>%
    summarize(across(c(size, weight), mean, .names = "{col}_mean"), .groups = "drop") %>%
    pivot_wider(names_from = all_of(cat), values_from = ends_with("_mean")) %>%
    mutate(
      across(ends_with("_mean"), 
             ~ .data[[paste0(substr(., 1, nchar(.)-5), "_mean_", pair[1])]] - 
               .data[[paste0(substr(., 1, nchar(.)-5), "_mean_", pair[2])]],
             .names = "{substr(.col, 1, nchar(.col)-5)}.diffmean.by{cat}")
    ) %>%
    select(fruit, ends_with(paste0(".by", cat)))
})

# 合并所有结果
final_result <- reduce(result_list, inner_join, by = "fruit")

print(final_result)

输出结果:

# A tibble: 2 × 5
  fruit size.diffmean.bycolor weight.diffmean.bycolor size.diffmean.bytaste weight.diffmean.bytaste
  <chr>                 <dbl>                    <dbl>                 <dbl>                    <dbl>
1 apple                 -0.25                      0.5                 -2.75                     -2.5
2 pear                   1.25                      1                   -0.75                      1  

方法2:使用data.table(性能优先)

适合超大规模数据集,运算速度显著优于tidyverse:

# 加载包
library(data.table)

setDT(df)

# 计算color分组的均值差
color_dt <- df[, lapply(.SD, mean), by = .(fruit, color), .SDcols = c("size", "weight")] %>%
  dcast(fruit ~ color, value.var = c("size", "weight")) %>%
  .[, `:=`(
    size.diffmean.bycolor = size_yellow - size_green,
    weight.diffmean.bycolor = weight_yellow - weight_green,
    size_green = NULL, size_yellow = NULL, weight_green = NULL, weight_yellow = NULL
  )]

# 计算taste分组的均值差
taste_dt <- df[, lapply(.SD, mean), by = .(fruit, taste), .SDcols = c("size", "weight")] %>%
  dcast(fruit ~ taste, value.var = c("size", "weight")) %>%
  .[, `:=`(
    size.diffmean.bytaste = size_sweet - size_sour,
    weight.diffmean.bytaste = weight_sweet - weight_sour,
    size_sweet = NULL, size_sour = NULL, weight_sweet = NULL, weight_sour = NULL
  )]

# 合并结果
final_dt <- merge(color_dt, taste_dt, by = "fruit")
print(final_dt)

核心逻辑说明

  1. 按fruit+类别变量分组,计算数值变量的均值
  2. 将类别变量的不同水平转成宽列,方便计算差值
  3. 按指定对比规则计算均值差
  4. 合并所有类别的结果,得到目标格式

内容的提问来源于stack exchange,提问作者David

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 14:40:57