You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中按组计算列内均值与中位数并生成绩效评价列

R数据分组统计与类别标记解决方案

问题需求

针对包含5万+行、51个不同course的DataFrame,需要实现:

  • 按group列分组,计算每组score列的均值(mean)和中位数(median)
  • 为原数据的每一行添加对应分组的均值、中位数信息
  • 根据个体score与所在组的均值/中位数的比较,将个体标记为weak(对应1)、competitive(对应2)或top(对应3)

示例数据

df <- structure(list(id = 1:20, age = c(18L, 21L, 20L, 19L, 20L, 20L, 
23L, 18L, 18L, 19L, 22L, 20L, 18L, 19L, 18L, 18L, 25L, 27L, 18L, 
18L), gender = c(0L, 0L, 1L, 1L, 1L, 1L, 0L, 1L, 0L, 0L, 0L, 
0L, 1L, 0L, 1L, 0L, 1L, 0L, 1L, 1L), school = c(0L, 0L, 1L, 2L, 
2L, 2L, 1L, 1L, 1L, 1L, 0L, 0L, 2L, 1L, 0L, 2L, 0L, 1L, 1L, 2L
), score = c(3.63, 18.77, 21.76, 12.57, 20.69, 13.94, 1.25, 13.07, 
12.94, 12.63, 13.01, 14.38, 21.9, 15.13, 5.76, 11.88, 12.51, 
17.69, 4.64, 5.77), course = c("Nursing", "Engineering", "Economy", 
"Medical", "Mathematics", "Economy", "Languages", "Literature", 
"Phiysics", "Biology", "Law", "Phiysics", "Engineering", "Law", 
"Journalism", "Languages", "Accounting", "Accounting", "Medical", 
"Journalism"), group = c(1L, 2L, 4L, 1L, 2L, 4L, 5L, 5L, 2L, 
1L, 6L, 2L, 2L, 6L, 6L, 5L, 4L, 4L, 1L, 6L), wage = c(2.8, 5, 
4.5, 6, 1.8, 4.5, 2.1, 2.3, 2, 2.5, 3.8, 2, 5, 3.8, 2.75, 2.1, 
3.9, 3.9, 6, 2.75)), class = "data.frame", row.names = c(NA, 
-20L))

解决方案代码

使用dplyr包可高效完成所有操作,适配大规模数据处理:

# 若未安装dplyr,先运行 install.packages("dplyr")
library(dplyr)

# 链式操作完成所有需求
df_processed <- df %>%
  # 按group列分组
  group_by(group) %>%
  # 新增分组的均值、中位数列
  mutate(
    group_score_mean = mean(score, na.rm = TRUE),
    group_score_median = median(score, na.rm = TRUE)
  ) %>%
  # 根据score与分组统计量的比较标记等级
  mutate(
    # 示例判断逻辑,可根据需求调整:
    # weak: score < 分组中位数
    # competitive: 分组中位数 ≤ score < 分组均值
    # top: score ≥ 分组均值
    performance_level = case_when(
      score < group_score_median ~ 1,
      score >= group_score_median & score < group_score_mean ~ 2,
      score >= group_score_mean ~ 3,
      TRUE ~ NA_integer_
    ),
    # 可选:添加对应文本标签
    performance_label = case_when(
      performance_level == 1 ~ "weak",
      performance_level == 2 ~ "competitive",
      performance_level == 3 ~ "top",
      TRUE ~ NA_character_
    )
  ) %>%
  # 取消分组(后续无需分组时建议执行)
  ungroup()

# 查看处理后的数据
head(df_processed)

代码说明

  1. 分组统计:通过group_by(group)指定分组维度,mutate直接在原数据中新增统计列,无需额外合并操作,效率更高。
  2. 类别标记:使用case_when实现多条件分支判断,示例以中位数和均值为分界点划分等级,你可根据业务需求调整逻辑(比如改用三分位数、均值±标准差等)。
  3. 异常处理:TRUE ~ NA_integer_确保所有情况被覆盖,避免出现未匹配的无效结果。

内容的提问来源于stack exchange,提问作者Anna Paula Gonçalves

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 22:07:01