You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用tidyverse按名称与抽取次数分组计算样本均值?

使用tidyverse按名称和抽取次数计算数值均值

现有一个包含以下字段的数据集:

  • nms:名称
  • draw:抽取次数
  • id:ID
  • val:数值

每个名称对应1至5个ID,每个ID拥有5个按抽取次数排序的数值。数据集可通过以下代码生成:

library(tidyverse)

set.seed(190)

dplyr::tibble(nms = c(rep("A", 5), rep("B",10), rep("C", 10)),
              draw = c(rep(1:5, times = 5)),
              id = c(rep(c("A1", "B1", "B2", "C1", "C2"), each = 5)),
              val = c(sample(1:4, replace = T, size = 25))) %>%  print(n=25)

需求

针对每个名称(nms),按抽取次数(draw)分组计算数值(val)的均值(mean_val):

  • 若名称仅对应单个ID,均值即为该ID的对应数值
  • 若名称对应多个ID,需对相同抽取次数下的所有ID数值求平均

解决方案

使用tidyverse的dplyr工具,通过分组汇总即可实现需求:

# 承接上述生成的数据集,假设数据集名为df
df %>%
  group_by(nms, draw) %>%
  summarise(mean_val = mean(val), .groups = "drop")

输出结果

# A tibble: 15 × 3
   nms    draw mean_val
   <chr> <int>    <dbl>
 1 A         1      3  
 2 A         2      4  
 3 A         3      2  
 4 A         4      4  
 5 A         5      1  
 6 B         1      1.5
 7 B         2      1.5
 8 B         3      2.5
 9 B         4      2  
10 B         5      2  
11 C         1      1.5
12 C         2      2.5
13 C         3      2.5
14 C         4      2.5
15 C         5      4  

内容的提问来源于stack exchange,提问作者jnat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 23:55:38