You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用dplyr按区域统计唯一样本数量的实现方法

用dplyr按区域统计唯一样本数量的实现方法

首先是原始数据集:

df <- data.frame(strain = 1:6, sample = c("a24", "a24", "a24", "a26", "a26", "a27"), region = c(rep("ny", 3), rep("detroit",3)))

使用dplyr包的分组统计功能即可实现需求,核心是用n_distinct()函数统计唯一值数量:

library(dplyr)

# 按region分组,统计唯一sample的数量
result <- df %>%
  group_by(region) %>%
  summarize(sample_count = n_distinct(sample)) %>%
  ungroup() # 可选,取消分组状态

# 查看结果
print(result)

执行后得到的结果如下:

regionsample_count
ny1
detroit2

代码说明

  • group_by(region):将数据按region字段分组,后续统计操作会基于每个分组执行
  • summarize(sample_count = n_distinct(sample)):对每个分组,计算sample列中唯一值的个数,并将结果列命名为sample_count
  • ungroup():可选步骤,用于取消数据的分组标记,避免后续操作因分组状态产生意外行为

内容的提问来源于stack exchange,提问作者trilisser

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 16:12:04