You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按owner分组统计存在detail的distinct unit_id数量?

统计对应detail非空的唯一unit_id数量

原始数据:

owner     unit_id   detail
abc123    002NH034  94847DT
abc123    002NH034  94868DT
abc123    002NH034  94889DT
abc123    112NH035  94899DT
abc123    112NH036
abc123    112NH037  

需求:按owner分组,统计有对应非空detail值的唯一unit_id数量,预期输出:

abc123  2

你之前的代码问题在于,sum(!is.na(detail))只是累加所有非空detail的行数(这里会得到4),没有对unit_id做去重处理。下面是两种正确的dplyr实现方式:

方法1:先筛选去重再统计

df %>%
  filter(!is.na(detail)) %>%  # 过滤出detail非空的行
  distinct(owner, unit_id) %>%  # 保留唯一的owner-unit_id组合
  group_by(owner) %>%
  summarise(unit_ids_with_detail = n())

方法2:直接在summarise中计算

利用n_distinct函数,仅统计detail非空对应的unit_id唯一值数量:

df %>%
  group_by(owner) %>%
  summarise(unit_ids_with_detail = n_distinct(unit_id[!is.na(detail)]))

两种方法都会得到预期结果:

# A tibble: 1 × 2
  owner  unit_ids_with_detail
  <chr>                 <int>
1 abc123                    2

内容的提问来源于stack exchange,提问作者simplycoding

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 16:05:41