You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R统计tibble/df中按日期分组的去重用户数量实现方法

R统计每日通话去重用户数解决方案

首先为你复现示例数据代码:

n <- 100
dates <- as.Date(c("2021-01-01", "2021-01-02", "2021-01-03", "2021-01-04"))

df <- data.frame( date = sample(dates, n, replace = TRUE),
                  user = sample(LETTERS, n, replace = TRUE)
                 )

核心需求为按日期分组,对用户列做去重计数,以下是三种常用实现方式:

1. 基础R实现

无需安装额外依赖,直接用基础函数即可完成统计:

# 统计并修改列名为要求的格式
result <- aggregate(
  user ~ date, 
  df, 
  FUN = function(x) length(unique(x))
)
colnames(result) <- c("date", "number_of_users_doing_phone_calls")

# 可选操作:补全所有统计日期,无通话的日期用户数填0
result <- merge(
  data.frame(date = dates), 
  result, 
  all.x = TRUE
)
result[is.na(result)] <- 0

2. dplyr包实现(语法最简洁)

适合习惯tidyverse生态的用户,代码可读性高:

library(dplyr)

result <- df %>%
  group_by(date) %>%
  summarise(
    number_of_users_doing_phone_calls = n_distinct(user)
  ) %>%
  # 可选操作:补全所有统计日期,无通话的日期用户数填0
  complete(date = dates, fill = list(number_of_users_doing_phone_calls = 0))

3. data.table包实现(处理大数据效率最高)

适合数据量较大的场景,运行速度远高于前两种方案:

library(data.table)
dt <- as.data.table(df)

result <- dt[, .(
  number_of_users_doing_phone_calls = uniqueN(user)
), by = date]

# 可选操作:补全所有统计日期,无通话的日期用户数填0
result <- setDT(data.frame(date = dates))[result, on = "date"]
result[is.na(number_of_users_doing_phone_calls), number_of_users_doing_phone_calls := 0]

以上三种方式输出的结果都符合要求,示例输出如下:

date number_of_users_doing_phone_calls
1 2021-01-01                                 10
2 2021-01-02                                 16
3 2021-01-03                                 26
4 2021-01-04                                 20

内容的提问来源于stack exchange,提问作者D. Studer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 07:45:02