R统计tibble/df中按日期分组的去重用户数量实现方法
R统计每日通话去重用户数解决方案
首先为你复现示例数据代码:
n <- 100 dates <- as.Date(c("2021-01-01", "2021-01-02", "2021-01-03", "2021-01-04")) df <- data.frame( date = sample(dates, n, replace = TRUE), user = sample(LETTERS, n, replace = TRUE) )
核心需求为按日期分组,对用户列做去重计数,以下是三种常用实现方式:
1. 基础R实现
无需安装额外依赖,直接用基础函数即可完成统计:
# 统计并修改列名为要求的格式 result <- aggregate( user ~ date, df, FUN = function(x) length(unique(x)) ) colnames(result) <- c("date", "number_of_users_doing_phone_calls") # 可选操作:补全所有统计日期,无通话的日期用户数填0 result <- merge( data.frame(date = dates), result, all.x = TRUE ) result[is.na(result)] <- 0
2. dplyr包实现(语法最简洁)
适合习惯tidyverse生态的用户,代码可读性高:
library(dplyr) result <- df %>% group_by(date) %>% summarise( number_of_users_doing_phone_calls = n_distinct(user) ) %>% # 可选操作:补全所有统计日期,无通话的日期用户数填0 complete(date = dates, fill = list(number_of_users_doing_phone_calls = 0))
3. data.table包实现(处理大数据效率最高)
适合数据量较大的场景,运行速度远高于前两种方案:
library(data.table) dt <- as.data.table(df) result <- dt[, .( number_of_users_doing_phone_calls = uniqueN(user) ), by = date] # 可选操作:补全所有统计日期,无通话的日期用户数填0 result <- setDT(data.frame(date = dates))[result, on = "date"] result[is.na(number_of_users_doing_phone_calls), number_of_users_doing_phone_calls := 0]
以上三种方式输出的结果都符合要求,示例输出如下:
date number_of_users_doing_phone_calls 1 2021-01-01 10 2 2021-01-02 16 3 2021-01-03 26 4 2021-01-04 20
内容的提问来源于stack exchange,提问作者D. Studer
相关产品推荐
相关产品推荐

