如何使用R的data.table或dplyr统计按日期累计的去重客户数量
data.table 实现
简洁写法(小数据量易理解)
# 如日期非YYYY-MM-DD格式,建议先转成Date类型保证比较逻辑正确 # dt2[, date := as.Date(date)] res_dt <- dt2[, .(date = unique(date))][ order(date), acts := sapply(date, function(cur_date) uniqueN(dt2[date <= cur_date, client])) ][]
高效写法(适合十万行以上大数据量)
res_dt <- dt2[, .(first_date = min(date)), by = client][ order(first_date), .(new_add = .N), by = first_date ][ dt2[, .(date = unique(date))], on = .(first_date = date) ][ is.na(new_add), new_add := 0 ][ order(first_date), .(date = first_date, acts = cumsum(new_add)) ][]
dplyr 实现
简洁写法
library(dplyr) # 转日期格式逻辑同上 # dt2 <- dt2 %>% mutate(date = as.Date(date)) res_dplyr <- dt2 %>% distinct(date) %>% arrange(date) %>% rowwise() %>% mutate(acts = n_distinct(dt2$client[dt2$date <= date])) %>% ungroup()
高效写法
res_dplyr <- dt2 %>% group_by(client) %>% summarise(first_date = min(date), .groups = "drop") %>% count(first_date, name = "new_add") %>% right_join(tibble(date = unique(dt2$date)), by = c("first_date" = "date")) %>% arrange(first_date) %>% mutate( new_add = tidyr::replace_na(new_add, 0), acts = cumsum(new_add) ) %>% select(date, acts)
内容的提问来源于stack exchange,提问作者Abraham Mathew
相关产品推荐
相关产品推荐

