You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言dataframe过滤分组后计算用户每日消息时间差均值问题

问题描述

我有一个存储消息交互记录的大型dataframe,结构示例如下:

structure(list(from = c(1, 8, 3, 3, 8, 1, 4, 5, 8, 3, 1, 8, 4, 
1, 4, 8, 1, 4, 5, 8, 3, 1, 8, 1, 4, 8), to = c(8, 3, 8, 54, 3, 
4, 1, 6, 7, 1, 4, 3, 8, 8, 1, 3, 4, 1, 6, 7, 1, 4, 3, 8, 1, 3
), time = c(63200, 81282, 81543, 81548, 81844, 82199, 82514, 
82711, 82739, 82814, 82936, 83889, 84207, 84427, 85523, 85545, 
86883, 87187, 87701, 89004, 89619, 92662, 93384, 93443, 94042, 
94203), month = c(2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 4, 4, 4, 4, 4, 
4, 4, 4, 4, 4, 6, 6, 6, 6, 6, 6), day = c(1, 1, 1, 1, 1, 1, 1, 
1, 1, 1, 2, 2, 2, 2, 2, 2, 2, 2, 2, 3, 15, 15, 15, 15, 15, 15
)), class = "data.frame", row.names = c(NA, -26L))

需求逻辑

计算每个用户每日收发的首条与末次消息时间差的平均值,具体规则如下:

  • 按指定索引过滤数据集,保留索引出现在from列或to列的所有记录
  • 按month和day字段分组,即按自然天分组
  • 对每个分组计算当日time字段最大值减最小值,作为当日时间差
  • 排除时间差为0的记录后,对所有单日时间差求平均,最终输出每个索引对应的日均时间差

预期输出示例:

index      avg
1     1 9429.333
2     3 2590.667
3     4 1982.000
4     8 7338.000

说明:索引1的均值是2月1日差值19164、4月2日差值4251、6月15日差值4423的平均值;如索引8在4月3日差值为0,该值需排除不参与均值计算。

现有尝试代码

运行失败的版本:

dur<-function(x)max(x)-min(x)  #The function to calculate the difference. In other cases I need to use other functions of my own

#index are the Names of the indexes for which I want the calculation
index <- c(1, 3, 4, 8)
names(index) <- index

index %>%
 map_dfr(~ df %>% filter(from == .x | to == .x) %>% group_by (month,day) %>% 
     summarize(result = dur(time)) %>% 
      summarize(mdur = mean(result)) ,.id = "index")

可计算全量时间差、但无法得到日均结果的版本:

index %>% 
  map_dfr(~ df %>% 
        filter(from == .x | to == .x) %>% 
        summarize(result = dur(time)),
        .id = "index")
解决方案

原有代码缺失了过滤时间差大于0的记录的步骤,同时补充summarize的分组配置避免警告,调整后的可用代码如下:

library(purrr)
library(dplyr)

# 自定义时间差计算函数保留
dur <- function(x) max(x) - min(x)
index <- c(1, 3, 4, 8)
names(index) <- index

result <- index %>%
  map_dfr(
    ~ df %>% 
      filter(from == .x | to == .x) %>% 
      group_by(month, day) %>% 
      summarize(result = dur(time), .groups = "drop") %>% 
      filter(result > 0) %>% # 新增:过滤时间差为0的记录
      summarize(avg = mean(result)),
    .id = "index"
  )

# 可选:将index列转为数值类型匹配预期输出格式
result$index <- as.numeric(result$index)

运行上述代码得到的result与预期输出完全一致。

内容的提问来源于stack exchange,提问作者Usuario5678

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 06:06:05