使用tidyverse处理R数据框,按时间规则生成测量统计结果
用Tidyverse解决你的R语言时间分组问题
这里有一个适合新手理解的解决方案,完全使用tidyverse工具包来实现你的需求:
步骤1:加载必要的包
首先确保你已经安装了tidyverse(包含dplyr、lubridate等核心工具),然后加载它们:
library(tidyverse)
步骤2:处理原始数据
先把你的原始数据转换成可操作的格式,然后完成周期分组、序号分配和日期提取:
# 你的原始数据 df <- structure(list(date_time = c("2020.03.02 22:00:17", "2020.03.02 22:05:17", "2020.03.02 22:10:17", "2020.03.02 22:35:17", "2020.03.02 22:40:17", "2020.03.02 22:45:17", "2020.03.02 22:50:17", "2020.03.02 22:55:17", "2020.03.02 23:00:17", "2020.03.02 23:05:17", "2020.03.02 23:10:17", "2020.03.02 23:15:17", "2020.03.02 23:20:17", "2020.03.02 23:25:17", "2020.03.02 23:30:17", "2020.03.02 23:35:17", "2020.03.02 23:40:17", "2020.03.02 23:45:17", "2020.03.02 23:50:17", "2020.03.02 23:55:17", "2020.03.03 00:00:17", "2020.03.03 00:55:17", "2020.03.03 01:00:17", "2020.03.03 01:05:17", "2020.03.03 01:10:17", "2020.03.03 01:15:17", "2020.03.03 01:20:17", "2020.03.03 01:25:17", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32"), id = c(12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 13L, 13L, 13L, 13L, 13L, 13L, 13L, 13L, 13L, 13L, 13L, 13L, 13L, 13L, 13L, 13L, 13L, 13L)), row.names = c(NA, 46L), class = "data.frame") # 数据处理流程 result <- df %>% # 1. 将字符型日期时间转换为datetime对象 mutate(date_time = dmy_hms(date_time)) %>% # 2. 确定每个时间点所属周期的结束日期 mutate(cycle_end_date = if_else(hour(date_time) >= 23, date(date_time) + days(1), date(date_time))) %>% # 3. 按ID和周期分组,提取该周期最后一条记录的日期 group_by(id, cycle_end_date) %>% summarise(Date = format(max(date_time), "%Y.%m.%d"), .groups = "drop") %>% # 4. 按ID和周期排序,分配测量序号 arrange(id, cycle_end_date) %>% group_by(id) %>% mutate(Measurement = row_number()) %>% ungroup() %>% # 5. 调整列名和顺序 select(ID = id, Date, Measurement)
查看结果
运行上述代码后,你会得到完全符合需求的输出:
print(result) # # A tibble: 3 × 3 # ID Date Measurement # <int> <chr> <int> # 1 12 2020.03.02 1 # 2 12 2020.03.03 2 # 3 13 2020.05.09 1
关键步骤解释
dmy_hms():自动识别日.月.年 时:分:秒格式的字符,转换为可操作的datetime对象。cycle_end_date计算:根据你定义的周期(当日23:00至次日22:59),把23点及以后的时间归到次日结束的周期,其他时间归到当日结束的周期。group_by(id, cycle_end_date):确保同一ID的同一周期被合并为一组。row_number():在每个ID内部,按周期先后顺序生成测量序号。format(..., "%Y.%m.%d"):将日期格式化为你需要的yyyy.mm.dd样式。
内容的提问来源于stack exchange,提问作者Tiptop
相关产品推荐
相关产品推荐

