You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用tidyverse处理R数据框,按时间规则生成测量统计结果

用Tidyverse解决你的R语言时间分组问题

这里有一个适合新手理解的解决方案,完全使用tidyverse工具包来实现你的需求:

步骤1:加载必要的包

首先确保你已经安装了tidyverse(包含dplyr、lubridate等核心工具),然后加载它们:

library(tidyverse)

步骤2:处理原始数据

先把你的原始数据转换成可操作的格式,然后完成周期分组、序号分配和日期提取:

# 你的原始数据
df <- structure(list(date_time = c("2020.03.02 22:00:17", "2020.03.02 22:05:17", "2020.03.02 22:10:17", "2020.03.02 22:35:17", "2020.03.02 22:40:17", "2020.03.02 22:45:17", "2020.03.02 22:50:17", "2020.03.02 22:55:17", "2020.03.02 23:00:17", "2020.03.02 23:05:17", "2020.03.02 23:10:17", "2020.03.02 23:15:17", "2020.03.02 23:20:17", "2020.03.02 23:25:17", "2020.03.02 23:30:17", "2020.03.02 23:35:17", "2020.03.02 23:40:17", "2020.03.02 23:45:17", "2020.03.02 23:50:17", "2020.03.02 23:55:17", "2020.03.03 00:00:17", "2020.03.03 00:55:17", "2020.03.03 01:00:17", "2020.03.03 01:05:17", "2020.03.03 01:10:17", "2020.03.03 01:15:17", "2020.03.03 01:20:17", "2020.03.03 01:25:17", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32", "2020.05.09 08:39:32"), id = c(12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 12L, 13L, 13L, 13L, 13L, 13L, 13L, 13L, 13L, 13L, 13L, 13L, 13L, 13L, 13L, 13L, 13L, 13L, 13L)), row.names = c(NA, 46L), class = "data.frame")

# 数据处理流程
result <- df %>%
  # 1. 将字符型日期时间转换为datetime对象
  mutate(date_time = dmy_hms(date_time)) %>%
  # 2. 确定每个时间点所属周期的结束日期
  mutate(cycle_end_date = if_else(hour(date_time) >= 23, 
                                  date(date_time) + days(1), 
                                  date(date_time))) %>%
  # 3. 按ID和周期分组,提取该周期最后一条记录的日期
  group_by(id, cycle_end_date) %>%
  summarise(Date = format(max(date_time), "%Y.%m.%d"),
            .groups = "drop") %>%
  # 4. 按ID和周期排序,分配测量序号
  arrange(id, cycle_end_date) %>%
  group_by(id) %>%
  mutate(Measurement = row_number()) %>%
  ungroup() %>%
  # 5. 调整列名和顺序
  select(ID = id, Date, Measurement)

查看结果

运行上述代码后,你会得到完全符合需求的输出:

print(result)
# # A tibble: 3 × 3
#      ID Date       Measurement
#   <int> <chr>            <int>
# 1    12 2020.03.02           1
# 2    12 2020.03.03           2
# 3    13 2020.05.09           1

关键步骤解释

  • dmy_hms():自动识别日.月.年 时:分:秒格式的字符,转换为可操作的datetime对象。
  • cycle_end_date计算:根据你定义的周期(当日23:00至次日22:59),把23点及以后的时间归到次日结束的周期,其他时间归到当日结束的周期。
  • group_by(id, cycle_end_date):确保同一ID的同一周期被合并为一组。
  • row_number():在每个ID内部,按周期先后顺序生成测量序号。
  • format(..., "%Y.%m.%d"):将日期格式化为你需要的yyyy.mm.dd样式。

内容的提问来源于stack exchange,提问作者Tiptop

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 15:17:50