R语言基于另一数据框时间区间为traffic_df新增Alert.level列的方法
实现方案
核心思路
先统一两个数据框的时间字段格式为POSIX时间类型,再通过遍历traffic_df的每个时间点,匹配Alert.Level中的时间区间得到对应预警等级,无需提前构造小时维度的关联表。
实现代码
1. 基于tidyverse的purrr map实现
# 加载依赖包 library(dplyr) library(purrr) # 统一转换为POSIX时间类型,注意时区保持一致 traffic_df <- traffic_df %>% mutate(Date_Time = as.POSIXct(Date_Time, tz = "UTC")) Alert.Level <- Alert.Level %>% mutate( Start = as.POSIXct(Start, format = "%d/%m/%Y %H:%M", tz = "UTC"), End = as.POSIXct(End, format = "%d/%m/%Y %H:%M", tz = "UTC") ) # 遍历每个时间点匹配区间 traffic_df <- traffic_df %>% mutate(Alert.Level = map_dbl(Date_Time, function(dt) { # 找到所有符合区间的下标 match_idx <- which(dt >= Alert.Level$Start & dt <= Alert.Level$End) # 存在匹配则返回第一个预警等级,无匹配返回NA ifelse(length(match_idx) > 0, Alert.Level$Alert.level[match_idx[1]], NA_real_) }))
2. 纯base R实现(无需额外安装包)
# 统一转换时间类型 traffic_df$Date_Time <- as.POSIXct(traffic_df$Date_Time, tz = "UTC") Alert.Level$Start <- as.POSIXct(Alert.Level$Start, format = "%d/%m/%Y %H:%M", tz = "UTC") Alert.Level$End <- as.POSIXct(Alert.Level$End, format = "%d/%m/%Y %H:%M", tz = "UTC") # 用sapply遍历匹配 traffic_df$Alert.Level <- sapply(traffic_df$Date_Time, function(dt) { match_idx <- which(dt >= Alert.Level$Start & dt <= Alert.Level$End) ifelse(length(match_idx) > 0, Alert.Level$Alert.level[match_idx[1]], NA) })
运行结果
最终得到的traffic_df如下:
| Date_Time | Traffic | Alert.Level |
|---|---|---|
| 2020-03-09 06:00:00 | 10 | NA |
| 2020-03-09 07:00:00 | 20 | NA |
| 2020-03-10 07:00:00 | 20 | NA |
| 2020-03-24 08:00:00 | 15 | 3 |
补充说明
如果存在多段区间重叠的情况,可自行调整匹配逻辑:比如要取最高预警等级,可将返回值改为max(Alert.Level$Alert.level[match_idx])即可。
内容的提问来源于stack exchange,提问作者csharpvsto
相关产品推荐
相关产品推荐

