在dplyr中如何通过列值索引外部列表并生成时间格式列
问题说明
我有一个存储分析机器设置的对象列表data.times,其中每个对象是包含时、分、秒的时间列表。另有一个directory.times数据框,包含data.times中对象的名称。需要为directory.times添加一列,当abif_type为"time"时,插入hh:mm:ss格式的时间。
示例数据
data.times <- list(RUNT.1 = list(hour = 17L, minute = 48L, second = 56L, hsecond = 0L), RUNT.2 = list(hour = 19L, minute = 0L, second = 47L, hsecond = 0L), RUNT.3 = list(hour = 18L, minute = 19L, second = 8L, hsecond = 0L), RUNT.4 = list(hour = 19L, minute = 0L, second = 48L, hsecond = 0L))
directory.times数据框示例
directory.times <- structure(list(data_label = c("RUNT.1", "RUNT.2", "RUNT.3", "RUNT.4" ), abif_type = c("time", "time", "time", "time"), description = c("Run start time", "Run stop time", "Data Collection start time", "Data Collection stop time" ), value = list(RUNT.1 = list(hour = 17L, minute = 48L, second = 56L, hsecond = 0L), RUNT.2 = list(hour = 19L, minute = 0L, second = 47L, hsecond = 0L), RUNT.3 = list(hour = 18L, minute = 19L, second = 8L, hsecond = 0L), RUNT.4 = list(hour = 19L, minute = 0L, second = 48L, hsecond = 0L))), row.names = c(NA, -4L), class = "data.frame")
单个时间提取正常运行代码
library(lubridate) hms(paste(data.times[["RUNT.1"]][["hour"]], data.times[["RUNT.1"]][["minute"]], data.times[["RUNT.1"]][["second"]], sep=":"))
报错代码及信息
尝试用dplyr实现时报错no such index at level 2,代码如下:
time.entry <- directory.times %>% mutate(time = case_when(abif_type == "time" ~ hms(paste(data.times[[paste0(data_label)]]["hour"], data.times[[paste0(data_label)]]["minute"], data.times[[paste0(data_label)]]["second"], sep="-"))))
解决方案
报错原因是data_label是向量,直接用[[索引列表时无法逐行匹配,需要用逐行处理或向量化映射的方式。
方法1:使用rowwise()逐行处理
library(dplyr) library(lubridate) time.entry <- directory.times %>% rowwise() %>% mutate( time = if_else( abif_type == "time", hms(paste(data.times[[data_label]]$hour, data.times[[data_label]]$minute, data.times[[data_label]]$second, sep = ":")), NA_character_ ) ) %>% ungroup()
方法2:使用purrr::map2()向量化映射
library(dplyr) library(lubridate) library(purrr) time.entry <- directory.times %>% mutate( time = map2_chr(data_label, abif_type, function(label, type) { if (type != "time") return(NA_character_) hms_str <- paste(data.times[[label]]$hour, data.times[[label]]$minute, data.times[[label]]$second, sep = ":") as.character(hms(hms_str)) }) )
方法3:直接利用directory.times自带的value列
注意到directory.times的value列已经包含时间数据,无需额外从data.times提取,效率更高:
library(dplyr) library(lubridate) time.entry <- directory.times %>% rowwise() %>% mutate( time = if_else( abif_type == "time", hms(paste(value[[1]]$hour, value[[1]]$minute, value[[1]]$second, sep = ":")), NA_character_ ) ) %>% ungroup()
内容的提问来源于stack exchange,提问作者Mike
相关产品推荐
相关产品推荐

