在R中如何提取字符串部分内容并转换为lubridate日期时间对象
解决方案
首先加载所需依赖包:
# 用到dplyr做数据处理、stringr做字符串提取、lubridate做时间转换 library(dplyr) library(stringr) library(lubridate)
核心处理逻辑:从timestamp字符串中提取所有数值,保留前6位(对应年、月、日、时、分、秒),传入make_datetime()完成转换即可,完整代码如下:
example_processed <- example %>% mutate( # 提取字符串中所有数字转为数值列表 time_parts = str_extract_all(timestamp, "\\d+", simplify = FALSE) %>% lapply(as.integer), # 取前6个值生成lubridate日期时间对象 timestamp = map_dbl(time_parts, ~make_datetime( year = .x[1], month = .x[2], day = .x[3], hour = .x[4], min = .x[5], sec = .x[6] )) %>% as_datetime(), # 删掉中间辅助列 time_parts = NULL )
你可以通过class(example_processed$timestamp)验证转换结果,输出为"POSIXct" "POSIXt",就是lubridate标准的日期时间格式。
如果想要更简洁的写法,也可以直接通过正则匹配前6组数字:
example %>% mutate( timestamp = str_match(timestamp, "\\[(\\d+), (\\d+), (\\d+), (\\d+), (\\d+), (\\d+)") %>% as.data.frame() %>% {make_datetime(.[[2]], .[[3]], .[[4]], .[[5]], .[[6]], .[[7]])} )
内容的提问来源于stack exchange,提问作者pgcudahy
相关产品推荐
相关产品推荐

