R语言如何提取字符串中最后两个日期之间的文本内容
R实现提取字符串中指定位置日期间文本的方法
依赖包
我们使用stringr做正则匹配,dplyr做数据框处理,可直接加载tidyverse套件,也可单独安装加载对应包:
# 安装命令(执行一次即可) install.packages(c("stringr", "dplyr")) # 加载包 library(stringr) library(dplyr)
实现逻辑
- 先定义日期匹配规则:匹配
YYYY-MM-DD格式的日期 - 逐行读取文本,提取所有符合规则的日期,统计日期数量
- 按规则返回目标文本:
- 只有1个日期时,返回该日期前的全部文本,去除首尾空白
- 有2个及以上日期时,匹配倒数第二个和倒数第一个日期中间的内容,去除开头多余的标点、空白后返回
完整代码
# 示例数据 df <- data.frame (f1 = c("Today is test 2021-09-15", "This is to be done today. 2020-04-05.Today is going tobe 2021-09-15", "Great Novel. 2018-08-09.This is to be done today. 2020-04-05.The lion is an animal 2021-09-15", "This is to be done today. 2020-04-05.Today is test 2021-09-01.Monday is the first day 2021-08-02" ) ) # 定义日期匹配正则 date_pattern <- "\\d{4}-\\d{2}-\\d{2}" # 处理数据 df_result <- df %>% rowwise() %>% # 逐行处理 mutate( date_list = list(str_extract_all(f1, date_pattern, simplify = TRUE)), date_num = length(date_list), target_text = case_when( date_num == 1 ~ str_remove(f1, date_pattern) %>% str_trim(), date_num >= 2 ~ { last_two_date <- tail(date_list, 2) # 匹配两个日期中间的内容 match_res <- str_match(f1, paste0(last_two_date[1], "([\\s\\S]*?)", last_two_date[2]))[,2] # 清理开头多余的标点和空白,再去除首尾空白 str_remove(match_res, "^[. ]+") %>% str_trim() } ) ) %>% ungroup() %>% select(f1, target_text) # 查看输出结果 print(df_result$target_text)
输出效果
运行后得到的target_text列结果如下,完全符合预期:
[1] "Today is test" "Today is going tobe" [3] "The lion is an animal" "Monday is the first day"
注:第二个结果中的
tobe为原文本输入拼写问题,若需统一修正为to be,可在清理步骤中添加str_replace("tobe", "to be")即可。
内容的提问来源于stack exchange,提问作者R Ban
相关产品推荐
相关产品推荐

