如何在R语言中提取Date:、City:、Note:后的对应内容?
问题描述
现有如下CSV格式的表格数据:
No, Memo 1, Date: 2020/10/22 City: UA Note: True mastery of any skill takes a lifetime. 2, Date: 2022/11/01 City: CH Note: Sweat is the lubricant of success. 3, Date: 2022y11m1d City: UA Note: Every noble work is at first impossible. 4, Date: 2022y2m15d City: AA Note: Live beautifully, dream passionately, love completely.
需要从Memo列中提取三个字段:
Date:与City:之间的日期内容City:与Note:之间的城市代码Note:之后的完整语句
最终期望输出格式如下:
No Date City Note 1 2020/10/22 UA True mastery of any skill takes a lifetime. 2 2022/11/01 CH Sweat is the lubricant of success. 3 2022y11m1d UA Every noble work is at first impossible. 4 2022y2m15d AA Live beautifully, dream passionately, love completely.
可行实现方法
方法1:使用tidyr::extract(简洁高效)
借助tidyr包的extract函数,通过正则表达式直接拆分Memo列,一步完成字段提取:
library(tidyr) library(readr) # 读取原始CSV数据(trim_ws = TRUE 自动去除单元格前后空格) df <- read_csv(""" No, Memo 1, Date: 2020/10/22 City: UA Note: True mastery of any skill takes a lifetime. 2, Date: 2022/11/01 City: CH Note: Sweat is the lubricant of success. 3, Date: 2022y11m1d City: UA Note: Every noble work is at first impossible. 4, Date: 2022y2m15d City: AA Note: Live beautifully, dream passionately, love completely. """, trim_ws = TRUE) # 拆分Memo列,提取目标字段 df_clean <- df %>% extract( Memo, into = c("Date", "City", "Note"), regex = "Date: (.*?) City: (.*?) Note: (.*)", remove = FALSE # 若无需保留原Memo列,设为TRUE即可 ) # 输出整理后的结果 print(df_clean[, c("No", "Date", "City", "Note")])
正则说明
Date: (.*?):非贪婪匹配Date:后到City:前的所有内容City: (.*?):非贪婪匹配City:后到Note:前的所有内容Note: (.*):贪婪匹配Note:后的所有剩余内容(因为是最后一段,无需担心过度匹配)
方法2:使用stringr包的正则提取
通过stringr提供的str_extract函数,结合正则断言精准提取每个字段:
library(stringr) # 构造数据框 df <- data.frame( No = 1:4, Memo = c( "Date: 2020/10/22 City: UA Note: True mastery of any skill takes a lifetime.", "Date: 2022/11/01 City: CH Note: Sweat is the lubricant of success.", "Date: 2022y11m1d City: UA Note: Every noble work is at first impossible.", "Date: 2022y2m15d City: AA Note: Live beautifully, dream passionately, love completely." ), stringsAsFactors = FALSE ) # 逐个提取字段 df$Date <- str_extract(df$Memo, "(?<=Date: ).*?(?= City:)") df$City <- str_extract(df$Memo, "(?<=City: ).*?(?= Note:)") df$Note <- str_extract(df$Memo, "(?<=Note: ).*") # 整理并输出结果 result <- df[, c("No", "Date", "City", "Note")] print(result)
正则断言说明
(?<=Date: ):正向后行断言,匹配前面是Date:的位置.*?(?= City:):正向先行断言,匹配到City:前的内容(非贪婪模式避免过度匹配)(?<=Note: ).*:匹配Note:后的所有剩余内容
方法3:Base R原生实现(无需额外包)
使用Base R的regexec和regmatches函数,不依赖第三方包完成提取:
# 构造数据框 df <- data.frame( No = 1:4, Memo = c( "Date: 2020/10/22 City: UA Note: True mastery of any skill takes a lifetime.", "Date: 2022/11/01 City: CH Note: Sweat is the lubricant of success.", "Date: 2022y11m1d City: UA Note: Every noble work is at first impossible.", "Date: 2022y2m15d City: AA Note: Live beautifully, dream passionately, love completely." ), stringsAsFactors = FALSE ) # 定义匹配正则 pattern <- "Date: (.*?) City: (.*?) Note: (.*)" # 提取所有匹配结果 matches <- regmatches(df$Memo, regexec(pattern, df$Memo)) # 将匹配结果转换为数据框并合并 extracted <- do.call(rbind, lapply(matches, function(x) x[-1])) colnames(extracted) <- c("Date", "City", "Note") result <- cbind(df["No"], extracted) # 输出结果 print(result)
逻辑说明
regexec:返回每个字符串中匹配正则的位置信息regmatches:根据位置信息提取对应内容- 通过
lapply和do.call(rbind, ...)将列表格式的匹配结果转换为数据框,最后与原No列合并
内容的提问来源于stack exchange,提问作者mashimena
相关产品推荐
相关产品推荐

