正则表达式匹配固定模式后的所有内容(含空格与标点)
提取字符串中特定模式后的内容
需求:提取字符串中特定模式: - (冒号+空格+连字符+空格)之后的所有内容,包含带连字符、撇号的文本。
示例数据
library(tibble) eg_data <- tibble( Text = c( "Please advise: - How to extract this part.", "Please advise: - I'm not sure how to extract the latter part of this, and I really need to.", "Please advise: - My co-workers can't help me."), IdealOutcome = c( "How to extract this part.", "I'm not sure how to extract the latter part of this, and I really need to.", "My co-workers can't help me.") )
解决方案
方法1:使用stringr包(tidyverse工具集)
借助str_remove移除模式前的所有内容:
library(stringr) # 提取目标内容 eg_data$Extracted <- str_remove(eg_data$Text, "^.*: - ")
正则说明:^.*: - 匹配从字符串开头到: - 的所有字符,str_remove会删除这部分,保留后续内容。
方法2:使用base R的sub函数
无需额外安装包,直接用基础正则替换:
# 提取目标内容 eg_data$Extracted <- sub("^.*: - ", "", eg_data$Text)
验证结果
检查提取结果与理想输出是否一致:
all.equal(eg_data$Extracted, eg_data$IdealOutcome) # 输出:TRUE
内容的提问来源于stack exchange,提问作者Adam_S
相关产品推荐
相关产品推荐

