You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

正则表达式匹配固定模式后的所有内容(含空格与标点)

提取字符串中特定模式后的内容

需求:提取字符串中特定模式: - (冒号+空格+连字符+空格)之后的所有内容,包含带连字符、撇号的文本。

示例数据

library(tibble)
eg_data <- tibble(
  Text = c(
    "Please advise: - How to extract this part.",
    "Please advise: - I'm not sure how to extract the latter part of this, and I really need to.", 
    "Please advise: - My co-workers can't help me."),
  IdealOutcome = c(
    "How to extract this part.",
    "I'm not sure how to extract the latter part of this, and I really need to.",
    "My co-workers can't help me.")
)

解决方案

方法1:使用stringr包(tidyverse工具集)

借助str_remove移除模式前的所有内容:

library(stringr)

# 提取目标内容
eg_data$Extracted <- str_remove(eg_data$Text, "^.*: - ")

正则说明:^.*: - 匹配从字符串开头到: - 的所有字符,str_remove会删除这部分,保留后续内容。

方法2:使用base R的sub函数

无需额外安装包,直接用基础正则替换:

# 提取目标内容
eg_data$Extracted <- sub("^.*: - ", "", eg_data$Text)

验证结果

检查提取结果与理想输出是否一致:

all.equal(eg_data$Extracted, eg_data$IdealOutcome)
# 输出:TRUE

内容的提问来源于stack exchange,提问作者Adam_S

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 19:11:02