You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R正则表达式提取文本指定信息并拆分至不同列?

解决方案:用R正则表达式提取文本字段到新列

没问题,我来帮你搞定这个文本提取需求!我们可以用stringr的正则匹配功能,搭配dplyr来高效处理数据,完美适配你给出的所有场景。

步骤1:准备示例数据

先把你提供的文本整理成数据框,方便后续批量处理:

library(dplyr)
library(stringr)

# 创建示例数据集
text_df <- tibble(
  text = c(
    "issue : there has been considerable changes and it is going on. finding : we need further investigation on this. resolution: please check the validity",
    "issue : there has been considerable changes and it is going on. resolution: please check the validity",
    "finding : we need further investigation on this. resolution: please check the validity",
    "please check the validity"
  )
)

步骤2:编写正则提取逻辑

我们用str_extract结合正向/反向预查的正则表达式,精准定位每个字段的内容:

  • 匹配issue:提取issue : 之后的内容,直到遇到下一个字段标识(finding : 或resolution:)或文本结尾
  • 匹配finding:提取finding : 之后的内容,直到遇到resolution:或文本结尾
  • 匹配resolution:既可以提取resolution: 之后的内容,也能直接匹配没有字段标识的纯resolution文本(比如Text-4)

具体处理代码如下:

text_processed <- text_df %>%
  mutate(
    issue = str_extract(text, "(?<=issue : ).*?(?= finding :| resolution:|$)"),
    finding = str_extract(text, "(?<=finding : ).*?(?= resolution:|$)"),
    # 通用版resolution匹配:适配带标识和纯内容两种情况
    resolution = str_extract(text, "(?<=resolution: ).*$|^(?!.*(issue|finding)).*$")
  )

# 查看处理结果
print(text_processed, width = Inf)

结果验证

运行后得到的结果完全符合你的预期:

  • Text-1的三个字段均提取到对应内容
  • Text-2的finding字段自动填充为NA
  • Text-3的issue字段自动填充为NA
  • Text-4的issue和finding为NA,resolution正确提取纯文本内容

如果你的实际数据中resolution的内容不固定,上面的通用正则也能完美适配,无需修改。

内容的提问来源于stack exchange,提问作者Rohit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:14:01