如何用R正则表达式提取文本指定信息并拆分至不同列?
解决方案:用R正则表达式提取文本字段到新列
没问题,我来帮你搞定这个文本提取需求!我们可以用stringr的正则匹配功能,搭配dplyr来高效处理数据,完美适配你给出的所有场景。
步骤1:准备示例数据
先把你提供的文本整理成数据框,方便后续批量处理:
library(dplyr) library(stringr) # 创建示例数据集 text_df <- tibble( text = c( "issue : there has been considerable changes and it is going on. finding : we need further investigation on this. resolution: please check the validity", "issue : there has been considerable changes and it is going on. resolution: please check the validity", "finding : we need further investigation on this. resolution: please check the validity", "please check the validity" ) )
步骤2:编写正则提取逻辑
我们用str_extract结合正向/反向预查的正则表达式,精准定位每个字段的内容:
- 匹配
issue:提取issue :之后的内容,直到遇到下一个字段标识(finding :或resolution:)或文本结尾 - 匹配
finding:提取finding :之后的内容,直到遇到resolution:或文本结尾 - 匹配
resolution:既可以提取resolution:之后的内容,也能直接匹配没有字段标识的纯resolution文本(比如Text-4)
具体处理代码如下:
text_processed <- text_df %>% mutate( issue = str_extract(text, "(?<=issue : ).*?(?= finding :| resolution:|$)"), finding = str_extract(text, "(?<=finding : ).*?(?= resolution:|$)"), # 通用版resolution匹配:适配带标识和纯内容两种情况 resolution = str_extract(text, "(?<=resolution: ).*$|^(?!.*(issue|finding)).*$") ) # 查看处理结果 print(text_processed, width = Inf)
结果验证
运行后得到的结果完全符合你的预期:
- Text-1的三个字段均提取到对应内容
- Text-2的
finding字段自动填充为NA - Text-3的
issue字段自动填充为NA - Text-4的
issue和finding为NA,resolution正确提取纯文本内容
如果你的实际数据中resolution的内容不固定,上面的通用正则也能完美适配,无需修改。
内容的提问来源于stack exchange,提问作者Rohit
相关产品推荐
相关产品推荐

