在R中如何根据subheader列移除full_speech列对应内容及后续文本?
问题:清理辩论发言中的子标题及后续内容
我需要清理一个包含辩论长篇发言的列,每行对应一位新发言者的内容,但类似子标题(subheader)的内容会留在每段发言的末尾,不符合需求。
示例数据
speeches <- tibble(subheader = c("3.Discussion", "8.Voting"), full_speech = c("I close this part. 3.Discussion Let's start with", "I think we can vote now") )
期望结果
subheader full_speech 3.Discussion I close this part. 8.Voting I think we can vote now
尝试过的代码(无效)
speeches %>% mutate(full_speech = str_remove(full_speech, subheader))
该代码仅能删除子标题本身,无法移除子标题之后的内容。
解决方案
使用str_remove结合正则表达式,精确匹配子标题并移除其后续所有内容:
library(dplyr) library(stringr) speeches %>% mutate(full_speech = str_remove(full_speech, paste0(fixed(subheader), ".*")))
说明
fixed(subheader)确保子标题中的特殊字符(比如.)被精确匹配,而非作为正则通配符.*匹配子标题之后的任意内容(包括空格、文字等)- 最终会保留子标题出现之前的发言内容,符合需求
内容的提问来源于stack exchange,提问作者stacksterppr
相关产品推荐
相关产品推荐

