移除HTML标签间指定内容并拆分字符串的R语言技术问询
R语言字符串清理与提取解决方案
步骤1:清除标签及其中内容
使用stringr包的str_replace_all()函数,通过正则表达式匹配所有<zt>与</zt>之间的内容(包括标签本身),并替换为空字符串:
library(stringr) example_str <- "<tx><zt>some title</zt><zt>some subtitle</zt> <p>Here comes the actual text that I want to keep.</p> <zt>some title</zt><zt>some subtitle</zt> <p>Here comes the actual text that I want to keep.</p></tx>" # 第一步清理:移除所有<zt>标签及内容 cleaned_step1 <- str_replace_all(example_str, "<zt>.*?</zt>", "")
执行后cleaned_step1的结果即为预期的第一次清理字符串:
"<tx> <p>Here comes the actual text that I want to keep.</p> <p>Here comes the actual text that I want to keep.</p></tx>"
步骤2:提取
标签内文本并转为字符串向量
使用str_extract_all()函数,结合正则捕获组提取<p>标签内的文本内容,再转为向量:
# 第二步清理:提取<p>标签内的文本 cleaned_step2 <- str_extract_all(cleaned_step1, "<p>(.*?)</p>", simplify = TRUE)[,2]
执行后cleaned_step2的结果即为预期的字符串向量:
c("Here comes the actual text that I want to keep.", "Here comes the actual text that I want to keep.")
关键说明
- 正则表达式
.*?是非贪婪匹配,确保只匹配单个<zt>或<p>标签对之间的内容,避免跨标签匹配。 - 如果未安装
stringr包,先执行install.packages("stringr")完成安装。
内容的提问来源于stack exchange,提问作者Nick Glättli
相关产品推荐
相关产品推荐

