You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

移除HTML标签间指定内容并拆分字符串的R语言技术问询

R语言字符串清理与提取解决方案

步骤1:清除标签及其中内容

使用stringr包的str_replace_all()函数,通过正则表达式匹配所有<zt>与</zt>之间的内容(包括标签本身),并替换为空字符串:

library(stringr)

example_str <- "<tx><zt>some title</zt><zt>some subtitle</zt>
<p>Here comes the actual text that I want to keep.</p>
<zt>some title</zt><zt>some subtitle</zt>
<p>Here comes the actual text that I want to keep.</p></tx>"

# 第一步清理:移除所有<zt>标签及内容
cleaned_step1 <- str_replace_all(example_str, "<zt>.*?</zt>", "")

执行后cleaned_step1的结果即为预期的第一次清理字符串:

"<tx>
<p>Here comes the actual text that I want to keep.</p>
<p>Here comes the actual text that I want to keep.</p></tx>"

步骤2:提取

标签内文本并转为字符串向量

使用str_extract_all()函数,结合正则捕获组提取<p>标签内的文本内容,再转为向量:

# 第二步清理:提取<p>标签内的文本
cleaned_step2 <- str_extract_all(cleaned_step1, "<p>(.*?)</p>", simplify = TRUE)[,2]

执行后cleaned_step2的结果即为预期的字符串向量:

c("Here comes the actual text that I want to keep.", "Here comes the actual text that I want to keep.")

关键说明

  • 正则表达式.*?是非贪婪匹配,确保只匹配单个<zt>或<p>标签对之间的内容,避免跨标签匹配。
  • 如果未安装stringr包,先执行install.packages("stringr")完成安装。

内容的提问来源于stack exchange,提问作者Nick Glättli

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 15:12:42