Python拆分带标记文本:移除正则捕获组冗余元素
Python正则拆分HTML标签时移除冗余捕获组元素
问题描述
我有如下文本:<code>stuff</code> and stuff and $\LaTeX$ and <pre class="mermaid">stuff</pre>,希望用Python将其拆分为目标列表:['<code>', 'stuff', '</code>', ' and stuff and $\LaTeX$ ', '<pre class="mermaid">', 'stuff', '</pre>']。当前使用正则表达式拆分后,结果中包含冗余的tag捕获组元素,如何在单行处理场景下移除这些冗余元素?
解决方案
核心思路
避免拆分后产生冗余元素的关键是:要么直接提取所有需要的片段,要么过滤掉拆分后产生的空值或无效匹配。
方法一:用re.findall直接提取目标片段(推荐)
放弃使用re.split,改用re.findall匹配所有符合要求的内容——包括完整的HTML标签(开始/结束)和标签外的文本,这样能直接得到无冗余的片段列表:
import re text = '<code>stuff</code> and stuff and $\\LaTeX$ and <pre class="mermaid">stuff</pre>' # 正则规则:匹配完整HTML标签,或标签外的所有文本 pattern = r'<[^>]+>|[^<]+' # 提取所有匹配片段 raw_result = re.findall(pattern, text) # 将标签中的< >转为HTML实体,得到目标列表 target_list = [item.replace('<', '<').replace('>', '>') if '<' in item else item for item in raw_result]
方法二:re.split结合列表推导式过滤冗余
如果坚持用re.split,可以在拆分后过滤掉空字符串和无效的捕获组内容:
import re text = '<code>stuff</code> and stuff and $\\LaTeX$ and <pre class="mermaid">stuff</pre>' # 带捕获组的拆分正则 pattern = r'(<[^>]+>)([^<]+)(<\/[^>]+>)' split_result = re.split(pattern, text) # 过滤空元素,得到有效片段 clean_result = [item for item in split_result if item] # 同样处理HTML实体转义 target_list = [item.replace('<', '<').replace('>', '>') if '<' in item else item for item in clean_result]
单行处理版本
将所有逻辑压缩为单行代码,适合快速场景:
import re text = '<code>stuff</code> and stuff and $\\LaTeX$ and <pre class="mermaid">stuff</pre>' target_list = [x.replace('<','<').replace('>','>') if '<' in x else x for x in re.findall(r'<[^>]+>|[^<]+', text)]
内容的提问来源于stack exchange,提问作者Aurélien Pierre
相关产品推荐
相关产品推荐

