You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python正则表达式re.findall匹配捕获标签对间内容

正则匹配标签对内容并按组输出

问题

需要匹配并捕获HTML标签对之间的有效文本,最终每行输出2组捕获到的内容。输入与预期输出如下:

输入

<a> </a> <b>hello hello 123</b> stuff to ignore here <i>123412bhje</i> <a>what???</a> stuff to ignore here <b>asd13asf</b> <i>who! Hooooo!</i> stuff to ignore here <i>df7887a</i>

预期输出

hello hello 123 123412bhje 
what??? asd13asf 
who! Hooooo! df7887a

要求使用Python的re.findall()实现匹配逻辑。


解决方案

步骤1:匹配所有非空标签内容

使用正则表达式捕获所有标签对中的文本,同时后续过滤掉空内容(比如示例中<a> </a>这类无意义的空标签)。

正则表达式:

<([a-z])>(.*?)</\1>
  • <([a-z])>:匹配小写字母命名的起始标签,捕获标签名用于匹配对应结束标签
  • (.*?):非贪婪匹配标签内的文本,避免跨标签匹配多余内容
  • </\1>:反向引用前面捕获的标签名,确保匹配对应的结束标签

步骤2:代码实现

import re

linein = '<a> </a> <b>hello hello 123</b> stuff to ignore here <i>123412bhje</i> <a>what???</a> stuff to ignore here <b>asd13asf</b> <i>who! Hooooo!</i> stuff to ignore here <i>df7887a</i>'

# 捕获所有标签对的标签名和内容
M = re.findall(r'<([a-z])>(.*?)</\1>', linein)
# 过滤空内容,提取有效文本
valid_content = [text for _, text in M if text.strip()]
# 按每2个一组拼接输出
for idx in range(0, len(valid_content), 2):
    print(' '.join(valid_content[idx:idx+2]))

运行结果

执行代码后会输出与预期一致的内容:

hello hello 123 123412bhje
what??? asd13asf
who! Hooooo! df7887a

内容的提问来源于stack exchange,提问作者JD13

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 22:21:37