使用Python正则表达式re.findall匹配捕获标签对间内容
正则匹配标签对内容并按组输出
问题
需要匹配并捕获HTML标签对之间的有效文本,最终每行输出2组捕获到的内容。输入与预期输出如下:
输入
<a> </a> <b>hello hello 123</b> stuff to ignore here <i>123412bhje</i> <a>what???</a> stuff to ignore here <b>asd13asf</b> <i>who! Hooooo!</i> stuff to ignore here <i>df7887a</i>
预期输出
hello hello 123 123412bhje what??? asd13asf who! Hooooo! df7887a
要求使用Python的re.findall()实现匹配逻辑。
解决方案
步骤1:匹配所有非空标签内容
使用正则表达式捕获所有标签对中的文本,同时后续过滤掉空内容(比如示例中<a> </a>这类无意义的空标签)。
正则表达式:
<([a-z])>(.*?)</\1>
<([a-z])>:匹配小写字母命名的起始标签,捕获标签名用于匹配对应结束标签(.*?):非贪婪匹配标签内的文本,避免跨标签匹配多余内容</\1>:反向引用前面捕获的标签名,确保匹配对应的结束标签
步骤2:代码实现
import re linein = '<a> </a> <b>hello hello 123</b> stuff to ignore here <i>123412bhje</i> <a>what???</a> stuff to ignore here <b>asd13asf</b> <i>who! Hooooo!</i> stuff to ignore here <i>df7887a</i>' # 捕获所有标签对的标签名和内容 M = re.findall(r'<([a-z])>(.*?)</\1>', linein) # 过滤空内容,提取有效文本 valid_content = [text for _, text in M if text.strip()] # 按每2个一组拼接输出 for idx in range(0, len(valid_content), 2): print(' '.join(valid_content[idx:idx+2]))
运行结果
执行代码后会输出与预期一致的内容:
hello hello 123 123412bhje what??? asd13asf who! Hooooo! df7887a
内容的提问来源于stack exchange,提问作者JD13
相关产品推荐
相关产品推荐

