You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python re.findall消除子组与空值,仅返回目标提取结果?

问题:re.findall提取链接内容返回含空值的元组列表

问题详情

用Python的re.findall从HTML文件中提取example.com/a/和example.com/b/后的内容时,当前输出是包含空字符串的元组列表:[('thing1', '', ''), ('thing2', '', ''), ('', '', 'thing3'), ('', 'thing4', '')],期望得到无空值的纯内容列表:['thing1', 'thing2', 'thing3', 'thing4'],要求仅使用纯正则表达式解决,不采用re.split或数组过滤等方法(认为这类方法不稳定且效率低)。

目标HTML内容

<!DOCTYPE html>
<html>
    <head></head>   
    <body></body>
        <a href="example.com/a/thing1"></a>
        <a href="example.com/a/thing2"></a>
        <a href="example.com/b/thing3"></a>
        <a href="example.com/b/thing4" ><img src="/thing4.png"></a>
    </body>
</html>

原代码

import re

html = open("help.html", "r").read()
links = re.findall('((?<=\.com\/a\/).*(?="))|((?<=\.com\/b\/).*(?=" >))|((?<=\.com\/b\/).*(?="></a))',html)

print(links)

解决方案

问题根源

原正则中使用了多个独立的捕获组(括号包裹的部分),re.findall在存在多个捕获组时,会返回每个捕获组的匹配结果组成的元组,未匹配到的组会返回空字符串,因此出现了含空值的元组列表。

修正后的正则与代码

将正则中的多个捕获组合并,改用非捕获组避免生成子组,并简化匹配逻辑:

import re

html = open("help.html", "r").read()
# 正则说明:
# (?<=\.com\/(?:a|b)\/) 正向预查,定位到.com/a/或.com/b/之后的位置,(?:a|b)是非捕获组,不生成子组结果
# [^"]+ 匹配所有非引号的字符,直到遇到引号(href属性的结束标记)
links = re.findall(r'(?<=\.com\/(?:a|b)\/)[^"]+', html)

print(links)

输出结果

['thing1', 'thing2', 'thing3', 'thing4']

补充说明

如果需要严格匹配原正则中的结尾场景(如部分链接后是" >),也可以使用更精确的正则:

links = re.findall(r'(?<=\.com\/(?:a|b)\/)[^"]+(?="(?: >)?>)', html)

不过[^"]+已经足够覆盖所有情况,因为引号是href属性的结束标记,后续的空格和标签符号不影响内容提取。

内容的提问来源于stack exchange,提问作者Pengalor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 07:15:34