You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何不依赖bs4等库 从HTML中用正则提取纯净图片链接

正则调整方案

你遇到的确实是贪婪匹配问题,同时需要适配两种链接包裹场景、过滤掉查询参数,调整后的实现如下:

完整代码

import re

html = """<img src="https://s2.example.com/path/image0.jpg?lastmod=1625296911"><br>
<div><a style="background-image:url(https://s2.example.com/path/image1.jpg?lastmod=1625296911)"</a><a style="background-image:url(https://s2.example.com/path/image2.jpg?lastmod=1625296912)"></a><a style="background-image:url(https://s2.example.com/path/image3.jpg?lastmod=1625296912)"></a></div>"""

images = []
for line in html.split('\n'):
    images.extend(re.findall(r'(?:src="|url\()(https://s2[^?]+)\?lastmod=\d+', line))

print(images)

调整说明

  1. 正则部分用(?:src="|url\()非捕获组匹配两种链接前缀,不会将前缀内容纳入匹配结果
  2. 用[^?]+替代原有的.*,既避免了贪婪匹配,又直接截取到查询参数?之前的内容,直接得到无参数的图片链接
  3. 列表操作改用extend替代append,避免匹配结果出现嵌套结构
    运行后输出结果和你预期的完全一致。

内容的提问来源于stack exchange,提问作者user16312732

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 14:45:05