You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup提取HTML中所有http/https开头的链接?

如何用BeautifulSoup提取页面中所有以http/https开头的有效链接?

你的问题很典型——当页面里的链接不只是在<a>标签里,还散落在<link>、<img>的属性或者脚本里时,宽泛的正则很容易匹配到无效内容。咱们一步步解决这个问题:

先分析你原来正则的问题

你用的正则r'(?:(?:https?|ftp)://)?[\w/\-?=%.]+\.[\w/\-?=%.]+'有几个明显的缺陷:

  • 开头的https?://是可选的,导致会匹配id1032680895.png这种非完整链接
  • 字符集[\w/\-?=%.]不包含+、|这类链接里常见的字符,所以会截断像https://fonts.googleapis.com/css?family=Open+Sans:600这样的链接
  • 没有限定匹配的上下文,所以会把脚本里的window.location、loc.href这类变量名误判成链接

解决方案:分场景精准提取

我们可以把提取分为两部分:标签属性里的链接和脚本里的链接,分别处理后合并去重。

完整代码示例

from bs4 import BeautifulSoup
import re

# 你的示例HTML内容
html_content = '''
<html>
<head>
</head>
<link href="https://fonts.googleapis.com/css?family=Open+Sans:600" rel="stylesheet"/>
<style>
 html, body {
 height: 100%;
 width: 100%;
 }

 body {
 background: #F5F6F8;
 font-size: 16px;
 font-family: 'Open Sans', sans-serif;
 color: #2C3E51;
 }
 .main {
 display: flex;
 align-items: center;
 justify-content: center;
 height: 100vh;
 }
 .main > div > div,
 .main > div > span {
 text-align: center;
 }
 .main span {
 display: block;
 padding: 80px 0 170px;
 font-size: 3rem;
 }
 .main .app img {
 width: 400px;
 }
</style>
<script type="text/javascript">
 var fallback_url = "null";
 var store_link = "itms-apps://itunes.apple.com/GB/app/id1032680895?ls=1&amp;mt=8";
 var web_store_link = "https://itunes.apple.com/GB/app/id1032680895?mt=8";
 var loc = window.location;
 function redirect_to_web_store(loc) {
 loc.href = web_store_link;
 }
 function redirect(loc) {
 loc.href = store_link;
 if (fallback_url.startsWith("http")) {
 setTimeout(function() {
 loc.href = fallback_url;
 },5000);
 }
 }
</script>
<body onload="redirect(loc)">
<div class="main">
<div class="workarea">
<div class="logo">
<img onclick="redirect_to_web_store(loc)" src="https://cdnappicons.appsflyer.com/app|id1032680895.png" style="width:200px;height:200px;border-radius:20px;"/>
</div>
<span>BetBull: Sports Betting &amp; Tips</span>
<div class="app">
<img onclick="redirect_to_web_store(loc)" src="https://cdn.appsflyer.com/af-statics/images/rta/app_store_badge.png"/>
</div>
</div>
</div>
</body>
</html>
'''

soup = BeautifulSoup(html_content, 'html.parser')
valid_links = set()  # 用集合自动去重

# 1. 提取所有标签的href、src属性中的http/https链接
for tag in soup.find_all(href=True):
    href = tag['href']
    if href.startswith(('http://', 'https://')):
        valid_links.add(href)

for tag in soup.find_all(src=True):
    src = tag['src']
    if src.startswith(('http://', 'https://')):
        valid_links.add(src)

# 2. 提取script标签中被双引号包裹的http/https链接
script_link_pattern = re.compile(r'"(https?://[^"]+)"')
for script in soup.find_all('script'):
    if script.string:
        matches = script_link_pattern.findall(script.string)
        valid_links.update(matches)

# 输出结果
print("提取到的有效链接:")
for link in sorted(valid_links):
    print(f"- {link}")

代码说明

  • 标签属性处理:直接遍历所有带href和src属性的标签,筛选以http://或https://开头的内容,这部分是页面中最常见的链接载体。
  • 脚本内容处理:用正则r'"(https?://[^"]+)"'精准匹配脚本中被双引号包裹的http/https链接,避免匹配变量名或代码语句。
  • 集合去重:用set存储链接,自动去除重复的内容。

运行结果

你会得到干净的有效链接列表:

提取到的有效链接:
- https://cdn.appsflyer.com/af-statics/images/rta/app_store_badge.png
- https://cdnappicons.appsflyer.com/app|id1032680895.png
- https://fonts.googleapis.com/css?family=Open+Sans:600
- https://itunes.apple.com/GB/app/id1032680895?mt=8

内容的提问来源于stack exchange,提问作者Kishan Mehta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 10:10:43