如何用BeautifulSoup提取HTML中所有http/https开头的链接?
如何用BeautifulSoup提取页面中所有以http/https开头的有效链接?
你的问题很典型——当页面里的链接不只是在<a>标签里,还散落在<link>、<img>的属性或者脚本里时,宽泛的正则很容易匹配到无效内容。咱们一步步解决这个问题:
先分析你原来正则的问题
你用的正则r'(?:(?:https?|ftp)://)?[\w/\-?=%.]+\.[\w/\-?=%.]+'有几个明显的缺陷:
- 开头的
https?://是可选的,导致会匹配id1032680895.png这种非完整链接 - 字符集
[\w/\-?=%.]不包含+、|这类链接里常见的字符,所以会截断像https://fonts.googleapis.com/css?family=Open+Sans:600这样的链接 - 没有限定匹配的上下文,所以会把脚本里的
window.location、loc.href这类变量名误判成链接
解决方案:分场景精准提取
我们可以把提取分为两部分:标签属性里的链接和脚本里的链接,分别处理后合并去重。
完整代码示例
from bs4 import BeautifulSoup import re # 你的示例HTML内容 html_content = ''' <html> <head> </head> <link href="https://fonts.googleapis.com/css?family=Open+Sans:600" rel="stylesheet"/> <style> html, body { height: 100%; width: 100%; } body { background: #F5F6F8; font-size: 16px; font-family: 'Open Sans', sans-serif; color: #2C3E51; } .main { display: flex; align-items: center; justify-content: center; height: 100vh; } .main > div > div, .main > div > span { text-align: center; } .main span { display: block; padding: 80px 0 170px; font-size: 3rem; } .main .app img { width: 400px; } </style> <script type="text/javascript"> var fallback_url = "null"; var store_link = "itms-apps://itunes.apple.com/GB/app/id1032680895?ls=1&mt=8"; var web_store_link = "https://itunes.apple.com/GB/app/id1032680895?mt=8"; var loc = window.location; function redirect_to_web_store(loc) { loc.href = web_store_link; } function redirect(loc) { loc.href = store_link; if (fallback_url.startsWith("http")) { setTimeout(function() { loc.href = fallback_url; },5000); } } </script> <body onload="redirect(loc)"> <div class="main"> <div class="workarea"> <div class="logo"> <img onclick="redirect_to_web_store(loc)" src="https://cdnappicons.appsflyer.com/app|id1032680895.png" style="width:200px;height:200px;border-radius:20px;"/> </div> <span>BetBull: Sports Betting & Tips</span> <div class="app"> <img onclick="redirect_to_web_store(loc)" src="https://cdn.appsflyer.com/af-statics/images/rta/app_store_badge.png"/> </div> </div> </div> </body> </html> ''' soup = BeautifulSoup(html_content, 'html.parser') valid_links = set() # 用集合自动去重 # 1. 提取所有标签的href、src属性中的http/https链接 for tag in soup.find_all(href=True): href = tag['href'] if href.startswith(('http://', 'https://')): valid_links.add(href) for tag in soup.find_all(src=True): src = tag['src'] if src.startswith(('http://', 'https://')): valid_links.add(src) # 2. 提取script标签中被双引号包裹的http/https链接 script_link_pattern = re.compile(r'"(https?://[^"]+)"') for script in soup.find_all('script'): if script.string: matches = script_link_pattern.findall(script.string) valid_links.update(matches) # 输出结果 print("提取到的有效链接:") for link in sorted(valid_links): print(f"- {link}")
代码说明
- 标签属性处理:直接遍历所有带
href和src属性的标签,筛选以http://或https://开头的内容,这部分是页面中最常见的链接载体。 - 脚本内容处理:用正则
r'"(https?://[^"]+)"'精准匹配脚本中被双引号包裹的http/https链接,避免匹配变量名或代码语句。 - 集合去重:用
set存储链接,自动去除重复的内容。
运行结果
你会得到干净的有效链接列表:
提取到的有效链接: - https://cdn.appsflyer.com/af-statics/images/rta/app_store_badge.png - https://cdnappicons.appsflyer.com/app|id1032680895.png - https://fonts.googleapis.com/css?family=Open+Sans:600 - https://itunes.apple.com/GB/app/id1032680895?mt=8
内容的提问来源于stack exchange,提问作者Kishan Mehta
相关产品推荐
相关产品推荐

