如何用Python提取HTML页面Script标签内的指定跳转链接?
如何从HTML的script标签中提取重定向URL?
你已经通过以下Python代码获取了HTML响应:
... resp = logout_session.get(logout_url, headers=headers, verify=False, allow_redirects=False) soup = BeautifulSoup(resp.content, "html.parser") print(soup.prettify())
需要从返回的HTML中提取script标签内的这个重定向链接:
https://idpftc.business.com/saml/Gy736KPK3v1aWDPECRZKAn/proxy_logout/?SAMLResponse=3VjJkuNIjv2VtKijLJObJIphlWnGfd93Xtoo7vsukvr6ZkRU
可以通过BeautifulSoup定位script标签+正则表达式匹配的方式实现,以下是两种可行方案:
方案1:结合BeautifulSoup与正则提取
先定位到目标script标签,再用正则匹配其中的URL:
import re from bs4 import BeautifulSoup # 你已有的获取响应代码 resp = logout_session.get(logout_url, headers=headers, verify=False, allow_redirects=False) soup = BeautifulSoup(resp.content, "html.parser") # 找到包含重定向逻辑的script标签 script_tag = soup.find("script", language="javascript") if script_tag: # 用正则匹配window.location.replace中的URL(匹配双引号包裹的https链接) url_pattern = re.compile(r'window\.location\.replace\("(https://[^"]+)"\)') match = url_pattern.search(script_tag.string) if match: target_url = match.group(1) print(target_url)
方案2:直接对响应文本用正则提取(无需BeautifulSoup)
如果不需要处理其他HTML内容,也可以直接对响应文本进行正则匹配:
import re resp = logout_session.get(logout_url, headers=headers, verify=False, allow_redirects=False) html_content = resp.text # 同样用正则匹配目标URL url_pattern = re.compile(r'window\.location\.replace\("(https://[^"]+)"\)') match = url_pattern.search(html_content) if match: target_url = match.group(1) print(target_url)
说明
- 正则表达式
r'window\.location\.replace\("(https://[^"]+)"\)'的作用:精准匹配window.location.replace("...")结构中的URL,其中https://[^"]+表示匹配以https开头、直到下一个双引号结束的内容,确保提取到完整的目标链接。 - 如果页面中有多个script标签,方案1通过
language="javascript"过滤更精准;如果只有一个script标签,两种方案都可以正常工作。
内容的提问来源于stack exchange,提问作者user3595231
相关产品推荐
相关产品推荐

