You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python提取HTML页面Script标签内的指定跳转链接?

如何从HTML的script标签中提取重定向URL?

你已经通过以下Python代码获取了HTML响应:

...
resp = logout_session.get(logout_url, headers=headers, verify=False, allow_redirects=False)
soup = BeautifulSoup(resp.content, "html.parser")
print(soup.prettify())

需要从返回的HTML中提取script标签内的这个重定向链接:

https://idpftc.business.com/saml/Gy736KPK3v1aWDPECRZKAn/proxy_logout/?SAMLResponse=3VjJkuNIjv2VtKijLJObJIphlWnGfd93Xtoo7vsukvr6ZkRU

可以通过BeautifulSoup定位script标签+正则表达式匹配的方式实现,以下是两种可行方案:

方案1:结合BeautifulSoup与正则提取

先定位到目标script标签,再用正则匹配其中的URL:

import re
from bs4 import BeautifulSoup

# 你已有的获取响应代码
resp = logout_session.get(logout_url, headers=headers, verify=False, allow_redirects=False)
soup = BeautifulSoup(resp.content, "html.parser")

# 找到包含重定向逻辑的script标签
script_tag = soup.find("script", language="javascript")
if script_tag:
    # 用正则匹配window.location.replace中的URL(匹配双引号包裹的https链接)
    url_pattern = re.compile(r'window\.location\.replace\("(https://[^"]+)"\)')
    match = url_pattern.search(script_tag.string)
    if match:
        target_url = match.group(1)
        print(target_url)

方案2:直接对响应文本用正则提取(无需BeautifulSoup)

如果不需要处理其他HTML内容,也可以直接对响应文本进行正则匹配:

import re

resp = logout_session.get(logout_url, headers=headers, verify=False, allow_redirects=False)
html_content = resp.text

# 同样用正则匹配目标URL
url_pattern = re.compile(r'window\.location\.replace\("(https://[^"]+)"\)')
match = url_pattern.search(html_content)
if match:
    target_url = match.group(1)
    print(target_url)

说明

  • 正则表达式r'window\.location\.replace\("(https://[^"]+)"\)'的作用:精准匹配window.location.replace("...")结构中的URL,其中https://[^"]+表示匹配以https开头、直到下一个双引号结束的内容,确保提取到完整的目标链接。
  • 如果页面中有多个script标签,方案1通过language="javascript"过滤更精准;如果只有一个script标签,两种方案都可以正常工作。

内容的提问来源于stack exchange,提问作者user3595231

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 18:03:35