You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

状态码200的链接浏览器跳转,Python Requests如何捕获最终URL及数据?

解决链接跳转捕获与200状态码跳转疑问

嘿,我来帮你搞定这个问题!你遇到的情况其实很常见,咱们分两部分来拆解:

一、如何捕获跳转后的最终URL及其数据

首先要明确:跳转分两种——服务器端重定向(HTTP 3xx状态码)和客户端跳转(页面通过JS/meta标签触发)。requests默认只自动处理服务器端的3xx重定向,但你的情况是服务器返回200,说明是客户端跳转,这时候得手动解析页面里的跳转逻辑。

下面是具体的代码方案:

1. 先检查是否存在服务器端隐性重定向

有时候服务器可能返回200但实际做了重定向(虽然不符合HTTP规范,但偶尔会遇到),先通过requests的内置属性排查:

import requests

url = 'http://www.afaqs.com/news/story/52344_The-target-is-to-get-advertisers-to-switch-from-print-to-TV-Ravish-Kumar-Viacom18'
r = requests.get(url)

# 查看请求历史(如果有服务器端重定向,这里会显示跳转记录)
print("请求历史记录:", r.history)
# 查看当前响应的URL(如果已经跳转,这里会是最终URL)
print("当前响应URL:", r.url)

2. 解析客户端跳转(Meta标签/JS跳转)

如果上面的结果显示没有服务器端重定向,那就是页面里的客户端跳转逻辑在起作用,咱们用BeautifulSoup解析页面:

from bs4 import BeautifulSoup
import re

# 先获取原始页面内容
soup = BeautifulSoup(r.text, 'html.parser')

# 情况1:捕获Meta Refresh跳转
meta_redirect = soup.find('meta', attrs={'http-equiv': 'refresh'})
if meta_redirect:
    content = meta_redirect.get('content')
    if 'url=' in content:
        # 提取跳转目标URL
        target_url = content.split('url=')[1].strip()
        # 处理相对URL,拼接成绝对地址
        if not target_url.startswith(('http://', 'https://')):
            target_url = requests.compat.urljoin(url, target_url)
        # 请求最终URL
        final_response = requests.get(target_url)
        print("最终跳转URL(Meta):", final_response.url)
        print("最终响应内容预览:", final_response.text[:500])

# 情况2:捕获JavaScript跳转(比如window.location.href)
js_redirect_match = re.search(r'window\.location\.href\s*=\s*["\'](.*?)["\']', r.text)
if js_redirect_match:
    target_url = js_redirect_match.group(1)
    if not target_url.startswith(('http://', 'https://')):
        target_url = requests.compat.urljoin(url, target_url)
    final_response = requests.get(target_url)
    print("最终跳转URL(JS):", final_response.url)
    print("最终响应内容预览:", final_response.text[:500])

二、为什么状态码200的链接也会跳转?

这是客户端侧跳转的典型特征:

  • 服务器正常返回200状态码和完整的HTML页面,没有返回3xx重定向状态码
  • 页面里嵌入了跳转逻辑:比如<meta http-equiv="refresh" content="0; url=目标地址">标签,或者JavaScript代码window.location.href = '目标地址'
  • 浏览器会自动解析这些跳转逻辑并执行,但requests是HTTP请求库,不会解析HTML或执行JavaScript,所以只会拿到原始页面的数据,不会自动跳转

简单说:服务器没让它跳,是页面自己让浏览器跳的,requests不“读”页面里的跳转指令,所以就停在原始链接了。

内容的提问来源于stack exchange,提问作者Nandesh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:25:52