You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取指定URL的重定向链接?解决请求返回503问题

提取重定向链接的解决方案

问题分析

请求返回<Response [503]>是因为目标网站的反爬机制拦截了你的请求——你当前的请求没有模拟正常浏览器的访问行为,被识别为非人类请求。

解决步骤

1. 模拟浏览器请求头

给requests.get添加常见的浏览器请求头,伪装成真实用户访问:

import requests

target_url = "https://www.forexfactory.com/news/403059-manufacturing-in-us-expands-after-reaching-three-year-low/hit"
# 模拟Chrome浏览器的请求头
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8",
    "Accept-Language": "zh-CN,zh;q=0.8,zh-TW;q=0.7,zh-HK;q=0.5,en-US;q=0.3,en;q=0.2"
}

# 默认allow_redirects为True,会自动跟随重定向
response = requests.get(target_url, headers=headers)

if response.status_code == 200:
    print("最终重定向目标链接:", response.url)
else:
    print(f"请求失败,状态码: {response.status_code}")

2. 应对更严格的反爬

如果加了基础请求头还是返回503,可以尝试:

  • 使用requests.Session()维持会话状态,模拟浏览器的Cookie留存行为
  • 添加随机延时,避免短时间内重复请求触发限流
  • 补充Referer字段,模拟从网站内部跳转过来的请求

3. 备选方案:解析页面跳转逻辑

如果直接请求无法跟随重定向,可先获取页面源码,解析页面内的跳转触发规则:

import requests
from bs4 import BeautifulSoup
import re

target_url = "https://www.forexfactory.com/news/403059-manufacturing-in-us-expands-after-reaching-three-year-low/hit"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

# 禁止自动重定向,先获取跳转前的页面
response = requests.get(target_url, headers=headers, allow_redirects=False)
if response.status_code == 200:
    soup = BeautifulSoup(response.text, "html.parser")
    # 查找meta标签跳转
    meta_redirect = soup.find("meta", attrs={"http-equiv": "refresh"})
    if meta_redirect:
        redirect_info = meta_redirect.get("content")
        redirect_url = redirect_info.split("url=")[-1]
        print("页面meta跳转链接:", redirect_url)
    
    # 查找JavaScript跳转(比如window.location.href)
    js_redirect = re.search(r'window\.location\.href\s*=\s*["\'](.*?)["\']', response.text)
    if js_redirect:
        print("JS跳转链接:", js_redirect.group(1))

核心要点

503状态码本质是网站的临时反爬拦截,解决的核心是让请求尽可能贴近真实用户的浏览器行为;如果网站启用了验证码、设备指纹验证等强反爬措施,可能需要结合selenium等工具模拟完整的浏览器交互。

内容的提问来源于stack exchange,提问作者backlog

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 20:00:58