Python Requests无法获取最终重定向URL问题求助
嘿,我太懂这种困惑了——明明浏览器里能跳两次到最终地址,用requests试遍了Stack Overflow里的方法却只返回原URL,问题出在重定向的类型上!
首先,浏览器里的跳转不一定都是HTTP协议层面的(也就是301/302这种带Location头的重定向)。你说的这个go.php链接很大概率用的是JS跳转或者HTML meta标签跳转,而requests库默认只会处理HTTP级别的重定向,对需要浏览器渲染执行的JS、meta跳转完全不处理,自然就只会返回原链接。
先验证一下我的猜测
你可以先运行这段代码,看看页面返回的内容到底是什么:
import requests url = "https://guangdiu.com/go.php?id=5347317" # 禁用自动重定向,直接看原页面内容 response = requests.get(url, allow_redirects=False) print(response.text)
如果输出的HTML里有类似 <meta http-equiv="refresh" content="0;url=目标地址"> 这种标签,或者有 window.location.href = 'xxx' 这类JS代码,那就实锤是非HTTP重定向了。
给你两个靠谱的解决方案
方案1:手动解析页面里的跳转地址
如果是meta标签跳转,用BeautifulSoup来提取很方便:
import requests from bs4 import BeautifulSoup url = "https://guangdiu.com/go.php?id=5347317" response = requests.get(url, allow_redirects=False) soup = BeautifulSoup(response.text, "html.parser") # 找meta刷新标签 meta_redirect = soup.find("meta", attrs={"http-equiv": "refresh"}) if meta_redirect: content = meta_redirect.get("content") # 拆分出url部分,比如content是"0;url=https://xxx.com" target_url = content.split("url=")[-1] print("最终跳转URL:", target_url)
如果是JS跳转,用正则匹配简单场景足够:
import requests import re url = "https://guangdiu.com/go.php?id=5347317" response = requests.get(url, allow_redirects=False) # 匹配常见的JS跳转语句 pattern = re.compile(r"window\.location\.href\s*=\s*['\"](.*?)['\"]|location\.replace\(['\"](.*?)['\"]\)") match = pattern.search(response.text) if match: # 取第一个不为空的匹配结果 target_url = match.group(1) or match.group(2) print("最终跳转URL:", target_url)
方案2:用模拟浏览器的工具(复杂场景首选)
如果页面跳转逻辑复杂(比如有延迟、JS判断),解析HTML可能搞不定,这时候用selenium或者playwright这类能模拟真实浏览器的工具就靠谱多了——它们会像你的Chrome/Firefox一样执行JS、完成跳转,自然能拿到最终URL。
以playwright为例(先装依赖:pip install playwright,然后运行playwright install装浏览器驱动):
from playwright.sync_api import sync_playwright url = "https://guangdiu.com/go.php?id=5347317" with sync_playwright() as p: # 无头模式启动浏览器,不弹出窗口 browser = p.chromium.launch(headless=True) page = browser.new_page() # 等待网络空闲,确保跳转全部完成 page.goto(url, wait_until="networkidle") final_url = page.url print("最终跳转URL:", final_url) browser.close()
最后总结一下
你之前用Stack Overflow的方法无效,核心原因就是这个链接的重定向不是requests能处理的HTTP重定向,而是需要浏览器渲染的JS/meta跳转。只要区分清楚重定向类型,选对应的方法就能搞定啦!
内容的提问来源于stack exchange,提问作者yong ho

