Python Requests部分HTTP转HTTPS重定向失败的原因与解决方法
Requests无法自动将http://www.google.com重定向到HTTPS的原因与解决方法
问题原因
你观察到的差异完全是**HSTS(HTTP严格传输安全)**机制导致的:
- Chromium浏览器会缓存HSTS域名列表(
www.google.com已包含在内),当你输入HTTP地址时,浏览器会直接在本地将请求转为HTTPS,根本不会发送HTTP请求到服务器——这就是你在开发者工具里看到的“预先转成HTTPS请求”的现象。 - Requests库没有内置HSTS缓存机制,会直接发送HTTP请求到
www.google.com服务器。而Google对已启用HSTS的域名,不会对HTTP请求返回301/302重定向(HSTS要求客户端强制使用HTTPS),所以Requests停留在HTTP地址,没有自动跳转。
对比github.com的情况:GitHub的HSTS配置允许对HTTP请求返回301重定向,因此Requests能正常跟随跳转。
测试代码
import requests r = requests.get("https://www.google.com", allow_redirects=True) print(r.url) print(r.history) r = requests.get("http://www.google.com", allow_redirects=True) print(r.url) print(r.history) r = requests.get("https://google.com", allow_redirects=True) print(r.url) print(r.history) r = requests.get("http://google.com", allow_redirects=True) print(r.url) print(r.history) r = requests.get("http://github.com", allow_redirects=True) print(r.url) print(r.history)
输出结果
https://www.google.com/ [] http://www.google.com/ [] https://www.google.com/ [<Response [301]>] http://www.google.com/ [<Response [301]>] https://github.com/ [<Response [301]>]
实现类浏览器的重定向效果
有几种方法可以让Requests模拟浏览器的HSTS行为:
直接使用HTTPS地址
最简单的方式是跳过HTTP,直接请求HTTPS版本的地址https://www.google.com,和浏览器的行为完全一致。手动实现HSTS检查逻辑
维护一个包含HSTS域名的列表,在发送请求前判断域名是否在列表中,自动将HTTP替换为HTTPS:import requests from urllib.parse import urlparse # 可根据实际需求扩展HSTS域名列表 HSTS_DOMAINS = {"www.google.com", "google.com"} def hsts_enforced_get(url): parsed_url = urlparse(url) if parsed_url.netloc in HSTS_DOMAINS and parsed_url.scheme == "http": url = url.replace("http://", "https://", 1) return requests.get(url, allow_redirects=True) # 测试 response = hsts_enforced_get("http://www.google.com") print(response.url) # 输出:https://www.google.com/使用支持HSTS的Requests扩展
借助requests-hsts这类第三方库,它会自动缓存服务器返回的HSTS头部,后续对同一域名的HTTP请求会自动转为HTTPS,无需手动维护域名列表。
内容的提问来源于stack exchange,提问作者Tom Lin
相关产品推荐
相关产品推荐

