urllib请求目标URL返回200,HTTPX请求却返回403,如何解决?
使用HTTPX请求目标URL返回403的解决方法
以下是几个可行的解决方案,针对网站识别HTTPX客户端特征的问题:
1. 强制使用HTTP/1.1(禁用HTTP/2)
HTTPX默认启用HTTP/2,而Python原生urllib仅支持HTTP/1.1,部分网站会对HTTP/2请求做拦截。强制切换到HTTP/1.1可以对齐urllib的请求行为:
import httpx headers = { 'User-Agent': 'Mozilla/5.0' } # 禁用HTTP/2,强制使用HTTP/1.1 with httpx.Client(http2=False) as client: response = client.get('https://www.csgodatabase.com/csgo-steam-status/', headers=headers) print(response.status_code)
2. 完全照搬urllib的请求头
urllib会自动添加一些默认请求头(比如Connection: close、Accept: */*),即使你添加了浏览器的全部请求头,也可能和urllib的头存在细微差异。可以通过抓包获取urllib发送的完整请求头,再在HTTPX中完全复用:
import httpx # 复制urllib实际发送的请求头 headers = { 'User-Agent': 'Mozilla/5.0', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate, br', 'Connection': 'close' } with httpx.Client(http2=False) as client: response = client.get('https://www.csgodatabase.com/csgo-steam-status/', headers=headers) print(response.status_code)
3. 禁用连接复用
HTTPX默认使用连接池保持连接复用,而urllib每次请求都会新建连接。网站可能通过连接复用特征识别HTTPX,禁用连接复用可以对齐urllib的行为:
import httpx headers = { 'User-Agent': 'Mozilla/5.0' } # 限制最大连接数为1,关闭连接复用 with httpx.Client(limits=httpx.Limits(max_connections=1, keepalive_expiry=0), http2=False) as client: response = client.get('https://www.csgodatabase.com/csgo-steam-status/', headers=headers) print(response.status_code)
核心原因
网站的反爬机制会通过HTTP版本、连接行为、请求头细节等特征识别不同的HTTP客户端。HTTPX作为现代库默认启用的HTTP/2、连接池等特性,与urllib的原生请求特征差异明显,因此被拦截返回403。通过对齐urllib的请求行为,可以绕过这类检测。
内容的提问来源于stack exchange,提问作者Oliver
相关产品推荐
相关产品推荐

