生产环境Python requests/urllib请求特定网站返回403问题求助
问题:Heroku生产环境请求Caixa房产图片返回403 Forbidden,本地运行正常
请求以下图片时遇到403错误:
- 图片1:
https://venda-imoveis.caixa.gov.br/fotos/F878770754273521.jpg - 图片2:
https://venda-imoveis.caixa.gov.br/fotos/F000001001212121.jpg
情况说明
- 代码在其他网站测试正常,仅针对
https://venda-imoveis.caixa.gov.br/出现问题; - 本地运行代码可正常请求上述网站,但部署到Heroku生产环境后返回HTTP错误403 - Forbidden;
- 已尝试开启和关闭SSL验证,两种方式均无效;
- 已尝试使用requests库替代urllib,但问题依旧。
疑问
- 为何本地正常但生产环境异常?
- 如何绕过拦截?
- 推测Heroku的IP段被该网站封禁,是否有解决办法?
测试代码
@app.route('/teste', methods=['GET', 'POST']) def teste(): import ssl # Define the headers as specified headers = { 'Accept-Encoding': 'gzip, deflate, br', 'Host': 'venda-imoveis.caixa.gov.br', 'Accept': '*/*', 'User-Agent': 'Mozilla/5.0 (compatible; filibot/1.0; +https://filibot.com/)', 'X-Request-Id': 'httpdev.b2e27aba-d204-417a-95fa-d8a9dadb88f5', 'Referer': 'https://http.app/' } # Create a request object with the headers url = 'https://venda-imoveis.caixa.gov.br/fotos/F000001001212121.jpg' req = urllib.request.Request(url, headers=headers) # Create an SSL context to disable SSL verification (for debugging purposes) context = ssl._create_unverified_context() status = 0 # Open the URL with the custom request and SSL context try: with urllib.request.urlopen(req, context=context) as response: # Print the response status code print('Response Status Code:', response.getcode()) status = response.getcode() # Get the response headers response_headers = response.getheaders() for header in response_headers: print(f'{header[0]}: {header[1]}') # Read and print a portion of the content content = response.read() print('Content Length:', len(content)) except urllib.error.HTTPError as e: print(f"HTTP error: {e.code} - {e.reason}") except urllib.error.URLError as e: print(f"URL error: {e.reason}") except Exception as e: print(f"An error occurred: {e}") status = str(status) return status
回答
原因分析
本地正常、Heroku异常的核心原因大概率是Heroku的IP段被目标网站的反爬/访问控制机制封禁:
- 目标网站(Caixa房产平台)可能对云服务商IP段做了拦截,避免批量爬虫或非人类访问;
- 本地IP属于普通民用网段,不在目标网站的封禁列表内,因此能正常请求;
- 你尝试的SSL验证开关、更换HTTP库(urllib/requests)都不影响IP层面的拦截,所以无效。
解决办法
要绕过这个限制,核心是让请求的源IP不在目标网站的封禁列表里,具体方案如下:
使用代理IP
- 配置第三方代理服务(如住宅代理),让Heroku上的请求通过代理IP发送,避开目标网站的IP封禁;
- 示例(用requests库配置代理):
proxies = { 'https': 'https://your-proxy-ip:port' } response = requests.get(url, headers=headers, proxies=proxies) - 优先选择住宅IP,这类IP更接近普通用户访问的网段,被拦截概率更低。
使用Heroku私人空间(Private Spaces)
- Heroku私人空间提供专属静态IP地址,你可以向目标网站申请将该IP加入白名单;
- 缺点是成本较高,适合长期稳定的业务需求。
调整请求头模拟真实浏览器
- 当前你的User-Agent为自定义的
filibot/1.0,目标网站可能识别为爬虫; - 替换为真实浏览器的User-Agent,比如:
Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36; - 补充
Accept-Language等常见请求头,让请求更接近真实用户行为。
- 当前你的User-Agent为自定义的
尝试目标网站官方API(若存在)
- Caixa作为巴西国有银行,可能提供官方房产数据API,通过合法渠道获取数据可避免被拦截;
- 可查询Caixa开发者文档或联系其支持团队确认是否有公开API可用。
内容的提问来源于stack exchange,提问作者Vinicius Silva
相关产品推荐
相关产品推荐

