You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

生产环境Python requests/urllib请求特定网站返回403问题求助

问题:Heroku生产环境请求Caixa房产图片返回403 Forbidden,本地运行正常

请求以下图片时遇到403错误:

  • 图片1:https://venda-imoveis.caixa.gov.br/fotos/F878770754273521.jpg
  • 图片2:https://venda-imoveis.caixa.gov.br/fotos/F000001001212121.jpg

情况说明

  • 代码在其他网站测试正常,仅针对https://venda-imoveis.caixa.gov.br/出现问题;
  • 本地运行代码可正常请求上述网站,但部署到Heroku生产环境后返回HTTP错误403 - Forbidden;
  • 已尝试开启和关闭SSL验证,两种方式均无效;
  • 已尝试使用requests库替代urllib,但问题依旧。

疑问

  1. 为何本地正常但生产环境异常?
  2. 如何绕过拦截?
  3. 推测Heroku的IP段被该网站封禁,是否有解决办法?

测试代码

@app.route('/teste', methods=['GET', 'POST'])
def teste():
    import ssl
    # Define the headers as specified
    headers = {
        'Accept-Encoding': 'gzip, deflate, br',
        'Host': 'venda-imoveis.caixa.gov.br',
        'Accept': '*/*',
        'User-Agent': 'Mozilla/5.0 (compatible; filibot/1.0; +https://filibot.com/)',
        'X-Request-Id': 'httpdev.b2e27aba-d204-417a-95fa-d8a9dadb88f5',
        'Referer': 'https://http.app/'
    }

# Create a request object with the headers
url = 'https://venda-imoveis.caixa.gov.br/fotos/F000001001212121.jpg'
req = urllib.request.Request(url, headers=headers)

# Create an SSL context to disable SSL verification (for debugging purposes)
context = ssl._create_unverified_context()
status = 0

# Open the URL with the custom request and SSL context
try:
    with urllib.request.urlopen(req, context=context) as response:
        # Print the response status code
        print('Response Status Code:', response.getcode())
        status = response.getcode()
        # Get the response headers
        response_headers = response.getheaders()
        for header in response_headers:
            print(f'{header[0]}: {header[1]}')

        # Read and print a portion of the content
        content = response.read()
        print('Content Length:', len(content))

except urllib.error.HTTPError as e:
    print(f"HTTP error: {e.code} - {e.reason}")

except urllib.error.URLError as e:
    print(f"URL error: {e.reason}")

except Exception as e:
    print(f"An error occurred: {e}")

status = str(status)
return status

回答

原因分析

本地正常、Heroku异常的核心原因大概率是Heroku的IP段被目标网站的反爬/访问控制机制封禁:

  • 目标网站(Caixa房产平台)可能对云服务商IP段做了拦截,避免批量爬虫或非人类访问;
  • 本地IP属于普通民用网段,不在目标网站的封禁列表内,因此能正常请求;
  • 你尝试的SSL验证开关、更换HTTP库(urllib/requests)都不影响IP层面的拦截,所以无效。

解决办法

要绕过这个限制,核心是让请求的源IP不在目标网站的封禁列表里,具体方案如下:

  1. 使用代理IP

    • 配置第三方代理服务(如住宅代理),让Heroku上的请求通过代理IP发送,避开目标网站的IP封禁;
    • 示例(用requests库配置代理):
      proxies = {
          'https': 'https://your-proxy-ip:port'
      }
      response = requests.get(url, headers=headers, proxies=proxies)
      
    • 优先选择住宅IP,这类IP更接近普通用户访问的网段,被拦截概率更低。
  2. 使用Heroku私人空间(Private Spaces)

    • Heroku私人空间提供专属静态IP地址,你可以向目标网站申请将该IP加入白名单;
    • 缺点是成本较高,适合长期稳定的业务需求。
  3. 调整请求头模拟真实浏览器

    • 当前你的User-Agent为自定义的filibot/1.0,目标网站可能识别为爬虫;
    • 替换为真实浏览器的User-Agent,比如:Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36;
    • 补充Accept-Language等常见请求头,让请求更接近真实用户行为。
  4. 尝试目标网站官方API(若存在)

    • Caixa作为巴西国有银行,可能提供官方房产数据API,通过合法渠道获取数据可避免被拦截;
    • 可查询Caixa开发者文档或联系其支持团队确认是否有公开API可用。

内容的提问来源于stack exchange,提问作者Vinicius Silva

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 13:31:22