You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

bs4 find_all报错:开发正常生产环境出现NoneType属性异常

解决Cloudflare拦截导致的爬虫失败问题

你的问题核心是生产服务器的请求被Cloudflare反爬机制拦截,返回了验证页面,导致无法定位目标DOM元素,抛出NoneType错误。以下是具体解决办法:

  • 改用专门处理Cloudflare的请求库
    推荐使用cloudscraper,它能自动处理Cloudflare的JS验证。先安装库:

    pip install cloudscraper
    

    替换原有requests代码:

    import cloudscraper
    import bs4
    
    # 创建scraper实例
    scraper = cloudscraper.create_scraper()
    # 发送请求,无需额外加User-Agent(库会自动模拟)
    exampleFile = scraper.get('https://statusinvest.com.br/acoes/petr4')
    exampleSoup = bs4.BeautifulSoup(exampleFile.text, features="html.parser")
    
    # 先判断容器是否存在,避免直接链式调用报错
    container_indicadores = exampleSoup.find('div', {'class': "indicator-today-container"})
    if container_indicadores:
        target_divs = container_indicadores.find_all('div', {'title': True})
        # 后续处理target_divs
    else:
        print("请求仍被拦截,未找到目标容器")
    
  • 增强请求头模拟真实浏览器(备选方案)
    如果不想换库,可以补充完整请求头,模拟真实用户的浏览器请求:

    headers = {
        'User-Agent': "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/92.0.4515.159 Safari/537.36",
        'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
        'Accept-Language': 'pt-BR,pt;q=0.8,en-US;q=0.5,en;q=0.3',
        'Referer': 'https://statusinvest.com.br/',
        'DNT': '1',
        'Connection': 'keep-alive',
        'Upgrade-Insecure-Requests': '1'
    }
    

    注意:这种方法对新版Cloudflare验证效果有限,优先用cloudscraper。

  • 控制请求频率
    在请求之间加入延迟,避免被判定为高频爬虫:

    import time
    # 请求前或请求后加延迟
    time.sleep(2)
    
  • 检查服务器IP状态
    Railway的部分IP段可能被Cloudflare标记为风险IP,可以尝试更换Railway服务器的部署地区,或者使用代理IP(付费代理稳定性更高)。

内容的提问来源于stack exchange,提问作者Felipe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 20:25:11