bs4 find_all报错:开发正常生产环境出现NoneType属性异常
解决Cloudflare拦截导致的爬虫失败问题
你的问题核心是生产服务器的请求被Cloudflare反爬机制拦截,返回了验证页面,导致无法定位目标DOM元素,抛出NoneType错误。以下是具体解决办法:
改用专门处理Cloudflare的请求库
推荐使用cloudscraper,它能自动处理Cloudflare的JS验证。先安装库:pip install cloudscraper替换原有requests代码:
import cloudscraper import bs4 # 创建scraper实例 scraper = cloudscraper.create_scraper() # 发送请求,无需额外加User-Agent(库会自动模拟) exampleFile = scraper.get('https://statusinvest.com.br/acoes/petr4') exampleSoup = bs4.BeautifulSoup(exampleFile.text, features="html.parser") # 先判断容器是否存在,避免直接链式调用报错 container_indicadores = exampleSoup.find('div', {'class': "indicator-today-container"}) if container_indicadores: target_divs = container_indicadores.find_all('div', {'title': True}) # 后续处理target_divs else: print("请求仍被拦截,未找到目标容器")增强请求头模拟真实浏览器(备选方案)
如果不想换库,可以补充完整请求头,模拟真实用户的浏览器请求:headers = { 'User-Agent': "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/92.0.4515.159 Safari/537.36", 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'Accept-Language': 'pt-BR,pt;q=0.8,en-US;q=0.5,en;q=0.3', 'Referer': 'https://statusinvest.com.br/', 'DNT': '1', 'Connection': 'keep-alive', 'Upgrade-Insecure-Requests': '1' }注意:这种方法对新版Cloudflare验证效果有限,优先用cloudscraper。
控制请求频率
在请求之间加入延迟,避免被判定为高频爬虫:import time # 请求前或请求后加延迟 time.sleep(2)检查服务器IP状态
Railway的部分IP段可能被Cloudflare标记为风险IP,可以尝试更换Railway服务器的部署地区,或者使用代理IP(付费代理稳定性更高)。
内容的提问来源于stack exchange,提问作者Felipe
相关产品推荐
相关产品推荐

