爬虫代码偶发AttributeError: 'NoneType'无find_all属性问题排查及解决
问题分析与修复方案
报错成因
- 未校验请求合法性:连续爬取时容易触发网站反爬机制,返回403/503等异常页面,此时解析后的HTML结构完全不符合预期,即便
catalog被找到,也可能是不具备find_all方法的非标签对象(如文本节点)。 - 依赖动态生成类名:代码中使用的
ProductListContainer-module__list___yHwue这类类名是前端打包工具生成的动态名称,网站更新或反爬策略调整后会失效,导致catalog定位错误。 - 无异常容错机制:DOM元素查找过程中未添加异常捕获,一旦某一步出现意外(如元素缺失),直接抛出错误中断程序。
修复方案
关键改进点
- 添加请求校验与防爬优化:检查响应状态码,模拟浏览器请求头,添加请求延迟避免频率限制。
- 替换稳定的元素定位方式:优先使用
data-test等固定属性定位元素,抛弃易变的动态类名。 - 增加异常捕获:对可能出错的DOM操作添加
try-except块,保证程序持续运行。
修改后的完整代码
import requests import bs4 import pandas as pd import time # 模拟浏览器请求头,降低被反爬拦截概率 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } url = 'https://www.hollandandbarrett.com/search/?query=prebiotic&page=1' try: resp = requests.get(url, headers=headers) # 校验请求是否成功 resp.raise_for_status() html = bs4.BeautifulSoup(resp.content, 'html.parser') row = {} # 改用data-test属性定位catalog,避免依赖动态类名 catalog = html.find('div', attrs={'data-test': 'list-Products'}) if catalog is not None: # 同样用data-test定位产品链接 products = catalog.find_all('a', attrs={'data-test': 'product-card'}) data = [] for product in products: if product is None: continue try: # 提取标题 t = product.find('div', attrs={'data-test': 'product-card-title'}) title = t.text.strip() if t else None # 提取评分 r = product.find('div', attrs={'class': 'RatingStars-module_star__3j5m8'}) rating = r['title'].strip() if r else None # 提取评论数 rr = product.find('div', attrs={'data-test': 'product-card-review-count'}) rating_review = rr.text.strip().replace('(', '').replace(')', '') if rr else None # 提取价格块,先判断价格块是否存在 price_block = product.find('div', attrs={'class': 'ProductCard-module__priceBlock___5GV3W'}) if price_block: p = price_block.findChild(attrs={'data-test': True}) price = p.text.strip() if p else None sp = price_block.findChild(attrs={'data-test': 'product-card-sale-price'}) sale_price = sp.text.strip() if sp else None else: price = None sale_price = None # 提取单价 pu = product.find('div', attrs={'data-test': 'price-per-unit'}) price_per_unit = pu.text.strip() if pu else None row = { 'title': title, 'rating': rating, 'rating_review': rating_review, 'price': price, 'sale_price': sale_price, 'price_per_unit': price_per_unit } data.append(row) except Exception as e: print(f"处理单个产品时出错: {str(e)}") continue if data: df = pd.DataFrame.from_dict(data, orient='columns') df.to_csv('hollandandbarrett_products.csv', encoding='utf-8', index=True, header=True ) print("数据已成功保存") else: print("未抓取到有效产品数据") else: print('未找到产品列表容器') except requests.exceptions.RequestException as e: print(f"请求出错: {str(e)}") except Exception as e: print(f"程序运行出错: {str(e)}") finally: # 添加请求延迟,避免频繁请求触发反爬 time.sleep(2)
代码说明
- 请求优化:添加浏览器UA头,用
raise_for_status()捕获请求异常,避免解析错误页面。 - 元素定位优化:将依赖动态类名的定位改为
data-test属性,这类属性通常是网站测试用的固定标识,不会轻易变更。 - 异常处理:外层捕获请求异常,内层捕获单个产品解析异常,保证某条产品出错不会中断整个爬取流程。
- 请求延迟:
finally块中添加2秒延迟,降低连续请求的频率,减少被反爬拦截的风险。
内容的提问来源于stack exchange,提问作者sifar
相关产品推荐
相关产品推荐

