You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬虫代码偶发AttributeError: 'NoneType'无find_all属性问题排查及解决

问题分析与修复方案

报错成因

  • 未校验请求合法性:连续爬取时容易触发网站反爬机制,返回403/503等异常页面,此时解析后的HTML结构完全不符合预期,即便catalog被找到,也可能是不具备find_all方法的非标签对象(如文本节点)。
  • 依赖动态生成类名:代码中使用的ProductListContainer-module__list___yHwue这类类名是前端打包工具生成的动态名称,网站更新或反爬策略调整后会失效,导致catalog定位错误。
  • 无异常容错机制:DOM元素查找过程中未添加异常捕获,一旦某一步出现意外(如元素缺失),直接抛出错误中断程序。

修复方案

关键改进点

  1. 添加请求校验与防爬优化:检查响应状态码,模拟浏览器请求头,添加请求延迟避免频率限制。
  2. 替换稳定的元素定位方式:优先使用data-test等固定属性定位元素,抛弃易变的动态类名。
  3. 增加异常捕获:对可能出错的DOM操作添加try-except块,保证程序持续运行。

修改后的完整代码

import requests
import bs4
import pandas as pd
import time

# 模拟浏览器请求头,降低被反爬拦截概率
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}

url = 'https://www.hollandandbarrett.com/search/?query=prebiotic&page=1'
try:
    resp = requests.get(url, headers=headers)
    # 校验请求是否成功
    resp.raise_for_status()
    html = bs4.BeautifulSoup(resp.content, 'html.parser')
    row = {}

    # 改用data-test属性定位catalog,避免依赖动态类名
    catalog = html.find('div', attrs={'data-test': 'list-Products'})
    if catalog is not None:
        # 同样用data-test定位产品链接
        products = catalog.find_all('a', attrs={'data-test': 'product-card'})
        data = []
        for product in products: 
            if product is None:
                continue
            try:
                # 提取标题
                t = product.find('div', attrs={'data-test': 'product-card-title'})
                title = t.text.strip() if t else None

                # 提取评分
                r = product.find('div', attrs={'class': 'RatingStars-module_star__3j5m8'})
                rating = r['title'].strip() if r else None
                
                # 提取评论数
                rr = product.find('div', attrs={'data-test': 'product-card-review-count'})
                rating_review = rr.text.strip().replace('(', '').replace(')', '') if rr else None

                # 提取价格块,先判断价格块是否存在
                price_block = product.find('div', attrs={'class': 'ProductCard-module__priceBlock___5GV3W'})
                if price_block:
                    p = price_block.findChild(attrs={'data-test': True})
                    price = p.text.strip() if p else None
                    
                    sp = price_block.findChild(attrs={'data-test': 'product-card-sale-price'})
                    sale_price = sp.text.strip() if sp else None
                else:
                    price = None
                    sale_price = None

                # 提取单价
                pu = product.find('div', attrs={'data-test': 'price-per-unit'})
                price_per_unit = pu.text.strip() if pu else None

                row = {
                    'title': title,
                    'rating': rating,
                    'rating_review': rating_review,
                    'price': price,
                    'sale_price': sale_price,
                    'price_per_unit': price_per_unit
                }
                data.append(row)
            except Exception as e:
                print(f"处理单个产品时出错: {str(e)}")
                continue

        if data:
            df = pd.DataFrame.from_dict(data, orient='columns')
            df.to_csv('hollandandbarrett_products.csv', encoding='utf-8', index=True, header=True )
            print("数据已成功保存")
        else:
            print("未抓取到有效产品数据")
    else:
        print('未找到产品列表容器')
except requests.exceptions.RequestException as e:
    print(f"请求出错: {str(e)}")
except Exception as e:
    print(f"程序运行出错: {str(e)}")
finally:
    # 添加请求延迟,避免频繁请求触发反爬
    time.sleep(2)

代码说明

  • 请求优化:添加浏览器UA头,用raise_for_status()捕获请求异常,避免解析错误页面。
  • 元素定位优化:将依赖动态类名的定位改为data-test属性,这类属性通常是网站测试用的固定标识,不会轻易变更。
  • 异常处理:外层捕获请求异常,内层捕获单个产品解析异常,保证某条产品出错不会中断整个爬取流程。
  • 请求延迟:finally块中添加2秒延迟,降低连续请求的频率,减少被反爬拦截的风险。

内容的提问来源于stack exchange,提问作者sifar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 04:58:15