You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy的CrawlSpider与LinkExtractor时如何解决Spider_error_processing_headers问题

爬虫报错排查:ERROR: Spider error processing

问题描述

终端报错信息:
ERROR: Spider error processing
报错位置:line 276, in aiter_errback yield await it.anext()

爬虫代码如下:

import scrapy
from scrapy.linkextractors import LinkExtractor
from scrapy.spiders import CrawlSpider, Rule


class CandywareCrawlspiderSpider(CrawlSpider):
    name = "candyware_crawlspider"
    allowed_domains = ["www.candywarehouse.com"]
    # start_urls = ["https://www.candywarehouse.com/collections/wedding?page=24"]

    user_agent = 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_0) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.114 Safari/537.36'

    # Editing the user-agent in the request sent
    def start_requests(self):
        yield scrapy.Request(url='https://www.candywarehouse.com/collections/wedding?page=24', headers={
            'user-agent': self.user_agent
        })

    # Setting rules for the crawler
    rules = (
        Rule(LinkExtractor(restrict_xpaths=('//ul[@class="pagination-custom"]//li/a[@title="Next »"]')), callback='parse_item', follow=True, process_request='set_user_agent'),)
    #
    # # Setting the user-agent
    def set_user_agent(self, request, spider):
        request.headers['User-Agent'] = self.user_agent
        return request

    def parse_item(self, response):

        product_list = response.xpath('//div[@class="js-grid"]/div')

        for product in product_list:
            product_name = product.xpath('.//p[@class="product__grid__title"]/text()').get().strip()
            price = product.xpath('.//span[@class="price"]/text()').get().strip()
            review_counts = product.xpath('.//span[@class="tt-product-block__rating"]/text()').get().replace('\n', '').replace('   ', '')

            yield {
                'product_name': product_name,
                'price': price,
                'review_counts': review_counts,
                'User-Agent': response.request.headers['User-Agent'],
            }

报错原因分析

  • 空值调用方法触发AttributeError:代码中使用.get()提取元素文本时,若目标元素不存在(比如部分商品无评论数),.get()会返回None,此时直接调用.strip()或.replace()会抛出异常,导致爬虫中断。
  • Rule逻辑存在小瑕疵:当前Rule将下一页链接的请求交给parse_item处理,但未显式指定初始请求的页面解析回调(不过这不是直接报错原因,核心问题为空值处理)。

修复方案

1. 增加空值判断,避免None调用方法

修改parse_item方法中的提取逻辑,对每个.get()的结果先做非空判断,再进行字符串处理:

def parse_item(self, response):
    product_list = response.xpath('//div[@class="js-grid"]/div')

    for product in product_list:
        # 处理商品名称
        product_name = product.xpath('.//p[@class="product__grid__title"]/text()').get()
        product_name = product_name.strip() if product_name else "无商品名称"
        
        # 处理价格
        price = product.xpath('.//span[@class="price"]/text()').get()
        price = price.strip() if price else "无价格"
        
        # 处理评论数
        review_counts = product.xpath('.//span[@class="tt-product-block__rating"]/text()').get()
        if review_counts:
            review_counts = review_counts.replace('\n', '').replace('   ', '').strip()
        else:
            review_counts = "0条评论"

        yield {
            'product_name': product_name,
            'price': price,
            'review_counts': review_counts,
            'User-Agent': response.request.headers['User-Agent'],
        }

2. 优化Rule逻辑(可选)

为确保初始请求的页面也被parse_item解析,可在start_requests中显式指定回调:

def start_requests(self):
    yield scrapy.Request(
        url='https://www.candywarehouse.com/collections/wedding?page=24',
        headers={'user-agent': self.user_agent},
        callback=self.parse_item
    )

3. 统一User-Agent设置(可选)

在项目的settings.py中全局设置User-Agent,避免重复配置:

USER_AGENT = 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_0) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.114 Safari/537.36'

设置后可删除set_user_agent方法和start_requests中的headers配置,简化代码。

内容的提问来源于stack exchange,提问作者Moniruzzaman Monir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 06:55:12