使用Scrapy的CrawlSpider与LinkExtractor时如何解决Spider_error_processing_headers问题
爬虫报错排查:ERROR: Spider error processing
问题描述
终端报错信息:
ERROR: Spider error processing
报错位置:line 276, in aiter_errback yield await it.anext()
爬虫代码如下:
import scrapy from scrapy.linkextractors import LinkExtractor from scrapy.spiders import CrawlSpider, Rule class CandywareCrawlspiderSpider(CrawlSpider): name = "candyware_crawlspider" allowed_domains = ["www.candywarehouse.com"] # start_urls = ["https://www.candywarehouse.com/collections/wedding?page=24"] user_agent = 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_0) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.114 Safari/537.36' # Editing the user-agent in the request sent def start_requests(self): yield scrapy.Request(url='https://www.candywarehouse.com/collections/wedding?page=24', headers={ 'user-agent': self.user_agent }) # Setting rules for the crawler rules = ( Rule(LinkExtractor(restrict_xpaths=('//ul[@class="pagination-custom"]//li/a[@title="Next »"]')), callback='parse_item', follow=True, process_request='set_user_agent'),) # # # Setting the user-agent def set_user_agent(self, request, spider): request.headers['User-Agent'] = self.user_agent return request def parse_item(self, response): product_list = response.xpath('//div[@class="js-grid"]/div') for product in product_list: product_name = product.xpath('.//p[@class="product__grid__title"]/text()').get().strip() price = product.xpath('.//span[@class="price"]/text()').get().strip() review_counts = product.xpath('.//span[@class="tt-product-block__rating"]/text()').get().replace('\n', '').replace(' ', '') yield { 'product_name': product_name, 'price': price, 'review_counts': review_counts, 'User-Agent': response.request.headers['User-Agent'], }
报错原因分析
- 空值调用方法触发AttributeError:代码中使用
.get()提取元素文本时,若目标元素不存在(比如部分商品无评论数),.get()会返回None,此时直接调用.strip()或.replace()会抛出异常,导致爬虫中断。 - Rule逻辑存在小瑕疵:当前Rule将下一页链接的请求交给
parse_item处理,但未显式指定初始请求的页面解析回调(不过这不是直接报错原因,核心问题为空值处理)。
修复方案
1. 增加空值判断,避免None调用方法
修改parse_item方法中的提取逻辑,对每个.get()的结果先做非空判断,再进行字符串处理:
def parse_item(self, response): product_list = response.xpath('//div[@class="js-grid"]/div') for product in product_list: # 处理商品名称 product_name = product.xpath('.//p[@class="product__grid__title"]/text()').get() product_name = product_name.strip() if product_name else "无商品名称" # 处理价格 price = product.xpath('.//span[@class="price"]/text()').get() price = price.strip() if price else "无价格" # 处理评论数 review_counts = product.xpath('.//span[@class="tt-product-block__rating"]/text()').get() if review_counts: review_counts = review_counts.replace('\n', '').replace(' ', '').strip() else: review_counts = "0条评论" yield { 'product_name': product_name, 'price': price, 'review_counts': review_counts, 'User-Agent': response.request.headers['User-Agent'], }
2. 优化Rule逻辑(可选)
为确保初始请求的页面也被parse_item解析,可在start_requests中显式指定回调:
def start_requests(self): yield scrapy.Request( url='https://www.candywarehouse.com/collections/wedding?page=24', headers={'user-agent': self.user_agent}, callback=self.parse_item )
3. 统一User-Agent设置(可选)
在项目的settings.py中全局设置User-Agent,避免重复配置:
USER_AGENT = 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_0) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.114 Safari/537.36'
设置后可删除set_user_agent方法和start_requests中的headers配置,简化代码。
内容的提问来源于stack exchange,提问作者Moniruzzaman Monir
相关产品推荐
相关产品推荐

