Scrapy多站点爬取时callback回调函数传参报错该如何解决?
错误原因
scrapy.Request的callback参数要求传入可调用的函数对象,当前代码遍历urls字典拿到的cb是字符串类型的函数名,不符合参数要求,直接触发类型报错- 类外定义的全局
AppItem实例存在并发风险,多请求同时处理时会出现数据相互覆盖的问题 - 原代码中
response.css('title')拿到的是选择器对象,需要额外处理才能提取到可存入Item的文本内容
代码修复方案
方案1:最小改动适配原有字典结构
仅修改start_requests方法,通过getattr从类实例中获取对应名称的方法作为回调对象,同时调整Item实例化逻辑避免数据冲突:
import scrapy from ..items import AppItem urls = { 'fun1': 'http://example1.com', 'fun2': 'https://example2.com', # 新增站点直接加对应键值对即可 } class Bot(scrapy.Spider): name = 'app' def start_requests(self): for cb_name, url in urls.items(): # 通过getattr获取类实例的可调用方法对象 callback_func = getattr(self, cb_name) yield scrapy.Request(url=url, callback=callback_func) def fun1(self, response): # 每个回调内单独实例化Item,避免数据覆盖 item = AppItem() # 提取title标签的文本内容 item['title'] = response.css('title::text').get() yield item def fun2(self, response): item = AppItem() item['title'] = response.css('title::text').get() yield item
方案2:更易维护的列表式回调管理
直接将URL和对应的处理函数绑定为列表项,新增站点时无需维护函数名字符串,减少出错概率:
import scrapy from ..items import AppItem class Bot(scrapy.Spider): name = 'app' # 列表形式绑定站点和对应处理函数,新增站点直接追加元组即可 site_configs = [ ('http://example1.com', 'fun1'), ('https://example2.com', 'fun2'), ] def start_requests(self): for url, cb_name in self.site_configs: callback_func = getattr(self, cb_name) yield scrapy.Request(url=url, callback=callback_func) def fun1(self, response): item = AppItem() item['title'] = response.css('title::text').get() yield item def fun2(self, response): item = AppItem() item['title'] = response.css('title::text').get() yield item
内容的提问来源于stack exchange,提问作者bootoo
相关产品推荐
相关产品推荐

