Scrapy多爬虫多Item导出CSV异常:仅page项写入其余为空
嘿,我碰到过好几个类似的问题了——你这情况是Scrapy多Item分CSV导出时,只有page类型的Item能写入,其他文件空,但Shell里明明能看到所有Item对吧?大概率是Pipeline的判断逻辑、配置或者Item实例化出了问题,我给你拆解几个最可能的排查点:
这是最常见的坑:你可能在Pipeline里判断Item类型时,要么导入错了Item类,要么判断条件写得不对。给你一个标准的正确实现参考,你对比下自己的代码:
import csv from scrapy.exceptions import DropItem from your_project_name.items import PageItem, OtherItem, AnotherItem # 务必替换成你项目的实际Item类 class MultiCSVItemPipeline: def open_spider(self, spider): # 初始化每个Item对应的文件和写入器 self.file_handles = {} self.csv_writers = {} # 处理PageItem page_file = open('page_items.csv', 'w', newline='', encoding='utf-8') self.file_handles['page'] = page_file self.csv_writers['page'] = csv.DictWriter( page_file, fieldnames=PageItem.fields.keys() ) self.csv_writers['page'].writeheader() # 处理OtherItem(替换成你的其他Item类) other_file = open('other_items.csv', 'w', newline='', encoding='utf-8') self.file_handles['other'] = other_file self.csv_writers['other'] = csv.DictWriter( other_file, fieldnames=OtherItem.fields.keys() ) self.csv_writers['other'].writeheader() # 同理添加其他Item的处理分支... def process_item(self, item, spider): # 精准匹配Item类型 if isinstance(item, PageItem): self.csv_writers['page'].writerow(dict(item)) elif isinstance(item, OtherItem): self.csv_writers['other'].writerow(dict(item)) # 其他Item的判断分支... else: raise DropItem(f"无法识别的Item类型: {type(item).__name__}") return item def close_spider(self, spider): # 关闭所有打开的文件句柄 for file in self.file_handles.values(): file.close()
要注意这几点:
- 必须从你的项目正确导入Item类,别写错项目名或Item类名
- 每个Item对应的
fieldnames要和Item类里定义的字段完全匹配,用Item.fields.keys()能避免手动写错
打开settings.py检查ITEM_PIPELINES,确保你的MultiCSVItemPipeline是唯一启用的CSV相关Pipeline,而且优先级设置合理:
ITEM_PIPELINES = { 'your_project_name.pipelines.MultiCSVItemPipeline': 300, # 务必注释掉其他默认的CSV导出Pipeline,比如scrapy自带的CsvItemExporter相关配置 }
如果同时启用了默认的CSV导出Pipeline,它会抢在你的自定义Pipeline之前处理Item,导致只有一种Item被导出。
有时候爬虫里生成Item时,可能不小心把其他类型的Item实例成了PageItem,或者压根没正确yield其他Item。你可以在爬虫的parse方法里加个日志,确认每个yield的Item类型:
import logging def parse(self, response): # ... 你的解析逻辑 ... item = OtherItem() # 替换成你的其他Item类 item['field_name'] = '对应的值' logging.info(f"正在导出Item类型: {type(item).__name__}") yield item
运行爬虫后看日志,确认是否有非page类型的Item被yield出来。
有时候其他CSV文件为空,可能是因为文件写到了其他目录,或者你的程序没有写入权限。你可以在open_spider方法里打印文件的绝对路径:
import os def open_spider(self, spider): # ... page_file_path = os.path.abspath('page_items.csv') other_file_path = os.path.abspath('other_items.csv') logging.info(f"PageItem将写入: {page_file_path}") logging.info(f"OtherItem将写入: {other_file_path}") # ...
运行后去对应的路径下检查文件是否存在,以及是否有内容。
你可以在process_item方法里加日志,确认每个Item是否进入了正确的写入分支:
def process_item(self, item, spider): item_type = type(item).__name__ logging.info(f"正在处理Item类型: {item_type}") if isinstance(item, PageItem): logging.info("写入page_items.csv") self.csv_writers['page'].writerow(dict(item)) elif isinstance(item, OtherItem): logging.info("写入other_items.csv") self.csv_writers['other'].writerow(dict(item)) else: logging.warning(f"未匹配到的Item类型: {item_type}") raise DropItem(f"无法识别的Item类型: {item_type}") return item
运行爬虫后,从日志里就能清楚看到每个Item的流向,很快就能定位到问题出在哪个环节。
内容的提问来源于stack exchange,提问作者OLF

