Scrapy输出列含空行,如何将胜负方值合并至同一行?
解决Scrapy输出空单元格&合并胜者败者到同一行的方案
兄弟,我来帮你搞定这个问题!从你的描述来看,大概率是同一赛事的胜者、败者被爬成了两个独立的Item,导致输出时各占一行有空单元格;或者Item里两个字段偶尔为空,想让非空值集中在同一行展示。下面分两种场景给你落地解决方案:
场景1:胜者败者分属不同Item,需合并到同一行
这种情况在赛事爬取里很常见——比如你从列表页分别抓取了胜者和败者的信息,需要通过一个唯一标识(比如赛事ID、赛事名称)把它们合并成一个完整的Item再输出。
第一步:完善items.py的字段
先给你的Item加一个用来关联同一赛事的字段,比如match_id:
import scrapy class MatchItem(scrapy.Item): match_id = scrapy.Field() # 同一赛事的唯一标识,比如赛事ID/名称 winner = scrapy.Field() loser = scrapy.Field() # 其他字段比如赛事时间、项目类型可以按需添加
第二步:重写Pipeline实现Item合并
写一个Pipeline来缓存未匹配的Item,等到同一match_id的两个Item都收集到后,合并成一个完整的Item再输出:
from scrapy.exceptions import DropItem class MergeMatchPipeline: def __init__(self): # 用字典缓存未匹配的Item,key是match_id,value是对应的Item self.match_cache = {} def process_item(self, item, spider): match_id = item.get('match_id') if not match_id: # 没有关联标识的Item直接放行 return item if match_id in self.match_cache: # 找到匹配的缓存Item,合并字段(优先保留非空值) cached_item = self.match_cache.pop(match_id) merged_item = scrapy.Item() merged_item['match_id'] = match_id merged_item['winner'] = cached_item.get('winner') or item.get('winner') merged_item['loser'] = cached_item.get('loser') or item.get('loser') # 其他字段也按这个逻辑合并 return merged_item else: # 当前Item暂时没有匹配项,存入缓存 self.match_cache[match_id] = item # 抛出DropItem让Scrapy暂时不输出这个Item raise DropItem(f"已缓存赛事{match_id}的部分数据,等待匹配项") def close_spider(self, spider): # 爬虫结束时,处理缓存中未匹配的剩余Item(比如只有胜者或只有败者的情况) for match_id, item in self.match_cache.items(): spider.logger.warning(f"赛事{match_id}只有部分数据(缺少胜者/败者),已单独输出") yield item
第三步:启用Pipeline
在项目的settings.py里添加这个Pipeline的配置(注意优先级数字越小,执行越靠前):
ITEM_PIPELINES = { '你的项目名.pipelines.MergeMatchPipeline': 300, # 后面可以添加CSV/JSON等导出Pipeline,比如官方的CSV导出 'scrapy.pipelines.export.CsvItemExporter': 400, }
场景2:Item同时包含胜者败者,想移除空单元格
如果你的Item本身就同时有winner和loser字段,只是偶尔其中一个为空,想让输出时只显示有值的字段(比如CSV里不生成空列),可以自定义一个导出器:
from scrapy.exporters import CsvItemExporter class CustomCsvExporter(CsvItemExporter): def __init__(self, file, include_empty_fields=False, **kwargs): # 关闭空字段的导出,只导出有值的字段 super().__init__(file, include_empty_fields=include_empty_fields, **kwargs) def export_item(self, item): # 过滤掉值为空的字段 filtered_item = {k: v for k, v in item.items() if v is not None and v != ''} super().export_item(filtered_item)
然后在Pipeline里使用这个自定义导出器:
from .your_exporter_file import CustomCsvExporter class CustomExportPipeline: def __init__(self): self.file = open('matches.csv', 'wb') self.exporter = CustomCsvExporter(self.file) self.exporter.start_exporting() def process_item(self, item, spider): self.exporter.export_item(item) return item def close_spider(self, spider): self.exporter.finish_exporting() self.file.close()
这样导出的CSV里,每行只会包含有值的字段,不会出现空单元格。
内容的提问来源于stack exchange,提问作者tomoc4
相关产品推荐
相关产品推荐

