Scrapy爬虫TXT输出自动化:如何添加仅生成一次的固定头尾
解决Scrapy爬虫TXT文件自动添加固定头尾的问题
我来帮你搞定这个Scrapy输出TXT的需求!你的目标是给每个分类的TXT文件只写一次固定头部和尾部,中间填充爬取的新闻数据对吧?先看看你现有代码里的几个关键问题,再给你一个符合Scrapy规范的解决方案,完全不需要额外外部模块~
现有代码的问题
- 字符串拼接错误:你现在把变量直接写在引号里了,比如
col = ('item['title'] ...')这只是个字符串字面量,并没有真正引用item里的字段;f.write('h \n + col \n + p')也是把h、col、p当成普通字符串写入,而不是用你定义的变量值。 - 文件打开模式问题:用
'a'(追加)模式每次处理一个item就打开文件写入,会导致头尾被重复写入多次,完全不符合你“头尾仅写一次”的需求。
推荐解决方案:使用Scrapy Item Pipeline
Scrapy的Pipeline是专门用来处理爬取后的数据输出的,用它可以完美实现“先收集所有item,再一次性写入头尾+所有数据”的逻辑,这也是Scrapy的标准做法。
步骤1:编写自定义Pipeline
在你的Scrapy项目的pipelines.py文件中,添加以下代码:
class NewsTxtPipeline: def open_spider(self, spider): # 初始化字典,用来按分类存储所有item self.category_items = {} def process_item(self, item, spider): category = item['cat'] # 如果当前分类还没在字典里,创建空列表 if category not in self.category_items: self.category_items[category] = [] # 将当前item添加到对应分类的列表中 self.category_items[category].append(item) return item def close_spider(self, spider): # 定义固定的头尾内容 header = "As últimas notícias" footer = "Você só encontra aqui" # 遍历每个分类,生成对应的TXT文件 for category, items in self.category_items.items(): filename = f"{category}.txt" # 使用with语句自动管理文件,避免手动关闭出错 with open(filename, 'w', encoding='utf-8') as f: # 先写入头部 f.write(f"{header}\n\n") # 遍历该分类的所有item,逐个写入内容 for item in items: # 按你需要的格式拼接每个item的内容 item_content = ( f"{item['title']}\n" f"{item['author']}\n" f"{item['img']}\n\n" f"{item['news']}\n\n" ) f.write(item_content) # 最后写入尾部 f.write(footer)
步骤2:启用Pipeline
打开项目的settings.py文件,找到ITEM_PIPELINES配置项,添加你的自定义Pipeline(替换成你的项目名):
ITEM_PIPELINES = { # 格式:'你的项目名.pipelines.NewsTxtPipeline': 优先级数字 'my_news_spider.pipelines.NewsTxtPipeline': 300, }
优先级数字越小,Pipeline执行越早,这里给300是比较合理的数值。
为什么这个方案好用?
- 头尾仅写一次:所有item都爬取完成后,才会一次性写入头尾和所有数据,完美符合需求。
- 自动管理文件:用
with open语句自动关闭文件,比手动调用f.close()更安全,避免资源泄漏。 - 支持多分类:自动按
item['cat']分类生成不同的TXT文件,每个文件都有独立的头尾和数据。 - 编码安全:指定
encoding='utf-8',避免葡萄牙语的重音字符或者其他特殊字符出现乱码。
备选方案:在Spider内处理(不推荐)
如果你不想用Pipeline,也可以在Spider类里用类变量收集item,在爬虫结束时写入文件,但这种方式不如Pipeline灵活(比如无法和其他数据处理Pipeline配合):
import scrapy from your_project.items import NewsItem class NewsSpider(scrapy.Spider): name = 'news_spider' start_urls = ['https://your-target-url.com'] def __init__(self, *args, **kwargs): super().__init__(*args, **kwargs) self.category_items = {} def parse(self, response): # 你的解析逻辑,生成item item = NewsItem() item['cat'] = response.css('.category::text').get().strip() item['title'] = response.css('h1.news-title::text').get().strip() item['author'] = response.css('.author::text').get().strip() item['img'] = response.css('.news-img::attr(src)').get() item['news'] = response.css('.news-content::text').getall() # 把item按分类存储 category = item['cat'] if category not in self.category_items: self.category_items[category] = [] self.category_items[category].append(item) yield item def closed(self, reason): header = "As últimas notícias" footer = "Você só encontra aqui" for category, items in self.category_items.items(): filename = f"{category}.txt" with open(filename, 'w', encoding='utf-8') as f: f.write(f"{header}\n\n") for item in items: item_content = ( f"{item['title']}\n" f"{item['author']}\n" f"{item['img']}\n\n" f"{''.join(item['news'])}\n\n" ) f.write(item_content) f.write(footer)
内容的提问来源于stack exchange,提问作者Antonio Oliveira
相关产品推荐
相关产品推荐

