You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫TXT输出自动化:如何添加仅生成一次的固定头尾

解决Scrapy爬虫TXT文件自动添加固定头尾的问题

我来帮你搞定这个Scrapy输出TXT的需求!你的目标是给每个分类的TXT文件只写一次固定头部和尾部,中间填充爬取的新闻数据对吧?先看看你现有代码里的几个关键问题,再给你一个符合Scrapy规范的解决方案,完全不需要额外外部模块~

现有代码的问题

  1. 字符串拼接错误:你现在把变量直接写在引号里了,比如col = ('item['title'] ...')这只是个字符串字面量,并没有真正引用item里的字段;f.write('h \n + col \n + p')也是把h、col、p当成普通字符串写入,而不是用你定义的变量值。
  2. 文件打开模式问题:用'a'(追加)模式每次处理一个item就打开文件写入,会导致头尾被重复写入多次,完全不符合你“头尾仅写一次”的需求。

推荐解决方案:使用Scrapy Item Pipeline

Scrapy的Pipeline是专门用来处理爬取后的数据输出的,用它可以完美实现“先收集所有item,再一次性写入头尾+所有数据”的逻辑,这也是Scrapy的标准做法。

步骤1:编写自定义Pipeline

在你的Scrapy项目的pipelines.py文件中,添加以下代码:

class NewsTxtPipeline:
    def open_spider(self, spider):
        # 初始化字典,用来按分类存储所有item
        self.category_items = {}

    def process_item(self, item, spider):
        category = item['cat']
        # 如果当前分类还没在字典里,创建空列表
        if category not in self.category_items:
            self.category_items[category] = []
        # 将当前item添加到对应分类的列表中
        self.category_items[category].append(item)
        return item

    def close_spider(self, spider):
        # 定义固定的头尾内容
        header = "As últimas notícias"
        footer = "Você só encontra aqui"
        
        # 遍历每个分类,生成对应的TXT文件
        for category, items in self.category_items.items():
            filename = f"{category}.txt"
            # 使用with语句自动管理文件,避免手动关闭出错
            with open(filename, 'w', encoding='utf-8') as f:
                # 先写入头部
                f.write(f"{header}\n\n")
                # 遍历该分类的所有item,逐个写入内容
                for item in items:
                    # 按你需要的格式拼接每个item的内容
                    item_content = (
                        f"{item['title']}\n"
                        f"{item['author']}\n"
                        f"{item['img']}\n\n"
                        f"{item['news']}\n\n"
                    )
                    f.write(item_content)
                # 最后写入尾部
                f.write(footer)

步骤2:启用Pipeline

打开项目的settings.py文件,找到ITEM_PIPELINES配置项,添加你的自定义Pipeline(替换成你的项目名):

ITEM_PIPELINES = {
    # 格式:'你的项目名.pipelines.NewsTxtPipeline': 优先级数字
    'my_news_spider.pipelines.NewsTxtPipeline': 300,
}

优先级数字越小,Pipeline执行越早,这里给300是比较合理的数值。

为什么这个方案好用?

  • 头尾仅写一次:所有item都爬取完成后,才会一次性写入头尾和所有数据,完美符合需求。
  • 自动管理文件:用with open语句自动关闭文件,比手动调用f.close()更安全,避免资源泄漏。
  • 支持多分类:自动按item['cat']分类生成不同的TXT文件,每个文件都有独立的头尾和数据。
  • 编码安全:指定encoding='utf-8',避免葡萄牙语的重音字符或者其他特殊字符出现乱码。

备选方案:在Spider内处理(不推荐)

如果你不想用Pipeline,也可以在Spider类里用类变量收集item,在爬虫结束时写入文件,但这种方式不如Pipeline灵活(比如无法和其他数据处理Pipeline配合):

import scrapy
from your_project.items import NewsItem

class NewsSpider(scrapy.Spider):
    name = 'news_spider'
    start_urls = ['https://your-target-url.com']

    def __init__(self, *args, **kwargs):
        super().__init__(*args, **kwargs)
        self.category_items = {}

    def parse(self, response):
        # 你的解析逻辑,生成item
        item = NewsItem()
        item['cat'] = response.css('.category::text').get().strip()
        item['title'] = response.css('h1.news-title::text').get().strip()
        item['author'] = response.css('.author::text').get().strip()
        item['img'] = response.css('.news-img::attr(src)').get()
        item['news'] = response.css('.news-content::text').getall()
        # 把item按分类存储
        category = item['cat']
        if category not in self.category_items:
            self.category_items[category] = []
        self.category_items[category].append(item)
        yield item

    def closed(self, reason):
        header = "As últimas notícias"
        footer = "Você só encontra aqui"
        for category, items in self.category_items.items():
            filename = f"{category}.txt"
            with open(filename, 'w', encoding='utf-8') as f:
                f.write(f"{header}\n\n")
                for item in items:
                    item_content = (
                        f"{item['title']}\n"
                        f"{item['author']}\n"
                        f"{item['img']}\n\n"
                        f"{''.join(item['news'])}\n\n"
                    )
                    f.write(item_content)
                f.write(footer)

内容的提问来源于stack exchange,提问作者Antonio Oliveira

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:20:10