You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy按输入文件动态配置FEEDS导出时Pipeline读取为空问题求助

问题根因
  • 配置修改时机错误:Scrapy 的 custom_settings 会在 Spider 类加载阶段就被框架读取、合并到全局配置中,之后才会初始化 Pipeline、导出器等核心组件,最后才执行 Spider 的 __init__ 方法。你在 __init__ 中修改 self.custom_settings 时,配置已经被冻结,所有组件都读取不到新增的 FEEDS 配置。
  • 拼写与逻辑错误:
    1. Pipeline 中判断导出格式时拼写错误,将 custom 写为了 cistom
    2. 遍历 FEEDS 时混淆了键值:FEEDS 的键是输出文件路径,值是格式配置,判断 k["format"] 属于逻辑错误
    3. Spider 代码中使用了 Path 但未导入 pathlib.Path
  • 导出器未注册:Scrapy 默认识别不到 custom 格式,需要在配置中注册你的自定义导出器。

修复方案

1. 调整Spider配置修改逻辑

将动态设置FEEDS的逻辑移到from_crawler类方法中,该方法会在组件初始化前执行,可以正常修改配置:

import scrapy
from pathlib import Path

class VideosSpider(scrapy.Spider):
    name = "videos"
    start_urls = []
    custom_settings = {
        "ITEM_PIPELINES": {
            "project.pipelines.VideosPipeline": 200,
            "project.pipelines.EpisodeImagesPipeline": 300,
        },
        "FILES_STORE": "episodes",
        "IMAGES_STORE": "episodes",
        # 注册自定义导出器,替换为你CustomExporter的实际导入路径
        "FEED_EXPORTERS": {
            "custom": "project.exporters.CustomExporter",
        }
    }

    @classmethod
    def from_crawler(cls, crawler, *args, **kwargs):
        spider = super().from_crawler(crawler, *args, **kwargs)
        urls_file = kwargs.get('urls_file')
        if urls_file:
            # 读取start_urls
            with open(urls_file, "r") as f:
                next(f)
                spider.start_urls = [url.strip() for url in f.readlines()]
            # 动态修改FEEDS配置
            crawler.settings.set("FEEDS", {
                Path(urls_file).with_suffix(".txt"): {
                    "format": "custom",
                },
            }, priority="spider")
        return spider

注意:原来__init__方法中读取start_urls和设置FEEDS的逻辑都移到了from_crawler中,不需要保留原来的__init__方法。

2. 修复Pipeline错误

修正拼写和遍历逻辑:

class VideosPipeline:
    def __init__(self, uri, files):
        self.uri = uri
        self.files = files

    @classmethod
    def from_crawler(cls, crawler):
        return cls(
            uri=crawler.settings.get("FEEDS"),
            files=crawler.settings.get("FILES_STORE"),
        )

    def open_spider(self, spider):
        # 修正遍历逻辑和拼写错误
        feeds = [path for path, conf in self.uri.items() if conf.get("format") == "custom"]
        self.exporter = CustomExporter(feeds[0])

额外说明

如果你不需要自定义Pipeline处理导出逻辑,直接用Scrapy自带的FEED功能即可,不需要自己在Pipeline中初始化Exporter,Scrapy会自动根据你配置的FEEDS导出数据。

内容的提问来源于stack exchange,提问作者devster

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 02:09:00