You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

XMLFeedSpider无法生成CSV输出及JSON解码错误求助

问题解决:XMLFeedSpider运行报错JSONDecodeError及爬取异常

核心错误原因

报错json.decoder.JSONDecodeError是因为Scrapy的FEEDS配置格式错误,你直接将FEEDS设为字符串'output_file.csv',但Scrapy要求FEEDS是一个字典,键为输出文件路径,值为包含格式等参数的子字典。

修正后的完整代码

蜘蛛代码

from ..items import TheItem
from scrapy.loader import ItemLoader
import scrapy
from scrapy.crawler import CrawlerProcess


class TheSpider(scrapy.spiders.XMLFeedSpider):
    name = 'stuff_spider'
    allowed_domains = ['www.website.net']
    start_urls = ['https://www.website.net/10016/stuff/otherstuff.xml']
    namespaces = [('xsi', 'https://schemas.website.net/xml/uslm'), ]
    itertag = 'xsi:item'
    iterator = 'xml'

    # XMLFeedSpider会自动处理请求并调用parse_node,无需自定义start_requests
    # def start_requests(self):
    #     yield scrapy.Request('https://www.website.net/10016/stuff/otherstuff.xml', callback=self.parse_node)

    def parse_node(self, response, node):
        l = ItemLoader(item=TheItem(), selector=node, response=response)

        just_want_something = 'just want the csv to show some output'

        # 使用相对路径xpath(.//而非//)避免全局匹配,且无需手动调用extract()
        l.add_xpath('title', './/xsi:title/text()')
        l.add_xpath('date', './/xsi:date/text()')
        l.add_xpath('category', './/xsi:cat1/text()')
        l.add_xpath('content', './/xsi:content/text()')
        l.add_value('manditory', just_want_something)

        yield l.load_item()


process = CrawlerProcess(settings={
    # 修正FEEDS为字典格式
    'FEEDS': {
        'output_file.csv': {
            'format': 'csv',
            'encoding': 'utf-8'
        }
    },
    'DOWNLOAD_DELAY': 1.25,
    'USER_AGENT': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:102.0) Gecko/20100101 Firefox/102.0'
})

process.crawl(TheSpider)
process.start()

Item代码(保留原代码即可)

from scrapy import Item, Field
from itemloaders.processors import Identity, Compose


def all_lower(value):
    return value.lower()


class TheItem(Item):
    title = Field(
        input_processor=Compose(all_lower),
        output_processor=Identity()
    )
    link = Field(
        input_processor=Compose(all_lower),
        output_processor=Identity()
    )
    date = Field(
        input_processor=Compose(all_lower),
        output_processor=Identity()
    )
    category = Field(
        input_processor=Compose(all_lower),
        output_processor=Identity()
    )
    manditory = Field(
        input_processor=Compose(all_lower),
        output_processor=Identity()
    )

额外优化说明

  1. 移除自定义start_requests:XMLFeedSpider会自动遍历start_urls,解析后将节点传入parse_node,无需手动指定callback。
  2. 修正XPath路径:用.//相对路径替代//全局路径,确保仅在当前节点下匹配内容,避免重复抓取。
  3. 规范ItemLoader用法:add_xpath无需手动调用extract(),ItemLoader会自动处理选择器结果。
  4. FEEDS配置优化:将格式、编码等参数整合到FEEDS字典中,无需单独设置FEED_FORMAT,避免冗余配置。

内容的提问来源于stack exchange,提问作者SC4RECROW

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 21:05:22