XMLFeedSpider无法生成CSV输出及JSON解码错误求助
问题解决:XMLFeedSpider运行报错JSONDecodeError及爬取异常
核心错误原因
报错json.decoder.JSONDecodeError是因为Scrapy的FEEDS配置格式错误,你直接将FEEDS设为字符串'output_file.csv',但Scrapy要求FEEDS是一个字典,键为输出文件路径,值为包含格式等参数的子字典。
修正后的完整代码
蜘蛛代码
from ..items import TheItem from scrapy.loader import ItemLoader import scrapy from scrapy.crawler import CrawlerProcess class TheSpider(scrapy.spiders.XMLFeedSpider): name = 'stuff_spider' allowed_domains = ['www.website.net'] start_urls = ['https://www.website.net/10016/stuff/otherstuff.xml'] namespaces = [('xsi', 'https://schemas.website.net/xml/uslm'), ] itertag = 'xsi:item' iterator = 'xml' # XMLFeedSpider会自动处理请求并调用parse_node,无需自定义start_requests # def start_requests(self): # yield scrapy.Request('https://www.website.net/10016/stuff/otherstuff.xml', callback=self.parse_node) def parse_node(self, response, node): l = ItemLoader(item=TheItem(), selector=node, response=response) just_want_something = 'just want the csv to show some output' # 使用相对路径xpath(.//而非//)避免全局匹配,且无需手动调用extract() l.add_xpath('title', './/xsi:title/text()') l.add_xpath('date', './/xsi:date/text()') l.add_xpath('category', './/xsi:cat1/text()') l.add_xpath('content', './/xsi:content/text()') l.add_value('manditory', just_want_something) yield l.load_item() process = CrawlerProcess(settings={ # 修正FEEDS为字典格式 'FEEDS': { 'output_file.csv': { 'format': 'csv', 'encoding': 'utf-8' } }, 'DOWNLOAD_DELAY': 1.25, 'USER_AGENT': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:102.0) Gecko/20100101 Firefox/102.0' }) process.crawl(TheSpider) process.start()
Item代码(保留原代码即可)
from scrapy import Item, Field from itemloaders.processors import Identity, Compose def all_lower(value): return value.lower() class TheItem(Item): title = Field( input_processor=Compose(all_lower), output_processor=Identity() ) link = Field( input_processor=Compose(all_lower), output_processor=Identity() ) date = Field( input_processor=Compose(all_lower), output_processor=Identity() ) category = Field( input_processor=Compose(all_lower), output_processor=Identity() ) manditory = Field( input_processor=Compose(all_lower), output_processor=Identity() )
额外优化说明
- 移除自定义start_requests:XMLFeedSpider会自动遍历start_urls,解析后将节点传入parse_node,无需手动指定callback。
- 修正XPath路径:用
.//相对路径替代//全局路径,确保仅在当前节点下匹配内容,避免重复抓取。 - 规范ItemLoader用法:
add_xpath无需手动调用extract(),ItemLoader会自动处理选择器结果。 - FEEDS配置优化:将格式、编码等参数整合到FEEDS字典中,无需单独设置FEED_FORMAT,避免冗余配置。
内容的提问来源于stack exchange,提问作者SC4RECROW
相关产品推荐
相关产品推荐

