Scrapy在Jupyter Notebook中无法保存JSON文件求助
解决Scrapy脚本完成后未生成JSON文件的问题
核心问题与修复步骤
1. 代码缩进错误(导致无数据输出)
你的代码存在多处缩进问题,这是爬虫未生成数据、进而无法创建JSON文件的核心原因:
def parse(self, response):前多了4个空格,需与类内其他属性(如name)对齐datos字典定义和yield datos语句都在for循环外,且yield缩进错误,导致爬虫根本不会输出任何爬取到的数据
修复后的parse方法:
def parse(self, response): events = response.xpath('//*[@id="PageContent"]/div[3]/table/tbody/tr') for event in events: dato1 = event.xpath('.//td[1]/text()').get() dato2 = event.xpath('.//td[2]/text()').get() datos = { 'Dato 1': dato1.strip() if dato1 else None, 'Dato 2': dato2.strip() if dato2 else None, } yield datos
2. FEEDS路径配置问题
你已挂载Google Drive,但代码中使用相对路径data.json,Jupyter Notebook默认工作目录不在Drive挂载路径下,导致文件生成在未挂载的本地目录(你无法找到),或因权限问题无法写入。
需修改为Drive的绝对路径,示例:
custom_settings = { 'FEEDS': { '/content/drive/MyDrive/data.json': { 'format': 'json', 'overwrite': True}} }
(注意替换为你实际的Drive挂载路径,通常挂载后路径为/content/drive/MyDrive/)
3. 可选:显式配置CrawlerProcess
若担心Spider的custom_settings未生效,可在创建CrawlerProcess时直接传入配置:
process = CrawlerProcess(settings={ 'FEEDS': { '/content/drive/MyDrive/data.json': { 'format': 'json', 'overwrite': True}} })
完整修复代码
import scrapy from scrapy.crawler import CrawlerProcess class DatosSpider(scrapy.Spider): name = 'spider_datos' start_urls = ['你的目标URL'] # 替换为实际爬取URL custom_settings = { 'FEEDS': { '/content/drive/MyDrive/data.json': { 'format': 'json', 'overwrite': True}} } def parse(self, response): events = response.xpath('//*[@id="PageContent"]/div[3]/table/tbody/tr') for event in events: dato1 = event.xpath('.//td[1]/text()').get() dato2 = event.xpath('.//td[2]/text()').get() datos = { 'Dato 1': dato1.strip() if dato1 else None, 'Dato 2': dato2.strip() if dato2 else None, } yield datos process = CrawlerProcess() process.crawl(DatosSpider) process.start()
内容的提问来源于stack exchange,提问作者Regue
相关产品推荐
相关产品推荐

