You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy在Jupyter Notebook中无法保存JSON文件求助

解决Scrapy脚本完成后未生成JSON文件的问题

核心问题与修复步骤

1. 代码缩进错误(导致无数据输出)

你的代码存在多处缩进问题,这是爬虫未生成数据、进而无法创建JSON文件的核心原因:

  • def parse(self, response): 前多了4个空格,需与类内其他属性(如name)对齐
  • datos字典定义和yield datos语句都在for循环外,且yield缩进错误,导致爬虫根本不会输出任何爬取到的数据

修复后的parse方法:

def parse(self, response):
    events = response.xpath('//*[@id="PageContent"]/div[3]/table/tbody/tr')
    for event in events:
        dato1 = event.xpath('.//td[1]/text()').get()
        dato2 = event.xpath('.//td[2]/text()').get()

        datos = {
            'Dato 1': dato1.strip() if dato1 else None,
            'Dato 2': dato2.strip() if dato2 else None,
        }

        yield datos

2. FEEDS路径配置问题

你已挂载Google Drive,但代码中使用相对路径data.json,Jupyter Notebook默认工作目录不在Drive挂载路径下,导致文件生成在未挂载的本地目录(你无法找到),或因权限问题无法写入。

需修改为Drive的绝对路径,示例:

custom_settings = {
    'FEEDS': { '/content/drive/MyDrive/data.json': { 'format': 'json', 'overwrite': True}}
}

(注意替换为你实际的Drive挂载路径,通常挂载后路径为/content/drive/MyDrive/)

3. 可选:显式配置CrawlerProcess

若担心Spider的custom_settings未生效,可在创建CrawlerProcess时直接传入配置:

process = CrawlerProcess(settings={
    'FEEDS': { '/content/drive/MyDrive/data.json': { 'format': 'json', 'overwrite': True}}
})

完整修复代码

import scrapy
from scrapy.crawler import CrawlerProcess


class DatosSpider(scrapy.Spider):
    name = 'spider_datos'
    start_urls = ['你的目标URL']  # 替换为实际爬取URL
    custom_settings = {
        'FEEDS': { '/content/drive/MyDrive/data.json': { 'format': 'json', 'overwrite': True}}
    }

    def parse(self, response):
        events = response.xpath('//*[@id="PageContent"]/div[3]/table/tbody/tr')
        for event in events:
            dato1 = event.xpath('.//td[1]/text()').get()
            dato2 = event.xpath('.//td[2]/text()').get()

            datos = {
                'Dato 1': dato1.strip() if dato1 else None,
                'Dato 2': dato2.strip() if dato2 else None,
            }

            yield datos
        
process = CrawlerProcess()
process.crawl(DatosSpider)
process.start()

内容的提问来源于stack exchange,提问作者Regue

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 20:51:10