You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy运行crawl命令无-o参数时仅输出每页首个结果如何解决

问题原因
  • 你自定义的逐页存文件逻辑存在变量覆盖问题:遍历单页10条quote的for循环中,你每次都给result变量重新赋值,没有把当前页所有结果收集起来,循环结束后写入JSON文件的只有最后一次循环得到的单条数据。
  • scrapy crawl ufcspider -o quotes.json是调用Scrapy自带的输出管道,会自动收集你代码中所有yield返回的结果,不受你自己写的存文件逻辑影响,所以输出是完整的。
修复方案

调整parse方法的逻辑,先收集当前页所有结果再写入文件即可,同时可以去掉冗余的f.close()语句(with open语法会自动关闭文件),新增的ensure_ascii=False和indent参数是为了优化JSON文件的可读性、避免特殊字符乱码。
修改后的完整代码如下:

import scrapy
import json

class ExampleSpider(scrapy.Spider):
    name = "ufcspider"

    start_urls = [
        'http://quotes.toscrape.com/page/1/',
    ]

    def parse(self, response):
        # 初始化列表存储当前页所有结果
        page_results = []
        for quote in response.css('div.quote'):
            result = {
                'text': quote.css('span.text::text').get(),
                'author': quote.css('small.author::text').get(),
                'link': 'http://quotes.toscrape.com' + quote.css("span a::attr(href)").get(),
                'tags': quote.css('div.tags a.tag::text').getall()
            }
            # 先把当前条结果加入列表
            page_results.append(result)
            yield result

        next_page = response.css("li.next a::attr(href)").get()
        if next_page is not None:
            yield response.follow(next_page, callback=self.parse)

        page = response.url.split("/")[-2]
        filename = f'quotes-{page}.json'
        with open(filename, 'w') as f:
            # 写入当前页所有结果组成的列表
            json.dump(page_results, f, ensure_ascii=False, indent=2)
        self.log(f'Saved file {filename}')

内容的提问来源于stack exchange,提问作者oyerohabib

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 11:18:03