You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python yield后赋值报SyntaxError,Scrapy爬虫爬取结果无法保存怎么解决?

问题原因与修复方案

1. 语法错误的直接原因

Python语法不允许在yield关键字后直接跟赋值语句,你想要同时把爬取结果赋值给变量和返回给Scrapy引擎,需要把赋值和yield拆成两行执行。

2. 其他需要修复的逻辑问题

  • 你在循环内给result赋值,循环结束后result只会保留最后一条爬取到的quote数据,要保存当前页所有数据,需要先把所有结果存到列表里。
  • 字典不能直接写入文件,你用了wb二进制写入模式,需要先把数据序列化为JSON格式的字节串。
  • scrapy.Request的回调参数名是callback不是callable,写错会导致翻页逻辑失效。

修正后的完整代码

import scrapy
import json

class ExampleSpider(scrapy.Spider):
    name = "ufcspider"

    start_urls = [
        'http://quotes.toscrape.com/page/1/',
    ]

    def parse(self, response):
        page_results = []
        for quote in response.css('div.quote'):
            # 先完成赋值
            result = {
                'text': quote.css('span.text::text').get(),
                'author': quote.css('small.author::text').get(),
                'link': 'http://quotes.toscrape.com' + quote.css("span a::attr(href)").get(),
                'tags': quote.css('div.tags a.tag::text').getall(),
            }
            # 再执行yield返回
            yield result
            page_results.append(result)

        next_page = response.css("li.next a::attr(href)").get()
        if next_page is not None:
            next_page = response.urljoin(next_page)
            # 修正回调参数名为callback
            yield scrapy.Request(next_page, callback=self.parse)

        page = response.url.split("/")[-2]
        filename = f'quotes-{page}.json'
        with open(filename, 'wb') as f:
            # 序列化为JSON字节串后写入
            f.write(json.dumps(page_results, ensure_ascii=False, indent=2).encode('utf-8'))
        self.log(f'Saved file {filename}')

内容的提问来源于stack exchange,提问作者oyerohabib

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 11:06:01