Python yield后赋值报SyntaxError,Scrapy爬虫爬取结果无法保存怎么解决?
问题原因与修复方案
1. 语法错误的直接原因
Python语法不允许在yield关键字后直接跟赋值语句,你想要同时把爬取结果赋值给变量和返回给Scrapy引擎,需要把赋值和yield拆成两行执行。
2. 其他需要修复的逻辑问题
- 你在循环内给
result赋值,循环结束后result只会保留最后一条爬取到的quote数据,要保存当前页所有数据,需要先把所有结果存到列表里。 - 字典不能直接写入文件,你用了
wb二进制写入模式,需要先把数据序列化为JSON格式的字节串。 scrapy.Request的回调参数名是callback不是callable,写错会导致翻页逻辑失效。
修正后的完整代码
import scrapy import json class ExampleSpider(scrapy.Spider): name = "ufcspider" start_urls = [ 'http://quotes.toscrape.com/page/1/', ] def parse(self, response): page_results = [] for quote in response.css('div.quote'): # 先完成赋值 result = { 'text': quote.css('span.text::text').get(), 'author': quote.css('small.author::text').get(), 'link': 'http://quotes.toscrape.com' + quote.css("span a::attr(href)").get(), 'tags': quote.css('div.tags a.tag::text').getall(), } # 再执行yield返回 yield result page_results.append(result) next_page = response.css("li.next a::attr(href)").get() if next_page is not None: next_page = response.urljoin(next_page) # 修正回调参数名为callback yield scrapy.Request(next_page, callback=self.parse) page = response.url.split("/")[-2] filename = f'quotes-{page}.json' with open(filename, 'wb') as f: # 序列化为JSON字节串后写入 f.write(json.dumps(page_results, ensure_ascii=False, indent=2).encode('utf-8')) self.log(f'Saved file {filename}')
内容的提问来源于stack exchange,提问作者oyerohabib
相关产品推荐
相关产品推荐

