Scrapy运行crawl命令无-o参数时仅输出每页首个结果如何解决
问题原因
- 你自定义的逐页存文件逻辑存在变量覆盖问题:遍历单页10条quote的for循环中,你每次都给
result变量重新赋值,没有把当前页所有结果收集起来,循环结束后写入JSON文件的只有最后一次循环得到的单条数据。 scrapy crawl ufcspider -o quotes.json是调用Scrapy自带的输出管道,会自动收集你代码中所有yield返回的结果,不受你自己写的存文件逻辑影响,所以输出是完整的。
修复方案
调整parse方法的逻辑,先收集当前页所有结果再写入文件即可,同时可以去掉冗余的f.close()语句(with open语法会自动关闭文件),新增的ensure_ascii=False和indent参数是为了优化JSON文件的可读性、避免特殊字符乱码。
修改后的完整代码如下:
import scrapy import json class ExampleSpider(scrapy.Spider): name = "ufcspider" start_urls = [ 'http://quotes.toscrape.com/page/1/', ] def parse(self, response): # 初始化列表存储当前页所有结果 page_results = [] for quote in response.css('div.quote'): result = { 'text': quote.css('span.text::text').get(), 'author': quote.css('small.author::text').get(), 'link': 'http://quotes.toscrape.com' + quote.css("span a::attr(href)").get(), 'tags': quote.css('div.tags a.tag::text').getall() } # 先把当前条结果加入列表 page_results.append(result) yield result next_page = response.css("li.next a::attr(href)").get() if next_page is not None: yield response.follow(next_page, callback=self.parse) page = response.url.split("/")[-2] filename = f'quotes-{page}.json' with open(filename, 'w') as f: # 写入当前页所有结果组成的列表 json.dump(page_results, f, ensure_ascii=False, indent=2) self.log(f'Saved file {filename}')
内容的提问来源于stack exchange,提问作者oyerohabib
相关产品推荐
相关产品推荐

