Scrapy中如何将birth_date合并到主parse函数的输出结果?
问题:Scrapy爬虫无法将birth_date字段与主数据一同输出
作为Python及Scrapy库的初学者,我尝试将爬取到的作者birth_date字段,与包含title、author、url_author等的主输出数据一同yield,但输出结果中始终缺少birth_date字段。
原爬虫代码
import scrapy class QuotespiderSpider(scrapy.Spider): name = "quotespider" allowed_domains = ["quotes.toscrape.com"] start_urls = ["https://quotes.toscrape.com"] def parse(self, response): quotes = response.css(".quote") all_url_tags = quotes.css(".tags a.tag::attr(href)").getall() for quote in quotes: url_tags = [] tags = (quote.css(".tags meta.keywords::attr(content)").get()).split(',') url_tags = [j for i in tags for j in all_url_tags if i in j] url_author = quote.css("span a::attr(href)").get() author_page = "https://quotes.toscrape.com/" + url_author items = { 'title' : quote.css("span.text::text").get(), 'author' : quote.css("span small.author::text").get(), 'url_author' : quote.css("span a::attr(href)").get(), 'tags' : tags, 'url_tags' : url_tags } yield scrapy.Request(url=author_page, callback=self.parse_author_page, meta={'items':items}) yield items next_page = response.css("li.next a::attr(href)").get() if next_page is not None: next_page_url = "https://quotes.toscrape.com/" + next_page yield response.follow(next_page_url, callback=self.parse) def parse_author_page(self, response): items = response.meta['items'] birth_date = response.css(".author-details p span.author-born-date::text").get() items['birth_date'] = birth_date return items
当前输出示例(缺少birth_date)
{"title": "“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”", "author": "Albert Einstein", "url_author": "/author/Albert-Einstein", "tags": ["change", "deep-thoughts", "thinking", "world"], "url_tags": ["/tag/change/page/1/", "/tag/deep-thoughts/page/1/", "/tag/thinking/page/1/", "/tag/world/page/1/"]}
解决方法
问题核心:你在parse函数里先yield了未包含birth_date的items,而parse_author_page中仅返回补充完birth_date的items,但未将其yield出去,导致最终输出只有原始的不完整数据。
修改方案:去掉parse里的yield items,转而在parse_author_page处理完birth_date后,统一yield完整的items。
修改后的代码
import scrapy class QuotespiderSpider(scrapy.Spider): name = "quotespider" allowed_domains = ["quotes.toscrape.com"] start_urls = ["https://quotes.toscrape.com"] def parse(self, response): quotes = response.css(".quote") all_url_tags = quotes.css(".tags a.tag::attr(href)").getall() for quote in quotes: url_tags = [] tags = (quote.css(".tags meta.keywords::attr(content)").get()).split(',') url_tags = [j for i in tags for j in all_url_tags if i in j] url_author = quote.css("span a::attr(href)").get() author_page = "https://quotes.toscrape.com/" + url_author items = { 'title' : quote.css("span.text::text").get(), 'author' : quote.css("span small.author::text").get(), 'url_author' : quote.css("span a::attr(href)").get(), 'tags' : tags, 'url_tags' : url_tags } # 仅发送请求到作者页,不在此处输出原始items yield scrapy.Request(url=author_page, callback=self.parse_author_page, meta={'items':items}) next_page = response.css("li.next a::attr(href)").get() if next_page is not None: next_page_url = "https://quotes.toscrape.com/" + next_page yield response.follow(next_page_url, callback=self.parse) def parse_author_page(self, response): items = response.meta['items'] birth_date = response.css(".author-details p span.author-born-date::text").get() items['birth_date'] = birth_date # 此处输出包含birth_date的完整数据 yield items
修改后输出示例(包含birth_date)
{"title": "“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”", "author": "Albert Einstein", "url_author": "/author/Albert-Einstein", "tags": ["change", "deep-thoughts", "thinking", "world"], "url_tags": ["/tag/change/page/1/", "/tag/deep-thoughts/page/1/", "/tag/thinking/page/1/", "/tag/world/page/1/"], "birth_date": "March 14, 1879"}
内容的提问来源于stack exchange,提问作者Siroj
相关产品推荐
相关产品推荐

