Scrapy爬虫技术问题:如何爬取作者页面的出生日期?
如何在Scrapy中爬取名言同时获取作者出生日期?
问题背景
学习Scrapy框架时,以http://quotes.toscrape.com为目标站点,通过scrapy genspider quotes命令创建了爬虫,希望在爬取名言内容的同时,进入每条名言对应的作者页面解析出生日期,但尝试的代码无法正常运行。
尝试的错误代码
import scrapy class QuotesSpider(scrapy.Spider): name = "quotes" allowed_domains = ["quotes.toscrape.com"] start_urls = ["http://quotes.toscrape.com/"] def parse(self, response): quotes=response.xpath('//div[@class="quote"]') item={} for quote in quotes: item['name']=quote.xpath('.//span[@class="text"]/text()').get() item['author']=quote.xpath('.//small[@class="author"]/text()').get() item['tags']=quote.xpath('.//div[@class="tags"]/a[@class="tag"]/text()').getall() url=quote.xpath('.//small[@class="author"]/../a/@href').get() response.follow(url, self.parse_additional_page, item) new_page=response.xpath('//li[@class="next"]/a/@href').get() if new_page is not None: yield response.follow(new_page,self.parse) def parse_additional_page(self, response, item): item['additional_data'] = response.xpath('//span[@class="author-born-date"]/text()').get() yield item
可正常爬取名言(无出生日期)的代码
import scrapy class QuotesSpiderSpider(scrapy.Spider): name = "quotes_spider" allowed_domains = ["quotes.toscrape.com"] start_urls = ["https://quotes.toscrape.com/"] def parse(self, response): quotes=response.xpath('//div[@class="quote"]') for quote in quotes: yield { 'name':quote.xpath('.//span[@class="text"]/text()').get(), 'author':quote.xpath('.//small[@class="author"]/text()').get(), 'tags':quote.xpath('.//div[@class="tags"]/a[@class="tag"]/text()').getall() } new_page=response.xpath('//li[@class="next"]/a/@href').get() if new_page is not None: yield response.follow(new_page,self.parse)
问题分析
错误代码存在两个核心问题:
- 字典复用导致数据覆盖:循环外定义
item={},每次迭代修改的是同一个字典对象,最终所有请求传递的item都会是最后一次迭代的内容。 response.follow参数传递错误:response.follow的第三个参数不是直接传递item,而是需要通过meta字典来传递额外数据,且回调函数不能直接接收item作为参数。
解决方案
修正后的代码如下:
import scrapy class QuotesSpider(scrapy.Spider): name = "quotes" allowed_domains = ["quotes.toscrape.com"] start_urls = ["http://quotes.toscrape.com/"] def parse(self, response): quotes = response.xpath('//div[@class="quote"]') for quote in quotes: # 循环内创建item字典,避免数据覆盖 item = {} item['name'] = quote.xpath('.//span[@class="text"]/text()').get() item['author'] = quote.xpath('.//small[@class="author"]/text()').get() item['tags'] = quote.xpath('.//div[@class="tags"]/a[@class="tag"]/text()').getall() author_url = quote.xpath('.//small[@class="author"]/../a/@href').get() # 通过meta参数传递item到回调函数 yield response.follow(author_url, self.parse_author, meta={'item': item}) new_page = response.xpath('//li[@class="next"]/a/@href').get() if new_page is not None: yield response.follow(new_page, self.parse) def parse_author(self, response): # 从meta中取出传递的item item = response.meta['item'] # 解析出生日期并添加到item item['birth_date'] = response.xpath('//span[@class="author-born-date"]/text()').get() yield item
关键改动说明
- 循环内创建item:每次迭代新建
item字典,确保每条名言的独立数据不会被后续迭代覆盖。 - 使用
meta传递数据:response.follow通过meta={'item': item}将当前item传递给回调函数,在回调中通过response.meta['item']取出。 - 回调函数参数调整:回调函数
parse_author只接收response参数,通过response.meta获取传递的item。
内容的提问来源于stack exchange,提问作者Alex Freeman
相关产品推荐
相关产品推荐

