Scrapy递归爬取时如何关联父页面节点数据与对应子页面信息
问题原因
你的问题根源是使用了类全局变量存储单条名言内容,Scrapy的请求是异步执行的,parse函数里的循环会瞬间执行完成,等所有作者页的请求返回触发parse_author回调时,AuthorSpider.var已经被覆盖为循环最后一次迭代的名言内容,所以所有条目quote字段都相同。
修复方案
用Scrapy请求的cb_kwargs参数(推荐,Scrapy 1.7+支持)将当前循环对应的名言内容作为独立参数传递给回调函数,不要用全局类变量存储。
修改后的完整代码
import scrapy class AuthorSpider(scrapy.Spider): name = 'quotes1' def start_requests(self): start_urls = ['http://quotes.toscrape.com/'] yield scrapy.Request(url=start_urls[0], callback=self.parse) def parse(self, response): for quote in response.css('div.quote'): # 提取当前循环对应的名言内容 current_quote = quote.css('div span.text::text').get() authShortLink = quote.css('small.author + a::attr(href)').get() authFullLink = response.urljoin(authShortLink) # 通过cb_kwargs传递名言内容到回调函数 yield scrapy.Request( url=authFullLink, callback=self.parse_author, cb_kwargs={'quote': current_quote} ) # # 遍历下一页 # nextPage = response.css('li.next a::attr(href)').get() # if nextPage is not None: # nextPage = response.urljoin(nextPage) # yield scrapy.Request(url=nextPage, callback=self.parse) # 回调函数新增参数接收传递的quote def parse_author(self, response, quote): def extract_with_css(query): return response.css(query).get(default='').strip() yield { 'name': extract_with_css('h3.author-title::text'), 'birthdate': extract_with_css('.author-born-date::text'), 'quote' : quote }
如果使用的是低于1.7版本的Scrapy,可以替换为meta参数传递:
- 发起请求时添加
meta={'quote': current_quote} - 在
parse_author中通过response.meta['quote']获取对应名言内容
内容的提问来源于stack exchange,提问作者beta green
相关产品推荐
相关产品推荐

