Scrapy爬虫输出缺失数据:目标页面10条仅爬取9条
问题原因分析
你遇到的问题核心是Scrapy默认的请求去重机制:同一个作者的多条名言会跳转至同一个作者详情页URL,Scrapy会自动过滤重复的URL请求,导致后面对该URL的请求被丢弃,对应的那条名言数据无法进入parse_page函数处理。
解决方案
1. 关闭重复请求过滤
修改response.follow的调用,添加dont_filter=True参数,让Scrapy不对重复的作者详情页URL做过滤:
yield response.follow(full_page_url, callback=self.parse_page, cb_kwargs={'item': n}, dont_filter=True)
(用cb_kwargs比meta更推荐,你之前的尝试方向是对的,只是缺了去重关闭的参数)
2. 修正字典值的元组问题
你代码里每个字典值后面都加了逗号,会把值变成元组(比如n['quote']会是('名言内容',)这种格式),虽然不影响请求传递,但会导致最终输出的数据格式不符合预期,建议去掉逗号:
# 修改前 n['quote'] = q.css('.text ::text').get(), n['tag'] = t, n['author'] = q.css('span .author ::text').get(), # 修改后 n['quote'] = q.css('.text ::text').get() n['tag'] = t n['author'] = q.css('span .author ::text').get()
完整修正后的代码示例
def parse(self, response): qs = response.css('.quote') for q in qs: n = {} page_url = q.css('span a').attrib['href'] full_page_url = 'https://quotes.toscrape.com' + page_url # tags t = [] tags = q.css('.tag') for tag in tags: t.append(tag.css('::text').get()) # items - 去掉逗号,避免元组 n['quote'] = q.css('.text ::text').get() n['tag'] = t n['author'] = q.css('span .author ::text').get() # 添加dont_filter=True关闭重复请求过滤 yield response.follow(full_page_url, callback=self.parse_page, cb_kwargs={'item': n}, dont_filter=True) def parse_page(self, response): q = response.css('.author-details') item = response.cb_kwargs.get('item') # 对应cb_kwargs的获取方式 yield { 'text': item['quote'], 'author': item['author'], 'tags': item['tag'], 'date': q.css('p .author-born-date ::text').get(), 'location': q.css('p .author-born-location ::text').get(), }
验证效果
修改后重新运行爬虫,就能获取到全部10条名言数据了,因为即使是同一个作者的详情页URL,Scrapy也会处理每一次请求,对应的名言数据都会被传递到parse_page生成结果。
内容的提问来源于stack exchange,提问作者Low LiFe
相关产品推荐
相关产品推荐

