You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫技术问题:如何爬取作者页面的出生日期?

如何在Scrapy中爬取名言同时获取作者出生日期?

问题背景

学习Scrapy框架时,以http://quotes.toscrape.com为目标站点,通过scrapy genspider quotes命令创建了爬虫,希望在爬取名言内容的同时,进入每条名言对应的作者页面解析出生日期,但尝试的代码无法正常运行。

尝试的错误代码

import scrapy


class QuotesSpider(scrapy.Spider):
    name = "quotes"
    allowed_domains = ["quotes.toscrape.com"]
    start_urls = ["http://quotes.toscrape.com/"]

    def parse(self, response):
        
        quotes=response.xpath('//div[@class="quote"]') 
        
        item={}

        for quote in quotes: 
            item['name']=quote.xpath('.//span[@class="text"]/text()').get()
            item['author']=quote.xpath('.//small[@class="author"]/text()').get()
            item['tags']=quote.xpath('.//div[@class="tags"]/a[@class="tag"]/text()').getall()
            url=quote.xpath('.//small[@class="author"]/../a/@href').get()
            response.follow(url, self.parse_additional_page, item) 
            

        new_page=response.xpath('//li[@class="next"]/a/@href').get() 

        if new_page is not None: 

            yield response.follow(new_page,self.parse) 
            
    def parse_additional_page(self, response, item): 
        item['additional_data'] = response.xpath('//span[@class="author-born-date"]/text()').get() 
        yield item

可正常爬取名言(无出生日期)的代码

import scrapy 



class QuotesSpiderSpider(scrapy.Spider): 
    name = "quotes_spider" 
    allowed_domains = ["quotes.toscrape.com"] 
    start_urls = ["https://quotes.toscrape.com/"] 
     
    def parse(self, response): 
        quotes=response.xpath('//div[@class="quote"]') 
        for quote in quotes: 
            yield { 
                'name':quote.xpath('.//span[@class="text"]/text()').get(), 
                'author':quote.xpath('.//small[@class="author"]/text()').get(), 
                'tags':quote.xpath('.//div[@class="tags"]/a[@class="tag"]/text()').getall() 
                } 
        new_page=response.xpath('//li[@class="next"]/a/@href').get() 
        if new_page is not None: 
            yield response.follow(new_page,self.parse) 

问题分析

错误代码存在两个核心问题:

  1. 字典复用导致数据覆盖:循环外定义item={},每次迭代修改的是同一个字典对象,最终所有请求传递的item都会是最后一次迭代的内容。
  2. response.follow参数传递错误:response.follow的第三个参数不是直接传递item,而是需要通过meta字典来传递额外数据,且回调函数不能直接接收item作为参数。

解决方案

修正后的代码如下:

import scrapy


class QuotesSpider(scrapy.Spider):
    name = "quotes"
    allowed_domains = ["quotes.toscrape.com"]
    start_urls = ["http://quotes.toscrape.com/"]

    def parse(self, response):
        quotes = response.xpath('//div[@class="quote"]') 
        
        for quote in quotes: 
            # 循环内创建item字典,避免数据覆盖
            item = {}
            item['name'] = quote.xpath('.//span[@class="text"]/text()').get()
            item['author'] = quote.xpath('.//small[@class="author"]/text()').get()
            item['tags'] = quote.xpath('.//div[@class="tags"]/a[@class="tag"]/text()').getall()
            author_url = quote.xpath('.//small[@class="author"]/../a/@href').get()
            
            # 通过meta参数传递item到回调函数
            yield response.follow(author_url, self.parse_author, meta={'item': item}) 

        new_page = response.xpath('//li[@class="next"]/a/@href').get() 
        if new_page is not None: 
            yield response.follow(new_page, self.parse) 
            
    def parse_author(self, response): 
        # 从meta中取出传递的item
        item = response.meta['item']
        # 解析出生日期并添加到item
        item['birth_date'] = response.xpath('//span[@class="author-born-date"]/text()').get() 
        yield item

关键改动说明

  • 循环内创建item:每次迭代新建item字典,确保每条名言的独立数据不会被后续迭代覆盖。
  • 使用meta传递数据:response.follow通过meta={'item': item}将当前item传递给回调函数,在回调中通过response.meta['item']取出。
  • 回调函数参数调整:回调函数parse_author只接收response参数,通过response.meta获取传递的item。

内容的提问来源于stack exchange,提问作者Alex Freeman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 14:17:49