You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy递归爬取时如何关联父页面节点数据与对应子页面信息

问题原因

你的问题根源是使用了类全局变量存储单条名言内容,Scrapy的请求是异步执行的,parse函数里的循环会瞬间执行完成,等所有作者页的请求返回触发parse_author回调时,AuthorSpider.var已经被覆盖为循环最后一次迭代的名言内容,所以所有条目quote字段都相同。

修复方案

用Scrapy请求的cb_kwargs参数(推荐,Scrapy 1.7+支持)将当前循环对应的名言内容作为独立参数传递给回调函数,不要用全局类变量存储。

修改后的完整代码
import scrapy 

class AuthorSpider(scrapy.Spider):
    name = 'quotes1'

    def start_requests(self):
        start_urls = ['http://quotes.toscrape.com/']
        yield scrapy.Request(url=start_urls[0], callback=self.parse)

    def parse(self, response):

        for quote in response.css('div.quote'):
            # 提取当前循环对应的名言内容
            current_quote = quote.css('div span.text::text').get()

            authShortLink = quote.css('small.author + a::attr(href)').get()
            authFullLink = response.urljoin(authShortLink)
            # 通过cb_kwargs传递名言内容到回调函数
            yield scrapy.Request(
                url=authFullLink, 
                callback=self.parse_author,
                cb_kwargs={'quote': current_quote}
            )

        # # 遍历下一页
        # nextPage = response.css('li.next a::attr(href)').get()
        # if nextPage is not None:
        #     nextPage = response.urljoin(nextPage)
        #     yield scrapy.Request(url=nextPage, callback=self.parse)

    # 回调函数新增参数接收传递的quote
    def parse_author(self, response, quote):
        def extract_with_css(query):
            return response.css(query).get(default='').strip()

        yield {
            'name': extract_with_css('h3.author-title::text'),
            'birthdate': extract_with_css('.author-born-date::text'),
            'quote' : quote
        }

如果使用的是低于1.7版本的Scrapy,可以替换为meta参数传递:

  • 发起请求时添加meta={'quote': current_quote}
  • 在parse_author中通过response.meta['quote']获取对应名言内容

内容的提问来源于stack exchange,提问作者beta green

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 00:06:03