You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Scrapy嵌套请求的结果合并为单个Item?

问题描述

我需要爬取一个包含多所大学的页面,每所大学对应其奖学金列表链接,列表内还有每个奖学金的详情链接。期望生成的Item格式如下:

{
  "name": "universityA",
  "scholarships": [
      {
          "name": "sch_A",
          "other_fields": "other_values"
      },
      {
          "name": "sch_B",
          "other_fields": "other_values"
      }
  ]
},
{
  "name": "universityB",
  ...
}

我尝试用Scrapy的meta传递Item,代码结构如下:

import scrapy

class UniversityItem(scrapy.Item):
    uni = scrapy.Field()
    scholarships = scrapy.Field()

class university_spider(scrapy.Spider):
    name = "test_scholarship_spider"
    start_urls = [
        "https://search.studyaustralia.gov.au/scholarship/search-results.html?pageno=1",
    ]

    def parse(self, response):
        for div in response.css("div.sr_p.brd_btm"):
            university_item = UniversityItem()
            university_item["scholarships"] = []
            uni_name = div.css("h2 a::text").get()
            university_item['uni'] = uni_name
            full_scholarship_detail_url = div.xpath('.//div[@class="rs_cnt"]/a/@href').get()
            if full_scholarship_detail_url:
                yield response.follow(url=full_scholarship_detail_url, callback=self.parse_all_scholarships, meta={ "uni_item": university_item })
            else:
                pass
            
            yield university_item
            
    
    def parse_all_scholarships(self, response):
        for div in response.css('div.rs_cnt'):
            scholarship_detail = div.css('h3 a::attr(href)').get()
            new_scholarship = {}
            resp_meta = response.request.meta
            resp_meta["scholarship_obj"] = new_scholarship

            yield response.follow(url=scholarship_detail, callback=self.parse_scholarship_detail, meta=resp_meta)

    
    def parse_scholarship_detail(self, response):
        university_item = response.request.meta['uni_item']
        scholarship_obj = response.request.meta['scholarship_obj']
        scholarship_obj['eligibility_requirements'] = "multiple requirements that will be scrapped using selectors."
        scholarship_obj['application_process'] = "multiple processes that will be scrapped using selectors."
        university_item['scholarships'].append(scholarship_obj)
        yield university_item

但实际运行后会生成多份重复的大学Item,每份包含的奖学金数量递增,推测是因为在parse_scholarship_detail方法中每次解析完一个奖学金就yield Item导致的,请问该如何解决?

解决方案

核心问题在于过早或重复yield大学Item:parse方法里提前yield了未填充奖学金的Item,且parse_scholarship_detail每解析一个奖学金就yield一次,导致同一所大学被输出多次。解决思路是等该大学的所有奖学金详情都爬取完成后,再一次性yield完整的大学Item。

修改步骤:

  1. 移除parse方法中的yield university_item:此时大学Item还未填充任何奖学金数据,提前输出没有意义。
  2. 追踪奖学金爬取进度:在parse_all_scholarships中统计该大学的奖学金总数,然后在parse_scholarship_detail中记录已完成的数量,当全部完成时再输出完整Item。
  3. 使用cb_kwargs替代meta(Scrapy官方推荐):传递数据更清晰,避免meta的潜在问题。

修改后的完整代码:

import scrapy

class UniversityItem(scrapy.Item):
    uni = scrapy.Field()
    scholarships = scrapy.Field()

class university_spider(scrapy.Spider):
    name = "test_scholarship_spider"
    start_urls = [
        "https://search.studyaustralia.gov.au/scholarship/search-results.html?pageno=1",
    ]

    def parse(self, response):
        for div in response.css("div.sr_p.brd_btm"):
            university_item = UniversityItem()
            university_item["scholarships"] = []
            uni_name = div.css("h2 a::text").get()
            university_item['uni'] = uni_name
            full_scholarship_detail_url = div.xpath('.//div[@class="rs_cnt"]/a/@href').get()
            if full_scholarship_detail_url:
                # 用cb_kwargs传递大学Item,替代meta
                yield response.follow(
                    url=full_scholarship_detail_url,
                    callback=self.parse_all_scholarships,
                    cb_kwargs={"uni_item": university_item}
                )
            else:
                # 没有奖学金链接的大学,直接输出空奖学金列表的Item
                yield university_item
            
    
    def parse_all_scholarships(self, response, uni_item):
        # 获取该大学的所有奖学金链接,统计总数
        scholarship_divs = response.css('div.rs_cnt')
        total_scholarships = len(scholarship_divs)
        # 初始化已完成计数
        uni_item['_completed'] = 0

        for div in scholarship_divs:
            scholarship_detail = div.css('h3 a::attr(href)').get()
            new_scholarship = {}
            # 传递大学Item、奖学金对象、总数到详情解析函数
            yield response.follow(
                url=scholarship_detail,
                callback=self.parse_scholarship_detail,
                cb_kwargs={
                    "uni_item": uni_item,
                    "scholarship_obj": new_scholarship,
                    "total": total_scholarships
                }
            )

    
    def parse_scholarship_detail(self, response, uni_item, scholarship_obj, total):
        # 填充奖学金详情字段
        scholarship_obj['name'] = response.css('h1::text').get()  # 示例:获取奖学金名称
        scholarship_obj['eligibility_requirements'] = "multiple requirements that will be scrapped using selectors."
        scholarship_obj['application_process'] = "multiple processes that will be scrapped using selectors."
        
        # 将当前奖学金添加到大学Item的列表中
        uni_item['scholarships'].append(scholarship_obj)
        # 更新已完成计数
        uni_item['_completed'] += 1

        # 当所有奖学金都解析完成时,才yield完整的大学Item
        if uni_item['_completed'] == total:
            # 移除临时计数字段
            del uni_item['_completed']
            yield uni_item

关键说明:

  • cb_kwargs的使用:Scrapy 1.7+推荐用cb_kwargs传递数据,比meta更直观,避免序列化问题。
  • 进度追踪:通过total_scholarships和_completed字段,确保只有当该大学的所有奖学金都解析完成后,才输出完整的Item,避免重复输出。
  • 无奖学金的情况:如果大学没有奖学金链接,直接输出空奖学金列表的Item,保证数据完整性。

内容的提问来源于stack exchange,提问作者aashish manandhar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 09:37:02