如何将Scrapy嵌套请求的结果合并为单个Item?
问题描述
我需要爬取一个包含多所大学的页面,每所大学对应其奖学金列表链接,列表内还有每个奖学金的详情链接。期望生成的Item格式如下:
{ "name": "universityA", "scholarships": [ { "name": "sch_A", "other_fields": "other_values" }, { "name": "sch_B", "other_fields": "other_values" } ] }, { "name": "universityB", ... }
我尝试用Scrapy的meta传递Item,代码结构如下:
import scrapy class UniversityItem(scrapy.Item): uni = scrapy.Field() scholarships = scrapy.Field() class university_spider(scrapy.Spider): name = "test_scholarship_spider" start_urls = [ "https://search.studyaustralia.gov.au/scholarship/search-results.html?pageno=1", ] def parse(self, response): for div in response.css("div.sr_p.brd_btm"): university_item = UniversityItem() university_item["scholarships"] = [] uni_name = div.css("h2 a::text").get() university_item['uni'] = uni_name full_scholarship_detail_url = div.xpath('.//div[@class="rs_cnt"]/a/@href').get() if full_scholarship_detail_url: yield response.follow(url=full_scholarship_detail_url, callback=self.parse_all_scholarships, meta={ "uni_item": university_item }) else: pass yield university_item def parse_all_scholarships(self, response): for div in response.css('div.rs_cnt'): scholarship_detail = div.css('h3 a::attr(href)').get() new_scholarship = {} resp_meta = response.request.meta resp_meta["scholarship_obj"] = new_scholarship yield response.follow(url=scholarship_detail, callback=self.parse_scholarship_detail, meta=resp_meta) def parse_scholarship_detail(self, response): university_item = response.request.meta['uni_item'] scholarship_obj = response.request.meta['scholarship_obj'] scholarship_obj['eligibility_requirements'] = "multiple requirements that will be scrapped using selectors." scholarship_obj['application_process'] = "multiple processes that will be scrapped using selectors." university_item['scholarships'].append(scholarship_obj) yield university_item
但实际运行后会生成多份重复的大学Item,每份包含的奖学金数量递增,推测是因为在parse_scholarship_detail方法中每次解析完一个奖学金就yield Item导致的,请问该如何解决?
解决方案
核心问题在于过早或重复yield大学Item:parse方法里提前yield了未填充奖学金的Item,且parse_scholarship_detail每解析一个奖学金就yield一次,导致同一所大学被输出多次。解决思路是等该大学的所有奖学金详情都爬取完成后,再一次性yield完整的大学Item。
修改步骤:
- 移除
parse方法中的yield university_item:此时大学Item还未填充任何奖学金数据,提前输出没有意义。 - 追踪奖学金爬取进度:在
parse_all_scholarships中统计该大学的奖学金总数,然后在parse_scholarship_detail中记录已完成的数量,当全部完成时再输出完整Item。 - 使用
cb_kwargs替代meta(Scrapy官方推荐):传递数据更清晰,避免meta的潜在问题。
修改后的完整代码:
import scrapy class UniversityItem(scrapy.Item): uni = scrapy.Field() scholarships = scrapy.Field() class university_spider(scrapy.Spider): name = "test_scholarship_spider" start_urls = [ "https://search.studyaustralia.gov.au/scholarship/search-results.html?pageno=1", ] def parse(self, response): for div in response.css("div.sr_p.brd_btm"): university_item = UniversityItem() university_item["scholarships"] = [] uni_name = div.css("h2 a::text").get() university_item['uni'] = uni_name full_scholarship_detail_url = div.xpath('.//div[@class="rs_cnt"]/a/@href').get() if full_scholarship_detail_url: # 用cb_kwargs传递大学Item,替代meta yield response.follow( url=full_scholarship_detail_url, callback=self.parse_all_scholarships, cb_kwargs={"uni_item": university_item} ) else: # 没有奖学金链接的大学,直接输出空奖学金列表的Item yield university_item def parse_all_scholarships(self, response, uni_item): # 获取该大学的所有奖学金链接,统计总数 scholarship_divs = response.css('div.rs_cnt') total_scholarships = len(scholarship_divs) # 初始化已完成计数 uni_item['_completed'] = 0 for div in scholarship_divs: scholarship_detail = div.css('h3 a::attr(href)').get() new_scholarship = {} # 传递大学Item、奖学金对象、总数到详情解析函数 yield response.follow( url=scholarship_detail, callback=self.parse_scholarship_detail, cb_kwargs={ "uni_item": uni_item, "scholarship_obj": new_scholarship, "total": total_scholarships } ) def parse_scholarship_detail(self, response, uni_item, scholarship_obj, total): # 填充奖学金详情字段 scholarship_obj['name'] = response.css('h1::text').get() # 示例:获取奖学金名称 scholarship_obj['eligibility_requirements'] = "multiple requirements that will be scrapped using selectors." scholarship_obj['application_process'] = "multiple processes that will be scrapped using selectors." # 将当前奖学金添加到大学Item的列表中 uni_item['scholarships'].append(scholarship_obj) # 更新已完成计数 uni_item['_completed'] += 1 # 当所有奖学金都解析完成时,才yield完整的大学Item if uni_item['_completed'] == total: # 移除临时计数字段 del uni_item['_completed'] yield uni_item
关键说明:
cb_kwargs的使用:Scrapy 1.7+推荐用cb_kwargs传递数据,比meta更直观,避免序列化问题。- 进度追踪:通过
total_scholarships和_completed字段,确保只有当该大学的所有奖学金都解析完成后,才输出完整的Item,避免重复输出。 - 无奖学金的情况:如果大学没有奖学金链接,直接输出空奖学金列表的Item,保证数据完整性。
内容的提问来源于stack exchange,提问作者aashish manandhar
相关产品推荐
相关产品推荐

