Scrapy解析器嵌套实现:画廊及馆长声明爬取问题
问题分析
这个错误的核心原因是你在处理馆长声明的逻辑中,试图对字符串类型的数据调用xpath()方法——只有Scrapy的Response对象才具备xpath()解析能力,显然你在数据传递或解析函数的处理上出现了对象类型混淆。
正确嵌套解析器实现方案
核心逻辑是拆分三个职责明确的解析函数,通过meta参数传递已爬取的画廊数据,实现链式的页面跟进与数据整合:
import scrapy class GallerySpider(scrapy.Spider): name = 'gallery_spider' start_urls = ['你的画廊列表页面URL'] def parse(self, response): # 提取前6个画廊链接并跟进 gallery_links = response.xpath('//匹配画廊链接的xpath表达式').getall()[:6] for link in gallery_links: yield response.follow(link, callback=self.parse_gallery) def parse_gallery(self, response): # 提取画廊页面的基础数据 gallery_data = { 'name': response.xpath('//匹配画廊名称的xpath').get(), 'location': response.xpath('//匹配画廊位置的xpath').get(), # 按需添加其他画廊字段 } # 定位馆长声明的链接 statement_link = response.xpath('//a[contains(text(), "Read the Curators\' Statement")]/@href').get() if statement_link: # 跟进声明页面,通过meta传递已爬取的画廊数据 yield response.follow( statement_link, callback=self.parse_curator_statement, meta={'gallery_data': gallery_data} ) else: # 无声明链接时直接返回画廊基础数据 yield gallery_data def parse_curator_statement(self, response): # 从meta中取出画廊基础数据 gallery_data = response.meta['gallery_data'] # 提取并格式化声明文本 statement_content = ' '.join(response.xpath('//匹配声明文本的xpath').getall()).strip() # 整合数据 gallery_data['curator_statement'] = statement_content # 输出最终完整数据 yield gallery_data
关键避坑点
- 必须通过
response.follow或scrapy.Request的meta参数传递已爬取数据,不能直接将非Response对象传入解析函数。 - 在
parse_curator_statement中,仅能对response对象调用xpath(),切勿误将gallery_data(字典类型)当成Response处理。 - 如果之前的错误是因直接传递字符串链接到解析函数导致,务必改为让Scrapy发起请求后再解析目标页面。
内容的提问来源于stack exchange,提问作者vonbecker
相关产品推荐
相关产品推荐

