Scrapy循环使用相对XPath爬取数据无返回结果问题排查
问题解决方案
核心错误原因
你在循环内调用XPath的语法错误:遍历得到的box本身就是Scrapy的Selector对象,直接调用xpath()方法即可,不需要额外添加.response属性,错误的调用链导致节点匹配失败,无法提取到数据。
另外你将NewContactItem实例的创建放在循环外部,会出现后续循环覆盖前一轮字段值的问题,也需要调整。
修正后可运行代码
import scrapy from NewContacts_Crawler.items import NewContactItem class XubioContadoresSpider(scrapy.Spider): name = "XubioContadores" # 用生成式简化分页URL生成,避免重复代码 start_urls = [ f'https://socios.xubio.com/ar/contadores/?pag={i}' for i in range(1, 10) ] def parse(self, response): for box in response.xpath('//div[@class="w-100 padding-15 "]'): # item实例移到循环内部,避免字段覆盖 items = NewContactItem() # 去掉冗余的.response,直接调用box的xpath方法 name = box.xpath('.//div[@class="col-11"]//p//text()').extract_first() phone = box.xpath('.//div[@class="col-md-2 col-12 h-60"]//p//text()').extract_first() address = box.xpath('.//div[@class="col-md-4 col-12 h-60"]//p//text()').extract_first() email = box.xpath('.//div[@class="col-md-3 col-12 h-60"]//p//text()').extract_first() # 可选:去除提取结果的多余空格和换行符 items['name'] = name.strip() if name else None items['phone'] = phone.strip() if phone else None items['address'] = address.strip() if address else None items['email'] = email.strip() if email else None yield items
内容的提问来源于stack exchange,提问作者maximodesousadias
相关产品推荐
相关产品推荐

