You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy循环使用相对XPath爬取数据无返回结果问题排查

问题解决方案

核心错误原因

你在循环内调用XPath的语法错误:遍历得到的box本身就是Scrapy的Selector对象,直接调用xpath()方法即可,不需要额外添加.response属性,错误的调用链导致节点匹配失败,无法提取到数据。
另外你将NewContactItem实例的创建放在循环外部,会出现后续循环覆盖前一轮字段值的问题,也需要调整。

修正后可运行代码

import scrapy
from NewContacts_Crawler.items import NewContactItem

class XubioContadoresSpider(scrapy.Spider):
    name = "XubioContadores"
    # 用生成式简化分页URL生成,避免重复代码
    start_urls = [
        f'https://socios.xubio.com/ar/contadores/?pag={i}' for i in range(1, 10)
    ]

    def parse(self, response):
        for box in response.xpath('//div[@class="w-100 padding-15 "]'):
            # item实例移到循环内部,避免字段覆盖
            items = NewContactItem()
            # 去掉冗余的.response,直接调用box的xpath方法
            name = box.xpath('.//div[@class="col-11"]//p//text()').extract_first()
            phone = box.xpath('.//div[@class="col-md-2 col-12 h-60"]//p//text()').extract_first()
            address = box.xpath('.//div[@class="col-md-4 col-12 h-60"]//p//text()').extract_first()
            email = box.xpath('.//div[@class="col-md-3 col-12 h-60"]//p//text()').extract_first()
            
            # 可选:去除提取结果的多余空格和换行符
            items['name'] = name.strip() if name else None
            items['phone'] = phone.strip() if phone else None
            items['address'] = address.strip() if address else None
            items['email'] = email.strip() if email else None

            yield items

内容的提问来源于stack exchange,提问作者maximodesousadias

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 00:48:04