You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy Parse函数无法生成输出,无法打印变量,求问题排查

你的Scrapy脚本问题分析及修正

核心问题点:

  • 变量作用域错误:print(name, phone, email, company)写在类外部,这些变量仅存在于parse函数的yield字典中,全局范围内根本不存在,自然无法打印。Scrapy的输出逻辑是通过yield生成Item,而非直接调用全局变量。
  • 未处理多条目数据:当前代码仅尝试提取单个代理信息,但目标页面包含多个代理条目,需先定位每个代理的容器节点,再循环提取每个节点内的内容。
  • 冗余依赖:导入了requests_html和HTMLSession,但Scrapy自带请求与解析能力,这些模块完全未被使用,属于多余代码。
  • 脆弱的绝对路径XPath://html/body/div[2]/div/section/div/main/div/div[3]/div[1]/div[2]/div/div/div[2]/h3/text()这类绝对路径,一旦页面结构微调就会失效,应改用基于类的相对路径。
  • 空值处理缺失:get()方法可能返回None,直接调用.strip()会触发AttributeError,需先判断值是否存在再处理。

修正后的代码:

import scrapy

class BillingsorgSpider(scrapy.Spider):
    name = "billingsorg"
    allowed_domains = ["billings.org"]
    start_urls = ['https://www.billings.org/agents/']

    def parse(self, response):
        # 遍历每个代理的容器节点
        for agent in response.xpath('//div[contains(@class, "staff-item")]'):
            # 处理空值,避免报错
            name = agent.xpath('.//h3/text()').get()
            name = name.strip() if name else None
            
            phone = agent.xpath('.//div[@class="staff-phone"]/i/following-sibling::text()').get()
            phone = phone.strip() if phone else None
            
            email = agent.css('a[href^="mailto"]::attr(href)').get()
            email = email.replace('mailto:', '') if email else None
            
            company = agent.xpath('.//div[@class="staff-company"]/i/following-sibling::text()').get()
            company = company.strip() if company else None
            
            yield {
                'name': name,
                'phone': phone,
                'email': email,
                'company': company
            }

查看结果的正确方式:

执行Scrapy命令 scrapy crawl billingsorg -o agents.json,数据会输出到agents.json文件中;或者直接查看终端的Scrapy日志输出(默认会打印生成的Item内容)。

内容的提问来源于stack exchange,提问作者Larry Hines

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 09:25:05