Scrapy Parse函数无法生成输出,无法打印变量,求问题排查
你的Scrapy脚本问题分析及修正
核心问题点:
- 变量作用域错误:
print(name, phone, email, company)写在类外部,这些变量仅存在于parse函数的yield字典中,全局范围内根本不存在,自然无法打印。Scrapy的输出逻辑是通过yield生成Item,而非直接调用全局变量。 - 未处理多条目数据:当前代码仅尝试提取单个代理信息,但目标页面包含多个代理条目,需先定位每个代理的容器节点,再循环提取每个节点内的内容。
- 冗余依赖:导入了
requests_html和HTMLSession,但Scrapy自带请求与解析能力,这些模块完全未被使用,属于多余代码。 - 脆弱的绝对路径XPath:
//html/body/div[2]/div/section/div/main/div/div[3]/div[1]/div[2]/div/div/div[2]/h3/text()这类绝对路径,一旦页面结构微调就会失效,应改用基于类的相对路径。 - 空值处理缺失:
get()方法可能返回None,直接调用.strip()会触发AttributeError,需先判断值是否存在再处理。
修正后的代码:
import scrapy class BillingsorgSpider(scrapy.Spider): name = "billingsorg" allowed_domains = ["billings.org"] start_urls = ['https://www.billings.org/agents/'] def parse(self, response): # 遍历每个代理的容器节点 for agent in response.xpath('//div[contains(@class, "staff-item")]'): # 处理空值,避免报错 name = agent.xpath('.//h3/text()').get() name = name.strip() if name else None phone = agent.xpath('.//div[@class="staff-phone"]/i/following-sibling::text()').get() phone = phone.strip() if phone else None email = agent.css('a[href^="mailto"]::attr(href)').get() email = email.replace('mailto:', '') if email else None company = agent.xpath('.//div[@class="staff-company"]/i/following-sibling::text()').get() company = company.strip() if company else None yield { 'name': name, 'phone': phone, 'email': email, 'company': company }
查看结果的正确方式:
执行Scrapy命令 scrapy crawl billingsorg -o agents.json,数据会输出到agents.json文件中;或者直接查看终端的Scrapy日志输出(默认会打印生成的Item内容)。
内容的提问来源于stack exchange,提问作者Larry Hines
相关产品推荐
相关产品推荐

