You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy Shell可正常提取数据,运行爬虫代码却返回空字典

问题原因及解决办法

核心问题:选择器作用域错误

你在循环中对section调用css()方法时,仍然使用了以#happy-hour开头的绝对选择器。section已经是#happy-hour .pt5匹配到的子节点,此时再用#happy-hour开头的选择器,会重新从整个HTML文档根节点开始查找,而不是从当前section节点下查找,自然找不到对应内容。

修复步骤

  • 去掉选择器中的#happy-hour前缀,改用相对路径,基于当前section节点进行查找
  • 优化选择器的精准度,避免重复选择相同内容(比如你现在食物和饮料的item选择器完全一致,会导致数据重复)

修复后的代码示例

import scrapy 

class Earls(scrapy.Spider): 
    name = 'earls_spider'   
    start_urls = [
        'https://earls.ca/locations/dalhousie/menu/'
    ]

    def parse(self, response):
        all_earls_happy_hour = response.css('#happy-hour .pt5')

        for section in all_earls_happy_hour:
            # 食物相关:基于section节点的相对选择器
            earls_food_item = section.css(".sub-section+ .sub-section .tl::text").extract()
            earls_food_price = section.css(".sub-section+ .sub-section .actual::text").extract()
            # 饮料相关:调整选择器区分食物和饮料
            earls_drink_items = section.css(".items-center+ .sub-section .tl::text").extract()
            earls_drink_price = section.css(".items-center+ .sub-section .actual::text").extract()
            yield {
                'earls_HH_food': earls_food_item,
                'earls_food_price': earls_food_price,
                'earls_drink_items': earls_drink_items,
                'earls_drink_price': earls_drink_price
            }

额外建议

  • 在Scrapy Shell中测试相对选择器时,可以先定位到section节点,再在该节点下测试子选择器,比如:
    section = response.css('#happy-hour .pt5')[0]
    section.css(".sub-section+ .sub-section .tl::text").extract()
    
    这样能确保选择器在爬虫代码中的行为和Shell中一致。
  • 如果页面是动态加载的,也可以检查下请求是否需要携带特定Headers(比如User-Agent),不过从你描述Shell能抓取到来看,动态加载的可能性较低。

内容的提问来源于stack exchange,提问作者Grant Rauser

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 15:35:39