You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫无法抓取YellowPages中2条商家数据的技术求助

问题排查与解决方法

可能原因

  • 动态加载延迟:这两个商家属于页面滚动后异步加载的内容,Scrapy默认抓取初始静态HTML,未获取到后续渲染的元素。
  • 选择器细微偏差:看似结构一致,但实际容器的class、层级存在细微差异(比如多了sponsored类、空格差异),导致原有选择器匹配失败。
  • 反爬隐藏机制:网站对部分商家采用特殊渲染逻辑,比如通过JS动态插入内容、将数据存储在data-*属性而非直接文本节点。

解决步骤

1. 验证静态HTML是否包含目标内容

用Scrapy Shell请求目标页面,检查原始响应中是否存在这两个商家的信息:

scrapy shell "https://www.yellowpages.com/search?search_terms=auto&geo_location_terms=02136"

在Shell中执行:

print(response.text.find("Roger's Services"))
print(response.text.find("Northeastern Bus Rebuilders"))

如果返回-1,说明是动态加载问题,需要引入渲染工具。

2. 集成Playwright处理动态渲染

确保已安装scrapy-playwright,修改爬虫代码实现页面渲染:

from scrapy_playwright.page import PageCoroutine

class YellowPagesSpider(scrapy.Spider):
    name = 'yellowpages'
    start_urls = ['https://www.yellowpages.com/search?search_terms=auto&geo_location_terms=02136']

    def start_requests(self):
        for url in self.start_urls:
            yield scrapy.Request(
                url,
                meta={
                    'playwright': True,
                    'playwright_page_coroutines': [
                        PageCoroutine('wait_for_selector', 'div.result'),
                        PageCoroutine('scroll_to_bottom'),  # 滚动触发加载
                    ],
                },
            )

    def parse(self, response):
        for result in response.css('div.result'):
            business_name = result.css('a.business-name::text').get()
            if business_name in ["Roger's Services", "Northeastern Bus Rebuilders"]:
                self.logger.info(f"成功抓取: {business_name}")
            # 其他字段提取逻辑
            yield {
                'name': business_name,
                # 补充其他字段
            }

3. 修正选择器精度

如果静态HTML中存在目标内容但选择器匹配失败,直接通过商家文本定位父容器,对比结构差异:

# 在Scrapy Shell中执行
roger_container = response.xpath('//div[contains(text(), "Roger\'s Services")]/ancestor::div[contains(@class, "result")]')
print(roger_container.attrib.get('class'))
# 对比正常商家的容器class
normal_container = response.xpath('//div[contains(text(), "其他可抓取商家名称")]/ancestor::div[contains(@class, "result")]')
print(normal_container.attrib.get('class'))

如果发现class有差异(比如多了featured),调整选择器为div.result[class*="result"]以兼容所有变体。

4. 检查隐藏数据属性

部分商家的信息可能存储在data-*属性中,而非直接文本:

# 提取data-name属性中的商家名称
result.css('div[data-name]::attr(data-name)').get()

内容的提问来源于stack exchange,提问作者user22460492

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 09:11:16