Scrapy爬虫无法抓取YellowPages中2条商家数据的技术求助
问题排查与解决方法
可能原因
- 动态加载延迟:这两个商家属于页面滚动后异步加载的内容,Scrapy默认抓取初始静态HTML,未获取到后续渲染的元素。
- 选择器细微偏差:看似结构一致,但实际容器的class、层级存在细微差异(比如多了
sponsored类、空格差异),导致原有选择器匹配失败。 - 反爬隐藏机制:网站对部分商家采用特殊渲染逻辑,比如通过JS动态插入内容、将数据存储在
data-*属性而非直接文本节点。
解决步骤
1. 验证静态HTML是否包含目标内容
用Scrapy Shell请求目标页面,检查原始响应中是否存在这两个商家的信息:
scrapy shell "https://www.yellowpages.com/search?search_terms=auto&geo_location_terms=02136"
在Shell中执行:
print(response.text.find("Roger's Services")) print(response.text.find("Northeastern Bus Rebuilders"))
如果返回-1,说明是动态加载问题,需要引入渲染工具。
2. 集成Playwright处理动态渲染
确保已安装scrapy-playwright,修改爬虫代码实现页面渲染:
from scrapy_playwright.page import PageCoroutine class YellowPagesSpider(scrapy.Spider): name = 'yellowpages' start_urls = ['https://www.yellowpages.com/search?search_terms=auto&geo_location_terms=02136'] def start_requests(self): for url in self.start_urls: yield scrapy.Request( url, meta={ 'playwright': True, 'playwright_page_coroutines': [ PageCoroutine('wait_for_selector', 'div.result'), PageCoroutine('scroll_to_bottom'), # 滚动触发加载 ], }, ) def parse(self, response): for result in response.css('div.result'): business_name = result.css('a.business-name::text').get() if business_name in ["Roger's Services", "Northeastern Bus Rebuilders"]: self.logger.info(f"成功抓取: {business_name}") # 其他字段提取逻辑 yield { 'name': business_name, # 补充其他字段 }
3. 修正选择器精度
如果静态HTML中存在目标内容但选择器匹配失败,直接通过商家文本定位父容器,对比结构差异:
# 在Scrapy Shell中执行 roger_container = response.xpath('//div[contains(text(), "Roger\'s Services")]/ancestor::div[contains(@class, "result")]') print(roger_container.attrib.get('class')) # 对比正常商家的容器class normal_container = response.xpath('//div[contains(text(), "其他可抓取商家名称")]/ancestor::div[contains(@class, "result")]') print(normal_container.attrib.get('class'))
如果发现class有差异(比如多了featured),调整选择器为div.result[class*="result"]以兼容所有变体。
4. 检查隐藏数据属性
部分商家的信息可能存储在data-*属性中,而非直接文本:
# 提取data-name属性中的商家名称 result.css('div[data-name]::attr(data-name)').get()
内容的提问来源于stack exchange,提问作者user22460492
相关产品推荐
相关产品推荐

