如何抓取XPath/CSS选择器无返回值的翻页按钮对应页面数据
问题根因
- 类名动态失效:该网站使用了CSS模块化技术,类名后缀的随机哈希字符串会随前端版本更新变化,你硬编码的
action-button--1O8tU这类类名无法稳定匹配到元素,这也是你在scrapy shell和splash里拿不到选择器的核心原因 - 方法调用逻辑错误:scrapy的XPath返回的是Selector对象,没有
click()方法,该方法是Selenium WebDriver对象的专属方法,原代码直接对Selector调用click()完全不会生效 - 翻页流程错乱:你模拟点击后直接把点击操作的返回值当URL传入SeleniumRequest,参数类型本身就不匹配,且翻页后的回调错误指定为
parse,会重复执行首页分类抓取逻辑,无法继续爬取列表数据
最优解决方案:直接构造翻页URL
该站点的翻页规则非常清晰,不需要模拟点击next按钮,完全可以规避选择器匹配问题:
所有分类列表页的翻页参数为page,第一页URL为https://tonaton.com/en/ads/ghana/electronics,第二页为https://tonaton.com/en/ads/ghana/electronics?page=2,以此类推,直接构造对应页码的URL发起请求即可,稳定性远高于模拟点击。
修正后的核心代码
# -*- coding: utf-8 -*- import scrapy from scrapy_selenium import SeleniumRequest from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC class VisionSpider(scrapy.Spider): name = 'vision' # 控制最大爬取页码,可根据需求调整 max_page = 50 def start_requests(self): yield SeleniumRequest( url='https://tonaton.com', wait_time=3, screenshot=True, callback=self.parse ) def parse(self, response): businesses = response.xpath("//a[contains(@class, 'gtm-home-category-link-click')]") for business in businesses: link = business.xpath(".//@href").get() category = business.xpath(".//div[2]/p/text()").get() # 分类页默认从第一页开始爬 yield response.follow( url=link, callback=self.parse_business, meta={'business_category': category, 'current_page': 1} ) def parse_business(self, response): category = response.meta['business_category'] current_page = response.meta['current_page'] rows = response.xpath("//a[contains(@class, 'gtm-ad-item')]") for row in rows: new_link = row.xpath(".//@href").get() yield response.follow( url=new_link, callback=self.next_parse, meta={'business_category': category} ) # 构造下一页URL if current_page < self.max_page: next_page = f"{response.url.split('?')[0]}?page={current_page+1}" yield SeleniumRequest( url=next_page, wait_time=3, callback=self.parse_business, meta={'business_category': category, 'current_page': current_page+1} ) def next_parse(self, response): category = response.meta['business_category'] lines = response.xpath("//a[contains(@class, 'gtm-visit-shop')]") for line in lines: next_link = line.xpath(".//@href").get() yield SeleniumRequest( url=next_link, callback=self.another_parse, meta={'business_category': category}, wait_time=2 ) def another_parse(self, response): category = response.meta['business_category'] # 从meta中取selenium的driver对象来执行点击操作 driver = response.meta['driver'] try: # 等待显示手机号按钮加载完成后点击 show_btn = WebDriverWait(driver, 5).until( EC.element_to_be_clickable((By.XPATH, "//button[contains(@class, 'gtm-show-number')]")) ) show_btn.click() # 等待手机号加载完成 WebDriverWait(driver, 3).until( EC.presence_of_element_located((By.XPATH, "//div[contains(@class, 'info-container--')]")) ) # 把更新后的页面内容传给scrapy的response new_response = scrapy.http.HtmlResponse( url=driver.current_url, body=driver.page_source, encoding='utf-8', meta={'business_category': category} ) yield from self.new_parse(new_response) except Exception as e: self.logger.error(f"点击显示手机号失败: {e}") def new_parse(self, response): category = response.meta['business_category'] times = response.xpath("//div[contains(@class, 'info-container--')]") for time in times: name = time.xpath(".//div/span/text()").get() location = time.xpath(".//div/div/div/span/text()").get() phone = time.xpath(".//div[3]/div/button/div[2]/div/text()").get() yield { 'business_category': category, 'business_name': name, 'phone': phone, 'location': location }
注意事项
- 所有类名匹配都用了
contains部分匹配规则,不再依赖类名的随机哈希后缀,不会出现前端更新后选择器失效的问题 - 点击显示手机号的操作改为通过Selenium Driver执行,点击后重新构造包含最新页面内容的Response传给后续解析函数
- 可根据实际需求调整
max_page参数控制爬取的最大页码,避免无意义的空请求 - 建议在settings.py中设置
DOWNLOAD_DELAY = 2,控制请求频率,避免被站点封禁IP
内容的提问来源于stack exchange,提问作者Danny Stringz
相关产品推荐
相关产品推荐

