You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何抓取XPath/CSS选择器无返回值的翻页按钮对应页面数据

问题根因
  • 类名动态失效:该网站使用了CSS模块化技术,类名后缀的随机哈希字符串会随前端版本更新变化,你硬编码的action-button--1O8tU这类类名无法稳定匹配到元素,这也是你在scrapy shell和splash里拿不到选择器的核心原因
  • 方法调用逻辑错误:scrapy的XPath返回的是Selector对象,没有click()方法,该方法是Selenium WebDriver对象的专属方法,原代码直接对Selector调用click()完全不会生效
  • 翻页流程错乱:你模拟点击后直接把点击操作的返回值当URL传入SeleniumRequest,参数类型本身就不匹配,且翻页后的回调错误指定为parse,会重复执行首页分类抓取逻辑,无法继续爬取列表数据
最优解决方案:直接构造翻页URL

该站点的翻页规则非常清晰,不需要模拟点击next按钮,完全可以规避选择器匹配问题:
所有分类列表页的翻页参数为page,第一页URL为https://tonaton.com/en/ads/ghana/electronics,第二页为https://tonaton.com/en/ads/ghana/electronics?page=2,以此类推,直接构造对应页码的URL发起请求即可,稳定性远高于模拟点击。

修正后的核心代码
# -*- coding: utf-8 -*-
import scrapy
from scrapy_selenium import SeleniumRequest
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

class VisionSpider(scrapy.Spider):
    name = 'vision'
    # 控制最大爬取页码,可根据需求调整
    max_page = 50

    def start_requests(self):
        yield SeleniumRequest(
            url='https://tonaton.com',
            wait_time=3,
            screenshot=True,
            callback=self.parse
        )

    def parse(self, response): 
        businesses = response.xpath("//a[contains(@class, 'gtm-home-category-link-click')]")
        for business in businesses:
            link = business.xpath(".//@href").get()
            category = business.xpath(".//div[2]/p/text()").get()
            # 分类页默认从第一页开始爬
            yield response.follow(
                url=link, 
                callback=self.parse_business, 
                meta={'business_category': category, 'current_page': 1}
            )

    def parse_business(self, response):
        category = response.meta['business_category']
        current_page = response.meta['current_page']
        rows = response.xpath("//a[contains(@class, 'gtm-ad-item')]")
        for row in rows:
            new_link = row.xpath(".//@href").get()
            yield response.follow(
                url=new_link, 
                callback=self.next_parse, 
                meta={'business_category': category}
            )
        # 构造下一页URL
        if current_page < self.max_page:
            next_page = f"{response.url.split('?')[0]}?page={current_page+1}"
            yield SeleniumRequest(
                url=next_page,
                wait_time=3,
                callback=self.parse_business,
                meta={'business_category': category, 'current_page': current_page+1}
            )

    def next_parse(self, response):
        category = response.meta['business_category']
        lines = response.xpath("//a[contains(@class, 'gtm-visit-shop')]")
        for line in lines:
            next_link = line.xpath(".//@href").get()
            yield SeleniumRequest(
                url=next_link, 
                callback=self.another_parse, 
                meta={'business_category': category},
                wait_time=2
            )

    def another_parse(self, response):
        category = response.meta['business_category']
        # 从meta中取selenium的driver对象来执行点击操作
        driver = response.meta['driver']
        try:
            # 等待显示手机号按钮加载完成后点击
            show_btn = WebDriverWait(driver, 5).until(
                EC.element_to_be_clickable((By.XPATH, "//button[contains(@class, 'gtm-show-number')]"))
            )
            show_btn.click()
            # 等待手机号加载完成
            WebDriverWait(driver, 3).until(
                EC.presence_of_element_located((By.XPATH, "//div[contains(@class, 'info-container--')]"))
            )
            # 把更新后的页面内容传给scrapy的response
            new_response = scrapy.http.HtmlResponse(
                url=driver.current_url,
                body=driver.page_source,
                encoding='utf-8',
                meta={'business_category': category}
            )
            yield from self.new_parse(new_response)
        except Exception as e:
            self.logger.error(f"点击显示手机号失败: {e}")

    def new_parse(self, response):
        category = response.meta['business_category']
        times = response.xpath("//div[contains(@class, 'info-container--')]")
        for time in times:
            name = time.xpath(".//div/span/text()").get()
            location = time.xpath(".//div/div/div/span/text()").get()
            phone = time.xpath(".//div[3]/div/button/div[2]/div/text()").get()
            yield {
                'business_category': category,
                'business_name': name,
                'phone': phone,
                'location': location
            }
注意事项
  • 所有类名匹配都用了contains部分匹配规则,不再依赖类名的随机哈希后缀,不会出现前端更新后选择器失效的问题
  • 点击显示手机号的操作改为通过Selenium Driver执行,点击后重新构造包含最新页面内容的Response传给后续解析函数
  • 可根据实际需求调整max_page参数控制爬取的最大页码,避免无意义的空请求
  • 建议在settings.py中设置DOWNLOAD_DELAY = 2,控制请求频率,避免被站点封禁IP

内容的提问来源于stack exchange,提问作者Danny Stringz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 17:57:04