You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy结合Selenium爬取分页网站仅能抓取首页问题求解

问题根因
  • Selenium实例与Scrapy调度逻辑完全脱节:Scrapy默认通过自身下载器获取第一页响应,初始化的Chrome驱动从未主动访问目标列表页,调用翻页点击时驱动实际处于空白初始页,无法定位下一页按钮。
  • 翻页后无新内容解析逻辑:原有代码的循环始终复用第一页的response对象提取详情链接,即使点击翻页成功,也不会读取新页面内容,所有提取操作仅对第一页生效。
  • 存在冗余无效逻辑:for k in range(1,10)循环无实际翻页作用,会重复提取9次第一页详情链接,造成重复请求。
  • 细节兼容问题:未导入scrapy.Request类,高版本Selenium已废弃find_element_by_xpath旧接口,点击翻页后无等待逻辑易触发元素不存在报错。
修正后完整代码
import scrapy
from scrapy import Request
from scrapy.http import HtmlResponse
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC


class TestSpider(scrapy.Spider):
    name = 'test'
    start_urls = ['https://www.ifep.ro/justice/lawyers/lawyerspanel.aspx']
    custom_settings = {
        'CONCURRENT_REQUESTS_PER_DOMAIN': 1,
        'DOWNLOAD_DELAY': 1,
        'USER_AGENT': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/79.0.3945.130 Safari/537.36'
    }

    def __init__(self):
        self.driver = webdriver.Chrome('C:\Program Files (x86)\chromedriver.exe')
        self.wait = WebDriverWait(self.driver, 10)
        self.driver.maximize_window()

    def parse(self, response):
        # 首次访问通过Selenium加载目标页,后续翻页由Selenium直接操作
        if not response.meta.get('is_next_page', False):
            self.driver.get(response.url)
            # 等待列表加载完成
            self.wait.until(EC.presence_of_element_located((By.XPATH, "//div[@class='list-group']//a")))

        # 基于Selenium获取的当前页源码构造可解析的响应对象
        current_response = HtmlResponse(
            url=self.driver.current_url,
            body=self.driver.page_source,
            encoding='utf-8'
        )
        # 提取当前页所有律师详情链接
        books = current_response.xpath("//div[@class='list-group']//@href").extract()
        for book in books:
            url = current_response.urljoin(book)
            if url.endswith('.ro') or url.endswith('.ro/'):
                continue
            yield Request(url, callback=self.parse_book)

        # 定位下一页按钮,不存在或禁用则终止爬取
        try:
            next_btn = self.driver.find_element(By.XPATH, "//a[@id='MainContent_PagerTop_NavNext']")
            if 'aspNetDisabled' in next_btn.get_attribute('class'):
                self.driver.quit()
                return
            next_btn.click()
            # 等待翻页后新列表加载完成
            self.wait.until(EC.staleness_of(self.driver.find_element(By.XPATH, "//div[@class='list-group']//a")))
            # 构造新页请求交回parse方法处理,实现翻页循环
            yield Request(
                self.driver.current_url,
                callback=self.parse,
                meta={'is_next_page': True}
            )
        except Exception:
            self.driver.quit()
            return

    def parse_book(self, response):
        title = response.xpath("//span[@id='HeadingContent_lblTitle']//text()").get()
        d1 = response.xpath("//div[@class='col-md-10']//p[1]//text()").get()
        d1 = d1.strip() if d1 else ''
        d2 = response.xpath("//div[@class='col-md-10']//p[2]//text()").get()
        d2 = d2.strip() if d2 else ''
        d3 = response.xpath("//div[@class='col-md-10']//p[3]//span//text()").get()
        d3 = d3.strip() if d3 else ''
        d4 = response.xpath("//div[@class='col-md-10']//p[4]//text()").get()
        d4 = d4.strip() if d4 else ''

        yield {
            "title1": title,
            "title2": d1,
            "title3": d2,
            "title4": d3,
            "title5": d4,
        }
关键修正点
  • 删除原代码中无效的9次空循环,避免重复生成第一页的无效请求。
  • 打通Selenium与Scrapy的数据链路:翻页后将Selenium获取到的新页面源码封装为Scrapy标准HtmlResponse对象,交回parse方法循环处理,实现翻页逻辑和爬取调度的联动。
  • 加入显式等待机制:翻页后等待旧列表元素失效、新列表加载完成再做内容提取,避免页面未加载完成导致的内容为空、元素定位失败问题。
  • 替换Selenium已废弃的旧版元素定位API,补全缺失的Request类导入。
  • 增加终止判断:检测到下一页按钮为禁用状态(即到达最后一页)时自动退出驱动、终止翻页循环,避免无意义报错。
  • 增加字段空值判断,避免XPath提取为空时调用strip()触发空指针报错。

内容的提问来源于stack exchange,提问作者Amen Aziz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 01:42:28