You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy结合Selenium爬取律师名录仅获单页数据问题求解

问题核心原因

代码无法爬取后续分页是三个硬伤导致的:

  • 分页相关的Selenium逻辑写在了Spider类的全局作用域,不在任何类方法的执行流里,Scrapy运行时根本不会执行这部分代码,且代码未导入WebDriverWait、expected_conditions、By这几个Selenium依赖,跑到对应位置会直接抛异常。
  • Selenium操作和Scrapy解析流程完全脱节:Scrapy默认用自带下载器拉取第一页静态源码交给parse方法解析,Selenium就算点击翻页,渲染出的新页面内容也没有转成Scrapy可解析的Response对象,解析逻辑拿不到后续页内容。
  • 详情页解析逻辑混乱:原parse_book方法每次运行都要操控浏览器跳回律师列表首页,既拖慢爬取速度,还会打乱浏览器当前的分页状态。
可行修复方案

不要把Selenium逻辑直接堆在Spider类里,把分页渲染、翻页操作统一放到下载中间件处理,翻页后等待列表内容刷新完成,再把渲染好的页面源码返回给Spider解析,流程打通后即可自动爬取全部分页。

第一步:修正Spider代码

把散在类全局的分页逻辑整合到解析流程里,去掉无效的重复跳转逻辑,修正后代码如下:

import scrapy
from scrapy.http import Request, HtmlResponse
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

class TestSpider(scrapy.Spider):
    name = 'test'
    start_urls = ['https://www.ifep.ro/justice/lawyers/lawyerspanel.aspx']
    custom_settings = {
        'CONCURRENT_REQUESTS_PER_DOMAIN': 1,
        'DOWNLOAD_DELAY': 1,
        'USER_AGENT': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/79.0.3945.130 Safari/537.36',
        # 注册自定义Selenium中间件
        'DOWNLOADER_MIDDLEWARES': {
            'your_project_name.middlewares.SeleniumPaginationMiddleware': 543,
        }
    }

    def __init__(self):
        # 初始化浏览器驱动,实例交给中间件统一管理
        self.driver = webdriver.Chrome('C:\Program Files (x86)\chromedriver.exe')
        self.current_page = 1
        self.max_page = None
        self.wait = WebDriverWait(self.driver, 10)
        # 首次加载列表首页
        self.driver.get(self.start_urls[0])

    def closed(self, reason):
        # 爬虫结束自动关闭浏览器,避免残留进程
        self.driver.quit()

    def parse(self, response):
        # 首次访问时获取总页数
        if not self.max_page:
            max_page_el = self.wait.until(
                EC.visibility_of_element_located((By.ID, "MainContent_PagerTop_lblPages"))
            )
            self.max_page = int(max_page_el.text.split("din")[-1].split(")")[0])
        
        # 解析当前页所有律师详情链接
        lawyer_links = response.xpath("//div[@class='list-group']//@href").extract()
        for link in lawyer_links:
            url = response.urljoin(link)
            if url.endswith('.ro') or url.endswith('.ro/'):
                continue
            yield Request(url, callback=self.parse_lawyer_detail)
        
        # 存在未爬取页码时构造下一页请求,带上翻页标记
        if self.current_page < self.max_page:
            self.current_page += 1
            yield Request(
                url=self.start_urls[0],
                callback=self.parse,
                meta={'next_page': self.current_page},
                dont_filter=True
            )

    def parse_lawyer_detail(self, response):
        # 解析律师详情页字段,无需跳转回列表页
        title = response.xpath("//span[@id='HeadingContent_lblTitle']//text()").get().strip()
        d1 = response.xpath("//div[@class='col-md-10']//p[1]//text()").get().strip()
        d2 = response.xpath("//div[@class='col-md-10']//p[2]//text()").get().strip()
        d3 = response.xpath("//div[@class='col-md-10']//p[3]//span//text()").get().strip()
        d4 = response.xpath("//div[@class='col-md-10']//p[4]//text()").get().strip()

        yield{
            "title1": title,
            "title2": d1,
            "title3": d2,
            "title4": d3,
            "title5": d4,
        }

第二步:编写Selenium下载中间件

在项目的middlewares.py文件中添加如下中间件代码,处理翻页点击和页面渲染:

from scrapy.http import HtmlResponse
import time

class SeleniumPaginationMiddleware:
    def process_request(self, request, spider):
        # 仅处理带翻页标记的列表页请求,详情页走默认下载逻辑
        if 'next_page' in request.meta:
            target_page = request.meta['next_page']
            # 点击对应页码按钮
            page_btn = spider.wait.until(
                EC.element_to_be_clickable((By.ID, f"MainContent_PagerTop_NavToPage{target_page}"))
            )
            page_btn.click()
            # 等待列表区域刷新完成,避免拿到上一页旧内容
            time.sleep(0.5)
            spider.wait.until(
                EC.presence_of_all_elements_located((By.CLASS_NAME, "list-group"))
            )
        
        # 将浏览器渲染后的页面源码转成HtmlResponse交给Spider解析
        if spider.driver.current_url == request.url:
            body = spider.driver.page_source
            return HtmlResponse(
                url=request.url,
                body=body,
                encoding='utf-8',
                request=request
            )

关键注意点

  • 将代码里的your_project_name替换为自己的Scrapy项目名,否则中间件会加载失败。
  • 翻页点击后必须加等待逻辑,等页面DOM刷新完成再取源码,避免出现爬取重复内容、漏数据的问题。
  • 不要在详情页解析逻辑里操控浏览器跳转回列表页,详情页直接用Scrapy默认下载器请求即可,速度更快也不会打乱分页状态。
  • 测试时可以先把总页数限制为3-4页验证逻辑正常,再放开全量爬取,避免被站点封禁IP。

内容的提问来源于stack exchange,提问作者Amen Aziz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 14:36:30