Scrapy结合Selenium爬取律师名录仅获单页数据问题求解
问题核心原因
代码无法爬取后续分页是三个硬伤导致的:
- 分页相关的Selenium逻辑写在了Spider类的全局作用域,不在任何类方法的执行流里,Scrapy运行时根本不会执行这部分代码,且代码未导入
WebDriverWait、expected_conditions、By这几个Selenium依赖,跑到对应位置会直接抛异常。 - Selenium操作和Scrapy解析流程完全脱节:Scrapy默认用自带下载器拉取第一页静态源码交给
parse方法解析,Selenium就算点击翻页,渲染出的新页面内容也没有转成Scrapy可解析的Response对象,解析逻辑拿不到后续页内容。 - 详情页解析逻辑混乱:原
parse_book方法每次运行都要操控浏览器跳回律师列表首页,既拖慢爬取速度,还会打乱浏览器当前的分页状态。
可行修复方案
不要把Selenium逻辑直接堆在Spider类里,把分页渲染、翻页操作统一放到下载中间件处理,翻页后等待列表内容刷新完成,再把渲染好的页面源码返回给Spider解析,流程打通后即可自动爬取全部分页。
第一步:修正Spider代码
把散在类全局的分页逻辑整合到解析流程里,去掉无效的重复跳转逻辑,修正后代码如下:
import scrapy from scrapy.http import Request, HtmlResponse from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time class TestSpider(scrapy.Spider): name = 'test' start_urls = ['https://www.ifep.ro/justice/lawyers/lawyerspanel.aspx'] custom_settings = { 'CONCURRENT_REQUESTS_PER_DOMAIN': 1, 'DOWNLOAD_DELAY': 1, 'USER_AGENT': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/79.0.3945.130 Safari/537.36', # 注册自定义Selenium中间件 'DOWNLOADER_MIDDLEWARES': { 'your_project_name.middlewares.SeleniumPaginationMiddleware': 543, } } def __init__(self): # 初始化浏览器驱动,实例交给中间件统一管理 self.driver = webdriver.Chrome('C:\Program Files (x86)\chromedriver.exe') self.current_page = 1 self.max_page = None self.wait = WebDriverWait(self.driver, 10) # 首次加载列表首页 self.driver.get(self.start_urls[0]) def closed(self, reason): # 爬虫结束自动关闭浏览器,避免残留进程 self.driver.quit() def parse(self, response): # 首次访问时获取总页数 if not self.max_page: max_page_el = self.wait.until( EC.visibility_of_element_located((By.ID, "MainContent_PagerTop_lblPages")) ) self.max_page = int(max_page_el.text.split("din")[-1].split(")")[0]) # 解析当前页所有律师详情链接 lawyer_links = response.xpath("//div[@class='list-group']//@href").extract() for link in lawyer_links: url = response.urljoin(link) if url.endswith('.ro') or url.endswith('.ro/'): continue yield Request(url, callback=self.parse_lawyer_detail) # 存在未爬取页码时构造下一页请求,带上翻页标记 if self.current_page < self.max_page: self.current_page += 1 yield Request( url=self.start_urls[0], callback=self.parse, meta={'next_page': self.current_page}, dont_filter=True ) def parse_lawyer_detail(self, response): # 解析律师详情页字段,无需跳转回列表页 title = response.xpath("//span[@id='HeadingContent_lblTitle']//text()").get().strip() d1 = response.xpath("//div[@class='col-md-10']//p[1]//text()").get().strip() d2 = response.xpath("//div[@class='col-md-10']//p[2]//text()").get().strip() d3 = response.xpath("//div[@class='col-md-10']//p[3]//span//text()").get().strip() d4 = response.xpath("//div[@class='col-md-10']//p[4]//text()").get().strip() yield{ "title1": title, "title2": d1, "title3": d2, "title4": d3, "title5": d4, }
第二步:编写Selenium下载中间件
在项目的middlewares.py文件中添加如下中间件代码,处理翻页点击和页面渲染:
from scrapy.http import HtmlResponse import time class SeleniumPaginationMiddleware: def process_request(self, request, spider): # 仅处理带翻页标记的列表页请求,详情页走默认下载逻辑 if 'next_page' in request.meta: target_page = request.meta['next_page'] # 点击对应页码按钮 page_btn = spider.wait.until( EC.element_to_be_clickable((By.ID, f"MainContent_PagerTop_NavToPage{target_page}")) ) page_btn.click() # 等待列表区域刷新完成,避免拿到上一页旧内容 time.sleep(0.5) spider.wait.until( EC.presence_of_all_elements_located((By.CLASS_NAME, "list-group")) ) # 将浏览器渲染后的页面源码转成HtmlResponse交给Spider解析 if spider.driver.current_url == request.url: body = spider.driver.page_source return HtmlResponse( url=request.url, body=body, encoding='utf-8', request=request )
关键注意点
- 将代码里的
your_project_name替换为自己的Scrapy项目名,否则中间件会加载失败。 - 翻页点击后必须加等待逻辑,等页面DOM刷新完成再取源码,避免出现爬取重复内容、漏数据的问题。
- 不要在详情页解析逻辑里操控浏览器跳转回列表页,详情页直接用Scrapy默认下载器请求即可,速度更快也不会打乱分页状态。
- 测试时可以先把总页数限制为3-4页验证逻辑正常,再放开全量爬取,避免被站点封禁IP。
内容的提问来源于stack exchange,提问作者Amen Aziz
相关产品推荐
相关产品推荐

