Scrapy结合Selenium爬取分页网站仅能抓取首页问题求解
问题根因
- Selenium实例与Scrapy调度逻辑完全脱节:Scrapy默认通过自身下载器获取第一页响应,初始化的Chrome驱动从未主动访问目标列表页,调用翻页点击时驱动实际处于空白初始页,无法定位下一页按钮。
- 翻页后无新内容解析逻辑:原有代码的循环始终复用第一页的response对象提取详情链接,即使点击翻页成功,也不会读取新页面内容,所有提取操作仅对第一页生效。
- 存在冗余无效逻辑:
for k in range(1,10)循环无实际翻页作用,会重复提取9次第一页详情链接,造成重复请求。 - 细节兼容问题:未导入
scrapy.Request类,高版本Selenium已废弃find_element_by_xpath旧接口,点击翻页后无等待逻辑易触发元素不存在报错。
修正后完整代码
import scrapy from scrapy import Request from scrapy.http import HtmlResponse from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC class TestSpider(scrapy.Spider): name = 'test' start_urls = ['https://www.ifep.ro/justice/lawyers/lawyerspanel.aspx'] custom_settings = { 'CONCURRENT_REQUESTS_PER_DOMAIN': 1, 'DOWNLOAD_DELAY': 1, 'USER_AGENT': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/79.0.3945.130 Safari/537.36' } def __init__(self): self.driver = webdriver.Chrome('C:\Program Files (x86)\chromedriver.exe') self.wait = WebDriverWait(self.driver, 10) self.driver.maximize_window() def parse(self, response): # 首次访问通过Selenium加载目标页,后续翻页由Selenium直接操作 if not response.meta.get('is_next_page', False): self.driver.get(response.url) # 等待列表加载完成 self.wait.until(EC.presence_of_element_located((By.XPATH, "//div[@class='list-group']//a"))) # 基于Selenium获取的当前页源码构造可解析的响应对象 current_response = HtmlResponse( url=self.driver.current_url, body=self.driver.page_source, encoding='utf-8' ) # 提取当前页所有律师详情链接 books = current_response.xpath("//div[@class='list-group']//@href").extract() for book in books: url = current_response.urljoin(book) if url.endswith('.ro') or url.endswith('.ro/'): continue yield Request(url, callback=self.parse_book) # 定位下一页按钮,不存在或禁用则终止爬取 try: next_btn = self.driver.find_element(By.XPATH, "//a[@id='MainContent_PagerTop_NavNext']") if 'aspNetDisabled' in next_btn.get_attribute('class'): self.driver.quit() return next_btn.click() # 等待翻页后新列表加载完成 self.wait.until(EC.staleness_of(self.driver.find_element(By.XPATH, "//div[@class='list-group']//a"))) # 构造新页请求交回parse方法处理,实现翻页循环 yield Request( self.driver.current_url, callback=self.parse, meta={'is_next_page': True} ) except Exception: self.driver.quit() return def parse_book(self, response): title = response.xpath("//span[@id='HeadingContent_lblTitle']//text()").get() d1 = response.xpath("//div[@class='col-md-10']//p[1]//text()").get() d1 = d1.strip() if d1 else '' d2 = response.xpath("//div[@class='col-md-10']//p[2]//text()").get() d2 = d2.strip() if d2 else '' d3 = response.xpath("//div[@class='col-md-10']//p[3]//span//text()").get() d3 = d3.strip() if d3 else '' d4 = response.xpath("//div[@class='col-md-10']//p[4]//text()").get() d4 = d4.strip() if d4 else '' yield { "title1": title, "title2": d1, "title3": d2, "title4": d3, "title5": d4, }
关键修正点
- 删除原代码中无效的9次空循环,避免重复生成第一页的无效请求。
- 打通Selenium与Scrapy的数据链路:翻页后将Selenium获取到的新页面源码封装为Scrapy标准
HtmlResponse对象,交回parse方法循环处理,实现翻页逻辑和爬取调度的联动。 - 加入显式等待机制:翻页后等待旧列表元素失效、新列表加载完成再做内容提取,避免页面未加载完成导致的内容为空、元素定位失败问题。
- 替换Selenium已废弃的旧版元素定位API,补全缺失的
Request类导入。 - 增加终止判断:检测到下一页按钮为禁用状态(即到达最后一页)时自动退出驱动、终止翻页循环,避免无意义报错。
- 增加字段空值判断,避免XPath提取为空时调用
strip()触发空指针报错。
内容的提问来源于stack exchange,提问作者Amen Aziz
相关产品推荐
相关产品推荐

