Selenium操作后页面源码未更新,Scrapy爬虫仅获初始内容的问题
问题描述
我用以下代码爬取jornaleconomico.pt的经济板块文章标题,预期点击10次「Ver mais artigos」(查看更多文章)按钮后获取所有加载出的标题,但实际只能拿到初始的9条标题。通过options.add_experimental_option("detach", True)冻结Selenium窗口后,查看页面源码发现和点击前的初始页面一致,但窗口里能正常看到所有加载出的文章,已经用了WebDriverWait还是没解决问题。
爬取代码如下:
import scrapy from selenium.webdriver.chrome.options import Options from selenium import webdriver from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By from selenium.common.exceptions import TimeoutException, NoSuchElementException, StaleElementReferenceException class JornaleconomicoSpider(scrapy.Spider): name = 'jornaleconomico' allowed_domains = ['jornaleconomico.pt'] start_urls = ['https://jornaleconomico.pt/categoria/economia'] def parse(self, response): options = Options() driver_path = '###' #Your Chrome Webdriver Path browser_path = '###' #Your Google Chrome Path options.binary_location = browser_path options.add_experimental_option("detach", True) self.driver = webdriver.Chrome(options=options, executable_path=driver_path) self.driver.get(response.url) ignored_exceptions=(NoSuchElementException,StaleElementReferenceException,) wait = WebDriverWait(self.driver, 120, ignored_exceptions=ignored_exceptions) self.new_src = None self.new_response = None i=0 while i<10: # click next link try: element = wait.until(EC.element_to_be_clickable((By.XPATH, '*//div[@class="je-btn je-btn-more"]'))) self.driver.execute_script("arguments[0].click();", element) self.new_src = self.driver.page_source self.new_response = response.replace(body=self.new_src) i += 1 except TimeoutException: self.logger.info('No more pages to load.') self.driver.quit() break # grab the data headlines = self.new_response.xpath('*//h1[@class="je-post-title"]/a/text()').extract() for headline in headlines: yield { 'text': headline }
问题分析与解决办法
核心原因
问题出在点击按钮后立刻获取page_source,此时新内容尚未完成渲染,导致拿到的还是旧页面源码;同时每次点击后都覆盖new_response,但如果某次点击后渲染未完成,最终的new_response还是旧内容。
具体修复方案
1. 点击后等待新内容加载完成
不要点击后立刻取源码,而是等待页面中新增的文章元素出现,确保内容渲染完成。比如通过监控文章标题元素的数量变化来判断:
# 点击后记录当前文章数量,等待数量增加 current_count = len(self.driver.find_elements(By.XPATH, '*//h1[@class="je-post-title"]')) wait.until(lambda driver: len(driver.find_elements(By.XPATH, '*//h1[@class="je-post-title"]')) > current_count)
2. 最后统一获取最新页面源码
不需要每次点击后都更新new_response,等所有点击操作完成后,再一次性获取完整的渲染后源码,生成最终的response对象。
3. 优化元素定位与关闭时机
- 精准定位「查看更多」按钮,避免定位到失效元素;
- 不要在异常分支里提前关闭浏览器,等所有操作完成后统一关闭,防止中断后续源码获取。
修复后的完整代码
import scrapy from selenium.webdriver.chrome.options import Options from selenium import webdriver from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By from selenium.common.exceptions import TimeoutException, NoSuchElementException, StaleElementReferenceException class JornaleconomicoSpider(scrapy.Spider): name = 'jornaleconomico' allowed_domains = ['jornaleconomico.pt'] start_urls = ['https://jornaleconomico.pt/categoria/economia'] def parse(self, response): options = Options() driver_path = '###' # 你的ChromeDriver路径 browser_path = '###' # 你的Chrome浏览器路径 options.binary_location = browser_path options.add_experimental_option("detach", True) self.driver = webdriver.Chrome(options=options, executable_path=driver_path) self.driver.get(response.url) ignored_exceptions=(NoSuchElementException, StaleElementReferenceException,) wait = WebDriverWait(self.driver, 120, ignored_exceptions=ignored_exceptions) i = 0 while i < 10: try: # 精准定位「查看更多」按钮 element = wait.until(EC.element_to_be_clickable( (By.XPATH, '//div[contains(text(), "Ver mais artigos") and @class="je-btn je-btn-more"]') )) self.driver.execute_script("arguments[0].click();", element) # 等待新文章加载完成:监控标题元素数量变化 current_headline_count = len(self.driver.find_elements(By.XPATH, '*//h1[@class="je-post-title"]')) wait.until( lambda d: len(d.find_elements(By.XPATH, '*//h1[@class="je-post-title"]')) > current_headline_count ) i += 1 except TimeoutException: self.logger.info('没有更多页面可加载。') break # 所有点击完成后,获取最新页面源码 self.new_src = self.driver.page_source self.new_response = response.replace(body=self.new_src) # 提取所有标题并去除多余空格 headlines = self.new_response.xpath('*//h1[@class="je-post-title"]/a/text()').extract() for headline in headlines: yield {'text': headline.strip()} # 最后统一关闭浏览器 self.driver.quit()
内容的提问来源于stack exchange,提问作者Higo Felipe Silva Pires
相关产品推荐
相关产品推荐

