You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium操作后页面源码未更新,Scrapy爬虫仅获初始内容的问题

问题描述

我用以下代码爬取jornaleconomico.pt的经济板块文章标题,预期点击10次「Ver mais artigos」(查看更多文章)按钮后获取所有加载出的标题,但实际只能拿到初始的9条标题。通过options.add_experimental_option("detach", True)冻结Selenium窗口后,查看页面源码发现和点击前的初始页面一致,但窗口里能正常看到所有加载出的文章,已经用了WebDriverWait还是没解决问题。

爬取代码如下:

import scrapy
from selenium.webdriver.chrome.options import Options
from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from selenium.common.exceptions import TimeoutException, NoSuchElementException, StaleElementReferenceException


class JornaleconomicoSpider(scrapy.Spider):
    name = 'jornaleconomico'
    allowed_domains = ['jornaleconomico.pt']
    start_urls = ['https://jornaleconomico.pt/categoria/economia']

    def parse(self, response):
        options = Options()
        driver_path = '###' #Your Chrome Webdriver Path
        browser_path = '###' #Your Google Chrome Path
        options.binary_location = browser_path
        options.add_experimental_option("detach", True)

        self.driver = webdriver.Chrome(options=options, executable_path=driver_path)
        self.driver.get(response.url)

        ignored_exceptions=(NoSuchElementException,StaleElementReferenceException,)
        wait = WebDriverWait(self.driver, 120, ignored_exceptions=ignored_exceptions)

        self.new_src = None
        self.new_response = None

        i=0

        while i<10:
            # click next link
            try:
                element = wait.until(EC.element_to_be_clickable((By.XPATH, '*//div[@class="je-btn je-btn-more"]')))
                self.driver.execute_script("arguments[0].click();", element)
                self.new_src = self.driver.page_source
                self.new_response = response.replace(body=self.new_src)
                i += 1
            except TimeoutException:
                self.logger.info('No more pages to load.')
                self.driver.quit()
                break
            
        # grab the data
        headlines = self.new_response.xpath('*//h1[@class="je-post-title"]/a/text()').extract()

        for headline in headlines:
            yield {
            'text': headline
        }
问题分析与解决办法

核心原因

问题出在点击按钮后立刻获取page_source,此时新内容尚未完成渲染,导致拿到的还是旧页面源码;同时每次点击后都覆盖new_response,但如果某次点击后渲染未完成,最终的new_response还是旧内容。

具体修复方案

1. 点击后等待新内容加载完成

不要点击后立刻取源码,而是等待页面中新增的文章元素出现,确保内容渲染完成。比如通过监控文章标题元素的数量变化来判断:

# 点击后记录当前文章数量,等待数量增加
current_count = len(self.driver.find_elements(By.XPATH, '*//h1[@class="je-post-title"]'))
wait.until(lambda driver: len(driver.find_elements(By.XPATH, '*//h1[@class="je-post-title"]')) > current_count)

2. 最后统一获取最新页面源码

不需要每次点击后都更新new_response,等所有点击操作完成后,再一次性获取完整的渲染后源码,生成最终的response对象。

3. 优化元素定位与关闭时机

  • 精准定位「查看更多」按钮,避免定位到失效元素;
  • 不要在异常分支里提前关闭浏览器,等所有操作完成后统一关闭,防止中断后续源码获取。

修复后的完整代码

import scrapy
from selenium.webdriver.chrome.options import Options
from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from selenium.common.exceptions import TimeoutException, NoSuchElementException, StaleElementReferenceException


class JornaleconomicoSpider(scrapy.Spider):
    name = 'jornaleconomico'
    allowed_domains = ['jornaleconomico.pt']
    start_urls = ['https://jornaleconomico.pt/categoria/economia']

    def parse(self, response):
        options = Options()
        driver_path = '###' # 你的ChromeDriver路径
        browser_path = '###' # 你的Chrome浏览器路径
        options.binary_location = browser_path
        options.add_experimental_option("detach", True)

        self.driver = webdriver.Chrome(options=options, executable_path=driver_path)
        self.driver.get(response.url)

        ignored_exceptions=(NoSuchElementException, StaleElementReferenceException,)
        wait = WebDriverWait(self.driver, 120, ignored_exceptions=ignored_exceptions)

        i = 0
        while i < 10:
            try:
                # 精准定位「查看更多」按钮
                element = wait.until(EC.element_to_be_clickable(
                    (By.XPATH, '//div[contains(text(), "Ver mais artigos") and @class="je-btn je-btn-more"]')
                ))
                self.driver.execute_script("arguments[0].click();", element)
                
                # 等待新文章加载完成:监控标题元素数量变化
                current_headline_count = len(self.driver.find_elements(By.XPATH, '*//h1[@class="je-post-title"]'))
                wait.until(
                    lambda d: len(d.find_elements(By.XPATH, '*//h1[@class="je-post-title"]')) > current_headline_count
                )
                i += 1
            except TimeoutException:
                self.logger.info('没有更多页面可加载。')
                break

        # 所有点击完成后,获取最新页面源码
        self.new_src = self.driver.page_source
        self.new_response = response.replace(body=self.new_src)
        
        # 提取所有标题并去除多余空格
        headlines = self.new_response.xpath('*//h1[@class="je-post-title"]/a/text()').extract()
        for headline in headlines:
            yield {'text': headline.strip()}
        
        # 最后统一关闭浏览器
        self.driver.quit()

内容的提问来源于stack exchange,提问作者Higo Felipe Silva Pires

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 01:55:35