You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的Selenium WebDriver实现分页元素的下一页跳转?

问题描述

尝试用Python的Selenium WebDriver操作网页分页元素跳转,分页的HTML结构如下:

<ul class="pagination">
    <li class="paginate_button previous" id="datatable_previous">
        <a href="#" aria-controls="datatable" data-dt-idx="0" tabindex="0">Précédent</a>
    </li>
    <li class="paginate_button ">
        <a href="#" aria-controls="datatable" data-dt-idx="1" tabindex="0">1</a>
    </li>
    <li class="paginate_button active">
        <a href="#" aria-controls="datatable" data-dt-idx="2" tabindex="0">2</a>
    </li>
    <!-- More page numbers... -->
    <li class="paginate_button next" id="datatable_next">
        <a href="#" aria-controls="datatable" data-dt-idx="8" tabindex="0">Suivant</a>
    </li>
</ul>

目标是点击标注为“Suivant”的下一页按钮,但该按钮的<a>标签href为#,点击不会改变URL。尝试过Selenium的click()方法和execute_script执行点击,页面内容都没更新。

使用的代码如下:

def scrape_bloc_files(soup):
    for i, pdf in enumerate(soup.select("tr a[target='_blank']")):
        pdf_link = pdf['href']
        driver.get(pdf_link)
        time.sleep(1)


options = webdriver.ChromeOptions()
options.add_experimental_option('prefs', {
    "download.default_directory": "/content/drive/MyDrive/Colab Notebooks/Files/resumes",
    "download.prompt_for_download": False,
    "download.directory_upgrade": True,
    "plugins.always_open_pdf_externally": True
})
options.add_argument('--headless')
options.add_argument('--no-sandbox')
options.add_argument('--disable-dev-shm-usage')
options.binary_location = '/usr/bin/chromium-browser'

driver = webdriver.Chrome(options=options)
driver.get('https://www.offre-emploi.tn/consulter-cv/')
time.sleep(1)

try:
    html = driver.page_source
except UnexpectedAlertPresentException:
    alert = driver.switch_to.alert
    alert.accept()
    html = driver.page_source

soup = BeautifulSoup(html, 'html.parser')

while True:
    try:
        next_button = WebDriverWait(driver, 10).until(
            EC.element_to_be_clickable((By.CSS_SELECTOR, "li#datatable_next a"))
        )
        driver.execute_script("arguments[0].click();", next_button)
        scrape_bloc_files(soup)
    except Exception as e:
        print(f"Couldn't click the next button: {e}")
        break

driver.quit()

疑问:为何点击按钮后页面内容不更新?如何成功实现分页跳转?


问题分析与解决方法

一、页面内容不更新的核心原因

  1. 静态Soup对象未更新:仅在页面初始加载时创建了一次soup,点击分页后页面内容动态刷新,但始终用旧的soup解析数据,相当于完全没处理新页面内容;且点击后未等待新内容加载完成就执行后续操作,元素状态未同步。
  2. 页面跳转丢失上下文:scrape_bloc_files中用driver.get(pdf_link)跳转到PDF页面,返回原分页页面后,原页面的JS事件绑定可能失效,导致分页按钮点击无反应。
  3. 未判断按钮可用性:到达最后一页时,“Suivant”按钮会被添加disabled类,此时点击不会触发任何操作,但代码未做判断,会反复尝试无效点击。

二、修正步骤与代码

关键修改点

  • 每次处理前重新获取当前页面的Soup对象,确保解析最新内容
  • 用新标签页打开PDF,避免丢失原分页页面的上下文
  • 检查分页按钮的disabled状态,到达最后一页时停止循环
  • 添加等待逻辑,确保分页后页面内容完全加载

修正后的代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import time
from selenium.common.exceptions import UnexpectedAlertPresentException

def scrape_bloc_files(driver):
    # 每次调用都重新解析当前页面的内容
    current_html = driver.page_source
    soup = BeautifulSoup(current_html, 'html.parser')
    for pdf in soup.select("tr a[target='_blank']"):
        pdf_link = pdf['href']
        # 用新标签页打开PDF,避免原页面上下文丢失
        driver.execute_script("window.open(arguments[0]);", pdf_link)
        # 切换到新标签页
        driver.switch_to.window(driver.window_handles[-1])
        time.sleep(1)
        # 关闭新标签页并切回原分页页面
        driver.close()
        driver.switch_to.window(driver.window_handles[0])

options = webdriver.ChromeOptions()
options.add_experimental_option('prefs', {
    "download.default_directory": "/content/drive/MyDrive/Colab Notebooks/Files/resumes",
    "download.prompt_for_download": False,
    "download.directory_upgrade": True,
    "plugins.always_open_pdf_externally": True
})
options.add_argument('--headless')
options.add_argument('--no-sandbox')
options.add_argument('--disable-dev-shm-usage')
options.binary_location = '/usr/bin/chromium-browser'

driver = webdriver.Chrome(options=options)
driver.get('https://www.offre-emploi.tn/consulter-cv/')
time.sleep(1)

# 处理页面可能出现的弹窗
try:
    alert = driver.switch_to.alert
    alert.accept()
except (UnexpectedAlertPresentException, Exception):
    pass

while True:
    try:
        # 先处理当前页面的PDF内容
        scrape_bloc_files(driver)
        
        # 检查下一页按钮是否可用(无disabled类)
        next_page_li = WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.ID, "datatable_next"))
        )
        if 'disabled' in next_page_li.get_attribute('class'):
            print("已到达最后一页,停止分页")
            break
        
        # 点击下一页按钮
        next_button = next_page_li.find_element(By.TAG_NAME, "a")
        driver.execute_script("arguments[0].click();", next_button)
        
        # 等待页面加载完成:等待原活跃页码元素失效,确保新页面已渲染
        WebDriverWait(driver, 10).until(
            EC.staleness_of(driver.find_element(By.CSS_SELECTOR, "li.paginate_button.active"))
        )
        time.sleep(0.5)  # 可选短暂等待,确保内容完全加载
        
    except Exception as e:
        print(f"操作出错: {e}")
        break

driver.quit()

内容的提问来源于stack exchange,提问作者Rami

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 09:06:11