如何用Python的Selenium WebDriver实现分页元素的下一页跳转?
问题描述
尝试用Python的Selenium WebDriver操作网页分页元素跳转,分页的HTML结构如下:
<ul class="pagination"> <li class="paginate_button previous" id="datatable_previous"> <a href="#" aria-controls="datatable" data-dt-idx="0" tabindex="0">Précédent</a> </li> <li class="paginate_button "> <a href="#" aria-controls="datatable" data-dt-idx="1" tabindex="0">1</a> </li> <li class="paginate_button active"> <a href="#" aria-controls="datatable" data-dt-idx="2" tabindex="0">2</a> </li> <!-- More page numbers... --> <li class="paginate_button next" id="datatable_next"> <a href="#" aria-controls="datatable" data-dt-idx="8" tabindex="0">Suivant</a> </li> </ul>
目标是点击标注为“Suivant”的下一页按钮,但该按钮的<a>标签href为#,点击不会改变URL。尝试过Selenium的click()方法和execute_script执行点击,页面内容都没更新。
使用的代码如下:
def scrape_bloc_files(soup): for i, pdf in enumerate(soup.select("tr a[target='_blank']")): pdf_link = pdf['href'] driver.get(pdf_link) time.sleep(1) options = webdriver.ChromeOptions() options.add_experimental_option('prefs', { "download.default_directory": "/content/drive/MyDrive/Colab Notebooks/Files/resumes", "download.prompt_for_download": False, "download.directory_upgrade": True, "plugins.always_open_pdf_externally": True }) options.add_argument('--headless') options.add_argument('--no-sandbox') options.add_argument('--disable-dev-shm-usage') options.binary_location = '/usr/bin/chromium-browser' driver = webdriver.Chrome(options=options) driver.get('https://www.offre-emploi.tn/consulter-cv/') time.sleep(1) try: html = driver.page_source except UnexpectedAlertPresentException: alert = driver.switch_to.alert alert.accept() html = driver.page_source soup = BeautifulSoup(html, 'html.parser') while True: try: next_button = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.CSS_SELECTOR, "li#datatable_next a")) ) driver.execute_script("arguments[0].click();", next_button) scrape_bloc_files(soup) except Exception as e: print(f"Couldn't click the next button: {e}") break driver.quit()
疑问:为何点击按钮后页面内容不更新?如何成功实现分页跳转?
问题分析与解决方法
一、页面内容不更新的核心原因
- 静态Soup对象未更新:仅在页面初始加载时创建了一次
soup,点击分页后页面内容动态刷新,但始终用旧的soup解析数据,相当于完全没处理新页面内容;且点击后未等待新内容加载完成就执行后续操作,元素状态未同步。 - 页面跳转丢失上下文:
scrape_bloc_files中用driver.get(pdf_link)跳转到PDF页面,返回原分页页面后,原页面的JS事件绑定可能失效,导致分页按钮点击无反应。 - 未判断按钮可用性:到达最后一页时,“Suivant”按钮会被添加
disabled类,此时点击不会触发任何操作,但代码未做判断,会反复尝试无效点击。
二、修正步骤与代码
关键修改点
- 每次处理前重新获取当前页面的Soup对象,确保解析最新内容
- 用新标签页打开PDF,避免丢失原分页页面的上下文
- 检查分页按钮的
disabled状态,到达最后一页时停止循环 - 添加等待逻辑,确保分页后页面内容完全加载
修正后的代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import time from selenium.common.exceptions import UnexpectedAlertPresentException def scrape_bloc_files(driver): # 每次调用都重新解析当前页面的内容 current_html = driver.page_source soup = BeautifulSoup(current_html, 'html.parser') for pdf in soup.select("tr a[target='_blank']"): pdf_link = pdf['href'] # 用新标签页打开PDF,避免原页面上下文丢失 driver.execute_script("window.open(arguments[0]);", pdf_link) # 切换到新标签页 driver.switch_to.window(driver.window_handles[-1]) time.sleep(1) # 关闭新标签页并切回原分页页面 driver.close() driver.switch_to.window(driver.window_handles[0]) options = webdriver.ChromeOptions() options.add_experimental_option('prefs', { "download.default_directory": "/content/drive/MyDrive/Colab Notebooks/Files/resumes", "download.prompt_for_download": False, "download.directory_upgrade": True, "plugins.always_open_pdf_externally": True }) options.add_argument('--headless') options.add_argument('--no-sandbox') options.add_argument('--disable-dev-shm-usage') options.binary_location = '/usr/bin/chromium-browser' driver = webdriver.Chrome(options=options) driver.get('https://www.offre-emploi.tn/consulter-cv/') time.sleep(1) # 处理页面可能出现的弹窗 try: alert = driver.switch_to.alert alert.accept() except (UnexpectedAlertPresentException, Exception): pass while True: try: # 先处理当前页面的PDF内容 scrape_bloc_files(driver) # 检查下一页按钮是否可用(无disabled类) next_page_li = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, "datatable_next")) ) if 'disabled' in next_page_li.get_attribute('class'): print("已到达最后一页,停止分页") break # 点击下一页按钮 next_button = next_page_li.find_element(By.TAG_NAME, "a") driver.execute_script("arguments[0].click();", next_button) # 等待页面加载完成:等待原活跃页码元素失效,确保新页面已渲染 WebDriverWait(driver, 10).until( EC.staleness_of(driver.find_element(By.CSS_SELECTOR, "li.paginate_button.active")) ) time.sleep(0.5) # 可选短暂等待,确保内容完全加载 except Exception as e: print(f"操作出错: {e}") break driver.quit()
内容的提问来源于stack exchange,提问作者Rami
相关产品推荐
相关产品推荐

