You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium使用XPATH遍历分页失败,持续出现NoSuchElementException错误

解决Selenium分页遍历的NoSuchElementException问题

你的代码报错核心原因是硬编码XPATH过于脆弱:依赖固定的tr[98]索引和修改字符串特定位置生成页码XPATH,一旦页面结果行数变化(导致分页栏位置偏移)、分页控件DOM结构变动,就会触发元素找不到的错误。另外,页面跳转后旧元素引用会失效,必须重新定位。

以下是基于XPATH的可靠分页实现,可处理包含“...”的分页场景:

from selenium import webdriver
from selenium.webdriver.support.ui import Select
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait 
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
import time

driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()))
url = 'https://www.sec.state.ma.us/LobbyistPublicSearch/Default.aspx'
wait = WebDriverWait(driver,10)

driver.get(url)

# 执行初始搜索流程
driver.find_element('id','ContentPlaceHolder1_rdbSearchByType').click()
select = Select(driver.find_element(By.CLASS_NAME,'p3'))
select.select_by_value('2020')
driver.find_element('id','ContentPlaceHolder1_btnSearch').click()

# 等待结果表格与分页栏加载完成
wait.until(EC.presence_of_element_located((By.ID, 'ContentPlaceHolder1_ucSearchResultByTypeAndCategory_grdvSearchResultByTypeAndCategory')))
# 动态定位分页栏(不依赖固定行索引)
pagination_table = wait.until(EC.presence_of_element_located((By.XPATH, '//table[@id="ContentPlaceHolder1_ucSearchResultByTypeAndCategory_grdvSearchResultByTypeAndCategory"]//td[@class="pager"]/table')))

# 提取总页数
pagination_text = pagination_table.text
total_pages = int(pagination_text.split('of')[-1].strip())

current_page = 1
while current_page < total_pages:
    # 页面跳转后重新定位分页栏(旧元素已失效)
    pagination_table = wait.until(EC.presence_of_element_located((By.XPATH, '//table[@id="ContentPlaceHolder1_ucSearchResultByTypeAndCategory_grdvSearchResultByTypeAndCategory"]//td[@class="pager"]/table')))
    page_elements = pagination_table.find_elements(By.TAG_NAME, 'a')
    
    next_page_found = False
    for elem in page_elements:
        # 直接定位下一页的页码按钮
        if elem.text.isdigit() and int(elem.text) == current_page + 1:
            wait.until(EC.element_to_be_clickable(elem)).click()
            current_page += 1
            next_page_found = True
            break
        # 处理“...”:当前页码列表中没有下一页时,点击“...”加载更多页码
        if elem.text == '...' and not next_page_found:
            wait.until(EC.element_to_be_clickable(elem)).click()
            time.sleep(1)
            break
    
    # 等待结果表格刷新完成
    wait.until(EC.staleness_of(driver.find_element(By.ID, 'ContentPlaceHolder1_ucSearchResultByTypeAndCategory_grdvSearchResultByTypeAndCategory')))
    wait.until(EC.presence_of_element_located((By.ID, 'ContentPlaceHolder1_ucSearchResultByTypeAndCategory_grdvSearchResultByTypeAndCategory')))

driver.quit()

关键优化点:

  • 动态定位分页栏:通过//td[@class="pager"]/table定位分页控件,摆脱固定行索引的限制,适配结果行数变化的场景。
  • 元素失效处理:每次分页跳转后,重新查找分页栏元素,避免旧DOM引用失效的问题。
  • “...”场景兼容:当目标页码不在当前显示列表中时,自动点击“...”加载更多页码,确保遍历所有页面。
  • 总页数驱动循环:从分页栏提取总页数,确保遍历逻辑覆盖全部页面,无遗漏或重复。

内容的提问来源于stack exchange,提问作者David Kahn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 01:15:31