You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python+Selenium获取href属性值并处理分页点击异常求助

解决Selenium分页点击时的NoSuchElementException问题

我来帮你搞定这个问题!你遇到的NoSuchElementException是因为直接用driver.find_element_by_link_text('»')的时候,一旦元素找不到,Selenium不会返回False,而是直接抛出异常——这就是你代码跑不起来的核心原因。

先分析你现有代码的问题

  • 第一行if not driver.find_element_by_link_text('»'): break:如果»元素不存在,这里直接抛异常,根本走不到break
  • 后面的xpath判断driver.find_element_by_xpath(...)同理,找不到元素就报错,不会进入else分支执行break

正确的解决方案

我们需要先安全地检查元素是否存在,再获取它的href属性,判断是否符合点击条件。这里有两种靠谱的写法:

写法一:先获取元素再判断href

from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException

while True:
    try:
        # 等待»链接加载出来,超时20秒
        next_link = WebDriverWait(driver, 20).until(
            EC.presence_of_element_located((By.LINK_TEXT, "»"))
        )
        # 获取链接的href属性值
        href_value = next_link.get_attribute("href")
        
        if href_value != "#":
            # 确保元素可点击后再点击
            clickable_link = WebDriverWait(driver, 20).until(
                EC.element_to_be_clickable((By.LINK_TEXT, "»"))
            )
            clickable_link.click()
            # 可选:等待页面切换完成,避免重复操作
            # WebDriverWait(driver, 20).until(EC.staleness_of(next_link))
        else:
            # href是#,说明没有下一页了,退出循环
            break
    except TimeoutException:
        # 找不到»元素,直接退出循环
        break

写法二:用XPath直接筛选有效按钮

这种方法更简洁,直接通过XPath定位href不等于#的»按钮,找不到就说明没有下一页了:

from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException

while True:
    try:
        # 直接定位符合条件的可点击按钮:文本是»且href不等于#
        next_link = WebDriverWait(driver, 20).until(
            EC.element_to_be_clickable((By.XPATH, "//a[.='»' and @href!='#']"))
        )
        next_link.click()
        # 可选:等待页面加载完成
        # WebDriverWait(driver, 20).until(EC.staleness_of(next_link))
    except TimeoutException:
        # 找不到有效按钮,退出循环
        break

额外提示

  • 点击后一定要等待页面加载完成,不然可能会出现重复点击或者元素定位错误的情况,用EC.staleness_of(next_link)等待旧元素失效是个不错的选择
  • 如果网页最后一页的»按钮会加上disabled类(比如你提供的源码里«按钮有disabled类),可以把XPath优化成//a[.='»' and @href!='#' and not(ancestor::li[@class='disabled'])],判断会更准确

内容的提问来源于stack exchange,提问作者Mohammed Benaou

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 06:49:37