You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页抓取:如何判断Next按钮是否存在并设置执行条件

解决TripAdvisor翻页爬取中判断Next按钮存在的问题

原代码的几个明显问题

  1. hasattr用法完全错误:你想用它判断按钮文本里有没有"Next",但hasattr是用来检查对象是否有某个属性的(比如判断字符串有没有strip方法),不是做子串匹配。而且目标站点是法语区TripAdvisor,Next按钮的实际文本是Suivant,就算用法正确也会判断失效。
  2. driver.find_element会直接抛异常:如果找不到Next按钮,这个方法会立刻抛出NoSuchElementException,根本走不到后续逻辑,没法实现“判断是否存在”的需求。
  3. 缩进错误:while True和try块的缩进不符合Python语法,会直接报错。
  4. links1未初始化:直接调用links1.append会触发NameError。

正确判断Next按钮是否存在的两种方法

方法1:用find_elements(推荐)

复数形式的find_elements找不到元素时会返回空列表,不会抛异常,直接判断列表长度就能知道按钮是否存在:

next_buttons = driver.find_elements(By.XPATH, './/a[@class="nav next rndBtn ui_button primary taLnk"]')
if len(next_buttons) > 0:
    # 存在Next按钮,执行翻页逻辑
    pass

方法2:用显式等待捕获超时异常

用WebDriverWait等待元素出现,超时则说明按钮不存在:

from selenium.common.exceptions import TimeoutException

try:
    next_button = WebDriverWait(driver, 5).until(
        EC.presence_of_element_located((By.XPATH, './/a[@class="nav next rndBtn ui_button primary taLnk"]'))
    )
    # 按钮存在,做后续操作
except TimeoutException:
    # 按钮不存在,跳过翻页
    pass

修正后的完整可运行代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException
import time

url = "https://www.tripadvisor.fr/Restaurants-g23084688-Beardsley_New_Brunswick.html"

# 初始化浏览器(这里用Chrome,确保ChromeDriver已配置到环境变量)
driver = webdriver.Chrome()
driver.get(url)
time.sleep(3)

# 初始化存储链接的列表,避免未定义错误
links1 = []

# 先爬取第一页的内容,避免遗漏
container = driver.find_elements(By.XPATH, "//a[@class='Lwqic Cj b']")
for item in container:
    link = item.get_attribute("href")
    links1.append(link)
    print(link)

# 循环翻页爬取
while True:
    try:
        # 等待Next按钮可点击,确保页面加载完成
        next_button = WebDriverWait(driver, 10).until(
            EC.element_to_be_clickable((By.XPATH, './/a[@class="nav next rndBtn ui_button primary taLnk"]'))
        )
        next_button.click()
        time.sleep(3)  # 等待新页面加载
        
        # 爬取当前页的链接
        container = driver.find_elements(By.XPATH, "//a[@class='Lwqic Cj b']")
        for item in container:
            link = item.get_attribute("href")
            links1.append(link)
            print(link)
    except TimeoutException:
        # 找不到Next按钮,退出循环
        break

# 将所有链接写入文件
with open("details.txt", "w") as filenam:
    for l in links1:
        filenam.write(l + "\n")

# 关闭浏览器,释放资源
driver.quit()

代码说明

  • 先爬第一页内容,避免因为直接进入翻页循环而漏掉初始页数据
  • 用WebDriverWait代替固定等待,更稳定地等待按钮可点击
  • 只捕获TimeoutException,精准处理“找不到Next按钮”的情况
  • 最后关闭浏览器,避免资源占用

内容的提问来源于stack exchange,提问作者Emma El

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 06:05:33