You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Selenium无法点击分页按钮,报错NoSuchElementException求助

问题:Selenium无法定位并点击分页按钮(NoSuchElementException)

Hey there, let's break down why you're hitting that NoSuchElementException and fix your pub-scraping script.

你的场景

You're trying to scrape pub links from the Good Pub Guide's London listings. You can pull links from the first page, but when you try to click the "next page" arrow at the bottom, Selenium can't find the element.

报错信息

selenium.common.exceptions.NoSuchElementException: Message: no such element: Unable to locate element: {"method":"xpath","selector":"//*[@id="search-results"]/div[2]/div/div/ul/li[8]/a"}

原代码

from selenium import webdriver 
from selenium.webdriver.common.keys import Keys 
import time 
from datetime import datetime 
import csv 
from selenium.webdriver.common.action_chains import ActionChains 

#User login info 
pagenum = 1 

#Creates link to Chrome Driver and shortens this to 'browser' 
path_to_chromedriver = '/Users/abc/Downloads/chromedriver 2' # change path as needed 
driver = webdriver.Chrome(executable_path = path_to_chromedriver) 

#Navigates Chrome to the specified page 
url = 'https://thegoodpubguide.co.uk/pubs/?paged=1&order_by=category&search=pubs&pub_name=&postal_code=&region=london' 

#Clicks Login 
def findlinks(address): 
    global pagenum 
    list = [] 
    driver.get(address) 
    #wait 
    while pagenum <= 2: 
        for i in range(20): 
            # Scrapes available links 
            xref = '//*[@id="search-results"]/div[1]/div[' + str(i+1) + ']/div/div/div[2]/div[1]/p/a' 
            link = driver.find_element_by_xpath(xref).get_attribute('href') 
            print(link) 
            list.append(link) 
        with open("links.csv", "a") as fp: 
            # Saves list to file 
            wr = csv.writer(fp, dialect='excel') 
            wr.writerow(list) 
        print(pagenum) 
        pagenum = pagenum + 1 
        element = driver.find_element_by_xpath('//*[@id="search-results"]/div[2]/div/div/ul/li[8]/a') 
        element.click() 
findlinks(url) 

为什么会出错?

  1. 固定索引的XPath太脆弱:你用li[8]定位下一页按钮,但分页列表的位置可能随页面数量、UI更新变化,硬编码索引很容易失效。
  2. 没有等待页面加载:你刚递增页码就尝试点击按钮,但分页栏可能还没完全加载完成。
  3. 递归调用的逻辑bug:末尾调用findlinks(url)会重新加载第一页,直接重置了你的爬取进度。
  4. 过时的Selenium语法:如果用的是Selenium 4+,executable_path已经被弃用,应该用Service类来配置驱动。

修复方案

这里是修改后的脚本,解决了所有问题:

核心改进点

  • 用稳定的元素定位方式(比如按文本或类名找下一页按钮)
  • 添加显式等待,确保元素加载完成后再操作
  • 移除递归,用清晰的循环逻辑遍历页面
  • 适配Selenium 4+的驱动配置

修订后的代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import csv
from selenium.webdriver.chrome.service import Service

# 配置Chrome驱动(兼容Selenium 4+)
driver_path = '/Users/abc/Downloads/chromedriver 2'
service = Service(driver_path)
driver = webdriver.Chrome(service=service)

base_url = 'https://thegoodpubguide.co.uk/pubs/?paged=1&order_by=category&search=pubs&pub_name=&postal_code=&region=london'
driver.get(base_url)

# 初始化CSV文件(添加表头)
with open("pub_links.csv", "w", newline="") as fp:
    writer = csv.writer(fp)
    writer.writerow(["Pub Link"])

pagenum = 1
max_pages = 5  # 设置你要爬取的最大页数,移除则爬取所有页面

while pagenum <= max_pages:
    print(f"正在爬取第{pagenum}页...")
    
    # 等待搜索结果加载完成
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.XPATH, '//*[@id="search-results"]/div[1]/div'))
    )
    
    # 动态获取当前页所有酒吧链接(避免硬编码数量)
    pub_links = driver.find_elements(By.XPATH, '//*[@id="search-results"]/div[1]/div/div/div/div[2]/div[1]/p/a')
    links_list = [link.get_attribute('href') for link in pub_links]
    
    # 将链接写入CSV
    with open("pub_links.csv", "a", newline="") as fp:
        writer = csv.writer(fp)
        for link in links_list:
            writer.writerow([link])
    
    # 尝试点击下一页
    try:
        # 等待下一页按钮可点击
        next_button = WebDriverWait(driver, 10).until(
            EC.element_to_be_clickable((By.LINK_TEXT, 'Next'))  # 如果文本不同,可替换为CSS选择器
            # 备选方案:EC.element_to_be_clickable((By.CSS_SELECTOR, 'li.pagination-next a'))
        )
        next_button.click()
        pagenum += 1
    except:
        # 没有更多页面或按钮找不到
        print("没有更多页面可爬取了。")
        break

driver.quit()
print("爬取完成!")

额外提示

  • 检查反爬机制:部分网站会拦截自动化工具,如果仍有问题,可以尝试添加自定义User-Agent,或者完成登录流程(你的原代码有登录注释但未实现)。
  • 避免硬编码范围:用find_elements动态获取所有链接,而不是range(20),这样能处理少于20个酒吧的页面。
  • 使用相对XPath:相对路径比绝对路径更稳定,比如//div[contains(@class,'pub-item')]//a可能比你用的长绝对路径更可靠。

内容的提问来源于stack exchange,提问作者lordf

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 10:06:40