You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium分页表格爬取循环无法终止及按钮不可点击检测问题

分页表格爬取:解决Next按钮循环无法终止的问题

我需要从分页网页提取表格,最初采用手动点击Next按钮的方式爬取,代码如下:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
from selenium.common.exceptions import TimeoutException
import requests
import pandas as pd


# Set up the web driver
driver = webdriver.Chrome()
driver.get('website_url')


tables =[]

soup = BeautifulSoup(driver.page_source, 'html.parser')
table = soup.find('table', {"class": "tablepress tablepress-id-2 dataTable no-footer"})
tables.append(table)

# Click the next button and extract the second table
next_button = WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.LINK_TEXT, 'Next')))
next_button.click()
soup = BeautifulSoup(driver.page_source, 'html.parser')
table = soup.find('table', {"class": "tablepress tablepress-id-2 dataTable no-footer"})
tables.append(table)

# Click the next button and extract the second table
next_button = WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.LINK_TEXT, 'Next')))
next_button.click()
soup = BeautifulSoup(driver.page_source, 'html.parser')
table = soup.find('table', {"class": "tablepress tablepress-id-2 dataTable no-footer"})
tables.append(table)


driver.quit()


# Convert each table to a DataFrame and concatenate them into a single DataFrame
dfs = []
for table in tables:
    df = pd.read_html(str(table))[0]
    dfs.append(df)
df = pd.concat(dfs)
# Save the DataFrame as an Excel file
df.to_excel('tables.xlsx', index=False)

之后我尝试用循环自动判断Next按钮是否可点击,从而持续爬取数据,但循环始终无法终止,先后尝试了两种写法:

第一种循环代码

while True:
    soup = BeautifulSoup(driver.page_source, 'html.parser')
    table = soup.find('table', {"class": "tablepress tablepress-id-2 dataTable no-footer"})
    tables.append(table)

    # Check if the "Next" button is clickable
    
    next_button = WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.LINK_TEXT, 'Next')))
    if not next_button.is_enabled():
        break

    # Click the "Next" button to go to the next page
    next_button.click()

# Close the web driver
driver.quit()

# Convert each table to a DataFrame and concatenate them into a single DataFrame
dfs = []
for table in tables:
    df = pd.read_html(str(table))[0]
    dfs.append(df)
df = pd.concat(dfs)

# Save the DataFrame as an Excel file
df.to_excel('tables.xlsx', index=False) 

第二种try/catch循环代码

while True:
    # Extract the table
    soup = BeautifulSoup(driver.page_source, 'html.parser')
    table = soup.find('table', {"class": "tablepress tablepress-id-2 dataTable no-footer"})
    tables.append(table)

    # Check if there is a next button
    try:
        next_button = WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.LINK_TEXT, 'Next')))
    except TimeoutException:
        break

    # Click the next button
    next_button.click()

driver.quit()



# Convert each table to a DataFrame and concatenate them into a single DataFrame
dfs = []
for table in tables:
    df = pd.read_html(str(table))[0]
    dfs.append(df)
df = pd.concat(dfs)
# Save the DataFrame as an Excel file
df.to_excel('tables.xlsx', index=False)

两种写法均无效,循环无限执行。同时我想了解:如何用Selenium检测网页上的按钮是否无法点击?

附Next按钮两种状态的HTML:

  • 可点击状态:
<a class="paginate_button next" aria-controls="tablepress-3" data-dt-idx="1" tabindex="0" id="tablepress-3_next">Next</a>
  • 不可点击状态:
<a class="paginate_button next disabled" aria-controls="tablepress-3" data-dt-idx="1" tabindex="-1" id="tablepress-3_next">Next</a>

问题原因分析

  1. 第一种写法失效原因:EC.element_to_be_clickable会等待元素可见且原生enabled,但这里的按钮禁用是通过添加disabled类实现的,并非原生HTML的disabled属性,所以next_button.is_enabled()始终返回True,无法触发break。
  2. 第二种写法失效原因:即使按钮不可点击,元素本身依然存在于DOM中,WebDriverWait不会触发TimeoutException,因此循环永远不会终止。

正确解决方案

核心是通过检查按钮的class属性是否包含disabled来判断是否不可点击,代码如下:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import pandas as pd

driver = webdriver.Chrome()
driver.get('website_url')

tables = []

while True:
    # 提取当前页表格
    soup = BeautifulSoup(driver.page_source, 'html.parser')
    table = soup.find('table', {"class": "tablepress tablepress-id-2 dataTable no-footer"})
    tables.append(table)
    
    # 等待Next按钮加载完成(不管是否可点击)
    next_button = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.ID, 'tablepress-3_next'))
    )
    
    # 检查按钮是否包含disabled类
    if 'disabled' in next_button.get_attribute('class'):
        break
    
    # 点击Next按钮
    next_button.click()
    # 等待页面刷新完成,避免提取重复表格
    WebDriverWait(driver, 10).until(
        EC.staleness_of(table)
    )

driver.quit()

# 合并表格并保存到Excel
dfs = []
for table in tables:
    df = pd.read_html(str(table))[0]
    dfs.append(df)
df = pd.concat(dfs)
df.to_excel('tables.xlsx', index=False)

如何检测按钮是否无法点击

根据提供的HTML结构,有两种可靠方式:

  1. 检查class是否包含disabled:
next_button = driver.find_element(By.ID, 'tablepress-3_next')
is_disabled = 'disabled' in next_button.get_attribute('class')
  1. 检查tabindex属性是否为-1:
next_button = driver.find_element(By.ID, 'tablepress-3_next')
is_disabled = next_button.get_attribute('tabindex') == '-1'

注意:不要使用is_enabled()方法,因为它仅检查原生HTML的disabled属性,对通过CSS类实现的禁用无效。

内容的提问来源于stack exchange,提问作者olaniyan oluwasegun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 20:05:01