Selenium分页表格爬取循环无法终止及按钮不可点击检测问题
分页表格爬取:解决Next按钮循环无法终止的问题
我需要从分页网页提取表格,最初采用手动点击Next按钮的方式爬取,代码如下:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup from selenium.common.exceptions import TimeoutException import requests import pandas as pd # Set up the web driver driver = webdriver.Chrome() driver.get('website_url') tables =[] soup = BeautifulSoup(driver.page_source, 'html.parser') table = soup.find('table', {"class": "tablepress tablepress-id-2 dataTable no-footer"}) tables.append(table) # Click the next button and extract the second table next_button = WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.LINK_TEXT, 'Next'))) next_button.click() soup = BeautifulSoup(driver.page_source, 'html.parser') table = soup.find('table', {"class": "tablepress tablepress-id-2 dataTable no-footer"}) tables.append(table) # Click the next button and extract the second table next_button = WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.LINK_TEXT, 'Next'))) next_button.click() soup = BeautifulSoup(driver.page_source, 'html.parser') table = soup.find('table', {"class": "tablepress tablepress-id-2 dataTable no-footer"}) tables.append(table) driver.quit() # Convert each table to a DataFrame and concatenate them into a single DataFrame dfs = [] for table in tables: df = pd.read_html(str(table))[0] dfs.append(df) df = pd.concat(dfs) # Save the DataFrame as an Excel file df.to_excel('tables.xlsx', index=False)
之后我尝试用循环自动判断Next按钮是否可点击,从而持续爬取数据,但循环始终无法终止,先后尝试了两种写法:
第一种循环代码
while True: soup = BeautifulSoup(driver.page_source, 'html.parser') table = soup.find('table', {"class": "tablepress tablepress-id-2 dataTable no-footer"}) tables.append(table) # Check if the "Next" button is clickable next_button = WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.LINK_TEXT, 'Next'))) if not next_button.is_enabled(): break # Click the "Next" button to go to the next page next_button.click() # Close the web driver driver.quit() # Convert each table to a DataFrame and concatenate them into a single DataFrame dfs = [] for table in tables: df = pd.read_html(str(table))[0] dfs.append(df) df = pd.concat(dfs) # Save the DataFrame as an Excel file df.to_excel('tables.xlsx', index=False)
第二种try/catch循环代码
while True: # Extract the table soup = BeautifulSoup(driver.page_source, 'html.parser') table = soup.find('table', {"class": "tablepress tablepress-id-2 dataTable no-footer"}) tables.append(table) # Check if there is a next button try: next_button = WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.LINK_TEXT, 'Next'))) except TimeoutException: break # Click the next button next_button.click() driver.quit() # Convert each table to a DataFrame and concatenate them into a single DataFrame dfs = [] for table in tables: df = pd.read_html(str(table))[0] dfs.append(df) df = pd.concat(dfs) # Save the DataFrame as an Excel file df.to_excel('tables.xlsx', index=False)
两种写法均无效,循环无限执行。同时我想了解:如何用Selenium检测网页上的按钮是否无法点击?
附Next按钮两种状态的HTML:
- 可点击状态:
<a class="paginate_button next" aria-controls="tablepress-3" data-dt-idx="1" tabindex="0" id="tablepress-3_next">Next</a>
- 不可点击状态:
<a class="paginate_button next disabled" aria-controls="tablepress-3" data-dt-idx="1" tabindex="-1" id="tablepress-3_next">Next</a>
问题原因分析
- 第一种写法失效原因:
EC.element_to_be_clickable会等待元素可见且原生enabled,但这里的按钮禁用是通过添加disabled类实现的,并非原生HTML的disabled属性,所以next_button.is_enabled()始终返回True,无法触发break。 - 第二种写法失效原因:即使按钮不可点击,元素本身依然存在于DOM中,
WebDriverWait不会触发TimeoutException,因此循环永远不会终止。
正确解决方案
核心是通过检查按钮的class属性是否包含disabled来判断是否不可点击,代码如下:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import pandas as pd driver = webdriver.Chrome() driver.get('website_url') tables = [] while True: # 提取当前页表格 soup = BeautifulSoup(driver.page_source, 'html.parser') table = soup.find('table', {"class": "tablepress tablepress-id-2 dataTable no-footer"}) tables.append(table) # 等待Next按钮加载完成(不管是否可点击) next_button = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, 'tablepress-3_next')) ) # 检查按钮是否包含disabled类 if 'disabled' in next_button.get_attribute('class'): break # 点击Next按钮 next_button.click() # 等待页面刷新完成,避免提取重复表格 WebDriverWait(driver, 10).until( EC.staleness_of(table) ) driver.quit() # 合并表格并保存到Excel dfs = [] for table in tables: df = pd.read_html(str(table))[0] dfs.append(df) df = pd.concat(dfs) df.to_excel('tables.xlsx', index=False)
如何检测按钮是否无法点击
根据提供的HTML结构,有两种可靠方式:
- 检查class是否包含disabled:
next_button = driver.find_element(By.ID, 'tablepress-3_next') is_disabled = 'disabled' in next_button.get_attribute('class')
- 检查tabindex属性是否为-1:
next_button = driver.find_element(By.ID, 'tablepress-3_next') is_disabled = next_button.get_attribute('tabindex') == '-1'
注意:不要使用is_enabled()方法,因为它仅检查原生HTML的disabled属性,对通过CSS类实现的禁用无效。
内容的提问来源于stack exchange,提问作者olaniyan oluwasegun
相关产品推荐
相关产品推荐

