使用Selenium爬取Web应用页面遇分页问题求助
问题解决:Selenium爬取Web应用分页数据失败
核心问题分析
- 分页元素
ag-paging-number可能对应多个DOM节点(比如当前页码、可点击的其他页码),直接用find_element_by_class_name只会定位到第一个匹配元素,大概率是当前页,点击自然无效。 - 固定
time.sleep(10)效率低且不可靠,页面加载完成时间不稳定,容易出现元素未就绪就执行点击的情况。 - 原代码仅在循环结束后保存最后一页的源码,没有实现每页数据的存储。
解决方案步骤
1. 精准定位可点击的分页元素
通过浏览器开发者工具查看分页区域的DOM结构,通常ag-paging-number中,非当前页的元素会带有可点击标识(比如tabindex="0"或特定CSS状态)。推荐直接定位专门的下一页按钮(如果存在),比定位页码更可靠:
# 定位下一页按钮(类名通常为ag-paging-button-next) next_btn = driver.find_element_by_class_name("ag-paging-button-next")
如果没有单独的下一页按钮,可定位当前页的下一个页码元素:
current_page = driver.find_element_by_css_selector(".ag-paging-number.ag-paging-number-active") next_page = current_page.find_element_by_xpath("./following-sibling::div[contains(@class, 'ag-paging-number')]")
2. 使用显式等待替代固定休眠
利用Selenium的WebDriverWait等待元素可点击,避免因页面加载慢导致的元素未就绪问题:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By # 等待下一页元素可点击,最长等待15秒 wait = WebDriverWait(driver, 15) next_btn = wait.until(EC.element_to_be_clickable((By.CLASS_NAME, "ag-paging-button-next"))) next_btn.click()
3. 修正循环逻辑,实现每页数据存储
将保存页面源码的代码放到循环内部,确保每页数据都被保存:
for page in range(1, 1497): # 遍历1到1496页 # 保存当前页源码 with codecs.open(f'page_{page}.htm', 'w', 'utf-8') as out: out.write(driver.page_source) # 最后一页跳过点击下一页 if page < 1496: try: next_btn = wait.until(EC.element_to_be_clickable((By.CLASS_NAME, "ag-paging-button-next"))) next_btn.click() # 等待页面数据加载完成(可根据页面实际状态调整) wait.until(EC.staleness_of(current_page)) # 等待当前页元素失效,说明页面已刷新 except Exception as e: print(f"第{page}页点击下一页失败: {e}") break
4. 目录设置建议
为避免文件混乱,建议按页码批次创建存储目录,方便后续管理:
import os # 创建基础存储目录 base_dir = "climate_projections_data" os.makedirs(base_dir, exist_ok=True) # 按每100页分一个子目录 for page in range(1, 1497): sub_dir = os.path.join(base_dir, f"batch_{(page-1)//100 + 1}") os.makedirs(sub_dir, exist_ok=True) # 保存到对应子目录 file_path = os.path.join(sub_dir, f'page_{page}.htm') with codecs.open(file_path, 'w', 'utf-8') as out: out.write(driver.page_source)
完整修改后的代码
from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By import codecs import os # 配置浏览器选项 options = Options() options.add_argument("--window-size=1920,1200") # 后续切换无头模式可添加:options.add_argument("--headless=new") driver = webdriver.Chrome(options=options) wait = WebDriverWait(driver, 15) url = "https://app.climatevaluation.com/apps/projections/table/index.html" driver.get(url) # 创建存储目录 base_dir = "climate_projections_data" os.makedirs(base_dir, exist_ok=True) try: for page in range(1, 1497): # 等待表格数据加载完成(可根据页面实际加载标识调整) wait.until(EC.presence_of_element_located((By.CLASS_NAME, "ag-row"))) # 按批次创建子目录 sub_dir = os.path.join(base_dir, f"batch_{(page-1)//100 + 1}") os.makedirs(sub_dir, exist_ok=True) file_path = os.path.join(sub_dir, f"page_{page}.htm") # 保存当前页源码 with codecs.open(file_path, 'w', 'utf-8') as out: out.write(driver.page_source) print(f"已保存第{page}页数据") # 最后一页无需点击下一页 if page == 1496: break # 点击下一页 try: next_btn = wait.until(EC.element_to_be_clickable((By.CLASS_NAME, "ag-paging-button-next"))) next_btn.click() # 短等待确保页面切换完成 wait.until(EC.text_to_be_present_in_element((By.CLASS_NAME, "ag-paging-number-active"), str(page+1))) except Exception as e: print(f"第{page}页点击下一页失败,终止循环: {str(e)}") break finally: driver.quit()
内容的提问来源于stack exchange,提问作者Matthew Oldham
相关产品推荐
相关产品推荐

