如何使用Selenium Python抓取ul li标签数据及实现多页爬取?
问题与解决方案
问题概述
需要抓取https://scan.multichain.org/#/tokens页面中红色标记区域的所有数据,数据共分布在33个页面,但无法通过点击分页按钮获取第2页及后续页面的数据,自行编写的Selenium代码无法正常运行。
原代码问题分析
原代码存在以下关键问题:
- 语法错误:
wait = WebDriverWait(driver, 10)行首存在多余空格,会触发Python语法报错 - 元素获取错误:使用
find_element而非find_elements,只能返回单个元素,后续循环操作会直接报错 - 分页控件定位错误:
li.primary并非分页按钮的正确选择器,无法定位到分页控件 - 缺少分页循环逻辑:没有实现点击下一页并重复抓取的完整流程
修正后的完整代码
from selenium import webdriver from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import NoSuchElementException, ElementClickInterceptedException # 初始化Chrome驱动 driver = webdriver.Chrome(r'C:\SeleniumDrivers\chromedriver.exe') wait = WebDriverWait(driver, 10) try: driver.get("https://scan.multichain.org/#/tokens") # 循环遍历所有分页 while True: # 等待当前页面数据区域加载完成(需根据红色标记区域结构调整选择器) wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "table tbody tr"))) # 抓取当前页红色标记区域数据(示例为抓取token链接,需根据实际需求修改) target_elements = driver.find_elements(By.CSS_SELECTOR, "table tbody tr td:nth-child(2) a") for elem in target_elements: print(elem.get_attribute('href')) # 尝试点击下一页按钮 try: # 定位可点击的下一页按钮(需根据页面实际结构调整选择器) next_page_btn = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, "li.page-item.next:not(.disabled) a"))) next_page_btn.click() # 等待页面切换完成,确保新页面元素加载 wait.until(EC.staleness_of(target_elements[0])) except (NoSuchElementException, ElementClickInterceptedException): # 无下一页按钮或按钮不可点击时,终止循环 break finally: # 关闭浏览器 driver.quit()
关键注意事项
- 元素定位适配:需通过浏览器开发者工具(F12)查看红色标记区域的实际DOM结构,修改代码中数据抓取和分页按钮的CSS选择器,确保定位准确。
- 页面等待逻辑:使用
staleness_of等待旧页面元素失效,保证新页面完全加载后再执行抓取操作。 - 异常处理:捕获分页按钮不存在、无法点击的异常,避免程序意外崩溃。
内容的提问来源于stack exchange,提问作者Đỗ Văn Thắng
相关产品推荐
相关产品推荐

