使用Selenium点击按钮翻页爬取多页面报错解决方案
Selenium多页面爬取实现方案(针对目标站点修复)
原代码报错核心原因
- 逻辑顺序错误:仅在第一页采集了律师列表链接,且把翻页逻辑写在了详情页遍历循环内,详情页不存在分页按钮,直接触发元素找不到的报错
- 使用废弃API:
find_elements_by_xpath、find_element_by_xpath在Selenium 4.x版本已被移除,运行会直接抛方法不存在的异常 - 翻页后未重新采集列表内容:仅能拿到第一页的律师链接,后续页面数据全部漏采
- 依赖固定等待时长:页面加载速度波动时,很容易出现元素还没加载就开始定位的报错
可直接运行的修正代码
import time from selenium import webdriver from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from selenium.webdriver.support.wait import WebDriverWait from selenium.common.exceptions import NoSuchElementException, TimeoutException from webdriver_manager.chrome import ChromeDriverManager options = webdriver.ChromeOptions() # 需要无头运行则取消下面这行的注释 # options.add_argument("--headless") options.add_argument("--no-sandbox") options.add_argument("--disable-gpu") options.add_argument("--window-size=1920x1080") options.add_argument("--disable-extensions") chrome_driver = webdriver.Chrome( service=Service(ChromeDriverManager().install()), options=options ) productlink=[] def supplyvan_scraper(): with chrome_driver as driver: wait = WebDriverWait(driver, 15) URL = 'https://www.ifep.ro/justice/lawyers/lawyerspanel.aspx' driver.get(URL) # 第一阶段:循环翻页采集所有列表页的律师详情链接 while True: # 等待当前页列表加载完成,采集当前页所有详情链接 wait.until(EC.presence_of_all_elements_located((By.XPATH, "//div[@class='list-group']//a"))) links = driver.find_elements(By.XPATH, "//div[@class='list-group']//a") for link in links: link_href = link.get_attribute("href") if link_href and link_href.startswith("https://www.ifep.ro/") and link_href not in productlink: productlink.append(link_href) # 检查是否存在可点击的下一页按钮 try: next_btn = wait.until(EC.element_to_be_clickable((By.ID, "MainContent_PagerTop_NextBtn"))) # 记录当前页第一个元素,用于判断翻页是否完成 first_list_item = driver.find_element(By.XPATH, "//div[@class='list-group']//a") next_btn.click() # 等待旧列表元素失效,确认新页面加载完成 wait.until(EC.staleness_of(first_list_item)) time.sleep(0.5) except (NoSuchElementException, TimeoutException): # 找不到下一页按钮说明到最后一页,退出翻页循环 break print(f"共采集到{len(productlink)}条律师详情链接") # 第二阶段:遍历所有详情链接,采集目标字段 for product_url in productlink: try: driver.get(product_url) wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, '#HeadingContent_lblTitle'))) title = driver.find_element(By.CSS_SELECTOR, '#HeadingContent_lblTitle').text d1 = driver.find_element(By.XPATH, "//div[@class='col-md-10']//p[1]").text.strip() d2 = driver.find_element(By.XPATH, "//div[@class='col-md-10']//p[2]").text.strip() d3 = driver.find_element(By.XPATH, "//div[@class='col-md-10']//p[3]//span").text.strip() d4 = driver.find_element(By.XPATH, "//div[@class='col-md-10']//p[4]").text.strip() print(title, d1, d2, d3, d4) time.sleep(0.3) except Exception as e: print(f"链接{product_url}采集失败:{str(e)}") continue driver.quit() supplyvan_scraper()
关键实现逻辑说明
- 流程拆分:先完成所有列表页的翻页和链接采集,再统一进入详情页抓字段,避免流程混乱找不到元素
- 自动识别分页:通过判断下一页按钮是否可点击决定是否继续翻页,不需要手动指定总页数,可自动爬完所有分页
- 显式等待替代固定sleep:通过元素状态判断页面加载进度,既减少无效等待时间,也避免加载慢导致的定位失败
- 去重+异常捕获:链接自动去重避免重复采集,单个详情页采集失败不会中断整个程序
- 适配新版Selenium API:所有元素定位方法都兼容4.x版本,不会出现方法废弃的报错
内容的提问来源于stack exchange,提问作者Amen Aziz
相关产品推荐
相关产品推荐

