使用Selenium爬取页面触发StaleElementReferenceException问题咨询
报错含义解释
StaleElementReferenceException 表示你之前定位到的元素已经和当前页面DOM树解绑,该引用已经失效,无法执行后续操作。通常在点击按钮后页面局部刷新、DOM重绘时容易触发,你这个场景里就是点击加载更多按钮后,页面更新列表内容时旧按钮被移除,下一轮查找元素时刚好命中旧元素的失效引用就会抛出该错误。
修复后的代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import TimeoutException, StaleElementReferenceException import time options = webdriver.ChromeOptions() options.add_argument('--no-sandbox') options.add_argument('--disable-dev-shm-usage') site = 'https://candidat.pole-emploi.fr/offres/emploi/horticulteur/s1m1' wd = webdriver.Chrome("C:\Program Files (x86)\chromedriver.exe", options=options) wd.get(site) wait = WebDriverWait(wd, 15) # 处理cookie弹窗 time.sleep(3) try: cookie_btn = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, '#description .tc-open-privacy-center'))) cookie_btn.click() except TimeoutException: # 没找到cookie弹窗就跳过 pass # 循环点击加载更多 while True: try: # 等待按钮可点击 more_button = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, 'a[title="AFFICHER LES 20 OFFRES SUIVANTES"]'))) more_button.click() # 点击后等待页面加载完成,避免马上进入下一轮循环拿到失效引用 time.sleep(2) # 等待新的列表项加载完成 wait.until(EC.presence_of_element_located((By.CLASS_NAME, 'result'))) except (TimeoutException, StaleElementReferenceException): # 超时说明没有更多按钮,元素过期则重试或退出 break time.sleep(5) print(wd.page_source) print("Complete") time.sleep(3) wd.quit()
主要修改点
- 新增了
StaleElementReferenceException异常捕获,出现该错误时直接退出循环或者重试 - 把
visibility_of_element_located替换为element_to_be_clickable,确保按钮可以点击的时候再操作 - 用更稳定的CSS选择器替换了LINK_TEXT定位,减少匹配失误
- 增加了点击后的加载等待逻辑,避免DOM还没更新完成就进入下一轮查找
- 优化了cookie弹窗的处理逻辑,增加异常捕获避免弹窗不存在时直接报错
内容的提问来源于stack exchange,提问作者user11278201
相关产品推荐
相关产品推荐

