Python Selenium爬虫触发StaleElementReferenceException问题求助
问题与解决:Selenium遍历元素触发StaleElementReferenceException
问题原因
StaleElementReferenceException 触发的核心原因是:你通过WebDriverWait获取到的元素列表medias,其中的元素引用已经和当前页面的DOM树脱离。可能是页面在遍历过程中发生了动态更新(比如异步加载、局部刷新),导致原本定位到的元素已经被移除或重新渲染,原有的元素引用失效。
你的代码仅在获取元素列表阶段添加了异常捕获,但在遍历元素调用get_attribute的阶段没有处理该异常,所以触发报错。
解决方案
- 在遍历元素的循环中添加
StaleElementReferenceException捕获逻辑,触发时关闭浏览器驱动、返回None并退出函数。 - 补充浏览器驱动的资源释放逻辑,避免未关闭的进程残留。
修改后的代码
from selenium import webdriver from webdriver_manager.chrome import ChromeDriverManager from selenium.webdriver.common.by import By from selenium.webdriver.chrome.service import Service from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common import exceptions def get_all_search_details(URL): SEARCH_RESULTS = {} driver = None try: options = Options() options.headless = True options.add_argument("--remote-debugging-port=9222") options.add_argument("--no-sandbox") options.add_argument("--disable-gpu") options.add_argument("--disable-dev-shm-usage") options.add_argument("--disable-extensions") options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option('useAutomationExtension', False) driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options) driver.get(URL) print(f"Scraping {driver.current_url}") medias = WebDriverWait(driver,timeout=5,).until(EC.presence_of_all_elements_located((By.CLASS_NAME, 'result-row'))) for media_idx, media_elem in enumerate(medias): try: outer_html = media_elem.get_attribute('outerHTML') result = scrap_newspaper(outer_html) # 外部处理函数 SEARCH_RESULTS[f"result_{media_idx}"] = result except exceptions.StaleElementReferenceException as e: print(f">> {type(e).__name__}: {e.args}") return None return SEARCH_RESULTS except exceptions.StaleElementReferenceException as e: print(f">> {type(e).__name__}: {e.args}") return None except exceptions.NoSuchElementException as e: print(f">> {type(e).__name__}: {e.args}") return None except exceptions.TimeoutException as e: print(f">> {type(e).__name__}: {e.args}") return None except exceptions.WebDriverException as e: print(f">> {type(e).__name__}: {e.args}") return None except exceptions.SessionNotCreatedException as e: print(f">> {type(e).__name__}: {e.args}") return None except Exception as e: print(f">> {type(e).__name__} line {e.__traceback__.tb_lineno} of {__file__}: {e.args}") return None except: print(f">> General Exception: {URL}") return None finally: # 确保无论是否异常,都关闭浏览器驱动 if driver: driver.quit() if __name__ == '__main__': in_url = "https://digi.kansalliskirjasto.fi/clippings?query=isokyr%C3%B6&categoryId=12&orderBy=RELEVANCE&page=3&resultMode=THUMB" my_res = get_all_search_details(in_url)
关键修改点
- 将
driver初始化移到外层try块,添加finally块确保驱动被关闭,避免资源泄漏。 - 在遍历元素的循环内部新增
StaleElementReferenceException捕获,触发时直接返回None退出函数。 - 所有异常分支统一返回
None,符合需求。
内容的提问来源于stack exchange,提问作者farid
相关产品推荐
相关产品推荐

