You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

解决Google Colab中Selenium爬ICML2024动态页面的StaleElementReferenceException

解决Selenium爬取OpenReview论文时的StaleElementReferenceException问题

问题场景

要从OpenReview的ICML 2024会议口头报告标签页提取1-6页的论文标题、作者和PDF链接,第一页能成功提取,但爬取到第6页时持续触发StaleElementReferenceException错误。页面为JavaScript动态加载,原代码无法解决该问题,原代码如下:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

options = Options()
options.add_argument('--headless')
options.add_argument('--no-sandbox')
options.add_argument('--disable-dev-shm-usage')

driver = webdriver.Chrome(options=options)
driver.get("https://openreview.net/group?id=ICML.cc/2024/Conference#tab-accept-oral")

try:
    for i in range(1, 7):  # repeat from 1 page to 6 page
        WebDriverWait(driver, 10).until(
            EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div#accept-oral div.note"))
        )
        notes = driver.find_elements(By.CSS_SELECTOR, "div#accept-oral div.note")
        for note in notes:
            title = note.find_element(By.CSS_SELECTOR, "h4 a").text
            authors = note.find_element(By.CSS_SELECTOR, "div.note-authors").text
            pdf_link = note.find_element(By.CSS_SELECTOR, "a.pdf-link").get_attribute('href')
            print("Title:", title)
            print("Authors:", authors)
            print("PDF Link:", pdf_link)

        # click next page
        if i < 6:
            next_page_button = WebDriverWait(driver, 10).until(
                EC.element_to_be_clickable((By.CSS_SELECTOR, f"ul.pagination li:nth-child({i+2}) a"))
            )
            driver.execute_script("arguments[0].click();", next_page_button)
finally:
    driver.quit()

错误原因分析

  1. 元素引用失效:翻页后页面DOM会重新渲染,原代码中notes列表是基于旧DOM的元素引用,后续循环中访问这些已失效的元素就会触发StaleElementReferenceException。
  2. 分页按钮定位逻辑错误:使用li:nth-child({i+2})定位下一页按钮,当页码超过页面默认显示范围(比如页面只显示前几个页码和Next按钮),该选择器会定位到错误元素,导致翻页失败或异常。

修复后的代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

options = Options()
options.add_argument('--headless')
options.add_argument('--no-sandbox')
options.add_argument('--disable-dev-shm-usage')

driver = webdriver.Chrome(options=options)
driver.get("https://openreview.net/group?id=ICML.cc/2024/Conference#tab-accept-oral")

try:
    target_pages = range(1, 7)
    for page_num in target_pages:
        # 等待当前页论文元素完全可见(确保可交互)
        WebDriverWait(driver, 15).until(
            EC.visibility_of_all_elements_located((By.CSS_SELECTOR, "div#accept-oral div.note"))
        )
        
        # 每次循环重新获取当前页的论文元素,避免旧引用失效
        notes = driver.find_elements(By.CSS_SELECTOR, "div#accept-oral div.note")
        for note in notes:
            title = note.find_element(By.CSS_SELECTOR, "h4 a").text
            authors = note.find_element(By.CSS_SELECTOR, "div.note-authors").text
            pdf_link = note.find_element(By.CSS_SELECTOR, "a.pdf-link").get_attribute('href')
            print(f"第{page_num}页 - 标题:", title)
            print("作者:", authors)
            print("PDF链接:", pdf_link)
            print("---")
        
        # 非最后一页时执行翻页
        if page_num < max(target_pages):
            # 定位"Next"按钮(不受页码显示范围影响)
            next_btn = WebDriverWait(driver, 10).until(
                EC.element_to_be_clickable((By.CSS_SELECTOR, "ul.pagination li.next a"))
            )
            next_btn.click()
            
            # 等待页面切换完成:确认目标页码处于激活状态
            WebDriverWait(driver, 10).until(
                EC.visibility_of_element_located((By.CSS_SELECTOR, f"ul.pagination li.active a[href*='page={page_num+1}']"))
            )
finally:
    driver.quit()

关键修改说明

  • 重新获取元素:将notes的获取逻辑放在每页循环内,确保每次都基于最新DOM获取元素引用,彻底避免 stale 错误。
  • 可靠的分页定位:改用li.next a定位下一页按钮,不管页面显示多少页码,都能准确找到Next按钮,避免原选择器的逻辑漏洞。
  • 增强等待条件:用visibility_of_all_elements_located替代presence_of_all_elements_located,确保元素可见且可交互;翻页后等待目标页码变为激活状态,保证页面完全切换完成。
  • 提升容错性:延长等待时间至15秒,适配OpenReview可能的加载延迟。

内容的提问来源于stack exchange,提问作者S H

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 12:53:20