Selenium元素在GitHub Actions测试中失效但本地正常的问题排查
问题描述
为开发更优质应用,需爬取官方日程数据(申请API遭拒)。本地测试无元素失效问题,但通过GitHub Actions执行爬取时会出现StaleElementReferenceException,且爬取15-20分钟后也会触发该元素失效异常。
已尝试延长等待时间、更换Chrome/Firefox浏览器,均无效。
补充:这段爬取代码会在主程序中多次运行,推测可能是切换课程时页面DOM修改导致元素失效。
爬取代码
if course_id not in output: output[course_id] = {} output[course_id][name] = {} course = course_dict[course_id] WebDriverWait(driver, 10).until(EC.presence_of_all_elements_located((By.XPATH,'//table[@class="as-content"]/tbody/tr'))) driver.implicitly_wait(0.5) tables = driver.find_elements(By.XPATH,'//table[@class="as-content"]/tbody/tr') color = getColor("") for i in range(1,len(tables),2): course_name = tables[i].find_elements(By.XPATH,'./td/table/tbody/tr/td/span')[0].text print(course_name) if course_name in course: course_name, color_ = course[course_name] color = getColor(color_) if course_name not in output[course_id][name]: output[course_id][name][course_name] = [] list_cursus = tables[i+1].find_elements(By.XPATH,'./td/div/table/tbody/tr') for cursus in list_cursus: info = cursus.find_elements(By.XPATH,'./td') driver.implicitly_wait(0.1) date = info[0].find_element(By.XPATH,'./div').get_attribute("innerHTML").replace(" ", " ") _time = info[1].get_attribute("innerHTML").replace(" ", " ") row_date = dateparser.parse(date + ' ' + _time[3:8]) if row_date != None: start = row_date.strftime("%Y-%m-%dT%H:%M:%S") end = dateparser.parse(date + ' ' + _time[10:17]).strftime("%Y-%m-%dT%H:%M:%S") teacher = info[3].get_attribute("innerHTML").replace(" ", " ") room = info[4].get_attribute("innerHTML").replace(" ", " ") title = course_name + '\n' + teacher + '\n' + room output[course_id][name][course_name].append({'title':title,'start':start,'end':end, 'color': color})
异常信息
Traceback (most recent call last): File "/home/runner/work/NovaPlanning/NovaPlanning/scraping/scraper.py", line 130, in <module> get_information(driver,"BAB1 INFO", "INFO") File "/home/runner/work/NovaPlanning/NovaPlanning/scraping/scraper.py", line 101, in get_information info = cursus.find_elements(By.XPATH,'./td') File "/opt/hostedtoolcache/Python/3.9.19/x64/lib/python3.9/site-packages/selenium/webdriver/remote/webelement.py", line 439, in find_elements return self._execute(Command.FIND_CHILD_ELEMENTS, {"using": by, "value": value})["value"] File "/opt/hostedtoolcache/Python/3.9.19/x64/lib/python3.9/site-packages/selenium/webdriver/remote/webelement.py", line 395, in _execute return self._parent.execute(command, params) File "/opt/hostedtoolcache/Python/3.9.19/x64/lib/python3.9/site-packages/selenium/webdriver/remote/webdriver.py", line 354, in execute self.error_handler.check_response(response) File "/opt/hostedtoolcache/Python/3.9.19/x64/lib/python3.9/site-packages/selenium/webdriver/remote/errorhandler.py", line 229, in check_response raise exception_class(message, screen, stacktrace) selenium.common.exceptions.StaleElementReferenceException: Message: The element with the reference 88f3af23-5a4b-4963-94ce-27c4de81849d is stale; either its node document is not the active document, or it is no longer connected to the DOM; For documentation on this error, please visit: https://www.selenium.dev/documentation/webdriver/troubleshooting/errors#stale-element-reference-exception Stacktrace: RemoteError@chrome://remote/content/shared/RemoteError.sys.mjs:8:8 WebDriverError@chrome://remote/content/shared/webdriver/Errors.sys.mjs:193:5 StaleElementReferenceError@chrome://remote/content/shared/webdriver/Errors.sys.mjs:725:5 getKnownElement@chrome://remote/content/marionette/json.sys.mjs:401:11 deserializeJSON@chrome://remote/content/marionette/json.sys.mjs:259:20 cloneObject@chrome://remote/content/marionette/json.sys.mjs:59:24 deserializeJSON@chrome://remote/content/marionette/json.sys.mjs:289:16 cloneObject@chrome://remote/content/marionette/json.sys.mjs:59:24 deserializeJSON@chrome://remote/content/marionette/json.sys.mjs:289:16 json.deserialize@chrome://remote/content/marionette/json.sys.mjs:293:10 receiveMessage@chrome://remote/content/marionette/actors/MarionetteCommandsChild.sys.mjs:73:30
核心原因
该异常是因为提前缓存的元素引用(比如tables、list_cursus中的元素)在页面DOM更新后,已与当前页面脱离关联——切换课程时页面会重新渲染部分内容,旧的元素引用自然失效。
解决办法
1. 避免缓存元素引用,每次操作重新定位
不要一次性把所有元素存入变量循环,而是每次需要时从根节点重新定位。比如不要提前获取tables再循环,而是每次循环时通过索引重新查找对应行:
修改后的关键代码片段:
# 先获取总行数,不缓存所有行 row_count = len(WebDriverWait(driver, 10).until(EC.presence_of_all_elements_located((By.XPATH,'//table[@class="as-content"]/tbody/tr')))) for i in range(1, row_count, 2): # 每次循环重新定位第i行(XPath索引从1开始) course_row = WebDriverWait(driver, 5).until(EC.presence_of_element_located((By.XPATH,f'//table[@class="as-content"]/tbody/tr[{i+1}]'))) course_name = course_row.find_element(By.XPATH,'./td/table/tbody/tr/td/span').text print(course_name) # ... 其他逻辑 ... # 重新定位对应的课程详情行 cursus_container = WebDriverWait(driver, 5).until(EC.presence_of_element_located((By.XPATH,f'//table[@class="as-content"]/tbody/tr[{i+2}]'))) list_cursus = cursus_container.find_elements(By.XPATH,'./td/div/table/tbody/tr') for cursus in list_cursus: # 确保元素可见后再操作 WebDriverWait(driver, 3).until(EC.visibility_of(cursus)) info = cursus.find_elements(By.XPATH,'./td') # ... 后续爬取逻辑 ...
2. 移除隐式等待,统一用显式等待
混合使用WebDriverWait显式等待和driver.implicitly_wait隐式等待会导致等待机制冲突,反而降低稳定性。建议删除所有driver.implicitly_wait()调用,全部用WebDriverWait等待元素出现/可见。
3. 增加异常捕获与重试机制
对容易失效的元素操作,添加重试逻辑,捕获StaleElementReferenceException后重试几次:
from selenium.common.exceptions import StaleElementReferenceException def retry_stale_element(func, retries=3): for _ in range(retries): try: return func() except StaleElementReferenceException: continue raise Exception("重试多次后仍无法定位元素") # 使用示例:获取课程名称 course_name = retry_stale_element(lambda: course_row.find_element(By.XPATH,'./td/table/tbody/tr/td/span').text)
4. 优化页面切换后的等待逻辑
每次切换课程后,不要只等待单个元素,而是等待页面关键DOM结构稳定——比如等待课程表格的tbody元素可见,且行数量不再变化后再开始爬取。
本地正常但GitHub Actions异常的原因
GitHub Actions运行环境的网络、页面加载速度与本地不同,页面渲染时机更不稳定,导致本地缓存的元素还未失效时,CI环境中页面已完成更新,触发了元素失效异常。
内容的提问来源于stack exchange,提问作者Florian Wauters

