You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决Selenium循环搜索时的StaleElementReferenceException异常

问题:Google搜索自动化循环时出现StaleElementReferenceException错误

我用Python结合Selenium开发自动化工具,用来在Google搜索特定链接,判断是否出现在顶部搜索结果里,标记已索引/未索引。首次搜索正常,但循环数据库里的下一个链接、定位搜索框时出错。

相关代码

parser/parser.py

from selenium import webdriver
from selenium.webdriver.support.wait import WebDriverWait
from selenium.webdriver.common.by import By
from selenium.common.exceptions import NoSuchElementException
from indexer.models import Link
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.keys import Keys

def main():
    indexed = 0
    not_indexed = 0

    db_links = Link.objects.all()

    options = webdriver.FirefoxOptions()
    options.add_argument("-headless")
    driver = webdriver.Firefox(options=options)
    wait = WebDriverWait(driver, timeout=10)

    visited_google = False

    for link in db_links:
        try:
            if not visited_google:
                driver.get(f"https://www.google.com")
                visited_google = True
            wait.until(EC.element_to_be_clickable((By.XPATH, "//textarea[@id='APjFqb']")))
            textarea = driver.find_element(By.XPATH, value="//textarea[@id='APjFqb']")
            textarea.send_keys(link.link)
            textarea.send_keys(Keys.ENTER)
            wait.until(EC.presence_of_element_located(locator=(By.XPATH, "//div[@id='search']")))
            search = driver.find_element(By.XPATH, value="//div[@id='search']")
            anchor = search.find_element(By.XPATH, value=f"//a[@href='{link}']")
            if anchor:
                indexed += 1
            else:
                not_indexed += 1
        except NoSuchElementException:
            not_indexed += 1
    driver.quit()
    print(indexed)
    print(not_indexed)

views.py

from django.shortcuts import HttpResponse
from django.http import Http404
from .parser.parser import main
def parse(req):
    if req.method == "POST":
        if not req.POST["links"]:
            return Http404("Bad Request")
        else:
            email = req.POST["email"]
            links = req.POST["links"]
            uid = math.floor(time.time())
            main()
            return HttpResponse("parsing")
    else:
        return Http404("Error")

错误信息

StaleElementReferenceException at /parse/
Message: The element reference of <textarea id="APjFqb" class="gLFyf" name="q" type="search"> is stale; either its node document is not the active document, or it is no longer connected to the DOM; For documentation on this error, please visit: https://www.selenium.dev/documentation/webdriver/troubleshooting/errors#stale-element-reference-exception
Stacktrace:
RemoteError@chrome://remote/content/shared/RemoteError.sys.mjs:8:8
WebDriverError@chrome://remote/content/shared/webdriver/Errors.sys.mjs:180:5
StaleElementReferenceError@chrome://remote/content/shared/webdriver/Errors.sys.mjs:461:5
element.resolveElement@chrome://remote/content/marionette/element.sys.mjs:509:11
deserializeJSON@chrome://remote/content/marionette/json.sys.mjs:206:33
cloneObject@chrome://remote/content/marionette/json.sys.mjs:56:24
deserializeJSON@chrome://remote/content/marionette/json.sys.mjs:213:16
json.deserialize@chrome://remote/content/marionette/json.sys.mjs:217:10
receiveMessage@chrome://remote/content/marionette/actors/MarionetteCommandsChild.sys.mjs:85:30

问题原因及解决方案

原因

StaleElementReferenceException 触发的核心原因是:页面跳转或刷新后,之前定位的DOM元素引用已经失效。你的代码第一次访问Google首页并搜索后,页面跳转到搜索结果页;第二次循环时,直接在当前结果页尝试定位搜索框,但此时的搜索框和初始首页的搜索框不是同一个DOM元素,或者页面渲染已更新,导致之前的元素引用过期。同时,即使结果页有搜索框,直接调用send_keys会追加内容,也会引发问题。

解决方案

  1. 每次循环重新访问Google首页,确保搜索框是最新的DOM元素,彻底避免元素引用失效问题
  2. 每次输入前清空搜索框,防止之前的搜索内容残留
  3. 优化异常处理逻辑,避免单个链接处理失败导致整个循环中断

修改后的代码

parser/parser.py

from selenium import webdriver
from selenium.webdriver.support.wait import WebDriverWait
from selenium.webdriver.common.by import By
from selenium.common.exceptions import NoSuchElementException
from indexer.models import Link
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.keys import Keys

def main():
    indexed = 0
    not_indexed = 0

    db_links = Link.objects.all()

    options = webdriver.FirefoxOptions()
    options.add_argument("-headless")
    driver = webdriver.Firefox(options=options)
    wait = WebDriverWait(driver, timeout=10)

    for link in db_links:
        try:
            # 每次循环重新访问Google首页,确保搜索框为最新DOM元素
            driver.get("https://www.google.com")
            # 等待搜索框可点击,直接返回元素对象
            textarea = wait.until(EC.element_to_be_clickable((By.XPATH, "//textarea[@id='APjFqb']")))
            # 清空搜索框再输入目标链接
            textarea.clear()
            textarea.send_keys(link.link)
            textarea.send_keys(Keys.ENTER)
            # 等待搜索结果区域加载完成
            wait.until(EC.presence_of_element_located((By.XPATH, "//div[@id='search']")))
            # 直接在搜索结果区域内查找目标链接
            driver.find_element(By.XPATH, f"//div[@id='search']//a[@href='{link.link}']")
            indexed += 1
        except NoSuchElementException:
            not_indexed += 1
        except Exception as e:
            # 捕获其他异常,避免循环中断
            not_indexed += 1
            print(f"处理链接 {link.link} 时出错: {str(e)}")
    driver.quit()
    print(f"已索引: {indexed}")
    print(f"未索引: {not_indexed}")

views.py(修复缺失的模块导入)

import math
import time
from django.shortcuts import HttpResponse
from django.http import Http404
from .parser.parser import main

def parse(req):
    if req.method == "POST":
        # 使用get方法避免KeyError
        if not req.POST.get("links"):
            return Http404("Bad Request")
        else:
            email = req.POST.get("email")
            links = req.POST.get("links")
            uid = math.floor(time.time())
            main()
            return HttpResponse("parsing")
    else:
        return Http404("Error")

内容的提问来源于stack exchange,提问作者Unnikrishnan G

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 16:24:51