You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium爬取time-mark.com遇StaleElementReferenceException求助

问题:Selenium爬取网站时触发StaleElementReferenceException错误

作为网页爬虫新手,编写了以下Selenium代码,意图爬取https://time-mark.com/网站,提取符合条件的href值并添加到product_model_list列表中,但运行时触发StaleElementReferenceException错误:

from selenium import webdriver
from selenium.common.exceptions import StaleElementReferenceException
from selenium.webdriver.chrome.service import Service
import time
from selenium.common import exceptions
ser_obj = Service("C:/Users/sunda/Downloads/Drivers/Chrome.exe")

driver = webdriver.Edge(executable_path="C:/Users/sunda/Downloads/Drivers/msedgedriver")
product_model_list = []
driver.get("https://time-mark.com/")
# Extract the value of the href attribute
a_tags = driver.find_elements_by_tag_name('a')

for a_tag in a_tags:

    product_name_url = a_tag.get_attribute('href')
    if "product-category" in product_name_url:
        print(product_name_url)
        driver.implicitly_wait(5)
        driver.get(product_name_url)
        a_tag_product_name = driver.find_elements_by_tag_name('a')
        for product in a_tag_product_name:
            product_model = product.get_attribute('href')
            # try:
            if "3-phase-monitors" in product_model and "page" not in product_model and "shop" in product_model:
                product_model_list.append(product_model)
                print("product name = ",product_model)

print(product_model_list)

报错信息:

File "C:\Users\sunda\anaconda3\lib\site-packages\selenium\webdriver\remote\errorhandler.py", line 242, in check_response
raise exception_class(message, screen, stacktrace)
selenium.common.exceptions.StaleElementReferenceException: Message: stale element reference: stale element not found
(Session info: MicrosoftEdge=122.0.2365.92)

错误原因

当调用driver.get(product_name_url)跳转到新页面后,原页面的DOM已被销毁,之前通过driver.find_elements_by_tag_name('a')获取的a_tags列表中的元素全部变成"过时元素",循环到下一个a_tag时,Selenium无法定位这些已失效的元素,因此抛出该异常。


修复方案

核心思路是先收集所有需要访问的分类页面URL,再遍历URL列表进行后续爬取,避免在遍历元素过程中跳转页面导致元素失效。同时优化等待策略和元素定位逻辑:

  • 先从首页提取所有含product-category的href,存入独立列表
  • 遍历URL列表逐个访问分类页面
  • 用显式等待替代隐式等待,提升爬取稳定性
  • 增加空值判断,避免href为None时触发报错
  • 用CSS选择器直接筛选符合条件的链接,减少无效遍历

修复后的代码
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 初始化浏览器与显式等待
driver = webdriver.Edge(executable_path="C:/Users/sunda/Downloads/Drivers/msedgedriver")
wait = WebDriverWait(driver, 10)
product_model_list = []

try:
    # 访问首页并收集所有分类页面URL
    driver.get("https://time-mark.com/")
    category_urls = []
    
    # 等待所有a标签加载完成
    a_tags = wait.until(EC.presence_of_all_elements_located((By.TAG_NAME, 'a')))
    for a_tag in a_tags:
        href = a_tag.get_attribute('href')
        if href and "product-category" in href:
            category_urls.append(href)
    
    # 去重避免重复访问
    category_urls = list(set(category_urls))
    
    # 遍历分类页面提取目标产品链接
    for url in category_urls:
        print(f"正在访问分类页面:{url}")
        driver.get(url)
        
        # 直接用CSS选择器筛选符合条件的产品链接
        product_links = wait.until(EC.presence_of_all_elements_located(
            (By.CSS_SELECTOR, 'a[href*="3-phase-monitors"][href*="shop"]:not([href*="page"])')
        ))
        
        for link in product_links:
            product_href = link.get_attribute('href')
            if product_href and product_href not in product_model_list:
                product_model_list.append(product_href)
                print(f"找到产品链接:{product_href}")

finally:
    # 确保浏览器关闭
    driver.quit()

print("\n最终提取的产品链接列表:")
print(product_model_list)

内容的提问来源于stack exchange,提问作者pscodes

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 16:33:22