Selenium爬取time-mark.com遇StaleElementReferenceException求助
问题:Selenium爬取网站时触发StaleElementReferenceException错误
作为网页爬虫新手,编写了以下Selenium代码,意图爬取https://time-mark.com/网站,提取符合条件的href值并添加到product_model_list列表中,但运行时触发StaleElementReferenceException错误:
from selenium import webdriver from selenium.common.exceptions import StaleElementReferenceException from selenium.webdriver.chrome.service import Service import time from selenium.common import exceptions ser_obj = Service("C:/Users/sunda/Downloads/Drivers/Chrome.exe") driver = webdriver.Edge(executable_path="C:/Users/sunda/Downloads/Drivers/msedgedriver") product_model_list = [] driver.get("https://time-mark.com/") # Extract the value of the href attribute a_tags = driver.find_elements_by_tag_name('a') for a_tag in a_tags: product_name_url = a_tag.get_attribute('href') if "product-category" in product_name_url: print(product_name_url) driver.implicitly_wait(5) driver.get(product_name_url) a_tag_product_name = driver.find_elements_by_tag_name('a') for product in a_tag_product_name: product_model = product.get_attribute('href') # try: if "3-phase-monitors" in product_model and "page" not in product_model and "shop" in product_model: product_model_list.append(product_model) print("product name = ",product_model) print(product_model_list)
报错信息:
File "C:\Users\sunda\anaconda3\lib\site-packages\selenium\webdriver\remote\errorhandler.py", line 242, in check_response raise exception_class(message, screen, stacktrace) selenium.common.exceptions.StaleElementReferenceException: Message: stale element reference: stale element not found (Session info: MicrosoftEdge=122.0.2365.92)
错误原因
当调用driver.get(product_name_url)跳转到新页面后,原页面的DOM已被销毁,之前通过driver.find_elements_by_tag_name('a')获取的a_tags列表中的元素全部变成"过时元素",循环到下一个a_tag时,Selenium无法定位这些已失效的元素,因此抛出该异常。
修复方案
核心思路是先收集所有需要访问的分类页面URL,再遍历URL列表进行后续爬取,避免在遍历元素过程中跳转页面导致元素失效。同时优化等待策略和元素定位逻辑:
- 先从首页提取所有含
product-category的href,存入独立列表 - 遍历URL列表逐个访问分类页面
- 用显式等待替代隐式等待,提升爬取稳定性
- 增加空值判断,避免
href为None时触发报错 - 用CSS选择器直接筛选符合条件的链接,减少无效遍历
修复后的代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 初始化浏览器与显式等待 driver = webdriver.Edge(executable_path="C:/Users/sunda/Downloads/Drivers/msedgedriver") wait = WebDriverWait(driver, 10) product_model_list = [] try: # 访问首页并收集所有分类页面URL driver.get("https://time-mark.com/") category_urls = [] # 等待所有a标签加载完成 a_tags = wait.until(EC.presence_of_all_elements_located((By.TAG_NAME, 'a'))) for a_tag in a_tags: href = a_tag.get_attribute('href') if href and "product-category" in href: category_urls.append(href) # 去重避免重复访问 category_urls = list(set(category_urls)) # 遍历分类页面提取目标产品链接 for url in category_urls: print(f"正在访问分类页面:{url}") driver.get(url) # 直接用CSS选择器筛选符合条件的产品链接 product_links = wait.until(EC.presence_of_all_elements_located( (By.CSS_SELECTOR, 'a[href*="3-phase-monitors"][href*="shop"]:not([href*="page"])') )) for link in product_links: product_href = link.get_attribute('href') if product_href and product_href not in product_model_list: product_model_list.append(product_href) print(f"找到产品链接:{product_href}") finally: # 确保浏览器关闭 driver.quit() print("\n最终提取的产品链接列表:") print(product_model_list)
内容的提问来源于stack exchange,提问作者pscodes
相关产品推荐
相关产品推荐

