You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium爬取YouTube视频标题遇StaleElementReferenceException错误求助

解决Selenium爬取YouTube标题时的StaleElementReferenceException错误

错误原因

StaleElementReferenceException的核心问题是:你提前获取的元素列表,在后续循环时已经和当前页面文档失去关联——比如YouTube滚动加载新内容、页面局部刷新时,旧的元素节点会被浏览器销毁,导致之前保存的元素引用失效。

解决方案

1. 优化元素定位逻辑,避免持有无效元素引用

不要一次性缓存所有元素对象再循环,而是优先提取文本,或在循环内保证元素有效性;同时缩小选择器范围,避免匹配无关元素:

from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

title_list = []
# 显式等待视频标题元素加载完成,用更精准的选择器定位视频标题
wait = WebDriverWait(driver, 10)
video_titles = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "yt-formatted-string#video-title")))

for title in video_titles:
    try:
        title_text = title.text.strip()
        if title_text:  # 过滤空文本
            title_list.append(title_text)  # 用append添加完整标题,原代码extend会拆分字符
    except Exception:
        continue  # 遇到失效元素直接跳过

print(title_list)

2. 处理滚动加载场景的动态更新

如果需要爬取更多视频(滚动加载的情况),每次滚动后重新定位元素,避免旧引用失效:

from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

title_list = []
last_height = driver.execute_script("return document.documentElement.scrollHeight")

while True:
    # 每次滚动后重新等待并定位元素
    wait = WebDriverWait(driver, 10)
    video_titles = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "yt-formatted-string#video-title")))
    
    # 提取文本并去重
    for title in video_titles:
        try:
            text = title.text.strip()
            if text and text not in title_list:
                title_list.append(text)
        except Exception:
            continue
    
    # 滚动加载更多内容
    driver.execute_script("window.scrollTo(0, document.documentElement.scrollHeight);")
    time.sleep(2)  # 也可用显式等待元素变化替代固定等待
    
    # 判断是否到达页面底部
    new_height = driver.execute_script("return document.documentElement.scrollHeight")
    if new_height == last_height:
        break
    last_height = new_height

print(title_list)

关键注意点

  • 原代码使用list.extend(text_title)是错误的,extend会将字符串拆分为单个字符添加到列表,应改为append添加完整标题文本。
  • 避免用list作为变量名,这是Python内置类型,会覆盖内置函数。
  • YouTube视频标题的精准选择器是yt-formatted-string#video-title,原选择器yt-formatted-string会匹配页面上大量无关文本(如频道名、描述),导致结果混乱。

内容的提问来源于stack exchange,提问作者Abhishek Kakadiya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 19:22:11