使用Selenium爬取YouTube视频标题遇StaleElementReferenceException错误求助
解决Selenium爬取YouTube标题时的StaleElementReferenceException错误
错误原因
StaleElementReferenceException的核心问题是:你提前获取的元素列表,在后续循环时已经和当前页面文档失去关联——比如YouTube滚动加载新内容、页面局部刷新时,旧的元素节点会被浏览器销毁,导致之前保存的元素引用失效。
解决方案
1. 优化元素定位逻辑,避免持有无效元素引用
不要一次性缓存所有元素对象再循环,而是优先提取文本,或在循环内保证元素有效性;同时缩小选择器范围,避免匹配无关元素:
from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC title_list = [] # 显式等待视频标题元素加载完成,用更精准的选择器定位视频标题 wait = WebDriverWait(driver, 10) video_titles = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "yt-formatted-string#video-title"))) for title in video_titles: try: title_text = title.text.strip() if title_text: # 过滤空文本 title_list.append(title_text) # 用append添加完整标题,原代码extend会拆分字符 except Exception: continue # 遇到失效元素直接跳过 print(title_list)
2. 处理滚动加载场景的动态更新
如果需要爬取更多视频(滚动加载的情况),每次滚动后重新定位元素,避免旧引用失效:
from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time title_list = [] last_height = driver.execute_script("return document.documentElement.scrollHeight") while True: # 每次滚动后重新等待并定位元素 wait = WebDriverWait(driver, 10) video_titles = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "yt-formatted-string#video-title"))) # 提取文本并去重 for title in video_titles: try: text = title.text.strip() if text and text not in title_list: title_list.append(text) except Exception: continue # 滚动加载更多内容 driver.execute_script("window.scrollTo(0, document.documentElement.scrollHeight);") time.sleep(2) # 也可用显式等待元素变化替代固定等待 # 判断是否到达页面底部 new_height = driver.execute_script("return document.documentElement.scrollHeight") if new_height == last_height: break last_height = new_height print(title_list)
关键注意点
- 原代码使用
list.extend(text_title)是错误的,extend会将字符串拆分为单个字符添加到列表,应改为append添加完整标题文本。 - 避免用
list作为变量名,这是Python内置类型,会覆盖内置函数。 - YouTube视频标题的精准选择器是
yt-formatted-string#video-title,原选择器yt-formatted-string会匹配页面上大量无关文本(如频道名、描述),导致结果混乱。
内容的提问来源于stack exchange,提问作者Abhishek Kakadiya
相关产品推荐
相关产品推荐

