Youtube Selenium爬虫异常:无报错却停止,未筛选出订阅量<500的用户
问题排查与修复方案
核心问题分析
你的代码出现无征兆停止、未筛选出目标用户的原因主要集中在元素定位失效、隐性异常未捕获、等待机制不可靠这几个方面,具体如下:
- 旧版Selenium语法兼容问题:
find_element_by_name这类旧API在Selenium 4+版本中已被废弃,隐性的兼容性问题会导致程序突然终止 - Stale元素异常未处理:点击视频返回搜索页后,原有的
video_links列表会因DOM刷新失效,触发StaleElementReferenceException,但你的异常捕获列表里没有这个类型,导致程序直接停止 - 硬编码sleep不可靠:固定的
time.sleep(2)无法适配不同网络速度,页面未加载完成就操作会导致元素找不到,直接被pass跳过 - 订阅量处理逻辑不全:未处理“隐藏订阅数”“0 subscribers”这类特殊情况,导致符合条件的用户被遗漏
- 元素选择器失效:YouTube页面元素的ID和结构会频繁更新,原选择器可能无法定位到正确元素
修复后的完整代码
import time from selenium import webdriver from selenium.webdriver.common.keys import Keys from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import ( NoSuchElementException, TimeoutException, StaleElementReferenceException ) from webdriver_manager.chrome import ChromeDriverManager # 接收用户关键词 keyword = input("Enter a keyword to search for: ") # 初始化Chrome浏览器,添加启动参数优化爬取 options = webdriver.ChromeOptions() options.add_argument("--start-maximized") options.add_argument("--disable-notifications") driver = webdriver.Chrome(ChromeDriverManager().install(), options=options) wait = WebDriverWait(driver, 10) try: # 打开YouTube首页,处理Cookie弹窗 driver.get("https://www.youtube.com") try: accept_cookie_btn = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[@aria-label='Accept all']"))) accept_cookie_btn.click() except (TimeoutException, NoSuchElementException): pass # 执行搜索操作 search_box = wait.until(EC.element_to_be_clickable((By.NAME, "search_query"))) search_box.send_keys(keyword) search_box.send_keys(Keys.RETURN) # 等待搜索结果加载完成 wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "a#video-title"))) # 获取视频链接的href,避免Stale元素问题 video_urls = [link.get_attribute("href") for link in driver.find_elements(By.CSS_SELECTOR, "a#video-title")] less_than_500_subs = set() # 用集合自动去重 for url in video_urls: driver.get(url) try: # 等待订阅数元素加载 sub_count_elem = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "yt-formatted-string#owner-sub-count"))) sub_text = sub_count_elem.text.strip() # 处理订阅数文本 if "hidden" in sub_text.lower() or "no subscribers" in sub_text.lower(): subscribers = 0 else: sub_num = sub_text.replace("subscribers", "").replace(",", "").strip() if "K" in sub_num: subscribers = float(sub_num.replace("K", "")) * 1000 elif "M" in sub_num: subscribers = float(sub_num.replace("M", "")) * 1000000 else: subscribers = float(sub_num) if sub_num else 0 # 判断订阅量并获取用户名 if subscribers < 500: channel_name_elem = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "a#channel-name"))) username = channel_name_elem.get_attribute("title") less_than_500_subs.add(username) except (NoSuchElementException, TimeoutException, StaleElementReferenceException): continue time.sleep(1) # 短延迟避免请求过于频繁 # 输出结果 print("订阅量少于500的创作者用户名:") for username in less_than_500_subs: print(username) finally: driver.quit()
关键修改说明
- 替换旧版Selenium API:全部改用
find_element(By.XXX, "value")的新版语法,避免兼容性问题 - 添加Stale元素异常捕获:将
StaleElementReferenceException加入异常捕获列表,防止DOM刷新导致程序终止 - 改用显式等待替代硬编码sleep:用
WebDriverWait等待元素可点击/存在,确保页面加载完成后再操作,提升稳定性 - 处理特殊订阅量情况:新增对“隐藏订阅数”“无订阅”的逻辑处理,避免遗漏符合条件的用户
- 用集合存储用户名:自动去重,比列表判断更高效
- 获取视频URL而非直接点击元素:避免循环中因DOM变化导致的元素失效问题
- 添加浏览器启动参数:最大化窗口、禁用通知,减少页面加载异常
内容的提问来源于stack exchange,提问作者Aleksandra25
相关产品推荐
相关产品推荐

