You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium Python爬取指定YouTube搜索页的Shorts数据

需求说明

需要在YouTube搜索页面(链接:https://www.youtube.com/results?search_query=birds+fly),用Selenium Python完成以下爬取任务:

  • YouTube Shorts标题
  • 播放量
  • 订阅人数
  • Shorts数量(手动统计为6个)

我尝试的代码

!pip install selenium==4.1.5
!pip install webdriver_manager
import selenium
from selenium import webdriver 
from selenium.webdriver.chrome.service import Service #chromedriver 
from webdriver_manager.chrome import ChromeDriverManager #chromedriver 
from selenium.webdriver.common.by import By 
service = Service(executable_path=ChromeDriverManager().install()) 
driver = webdriver.Chrome(service=service)

driver.get('https://www.youtube.com/results?search_query=birds+fly')

# get the elements
elements = driver.find_elements(By.CSS_SELECTOR, '.style-scope.yt-horizontal-list-renderer[id="items"]')

# print the count of elements
print(len(elements))

正确实现代码

YouTube页面为动态加载,直接查找元素易因页面未渲染完全失败,需使用显式等待定位目标元素,以下是完整可运行代码:

# 首次运行执行安装依赖
!pip install selenium==4.1.5
!pip install webdriver_manager

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 初始化Chrome驱动
service = Service(executable_path=ChromeDriverManager().install())
driver = webdriver.Chrome(service=service)
driver.get('https://www.youtube.com/results?search_query=birds+fly')

try:
    # 显式等待Shorts横向列表加载完成,最长等待10秒
    shorts_container = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CSS_SELECTOR, 'ytd-horizontal-card-list-renderer[is-shorts="true"]'))
    )
    
    # 获取所有Shorts条目
    shorts_items = shorts_container.find_elements(By.CSS_SELECTOR, 'ytd-reel-video-renderer')
    print(f"Shorts数量: {len(shorts_items)}")
    
    # 遍历每个Shorts提取信息
    for index, short in enumerate(shorts_items, 1):
        # 提取标题
        title = short.find_element(By.CSS_SELECTOR, '#video-title').text
        # 提取播放量
        views = short.find_element(By.CSS_SELECTOR, '#metadata-line span:nth-child(1)').text
        # 点击进入Shorts页面获取订阅人数(搜索页不直接显示该数据)
        short.click()
        
        # 等待频道订阅信息加载
        channel_sub = WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.CSS_SELECTOR, '#owner-sub-count'))
        ).text
        
        print(f"\n第{index}条Shorts信息:")
        print(f"标题: {title}")
        print(f"播放量: {views}")
        print(f"订阅人数: {channel_sub}")
        
        # 返回搜索页继续处理下一条
        driver.back()
        # 返回后重新定位Shorts容器,确保后续元素可正常获取
        shorts_container = WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.CSS_SELECTOR, 'ytd-horizontal-card-list-renderer[is-shorts="true"]'))
        )
        shorts_items = shorts_container.find_elements(By.CSS_SELECTOR, 'ytd-reel-video-renderer')
        
finally:
    # 任务完成后关闭浏览器
    driver.quit()

关键说明

  1. 显式等待:通过WebDriverWait等待元素加载,避免页面未渲染完成导致的元素定位失败问题
  2. 精准选择器:用ytd-horizontal-card-list-renderer[is-shorts="true"]定位Shorts专属容器,再通过ytd-reel-video-renderer获取单个Shorts条目
  3. 订阅人数获取:搜索页的Shorts条目不直接显示订阅数据,需进入详情页提取频道的订阅信息
  4. 页面切换处理:从详情页返回搜索页后,需重新定位Shorts容器,保证后续元素定位正常

内容的提问来源于stack exchange,提问作者ann25

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 16:35:04