You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup抓取YouTube频道的视频标题及相关数据

YouTube频道视频数据爬虫实现方案

首先明确核心问题:YouTube页面内容为JavaScript动态渲染,直接用requests请求静态页面返回的HTML骨架中没有实际的视频数据,这是你直接用BeautifulSoup调试失败的主要原因。

一、获取指定频道所有视频的标题与访问链接

推荐搭配selenium模拟浏览器渲染页面,再用BeautifulSoup解析内容,适配你现有的技术栈,实现步骤如下:

  • 先安装依赖库:
pip install selenium beautifulsoup4 webdriver-manager
  • 示例代码:
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
import time
from bs4 import BeautifulSoup

# 初始化浏览器
driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()))
# 目标频道视频页地址
url = 'https://www.youtube.com/c/AlexTheAnalyst/videos'
driver.get(url)
# 模拟滚动加载全部视频,可根据视频数量调整滚动次数
scroll_pause_time = 2
last_height = driver.execute_script("return document.documentElement.scrollHeight")
while True:
    # 滚动到页面底部
    driver.execute_script("window.scrollTo(0, document.documentElement.scrollHeight);")
    time.sleep(scroll_pause_time)
    # 计算新的滚动高度,和上次高度对比判断是否加载完所有内容
    new_height = driver.execute_script("return document.documentElement.scrollHeight")
    if new_height == last_height:
        break
    last_height = new_height

# 拿到渲染后的页面源码交给BeautifulSoup解析
soup = BeautifulSoup(driver.page_source, 'html.parser')
video_items = soup.find_all('ytd-grid-video-renderer', class_='style-scope ytd-grid-renderer')
video_list = []
for item in video_items:
    title_tag = item.find('a', id='video-title')
    video_title = title_tag.get('title')
    video_url = 'https://www.youtube.com' + title_tag.get('href')
    video_list.append({'title': video_title, 'url': video_url})
# 打印测试
print(f"共获取到{len(video_list)}条视频")
for v in video_list[:5]:
    print(v)

二、单视频详细数据爬取实现

拿到视频链接后,同样可以基于渲染后的页面解析播放量、点赞数、发布日期等数据,示例代码如下:

def get_video_detail(video_url):
    driver.get(video_url)
    # 等待页面加载完成
    time.sleep(3)
    soup = BeautifulSoup(driver.page_source, 'html.parser')
    # 解析结构化数据,比直接找元素更稳定
    json_data = soup.find('script', type='application/ld+json').string
    import json
    data = json.loads(json_data)[0]
    # 提取所需字段
    detail = {
        'title': data.get('name'),
        'publish_date': data.get('uploadDate'),
        'view_count': data.get('interactionCount'),
        'duration': data.get('duration'),
        'description': data.get('description')
    }
    # 单独解析点赞数(结构化数据中无此字段,需从页面元素提取)
    like_tag = soup.find('yt-formatted-string', class_='style-scope ytd-toggle-button-renderer style-text')
    if like_tag:
        detail['like_count'] = like_tag.get('aria-label').replace('个赞','').strip()
    # 评论数提取
    comment_tag = soup.find('yt-formatted-string', class_='count-text style-scope ytd-comments-header-renderer')
    if comment_tag:
        detail['comment_count'] = comment_tag.text.strip()
    return detail

# 测试单视频抓取
test_video = video_list[0]['url']
print(get_video_detail(test_video))

# 全部爬取完成后关闭浏览器
driver.quit()

注意事项

  • 元素选择器可能随YouTube页面迭代更新,若解析失败可自行检查对应元素的标签、属性后调整代码
  • 频繁发起请求会触发YouTube反爬机制,出现人机验证,建议请求间隙添加随机延迟,控制爬取频率
  • 若不需要模拟浏览器,也可直接抓频道页加载视频的XHR异步接口,返回的JSON格式数据解析效率更高,无需处理HTML结构

内容的提问来源于stack exchange,提问作者Donatus Prince

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 12:18:03