You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:如何用Python的BS4正确爬取YouTube标题?

解决YouTube标题爬取失效问题

问题根源

  • 动态内容渲染:YouTube首页的视频列表依赖JavaScript动态加载,requests.get()仅能获取初始静态HTML,无法捕获JS渲染后的标题元素。
  • 反爬拦截:未携带浏览器请求头的直接请求会被YouTube识别为非合法访问,返回的页面内容不完整或被拦截。
  • 元素选择器过时:YouTube前端元素的class属性会频繁更新,原代码依赖的yt-simple-endpoint focus-on-expand style-scope ytd-rich-grid-media类名已失效。

解决方案1:优化请求头并调整元素选择器

通过模拟浏览器请求头绕过基础反爬,同时使用当前有效的元素选择器尝试抓取:

import requests
from bs4 import BeautifulSoup

def get_youtube_titles():
    url = 'https://www.youtube.com/'
    # 模拟Chrome浏览器请求头
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
    }

    try:
        response = requests.get(url, headers=headers)
        response.raise_for_status()
        soup = BeautifulSoup(response.text, 'html.parser')
        
        # 定位当前有效的标题元素选择器
        title_elements = soup.find_all('yt-formatted-string', class_='style-scope ytd-rich-grid-media')
        
        if not title_elements:
            print("未找到标题元素,可能页面仍为动态渲染或选择器已更新")
        
        for title_element in title_elements:
            title = title_element.text.strip()
            if title:
                print(title)
    
    except requests.exceptions.RequestException as e:
        print('网络请求错误:', e)

get_youtube_titles()

解决方案2:使用Selenium获取动态渲染内容

若上述方法仍失效,可通过Selenium模拟浏览器加载完整页面,确保获取JS渲染后的内容:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

def get_youtube_titles():
    url = 'https://www.youtube.com/'
    
    # 配置Chrome浏览器无头模式
    options = webdriver.ChromeOptions()
    options.add_argument('--headless=new')
    options.add_argument('--disable-gpu')
    options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')
    
    try:
        driver = webdriver.Chrome(options=options)
        driver.get(url)
        
        # 等待标题元素加载完成(最长等待10秒)
        WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.CSS_SELECTOR, 'yt-formatted-string.style-scope.ytd-rich-grid-media'))
        )
        
        # 获取渲染后的页面源码
        page_source = driver.page_source
        soup = BeautifulSoup(page_source, 'html.parser')
        
        # 提取并打印标题
        title_elements = soup.find_all('yt-formatted-string', class_='style-scope ytd-rich-grid-media')
        for title_element in title_elements:
            title = title_element.text.strip()
            if title:
                print(title)
    
    except Exception as e:
        print('爬取错误:', e)
    finally:
        driver.quit()

get_youtube_titles()

注意事项

  • 执行Selenium方案需先安装依赖:pip install selenium
  • 需下载对应浏览器版本的驱动(如ChromeDriver),并确保其路径配置在系统PATH中或直接指定驱动路径。

内容的提问来源于stack exchange,提问作者Peter

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 12:47:44