求助:如何用Python的BS4正确爬取YouTube标题?
解决YouTube标题爬取失效问题
问题根源
- 动态内容渲染:YouTube首页的视频列表依赖JavaScript动态加载,
requests.get()仅能获取初始静态HTML,无法捕获JS渲染后的标题元素。 - 反爬拦截:未携带浏览器请求头的直接请求会被YouTube识别为非合法访问,返回的页面内容不完整或被拦截。
- 元素选择器过时:YouTube前端元素的class属性会频繁更新,原代码依赖的
yt-simple-endpoint focus-on-expand style-scope ytd-rich-grid-media类名已失效。
解决方案1:优化请求头并调整元素选择器
通过模拟浏览器请求头绕过基础反爬,同时使用当前有效的元素选择器尝试抓取:
import requests from bs4 import BeautifulSoup def get_youtube_titles(): url = 'https://www.youtube.com/' # 模拟Chrome浏览器请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } try: response = requests.get(url, headers=headers) response.raise_for_status() soup = BeautifulSoup(response.text, 'html.parser') # 定位当前有效的标题元素选择器 title_elements = soup.find_all('yt-formatted-string', class_='style-scope ytd-rich-grid-media') if not title_elements: print("未找到标题元素,可能页面仍为动态渲染或选择器已更新") for title_element in title_elements: title = title_element.text.strip() if title: print(title) except requests.exceptions.RequestException as e: print('网络请求错误:', e) get_youtube_titles()
解决方案2:使用Selenium获取动态渲染内容
若上述方法仍失效,可通过Selenium模拟浏览器加载完整页面,确保获取JS渲染后的内容:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup def get_youtube_titles(): url = 'https://www.youtube.com/' # 配置Chrome浏览器无头模式 options = webdriver.ChromeOptions() options.add_argument('--headless=new') options.add_argument('--disable-gpu') options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36') try: driver = webdriver.Chrome(options=options) driver.get(url) # 等待标题元素加载完成(最长等待10秒) WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, 'yt-formatted-string.style-scope.ytd-rich-grid-media')) ) # 获取渲染后的页面源码 page_source = driver.page_source soup = BeautifulSoup(page_source, 'html.parser') # 提取并打印标题 title_elements = soup.find_all('yt-formatted-string', class_='style-scope ytd-rich-grid-media') for title_element in title_elements: title = title_element.text.strip() if title: print(title) except Exception as e: print('爬取错误:', e) finally: driver.quit() get_youtube_titles()
注意事项
- 执行Selenium方案需先安装依赖:
pip install selenium - 需下载对应浏览器版本的驱动(如ChromeDriver),并确保其路径配置在系统PATH中或直接指定驱动路径。
内容的提问来源于stack exchange,提问作者Peter
相关产品推荐
相关产品推荐

