Python编写YouTube频道视频标题爬虫无结果问题排查
问题出在哪?
1. YouTube视频列表是动态渲染的
你用requests.get()只能拿到页面的静态HTML,但YouTube的视频列表是页面加载完成后靠JavaScript动态拉取数据渲染出来的,所以你拿到的响应里根本没有那些视频元素,自然搜不到结果。
2. CSS选择器写法错误
你的选择器.yt-simple-endpoint style-scope ytd-grid-video-renderer格式不对,style-scope和ytd-grid-video-renderer都是类名,多个类名应该用点连接,正确写法是.yt-simple-endpoint.style-scope.ytd-grid-video-renderer——不过这个问题在动态内容面前,改对了也没用,因为静态页面里根本不存在这些元素。
解决办法
方法一:用Selenium模拟浏览器爬取(快速上手)
Selenium能模拟真实浏览器打开页面,等待JavaScript加载完视频列表后再获取内容:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup url = 'https://www.youtube.com/c/DanTDM/videos' # 启动Chrome浏览器(需提前安装对应版本的ChromeDriver并配置环境变量) driver = webdriver.Chrome() driver.get(url) # 等待10秒,直到视频列表元素加载完成 wait = WebDriverWait(driver, 10) wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, ".yt-simple-endpoint.style-scope.ytd-grid-video-renderer"))) # 获取渲染后的页面源码,关闭浏览器 page_source = driver.page_source driver.quit() # 解析页面并提取视频标题 soup = BeautifulSoup(page_source, 'html.parser') video_items = soup.select(".yt-simple-endpoint.style-scope.ytd-grid-video-renderer") for item in video_items: print(item.get('title'))
方法二:使用YouTube官方API(稳定合规)
如果打算长期爬取,推荐用官方API,避免被反爬机制拦截,步骤如下:
- 前往Google Cloud平台创建项目,启用YouTube Data API v3
- 生成个人API密钥
- 调用接口获取视频数据(DanTDM的频道ID为
UCX6OQ3DkcsbYNE6H8uQQuVA)
示例代码(先安装依赖:pip install google-api-python-client):
from googleapiclient.discovery import build API_KEY = '你的API密钥' CHANNEL_ID = 'UCX6OQ3DkcsbYNE6H8uQQuVA' # 初始化API服务 youtube = build('youtube', 'v3', developerKey=API_KEY) # 获取频道的上传列表ID channel_info = youtube.channels().list(part='contentDetails', id=CHANNEL_ID).execute() upload_playlist_id = channel_info['items'][0]['contentDetails']['relatedPlaylists']['uploads'] # 拉取上传列表中的视频数据 video_list = youtube.playlistItems().list( part='snippet', playlistId=upload_playlist_id, maxResults=50 # 单次请求最多返回50条,需更多内容可分页获取 ).execute() # 打印所有视频标题 for video in video_list['items']: print(video['snippet']['title'])
注意事项
- 使用Selenium时,ChromeDriver版本必须与本地Chrome浏览器版本匹配,否则会出现兼容性问题
- 爬取YouTube时不要过于频繁,建议添加合理延时,模拟人类操作节奏,避免触发反爬机制
- YouTube Data API有免费配额,个人使用完全足够,注意不要超出配额限制
内容的提问来源于stack exchange,提问作者blackscratch22
相关产品推荐
相关产品推荐

