使用BeautifulSoup爬取YouTube频道视频信息,find_all返回空列表求解决
解决YouTube频道视频爬取返回空列表的问题
问题原因
YouTube的视频列表是通过JavaScript动态渲染加载的,直接用requests.get()获取的只是静态HTML骨架,里面根本没有你要查找的yt-simple-endpoint inline-block style-scope ytd-thumbnail这类元素,所以find_all()自然返回空列表。
解决方案1:用Selenium模拟浏览器加载动态内容
Selenium可以模拟真实浏览器的行为,等待页面JS执行完成后再获取渲染后的HTML内容,代码示例:
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.options import Options import time url = "https://www.youtube.com/@PW-Foundation/videos" # 配置无头Chrome(不显示浏览器窗口) chrome_options = Options() chrome_options.add_argument("--headless=new") driver = webdriver.Chrome(options=chrome_options) driver.get(url) # 等待页面加载,根据网络情况调整时间 time.sleep(3) # 获取渲染后的页面源码 soup = BeautifulSoup(driver.page_source, 'html.parser') driver.quit() # 提取前五条视频信息 video_items = soup.find_all('ytd-grid-video-renderer', limit=5) for idx, item in enumerate(video_items, 1): # 视频URL video_url = "https://www.youtube.com" + item.find('a', id='thumbnail')['href'] # 缩略图URL thumbnail_url = item.find('img')['src'] # 视频标题 title = item.find('a', id='video-title').text.strip() # 播放量和发布时间 metadata = item.find_all('span', class_='inline-metadata-item style-scope ytd-video-meta-block') views = metadata[0].text.strip() if len(metadata) > 0 else '未知' publish_time = metadata[1].text.strip() if len(metadata) > 1 else '未知' print(f"第{idx}条视频:") print(f"标题: {title}") print(f"URL: {video_url}") print(f"缩略图: {thumbnail_url}") print(f"播放量: {views}") print(f"发布时间: {publish_time}\n")
解决方案2:使用YouTube Data API(合规且稳定)
直接调用官方API是更可靠的方式,避免反爬限制,但需要先在Google开发者控制台申请API密钥,代码示例:
import requests # 替换为你的API密钥 API_KEY = "你的Google API密钥" # PW-Foundation的频道ID CHANNEL_ID = "UCphU2bAGmw304CFAzy0Enuw" # 获取最新5条视频的基础信息 search_url = f"https://www.googleapis.com/youtube/v3/search?key={API_KEY}&channelId={CHANNEL_ID}&part=snippet,id&order=date&maxResults=5" search_response = requests.get(search_url) search_data = search_response.json() for idx, item in enumerate(search_data['items'], 1): if item['id']['kind'] != 'youtube#video': continue video_id = item['id']['videoId'] video_url = f"https://www.youtube.com/watch?v={video_id}" thumbnail_url = item['snippet']['thumbnails']['high']['url'] title = item['snippet']['title'] publish_time = item['snippet']['publishedAt'] # 调用接口获取播放量 stats_url = f"https://www.googleapis.com/youtube/v3/videos?key={API_KEY}&id={video_id}&part=statistics" stats_response = requests.get(stats_url) stats_data = stats_response.json() views = stats_data['items'][0]['statistics']['viewCount'] print(f"第{idx}条视频:") print(f"标题: {title}") print(f"URL: {video_url}") print(f"缩略图: {thumbnail_url}") print(f"播放量: {views}次") print(f"发布时间: {publish_time}\n")
注意事项
- 使用Selenium时,需确保ChromeDriver版本与本地Chrome浏览器版本一致,否则会报错。
- 使用YouTube Data API时,注意免费额度限制(每天默认10000单位调用量),避免超出配额导致请求失败。
内容的提问来源于stack exchange,提问作者Aqib Ansari
相关产品推荐
相关产品推荐

