You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取YouTube频道视频信息,find_all返回空列表求解决

解决YouTube频道视频爬取返回空列表的问题

问题原因

YouTube的视频列表是通过JavaScript动态渲染加载的,直接用requests.get()获取的只是静态HTML骨架,里面根本没有你要查找的yt-simple-endpoint inline-block style-scope ytd-thumbnail这类元素,所以find_all()自然返回空列表。

解决方案1:用Selenium模拟浏览器加载动态内容

Selenium可以模拟真实浏览器的行为,等待页面JS执行完成后再获取渲染后的HTML内容,代码示例:

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
import time

url = "https://www.youtube.com/@PW-Foundation/videos"

# 配置无头Chrome(不显示浏览器窗口)
chrome_options = Options()
chrome_options.add_argument("--headless=new")
driver = webdriver.Chrome(options=chrome_options)

driver.get(url)
# 等待页面加载,根据网络情况调整时间
time.sleep(3)

# 获取渲染后的页面源码
soup = BeautifulSoup(driver.page_source, 'html.parser')
driver.quit()

# 提取前五条视频信息
video_items = soup.find_all('ytd-grid-video-renderer', limit=5)
for idx, item in enumerate(video_items, 1):
    # 视频URL
    video_url = "https://www.youtube.com" + item.find('a', id='thumbnail')['href']
    # 缩略图URL
    thumbnail_url = item.find('img')['src']
    # 视频标题
    title = item.find('a', id='video-title').text.strip()
    # 播放量和发布时间
    metadata = item.find_all('span', class_='inline-metadata-item style-scope ytd-video-meta-block')
    views = metadata[0].text.strip() if len(metadata) > 0 else '未知'
    publish_time = metadata[1].text.strip() if len(metadata) > 1 else '未知'
    
    print(f"第{idx}条视频:")
    print(f"标题: {title}")
    print(f"URL: {video_url}")
    print(f"缩略图: {thumbnail_url}")
    print(f"播放量: {views}")
    print(f"发布时间: {publish_time}\n")

解决方案2:使用YouTube Data API(合规且稳定)

直接调用官方API是更可靠的方式,避免反爬限制,但需要先在Google开发者控制台申请API密钥,代码示例:

import requests

# 替换为你的API密钥
API_KEY = "你的Google API密钥"
# PW-Foundation的频道ID
CHANNEL_ID = "UCphU2bAGmw304CFAzy0Enuw"

# 获取最新5条视频的基础信息
search_url = f"https://www.googleapis.com/youtube/v3/search?key={API_KEY}&channelId={CHANNEL_ID}&part=snippet,id&order=date&maxResults=5"
search_response = requests.get(search_url)
search_data = search_response.json()

for idx, item in enumerate(search_data['items'], 1):
    if item['id']['kind'] != 'youtube#video':
        continue
    
    video_id = item['id']['videoId']
    video_url = f"https://www.youtube.com/watch?v={video_id}"
    thumbnail_url = item['snippet']['thumbnails']['high']['url']
    title = item['snippet']['title']
    publish_time = item['snippet']['publishedAt']
    
    # 调用接口获取播放量
    stats_url = f"https://www.googleapis.com/youtube/v3/videos?key={API_KEY}&id={video_id}&part=statistics"
    stats_response = requests.get(stats_url)
    stats_data = stats_response.json()
    views = stats_data['items'][0]['statistics']['viewCount']
    
    print(f"第{idx}条视频:")
    print(f"标题: {title}")
    print(f"URL: {video_url}")
    print(f"缩略图: {thumbnail_url}")
    print(f"播放量: {views}次")
    print(f"发布时间: {publish_time}\n")

注意事项

  • 使用Selenium时,需确保ChromeDriver版本与本地Chrome浏览器版本一致,否则会报错。
  • 使用YouTube Data API时,注意免费额度限制(每天默认10000单位调用量),避免超出配额导致请求失败。

内容的提问来源于stack exchange,提问作者Aqib Ansari

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 06:12:49