You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup4爬取YouTube视频信息失败,寻求解决办法

解决YouTube视频信息爬取失败的问题

核心原因

YouTube采用客户端动态渲染机制,你用requests获取的只是页面静态骨架,实际的浏览量、发布时间等核心数据是通过JavaScript动态加载的,所以BeautifulSoup只能解析到占位元素,无法提取真实内容。

可行解决方案

方法1:提取页面内嵌的ytInitialData(无需模拟浏览器,推荐)

YouTube会把视频核心数据以JSON形式嵌在页面的<script>标签中,直接提取解析即可:

import requests
import json
from bs4 import BeautifulSoup

url = "https://www.youtube.com/watch?v=S4E4yAktjug"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.text, "html.parser")

# 定位包含ytInitialData的script标签
for script in soup.find_all("script"):
    if "ytInitialData" in script.text:
        # 截取并解析JSON数据
        json_str = script.text.split("var ytInitialData = ")[1].split(";")[0]
        data = json.loads(json_str)
        break

# 提取浏览量与发布时间
video_info = data["contents"]["twoColumnWatchNextResults"]["results"]["results"]["contents"][0]["videoPrimaryInfoRenderer"]
view_count = video_info["viewCount"]["videoViewCountRenderer"]["viewCount"]["simpleText"]
publish_date = video_info["dateText"]["simpleText"]

print(f"浏览量: {view_count}")
print(f"发布时间: {publish_date}")

注意:该方案依赖当前YouTube页面结构,若页面更新可能需要调整JSON路径,但目前是最轻量的实现方式。

方法2:用Selenium模拟浏览器加载完整页面

如果内嵌数据方案失效,可通过Selenium模拟真实浏览器加载动态内容后再解析:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
import time

url = "https://www.youtube.com/watch?v=S4E4yAktjug"

# 配置无头模式(可选,无需打开可视化浏览器窗口)
chrome_options = Options()
chrome_options.add_argument("--headless=new")
chrome_options.add_argument("--disable-gpu")

driver = webdriver.Chrome(options=chrome_options)
driver.get(url)
time.sleep(2)  # 等待页面动态内容加载完成

# 定位并提取目标数据
view_count = driver.find_element(By.CSS_SELECTOR, ".view-count").text
publish_date = driver.find_element(By.CSS_SELECTOR, "#info-strings yt-formatted-string").text

print(f"浏览量: {view_count}")
print(f"发布时间: {publish_date}")

driver.quit()

需提前安装Selenium库及对应浏览器驱动(如ChromeDriver)。

方法3:使用YouTube Data API(最稳定合规)

若需长期稳定爬取,推荐使用官方API,数据准确且符合平台规则:

import googleapiclient.discovery

api_key = "你的API密钥"  # 从Google Cloud控制台申请获取
youtube = googleapiclient.discovery.build("youtube", "v3", developerKey=api_key)

request = youtube.videos().list(
    part="statistics,snippet",
    id="S4E4yAktjug"
)
response = request.execute()

view_count = response["items"][0]["statistics"]["viewCount"]
publish_date = response["items"][0]["snippet"]["publishedAt"]

print(f"浏览量: {view_count}")
print(f"发布时间: {publish_date}")

注意:API有调用次数配额限制,需根据需求申请对应配额。

关键提示

  • 所有方案都需设置合理的User-Agent,避免被YouTube封禁IP。
  • 控制请求频率,遵守网站robots.txt规则。

内容的提问来源于stack exchange,提问作者Rayaan khan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 14:30:14