使用BeautifulSoup4爬取YouTube视频信息失败,寻求解决办法
解决YouTube视频信息爬取失败的问题
核心原因
YouTube采用客户端动态渲染机制,你用requests获取的只是页面静态骨架,实际的浏览量、发布时间等核心数据是通过JavaScript动态加载的,所以BeautifulSoup只能解析到占位元素,无法提取真实内容。
可行解决方案
方法1:提取页面内嵌的ytInitialData(无需模拟浏览器,推荐)
YouTube会把视频核心数据以JSON形式嵌在页面的<script>标签中,直接提取解析即可:
import requests import json from bs4 import BeautifulSoup url = "https://www.youtube.com/watch?v=S4E4yAktjug" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, "html.parser") # 定位包含ytInitialData的script标签 for script in soup.find_all("script"): if "ytInitialData" in script.text: # 截取并解析JSON数据 json_str = script.text.split("var ytInitialData = ")[1].split(";")[0] data = json.loads(json_str) break # 提取浏览量与发布时间 video_info = data["contents"]["twoColumnWatchNextResults"]["results"]["results"]["contents"][0]["videoPrimaryInfoRenderer"] view_count = video_info["viewCount"]["videoViewCountRenderer"]["viewCount"]["simpleText"] publish_date = video_info["dateText"]["simpleText"] print(f"浏览量: {view_count}") print(f"发布时间: {publish_date}")
注意:该方案依赖当前YouTube页面结构,若页面更新可能需要调整JSON路径,但目前是最轻量的实现方式。
方法2:用Selenium模拟浏览器加载完整页面
如果内嵌数据方案失效,可通过Selenium模拟真实浏览器加载动态内容后再解析:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options import time url = "https://www.youtube.com/watch?v=S4E4yAktjug" # 配置无头模式(可选,无需打开可视化浏览器窗口) chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("--disable-gpu") driver = webdriver.Chrome(options=chrome_options) driver.get(url) time.sleep(2) # 等待页面动态内容加载完成 # 定位并提取目标数据 view_count = driver.find_element(By.CSS_SELECTOR, ".view-count").text publish_date = driver.find_element(By.CSS_SELECTOR, "#info-strings yt-formatted-string").text print(f"浏览量: {view_count}") print(f"发布时间: {publish_date}") driver.quit()
需提前安装Selenium库及对应浏览器驱动(如ChromeDriver)。
方法3:使用YouTube Data API(最稳定合规)
若需长期稳定爬取,推荐使用官方API,数据准确且符合平台规则:
import googleapiclient.discovery api_key = "你的API密钥" # 从Google Cloud控制台申请获取 youtube = googleapiclient.discovery.build("youtube", "v3", developerKey=api_key) request = youtube.videos().list( part="statistics,snippet", id="S4E4yAktjug" ) response = request.execute() view_count = response["items"][0]["statistics"]["viewCount"] publish_date = response["items"][0]["snippet"]["publishedAt"] print(f"浏览量: {view_count}") print(f"发布时间: {publish_date}")
注意:API有调用次数配额限制,需根据需求申请对应配额。
关键提示
- 所有方案都需设置合理的
User-Agent,避免被YouTube封禁IP。 - 控制请求频率,遵守网站
robots.txt规则。
内容的提问来源于stack exchange,提问作者Rayaan khan
相关产品推荐
相关产品推荐

