You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python提取YouTube播放列表指定日期后视频描述遇请求过多错误求解决方案

解决方案

一、使用YouTube Data API(最稳定可靠)

这是官方提供的接口,请求限制宽松(每天10000单位配额,单个视频请求占1单位,1000条完全够用),不会轻易触发"too many requests"。

步骤:

  1. 去Google Cloud控制台创建项目,启用YouTube Data API v3,生成API密钥。
  2. 用google-api-python-client库调用接口,先获取播放列表所有视频ID,再批量获取视频信息。

示例代码:

from googleapiclient.discovery import build
from datetime import datetime

API_KEY = "你的API密钥"
PLAYLIST_ID = "PLG8IrydigQfcRNrWVqNkeZiCJ_DWgXDVX"
start_date = datetime(2023, 4, 1).isoformat() + "Z"  # 转成ISO 8601格式,带时区

# 初始化API客户端
youtube = build('youtube', 'v3', developerKey=API_KEY)

# 获取播放列表所有视频ID
video_ids = []
next_page_token = None
while True:
    pl_request = youtube.playlistItems().list(
        part="contentDetails",
        playlistId=PLAYLIST_ID,
        maxResults=50,
        pageToken=next_page_token
    )
    pl_response = pl_request.execute()
    video_ids.extend([item['contentDetails']['videoId'] for item in pl_response['items']])
    
    next_page_token = pl_response.get('nextPageToken')
    if not next_page_token:
        break

# 批量获取视频的上传日期和描述,每次最多50个ID
descriptions = []
for i in range(0, len(video_ids), 50):
    batch_ids = video_ids[i:i+50]
    vid_request = youtube.videos().list(
        part="snippet",
        id=",".join(batch_ids)
    )
    vid_response = vid_request.execute()
    
    for item in vid_response['items']:
        upload_date = item['snippet']['publishedAt']
        if datetime.fromisoformat(upload_date.replace("Z", "+00:00")) >= start_date:
            descriptions.append(item['snippet']['description'])

print(f"提取到{len(descriptions)}条符合条件的视频描述")

二、优化pytube的请求策略

如果不想用API,可以调整pytube的请求方式,降低触发反爬的概率:

  • 添加自定义请求头(模拟浏览器),避免被识别为爬虫
  • 增加随机延迟,避免短时间内大量请求
  • 批量获取视频URL后再逐个处理,而不是直接遍历BBC.videos(这个属性会实时请求每个视频的信息,容易触发限制)

示例代码:

from pytube import Playlist, YouTube
from datetime import datetime
import time
import random
from pytube.request import set_header

# 设置浏览器请求头
set_header("User-Agent", "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")

BBC = Playlist("https://www.youtube.com/playlist?list=PLG8IrydigQfcRNrWVqNkeZiCJ_DWgXDVX")
start_date = datetime(year=2023, month=4, day=1)

# 先获取所有视频URL,避免提前触发大量请求
video_urls = list(BBC.video_urls)
descriptions = []

for url in video_urls:
    try:
        video = YouTube(url)
        upload_date = video.upload_date
        if upload_date >= start_date:
            # 直接用video.description获取描述,比vid_info更稳定
            descriptions.append(video.description)
        # 添加随机延迟,1-3秒
        time.sleep(random.uniform(1, 3))
    except Exception as e:
        print(f"处理视频{url}时出错: {str(e)}")
        # 出错时延迟更久一点再重试
        time.sleep(random.uniform(5, 10))

print(f"提取到{len(descriptions)}条符合条件的视频描述")

三、爬虫方案(使用Requests+BeautifulSoup)

如果上述方法都不行,可以用爬虫模拟浏览器请求,但需要处理YouTube的反爬机制(比如Cookie、动态渲染的内容):

  • 使用requests发送请求,带上有效的Cookie和User-Agent
  • 用BeautifulSoup解析页面,提取上传日期和描述
  • 同样需要添加延迟,避免频繁请求

示例代码(注意:YouTube页面结构可能会变化,需根据实际情况调整选择器):

import requests
from bs4 import BeautifulSoup
from datetime import datetime
import time
import random

PLAYLIST_URL = "https://www.youtube.com/playlist?list=PLG8IrydigQfcRNrWVqNkeZiCJ_DWgXDVX"
start_date = datetime(2023, 4, 1)
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
    "Cookie": "你的YouTube登录Cookie(可选,未登录可能限制更多)"
}

# 获取播放列表所有视频链接
video_urls = []
next_page_url = PLAYLIST_URL
while next_page_url:
    response = requests.get(next_page_url, headers=headers)
    soup = BeautifulSoup(response.text, 'html.parser')
    
    # 提取视频链接
    for a_tag in soup.select('a#video-title'):
        video_url = "https://www.youtube.com" + a_tag['href']
        video_urls.append(video_url)
    
    # 获取下一页链接
    next_page = soup.select_one('a#pagination-next')
    next_page_url = "https://www.youtube.com" + next_page['href'] if next_page else None
    
    time.sleep(random.uniform(2, 4))

# 逐个处理视频页面
descriptions = []
for url in video_urls:
    try:
        response = requests.get(url, headers=headers)
        soup = BeautifulSoup(response.text, 'html.parser')
        
        # 提取上传日期(格式示例:"Apr 5, 2023")
        upload_date_str = soup.select_one('div#info-strings yt-formatted-string').text
        upload_date = datetime.strptime(upload_date_str, "%b %d, %Y")
        
        if upload_date >= start_date:
            # 提取描述
            desc_element = soup.select_one('div#description-inline-expander yt-formatted-string')
            description = desc_element.text.strip() if desc_element else ""
            descriptions.append(description)
        
        time.sleep(random.uniform(2, 4))
    except Exception as e:
        print(f"处理视频{url}时出错: {str(e)}")
        time.sleep(random.uniform(5, 10))

print(f"提取到{len(descriptions)}条符合条件的视频描述")

注意事项

  • 爬虫方案的页面选择器可能会随YouTube页面更新而失效,需要定期检查调整
  • 未登录状态下YouTube的请求限制更严格,建议登录后获取Cookie使用(但不要分享Cookie)
  • 所有方案都要控制请求频率,避免给服务器造成过大压力,同时降低被封禁的风险

内容的提问来源于stack exchange,提问作者Physics_Student

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 22:25:00