You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Flask的YouTube播放列表计算器输出异常问题及部署指导

YouTube播放列表时长计算器问题修复、GCP部署及YouTube API迁移指南

一、修复抓取返回0时长的问题

问题原因

当前YouTube页面采用JavaScript动态渲染视频内容,requests.get只能获取静态HTML源码,视频时长元素是页面加载后异步生成的,导致BeautifulSoup无法定位到目标span.ytd-thumbnail-overlay-time-status-renderer标签,最终总时长始终为0。

解决方案1:改用Selenium动态加载页面(临时快速修复)

  • 先安装依赖:
    pip install selenium
    
  • 下载对应浏览器的驱动(如ChromeDriver),确保与浏览器版本匹配
  • 修改calculate_playlist_length函数:
    from selenium import webdriver
    from selenium.webdriver.chrome.options import Options
    import re
    import time
    
    def calculate_playlist_length(playlist_url):
        # 配置无头浏览器模式
        chrome_options = Options()
        chrome_options.add_argument("--headless=new")
        chrome_options.add_argument("--disable-gpu")
        chrome_options.add_argument("--no-sandbox")
    
        driver = webdriver.Chrome(options=chrome_options)
        driver.get(playlist_url)
        # 等待页面加载完成(可替换为显式等待提升稳定性)
        time.sleep(5)
    
        total_seconds = 0
        # 定位所有时长元素
        spans = driver.find_elements("css selector", "span.ytd-thumbnail-overlay-time-status-renderer")
        for span in spans:
            time_text = span.text.strip()
            # 处理分:秒格式
            min_sec_match = re.search(r'(\d+):(\d+)', time_text)
            if min_sec_match:
                minutes, seconds = min_sec_match.groups()
                total_seconds += int(minutes)*60 + int(seconds)
            # 处理时:分:秒格式
            hour_match = re.search(r'(\d+):(\d+):(\d+)', time_text)
            if hour_match:
                hours, minutes, seconds = hour_match.groups()
                total_seconds += int(hours)*3600 + int(minutes)*60 + int(seconds)
        
        driver.quit()
    
        hours = total_seconds // 3600
        minutes = (total_seconds % 3600) // 60
        seconds = total_seconds % 60
        return hours, minutes, seconds
    

解决方案2:迁移到YouTube API(长期稳定方案,推荐)

网页抓取易触发YouTube反爬机制,API方案更可靠,具体步骤见第三部分。

二、GCP部署步骤

1. 前期准备

  • 登录GCP控制台,创建新项目
  • 启用App Engine服务(Flask应用常用部署载体)
  • 安装gcloud CLI并完成初始化(登录账号、关联目标项目)

2. 项目配置

  • 在项目根目录创建requirements.txt,列出所有依赖:
    flask
    # 若用Selenium需添加:selenium
    # 若用API需添加:google-api-python-client
    
  • 创建app.yaml配置文件:
    runtime: python310 # 匹配你的Python版本,如python39、python311
    entrypoint: gunicorn -b :$PORT app:app
    

3. 部署执行

在终端运行部署命令:

gcloud app deploy

部署完成后,访问控制台提示的URL即可使用应用。

注意事项

  • 若使用Selenium,需配置无头模式,且GCP App Engine标准环境可能需要额外配置驱动权限,推荐优先用API方案
  • 敏感信息(如API密钥)不要硬编码,通过GCP环境变量管理

三、迁移到YouTube API的指导

1. 申请API密钥

  • 进入GCP控制台API库,搜索并启用YouTube Data API v3
  • 创建API密钥,建议设置使用限制(仅允许YouTube Data API调用),避免密钥滥用

2. 安装依赖

pip install google-api-python-client

3. 修改代码逻辑

替换原calculate_playlist_length函数,通过API批量获取视频时长:

from googleapiclient.discovery import build
import os

def calculate_playlist_length(playlist_url):
    # 从环境变量读取API密钥
    api_key = os.getenv('YOUTUBE_API_KEY')
    youtube = build('youtube', 'v3', developerKey=api_key)

    # 从URL中提取播放列表ID
    playlist_id = playlist_url.split('list=')[1].split('&')[0]

    total_seconds = 0
    next_page_token = None

    # 处理播放列表分页(API单次最多返回50条数据)
    while True:
        # 获取播放列表中的视频ID列表
        pl_request = youtube.playlistItems().list(
            part='contentDetails',
            playlistId=playlist_id,
            maxResults=50,
            pageToken=next_page_token
        )
        pl_response = pl_request.execute()

        # 收集视频ID
        video_ids = [item['contentDetails']['videoId'] for item in pl_response['items']]

        # 批量获取视频时长信息
        vid_request = youtube.videos().list(
            part='contentDetails',
            id=','.join(video_ids)
        )
        vid_response = vid_request.execute()

        # 解析ISO 8601格式的时长并累加
        for item in vid_response['items']:
            duration = item['contentDetails']['duration']
            hours, minutes, seconds = 0, 0, 0
            if 'H' in duration:
                hours = int(duration.split('H')[0].replace('PT', ''))
                duration = duration.split('H')[1]
            if 'M' in duration:
                minutes = int(duration.split('M')[0])
                duration = duration.split('M')[1]
            if 'S' in duration:
                seconds = int(duration.split('S')[0])
            total_seconds += hours*3600 + minutes*60 + seconds

        # 检查是否有下一页数据
        next_page_token = pl_response.get('nextPageToken')
        if not next_page_token:
            break

    hours = total_seconds // 3600
    minutes = (total_seconds % 3600) // 60
    seconds = total_seconds % 60
    return hours, minutes, seconds

4. 注意事项

  • YouTube API有每日配额限制(免费配额为10000单位),获取播放列表视频列表单次请求消耗1单位,批量获取视频详情单次请求消耗1单位,需合理控制请求频率
  • API密钥需通过GCP环境变量注入,不要硬编码在代码中

内容的提问来源于stack exchange,提问作者DevOpsnoob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 22:24:57