You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何配置yt-dlp动态选可用格式解决YouTube视频下载格式错误?

yt-dlp动态选择最佳可用音频格式解决下载失败问题

问题描述

在macOS的Python 3.13虚拟环境中使用yt-dlp下载特定视频(ID: QageNN-V8rY)时抛出DownloadError,错误信息如下:

File "/Users/ashazakhtar/coding/project-idea/.venv/lib/python3.13/site-packages/yt_dlp/YoutubeDL.py", line 1093, in trouble
    raise DownloadError(message, exc_info)
yt_dlp.utils.DownloadError: ERROR: [youtube] QageNN-V8rY: Requested format is not available. Use --list-formats for a list of available formats

当前使用的代码:

def transcript_video(self, url):
    with tempfile.TemporaryDirectory() as temp_dir:
        output_path = os.path.join(temp_dir, "audio.m4a")
        cookie_path = os.path.join(settings.BASE_DIR, "www.youtube.com_cookies.txt")
        ydl_opts = {
            "format": "bestaudio[ext=m4a]/bestaudio[ext=webm]/bestaudio",
            "quiet": True,
            "outtmpl": output_path,
            "cookiefile": cookie_path,
            "nocheckcertificate": True,
            "user_agent": "Mozilla/5.0",
            "cookies_readonly": True,
        }

        with YoutubeDL(ydl_opts) as ydl:  # type: ignore
            ydl.download([url])  # type: ignore

        print(f"Audio file size: {os.path.getsize(output_path)} bytes")

        # Send audio file to local FastAPI
        with open(output_path, "rb") as audio_file:
            response = requests.post(
                f"{FASTAPI_BASE_URL}/transcribe",
                files={"file": ("audio.m4a", audio_file, "audio/m4a")},
                timeout=300,  # whisper can be slow on CPU
            )

        response.raise_for_status()
        result = response.json()
        print(f"Transcription done: {result['text'][:100]}...")
        return result["text"]

已尝试的操作:

  • 验证视频URL正确且可在浏览器中访问
  • 将虚拟环境中的yt-dlp更新至最新版本(pip install -U yt-dlp)
  • 通过命令行执行yt-dlp --list-formats https://www.youtube.com/watch?v=QageNN-V8rY可查看可用格式,但Python脚本仍报错

解决方案

1. 调整格式选择规则,实现动态匹配

当前的format规则虽然覆盖了m4a、webm,但可能该视频的最佳音频格式不在指定优先级内,或是筛选逻辑不够灵活。可以简化规则,让yt-dlp自动降级选择可用的最佳音频:

将ydl_opts中的format参数修改为:

"format": "bestaudio[ext=m4a]/bestaudio",

这种写法会优先尝试m4a格式,不可用时自动选择其他可用的最佳音频,避免直接抛出错误。如果不需要偏好m4a,也可以直接用"format": "bestaudio/best"让yt-dlp完全自动选择。

2. 处理动态文件格式匹配问题

原来的代码固定使用audio.m4a作为输出路径,若下载的是其他格式(如webm)会导致文件后缀与实际格式不匹配,进而影响后续FastAPI上传的兼容性。可以通过两种方式解决:

方法一:自动获取实际下载的文件路径

修改outtmpl为带扩展名占位符的格式,再通过yt-dlp的返回结果获取真实文件名:

"outtmpl": os.path.join(temp_dir, "audio.%(ext)s"),

下载时通过extract_info获取视频信息,进而拿到实际路径:

with YoutubeDL(ydl_opts) as ydl:
    info_dict = ydl.extract_info(url, download=True)
    output_path = ydl.prepare_filename(info_dict)

方法二:统一转换为m4a格式

如果FastAPI服务仅支持m4a,可以配置yt-dlp调用ffmpeg自动转换格式(需先通过brew install ffmpeg在macOS上安装ffmpeg):

ydl_opts = {
    "format": "bestaudio",
    "quiet": True,
    "outtmpl": os.path.join(temp_dir, "audio.m4a"),
    "cookiefile": cookie_path,
    "nocheckcertificate": True,
    "user_agent": "Mozilla/5.0",
    "cookies_readonly": True,
    "postprocessors": [{
        'key': 'FFmpegExtractAudio',
        'preferredcodec': 'm4a',
        'preferredquality': '192',
    }],
}

3. 完整修改后的代码示例(方法一)

def transcript_video(self, url):
    with tempfile.TemporaryDirectory() as temp_dir:
        cookie_path = os.path.join(settings.BASE_DIR, "www.youtube.com_cookies.txt")
        ydl_opts = {
            "format": "bestaudio[ext=m4a]/bestaudio",
            "quiet": True,
            "outtmpl": os.path.join(temp_dir, "audio.%(ext)s"),
            "cookiefile": cookie_path,
            "nocheckcertificate": True,
            "user_agent": "Mozilla/5.0",
            "cookies_readonly": True,
        }

        with YoutubeDL(ydl_opts) as ydl:  # type: ignore
            info_dict = ydl.extract_info(url, download=True)
            output_path = ydl.prepare_filename(info_dict)

        print(f"Audio file size: {os.path.getsize(output_path)} bytes")
        
        # 获取实际文件扩展名,设置正确的MIME类型
        file_ext = os.path.splitext(output_path)[1].lstrip('.')
        mime_type = f"audio/{file_ext}" if file_ext in ['m4a', 'webm', 'mp3'] else "audio/mpeg"

        # Send audio file to local FastAPI
        with open(output_path, "rb") as audio_file:
            response = requests.post(
                f"{FASTAPI_BASE_URL}/transcribe",
                files={"file": (f"audio.{file_ext}", audio_file, mime_type)},
                timeout=300,  # whisper can be slow on CPU
            )

        response.raise_for_status()
        result = response.json()
        print(f"Transcription done: {result['text'][:100]}...")
        return result["text"]

内容的提问来源于stack exchange,提问作者Ashaz Akhtar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.02 04:24:51