如何配置yt-dlp动态选可用格式解决YouTube视频下载格式错误?
yt-dlp动态选择最佳可用音频格式解决下载失败问题
问题描述
在macOS的Python 3.13虚拟环境中使用yt-dlp下载特定视频(ID: QageNN-V8rY)时抛出DownloadError,错误信息如下:
File "/Users/ashazakhtar/coding/project-idea/.venv/lib/python3.13/site-packages/yt_dlp/YoutubeDL.py", line 1093, in trouble raise DownloadError(message, exc_info) yt_dlp.utils.DownloadError: ERROR: [youtube] QageNN-V8rY: Requested format is not available. Use --list-formats for a list of available formats
当前使用的代码:
def transcript_video(self, url): with tempfile.TemporaryDirectory() as temp_dir: output_path = os.path.join(temp_dir, "audio.m4a") cookie_path = os.path.join(settings.BASE_DIR, "www.youtube.com_cookies.txt") ydl_opts = { "format": "bestaudio[ext=m4a]/bestaudio[ext=webm]/bestaudio", "quiet": True, "outtmpl": output_path, "cookiefile": cookie_path, "nocheckcertificate": True, "user_agent": "Mozilla/5.0", "cookies_readonly": True, } with YoutubeDL(ydl_opts) as ydl: # type: ignore ydl.download([url]) # type: ignore print(f"Audio file size: {os.path.getsize(output_path)} bytes") # Send audio file to local FastAPI with open(output_path, "rb") as audio_file: response = requests.post( f"{FASTAPI_BASE_URL}/transcribe", files={"file": ("audio.m4a", audio_file, "audio/m4a")}, timeout=300, # whisper can be slow on CPU ) response.raise_for_status() result = response.json() print(f"Transcription done: {result['text'][:100]}...") return result["text"]
已尝试的操作:
- 验证视频URL正确且可在浏览器中访问
- 将虚拟环境中的yt-dlp更新至最新版本(
pip install -U yt-dlp) - 通过命令行执行
yt-dlp --list-formats https://www.youtube.com/watch?v=QageNN-V8rY可查看可用格式,但Python脚本仍报错
解决方案
1. 调整格式选择规则,实现动态匹配
当前的format规则虽然覆盖了m4a、webm,但可能该视频的最佳音频格式不在指定优先级内,或是筛选逻辑不够灵活。可以简化规则,让yt-dlp自动降级选择可用的最佳音频:
将ydl_opts中的format参数修改为:
"format": "bestaudio[ext=m4a]/bestaudio",
这种写法会优先尝试m4a格式,不可用时自动选择其他可用的最佳音频,避免直接抛出错误。如果不需要偏好m4a,也可以直接用"format": "bestaudio/best"让yt-dlp完全自动选择。
2. 处理动态文件格式匹配问题
原来的代码固定使用audio.m4a作为输出路径,若下载的是其他格式(如webm)会导致文件后缀与实际格式不匹配,进而影响后续FastAPI上传的兼容性。可以通过两种方式解决:
方法一:自动获取实际下载的文件路径
修改outtmpl为带扩展名占位符的格式,再通过yt-dlp的返回结果获取真实文件名:
"outtmpl": os.path.join(temp_dir, "audio.%(ext)s"),
下载时通过extract_info获取视频信息,进而拿到实际路径:
with YoutubeDL(ydl_opts) as ydl: info_dict = ydl.extract_info(url, download=True) output_path = ydl.prepare_filename(info_dict)
方法二:统一转换为m4a格式
如果FastAPI服务仅支持m4a,可以配置yt-dlp调用ffmpeg自动转换格式(需先通过brew install ffmpeg在macOS上安装ffmpeg):
ydl_opts = { "format": "bestaudio", "quiet": True, "outtmpl": os.path.join(temp_dir, "audio.m4a"), "cookiefile": cookie_path, "nocheckcertificate": True, "user_agent": "Mozilla/5.0", "cookies_readonly": True, "postprocessors": [{ 'key': 'FFmpegExtractAudio', 'preferredcodec': 'm4a', 'preferredquality': '192', }], }
3. 完整修改后的代码示例(方法一)
def transcript_video(self, url): with tempfile.TemporaryDirectory() as temp_dir: cookie_path = os.path.join(settings.BASE_DIR, "www.youtube.com_cookies.txt") ydl_opts = { "format": "bestaudio[ext=m4a]/bestaudio", "quiet": True, "outtmpl": os.path.join(temp_dir, "audio.%(ext)s"), "cookiefile": cookie_path, "nocheckcertificate": True, "user_agent": "Mozilla/5.0", "cookies_readonly": True, } with YoutubeDL(ydl_opts) as ydl: # type: ignore info_dict = ydl.extract_info(url, download=True) output_path = ydl.prepare_filename(info_dict) print(f"Audio file size: {os.path.getsize(output_path)} bytes") # 获取实际文件扩展名,设置正确的MIME类型 file_ext = os.path.splitext(output_path)[1].lstrip('.') mime_type = f"audio/{file_ext}" if file_ext in ['m4a', 'webm', 'mp3'] else "audio/mpeg" # Send audio file to local FastAPI with open(output_path, "rb") as audio_file: response = requests.post( f"{FASTAPI_BASE_URL}/transcribe", files={"file": (f"audio.{file_ext}", audio_file, mime_type)}, timeout=300, # whisper can be slow on CPU ) response.raise_for_status() result = response.json() print(f"Transcription done: {result['text'][:100]}...") return result["text"]
内容的提问来源于stack exchange,提问作者Ashaz Akhtar
相关产品推荐
相关产品推荐

