You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用requests模块爬取YouTube音视频遇403或无响应问题求助

解决YouTube音视频流请求403/无响应问题

核心问题分析

你遇到的问题主要源于两个关键错误,以及YouTube的反爬机制限制:

  • 错误使用POST请求媒体流链接:YouTube的音视频流仅支持GET请求,POST会被服务器判定为非法请求,直接返回403。
  • 错误添加Content-Length请求头:该头用于标识请求体的长度,GET请求没有请求体,添加后服务器会一直等待请求体数据,导致无响应。
  • 额外限制:YouTube的cookies会过期,且部分流链接带有签名验证,过期或未验证的链接也会触发403。

修正后的代码

import requests
import re
import json

url = 'https://www.youtube.com/watch?v=7LkIUfpX-k0'
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/115.0.0.0 Safari/537.36",
    "Accept-Language": "zh-TW,zh;q=0.9",
    "Accept-Encoding": "gzip, deflate, br, zstd",
    "Accept": "*/*",
    "Connection": "keep-alive",
    "Referer": "https://www.youtube.com/",
    "Origin": "https://www.youtube.com"  # 增强请求合法性的必要头
}
# 注意:cookies会过期,建议从当前登录YouTube的浏览器中抓取最新值替换
cookies = {
    "YSC": "e2LSE4IOe5g",
    "VISITOR_INFO1_LIVE": "BHEJagtnezo",
    "VISITOR_PRIVACY_METADATA": "CgJUVxIEGgAgag%3D%3D",
    "PREF": "f4=4000000&tz=Asia.Taipei",
    "GPS": "1"
}

# 提取页面中的播放器响应数据
response = requests.get(url=url, headers=headers, cookies=cookies)
player_response_match = re.findall('var ytInitialPlayerResponse = (.*?);var', response.text)
if not player_response_match:
    print("无法提取播放器响应数据")
    exit()
ans = json.loads(player_response_match[0])

# 筛选有效音频流链接
audio_url = None
for fmt in ans['streamingData']['adaptiveFormats']:
    if 'audio' in fmt['mimeType'] and 'url' in fmt:
        audio_url = fmt['url']
        break

if not audio_url:
    print("未找到有效音频流链接")
    exit()

# 用GET请求流链接,允许重定向并分块下载
try:
    audio_response = requests.get(
        url=audio_url, 
        headers=headers, 
        cookies=cookies, 
        allow_redirects=True, 
        stream=True  # 大文件分块下载,避免内存溢出
    )
    print(f"请求状态码:{audio_response.status_code}")
    if audio_response.status_code == 200:
        with open('video.mp3', mode='wb') as file:
            for chunk in audio_response.iter_content(chunk_size=1024*1024):
                file.write(chunk)
        print("音频文件下载完成")
    else:
        print(f"请求失败,状态码:{audio_response.status_code}")
except Exception as e:
    print(f"请求出错:{str(e)}")

关键优化点说明

  1. 请求方法修正:将requests.post改为requests.get,符合YouTube流链接的请求规范。
  2. 移除无效请求头:删除Content-Length,避免服务器等待无效的请求体。
  3. 增强请求合法性:添加Origin头,模拟浏览器的同源请求规则。
  4. 分块下载:使用stream=True和iter_content分块写入文件,适合处理大体积音视频文件。
  5. cookies有效性:如果仍返回403,需从当前浏览器的YouTube会话中抓取最新cookies替换代码中的值,过期cookies会被服务器拦截。
  6. 签名处理(进阶):若遇到无url字段的流格式,说明链接需要签名验证,需解析页面中的签名函数生成有效链接(该逻辑较复杂,需额外处理JavaScript代码)。

内容的提问来源于stack exchange,提问作者zaber8787利巴

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 10:52:24