Python:如何从无文件名的Content-Disposition响应头下载文件?
解决YouTube无限滚动Ajax响应的JSON下载问题
我明白你遇到的问题——YouTube的无限滚动Ajax请求返回200状态,但直接用requests.get().json()会报错,而且响应头里的content-disposition: attachment提示这是一个可下载的附件。让我一步步帮你解决这个问题:
先分析你代码里的几个坑
- Brotli压缩未处理:响应头里的
content-encoding: br说明内容是用Brotli压缩的,requests默认不会自动解码这种压缩格式,直接调用.json()肯定会失败。 - URL构造错误:你用了
&,这是HTML的转义字符,实际URL里应该用&来分隔参数,否则YouTube无法正确识别你的ctoken和continuation参数。 - 反爬拦截:YouTube对非浏览器请求非常敏感,不带正确的请求头(比如User-Agent、Cookie)会返回无效内容,哪怕状态码是200。
解决方案:手动解码压缩+模拟浏览器请求+保存JSON
第一步:安装依赖
首先需要安装Brotli解码库:
pip install requests brotli
第二步:完整的Python代码(用requests)
import requests import brotli import json # 替换成你实际的搜索链接和提取到的ctoken base_search_url = "https://www.youtube.com/results?search_query=tamil" ctoken = "xyxxxxxxxx" # 正确构造Ajax请求URL(注意用&而不是&) ajax_url = f"{base_search_url}&ctoken={ctoken}&continuation={ctoken}" # 模拟浏览器的请求头,必须带上,否则会被YouTube拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Accept-Encoding": "gzip, deflate, br", "Accept-Language": "en-US,en;q=0.9", # 从Chrome/Firefox调试器里复制你的YouTube Cookie,这很关键 "Cookie": "YOUR_YOUTUBE_COOKIE_FROM_DEVTOOLS" } # 发起请求 response = requests.get(ajax_url, headers=headers) # 解码Brotli压缩的响应内容 if response.headers.get("content-encoding") == "br": decoded_content = brotli.decompress(response.content) else: decoded_content = response.content # 解析JSON数据 try: youtube_data = json.loads(decoded_content) except json.JSONDecodeError as e: print(f"JSON解析失败:{e}") # 如果解析失败,先保存原始内容排查问题 with open("raw_response.txt", "wb") as f: f.write(decoded_content) exit() # 保存为JSON文件(对应content-disposition的下载提示) output_filename = "youtube_continuation_results.json" with open(output_filename, "w", encoding="utf-8") as f: json.dump(youtube_data, f, indent=2, ensure_ascii=False) print(f"✅ JSON数据已成功保存到 {output_filename}")
如果你用Scrapy(你最初提到的框架)
这里也给你Scrapy版本的实现,逻辑和requests一致:
import scrapy import brotli import json class YouTubeScrollSpider(scrapy.Spider): name = "youtube_scroll" start_urls = ["https://www.youtube.com/results?search_query=tamil"] def parse(self, response): # 从初始搜索页面提取ctoken(需要你根据页面结构实现,比如用XPath/CSS选择器) ctoken = response.xpath('//script[contains(text(), "continuation")]/text()').re_first(r'"ctoken":"(.*?)"') if ctoken: ajax_url = f"{response.url}&ctoken={ctoken}&continuation={ctoken}" yield scrapy.Request( url=ajax_url, headers={ "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Accept-Encoding": "gzip, deflate, br", "Cookie": "YOUR_YOUTUBE_COOKIE_FROM_DEVTOOLS" }, callback=self.parse_ajax_response ) def parse_ajax_response(self, response): # 处理Brotli压缩 if response.headers.get("Content-Encoding") == "br": decoded_body = brotli.decompress(response.body) else: decoded_body = response.body try: json_data = json.loads(decoded_body) except json.JSONDecodeError as e: self.logger.error(f"JSON解析错误:{str(e)}") with open("raw_scrapy_response.txt", "wb") as f: f.write(decoded_body) return # 保存JSON文件 with open("youtube_scrapy_results.json", "w", encoding="utf-8") as f: json.dump(json_data, f, indent=2, ensure_ascii=False) self.logger.info("✅ JSON数据已成功保存")
关键注意事项
- 动态提取ctoken:ctoken是临时的,每次滚动都会生成新的,你需要从之前的页面/响应中动态提取,不能硬编码。
- 反爬应对:YouTube会频繁检测爬虫,建议轮换User-Agent、使用代理,不要频繁请求。
- 多部分响应:如果响应头有
x-spf-response-type: multipart,说明内容是多部分的,你需要拆分后再解析JSON(不过大部分情况下解码后直接是有效JSON)。
内容的提问来源于stack exchange,提问作者Stephen Paulraj
相关产品推荐
相关产品推荐

