如何用Python直接将文件下载至SSD而非先存RAM?支持Cookies
针对你遇到的大文件下载先占内存再写磁盘的问题,以下是几种直接流式写入磁盘的解决方案,均支持Cookies:
优化requests库的用法
默认的iter_content可能存在额外缓冲,改用底层响应流读取可以更直接地控制写入逻辑:
import requests URL = "https://example.com/video.mp4" local_filename = "video.mp4" req_cookies = {"your_cookie_key": "cookie_value"} with requests.get(URL, cookies=req_cookies, stream=True) as r: r.raise_for_status() with open(local_filename, 'wb') as f: # 直接读取原始响应流,减少中间内存缓冲 for chunk in r.raw.stream(8192, decode_content=False): if chunk: f.write(chunk)
使用urllib3库(更底层的控制)
urllib3的流式读取可以完全避免预加载内容到内存:
import urllib3 URL = "https://example.com/video.mp4" local_filename = "video.mp4" req_cookies = {"your_cookie_key": "cookie_value"} # 将Cookies转换为请求头格式 headers = {"Cookie": "; ".join([f"{k}={v}" for k, v in req_cookies.items()])} http = urllib3.PoolManager() with http.request('GET', URL, headers=headers, preload_content=False) as r: if r.status != 200: raise Exception(f"请求失败,状态码:{r.status}") with open(local_filename, 'wb') as f: while True: chunk = r.read(8192) if not chunk: break f.write(chunk)
多线程分块下载(解决大文件+并行下载的内存问题)
如果需要多线程并行下载,采用分块写入的方式,每个线程只处理对应字节区间,直接写入文件指定位置:
import requests from threading import Thread import os URL = "https://example.com/video.mp4" local_filename = "video.mp4" req_cookies = {"your_cookie_key": "cookie_value"} num_threads = 4 def download_chunk(start, end): headers = {"Range": f"bytes={start}-{end}"} with requests.get(URL, cookies=req_cookies, headers=headers, stream=True) as r: r.raise_for_status() with open(local_filename, 'r+b') as f: f.seek(start) for chunk in r.raw.stream(8192, decode_content=False): if chunk: f.write(chunk) # 获取文件总大小 with requests.head(URL, cookies=req_cookies) as r: r.raise_for_status() total_size = int(r.headers.get('content-length', 0)) # 创建对应大小的空文件 with open(local_filename, 'wb') as f: f.truncate(total_size) # 分配每个线程的下载区间 chunk_size = total_size // num_threads threads = [] for i in range(num_threads): start = i * chunk_size end = start + chunk_size - 1 if i != num_threads - 1 else total_size - 1 thread = Thread(target=download_chunk, args=(start, end)) threads.append(thread) thread.start() # 等待所有线程完成 for thread in threads: thread.join()
内容的提问来源于stack exchange,提问作者one_hell_of_a_guy
相关产品推荐
相关产品推荐

