无需下载本地,通过URL直接加载服务器Pickle文件至FastAPI
直接从HTTP URL加载大型Pickle文件至FastAPI(无需本地下载)
我服务器上有若干可通过静态HTTP URL访问的大型Pickle文件,目前需要手动下载到本地后才能加载到FastAPI中使用,现有下载代码如下:
import requests import tqdm import time # 补全原代码缺失的导入 def download_file(url, save_path): block_size = 1024 file_name = url.split("/")[-1] try: headers = {} response = requests.request( "HEAD", url, headers=headers, ) if response.status_code != 200: return None else: time.sleep(1) total = int(response.headers["Content-Length"], 0) resp = requests.get( url, headers=headers, stream=True, ) with open(save_path, "wb") as file, tqdm( desc=save_path, total=total, unit="B", unit_scale=True, unit_divisor=1024, ) as bar: for data in resp.iter_content(chunk_size=block_size): size = file.write(data) bar.update(size) return total except requests.exceptions.RequestException as e: raise e
以下是无需本地下载、直接从服务器加载Pickle文件的优化方案:
方案1:内存流式加载(适合多数场景)
直接通过requests获取流式响应,用io.BytesIO将响应内容包装成类文件对象,跳过磁盘写入步骤,直接用pickle.load加载。这种方法省去磁盘IO开销,速度更快且不占用本地存储。
代码示例:
import requests import pickle import io from tqdm import tqdm def load_pickle_from_url(url): try: # HEAD请求确认文件存在并获取总大小 head_resp = requests.head(url) head_resp.raise_for_status() total_size = int(head_resp.headers.get("Content-Length", 0)) # 流式获取文件内容 with requests.get(url, stream=True) as resp: resp.raise_for_status() buffer = io.BytesIO() # 可选:显示下载进度条 with tqdm(total=total_size, unit="B", unit_scale=True, unit_divisor=1024) as bar: # 用1MB块大小提升传输效率 for chunk in resp.iter_content(chunk_size=1024*1024): if chunk: buffer.write(chunk) bar.update(len(chunk)) # 将缓冲区指针移至开头,供pickle读取 buffer.seek(0) return pickle.load(buffer) except requests.exceptions.RequestException as e: raise e except pickle.UnpicklingError as e: raise RuntimeError("Pickle文件解析失败") from e
在FastAPI中使用(推荐启动时预加载):
from fastapi import FastAPI app = FastAPI() pickle_data = None # 启动时加载一次,避免每次请求重复拉取文件 @app.on_event("startup") async def preload_pickle(): global pickle_data pickle_data = load_pickle_from_url("http://your-server-url/target-file.pkl") @app.get("/use-pickle-data") async def use_pickle(): # 直接使用预加载的pickle数据 return {"data_preview": str(pickle_data[:10])} # 示例:返回前10条数据预览
方案2:超大文件分块加载(内存有限时)
如果Pickle文件体积远超可用内存,可利用HTTP范围请求分块读取(需服务器支持Range请求头),将内容逐步写入内存缓冲区后再加载:
import requests import pickle import io def load_large_pickle_from_url(url, chunk_size=1024*1024*100): # 100MB每块 try: head_resp = requests.head(url) head_resp.raise_for_status() total_size = int(head_resp.headers.get("Content-Length", 0)) buffer = io.BytesIO() # 分块请求文件内容 for start in range(0, total_size, chunk_size): end = min(start + chunk_size - 1, total_size - 1) headers = {"Range": f"bytes={start}-{end}"} with requests.get(url, headers=headers) as resp: resp.raise_for_status() buffer.write(resp.content) buffer.seek(0) return pickle.load(buffer) except requests.exceptions.RequestException as e: raise e
注意事项
- 确保HTTP服务器支持
HEAD请求和流式响应(多数静态文件服务器默认支持) - 敏感文件请使用HTTPS协议传输,避免数据泄露
- 大文件预加载时,需调整FastAPI启动超时时间(可通过
uvicorn --timeout-keep-alive 60设置) - 频繁使用的Pickle文件建议预加载到全局变量,减少重复请求开销
内容的提问来源于stack exchange,提问作者8589934592
相关产品推荐
相关产品推荐

