You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在FastAPI中流式处理大文件上传?

FastAPI流式处理大文件上传方案

FastAPI完全可以实现流式处理大文件上传,无需依赖临时文件或加载整个文件到内存,下面是具体方案和示例:

方案一:基于python-multipart流式解析multipart/form-data

FastAPI底层依赖python-multipart处理表单数据,我们可以直接用它的流式API解析请求体,避免使用UploadFile带来的临时文件机制。

代码示例

from fastapi import FastAPI, Request, HTTPException
import multipart

app = FastAPI()

@app.post("/stream-upload")
async def stream_upload(request: Request):
    content_type = request.headers.get("Content-Type")
    if not content_type or "multipart/form-data" not in content_type:
        raise HTTPException(status_code=400, detail="请求必须为multipart/form-data格式")
    
    # 解析multipart边界符
    boundary = content_type.split("boundary=")[-1]
    parser = multipart.MultipartParser(request.stream(), boundary.encode())
    
    async for part in parser:
        if part.filename:
            # 分块读取文件并处理,此处以本地存储为例
            chunk_size = 1024 * 1024  # 按1MB分块
            total_bytes = 0
            async with open(f"./uploads/{part.filename}", "wb") as f:
                while chunk := await part.read(chunk_size):
                    f.write(chunk)
                    total_bytes += len(chunk)
                    # 可在此添加自定义分块逻辑:比如校验、同步到云存储等
                    print(f"已处理 {total_bytes / (1024*1024):.2f} MB")
            return {"filename": part.filename, "total_size": total_bytes}
    raise HTTPException(status_code=400, detail="未找到上传的文件")

方案二:直接读取原始请求流(适合非multipart场景)

如果你的上传场景不需要额外表单字段,仅需上传原始文件,可以直接读取请求的原始字节流:

代码示例

from fastapi import FastAPI, Request

app = FastAPI()

@app.post("/raw-stream-upload")
async def raw_stream_upload(request: Request):
    chunk_size = 1024 * 1024  # 1MB分块
    total_bytes = 0
    # 替换为你的目标存储/业务处理逻辑
    async with open("./uploads/raw_file.bin", "wb") as f:
        async for chunk in request.stream():
            f.write(chunk)
            total_bytes += len(chunk)
            print(f"已处理 {total_bytes / (1024*1024):.2f} MB")
    return {"total_size": total_bytes}

关键说明

  • 舍弃UploadFile的原因:UploadFile默认使用tempfile.SpooledTemporaryFile,会先将文件缓存到内存,超过阈值(默认1MB)才写入磁盘,无法实现真正的流式处理。
  • 分块大小可根据业务需求调整,比如结合网络带宽、存储性能灵活设置。
  • 流式处理时,需确保下游逻辑(如存储、数据校验)也支持分块操作,避免阻塞流程。

内容的提问来源于stack exchange,提问作者LtGenFlower

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 10:17:25