You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python大文件流式上传失败求助:带额外数据的正确实现方式

问题:Python大文件流式上传并传递额外数据的正确方式?

我尝试通过以下代码上传文件:

data = {'type': 'gallary', 'username': 'a457'}
url = "url/to/upload"

with open('filepath', "rb") as file:
    files = {
        "file": (
            'gallary.zip',
            file,
        )
    }
    response = requests.post(url, files=files, data=data)

部分用户的文件体积较大,但我的设备内存有限,发送POST请求时程序因out of memory错误被终止。我原以为使用with open('filepath', "rb") as file获取类文件对象可实现流式上传。

之后我尝试了许多方案中推荐的如下代码:

encoder = MultipartEncoder(
    fields={
        "file": (
            'gallary.zip',
             file,
        ),
        **data,
    },
)
response = requests.post(url, data=encoder)

此时出现415 Client Error: Unsupported Media Type错误。请问在Python中,需要传递额外数据时,实现大文件流式上传的正确方式是什么?


解决方案

1. 问题根源

  • 最初的requests.post(files=...)会默认把整个请求体(含文件)加载到内存,哪怕用了类文件对象,无法实现真正流式上传,导致大文件OOM。
  • 使用MultipartEncoder时未设置正确的Content-Type头,服务器无法识别多部分表单数据,触发415错误。

2. 正确实现代码

结合MultipartEncoder并配置正确请求头,同时确保文件流式读取:

from requests_toolbelt.multipart.encoder import MultipartEncoder
import requests

data = {'type': 'gallary', 'username': 'a457'}
url = "url/to/upload"

with open('filepath', "rb") as file:
    # 构建多部分表单编码器,指定文件MIME类型更稳妥
    encoder = MultipartEncoder(
        fields={
            "file": ('gallary.zip', file, 'application/zip'),
            **data,
        }
    )
    # 必须设置Content-Type为编码器生成的类型
    headers = {'Content-Type': encoder.content_type}
    # 发送请求,编码器会逐块读取文件,避免内存溢出
    response = requests.post(url, data=encoder, headers=headers)
    
    # 验证响应
    response.raise_for_status()
    print(response.text)

3. 可选:上传进度跟踪

如果需要监控上传进度,可使用MultipartEncoderMonitor:

from requests_toolbelt.multipart.encoder import MultipartEncoderMonitor

def progress_callback(monitor):
    print(f"上传进度: {monitor.bytes_read / monitor.len * 100:.2f}%")

with open('filepath', "rb") as file:
    encoder = MultipartEncoder(
        fields={
            "file": ('gallary.zip', file, 'application/zip'),
            **data,
        }
    )
    # 包装编码器添加进度回调
    monitor = MultipartEncoderMonitor(encoder, progress_callback)
    headers = {'Content-Type': monitor.content_type}
    response = requests.post(url, data=monitor, headers=headers)

核心注意事项

  • 必须设置Content-Type为encoder.content_type,这是解决415错误的关键。
  • MultipartEncoder会逐块读取文件内容,不会一次性加载到内存,彻底解决OOM问题。
  • 指定文件MIME类型(如application/zip)可帮助服务器正确识别文件类型,避免额外解析错误。

内容的提问来源于stack exchange,提问作者StaticName

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 19:27:15