You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Python并行调用Dropbox API的files_upload_session_append_v2?

实现Dropbox API并行大文件上传的问题与修复

我希望通过Dropbox API并行上传大文件(比顺序上传更快),官方文档说明如下:

默认情况下,上传会话要求你通过连续的files_upload_session_start、files_upload_session_append_v2、files_upload_session_finish调用按顺序发送文件内容。为了获得更好的性能,你可以选择使用UploadSessionType.concurrent上传会话。要启动新的并发会话,请将UploadSessionStartArg.session_type设置为UploadSessionType.concurrent。之后,你可以通过并发的files_upload_session_append_v2请求发送文件数据。最后用files_upload_session_finish结束会话。并发会话有一些约束:你不能在files_upload_session_start或files_upload_session_finish调用中发送数据,只能用files_upload_session_append_v2调用。此外,files_upload_session_append_v2上传的数据必须是4194304字节的倍数(最后一个设置了UploadSessionStartArg.close=True的files_upload_session_append_v2除外,它可以包含任何剩余数据)。

但我不确定如何实现(比如没有files_upload_async_session_append_v2()方法),也找不到相关示例。我尝试了以下代码,但相比顺序上传没有速度提升:

import asyncio
import pathlib
from dropbox import DropboxBase, UploadSessionType, UploadSessionCursor, CommitInfo

async def upload_file(local_file_path: str, remote_folder_path: str, client: DropboxBase):
    """
        Uploads a file to Dropbox by chunks. This method uses v2 methods of Dropbox API.

        Example:
            upload_file('test.txt', '/Builds/', dropbox_client)

        :param local_file_path:
        :param remote_folder_path: A path to a folder on Dropbox, must end with a slash.
        :param client: Authorized Dropbox client.
        :return:
    """
    with open(local_file_path, 'rb') as file_stream:
        await __upload_file_by_chunks(file_stream, local_file_path, remote_folder_path, client)

async def test(data: bytes, cursor: UploadSessionCursor, client: DropboxBase, close: bool = False):
    client.files_upload_session_append_v2(data, cursor, close=close)


async def __upload_file_by_chunks(file_stream: BinaryIO, local_file_path: str, remote_folder_path: str, client: DropboxBase):
    # As default size for a chunk 4 MB were chosen. I think it's a good compromise between speed and reliability.
    # Also, Dropbox API guide says "Consider uploading chunks in multiples of 4 MBs."
    # ATTENTION: The maximum value can be placed here is 150 MB.
    chunk_size_bytes = 4 * 1024 * 1024

    session_id = __start_upload_session(client)
    cursor = __create_upload_session_cursor(file_stream, session_id)
    file_length = pathlib.Path(local_file_path).stat().st_size

    test_pool = set()

    # TODO: In theory this can be done in parallel, that should speed up the file upload.
    #  Maybe instead of while loop we can precalculate all chunks and then upload them in parallel.
    while file_stream.tell() < file_length:
        if __chunk_size_is_bigger_than_left_data(file_stream.tell(), file_length, chunk_size_bytes):
            chunk_size_bytes = file_length - file_stream.tell()
            test_pool.add(asyncio.create_task(test(file_stream.read(chunk_size_bytes), cursor, client, close=True)))
            continue

        test_pool.add(asyncio.create_task(test(file_stream.read(chunk_size_bytes), cursor, client)))
        cursor = __create_upload_session_cursor(file_stream, session_id)

    await asyncio.wait(test_pool)

    client.files_upload_session_finish(bytes(), cursor, commit=CommitInfo(
        path=__construct_remote_file_path(local_file_path, remote_folder_path)))


def __start_upload_session(client: DropboxBase) -> str:
    session_start_response = client.files_upload_session_start(bytes(), session_type=UploadSessionType.concurrent)
    return session_start_response.session_id

def __create_upload_session_cursor(file_stream, session_id):
    return UploadSessionCursor(session_id=session_id, offset=file_stream.tell())

def __chunk_size_is_bigger_than_left_data(current_offset, file_length, chunk_size):
    return (current_offset + chunk_size) > file_length

def __construct_remote_file_path(local_file_path, remote_folder_path):
    return remote_folder_path + pathlib.Path(local_file_path).name

# 调用示例
# asyncio.run(upload_file(file_name, DROPBOX_TEST_FOLDER, client))

问题分析

你的代码没有实现真正的并行上传,核心问题如下:

  1. 同步客户端阻塞异步执行:使用的是同步版Dropbox客户端,异步任务里调用同步API会阻塞事件循环,所有任务实际是串行执行,无法实现并行加速。
  2. Cursor偏移量逻辑错误:依赖文件流的当前位置生成Cursor,并行任务会拿到错误的offset,导致Dropbox无法正确拼接chunk,甚至上传失败。
  3. 错误设置close参数:并行上传时无需在最后一个chunk设置close=True,所有chunk上传完成后统一调用files_upload_session_finish即可。

修复后的代码

要实现真正的并行上传,必须使用异步Dropbox客户端,并预计算每个chunk的偏移量:

import asyncio
import pathlib
from dropbox import AsyncDropbox, UploadSessionType, UploadSessionCursor, CommitInfo

async def upload_file(local_file_path: str, remote_folder_path: str, client: AsyncDropbox):
    chunk_size = 4 * 1024 * 1024  # 4MB,符合API要求的倍数
    file_path = pathlib.Path(local_file_path)
    file_size = file_path.stat().st_size
    remote_path = remote_folder_path + file_path.name

    # 启动并发上传会话
    session_start = await client.files_upload_session_start(
        bytes(),
        session_type=UploadSessionType.concurrent
    )
    session_id = session_start.session_id

    # 预计算所有chunk的偏移量和数据
    chunks = []
    with open(local_file_path, 'rb') as f:
        offset = 0
        while offset < file_size:
            read_size = chunk_size if (offset + chunk_size) <= file_size else (file_size - offset)
            data = f.read(read_size)
            chunks.append( (offset, data) )
            offset += read_size

    # 并行上传所有chunk
    async def upload_chunk(offset: int, data: bytes):
        cursor = UploadSessionCursor(session_id=session_id, offset=offset)
        await client.files_upload_session_append_v2(data, cursor)

    tasks = [asyncio.create_task(upload_chunk(offset, data)) for offset, data in chunks]
    await asyncio.gather(*tasks)

    # 完成上传会话
    final_cursor = UploadSessionCursor(session_id=session_id, offset=file_size)
    await client.files_upload_session_finish(
        bytes(),
        final_cursor,
        commit=CommitInfo(path=remote_path)
    )

# 调用示例
# asyncio.run(upload_file("large_file.bin", "/uploads/", AsyncDropbox("YOUR_ACCESS_TOKEN")))

关键改进点

  • 使用AsyncDropbox异步客户端,所有API调用用await,避免阻塞事件循环,真正实现并行请求。
  • 预读取所有chunk并记录对应偏移量,确保每个并行任务能正确指定自己的上传位置。
  • 移除错误的close=True设置,统一在最后调用files_upload_session_finish完成上传。

内容的提问来源于stack exchange,提问作者user9815351

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 19:10:28