You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过pygdrive3或google-api-python-client批量获取Google Drive文件?

优化Google Drive小文件批量获取速度的方案

首先得明确:Google Drive API本身没有提供单次请求直接批量获取多个文件媒体内容的接口,但我们可以通过「批量请求」机制把多个文件的获取请求打包在同一个HTTP会话中发送,避免多次TCP握手的开销,这对小文件来说能显著提升效率。另外,并行请求也是优化这类场景的常用手段,下面给你具体实现思路和代码示例:

方法1:使用Google API的批量请求功能

pygdrive3底层基于Google官方的google-api-python-client,而这个库支持批量请求功能。我们可以把100个get_media请求打包成一个批量任务执行,大幅减少网络往返的耗时。

示例代码:

from pydrive3.auth import GoogleAuth
from pydrive3.drive import GoogleDrive
from googleapiclient.http import BatchHttpRequest

def handle_download(request_id, response, exception):
    # 处理每个文件的下载结果
    if exception is not None:
        print(f"下载文件[{request_id}]出错: {exception}")
    else:
        # 这里可根据需求将二进制内容保存到本地或做其他处理
        with open(f"downloaded_file_{request_id}", "wb") as f:
            f.write(response)
        print(f"文件[{request_id}]下载完成")

# 完成授权流程
gauth = GoogleAuth()
gauth.LocalWebserverAuth()
drive = GoogleDrive(gauth)

# 假设你已经有100个目标文件的ID列表
target_file_ids = ["file_id_001", "file_id_002", ..., "file_id_100"]

# 创建批量请求对象
batch_request = BatchHttpRequest(callback=handle_download)

# 将所有文件的获取请求添加到批量任务中
for idx, file_id in enumerate(target_file_ids):
    batch_request.add(drive.auth.service.files().get_media(fileId=file_id), request_id=str(idx))

# 执行批量请求
batch_request.execute()

这个方法的核心是复用同一个HTTP连接发送所有请求,避免了多次建立连接的额外开销,比循环逐个请求快很多。

方法2:使用线程池并行下载

小文件的下载瓶颈通常是网络延迟而非CPU性能,所以用多线程并行处理多个下载请求,能充分利用网络带宽,同样能大幅提升整体速度。

示例代码:

from pydrive3.auth import GoogleAuth
from pydrive3.drive import GoogleDrive
from concurrent.futures import ThreadPoolExecutor

def download_single_file(file_id):
    try:
        drive_file = drive.CreateFile({'id': file_id})
        # 若需直接获取二进制内容,用GetContentFile()会自动保存到本地
        drive_file.GetContentFile(f"downloaded_{file_id}")
        print(f"文件[{file_id}]下载完成")
        return True
    except Exception as e:
        print(f"文件[{file_id}]下载失败: {e}")
        return False

# 完成授权流程
gauth = GoogleAuth()
gauth.LocalWebserverAuth()
drive = GoogleDrive(gauth)

target_file_ids = ["file_id_001", "file_id_002", ..., "file_id_100"]

# 线程数建议设为10-20(可根据自身网络情况调整)
with ThreadPoolExecutor(max_workers=15) as executor:
    # 并行执行所有下载任务
    download_results = list(executor.map(download_single_file, target_file_ids))

注意事项

  • 两种方法都会消耗API调用次数,100个文件对应100次API调用,只要不超过你的Google Drive API配额就没问题。
  • 如果文件数量极大(比如上千个),建议把批量请求拆分成多个小批次,或者控制线程池大小,避免触发API的限流机制。

内容的提问来源于stack exchange,提问作者user2194805

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:33:12