使用Python下载Google Drive大文件夹遇限制的替代方案求解
大容量Google Drive文件夹下载解决方案
方案1:修改gdown参数快速适配
50文件限制可先通过调整gdown原生参数解决,修改代码如下:
import gdown url = 'https://drive.google.com/drive/folders/135hTTURfjn43fo4f?usp=sharing' gdown.download_folder( url, remaining_ok=True, # 跳过50文件数量限制的报错 use_cookies=True, # 复用浏览器Google Drive登录态,绕过临时配额限制 quiet=False )
如果触发下载配额限制,只需要在浏览器登录Google账号访问一次目标文件夹,gdown会自动读取本地浏览器的cookie完成鉴权,无需额外配置。
方案2:Google Drive官方API稳定下载(推荐5万文件场景使用)
该方案没有文件数量限制,自带断点续传能力,稳定性最高:
前置准备
- 到Google Cloud Console创建项目,启用Google Drive API,生成OAuth 2.0客户端密钥,保存为
credentials.json放在工作目录 - 安装依赖包:
pip install google-api-python-client google-auth-httplib2 google-auth-oauthlib
实现代码
import os import io from google.auth.transport.requests import Request from google.oauth2.credentials import Credentials from google_auth_oauthlib.flow import InstalledAppFlow from googleapiclient.discovery import build from googleapiclient.errors import HttpError from googleapiclient.http import MediaIoBaseDownload # 权限范围,修改后需删除本地旧token.json重新授权 SCOPES = ['https://www.googleapis.com/auth/drive.readonly'] # 替换为你的目标文件夹ID,从分享URL中提取 TARGET_FOLDER_ID = '135hTTURfjn43fo4f' # 本地保存目录 LOCAL_SAVE_PATH = './drive_download_images' os.makedirs(LOCAL_SAVE_PATH, exist_ok=True) def main(): creds = None # 复用已有授权token,避免每次都要登录 if os.path.exists('token.json'): creds = Credentials.from_authorized_user_file('token.json', SCOPES) if not creds or not creds.valid: if creds and creds.expired and creds.refresh_token: creds.refresh(Request()) else: flow = InstalledAppFlow.from_client_secrets_file('credentials.json', SCOPES) creds = flow.run_local_server(port=0) # 保存授权token供后续使用 with open('token.json', 'w') as token: token.write(creds.to_json()) try: service = build('drive', 'v3', credentials=creds) page_token = None # 分页拉取所有文件,突破单页数量限制 while True: resp = service.files().list( q=f"'{TARGET_FOLDER_ID}' in parents and trashed=false", spaces='drive', fields='nextPageToken, files(id, name)', pageToken=page_token, pageSize=1000 ).execute() for file in resp.get('files', []): save_file_path = os.path.join(LOCAL_SAVE_PATH, file.get('name')) # 跳过已下载文件,实现断点续传 if os.path.exists(save_file_path): continue file_req = service.files().get_media(fileId=file.get('id')) fh = io.FileIO(save_file_path, 'wb') downloader = MediaIoBaseDownload(fh, file_req, chunksize=10*1024*1024) done = False while not done: status, done = downloader.next_chunk() print(f"文件{file.get('name')}下载进度:{int(status.progress()*100)}%") page_token = resp.get('nextPageToken', None) if not page_token: break except HttpError as e: print(f"接口错误:{e}") if __name__ == '__main__': main()
首次运行会自动弹出浏览器授权页面,登录拥有文件夹访问权限的Google账号即可,后续运行无需重复授权。
方案3:无代码rclone工具下载
如果不需要自定义开发,可以直接使用rclone命令行工具完成下载:
- 安装rclone后执行
rclone config按照提示完成Google Drive挂载配置,命名为mygdrive - 执行下载命令:
rclone copy mygdrive:目标文件夹路径 ./local_save_path --transfers 8 --checkers 15 - 可根据服务器带宽调整
transfers参数控制同时下载的文件数,工具自动支持断点续传、错误重试。
注意事项
- 5万张图片总大小如果超过100G,建议拆分批次下载,避免触发Google Drive单日流量配额限制
- 如果遇到403配额错误,暂停24小时后重新运行即可,或者切换其他Google账号授权恢复下载
内容的提问来源于stack exchange,提问作者Vahid the Great
相关产品推荐
相关产品推荐

