如何在Kaggle环境中访问Google Drive中的图像数据集文件夹?
解决Google Drive 5GB数据集在Kaggle中访问的问题
以下是几种无需遍历文件链接、避开Cookie问题的可行方案:
1. 直接挂载Google Drive到Kaggle
Kaggle内置了Google Drive集成,操作简单且稳定:
- 打开你的Kaggle Notebook,在右侧边栏找到Add-ons > Google Drive,点击Connect。
- 按照弹窗提示完成Google账号授权,授权成功后,你的Drive会被挂载到
/kaggle/input/google-drive/路径。 - 之后直接通过文件路径读取数据集即可,示例代码:
import os dataset_dir = "/kaggle/input/google-drive/你的数据集文件夹路径" for img_name in os.listdir(dataset_dir): # 处理图像逻辑
2. 打包文件夹为Zip后下载
把整个数据集文件夹压缩成单个Zip文件,避免处理多个文件链接:
- 在Google Drive中右键点击目标文件夹,选择压缩为ZIP文件(5GB文件夹压缩后体积变化不大,确保Drive有足够空间)。
- 将Zip文件设置为公开共享(或共享给你的Kaggle关联邮箱),复制其共享链接。
- 在Kaggle Notebook中用
wget下载并解压:!wget -O dataset.zip "你的Zip文件共享链接" !unzip dataset.zip -d /kaggle/working/dataset/
3. 用Google Drive服务账号无交互访问
如果需要自动化访问(比如批量运行Notebook),可以用服务账号:
- 登录Google Cloud Console,创建一个服务账号,下载对应的JSON密钥文件。
- 将JSON密钥上传到Kaggle Notebook(或存入Notebook的Secrets中)。
- 安装依赖库并编写代码访问Drive:
!pip install google-api-python-client google-auth-httplib2 google-auth-oauthlib from google.oauth2 import service_account from googleapiclient.discovery import build from googleapiclient.http import MediaIoBaseDownload import io import os # 加载服务账号密钥 creds = service_account.Credentials.from_service_account_file('service-account-key.json') drive_service = build('drive', 'v3', credentials=creds) # 指定目标文件夹ID(从Drive文件夹URL中提取) folder_id = "你的文件夹ID" # 获取文件夹内所有文件 files = drive_service.files().list(q=f"'{folder_id}' in parents", fields="files(id, name)").execute().get('files', []) # 创建本地存储目录 os.makedirs('/kaggle/working/dataset', exist_ok=True) # 批量下载文件 for file in files: request = drive_service.files().get_media(fileId=file['id']) fh = io.FileIO(f"/kaggle/working/dataset/{file['name']}", 'wb') downloader = MediaIoBaseDownload(fh, request) done = False while not done: status, done = downloader.next_chunk()
注意要把目标文件夹共享给服务账号的邮箱地址,确保权限正确。
内容的提问来源于stack exchange,提问作者harshmangalamv
相关产品推荐
相关产品推荐

