You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Google Colab中提取Google Drive的2K图像数据集压缩包用于ML训练

从Google Drive提取压缩数据集用于机器学习训练

嘿,我帮你把代码补全并整理好了,刚好能解决你从Google Drive提取2000张图像的压缩包、用于ML训练的需求:

from pydrive.auth import GoogleAuth
from pydrive.drive import GoogleDrive
from google.colab import auth
from oauth2client.client import GoogleCredentials
import zipfile
import os

# 认证并创建PyDrive客户端(每个Notebook只需执行一次)
auth.authenticate_user()
gauth = GoogleAuth()
gauth.credentials = GoogleCredentials.get_application_default()
drive = GoogleDrive(gauth)

# 👇 替换成你Google Drive中压缩包的文件ID
# 获取方式:打开压缩包的共享链接,链接格式是https://drive.google.com/file/d/[这里就是文件ID]/view
file_id = "你的压缩包文件ID"

# 把Drive上的压缩包下载到Colab当前工作目录
downloaded = drive.CreateFile({'id': file_id})
downloaded.GetContentFile('dataset.zip')

# 提取压缩包内容到指定文件夹(这里命名为ml_dataset,你可以按需修改)
with zipfile.ZipFile('dataset.zip', 'r') as zip_ref:
    zip_ref.extractall('ml_dataset')

# 可选:验证提取结果,确认文件数量是否符合预期
print(f"提取完成!当前ml_dataset文件夹中有 {len(os.listdir('ml_dataset'))} 个文件")

几个关键注意点:

  • 文件ID别填错:这是最容易踩坑的地方,一定要把file_id替换成你自己压缩包的ID,不然代码会找不到目标文件。
  • 授权步骤:运行auth.authenticate_user()时,会弹出一个授权窗口,跟着提示完成登录授权就行,这一步是为了让Colab能访问你的Drive文件。
  • 提取路径:extractall('ml_dataset')会把所有图像解压到这个文件夹里,之后你训练模型的时候,直接从这个路径读取数据就可以了(比如用TensorFlow的ImageDataGenerator或者PyTorch的ImageFolder)。

如果你的压缩包有多层目录结构,提取后会完全保留原结构,不会影响后续的数据集加载~

内容的提问来源于stack exchange,提问作者Laxmikant

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:14:14