You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Colab中从公开Google Drive提取Zip数据集(无需挂载Drive)

问题:从公开Google Drive文件夹直接下载Zip数据集(无需挂载/复制到个人Drive)

我需要从公开Google Drive文件夹下载Zip格式的数据集,文件夹链接为:url = https://drive.google.com/drive/folders/1TzwfNA5JRFTPO-kHMU___kILmOEodoBo
要求其他人可复现操作,不复制到自己Drive,也不挂载Drive,请问该如何实现?


已尝试的两种失败方法

方法一:使用requests直接请求文件链接

执行代码:

import requests
import io
import zipfile

zip_url = 'https://drive.google.com/file/d/1fdFu5NGXe4rTLYKD5wOqk9dl-eJOefXo'

response = requests.get(zip_url)
file_contents = io.BytesIO(response.content)
print(file_contents)
with zipfile.ZipFile(file_contents, 'r') as zip_ref:
    zip_ref.extractall('/content/')  # Replace with your desired extraction path

返回错误:

<_io.BytesIO object at 0x7ad7efbf27f0>
---------------------------------------------------------------------------
BadZipFile                                Traceback (most recent call last)
<ipython-input-18-56d2c8f2bfe8> in <cell line: 14>()
     12 print(file_contents)
     13 # Extract the zip file (if needed)
---> 14 with zipfile.ZipFile(file_contents, 'r') as zip_ref:
     15     zip_ref.extractall('/content/')  # Replace with your desired extraction path

1 frames
/usr/lib/python3.10/zipfile.py in _RealGetContents(self)
   1334             raise BadZipFile("File is not a zip file")
   1335         if not endrec:
-> 1336             raise BadZipFile("File is not a zip file")
   1337         if self.debug > 1:
   1338             print(endrec)

BadZipFile: File is not a zip file

方法二:使用wget请求导出链接

执行代码:

file_id = '1fdFu5NGXe4rTLYKD5wOqk9dl-eJOefXo'
download_url = f'https://drive.google.com/uc?export=download&id={file_id}'
!wget --no-check-certificate -O '/content/file.zip' 'https://drive.google.com/uc?export=download&id=1fdFu5NGXe4rTLYKD5wOqk9dl-eJOefXo'

结果:生成空的Zip文件。


可行解决方案

方案1:使用gdown库(推荐,易复现)

gdown是Google Drive文件下载专用Python库,无需挂载Drive,直接通过文件ID完成下载。

  1. 安装依赖:
!pip install gdown -q
  1. 下载并解压:
import gdown
import zipfile

file_id = '1fdFu5NGXe4rTLYKD5wOqk9dl-eJOefXo'
output_path = '/content/dataset.zip'

# 下载文件
gdown.download(id=file_id, output=output_path, quiet=False)

# 解压到目标目录
with zipfile.ZipFile(output_path, 'r') as zip_ref:
    zip_ref.extractall('/content/')

方案2:改进wget命令,处理Drive下载确认

Google Drive对部分文件会要求下载确认,直接请求会返回空文件,需获取确认token并附加到请求中:

!wget --load-cookies /tmp/cookies.txt "https://docs.google.com/uc?export=download&confirm=$(wget --quiet --save-cookies /tmp/cookies.txt --keep-session-cookies --no-check-certificate 'https://docs.google.com/uc?export=download&id=1fdFu5NGXe4rTLYKD5wOqk9dl-eJOefXo' -O- | sed -rn 's/.*confirm=([0-9A-Za-z_]+).*/\1\n/p')&id=1fdFu5NGXe4rTLYKD5wOqk9dl-eJOefXo" -O /content/dataset.zip && rm -rf /tmp/cookies.txt

解压文件:

!unzip /content/dataset.zip -d /content/

方案3:手动用requests处理确认逻辑

无需第三方库,通过requests会话管理Cookie并提取确认token:

import requests
import zipfile
import re

file_id = '1fdFu5NGXe4rTLYKD5wOqk9dl-eJOefXo'
base_url = f'https://docs.google.com/uc?export=download&id={file_id}'

session = requests.Session()
response = session.get(base_url, stream=True)

# 提取下载确认token
confirm_token = None
for cookie_name, cookie_value in response.cookies.items():
    if cookie_name.startswith('download_warning'):
        confirm_token = cookie_value
        break

# 如果需要确认,追加token到请求URL
if confirm_token:
    base_url += f'&confirm={confirm_token}'
    response = session.get(base_url, stream=True)

# 保存文件到本地
with open('/content/dataset.zip', 'wb') as f:
    for chunk in response.iter_content(chunk_size=1024*1024):
        f.write(chunk)

# 解压文件
with zipfile.ZipFile('/content/dataset.zip', 'r') as zip_ref:
    zip_ref.extractall('/content/')

内容的提问来源于stack exchange,提问作者ArieAI

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 21:17:23