You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何直接将网站Zip文件读取至Jupyter Notebook?Kaggle数据集解压报错求助

Fixing "End-of-central-directory signature not found" When Unzipping Kaggle Dataset

Hey, I’ve dealt with this exact headache before—let’s break down what’s going wrong and how to fix it.

Why You’re Getting This Error

That error means the file you downloaded (train.csv.zip) isn’t actually a valid zip archive. Chances are, when you used wget on the Kaggle dataset page URL, you didn’t authenticate with Kaggle, so it just downloaded the login page HTML instead of the real zip file. That’s why unzip can’t recognize it as a valid archive.

Step 1: Download the Real Dataset (Two Reliable Ways)

Option 1: Use the Kaggle API (Recommended for Notebooks)

This is the easiest method for data science notebooks (like Colab, AWS SageMaker, etc.):

  1. Go to your Kaggle account page, scroll to the API section, and click "Create New API Token". This will download a kaggle.json file with your credentials.
  2. Upload kaggle.json to your notebook environment, or move it to the ~/.kaggle/ directory (run this command to set permissions correctly):
    mkdir -p ~/.kaggle && mv kaggle.json ~/.kaggle/ && chmod 600 ~/.kaggle/kaggle.json
    
  3. Run this command to download the exact file you need:
    kaggle competitions download -c talkingdata-adtracking-fraud-detection -f train.csv.zip
    
    The API handles authentication automatically, so you’ll get the real zip file every time.

Option 2: Use an Authenticated Download Link

If you don’t want to set up the API:

  1. Log into Kaggle, navigate to the TalkingData competition data page.
  2. Right-click the "Download" button next to train.csv.zip and select "Copy link address". This link includes a temporary auth token (valid for a few hours).
  3. Use wget to download with this link (wrap it in quotes to handle special characters):
    wget "PASTE_YOUR_AUTHENTICATED_LINK_HERE" -O train.csv.zip
    

Step 2: Verify the File Before Unzipping

To make sure you have a valid zip file, run:

file train.csv.zip

If the output says Zip archive data, you’re good to go. If it says HTML document, repeat the download step—you still have the login page.

Step 3: Unzip (Or Skip Unzipping Entirely!)

If you need to unzip it:

unzip train.csv.zip

But since the file is huge (the uncompressed train.csv is several GB), you can save space and time by reading the zip directly into pandas without unzipping:

import pandas as pd
# Read directly from the zip file
df = pd.read_csv('train.csv.zip', compression='zip')

That should get you up and running with the dataset in your notebook without local storage issues!

内容的提问来源于stack exchange,提问作者madsthaks

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:32:13