如何直接将网站Zip文件读取至Jupyter Notebook?Kaggle数据集解压报错求助
Hey, I’ve dealt with this exact headache before—let’s break down what’s going wrong and how to fix it.
Why You’re Getting This Error
That error means the file you downloaded (train.csv.zip) isn’t actually a valid zip archive. Chances are, when you used wget on the Kaggle dataset page URL, you didn’t authenticate with Kaggle, so it just downloaded the login page HTML instead of the real zip file. That’s why unzip can’t recognize it as a valid archive.
Step 1: Download the Real Dataset (Two Reliable Ways)
Option 1: Use the Kaggle API (Recommended for Notebooks)
This is the easiest method for data science notebooks (like Colab, AWS SageMaker, etc.):
- Go to your Kaggle account page, scroll to the API section, and click "Create New API Token". This will download a
kaggle.jsonfile with your credentials. - Upload
kaggle.jsonto your notebook environment, or move it to the~/.kaggle/directory (run this command to set permissions correctly):mkdir -p ~/.kaggle && mv kaggle.json ~/.kaggle/ && chmod 600 ~/.kaggle/kaggle.json - Run this command to download the exact file you need:
The API handles authentication automatically, so you’ll get the real zip file every time.kaggle competitions download -c talkingdata-adtracking-fraud-detection -f train.csv.zip
Option 2: Use an Authenticated Download Link
If you don’t want to set up the API:
- Log into Kaggle, navigate to the TalkingData competition data page.
- Right-click the "Download" button next to
train.csv.zipand select "Copy link address". This link includes a temporary auth token (valid for a few hours). - Use
wgetto download with this link (wrap it in quotes to handle special characters):wget "PASTE_YOUR_AUTHENTICATED_LINK_HERE" -O train.csv.zip
Step 2: Verify the File Before Unzipping
To make sure you have a valid zip file, run:
file train.csv.zip
If the output says Zip archive data, you’re good to go. If it says HTML document, repeat the download step—you still have the login page.
Step 3: Unzip (Or Skip Unzipping Entirely!)
If you need to unzip it:
unzip train.csv.zip
But since the file is huge (the uncompressed train.csv is several GB), you can save space and time by reading the zip directly into pandas without unzipping:
import pandas as pd # Read directly from the zip file df = pd.read_csv('train.csv.zip', compression='zip')
That should get you up and running with the dataset in your notebook without local storage issues!
内容的提问来源于stack exchange,提问作者madsthaks

