You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas读取URL指向的ZIP包内指定CSV文件?

How to Load Summary_stats_all_locs.csv from the ZIP Archive into a Pandas DataFrame

Great question! Since your target CSV is nested inside a ZIP file at that URL, you can’t use pd.read_csv() directly—Pandas can’t parse ZIP archives out of the box for remote URLs. Instead, you’ll need to fetch the ZIP, extract the specific CSV in memory, and then load it into a DataFrame. Here’s a clean, efficient way to do this without saving any files to your local system:

Step-by-Step Solution

First, import the required libraries:

import requests
import zipfile
import io
import pandas as pd

Then, use this code to fetch the ZIP, extract the CSV, and load it:

# Define the URL to the daily-updated ZIP archive
zip_url = "https://ihmecovid19storage.blob.core.windows.net/latest/ihme-covid19.zip"

try:
    # Fetch the ZIP file from the URL
    response = requests.get(zip_url)
    response.raise_for_status()  # Trigger an error if the request fails (e.g., 404, 500)

    # Convert the downloaded ZIP content into an in-memory file object
    zip_in_memory = zipfile.ZipFile(io.BytesIO(response.content))

    # Define the target CSV filename we want to extract
    target_csv = "Summary_stats_all_locs.csv"

    # Make sure the CSV exists in the ZIP archive
    if target_csv not in zip_in_memory.namelist():
        raise FileNotFoundError(f"Error: {target_csv} was not found in the ZIP archive")

    # Open the CSV directly from the ZIP and load it into a Pandas DataFrame
    with zip_in_memory.open(target_csv) as csv_file:
        covid_df = pd.read_csv(csv_file)

    # Verify the DataFrame loaded correctly (optional)
    print("Success! Here's a preview of the data:")
    print(covid_df.head())

except requests.exceptions.RequestException as e:
    print(f"Failed to download the ZIP file: {str(e)}")
except zipfile.BadZipFile:
    print("Error: The downloaded file is not a valid ZIP archive")
except FileNotFoundError as e:
    print(str(e))

Key Details Explained

  • requests.get(): Fetches the ZIP file from the remote URL. response.raise_for_status() ensures we catch any HTTP errors (like broken links or server issues).
  • io.BytesIO(): Treats the downloaded ZIP content as an in-memory file, so we don’t need to save it to your hard drive first.
  • zipfile.ZipFile(): Lets us interact with the ZIP archive without extracting all files. We only access the specific CSV we need, which saves memory and time.
  • Error Handling: The try/except blocks catch common issues like failed downloads, invalid ZIP files, or missing CSV files—making the script more robust for daily use.

This approach works perfectly for the daily-updated ZIP, as it always pulls the latest version each time you run the code.

内容的提问来源于stack exchange,提问作者kspr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 17:22:58