如何用Pandas读取URL指向的ZIP包内指定CSV文件?
How to Load
Summary_stats_all_locs.csv from the ZIP Archive into a Pandas DataFrame Great question! Since your target CSV is nested inside a ZIP file at that URL, you can’t use pd.read_csv() directly—Pandas can’t parse ZIP archives out of the box for remote URLs. Instead, you’ll need to fetch the ZIP, extract the specific CSV in memory, and then load it into a DataFrame. Here’s a clean, efficient way to do this without saving any files to your local system:
Step-by-Step Solution
First, import the required libraries:
import requests import zipfile import io import pandas as pd
Then, use this code to fetch the ZIP, extract the CSV, and load it:
# Define the URL to the daily-updated ZIP archive zip_url = "https://ihmecovid19storage.blob.core.windows.net/latest/ihme-covid19.zip" try: # Fetch the ZIP file from the URL response = requests.get(zip_url) response.raise_for_status() # Trigger an error if the request fails (e.g., 404, 500) # Convert the downloaded ZIP content into an in-memory file object zip_in_memory = zipfile.ZipFile(io.BytesIO(response.content)) # Define the target CSV filename we want to extract target_csv = "Summary_stats_all_locs.csv" # Make sure the CSV exists in the ZIP archive if target_csv not in zip_in_memory.namelist(): raise FileNotFoundError(f"Error: {target_csv} was not found in the ZIP archive") # Open the CSV directly from the ZIP and load it into a Pandas DataFrame with zip_in_memory.open(target_csv) as csv_file: covid_df = pd.read_csv(csv_file) # Verify the DataFrame loaded correctly (optional) print("Success! Here's a preview of the data:") print(covid_df.head()) except requests.exceptions.RequestException as e: print(f"Failed to download the ZIP file: {str(e)}") except zipfile.BadZipFile: print("Error: The downloaded file is not a valid ZIP archive") except FileNotFoundError as e: print(str(e))
Key Details Explained
requests.get(): Fetches the ZIP file from the remote URL.response.raise_for_status()ensures we catch any HTTP errors (like broken links or server issues).io.BytesIO(): Treats the downloaded ZIP content as an in-memory file, so we don’t need to save it to your hard drive first.zipfile.ZipFile(): Lets us interact with the ZIP archive without extracting all files. We only access the specific CSV we need, which saves memory and time.- Error Handling: The try/except blocks catch common issues like failed downloads, invalid ZIP files, or missing CSV files—making the script more robust for daily use.
This approach works perfectly for the daily-updated ZIP, as it always pulls the latest version each time you run the code.
内容的提问来源于stack exchange,提问作者kspr
相关产品推荐
相关产品推荐

