Python3指定链接批量下载HDF文件及认证问题求助
Hey Lucas, let's work through your MODIS HDF download problems step by step—you're close, just a few tweaks needed to get past the 401 error and automate the process properly.
First: Fix the 401 Authentication Error
Your original script used a POST request to download the file, which is incorrect—file downloads rely on GET requests. Plus, using requests.Session() will persist your authentication across requests, making the process more reliable than passing credentials each time. LP DAAC servers accept Earthdata Login credentials via HTTP Basic Auth, so here's a corrected authentication flow:
import requests from bs4 import BeautifulSoup import os # Initialize a session to keep auth state across requests session = requests.Session() # Replace with your Earthdata Login credentials session.auth = ("lucas", "xxxx") # Test access to a target directory target_dir_url = "https://e4ftl01.cr.usgs.gov/MOLT/MOD09A1.005/2013.08.21/" response = session.get(target_dir_url) # This will throw an error if authentication fails (helps catch issues early) response.raise_for_status()
Second: Automate Variable Date/File Matching
MOD09A1 files are organized by date in YYYY.MM.DD directories, and filenames follow the pattern MOD09A1.<variable_part>.h13v12.005.<variable_part>.hdf. To target your specific tile (h13v12) and generate the date ranges you need, use these steps:
Generate Date Ranges: Use the
datetimemodule to create a list of dates (e.g., 3 per month) across your target years:from datetime import datetime, timedelta def get_monthly_dates(start_year, end_year): """Generate ~3 dates per month across the target years""" dates = [] for year in range(start_year, end_year + 1): for month in range(1, 13): # Pick 1st, 11th, 21st of each month (adjust as needed) for day in [1, 11, 21]: try: date = datetime(year, month, day) dates.append(date.strftime("%Y.%m.%d")) except ValueError: # Skip invalid dates (like Feb 31) continue return dates # Example: Get dates from 2013 to 2015 date_list = get_monthly_dates(2013, 2015)Filter Target Files: When parsing the directory HTML, filter links to only include your tile and valid .hdf files:
def get_hdf_links(session, dir_url): response = session.get(dir_url) response.raise_for_status() soup = BeautifulSoup(response.text, "lxml") # Filter for links ending with .hdf and containing your tile (h13v12) return [a["href"] for a in soup.find_all("a") if a["href"].endswith(".hdf") and "h13v12" in a["href"]]
Third: Efficiently Download HDF Files
Use streaming to handle large HDF files without loading them entirely into memory, and optionally add multithreading to speed up downloads (be mindful of server rate limits—don't set too many workers):
from concurrent.futures import ThreadPoolExecutor def download_hdf(session, file_url, save_path): """Stream-download a single HDF file, skipping already downloaded files""" if os.path.exists(save_path): print(f"Skipping {os.path.basename(save_path)} (already exists)") return try: with session.get(file_url, stream=True) as r: r.raise_for_status() with open(save_path, "wb") as fd: # Chunk size balances speed and memory usage for chunk in r.iter_content(chunk_size=8192): fd.write(chunk) print(f"Downloaded: {os.path.basename(save_path)}") except Exception as e: print(f"Failed to download {file_url}: {str(e)}") # Main workflow save_directory = "./modis_hdf_downloads" os.makedirs(save_directory, exist_ok=True) # Use multithreading (adjust max_workers based on your bandwidth) with ThreadPoolExecutor(max_workers=3) as executor: for date_str in date_list: dir_url = f"https://e4ftl01.cr.usgs.gov/MOLT/MOD09A1.005/{date_str}/" print(f"Processing directory: {dir_url}") hdf_links = get_hdf_links(session, dir_url) for link in hdf_links: file_url = dir_url + link save_path = os.path.join(save_directory, link) executor.submit(download_hdf, session, file_url, save_path)
Key Fixes from Your Original Script
- Swapped
POSTforGETwhen downloading files (POST is for submitting data, not retrieving files) - Used
requests.Session()to persist authentication, which resolves the 401 error - Added filtering to target only your desired
h13v12tile files - Automated date range generation to avoid manual URL entry
- Added streaming and optional multithreading for faster, more efficient downloads
内容的提问来源于stack exchange,提问作者Lucas

