You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python3指定链接批量下载HDF文件及认证问题求助

Hey Lucas, let's work through your MODIS HDF download problems step by step—you're close, just a few tweaks needed to get past the 401 error and automate the process properly.

First: Fix the 401 Authentication Error

Your original script used a POST request to download the file, which is incorrect—file downloads rely on GET requests. Plus, using requests.Session() will persist your authentication across requests, making the process more reliable than passing credentials each time. LP DAAC servers accept Earthdata Login credentials via HTTP Basic Auth, so here's a corrected authentication flow:

import requests
from bs4 import BeautifulSoup
import os

# Initialize a session to keep auth state across requests
session = requests.Session()
# Replace with your Earthdata Login credentials
session.auth = ("lucas", "xxxx")

# Test access to a target directory
target_dir_url = "https://e4ftl01.cr.usgs.gov/MOLT/MOD09A1.005/2013.08.21/"
response = session.get(target_dir_url)
# This will throw an error if authentication fails (helps catch issues early)
response.raise_for_status()

Second: Automate Variable Date/File Matching

MOD09A1 files are organized by date in YYYY.MM.DD directories, and filenames follow the pattern MOD09A1.<variable_part>.h13v12.005.<variable_part>.hdf. To target your specific tile (h13v12) and generate the date ranges you need, use these steps:

  1. Generate Date Ranges: Use the datetime module to create a list of dates (e.g., 3 per month) across your target years:

    from datetime import datetime, timedelta
    
    def get_monthly_dates(start_year, end_year):
        """Generate ~3 dates per month across the target years"""
        dates = []
        for year in range(start_year, end_year + 1):
            for month in range(1, 13):
                # Pick 1st, 11th, 21st of each month (adjust as needed)
                for day in [1, 11, 21]:
                    try:
                        date = datetime(year, month, day)
                        dates.append(date.strftime("%Y.%m.%d"))
                    except ValueError:
                        # Skip invalid dates (like Feb 31)
                        continue
        return dates
    
    # Example: Get dates from 2013 to 2015
    date_list = get_monthly_dates(2013, 2015)
    
  2. Filter Target Files: When parsing the directory HTML, filter links to only include your tile and valid .hdf files:

    def get_hdf_links(session, dir_url):
        response = session.get(dir_url)
        response.raise_for_status()
        soup = BeautifulSoup(response.text, "lxml")
        # Filter for links ending with .hdf and containing your tile (h13v12)
        return [a["href"] for a in soup.find_all("a") if a["href"].endswith(".hdf") and "h13v12" in a["href"]]
    

Third: Efficiently Download HDF Files

Use streaming to handle large HDF files without loading them entirely into memory, and optionally add multithreading to speed up downloads (be mindful of server rate limits—don't set too many workers):

from concurrent.futures import ThreadPoolExecutor

def download_hdf(session, file_url, save_path):
    """Stream-download a single HDF file, skipping already downloaded files"""
    if os.path.exists(save_path):
        print(f"Skipping {os.path.basename(save_path)} (already exists)")
        return
    try:
        with session.get(file_url, stream=True) as r:
            r.raise_for_status()
            with open(save_path, "wb") as fd:
                # Chunk size balances speed and memory usage
                for chunk in r.iter_content(chunk_size=8192):
                    fd.write(chunk)
        print(f"Downloaded: {os.path.basename(save_path)}")
    except Exception as e:
        print(f"Failed to download {file_url}: {str(e)}")

# Main workflow
save_directory = "./modis_hdf_downloads"
os.makedirs(save_directory, exist_ok=True)

# Use multithreading (adjust max_workers based on your bandwidth)
with ThreadPoolExecutor(max_workers=3) as executor:
    for date_str in date_list:
        dir_url = f"https://e4ftl01.cr.usgs.gov/MOLT/MOD09A1.005/{date_str}/"
        print(f"Processing directory: {dir_url}")
        hdf_links = get_hdf_links(session, dir_url)
        for link in hdf_links:
            file_url = dir_url + link
            save_path = os.path.join(save_directory, link)
            executor.submit(download_hdf, session, file_url, save_path)

Key Fixes from Your Original Script

  • Swapped POST for GET when downloading files (POST is for submitting data, not retrieving files)
  • Used requests.Session() to persist authentication, which resolves the 401 error
  • Added filtering to target only your desired h13v12 tile files
  • Automated date range generation to avoid manual URL entry
  • Added streaming and optional multithreading for faster, more efficient downloads

内容的提问来源于stack exchange,提问作者Lucas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:44:39