如何用Python通过API下载文件?附现有文件查询代码求助
Hey there! Let's walk through how to download the actual file using the API you're already working with. Your current code successfully fetches the publication metadata, so we just need to extend it to grab the file itself.
Step 1: Understand the API Structure
The metadata you retrieved includes the publication ID (ws859f67m), which we'll use to target the download endpoint. Most REST APIs for publications follow a consistent pattern for download URLs—usually something like /api/v1/publication/{publication_id}/download.
Step 2: Update Your Code to Handle Downloads
Here's a modified version of your code that adds download functionality, including best practices like streaming large files and error handling:
import requests import pprint # 1. Authenticate and get your Bearer token token_response = requests.post( "https://usda.library.cornell.edu/user_token", json={'auth': {'email': 'my_mail', 'password':'my_pass'}} ) bearer_token = token_response.json()['jwt'] headers = {'Authorization': f'Bearer {bearer_token}'} # 2. Fetch publication metadata (your existing code) metadata_url = 'https://usda.library.cornell.edu/api/v1/publication/findById/ws859f67m' metadata_response = requests.get(metadata_url, headers=headers) metadata = metadata_response.json() pprint.pprint(metadata) # 3. Build the download URL and save the file publication_id = metadata[0]['id'] download_url = f'https://usda.library.cornell.edu/api/v1/publication/{publication_id}/download' # Stream the download to handle large files efficiently with requests.get(download_url, headers=headers, stream=True) as download_response: # Raise an error if the request fails (e.g., 401 unauthorized, 404 not found) download_response.raise_for_status() # Get a meaningful filename: either from the API's response header or metadata filename = None content_disposition = download_response.headers.get('Content-Disposition') if content_disposition and 'filename=' in content_disposition: filename = content_disposition.split('filename=')[1].strip('"') else: # Fallback to using the publication title from metadata filename = f"{metadata[0]['title'][0].replace('/', '_')}.pdf" # Write the file to disk with open(filename, 'wb') as file: for chunk in download_response.iter_content(chunk_size=8192): file.write(chunk) print(f"Successfully downloaded: {filename}")
Key Notes:
- Streaming Downloads: Using
stream=Trueanditer_content()ensures we don't load the entire file into memory at once—critical for large reports. - Error Handling:
raise_for_status()will alert you if there's an issue with the request (like invalid credentials or a missing file). - Filename Logic: The code first tries to extract the official filename from the
Content-Dispositionheader sent by the API. If that's missing, it uses the publication title (with slashes replaced to avoid OS errors) and assumes a PDF extension—adjust the extension if you know the file is a different type (like CSV).
If the Download URL Doesn't Work:
If the /download endpoint returns an error, try these common alternative URLs:
https://usda.library.cornell.edu/api/v1/publication/{publication_id}/filehttps://usda.library.cornell.edu/api/v1/publication/{publication_id}/resource
You can also double-check the metadata for any fields that might contain a direct file URL (like file_url or download_link—though your current metadata output doesn't show these).
内容的提问来源于stack exchange,提问作者Alex Riabukha

