使用Python Requests无法获取带Content-Disposition: attachment的文件
Let’s walk through this problem and figure out why your Python Requests script isn’t triggering the CSV download like Firefox does. I’ll break down the likely issues and give you actionable fixes.
Key Observations from Your Setup
First, let’s recap what’s working vs. not:
- Firefox successfully triggers the
Opening report1.csvdialog because the server returns aContent-Disposition: attachment; filename="report1.csv"header. - Your Requests script logs in successfully (200 status) but gets redirected back to the starting page instead of the download response—no
Content-Dispositionheader in sight.
Likely Causes & Fixes
1. Session Cookie Management (Yes, Requests Handles It—But Maybe Not Correctly)
Requests’ Session does automatically manage cookies, but there are a few gotchas:
- Verify login actually succeeded: A 200 status code doesn’t always mean you’re authenticated. Check if the response content includes a welcome message or use
login_response.historyto see if you were redirected to a dashboard. - Compare cookies with Firefox: After logging in, print
session.cookiesand cross-reference with Firefox’s Developer Tools (Application → Cookies). If critical cookies are missing (like a session ID or auth token), your login request didn’t capture them—maybe the login form requires hidden fields (like a CSRF token) you didn’t include?
2. Your Request Doesn’t Match Firefox’s Request
Servers often block or alter responses if the request doesn’t match what a typical browser sends. Focus on these headers:
- User-Agent: Many servers reject requests with Requests’ default UA string. Copy Firefox’s UA from the Network tab (Request Headers → User-Agent) and add it to your requests.
- Referer: Some servers check that the download request comes from the query page (not a direct URL). Include the URL of the page where you clicked the "Query" button as the
Refererheader. - Accept: Match Firefox’s Accept header to signal you’re expecting a CSV file.
3. You Might Be Using the Wrong Request Method
When you click the "Query" button in Firefox, is it sending a GET or POST request? Live HTTP Headers will show this. If it’s a POST, your script’s GET request to the download URL won’t work—you need to replicate the POST to the query endpoint first, which then triggers the download redirect.
4. Missing CSRF Token
Most modern sites require a CSRF token for form submissions (including login and query actions). You need to extract this token from the login/query page’s HTML and include it in your POST data.
Example Fixed Script
Here’s a revised script that addresses these issues:
import requests from bs4 import BeautifulSoup # Initialize session session = requests.Session() session.verify = True # Keep this unless you have a specific reason to disable # Step 1: Fetch login page to get CSRF token and set initial cookies login_url = "https://myserver/login" login_page = session.get(login_url) soup = BeautifulSoup(login_page.text, "html.parser") csrf_token = soup.find("input", {"name": "csrfmiddlewaretoken"})["value"] # Adjust name to match your site # Step 2: Authenticate with correct headers and token login_payload = { "username": "your_username", "password": "your_password", "csrfmiddlewaretoken": csrf_token } browser_headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:117.0) Gecko/20100101 Firefox/117.0", "Referer": login_url } login_response = session.post(login_url, data=login_payload, headers=browser_headers) # Verify login success (adjust check to match your site) if "Welcome" not in login_response.text: print("Login failed—check credentials or CSRF token") exit() # Step 3: Trigger download with matching headers and session download_url = "https://myserver/ReportPage?download&NAME=ALL&DATE=THISYEAR" download_headers = { "User-Agent": browser_headers["User-Agent"], "Referer": "https://myserver/query-page-url" # Replace with the page where you clicked "Query" } # Use stream=True for large files download_response = session.get(download_url, headers=download_headers, stream=True) # Check for Content-Disposition header if "Content-Disposition" in download_response.headers: # Extract filename from header disp_header = download_response.headers["Content-Disposition"] filename = disp_header.split('filename="')[1].split('"')[0] # Save the file with open(filename, "wb") as f: for chunk in download_response.iter_content(chunk_size=8192): f.write(chunk) print(f"Successfully downloaded {filename}") else: print("No download response received—check request headers or session") # Print first 500 chars of response to debug what page you're getting print(download_response.text[:500])
Final Checks
- Double-check the download URL from Live HTTP Headers—make sure you’re using the exact same parameters and URL structure.
- If the query button uses
POST, replace theGETrequest todownload_urlwith aPOSTto the query endpoint, including all form parameters from the query page.
内容的提问来源于stack exchange,提问作者Ingo

