You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Requests无法获取带Content-Disposition: attachment的文件

Troubleshooting CSV Download with Python Requests

Let’s walk through this problem and figure out why your Python Requests script isn’t triggering the CSV download like Firefox does. I’ll break down the likely issues and give you actionable fixes.

Key Observations from Your Setup

First, let’s recap what’s working vs. not:

  • Firefox successfully triggers the Opening report1.csv dialog because the server returns a Content-Disposition: attachment; filename="report1.csv" header.
  • Your Requests script logs in successfully (200 status) but gets redirected back to the starting page instead of the download response—no Content-Disposition header in sight.

Likely Causes & Fixes

Requests’ Session does automatically manage cookies, but there are a few gotchas:

  • Verify login actually succeeded: A 200 status code doesn’t always mean you’re authenticated. Check if the response content includes a welcome message or use login_response.history to see if you were redirected to a dashboard.
  • Compare cookies with Firefox: After logging in, print session.cookies and cross-reference with Firefox’s Developer Tools (Application → Cookies). If critical cookies are missing (like a session ID or auth token), your login request didn’t capture them—maybe the login form requires hidden fields (like a CSRF token) you didn’t include?

2. Your Request Doesn’t Match Firefox’s Request

Servers often block or alter responses if the request doesn’t match what a typical browser sends. Focus on these headers:

  • User-Agent: Many servers reject requests with Requests’ default UA string. Copy Firefox’s UA from the Network tab (Request Headers → User-Agent) and add it to your requests.
  • Referer: Some servers check that the download request comes from the query page (not a direct URL). Include the URL of the page where you clicked the "Query" button as the Referer header.
  • Accept: Match Firefox’s Accept header to signal you’re expecting a CSV file.

3. You Might Be Using the Wrong Request Method

When you click the "Query" button in Firefox, is it sending a GET or POST request? Live HTTP Headers will show this. If it’s a POST, your script’s GET request to the download URL won’t work—you need to replicate the POST to the query endpoint first, which then triggers the download redirect.

4. Missing CSRF Token

Most modern sites require a CSRF token for form submissions (including login and query actions). You need to extract this token from the login/query page’s HTML and include it in your POST data.

Example Fixed Script

Here’s a revised script that addresses these issues:

import requests
from bs4 import BeautifulSoup

# Initialize session
session = requests.Session()
session.verify = True  # Keep this unless you have a specific reason to disable

# Step 1: Fetch login page to get CSRF token and set initial cookies
login_url = "https://myserver/login"
login_page = session.get(login_url)
soup = BeautifulSoup(login_page.text, "html.parser")
csrf_token = soup.find("input", {"name": "csrfmiddlewaretoken"})["value"]  # Adjust name to match your site

# Step 2: Authenticate with correct headers and token
login_payload = {
    "username": "your_username",
    "password": "your_password",
    "csrfmiddlewaretoken": csrf_token
}

browser_headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:117.0) Gecko/20100101 Firefox/117.0",
    "Referer": login_url
}

login_response = session.post(login_url, data=login_payload, headers=browser_headers)

# Verify login success (adjust check to match your site)
if "Welcome" not in login_response.text:
    print("Login failed—check credentials or CSRF token")
    exit()

# Step 3: Trigger download with matching headers and session
download_url = "https://myserver/ReportPage?download&NAME=ALL&DATE=THISYEAR"
download_headers = {
    "User-Agent": browser_headers["User-Agent"],
    "Referer": "https://myserver/query-page-url"  # Replace with the page where you clicked "Query"
}

# Use stream=True for large files
download_response = session.get(download_url, headers=download_headers, stream=True)

# Check for Content-Disposition header
if "Content-Disposition" in download_response.headers:
    # Extract filename from header
    disp_header = download_response.headers["Content-Disposition"]
    filename = disp_header.split('filename="')[1].split('"')[0]
    
    # Save the file
    with open(filename, "wb") as f:
        for chunk in download_response.iter_content(chunk_size=8192):
            f.write(chunk)
    print(f"Successfully downloaded {filename}")
else:
    print("No download response received—check request headers or session")
    # Print first 500 chars of response to debug what page you're getting
    print(download_response.text[:500])

Final Checks

  • Double-check the download URL from Live HTTP Headers—make sure you’re using the exact same parameters and URL structure.
  • If the query button uses POST, replace the GET request to download_url with a POST to the query endpoint, including all form parameters from the query page.

内容的提问来源于stack exchange,提问作者Ingo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:53:19