You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python requests登录网站后下载PDF失败,请求排查问题

Troubleshooting Your PDF Download Script with Requests

Hey there, let's break down why your script isn't working and fix it step by step. The core issue here is how you're handling the authenticated session, plus a few other small gotchas that are throwing things off.

Key Issues in Your Current Code

  • You're not using the authenticated session to fetch the PDF: After logging in with s.post(), you're making a fresh requests.get() call instead of reusing the same s session object. This means the login cookies aren't carried over, so the website still sees you as an unauthenticated user.
  • No check for successful login: You don't verify if the login actually worked before trying to download the PDF. If the login fails (wrong credentials, invalid token), your script will still attempt to fetch the file, which will inevitably fail.
  • Unsafe handling of the RequestVerificationToken: If the script can't find the token in the login page, it'll throw a KeyError and crash immediately—there's no error handling to catch this scenario.

Fixed Version of Your Script

Here's the revised code with all these issues addressed, plus some extra safeguards:

import requests
import sys
from bs4 import BeautifulSoup

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}
login_data = {
    'Email': 'My-email',
    'Password': 'My-password',
    'login': 'Login'
}
pdf_url = 'https://download-website'  # URL to your target PDF
filename = 'filename.pdf'
login_url = 'https://login-website/'  # Login page URL

print("Creating the connection ...")
with requests.session() as s:
    # 1. Fetch the login page to grab the CSRF token
    r = s.get(login_url, headers=headers)
    r.raise_for_status()  # Crash immediately if we can't load the login page
    
    soup = BeautifulSoup(r.content, 'html5lib')
    token_element = soup.find('input', attrs={'name':'__RequestVerificationToken'})
    if not token_element:
        print("Error: Could not find RequestVerificationToken on login page", file=sys.stderr)
        sys.exit(1)
    login_data['__RequestVerificationToken'] = token_element['value']
    
    # 2. Submit the login form
    print("Logging in ...")
    login_response = s.post(login_url, data=login_data, headers=headers)
    login_response.raise_for_status()
    
    # Optional: Verify login success (adjust this based on the website's behavior)
    # For example, check if the response contains a logged-in indicator or redirects to a dashboard
    if "Invalid email or password" in login_response.text:  # Replace with actual failure text from the site
        print("Error: Login failed - check your credentials", file=sys.stderr)
        sys.exit(1)
    
    # 3. Use the SAME authenticated session to download the PDF
    print("Downloading the PDF ...")
    with s.get(pdf_url, stream=True, headers=headers) as r:
        r.raise_for_status()  # Crash if the download request fails
        
        print("Saving the pdf file:\n\"{}\" ...".format(filename))
        with open(filename, 'wb') as f:
            try:
                total_size = int(r.headers.get('Content-Length', 0))
                saved_size = 0
                chunk_size = 8192
                
                for chunk in r.iter_content(chunk_size=chunk_size):
                    if chunk:
                        f.write(chunk)
                        saved_size += len(chunk)
                        if total_size > 0:
                            progress = (saved_size / total_size) * 100
                            print("\r=> %.2f%%" % min(progress, 100.0), end='')
                print(end='\n\n')
            except Exception as e:
                print(f"==> Couldn't save : {filename} - Error: {str(e)}", file=sys.stderr)
            finally:
                f.flush()

Additional Tips to Debug Further

  • Double-check login form fields: Use your browser's developer tools (F12) to inspect the login form. Make sure the name attributes of the email/password fields match what's in your login_data (some sites use username instead of Email, or pass instead of Password).
  • Inspect cookies and headers: After logging in, print s.cookies to see if session cookies are set. You can also compare the headers your browser sends when downloading the PDF to what your script sends.
  • Update your User-Agent: Chrome 87 is pretty old—try using a newer UA string to avoid being flagged as a bot.
  • Handle redirects: Some sites redirect after login; the requests.Session handles this automatically, but you can check login_response.history to see if redirects happened.

内容的提问来源于stack exchange,提问作者DmarZX

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:26:40