使用Python requests登录网站后下载PDF失败,请求排查问题
Troubleshooting Your PDF Download Script with Requests
Hey there, let's break down why your script isn't working and fix it step by step. The core issue here is how you're handling the authenticated session, plus a few other small gotchas that are throwing things off.
Key Issues in Your Current Code
- You're not using the authenticated session to fetch the PDF: After logging in with
s.post(), you're making a freshrequests.get()call instead of reusing the samessession object. This means the login cookies aren't carried over, so the website still sees you as an unauthenticated user. - No check for successful login: You don't verify if the login actually worked before trying to download the PDF. If the login fails (wrong credentials, invalid token), your script will still attempt to fetch the file, which will inevitably fail.
- Unsafe handling of the RequestVerificationToken: If the script can't find the token in the login page, it'll throw a
KeyErrorand crash immediately—there's no error handling to catch this scenario.
Fixed Version of Your Script
Here's the revised code with all these issues addressed, plus some extra safeguards:
import requests import sys from bs4 import BeautifulSoup headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } login_data = { 'Email': 'My-email', 'Password': 'My-password', 'login': 'Login' } pdf_url = 'https://download-website' # URL to your target PDF filename = 'filename.pdf' login_url = 'https://login-website/' # Login page URL print("Creating the connection ...") with requests.session() as s: # 1. Fetch the login page to grab the CSRF token r = s.get(login_url, headers=headers) r.raise_for_status() # Crash immediately if we can't load the login page soup = BeautifulSoup(r.content, 'html5lib') token_element = soup.find('input', attrs={'name':'__RequestVerificationToken'}) if not token_element: print("Error: Could not find RequestVerificationToken on login page", file=sys.stderr) sys.exit(1) login_data['__RequestVerificationToken'] = token_element['value'] # 2. Submit the login form print("Logging in ...") login_response = s.post(login_url, data=login_data, headers=headers) login_response.raise_for_status() # Optional: Verify login success (adjust this based on the website's behavior) # For example, check if the response contains a logged-in indicator or redirects to a dashboard if "Invalid email or password" in login_response.text: # Replace with actual failure text from the site print("Error: Login failed - check your credentials", file=sys.stderr) sys.exit(1) # 3. Use the SAME authenticated session to download the PDF print("Downloading the PDF ...") with s.get(pdf_url, stream=True, headers=headers) as r: r.raise_for_status() # Crash if the download request fails print("Saving the pdf file:\n\"{}\" ...".format(filename)) with open(filename, 'wb') as f: try: total_size = int(r.headers.get('Content-Length', 0)) saved_size = 0 chunk_size = 8192 for chunk in r.iter_content(chunk_size=chunk_size): if chunk: f.write(chunk) saved_size += len(chunk) if total_size > 0: progress = (saved_size / total_size) * 100 print("\r=> %.2f%%" % min(progress, 100.0), end='') print(end='\n\n') except Exception as e: print(f"==> Couldn't save : {filename} - Error: {str(e)}", file=sys.stderr) finally: f.flush()
Additional Tips to Debug Further
- Double-check login form fields: Use your browser's developer tools (F12) to inspect the login form. Make sure the
nameattributes of the email/password fields match what's in yourlogin_data(some sites useusernameinstead ofEmail, orpassinstead ofPassword). - Inspect cookies and headers: After logging in, print
s.cookiesto see if session cookies are set. You can also compare the headers your browser sends when downloading the PDF to what your script sends. - Update your User-Agent: Chrome 87 is pretty old—try using a newer UA string to avoid being flagged as a bot.
- Handle redirects: Some sites redirect after login; the
requests.Sessionhandles this automatically, but you can checklogin_response.historyto see if redirects happened.
内容的提问来源于stack exchange,提问作者DmarZX
相关产品推荐
相关产品推荐

