使用requests模块批量爬取高校网站PDF的登录表单问题咨询
Hey there! Let's work through that tricky campus login form issue you're hitting with your PDF crawler. I've messed around with a few finicky university authentication systems before, so here's a step-by-step approach to get past this:
Most university login pages throw in hidden fields (like __VIEWSTATE, __EVENTVALIDATION, or custom CSRF tokens) that change every time you load the page. You can't skip these—you need to grab them first before sending your login request.
Here's how to do it with requests and BeautifulSoup:
import requests from bs4 import BeautifulSoup # Initialize a session to persist cookies automatically session = requests.Session() login_url = "https://your-campus-login-page-url.com" # First, fetch the login page to get all form fields login_page = session.get(login_url) soup = BeautifulSoup(login_page.text, "html.parser") # Extract every input field from the form (including hidden ones) form_payload = {} for input_field in soup.find_all("input"): field_name = input_field.get("name") field_value = input_field.get("value", "") if field_name: form_payload[field_name] = field_value # Add your actual login credentials to the payload form_payload["username"] = "your-student-id-or-email" form_payload["password"] = "your-campus-password"
Schools often block requests that don't look like they're coming from a real browser. Plus, the form might submit to a specific endpoint (not always the same as the login page URL).
- Check the
<form>tag'sactionattribute to find the correct POST URL (it might be relative, so userequests.compat.urljointo fix that) - Add realistic headers to avoid being flagged:
# Mimic a modern browser's headers headers = { "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 13_4) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36", "Referer": login_url # Tell the server you came from the login page } # Get the form's submission URL form_action = soup.find("form").get("action", "") post_url = login_url if form_action == "" else requests.compat.urljoin(login_url, form_action) # Send the login request login_response = session.post(post_url, data=form_payload, headers=headers) # Verify login worked—look for a unique string that appears only when logged in if "Welcome back," in login_response.text: print("Login successful! Time to grab those PDFs.") else: print("Login failed—double-check your credentials or form fields.")
If the above doesn't work, here are some common fixes:
- JavaScript-generated tokens: If the form uses JS to create tokens (you'll see this in DevTools Network tab),
requestscan't execute JS. Switch toseleniumorplaywrightto render the page and grab dynamic values. - Check DevTools: Open your browser's DevTools > Network tab, submit the login manually, and copy the exact Form Data and Headers from the POST request. Match those in your code—even small differences can break login.
- Session persistence: Always use
requests.Session()—it keeps track of login cookies so you don't have to manually pass them around for every PDF request.
Once you're logged in, downloading the PDFs is straightforward with the same session:
# Replace with your list of PDF URLs pdf_urls = [ "https://your-campus-site.com/docs/report1.pdf", "https://your-campus-site.com/docs/report2.pdf" ] for i, pdf_link in enumerate(pdf_urls, 1): pdf_response = session.get(pdf_link) # Save the PDF to your local machine with open(f"campus_pdf_{i}.pdf", "wb") as file: file.write(pdf_response.content) print(f"Downloaded: {pdf_link}")
内容的提问来源于stack exchange,提问作者Brand0n

