You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用requests模块批量爬取高校网站PDF的登录表单问题咨询

Hey there! Let's work through that tricky campus login form issue you're hitting with your PDF crawler. I've messed around with a few finicky university authentication systems before, so here's a step-by-step approach to get past this:

Step 1: Scrape Dynamic Hidden Form Fields

Most university login pages throw in hidden fields (like __VIEWSTATE, __EVENTVALIDATION, or custom CSRF tokens) that change every time you load the page. You can't skip these—you need to grab them first before sending your login request.

Here's how to do it with requests and BeautifulSoup:

import requests
from bs4 import BeautifulSoup

# Initialize a session to persist cookies automatically
session = requests.Session()
login_url = "https://your-campus-login-page-url.com"

# First, fetch the login page to get all form fields
login_page = session.get(login_url)
soup = BeautifulSoup(login_page.text, "html.parser")

# Extract every input field from the form (including hidden ones)
form_payload = {}
for input_field in soup.find_all("input"):
    field_name = input_field.get("name")
    field_value = input_field.get("value", "")
    if field_name:
        form_payload[field_name] = field_value

# Add your actual login credentials to the payload
form_payload["username"] = "your-student-id-or-email"
form_payload["password"] = "your-campus-password"
Step 2: Mimic Browser Behavior for Form Submission

Schools often block requests that don't look like they're coming from a real browser. Plus, the form might submit to a specific endpoint (not always the same as the login page URL).

  • Check the <form> tag's action attribute to find the correct POST URL (it might be relative, so use requests.compat.urljoin to fix that)
  • Add realistic headers to avoid being flagged:
# Mimic a modern browser's headers
headers = {
    "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 13_4) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36",
    "Referer": login_url  # Tell the server you came from the login page
}

# Get the form's submission URL
form_action = soup.find("form").get("action", "")
post_url = login_url if form_action == "" else requests.compat.urljoin(login_url, form_action)

# Send the login request
login_response = session.post(post_url, data=form_payload, headers=headers)

# Verify login worked—look for a unique string that appears only when logged in
if "Welcome back," in login_response.text:
    print("Login successful! Time to grab those PDFs.")
else:
    print("Login failed—double-check your credentials or form fields.")
Step 3: Troubleshooting Sticky Cases

If the above doesn't work, here are some common fixes:

  • JavaScript-generated tokens: If the form uses JS to create tokens (you'll see this in DevTools Network tab), requests can't execute JS. Switch to selenium or playwright to render the page and grab dynamic values.
  • Check DevTools: Open your browser's DevTools > Network tab, submit the login manually, and copy the exact Form Data and Headers from the POST request. Match those in your code—even small differences can break login.
  • Session persistence: Always use requests.Session()—it keeps track of login cookies so you don't have to manually pass them around for every PDF request.

Once you're logged in, downloading the PDFs is straightforward with the same session:

# Replace with your list of PDF URLs
pdf_urls = [
    "https://your-campus-site.com/docs/report1.pdf",
    "https://your-campus-site.com/docs/report2.pdf"
]

for i, pdf_link in enumerate(pdf_urls, 1):
    pdf_response = session.get(pdf_link)
    # Save the PDF to your local machine
    with open(f"campus_pdf_{i}.pdf", "wb") as file:
        file.write(pdf_response.content)
    print(f"Downloaded: {pdf_link}")

内容的提问来源于stack exchange,提问作者Brand0n

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:11:46