如何用urllib/urllib2/requests登录Box指定页面并获取重定向地址?
Alright, let's break down how to automate this Box login flow using Python's requests, urllib, or urllib2 libraries. The key here is handling hidden form parameters (like CSRF tokens) and properly capturing the redirect URL after submitting the "Continue" button.
Using the requests Library (Recommended)
requests simplifies session management and form handling, making it the easiest option for this task. We'll use a session to persist cookies and parse the login page to extract required form fields.
import requests from bs4 import BeautifulSoup # Core URLs for the login flow LOGIN_PAGE_URL = "https://pg.account.box.com/login" SUBMIT_URL = "https://pg.account.box.com/api/v2/login/submit" # Replace with your actual email address USER_EMAIL = "your_email@example.com" # Step 1: Fetch the login page to get hidden form parameters and initialize session session = requests.Session() # Add a realistic User-Agent to avoid being blocked by anti-scraping measures headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } login_page_response = session.get(LOGIN_PAGE_URL, headers=headers) soup = BeautifulSoup(login_page_response.text, "html.parser") # Extract all hidden and required form fields form_data = {} for input_tag in soup.find_all("input"): field_name = input_tag.get("name") field_value = input_tag.get("value") if field_name: form_data[field_name] = field_value # Add the email address to the form data (this is what the "Continue" button submits) form_data["login"] = USER_EMAIL # Step 2: Submit the form and capture the redirect URL # Disable auto-redirects to directly access the Location header submit_response = session.post(SUBMIT_URL, data=form_data, headers=headers, allow_redirects=False) # Check for redirect status codes (301/302) if submit_response.status_code in (301, 302): redirect_url = submit_response.headers.get("Location") print(f"Redirect URL after clicking 'Continue': {redirect_url}") else: # If no redirect, print the response to debug (e.g., invalid email, blocked request) print("No redirect occurred. Response content for debugging:") print(submit_response.text)
Key Notes for requests:
- Using
requests.Session()ensures cookies are persisted between requests, which is critical for maintaining login state. - We extract all form fields (not just the email) because Box includes hidden CSRF tokens and client identifiers that are required for valid requests.
- Adding a valid
User-Agentheader helps avoid being flagged as a bot.
Using urllib (Python 3)
If you prefer using Python's built-in libraries, urllib can handle this flow, though it requires more manual setup for cookies and form encoding.
import urllib.request from urllib.parse import urlencode from bs4 import BeautifulSoup LOGIN_PAGE_URL = "https://pg.account.box.com/login" SUBMIT_URL = "https://pg.account.box.com/api/v2/login/submit" USER_EMAIL = "your_email@example.com" # Step 1: Set up a cookie handler to persist session cookies cookie_handler = urllib.request.HTTPCookieProcessor() opener = urllib.request.build_opener(cookie_handler) urllib.request.install_opener(opener) # Fetch the login page headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } login_request = urllib.request.Request(LOGIN_PAGE_URL, headers=headers) login_response = opener.open(login_request) html_content = login_response.read().decode("utf-8") # Parse form fields soup = BeautifulSoup(html_content, "html.parser") form_data = {} for input_tag in soup.find_all("input"): field_name = input_tag.get("name") field_value = input_tag.get("value") if field_name: form_data[field_name] = field_value form_data["login"] = USER_EMAIL # Step 2: Encode form data and submit the request encoded_form_data = urlencode(form_data).encode("utf-8") submit_request = urllib.request.Request( SUBMIT_URL, data=encoded_form_data, headers=headers ) # Handle redirects (urllib throws an HTTPError for 3xx status codes) try: submit_response = opener.open(submit_request) except urllib.error.HTTPError as e: if e.code in (301, 302): redirect_url = e.headers.get("Location") print(f"Redirect URL after clicking 'Continue': {redirect_url}") else: # Re-raise the error if it's not a redirect raise else: print("No redirect occurred. Response content for debugging:") print(submit_response.read().decode("utf-8"))
Using urllib2 (Python 2)
For Python 2 environments, urllib2 is the equivalent built-in library. The flow is similar to Python 3's urllib, but with minor syntax differences:
import urllib2 from urllib import urlencode from bs4 import BeautifulSoup LOGIN_PAGE_URL = "https://pg.account.box.com/login" SUBMIT_URL = "https://pg.account.box.com/api/v2/login/submit" USER_EMAIL = "your_email@example.com" # Set up cookie handler cookie_handler = urllib2.HTTPCookieProcessor() opener = urllib2.build_opener(cookie_handler) urllib2.install_opener(opener) # Fetch login page headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } login_request = urllib2.Request(LOGIN_PAGE_URL, headers=headers) login_response = opener.open(login_request) html_content = login_response.read() # Parse form fields soup = BeautifulSoup(html_content, "html.parser") form_data = {} for input_tag in soup.find_all("input"): field_name = input_tag.get("name") field_value = input_tag.get("value") if field_name: form_data[field_name] = field_value form_data["login"] = USER_EMAIL # Submit form encoded_form_data = urlencode(form_data) submit_request = urllib2.Request( SUBMIT_URL, data=encoded_form_data, headers=headers ) try: submit_response = opener.open(submit_request) except urllib2.HTTPError as e: if e.code in (301, 302): redirect_url = e.headers.get("Location") print("Redirect URL after clicking 'Continue': {}".format(redirect_url)) else: raise else: print("No redirect occurred. Response content for debugging:") print(submit_response.read())
General Troubleshooting Tips:
- If you get a "403 Forbidden" error, double-check your
User-Agentheader or try adding additional headers (likeReferer) to mimic a real browser. - Ensure you're using the correct
SUBMIT_URL– you can verify this by inspecting the "Continue" button's form action in your browser's dev tools. - If the redirect URL leads to a password entry page, that's expected behavior (Box's flow first asks for an email, then redirects to password entry if the email is recognized).
内容的提问来源于stack exchange,提问作者Mariah Akinbi

