使用Python Requests模块实现网页登录爬取时的故障排查求助
Hey Marvin, let’s break down your three issues and fix them one by one—here’s practical, actionable solutions tailored to your code:
Right now, your login_values() function returns raw plaintext username/password data, which means anyone with access to your script’s runtime (or if you accidentally print the data dict) can easily see sensitive credentials. Here’s how to harden this:
- Use system keyring for secure password storage: Instead of letting passwords live in memory as plain strings, use the
keyringlibrary to store credentials in your system’s secure keychain (no plaintext ever touches your script’s runtime after the first setup):import keyring import getpass def login_values(): login = "https://www.****.com/login" host = "www.****.com" username = input("Please insert your username: ") # Check if password is already saved in the system keyring password = keyring.get_password(host, username) if not password: password = getpass.getpass("Please type in your password: ") # Save it to the keyring for future use (no more typing!) keyring.set_password(host, username, password) data = { "username": username, "password": password, } return login, host, data - Limit credential scope: Avoid returning the plaintext
datadict if possible—handle the login request directly inside the function so the password never leaves the function’s scope unless absolutely necessary. - Block accidental credential leaks: Double-check all debug
print()statements to ensure they never include thedatadict or password variable.
Your webscrape function has a critical cookie management mistake: you’re creating a CookieJar but not attaching it to your requests.Session, then passing an empty jar to the get request. requests.Session automatically handles cookie persistence—you don’t need to manually manage a CookieJar here. Here’s the fixed version, plus other fixes for session persistence:
import requests import random import socket from bs4 import BeautifulSoup def webscrape(login_url, host_url, login_data, target_url): user_agents = [ # Your list of user agents here ] agent = random.choice(user_agents) headers = { 'User-agent': agent, 'Accept': '*/*', 'Accept-Language': 'en-US,en;q=0.9;zh-cmn-Hans', 'Host': host_url, 'charset': 'utf-8', } socket.setdefaulttimeout(20) s = requests.Session() # Attach headers to the session so all requests use the same identity s.headers.update(headers) # First, load the login page to capture initial cookies (like CSRF tokens) login_page = s.get(login_url) # Many sites require a CSRF token from the login page—extract it if needed: # csrf_token = BeautifulSoup(login_page.text, "lxml").find('input', {'name': 'csrf_token'})['value'] # login_data['csrf_token'] = csrf_token # Send the login POST request login_response = s.post(login_url, data=login_data) # Validate login success (don't assume it worked!) if login_response.status_code != 200 or "login" in login_response.url: raise Exception("Login failed! Check credentials or site changes.") # Fetch the target URL using the same logged-in session res = s.get(target_url) # Double-check we didn't get redirected back to login if "login" in res.url: raise Exception("Session expired or login not persisted.") return res.text
Additional tips to prevent random failures:
- Always validate login success: Check the response URL, status code, or content for signs you’re logged in (e.g., a welcome message or user profile link).
- Reuse sessions across requests: Don’t create a new
Sessionfor everywebscrapecall—reuse one session for login, search, and pagination to maintain consistent login state. - Mimic real browser behavior: Use the same User-Agent and headers for all requests, just like a real browser would.
To handle server-side errors (like 500s) and avoid rate limits, add retry logic with exponential backoff and random delays. We’ll use the tenacity library for clean retry handling:
First, install it:
pip install tenacity
Then update your webscrape function with retries:
import time from tenacity import retry, stop_after_attempt, wait_exponential_jitter @retry(stop=stop_after_attempt(5), wait=wait_exponential_jitter(initial=1, max=10)) def webscrape(login_url, host_url, login_data, target_url): user_agents = [ # Your user agents here ] agent = random.choice(user_agents) headers = { 'User-agent': agent, 'Accept': '*/*', 'Accept-Language': 'en-US,en;q=0.9;zh-cmn-Hans', 'Host': host_url, 'charset': 'utf-8', } socket.setdefaulttimeout(20) s = requests.Session() s.headers.update(headers) login_page = s.get(login_url) login_page.raise_for_status() # Trigger retry on HTTP errors # Add CSRF token if needed (uncomment below) # csrf_token = BeautifulSoup(login_page.text, "lxml").find('input', {'name': 'csrf_token'})['value'] # login_data['csrf_token'] = csrf_token login_response = s.post(login_url, data=login_data) login_response.raise_for_status() if "login" in login_response.url: raise Exception("Login failed, aborting retry.") res = s.get(target_url) res.raise_for_status() if "login" in res.url: raise Exception("Session lost, aborting retry.") # Add a small random delay to mimic human behavior time.sleep(random.uniform(1, 3)) return res.text
What this does:
- Retries up to 5 times: If you hit server errors or timeouts, it retries after a random exponential delay (starts at 1s, max 10s) to avoid overwhelming the server.
- Aborts on permanent failures: If login fails, it stops retrying (since that’s a credential issue, not a temporary server problem).
- Mimics human behavior: The random delay between requests helps avoid triggering rate limits.
Bonus: For your getlist function, reuse the same logged-in session instead of calling webscrape repeatedly (which redoes login every time). This will make pagination more reliable and efficient.
内容的提问来源于stack exchange,提问作者Marvin

