Python新手求助:使用requests模块构建请求负载遇到问题
Hey there! As a fellow Python learner, I get how tricky request payloads can feel at first—let’s break this down step by step, using your existing URL array as a starting point.
First, let’s clarify: "request payload" can mean a few different things depending on whether you’re making a GET or POST request. I’ll cover the most common scenarios with code examples tailored to your workflow.
First, Let’s Set Up Your Base Code
Assuming your generated URL array looks like this (your 2-item example):
import requests from bs4 import BeautifulSoup import time # Your pre-generated full URL array target_urls = [ "https://example.com/target-page-1", "https://example.com/target-page-2" ]
Scenario 1: GET Requests with Query Parameters
If your "payload" refers to the query parameters (the ?key=value part at the end of URLs), use the params argument in requests.get():
# Define your query parameters as a dictionary query_params = { "category": "tutorials", "sort": "popular", "limit": 5 } for url in target_urls: try: # Send GET request with parameters response = requests.get( url, params=query_params, headers={"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"} ) response.raise_for_status() # Catch HTTP errors (404, 500, etc.) # Process the response (e.g., parse with Beautiful Soup) soup = BeautifulSoup(response.text, "html.parser") # Add your data extraction logic here print(f"Successfully processed: {url}") except requests.exceptions.RequestException as e: print(f"Error accessing {url}: {str(e)}") # Add delay to avoid overwhelming the server time.sleep(2)
Scenario 2: POST Requests with Form Data
If you need to submit a form (like a login or search form), use the data argument for application/x-www-form-urlencoded payloads:
# Define your form fields as a dictionary form_payload = { "username": "demo-user", "password": "demo-pass", "submit_button": "Login" } for url in target_urls: try: response = requests.post( url, data=form_payload, headers={"User-Agent": "Mozilla/5.0"} ) response.raise_for_status() # Check if the POST was successful (e.g., look for a success message) if "Welcome" in response.text: print(f"Login successful for {url}") else: print(f"Login failed for {url}") except requests.exceptions.RequestException as e: print(f"Error submitting form to {url}: {str(e)}") time.sleep(2)
Scenario 3: POST Requests with JSON Payloads
For APIs or modern web apps that accept JSON, use the json argument (requests automatically sets the correct Content-Type header):
# Define your JSON payload as a dictionary json_payload = { "action": "fetch_content", "filters": { "type": "article", "date_range": "last_week" } } for url in target_urls: try: response = requests.post( url, json=json_payload, headers={"User-Agent": "Mozilla/5.0"} ) response.raise_for_status() # Parse the JSON response response_data = response.json() print(f"Received data from {url}: {response_data['results'][:2]}") # Print first 2 results except requests.exceptions.RequestException as e: print(f"Error sending JSON payload to {url}: {str(e)}") time.sleep(2)
Quick Tips for Smooth Requests
- Always include a User-Agent: Many websites block requests without a valid user agent, so adding one mimics a real browser.
- Use
requests.Session()for persistent sessions: If you need to maintain cookies (like after logging in), create a session object once and reuse it for all requests. - Adjust the delay: Depending on the website’s
robots.txtrules, you might want to increase thetime.sleep()value to 3-5 seconds for extra caution. - Handle exceptions: The
try/exceptblocks ensure your script doesn’t crash if one request fails.
内容的提问来源于stack exchange,提问作者pauljohn32

