表单式网站数据抓取问题:批量遍历邮编获取工党站点数据
Hey there! Let's tackle this council cuts calculator data scraping problem step by step. First, let's fill in the gaps in your existing code and build a complete solution for extracting data across your postcode list.
Step 1: Validate Request Success & Inspect Response Structure
First, let's confirm your POST request works and figure out what kind of response we're dealing with. Add these lines right after your response = requests.post(...) call to debug:
print(response.status_code) # Should return 200 if the request succeeds print(response.headers['Content-Type']) # Tells us if it's HTML, JSON, etc. print(response.text[:500]) # Print the first 500 characters to spot data patterns
Most likely, the response will be HTML (since the URL points to a web page). We'll use BeautifulSoup to parse this HTML and pull out the relevant data.
Step 2: Install & Integrate BeautifulSoup for HTML Parsing
First, install the package if you haven't already:
pip install beautifulsoup4
Then, update your code to loop through your postcode list, parse responses, and extract data. You'll need to tweak the selectors based on the actual page structure (use your browser's dev tools to inspect elements):
import requests from bs4 import BeautifulSoup import time # Replace with your full list of postcodes postcode_list = ['WS13 6QG', 'SW1A 0AA', 'M1 1AE'] target_url = 'https://action.labour.org.uk/page/content/council-cuts-calculator' # Loop through each postcode for postcode in postcode_list: # Define the payload with the current postcode payload = { 'firstname': 'james', 'email': 'myemailaddress', 'zip': postcode, 'custom_15452': 'no' } try: # Send the POST request response = requests.post(target_url, data=payload) response.raise_for_status() # Trigger error if request fails (e.g., 404, 500) # Parse the HTML response soup = BeautifulSoup(response.text, 'html.parser') # Example: Extract data (adjust selectors to match the actual page) # Use browser dev tools to find the right classes/IDs for your target data results_container = soup.find('div', class_='council-cuts-results') if results_container: council_name = results_container.find('h3', class_='council-title').text.strip() total_cuts = results_container.find('span', class_='total-cuts-amount').text.strip() service_impacts = [item.text.strip() for item in results_container.find_all('li', class_='impact-item')] # Print or store the extracted data print(f"=== Postcode: {postcode} ===") print(f"Council: {council_name}") print(f"Total Cuts: {total_cuts}") print("Service Impacts:") for impact in service_impacts: print(f"- {impact}") print("\n") else: print(f"No results found for postcode: {postcode}\n") # Add a small delay to avoid overwhelming the server (prevents getting blocked) time.sleep(2) except requests.exceptions.RequestException as e: print(f"Error processing postcode {postcode}: {str(e)}\n")
Step 3: Tune Selectors to Match the Actual Page
The selectors (like 'div', class_='council-cuts-results') are just examples. To get the correct ones:
- Open the calculator page in your browser
- Submit a postcode manually
- Right-click the data you want to extract and select "Inspect"
- Copy the element's class, ID, or parent container details to use in BeautifulSoup's
find()/find_all()methods.
Key Best Practices
- Respect Rate Limits: The
time.sleep(2)ensures you don't send requests too quickly—adjust the delay if needed to stay compliant. - Check Terms of Service: Make sure scraping this data aligns with the site's terms of use.
- Handle Edge Cases: The try/except block catches request errors (like timeouts or server issues) so your script doesn't crash unexpectedly.
内容的提问来源于stack exchange,提问作者TaxpayersMoney

