多页ASPX表格爬取问题:赛狗比赛数据全量获取方案咨询
Hey there! Let's break down how to scrape all race results for Hardwick Serena from that ASP.NET-powered GBGB page—those tricky __VIEWSTATE tokens and submit-button pagination are common hurdles, but we’ve got two solid solutions for you.
Option 1: Load All Results in One Page (Simplest Approach)
Most ASP.NET grid controls let you adjust the number of results per page via a dropdown. If you set this to a large value (like 100 or even 999, if the system allows), you can pull all results in a single request instead of iterating pages. Here's how to do it:
Step-by-Step Implementation
- Initial GET Request: Fetch the page first to capture all required form parameters (
__VIEWSTATE,__VIEWSTATEGENERATOR,__EVENTVALIDATION) and identify the page-size dropdown control ID. - POST with Large Page Size: Submit a form request that sets the page size to a maximum value, forcing all results to load on one page.
Code Example
import requests from bs4 import BeautifulSoup # Persist cookies with a session (critical for ASP.NET) session = requests.Session() target_url = "http://www.gbgb.org.uk/RaceCard.aspx?dogName=Hardwick%20Serena" # First pass to grab form parameters initial_response = session.get(target_url) soup = BeautifulSoup(initial_response.text, "html.parser") # Extract hidden ASP.NET form fields view_state = soup.find("input", {"id": "__VIEWSTATE"})["value"] view_state_gen = soup.find("input", {"id": "__VIEWSTATEGENERATOR"})["value"] event_validation = soup.find("input", {"id": "__EVENTVALIDATION"})["value"] # Find the page-size dropdown ID (check your browser's dev tools for the real ID) page_size_control = "ctl00$MainContent$ucRaceCard$gvResults$ctl13$ddlPageSize" # Build POST data to set page size to 100 (adjust if 100 isn't enough) post_data = { "__VIEWSTATE": view_state, "__VIEWSTATEGENERATOR": view_state_gen, "__EVENTVALIDATION": event_validation, page_size_control: "100", # Use a value larger than the total number of results "ctl00$MainContent$ucRaceCard$btnSearch": "Search" # Match the form's submit button name } # Submit the request to load all results full_results_response = session.post(target_url, data=post_data) results_soup = BeautifulSoup(full_results_response.text, "html.parser") # Scrape the results table results_table = results_soup.find("table", {"id": "ctl00_MainContent_ucRaceCard_gvResults"}) if results_table: rows = results_table.find_all("tr")[1:] # Skip header row for row in rows: columns = row.find_all("td") # Extract data from each column (adjust indices based on table structure) race_date = columns[0].text.strip() track = columns[1].text.strip() race_num = columns[2].text.strip() print(f"Date: {race_date}, Track: {track}, Race #: {race_num}")
Option 2: Iterate Pages Using __VIEWSTATE (More Robust)
If the site limits the maximum page size, you’ll need to handle the "Next Page" submit button by reusing the ASP.NET form tokens with each request. Here's how:
Step-by-Step Implementation
- Initial GET: Grab the initial form tokens and check if a "Next Page" button exists.
- Loop Through Pages: For each page, submit a POST request with the current
__VIEWSTATEtokens and the "Next Page" event parameters. Extract new tokens from the response for the next iteration, until no more pages exist.
Code Example
import requests from bs4 import BeautifulSoup session = requests.Session() target_url = "http://www.gbgb.org.uk/RaceCard.aspx?dogName=Hardwick%20Serena" all_race_results = [] def extract_race_data(soup): """Helper to pull race details from the results table""" table = soup.find("table", {"id": "ctl00_MainContent_ucRaceCard_gvResults"}) if not table: return [] rows = table.find_all("tr")[1:] results = [] for row in rows: cols = row.find_all("td") results.append({ "date": cols[0].text.strip(), "track": cols[1].text.strip(), "race_number": cols[2].text.strip(), "position": cols[3].text.strip(), # Add other fields based on your needs }) return results def get_next_page_form_data(soup): """Extract form tokens and check if next page is available""" # Grab hidden ASP.NET fields form_data = { "__VIEWSTATE": soup.find("input", {"id": "__VIEWSTATE"})["value"], "__VIEWSTATEGENERATOR": soup.find("input", {"id": "__VIEWSTATEGENERATOR"})["value"], "__EVENTVALIDATION": soup.find("input", {"id": "__EVENTVALIDATION"})["value"], "__EVENTTARGET": "ctl00$MainContent$ucRaceCard$gvResults$ctl13$lnkNext", "__EVENTARGUMENT": "" } # Check if next page button is enabled next_button = soup.find("a", {"id": "ctl00_MainContent_ucRaceCard_gvResults_ctl13_lnkNext"}) if not next_button or "disabled" in next_button.get("class", []): return None # No more pages return form_data # Start with the first page current_response = session.get(target_url) current_soup = BeautifulSoup(current_response.text, "html.parser") all_race_results.extend(extract_race_data(current_soup)) # Iterate through subsequent pages while True: next_form_data = get_next_page_form_data(current_soup) if not next_form_data: break # Submit request for next page current_response = session.post(target_url, data=next_form_data) current_soup = BeautifulSoup(current_response.text, "html.parser") new_results = extract_race_data(current_soup) if not new_results: break all_race_results.extend(new_results) # Print total results print(f"Successfully scraped {len(all_race_results)} race results for Hardwick Serena!") for result in all_race_results: print(result)
Key Notes for Both Approaches
- Check Control IDs: The control IDs (like the page-size dropdown or next button) might vary—use your browser's DevTools to inspect the HTML and confirm the exact IDs.
- Session Persistence: Always use a
requests.Session()to keep cookies alive, as ASP.NET uses session cookies to track your request context. - User-Agent Header: Add a
User-Agentto your requests to mimic a real browser (e.g.,headers={"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"}) to avoid being blocked. - Rate Limiting: Add small delays between requests if the site has anti-scraping measures (use
time.sleep(1)between POSTs).
内容的提问来源于stack exchange,提问作者Alex

