Selenium动态站点抓取失败:同一URL多供水系统联系信息抓取解决方案咨询
Selenium动态站点抓取失败:同一URL多供水系统联系信息抓取解决方案咨询
Hey there, I totally get how frustrating this is—dealing with dynamic government pages that reuse the same URL, Selenium timeouts, and frameset confusion can make you pull your hair out. Let’s break down what’s wrong with your code and fix it step by step.
🔍 First, Let’s Diagnose the Core Issues
Your script has a few key problems that are causing timeouts and failures, plus the unique challenge of this site’s dynamic content:
- Dynamic Content Limitation: All system details load into the same URL via JavaScript—no unique URLs per system, so you can’t just construct links; you have to trigger JS events to load the right data.
- Incorrect Frame Targeting: You tried switching to frame index 2, but hardcoded indexes are unreliable. Also, since this uses
<frame>tags from a frameset (not<iframe>), checking foriframeelements will always return 0—that’s normal here. - Stale Element References: After clicking a link and navigating back, your original
rowselements become "stale" (the DOM reloads, so old element references no longer work). - Unreliable Wait Conditions: Government sites are often slow, and using just
presence_of_element_locateddoesn’t guarantee the element is ready to interact with.
🛠️ Fixed Code to Scrape Alternate State Numbers
Here’s a revised, robust version of your script that addresses all these issues:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import csv import time import random # --- Set up Chrome with stable options --- driver = webdriver.Chrome() # Keep non-headless for testing driver.maximize_window() # Target URL for Anderson County water systems base_url = ( "https://dww.kdhe.ks.gov/DWW/JSP/WaterSystems.jsp?" "PointOfContactType=none&RegulatingAgency=All&number=&name=&county=Anderson" ) driver.get(base_url) results = [] try: # 1️⃣ Switch to the content frame (use name for reliability) # Verify frame name via inspect: look for <frame name="content"> in the frameset WebDriverWait(driver, 15).until( EC.frame_to_be_available_and_switch_to_it((By.NAME, "content")) ) print("✅ Switched to content frame successfully") # 2️⃣ Wait for the water systems table to load fully table = WebDriverWait(driver, 15).until( EC.presence_of_element_located((By.XPATH, "//table[contains(@class,'Data')]")) ) # 3️⃣ Loop through each water system row total_rows = len(WebDriverWait(driver, 15).until( EC.presence_of_all_elements_located((By.XPATH, "//table[contains(@class,'Data')]/tbody/tr[position()>1]")) )) for idx in range(total_rows): # Re-fetch rows every iteration to avoid stale elements after back navigation rows = WebDriverWait(driver, 15).until( EC.presence_of_all_elements_located((By.XPATH, "//table[contains(@class,'Data')]/tbody/tr[position()>1]")) ) row = rows[idx] # Extract basic system info ws_link = row.find_element(By.TAG_NAME, "a") ws_number = ws_link.text.strip() ws_name = row.find_elements(By.TAG_NAME, "td")[1].text.strip() print(f"\n🔍 Processing: {ws_name} ({ws_number})") # Open detail page: use JS execution instead of click for better stability onclick_action = ws_link.get_attribute("onclick") driver.execute_script(onclick_action) time.sleep(random.uniform(1.5, 3)) # Random delay to avoid anti-scraping blocks # 4️⃣ Extract Alternate State No. from the detail table try: # Wait for the target label to be visible (not just present) alt_label = WebDriverWait(driver, 15).until( EC.visibility_of_element_located((By.XPATH, "//td[contains(text(), 'Alternate State No.')]")) ) alt_state_no = alt_label.find_element(By.XPATH, "./following-sibling::td").text.strip() print(f" → Alternate State No.: {alt_state_no}") results.append([ws_name, ws_number, alt_state_no]) except Exception as e: print(f" ❌ Failed to extract data for {ws_number}: {str(e)}") results.append([ws_name, ws_number, "N/A"]) # 5️⃣ Navigate back and re-switch to content frame driver.back() WebDriverWait(driver, 15).until( EC.frame_to_be_available_and_switch_to_it((By.NAME, "content")) ) time.sleep(random.uniform(1, 2)) # --- Handle critical errors that stop the script --- except Exception as e: print(f"\n❌ Critical script failure: {str(e)}") # --- Save results and clean up --- finally: with open("anderson_alternate_state_numbers.csv", "w", newline="", encoding="utf-8") as f: writer = csv.writer(f) writer.writerow(["Water System Name", "System No.", "Alternate State No."]) writer.writerows(results) print(f"\n✅ Saved {len(results)} records to anderson_alternate_state_numbers.csv") driver.quit()
📌 Key Fixes & Critical Explanations
Let’s go over the most important changes that make this script work:
- Reliable Frame Switching: Instead of a hardcoded index, we switch to the frame by its
nameattribute ("content"). This is far more stable because frame indexes can change if the site updates its layout. - Avoid Stale Elements: We re-fetch the
rowslist every loop iteration after navigating back. This ensures we always use fresh DOM elements that aren’t stale. - Stable Detail Page Loading: Instead of simulating a mouse click, we execute the link’s native JavaScript
onclickaction directly. This is more reliable for sites with finicky event handlers. - Robust Wait & Delays: We increased timeouts to 15 seconds (government sites are often slow) and added random delays to avoid triggering anti-scraping measures. We also use
visibility_of_element_locatedto ensure the element is actually visible before extracting text. - Error Handling: Added try/except blocks to handle individual row failures without crashing the entire script, so you can see exactly which systems had issues.
内容来源于stack exchange
相关产品推荐
相关产品推荐

