You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium动态站点抓取失败:同一URL多供水系统联系信息抓取解决方案咨询

Selenium动态站点抓取失败:同一URL多供水系统联系信息抓取解决方案咨询

Hey there, I totally get how frustrating this is—dealing with dynamic government pages that reuse the same URL, Selenium timeouts, and frameset confusion can make you pull your hair out. Let’s break down what’s wrong with your code and fix it step by step.

🔍 First, Let’s Diagnose the Core Issues

Your script has a few key problems that are causing timeouts and failures, plus the unique challenge of this site’s dynamic content:

  1. Dynamic Content Limitation: All system details load into the same URL via JavaScript—no unique URLs per system, so you can’t just construct links; you have to trigger JS events to load the right data.
  2. Incorrect Frame Targeting: You tried switching to frame index 2, but hardcoded indexes are unreliable. Also, since this uses <frame> tags from a frameset (not <iframe>), checking for iframe elements will always return 0—that’s normal here.
  3. Stale Element References: After clicking a link and navigating back, your original rows elements become "stale" (the DOM reloads, so old element references no longer work).
  4. Unreliable Wait Conditions: Government sites are often slow, and using just presence_of_element_located doesn’t guarantee the element is ready to interact with.

🛠️ Fixed Code to Scrape Alternate State Numbers

Here’s a revised, robust version of your script that addresses all these issues:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import csv
import time
import random

# --- Set up Chrome with stable options ---
driver = webdriver.Chrome()  # Keep non-headless for testing
driver.maximize_window()

# Target URL for Anderson County water systems
base_url = (
    "https://dww.kdhe.ks.gov/DWW/JSP/WaterSystems.jsp?"
    "PointOfContactType=none&RegulatingAgency=All&number=&name=&county=Anderson"
)
driver.get(base_url)

results = []

try:
    # 1️⃣ Switch to the content frame (use name for reliability)
    # Verify frame name via inspect: look for <frame name="content"> in the frameset
    WebDriverWait(driver, 15).until(
        EC.frame_to_be_available_and_switch_to_it((By.NAME, "content"))
    )
    print("✅ Switched to content frame successfully")

    # 2️⃣ Wait for the water systems table to load fully
    table = WebDriverWait(driver, 15).until(
        EC.presence_of_element_located((By.XPATH, "//table[contains(@class,'Data')]"))
    )

    # 3️⃣ Loop through each water system row
    total_rows = len(WebDriverWait(driver, 15).until(
        EC.presence_of_all_elements_located((By.XPATH, "//table[contains(@class,'Data')]/tbody/tr[position()>1]"))
    ))

    for idx in range(total_rows):
        # Re-fetch rows every iteration to avoid stale elements after back navigation
        rows = WebDriverWait(driver, 15).until(
            EC.presence_of_all_elements_located((By.XPATH, "//table[contains(@class,'Data')]/tbody/tr[position()>1]"))
        )
        row = rows[idx]

        # Extract basic system info
        ws_link = row.find_element(By.TAG_NAME, "a")
        ws_number = ws_link.text.strip()
        ws_name = row.find_elements(By.TAG_NAME, "td")[1].text.strip()
        print(f"\n🔍 Processing: {ws_name} ({ws_number})")

        # Open detail page: use JS execution instead of click for better stability
        onclick_action = ws_link.get_attribute("onclick")
        driver.execute_script(onclick_action)
        time.sleep(random.uniform(1.5, 3))  # Random delay to avoid anti-scraping blocks

        # 4️⃣ Extract Alternate State No. from the detail table
        try:
            # Wait for the target label to be visible (not just present)
            alt_label = WebDriverWait(driver, 15).until(
                EC.visibility_of_element_located((By.XPATH, "//td[contains(text(), 'Alternate State No.')]"))
            )
            alt_state_no = alt_label.find_element(By.XPATH, "./following-sibling::td").text.strip()
            print(f" → Alternate State No.: {alt_state_no}")
            results.append([ws_name, ws_number, alt_state_no])
        except Exception as e:
            print(f" ❌ Failed to extract data for {ws_number}: {str(e)}")
            results.append([ws_name, ws_number, "N/A"])

        # 5️⃣ Navigate back and re-switch to content frame
        driver.back()
        WebDriverWait(driver, 15).until(
            EC.frame_to_be_available_and_switch_to_it((By.NAME, "content"))
        )
        time.sleep(random.uniform(1, 2))

# --- Handle critical errors that stop the script ---
except Exception as e:
    print(f"\n❌ Critical script failure: {str(e)}")

# --- Save results and clean up ---
finally:
    with open("anderson_alternate_state_numbers.csv", "w", newline="", encoding="utf-8") as f:
        writer = csv.writer(f)
        writer.writerow(["Water System Name", "System No.", "Alternate State No."])
        writer.writerows(results)
    print(f"\n✅ Saved {len(results)} records to anderson_alternate_state_numbers.csv")
    driver.quit()

📌 Key Fixes & Critical Explanations

Let’s go over the most important changes that make this script work:

  1. Reliable Frame Switching: Instead of a hardcoded index, we switch to the frame by its name attribute ("content"). This is far more stable because frame indexes can change if the site updates its layout.
  2. Avoid Stale Elements: We re-fetch the rows list every loop iteration after navigating back. This ensures we always use fresh DOM elements that aren’t stale.
  3. Stable Detail Page Loading: Instead of simulating a mouse click, we execute the link’s native JavaScript onclick action directly. This is more reliable for sites with finicky event handlers.
  4. Robust Wait & Delays: We increased timeouts to 15 seconds (government sites are often slow) and added random delays to avoid triggering anti-scraping measures. We also use visibility_of_element_located to ensure the element is actually visible before extracting text.
  5. Error Handling: Added try/except blocks to handle individual row failures without crashing the entire script, so you can see exactly which systems had issues.

内容来源于stack exchange

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.08 07:29:35