You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

VBA实现HTML爬取:多页列表全数据获取问题

Solution for Crawling All 7000+ Listings from Tax Sales Site

Hey Derek, let's tackle this pagination problem for your tax sales property crawl. You've already nailed handling the initial popup and extracting single-page property details—awesome progress! Here's how to scale this to capture all listings:

Step 1: Identify the Site's Pagination Mechanism

First, figure out how the site loads additional pages. There are 3 common patterns you’ll encounter:

  • Pagination Controls: Visible "Next Page"/page number buttons at the bottom of listings
  • URL Parameter Pagination: The page number is part of the URL (e.g., http://taxsales.lgbs.com/?page=2)
  • Infinite Scroll: New listings load automatically when you scroll to the bottom

Quick Check Steps:

  • For pagination controls: Scroll to the bottom of the listings page and look for navigation buttons.
  • For URL parameters: Click "Next Page" once and check if the URL updates with a page number parameter.
  • For infinite scroll: Scroll slowly to the bottom—if new listings pop up without a URL change, it’s infinite scroll.

Step 2: Implement the Corresponding Crawl Logic

Below are tailored solutions for each pattern, assuming you’re using Selenium (since you already used it to click the "Agree" button):

Case 1: Pagination Controls (Next Page Button)

Loop through pages by clicking the "Next Page" button until it’s no longer clickable (e.g., disabled or hidden):

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import NoSuchElementException

driver = webdriver.Chrome()
driver.get("http://taxsales.lgbs.com/")

# Handle initial agree popup (your existing code)
agree_button = WebDriverWait(driver, 10).until(
    EC.element_to_be_clickable((By.XPATH, "//button[text()='Agree']"))
)
agree_button.click()

all_properties = []

while True:
    # Extract current page's properties (replace selectors with your own)
    properties = driver.find_elements(By.CSS_SELECTOR, ".property-listing")
    for prop in properties:
        name = prop.find_element(By.CSS_SELECTOR, ".property-name").text
        details = prop.find_element(By.CSS_SELECTOR, ".property-details").text
        all_properties.append({"name": name, "details": details})
    
    # Try to click Next Page
    try:
        next_button = WebDriverWait(driver, 10).until(
            EC.element_to_be_clickable((By.XPATH, "//button[contains(text(), 'Next')]"))
        )
        # Stop if button is disabled (adjust class check to match the site)
        if "disabled" in next_button.get_attribute("class"):
            break
        next_button.click()
        # Wait for new page content to load
        WebDriverWait(driver, 10).until(
            EC.staleness_of(properties[0])  # Wait until old listings disappear
        )
    except NoSuchElementException:
        # No more pages to load
        break

# Save data to CSV (adjust fields as needed)
import csv
with open("tax_sales_properties.csv", "w", newline="", encoding="utf-8") as f:
    writer = csv.DictWriter(f, fieldnames=["name", "details"])
    writer.writeheader()
    writer.writerows(all_properties)

driver.quit()

Case 2: URL Parameter Pagination

If the URL uses a page parameter, calculate the total number of pages and loop through each URL:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

driver = webdriver.Chrome()
driver.get("http://taxsales.lgbs.com/")

# Handle agree popup
agree_button = WebDriverWait(driver, 10).until(
    EC.element_to_be_clickable((By.XPATH, "//button[text()='Agree']"))
)
agree_button.click()

# Get total number of pages (adjust selector to match the site's total page display)
total_pages_text = WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.CSS_SELECTOR, ".total-pages"))
).text
total_pages = int(total_pages_text.split()[-1])  # Example: extracts 350 from "Page 1 of 350"

all_properties = []

for page_num in range(1, total_pages + 1):
    # Construct page URL with parameter
    page_url = f"http://taxsales.lgbs.com/?page={page_num}"
    driver.get(page_url)
    
    # Wait for listings to load
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CSS_SELECTOR, ".property-listing"))
    )
    
    # Extract properties (replace selectors with your own)
    properties = driver.find_elements(By.CSS_SELECTOR, ".property-listing")
    for prop in properties:
        name = prop.find_element(By.CSS_SELECTOR, ".property-name").text
        details = prop.find_element(By.CSS_SELECTOR, ".property-details").text
        all_properties.append({"name": name, "details": details})

# Save data to CSV
import csv
with open("tax_sales_properties.csv", "w", newline="", encoding="utf-8") as f:
    writer = csv.DictWriter(f, fieldnames=["name", "details"])
    writer.writeheader()
    writer.writerows(all_properties)

driver.quit()

Case 3: Infinite Scroll

For sites that load new listings on scroll, repeatedly scroll to the bottom until no new content loads:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

driver = webdriver.Chrome()
driver.get("http://taxsales.lgbs.com/")

# Handle agree popup
agree_button = WebDriverWait(driver, 10).until(
    EC.element_to_be_clickable((By.XPATH, "//button[text()='Agree']"))
)
agree_button.click()

all_properties = []
last_height = driver.execute_script("return document.body.scrollHeight")

while True:
    # Extract only new properties added since last scroll
    properties = driver.find_elements(By.CSS_SELECTOR, ".property-listing")
    for prop in properties[len(all_properties):]:
        name = prop.find_element(By.CSS_SELECTOR, ".property-name").text
        details = prop.find_element(By.CSS_SELECTOR, ".property-details").text
        all_properties.append({"name": name, "details": details})
    
    # Scroll to bottom of page
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    # Wait for new content to load (adjust time based on site speed)
    time.sleep(2)
    new_height = driver.execute_script("return document.body.scrollHeight")
    
    # Stop if no new content loaded
    if new_height == last_height:
        break
    last_height = new_height

# Save data to CSV
import csv
with open("tax_sales_properties.csv", "w", newline="", encoding="utf-8") as f:
    writer = csv.DictWriter(f, fieldnames=["name", "details"])
    writer.writeheader()
    writer.writerows(all_properties)

driver.quit()

Key Tips to Avoid Headaches

  • Wait for Elements: Use WebDriverWait instead of time.sleep whenever possible—it’s more reliable for dynamic content.
  • Anti-Crawling Protections: Add small delays between page requests, or rotate user agents if the site blocks you.
  • Incremental Saving: Save data after each page instead of waiting until the end—this prevents losing all progress if the script crashes.

内容的提问来源于stack exchange,提问作者Derek Schilling

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:27:22