VBA实现HTML爬取:多页列表全数据获取问题
Hey Derek, let's tackle this pagination problem for your tax sales property crawl. You've already nailed handling the initial popup and extracting single-page property details—awesome progress! Here's how to scale this to capture all listings:
Step 1: Identify the Site's Pagination Mechanism
First, figure out how the site loads additional pages. There are 3 common patterns you’ll encounter:
- Pagination Controls: Visible "Next Page"/page number buttons at the bottom of listings
- URL Parameter Pagination: The page number is part of the URL (e.g.,
http://taxsales.lgbs.com/?page=2) - Infinite Scroll: New listings load automatically when you scroll to the bottom
Quick Check Steps:
- For pagination controls: Scroll to the bottom of the listings page and look for navigation buttons.
- For URL parameters: Click "Next Page" once and check if the URL updates with a page number parameter.
- For infinite scroll: Scroll slowly to the bottom—if new listings pop up without a URL change, it’s infinite scroll.
Step 2: Implement the Corresponding Crawl Logic
Below are tailored solutions for each pattern, assuming you’re using Selenium (since you already used it to click the "Agree" button):
Case 1: Pagination Controls (Next Page Button)
Loop through pages by clicking the "Next Page" button until it’s no longer clickable (e.g., disabled or hidden):
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import NoSuchElementException driver = webdriver.Chrome() driver.get("http://taxsales.lgbs.com/") # Handle initial agree popup (your existing code) agree_button = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, "//button[text()='Agree']")) ) agree_button.click() all_properties = [] while True: # Extract current page's properties (replace selectors with your own) properties = driver.find_elements(By.CSS_SELECTOR, ".property-listing") for prop in properties: name = prop.find_element(By.CSS_SELECTOR, ".property-name").text details = prop.find_element(By.CSS_SELECTOR, ".property-details").text all_properties.append({"name": name, "details": details}) # Try to click Next Page try: next_button = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, "//button[contains(text(), 'Next')]")) ) # Stop if button is disabled (adjust class check to match the site) if "disabled" in next_button.get_attribute("class"): break next_button.click() # Wait for new page content to load WebDriverWait(driver, 10).until( EC.staleness_of(properties[0]) # Wait until old listings disappear ) except NoSuchElementException: # No more pages to load break # Save data to CSV (adjust fields as needed) import csv with open("tax_sales_properties.csv", "w", newline="", encoding="utf-8") as f: writer = csv.DictWriter(f, fieldnames=["name", "details"]) writer.writeheader() writer.writerows(all_properties) driver.quit()
Case 2: URL Parameter Pagination
If the URL uses a page parameter, calculate the total number of pages and loop through each URL:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC driver = webdriver.Chrome() driver.get("http://taxsales.lgbs.com/") # Handle agree popup agree_button = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, "//button[text()='Agree']")) ) agree_button.click() # Get total number of pages (adjust selector to match the site's total page display) total_pages_text = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, ".total-pages")) ).text total_pages = int(total_pages_text.split()[-1]) # Example: extracts 350 from "Page 1 of 350" all_properties = [] for page_num in range(1, total_pages + 1): # Construct page URL with parameter page_url = f"http://taxsales.lgbs.com/?page={page_num}" driver.get(page_url) # Wait for listings to load WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, ".property-listing")) ) # Extract properties (replace selectors with your own) properties = driver.find_elements(By.CSS_SELECTOR, ".property-listing") for prop in properties: name = prop.find_element(By.CSS_SELECTOR, ".property-name").text details = prop.find_element(By.CSS_SELECTOR, ".property-details").text all_properties.append({"name": name, "details": details}) # Save data to CSV import csv with open("tax_sales_properties.csv", "w", newline="", encoding="utf-8") as f: writer = csv.DictWriter(f, fieldnames=["name", "details"]) writer.writeheader() writer.writerows(all_properties) driver.quit()
Case 3: Infinite Scroll
For sites that load new listings on scroll, repeatedly scroll to the bottom until no new content loads:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time driver = webdriver.Chrome() driver.get("http://taxsales.lgbs.com/") # Handle agree popup agree_button = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, "//button[text()='Agree']")) ) agree_button.click() all_properties = [] last_height = driver.execute_script("return document.body.scrollHeight") while True: # Extract only new properties added since last scroll properties = driver.find_elements(By.CSS_SELECTOR, ".property-listing") for prop in properties[len(all_properties):]: name = prop.find_element(By.CSS_SELECTOR, ".property-name").text details = prop.find_element(By.CSS_SELECTOR, ".property-details").text all_properties.append({"name": name, "details": details}) # Scroll to bottom of page driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # Wait for new content to load (adjust time based on site speed) time.sleep(2) new_height = driver.execute_script("return document.body.scrollHeight") # Stop if no new content loaded if new_height == last_height: break last_height = new_height # Save data to CSV import csv with open("tax_sales_properties.csv", "w", newline="", encoding="utf-8") as f: writer = csv.DictWriter(f, fieldnames=["name", "details"]) writer.writeheader() writer.writerows(all_properties) driver.quit()
Key Tips to Avoid Headaches
- Wait for Elements: Use
WebDriverWaitinstead oftime.sleepwhenever possible—it’s more reliable for dynamic content. - Anti-Crawling Protections: Add small delays between page requests, or rotate user agents if the site blocks you.
- Incremental Saving: Save data after each page instead of waiting until the end—this prevents losing all progress if the script crashes.
内容的提问来源于stack exchange,提问作者Derek Schilling

