You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R从Zillow抓取更多数据点?求助扩展现有代码

Expanding Your Zillow Scraper to Grab More Property Data

Hey Scott, totally get where you're coming from—inheriting code when you're still learning can feel like trying to solve a puzzle with half the pieces missing. Let's break this down step by step to expand your Zillow scraper to grab more useful data points, and I’ll explain everything so you can follow along easily.

First, Let’s Outline Common Data Points You Can Grab

Most Zillow property pages have these key details available (depending on the property type):

  • Full property address (street, city, state, zip)
  • Number of bedrooms and bathrooms
  • Total square footage
  • Year the property was built
  • Property type (single-family home, townhouse, condo, etc.)
  • Estimated annual property taxes
  • HOA fees (if applicable)
  • Lot size (for single-family homes)

Revised Scraper Code (With Clear Comments!)

I’ll assume your original code uses rvest (the standard R package for web scraping) since it’s the go-to tool for this kind of task. Here’s an expanded version with explanations for every part:

# Install required packages first if you haven't (run this once)
# install.packages(c("rvest", "dplyr", "readr", "stringr", "purrr"))

# Load the packages we need
library(rvest)
library(dplyr)
library(readr)
library(stringr)
library(purrr)

# Define a function to scrape data from a single Zillow property URL
scrape_zillow_property <- function(url) {
  # Add a 2-second delay between requests to avoid getting blocked by Zillow
  Sys.sleep(2)
  
  # Try to load the page and handle errors gracefully (no crashing if one URL fails!)
  page <- tryCatch(
    read_html(url),
    error = function(e) {
      message(paste("Oops, failed to load URL:", url))
      return(NULL)
    }
  )
  
  if (is.null(page)) return(NULL)
  
  # Extract each data point using CSS selectors (I'll explain how to find these below)
  property_data <- tibble(
    url = url,
    zestimate = page %>% 
      html_node(".Text-c11n-8-84-3__sc-aiai24-0.juCZje") %>% 
      html_text() %>% 
      str_remove_all("\\$|,") %>%  # Remove $ and commas to convert to number
      as.numeric(),
    rent_zestimate = page %>% 
      html_node(".Text-c11n-8-84-3__sc-aiai24-0.hOqYiw") %>% 
      html_text() %>% 
      str_remove_all("\\$|,") %>% 
      as.numeric(),
    full_address = page %>% 
      html_node(".Text-c11n-8-84-3__sc-aiai24-0.hRqIYX") %>% 
      html_text(),
    bedrooms = page %>% 
      html_node(".ds-bed-bath-living-area-container .ds-bed-bath-item:nth-child(1) .ds-bed-bath-value") %>% 
      html_text() %>% 
      as.integer(),
    bathrooms = page %>% 
      html_node(".ds-bed-bath-living-area-container .ds-bed-bath-item:nth-child(2) .ds-bed-bath-value") %>% 
      html_text() %>% 
      as.numeric(),
    sqft = page %>% 
      html_node(".ds-bed-bath-living-area-container .ds-bed-bath-item:nth-child(3) .ds-bed-bath-value") %>% 
      html_text() %>% 
      str_remove_all(",") %>% 
      as.integer(),
    year_built = page %>% 
      html_node(".ds-home-fact-list li:nth-child(1) .ds-home-fact-value") %>% 
      html_text() %>% 
      as.integer(),
    property_type = page %>% 
      html_node(".ds-home-fact-list li:nth-child(2) .ds-home-fact-value") %>% 
      html_text(),
    annual_taxes = page %>% 
      html_node(".ds-home-fact-list li:nth-child(3) .ds-home-fact-value") %>% 
      html_text() %>% 
      str_remove_all("\\$|,") %>% 
      as.numeric(),
    hoa_fees_monthly = page %>% 
      html_node(".ds-home-fact-list li:nth-child(4) .ds-home-fact-value") %>% 
      html_text() %>% 
      str_remove_all("\\$|,/mo") %>% 
      as.numeric()
  )
  
  # Replace missing values with a clear label so your CSV is clean
  property_data <- property_data %>% 
    mutate(across(everything(), ~replace_na(.x, "Not Available")))
  
  return(property_data)
}

# Example: Create a list of Zillow URLs you want to scrape
zillow_urls <- c(
  "https://www.zillow.com/homedetails/123-Main-St-Anytown-USA-12345/123456789_zpid/",
  "https://www.zillow.com/homedetails/456-Oak-St-Sometown-USA-67890/987654321_zpid/"
)

# Scrape all URLs and combine results into one data frame
scraped_data <- map_dfr(zillow_urls, scrape_zillow_property)

# Export the final data to a CSV file
write_csv(scraped_data, "zillow_property_data.csv")

Key Tips for Success

  1. Updating CSS Selectors: Zillow sometimes changes their page structure, which can break selectors. If a data point stops loading:

    • Right-click the element on the Zillow page (e.g., the bedroom count)
    • Select "Inspect" to open your browser’s developer tools
    • Right-click the highlighted HTML element and choose "Copy > Copy selector"
    • Replace the old selector in the code with this new one
  2. Avoid Getting Blocked: Zillow restricts frequent scraping, so Sys.sleep(2) adds a 2-second wait between requests. For large lists of URLs, increase this to 3-5 seconds.

  3. Handling Missing Data: Some properties won’t have all data points (e.g., no HOA fees for single-family homes). The replace_na() function turns those gaps into "Not Available" so your CSV is easy to read.

If you run into issues with specific data points or URLs, just let me know—we can tweak the code together!

内容的提问来源于stack exchange,提问作者ScottB

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 10:05:15