You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用rvest进行网页爬取:实现Zillow多页房屋数据获取

Hey there, fellow R newbie! I’ve been right where you are—trying to scrape Zillow’s housing data with rvest and hitting that single-page wall. Let’s fix that and build a scraper that grabs all the listings you need: prices, bedroom counts, bathroom counts, and square footage.

Step 1: Understand Zillow’s Pagination

First, Zillow’s search results are paginated, usually with a page parameter in the URL (like ?page=2 for the second page). Some searches use offset instead (each offset increments by the number of listings per page, usually 40). We’ll use the page parameter here, but you can adjust if your URL uses offset.

Step 2: Build a Reusable Single-Page Scraper

Let’s wrap the scraping logic into a function so we can loop through pages easily. We’ll also add a user-agent header to avoid getting blocked immediately—Zillow doesn’t love bots!

library(rvest)
library(dplyr)
library(stringr)
library(httr)

scrape_single_page <- function(page_number) {
  # Replace this base URL with your actual Zillow search URL
  base_url <- "https://www.zillow.com/homes/for_sale/"
  full_url <- paste0(base_url, "?page=", page_number)
  
  # Mimic a real browser request
  browser_headers <- c(
    `User-Agent` = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
  )
  
  # Fetch and parse the page
  page_html <- read_html(GET(full_url, add_headers(.headers = browser_headers)))
  
  # Grab each data point—note: Zillow's class names change sometimes, so double-check these!
  prices <- page_html %>%
    html_nodes(".ListItem-c11n-8-84-3__sc-10e22w8-0 .StyledPropertyCardDataArea-c11n-8-84-3__sc-yipmu-0 span") %>%
    html_text()
  
  bedrooms <- page_html %>%
    html_nodes(".StyledPropertyCardHomeDetailsList-c11n-8-84-3__sc-1xvdaej-0 li:nth-child(1)") %>%
    html_text() %>%
    str_extract("\\d+")  # Pull only the number from the text
  
  bathrooms <- page_html %>%
    html_nodes(".StyledPropertyCardHomeDetailsList-c11n-8-84-3__sc-1xvdaej-0 li:nth-child(2)") %>%
    html_text() %>%
    str_extract("\\d+")
  
  square_footage <- page_html %>%
    html_nodes(".StyledPropertyCardHomeDetailsList-c11n-8-84-3__sc-1xvdaej-0 li:nth-child(3)") %>%
    html_text() %>%
    str_extract("\\d+,?\\d+") %>%
    str_remove(",")  # Clean up commas for numeric conversion later
  
  # Combine into a data frame—handle cases where some listings might miss fields
  page_data <- tibble(
    price = prices,
    bedrooms = bedrooms,
    bathrooms = bathrooms,
    square_footage = square_footage
  )
  
  return(page_data)
}

Step 3: Loop Through All Pages

Now we’ll loop through pages until we hit one with no listings (that means we’ve grabbed everything). We’ll add a small delay between requests to be nice to Zillow’s servers.

# Start with an empty data frame to store all results
all_zillow_data <- tibble()

current_page <- 1
max_pages_to_check <- 20  # Set a reasonable upper limit, or let the loop stop automatically

while(current_page <= max_pages_to_check) {
  cat("Scraping page", current_page, "...\n")
  page_results <- scrape_single_page(current_page)
  
  # If the page has no data, we're done
  if(nrow(page_results) == 0) {
    cat("No more listings found on page", current_page, ". Stopping scrape.\n")
    break
  }
  
  # Add the page's data to our full dataset
  all_zillow_data <- bind_rows(all_zillow_data, page_results)
  
  # Wait 2 seconds before the next request—adjust if needed
  Sys.sleep(2)
  
  current_page <- current_page + 1
}

# Check out your full dataset!
head(all_zillow_data)

Important Notes to Avoid Headaches

  • Zillow changes its HTML often: If your selectors stop working, right-click the element on Zillow’s page, choose "Inspect", and copy the updated CSS selector to replace the ones in the code.
  • Anti-scraping measures: If you get blocked or see a CAPTCHA, pause the script for a bit, or try changing the user-agent string. For more stubborn blocks, you might need to use RSelenium to simulate a real browser (but that’s a bit more advanced).
  • Data cleaning: The raw data will have symbols like $ and commas—use str_remove or parse_number from readr to clean it up for analysis.

内容的提问来源于stack exchange,提问作者alex_fields1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:51:35