使用rvest进行网页爬取:实现Zillow多页房屋数据获取
Hey there, fellow R newbie! I’ve been right where you are—trying to scrape Zillow’s housing data with rvest and hitting that single-page wall. Let’s fix that and build a scraper that grabs all the listings you need: prices, bedroom counts, bathroom counts, and square footage.
Step 1: Understand Zillow’s Pagination
First, Zillow’s search results are paginated, usually with a page parameter in the URL (like ?page=2 for the second page). Some searches use offset instead (each offset increments by the number of listings per page, usually 40). We’ll use the page parameter here, but you can adjust if your URL uses offset.
Step 2: Build a Reusable Single-Page Scraper
Let’s wrap the scraping logic into a function so we can loop through pages easily. We’ll also add a user-agent header to avoid getting blocked immediately—Zillow doesn’t love bots!
library(rvest) library(dplyr) library(stringr) library(httr) scrape_single_page <- function(page_number) { # Replace this base URL with your actual Zillow search URL base_url <- "https://www.zillow.com/homes/for_sale/" full_url <- paste0(base_url, "?page=", page_number) # Mimic a real browser request browser_headers <- c( `User-Agent` = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" ) # Fetch and parse the page page_html <- read_html(GET(full_url, add_headers(.headers = browser_headers))) # Grab each data point—note: Zillow's class names change sometimes, so double-check these! prices <- page_html %>% html_nodes(".ListItem-c11n-8-84-3__sc-10e22w8-0 .StyledPropertyCardDataArea-c11n-8-84-3__sc-yipmu-0 span") %>% html_text() bedrooms <- page_html %>% html_nodes(".StyledPropertyCardHomeDetailsList-c11n-8-84-3__sc-1xvdaej-0 li:nth-child(1)") %>% html_text() %>% str_extract("\\d+") # Pull only the number from the text bathrooms <- page_html %>% html_nodes(".StyledPropertyCardHomeDetailsList-c11n-8-84-3__sc-1xvdaej-0 li:nth-child(2)") %>% html_text() %>% str_extract("\\d+") square_footage <- page_html %>% html_nodes(".StyledPropertyCardHomeDetailsList-c11n-8-84-3__sc-1xvdaej-0 li:nth-child(3)") %>% html_text() %>% str_extract("\\d+,?\\d+") %>% str_remove(",") # Clean up commas for numeric conversion later # Combine into a data frame—handle cases where some listings might miss fields page_data <- tibble( price = prices, bedrooms = bedrooms, bathrooms = bathrooms, square_footage = square_footage ) return(page_data) }
Step 3: Loop Through All Pages
Now we’ll loop through pages until we hit one with no listings (that means we’ve grabbed everything). We’ll add a small delay between requests to be nice to Zillow’s servers.
# Start with an empty data frame to store all results all_zillow_data <- tibble() current_page <- 1 max_pages_to_check <- 20 # Set a reasonable upper limit, or let the loop stop automatically while(current_page <= max_pages_to_check) { cat("Scraping page", current_page, "...\n") page_results <- scrape_single_page(current_page) # If the page has no data, we're done if(nrow(page_results) == 0) { cat("No more listings found on page", current_page, ". Stopping scrape.\n") break } # Add the page's data to our full dataset all_zillow_data <- bind_rows(all_zillow_data, page_results) # Wait 2 seconds before the next request—adjust if needed Sys.sleep(2) current_page <- current_page + 1 } # Check out your full dataset! head(all_zillow_data)
Important Notes to Avoid Headaches
- Zillow changes its HTML often: If your selectors stop working, right-click the element on Zillow’s page, choose "Inspect", and copy the updated CSS selector to replace the ones in the code.
- Anti-scraping measures: If you get blocked or see a CAPTCHA, pause the script for a bit, or try changing the user-agent string. For more stubborn blocks, you might need to use
RSeleniumto simulate a real browser (but that’s a bit more advanced). - Data cleaning: The raw data will have symbols like
$and commas—usestr_removeorparse_numberfromreadrto clean it up for analysis.
内容的提问来源于stack exchange,提问作者alex_fields1

