如何修改R代码爬取Zillow所有页面的Agent listings与Other listings
问题修复方案
核心问题说明
- 页码参数未传入请求:代码里虽然写了1到40的循环,但URL中的分页参数
pagination始终为空,每次请求的都是第一页,自然只能拿到前20条代理房源 - 选择器覆盖范围不足:当前用的
.list-card-前缀的CSS选择器仅匹配代理发布的房源,Other listings的卡片类名规则不同,没有被采集到 - 缺少反爬配置:Zillow对无标识高频请求拦截严格,即使改了分页参数,后续请求也大概率被屏蔽返回空数据
修改后可运行代码
首先确保你已经安装加载了必要依赖包:
library(rvest) library(tibble) library(dplyr) library(httr) library(jsonlite)
修改后的采集逻辑:
res_all <- NULL # 基础查询参数,直接解码原URL里的searchQueryState得到 base_query <- list( pagination = list(), usersSearchTerm = "Providence, RI", mapBounds = list(west = -71.48892251635742, east = -71.36017648364258, south = 41.77131876826507, north = 41.862664689400106), regionSelection = list(list(regionId = 26637, regionType = 6)), isMapVisible = TRUE, filterState = list( sort = list(value = "globalrelevanceex"), ah = list(value = TRUE), sf = list(value = FALSE), tow = list(value = FALSE), con = list(value = FALSE), apco = list(value = FALSE), land = list(value = FALSE), apa = list(value = FALSE), manu = list(value = FALSE) ), isListVisible = TRUE, mapZoom = 13 ) # 最多爬10页足够覆盖你要的65条数据,不用跑40页 for (page_result in 1:10) { # 动态设置当前页码 base_query$pagination$currentPage <- page_result # 编码查询参数拼接到URL query_encode <- URLencode(toJSON(base_query, auto_unbox = TRUE)) zillow_url <- paste0("https://www.zillow.com/providence-ri/duplex/?searchQueryState=", query_encode) # 加请求头模拟浏览器,避免被反爬 zpg <- GET(zillow_url, user_agent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") ) %>% read_html() # 用覆盖两类房源的通用选择器提取数据 property_cards <- zpg %>% html_nodes(".property-card") # 当前页无数据直接跳出循环 if(length(property_cards) == 0) break zillow_pg <- tibble( addr = property_cards %>% html_node(".property-card-addr") %>% html_text(trim = TRUE), price = property_cards %>% html_node(".property-card-price") %>% html_text(trim = TRUE), details = property_cards %>% html_node(".property-card-details") %>% html_text(trim = TRUE), heading = property_cards %>% html_node(".property-card-info a") %>% html_text(trim = TRUE), type = property_cards %>% html_node(".property-card-statusText") %>% html_text(trim = TRUE), source = ifelse(grepl("list-card", html_attr(property_cards, "class")), "Agent listings", "Other listings") ) res_all <- distinct(bind_rows(res_all, zillow_pg)) # 加延时2-5秒,避免被封 Sys.sleep(runif(1, 2, 5)) }
关键修改说明
- 动态生成分页URL:将原固定的查询参数转为列表,每次循环修改
currentPage字段后重新编码,保证每次请求对应不同页码 - 更换通用选择器:改用
.property-card前缀的选择器,同时覆盖代理房源和其他房源,额外加了source字段区分两类数据 - 新增反爬配置:添加浏览器UA请求头,每页采集后加随机延时,降低被拦截概率
- 增加跳出逻辑:当前页无房源返回时直接终止循环,避免无效请求
内容的提问来源于stack exchange,提问作者DJS
相关产品推荐
相关产品推荐

