如何用R直接爬取realtor.ca中cardcon区块的房源链接
问题:提取Realtor.ca页面中房源详情链接
我需要从以下页面提取所有<div class="cardcon">区块内的房源详情链接:
https://www.realtor.ca/map#ZoomLevel=4&Center=58.695434%2C-96.000000&LatitudeMax=72.60462&LongitudeMax=-26.39063&LatitudeMin=35.66836&LongitudeMin=-165.60938&Sort=6-D&PropertyTypeGroupID=1&PropertySearchTypeId=1&TransactionTypeId=2&Currency=CAD
预期输出示例:
- https://www.realtor.ca/real-estate/25054113/4918-lafontaine-hanmer
- https://www.realtor.ca/real-estate/25054111/77-shady-shores-drive-w-winnipeg-waterside-estates
- ...
此前尝试API爬取失败,转而直接网页爬取,编写了以下代码但报错:
library(rvest) library(httr) library(XML) url<-"https://www.realtor.ca/map#ZoomLevel=4&Center=58.695434%2C-96.000000&LatitudeMax=72.60462&LongitudeMax=-26.39063&LatitudeMin=35.66836&LongitudeMin=-165.60938&Sort=6-D&PropertyTypeGroupID=1&PropertySearchTypeId=1&TransactionTypeId=2&Currency=CAD" # making http request resource <- GET(url) # converting all the data to HTML format parse <- htmlParse(resource) # scrapping all the href tags links <- xpathSApply(parse, path="//a", xmlGetAttr, "href") page <-read_html(links)
报错信息:
Error in UseMethod("read_xml") : no applicable method for 'read_xml' applied to an object of class "list"
问题分析与解决方案
核心问题
- 动态页面加载限制:直接用
GET()获取的是静态HTML,该页面的房源数据是通过JavaScript动态渲染的,静态页面中不存在cardcon区块的内容。 - 代码逻辑错误:
links是提取到的href列表,而read_html()仅支持传入HTML内容或URL,无法处理列表类型,导致报错。
可行方案(使用动态渲染工具)
方案一:RSelenium(需浏览器驱动支持)
# 安装依赖 install.packages("RSelenium") library(RSelenium) library(rvest) # 启动Chrome驱动(需提前安装ChromeDriver并配置系统路径) driver <- rsDriver(browser = "chrome", port = 4567L) remDr <- driver[["client"]] # 访问目标页面并等待加载 url <- "https://www.realtor.ca/map#ZoomLevel=4&Center=58.695434%2C-96.000000&LatitudeMax=72.60462&LongitudeMax=-26.39063&LatitudeMin=35.66836&LongitudeMin=-165.60938&Sort=6-D&PropertyTypeGroupID=1&PropertySearchTypeId=1&TransactionTypeId=2&Currency=CAD" remDr$navigate(url) Sys.sleep(5) # 根据网络情况调整等待时间 # 获取渲染后的页面HTML page_source <- remDr$getPageSource()[[1]] page <- read_html(page_source) # 提取cardcon区块内的房源链接 property_links <- page %>% html_elements("div.cardcon a") %>% html_attr("href") %>% paste0("https://www.realtor.ca", .) %>% # 拼接完整URL unique() %>% grep("real-estate", ., value = TRUE) # 过滤有效房源链接 # 查看结果 head(property_links) # 关闭浏览器与驱动 remDr$close() driver$server$stop()
方案二:Playwright(轻量动态渲染工具)
# 安装依赖 install.packages("playwright") library(playwright) playwright::install_browsers() # 启动浏览器并访问页面 pw <- playwright$firefox() browser <- pw$launch(headless = TRUE) page <- browser$new_page() page$goto(url, wait_until = "networkidle") # 等待网络空闲确保加载完成 page_source <- page$content() # 解析页面并提取链接 page_html <- read_html(page_source) property_links <- page_html %>% html_elements("div.cardcon a") %>% html_attr("href") %>% paste0("https://www.realtor.ca", .) %>% unique() %>% grep("real-estate", ., value = TRUE) # 查看结果 head(property_links) # 关闭资源 browser$close() pw$stop()
内容的提问来源于stack exchange,提问作者stats_noob
相关产品推荐
相关产品推荐

