R爬取房产网站报错:read_xml无法处理list对象,求提取指定链接
问题
需要爬取链接https://www.realtor.ca/map#ZoomLevel=4&Center=58.695434%2C-96.000000&LatitudeMax=72.60462&LongitudeMax=-26.39063&LatitudeMin=35.66836&LongitudeMin=-165.60938&Sort=6-D&PropertyTypeGroupID=1&PropertySearchTypeId=1&TransactionTypeId=2&Currency=CAD中,<div class="cardcon">区域内符合<a href="..." data-binding="href=DetailsURL" class="blockLink listingDetailsLink" target="_blank">结构的房屋详情链接。
尝试的R代码如下:
library(rvest) library(httr) library(XML) url<-"https://www.realtor.ca/map#ZoomLevel=4&Center=58.695434%2C-96.000000&LatitudeMax=72.60462&LongitudeMax=-26.39063&LatitudeMin=35.66836&LongitudeMin=-165.60938&Sort=6-D&PropertyTypeGroupID=1&PropertySearchTypeId=1&TransactionTypeId=2&Currency=CAD" # making http request resource <- GET(url) # converting all the data to HTML format parse <- htmlParse(resource) # scrapping all the href tags links <- xpathSApply(parse, path="//a", xmlGetAttr, "href") page <-read_html(links)
运行时出现错误:Error in UseMethod("read_xml") : no applicable method for 'read_xml' applied to an object of class "list"
解决方案
错误原因
read_html()仅接受单个HTML资源(URL、响应对象或HTML字符串),但你传入的是xpathSApply()返回的链接列表,类型不匹配导致报错。此外原代码还有两个核心问题:
- 目标页面是动态渲染的,直接用
GET()只能获取静态HTML框架,无法拿到JavaScript加载的房屋列表数据。 - 原XPath未限定
cardcon区域,会抓取所有<a>标签的href,不符合需求。
修正方案
使用RSelenium模拟浏览器加载动态内容,配合rvest精准提取目标链接:
1. 安装依赖包
install.packages("rvest") install.packages("RSelenium") install.packages("wdman")
2. 完整代码
library(rvest) library(RSelenium) library(wdman) # 启动Chrome驱动 driver <- rsDriver(browser = "chrome", port = 4567L) remDr <- driver[["client"]] # 访问目标页面 url <- "https://www.realtor.ca/map#ZoomLevel=4&Center=58.695434%2C-96.000000&LatitudeMax=72.60462&LongitudeMax=-26.39063&LatitudeMin=35.66836&LongitudeMin=-165.60938&Sort=6-D&PropertyTypeGroupID=1&PropertySearchTypeId=1&TransactionTypeId=2&Currency=CAD" remDr$navigate(url) # 等待页面加载(根据网络情况调整时长) Sys.sleep(5) # 获取渲染后的页面HTML并解析 page_source <- remDr$getPageSource()[[1]] html <- read_html(page_source) # 精准提取目标链接:限定cardcon容器内的指定a标签 house_links <- html %>% html_elements("div.cardcon a.blockLink.listingDetailsLink[data-binding='href=DetailsURL']") %>% html_attr("href") # 补全域名(若href为相对路径) house_links <- paste0("https://www.realtor.ca", house_links) # 输出结果 print(house_links) # 关闭浏览器和驱动 remDr$close() driver$server$stop()
关键说明
- RSelenium模拟真实浏览器行为,能加载动态生成的内容,解决静态请求拿不到数据的问题。
- 使用
html_elements()精准定位:先锁定div.cardcon容器,再筛选符合class和data-binding属性的<a>标签,确保只提取目标房屋链接。 - 注意控制请求频率,避免触发网站反爬机制导致IP被封禁。
内容的提问来源于stack exchange,提问作者stats_noob
相关产品推荐
相关产品推荐

