You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R爬取房产网站报错:read_xml无法处理list对象,求提取指定链接

问题

需要爬取链接https://www.realtor.ca/map#ZoomLevel=4&Center=58.695434%2C-96.000000&LatitudeMax=72.60462&LongitudeMax=-26.39063&LatitudeMin=35.66836&LongitudeMin=-165.60938&Sort=6-D&PropertyTypeGroupID=1&PropertySearchTypeId=1&TransactionTypeId=2&Currency=CAD中,<div class="cardcon">区域内符合<a href="..." data-binding="href=DetailsURL" class="blockLink listingDetailsLink" target="_blank">结构的房屋详情链接。

尝试的R代码如下:

library(rvest)
library(httr)
library(XML)

url<-"https://www.realtor.ca/map#ZoomLevel=4&Center=58.695434%2C-96.000000&LatitudeMax=72.60462&LongitudeMax=-26.39063&LatitudeMin=35.66836&LongitudeMin=-165.60938&Sort=6-D&PropertyTypeGroupID=1&PropertySearchTypeId=1&TransactionTypeId=2&Currency=CAD"

# making http request
resource <- GET(url)

# converting all the data to HTML format
parse <- htmlParse(resource)

# scrapping all the href tags
links <- xpathSApply(parse, path="//a", xmlGetAttr, "href")

page <-read_html(links)

运行时出现错误:Error in UseMethod("read_xml") : no applicable method for 'read_xml' applied to an object of class "list"


解决方案

错误原因

read_html()仅接受单个HTML资源(URL、响应对象或HTML字符串),但你传入的是xpathSApply()返回的链接列表,类型不匹配导致报错。此外原代码还有两个核心问题:

  1. 目标页面是动态渲染的,直接用GET()只能获取静态HTML框架,无法拿到JavaScript加载的房屋列表数据。
  2. 原XPath未限定cardcon区域,会抓取所有<a>标签的href,不符合需求。

修正方案

使用RSelenium模拟浏览器加载动态内容,配合rvest精准提取目标链接:

1. 安装依赖包

install.packages("rvest")
install.packages("RSelenium")
install.packages("wdman")

2. 完整代码

library(rvest)
library(RSelenium)
library(wdman)

# 启动Chrome驱动
driver <- rsDriver(browser = "chrome", port = 4567L)
remDr <- driver[["client"]]

# 访问目标页面
url <- "https://www.realtor.ca/map#ZoomLevel=4&Center=58.695434%2C-96.000000&LatitudeMax=72.60462&LongitudeMax=-26.39063&LatitudeMin=35.66836&LongitudeMin=-165.60938&Sort=6-D&PropertyTypeGroupID=1&PropertySearchTypeId=1&TransactionTypeId=2&Currency=CAD"
remDr$navigate(url)

# 等待页面加载(根据网络情况调整时长)
Sys.sleep(5)

# 获取渲染后的页面HTML并解析
page_source <- remDr$getPageSource()[[1]]
html <- read_html(page_source)

# 精准提取目标链接:限定cardcon容器内的指定a标签
house_links <- html %>%
  html_elements("div.cardcon a.blockLink.listingDetailsLink[data-binding='href=DetailsURL']") %>%
  html_attr("href")

# 补全域名(若href为相对路径)
house_links <- paste0("https://www.realtor.ca", house_links)

# 输出结果
print(house_links)

# 关闭浏览器和驱动
remDr$close()
driver$server$stop()

关键说明

  • RSelenium模拟真实浏览器行为,能加载动态生成的内容,解决静态请求拿不到数据的问题。
  • 使用html_elements()精准定位:先锁定div.cardcon容器,再筛选符合class和data-binding属性的<a>标签,确保只提取目标房屋链接。
  • 注意控制请求频率,避免触发网站反爬机制导致IP被封禁。

内容的提问来源于stack exchange,提问作者stats_noob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 23:50:37