You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Rvest与RSelenium批量爬Airbnb返回空值,求排查代码错误

问题:Airbnb爬虫批量处理返回空值,单个链接测试正常

我使用Rvest和RSelenium编写了爬虫代码,用于获取Airbnb房源名称。单个链接测试时功能正常,但通过for循环批量处理链接时返回空值。以下是数据结构及代码,请帮忙指出错误所在:

数据结构

structure(list(Property_sku = c("1B - Anantara - 316", "1B - Mag540 - 306", 
"1B- Downtown Views - 3109", "1B- Tiara Tanzanite- 504", "1B-Address JBR - 1107"
), Airbnb_link = c("https://www.airbnb.co.in/rooms/552037226634913505?preview_for_ml=true&source_impression_id=p3_1654086364_RJjGWicrEoR%2FB%2Bgu", 
"https://www.airbnb.co.in/rooms/54045333?preview_for_ml=true&source_impression_id=p3_1644216409_ftDpMWrY34gbixtv", 
"https://www.airbnb.co.in/rooms/54360731?preview_for_ml=true&source_impression_id=p3_1649243904_EjWEoEoKTYpW1zaT", 
"https://www.airbnb.co.in/rooms/565630731118783569?preview_for_ml=true&source_impression_id=p3_1649245563_mMhnLLQhlqTS26sb", 
"https://www.airbnb.co.in/rooms/53245239?preview_for_ml=true&source_impression_id=p3_1644215345_i3xkL5TcGvenCy2j"
)), row.names = c(NA, -5L), class = c("tbl_df", "tbl", "data.frame"
))

我的代码

library(rvest)
library(tidyverse)
library(RSelenium)
rD <-  rsDriver(browser="chrome",port=4234L,chromever="104.0.5112.79")
remDr <- rD$client
for(i in 1:length(Aug_Active_property_airbnb_review$Airbnb_link)){
  remDr$navigate(paste0(Aug_Active_property_airbnb_review$Airbnb_link[[i]]))
  Sys.sleep(1)
  Aug_Active_property_airbnb_review$listing_name <- sapply(Aug_Active_property_airbnb_review$Airbnb_link[[i]],function(url){
    read_html(remDr$getPageSource()[[1]]) %>% html_nodes("h1._fecoyn4") %>% html_text2()
    Sys.sleep(1)
  }, USE.NAMES = FALSE)}

错误分析

  • sapply使用逻辑错误:在for循环中,你传入sapply的是单个链接字符串Aug_Active_property_airbnb_review$Airbnb_link[[i]],sapply会遍历该字符串的每个字符而非链接本身,导致每次循环都在处理无效的字符输入,自然无法获取房源名称。同时,每次循环都会覆盖listing_name整列,最终仅保留最后一次错误的结果。
  • 页面等待时间不足:Sys.sleep(1)的固定休眠时间过短,Airbnb页面加载(尤其是动态渲染)需要更长时间,批量访问时加载速度可能更慢,导致解析时目标节点尚未渲染完成,返回空值。
  • 列赋值逻辑错误:循环中直接给Aug_Active_property_airbnb_review$listing_name整列赋值,而非对应行的位置,导致之前的结果被反复覆盖,最终数据混乱。

修正后的代码

library(rvest)
library(tidyverse)
library(RSelenium)

# 启动Chrome驱动
rD <- rsDriver(browser="chrome", port=4234L, chromever="104.0.5112.79")
remDr <- rD$client

# 初始化房源名称列
Aug_Active_property_airbnb_review$listing_name <- NA_character_

# 循环处理每个链接
for(i in seq_along(Aug_Active_property_airbnb_review$Airbnb_link)){
  # 跳转到目标链接
  remDr$navigate(Aug_Active_property_airbnb_review$Airbnb_link[[i]])
  
  # 延长等待时间,确保页面完全加载
  Sys.sleep(3)
  
  # 获取页面源码并解析房源名称
  page_source <- remDr$getPageSource()[[1]]
  listing_name <- read_html(page_source) %>% 
    html_nodes("h1._fecoyn4") %>% 
    html_text2()
  
  # 给对应行赋值,若未获取到则设为NA
  Aug_Active_property_airbnb_review$listing_name[i] <- ifelse(length(listing_name) > 0, listing_name, NA)
  
  # 添加随机延迟,避免触发反爬机制
  Sys.sleep(sample(1:2, 1))
}

# 关闭浏览器和驱动
remDr$close()
rD$server$stop()

额外优化建议

  • 使用显式等待替代固定休眠:显式等待可以等到目标元素出现后再操作,比固定休眠更高效可靠:
# 等待h1元素出现,最多等待10秒
webElem <- remDr$waitForElement(using = "css selector", value = "h1._fecoyn4", timeout = 10000)
  • 添加错误捕获机制:用tryCatch处理单个链接的异常,避免循环因某一个链接出错而中断:
for(i in seq_along(Aug_Active_property_airbnb_review$Airbnb_link)){
  tryCatch({
    remDr$navigate(Aug_Active_property_airbnb_review$Airbnb_link[[i]])
    Sys.sleep(3)
    page_source <- remDr$getPageSource()[[1]]
    listing_name <- read_html(page_source) %>% 
      html_nodes("h1._fecoyn4") %>% 
      html_text2()
    Aug_Active_property_airbnb_review$listing_name[i] <- ifelse(length(listing_name) > 0, listing_name, NA)
    Sys.sleep(sample(1:2, 1))
  }, error = function(e) {
    message(paste("处理第", i, "个链接时出错:", e$message))
    Aug_Active_property_airbnb_review$listing_name[i] <- NA
  })
}

内容的提问来源于stack exchange,提问作者shubham tiwari

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 12:12:24