You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R抓取Booking.com数据时,如何为无值字段填充NA?

问题描述

我用R抓取Booking.com的酒店数据,构建DataFrame时发现不是所有酒店都有评分。尝试了以下代码,但还没实现无对应值时自动填充NA的功能:

# 从页面检查代码获取元素
titles_page <- page %>% html_elements("div[data-testid='title'][class='fcab3ed991 a23c043802']") %>% html_text()
prices_page <- page %>% html_elements("span[data-testid='price-and-discounted-price']") %>% html_text()
ratings_page <- page %>% html_elements("div[aria-label^='Punteggio di']") %>% html_text()

# 处理评分变量
tryCatch(expr ={
      ratings_page <- remDr$findElements(using = "xpath", value = "div[aria-label^='Punteggio di']")$getElementAttribute('value')
    },   
    # 若信息不存在则给评分元素写入NA
    error = function(e){          
      ratings_page <-NA
    })

请问如何让没有对应值的对象自动返回NA?


解决方案

核心思路:按酒店容器遍历提取,确保每个酒店的字段一一对应

直接单独提取所有标题、价格、评分容易出现长度不匹配的问题(比如部分酒店无评分,导致ratings_page长度比titles_page短),更稳妥的方式是先定位每个酒店的父容器,再逐个提取每个酒店的信息,缺失时自动填充NA。


方法1:使用rvest的html_element(单数)替代html_elements(复数)

html_element会为每个节点返回一个结果,不存在则返回NA,完美适配需求:

# 先定位所有酒店的父容器(需根据实际页面结构调整选择器,以下为Booking常见容器示例)
hotels <- page %>% html_elements("div[data-testid='property-card']")

# 逐个提取每个酒店的字段
titles_page <- hotels %>% html_element("div[data-testid='title'][class='fcab3ed991 a23c043802']") %>% html_text2()
prices_page <- hotels %>% html_element("span[data-testid='price-and-discounted-price']") %>% html_text2()
# 提取评分,无评分时自动返回NA
ratings_page <- hotels %>% html_element("div[aria-label^='Punteggio di']") %>% html_text2()

# 构建DataFrame
hotel_df <- tibble(
  title = titles_page,
  price = prices_page,
  rating = ratings_page
)

注:html_text2()比html_text()处理空格更友好,可按需替换;酒店容器选择器需根据实际页面结构调整,确保能选中每个独立的酒店卡片


方法2:修复原有tryCatch逻辑(针对Selenium场景)

如果必须用Selenium的remDr操作,需注意findElements返回的是列表,要逐个处理而非直接调用getElementAttribute:

# 先获取所有酒店元素(对应rvest的hotels节点)
hotel_elements <- remDr$findElements(using = "css", value = "div[data-testid='property-card']")

# 定义函数提取单酒店的评分
extract_rating <- function(hotel_elem) {
  tryCatch({
    rating_elem <- hotel_elem$findElement(using = "xpath", value = "./div[aria-label^='Punteggio di']")
    rating_elem$getElementAttribute("value")[[1]]
  }, error = function(e) {
    NA_character_
  })
}

# 遍历所有酒店提取评分
ratings_page <- sapply(hotel_elements, extract_rating)

# 同理提取标题和价格,确保长度一致
# titles_page <- sapply(hotel_elements, function(x) x$findElement(...)...)
# prices_page <- sapply(hotel_elements, function(x) x$findElement(...)...)

关键说明

  • 避免单独提取所有字段:单独提取会因缺失值导致数组长度不一致,无法直接合并成DataFrame
  • 优先按酒店容器遍历:每个酒店作为独立单元处理,保证每个字段的数量和顺序完全匹配
  • 缺失值自动填充:html_element(rvest)或tryCatch(Selenium)都能在元素不存在时返回NA

内容的提问来源于stack exchange,提问作者Anisa B.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 22:47:26