使用R抓取Booking.com数据时,如何为无值字段填充NA?
问题描述
我用R抓取Booking.com的酒店数据,构建DataFrame时发现不是所有酒店都有评分。尝试了以下代码,但还没实现无对应值时自动填充NA的功能:
# 从页面检查代码获取元素 titles_page <- page %>% html_elements("div[data-testid='title'][class='fcab3ed991 a23c043802']") %>% html_text() prices_page <- page %>% html_elements("span[data-testid='price-and-discounted-price']") %>% html_text() ratings_page <- page %>% html_elements("div[aria-label^='Punteggio di']") %>% html_text() # 处理评分变量 tryCatch(expr ={ ratings_page <- remDr$findElements(using = "xpath", value = "div[aria-label^='Punteggio di']")$getElementAttribute('value') }, # 若信息不存在则给评分元素写入NA error = function(e){ ratings_page <-NA })
请问如何让没有对应值的对象自动返回NA?
解决方案
核心思路:按酒店容器遍历提取,确保每个酒店的字段一一对应
直接单独提取所有标题、价格、评分容易出现长度不匹配的问题(比如部分酒店无评分,导致ratings_page长度比titles_page短),更稳妥的方式是先定位每个酒店的父容器,再逐个提取每个酒店的信息,缺失时自动填充NA。
方法1:使用rvest的html_element(单数)替代html_elements(复数)
html_element会为每个节点返回一个结果,不存在则返回NA,完美适配需求:
# 先定位所有酒店的父容器(需根据实际页面结构调整选择器,以下为Booking常见容器示例) hotels <- page %>% html_elements("div[data-testid='property-card']") # 逐个提取每个酒店的字段 titles_page <- hotels %>% html_element("div[data-testid='title'][class='fcab3ed991 a23c043802']") %>% html_text2() prices_page <- hotels %>% html_element("span[data-testid='price-and-discounted-price']") %>% html_text2() # 提取评分,无评分时自动返回NA ratings_page <- hotels %>% html_element("div[aria-label^='Punteggio di']") %>% html_text2() # 构建DataFrame hotel_df <- tibble( title = titles_page, price = prices_page, rating = ratings_page )
注:html_text2()比html_text()处理空格更友好,可按需替换;酒店容器选择器需根据实际页面结构调整,确保能选中每个独立的酒店卡片
方法2:修复原有tryCatch逻辑(针对Selenium场景)
如果必须用Selenium的remDr操作,需注意findElements返回的是列表,要逐个处理而非直接调用getElementAttribute:
# 先获取所有酒店元素(对应rvest的hotels节点) hotel_elements <- remDr$findElements(using = "css", value = "div[data-testid='property-card']") # 定义函数提取单酒店的评分 extract_rating <- function(hotel_elem) { tryCatch({ rating_elem <- hotel_elem$findElement(using = "xpath", value = "./div[aria-label^='Punteggio di']") rating_elem$getElementAttribute("value")[[1]] }, error = function(e) { NA_character_ }) } # 遍历所有酒店提取评分 ratings_page <- sapply(hotel_elements, extract_rating) # 同理提取标题和价格,确保长度一致 # titles_page <- sapply(hotel_elements, function(x) x$findElement(...)...) # prices_page <- sapply(hotel_elements, function(x) x$findElement(...)...)
关键说明
- 避免单独提取所有字段:单独提取会因缺失值导致数组长度不一致,无法直接合并成DataFrame
- 优先按酒店容器遍历:每个酒店作为独立单元处理,保证每个字段的数量和顺序完全匹配
- 缺失值自动填充:
html_element(rvest)或tryCatch(Selenium)都能在元素不存在时返回NA
内容的提问来源于stack exchange,提问作者Anisa B.
相关产品推荐
相关产品推荐

